A line array camera image matching method based on particle swarm feedback

CN122434729BActive Publication Date: 2026-09-25CRRC HANGZHOU DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610913985.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

匹配过程中产生的误匹配点对比例较高,而传统的随机抽样一致算法采用固定距离阈值和迭代次数,无法适应不同图像对的质量波动,剔除效果不稳定

Benefits of technology

[0018]本发明的有益效果是:本发明通过几何校正与运动补偿消除了图像采集中的畸变与拉伸压缩,利用双分支深度网络和组合特征描述子提升了特征区分力,采用两阶段误匹配剔除算法获得稳定内点集,改进粒子群算法实现了亚像素级位移优化并输出质量评价指数,再通过分级反馈闭环动态调节误匹配剔除参数,显著提高了线阵相机在高速运动、抖动及纹理重复等复杂条件下的匹配精度与拼接质量,实现了自适应、高鲁棒的图像匹配。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434729B_ABST
    Figure CN122434729B_ABST
Patent Text Reader

Abstract

The application discloses a kind of line array camera image matching methods based on particle swarm feedback, collects real-time image and speed signal, utilizes camera internal and external parameter to construct geometric correction model, eliminates distortion and installation deviation;Motion compensation model is constructed, and according to speed signal and image blurring feedback dynamic adjustment scanning frequency and reacquire image, adopt double branch deep network to extract space and time sequence feature, combination gradient histogram, grey statistics and local binary pattern texture form feature descriptor, obtain initial point pair by nearest neighbor matching, and obtain inner point set and initial transformation model by two-stage false matching rejection, according to quality index grading adjustment multiple threshold parameters of false matching rejection algorithm, and trigger re-elimination or particle reset, form feedback closed loop.When transformation model change amount is less than threshold value, image stitching is completed.The application realizes the high-precision adaptive matching of line array camera image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing for line scan cameras, and more specifically to an image matching method for line scan cameras based on particle swarm feedback. Background Technology

[0002] Line scan cameras acquire images of objects through line-by-line scanning and are widely used in visual inspection and quality monitoring of high-speed moving objects. In practical applications, there is a relatively high-speed motion between the object and the camera, which raises several technical challenges.

[0003] Inherent distortion of camera lenses, angular deviations during installation, and minute displacements of mechanical structures can cause geometric distortions in acquired images, such as radial barrel distortion, tangential trapezoidal distortion, and scan line tilting or misalignment. These distortions disrupt the spatial consistency of the image, directly affecting the accuracy of subsequent matching and stitching.

[0004] The scanning frequency of a line scan camera needs to be precisely matched to the speed of the moving object. When the scanning frequency is lower than the object's speed, the image is compressed; when the scanning frequency is higher than the object's speed, the image is stretched. Traditional methods often use a fixed frequency or a simple encoder proportional mapping, ignoring the feedback from the image content itself and the instantaneous fluctuations in speed, resulting in limited compensation accuracy.

[0005] In image matching, commonly used feature extraction methods such as scale-invariant feature transform or oriented fast rotation descriptors lack discriminative power when dealing with scenes with repetitive textures or drastic lighting changes. The proportion of mismatched point pairs generated during the matching process is relatively high, and traditional random sampling consensus algorithms, which use fixed distance thresholds and iteration counts, cannot adapt to the quality fluctuations of different image pairs, resulting in unstable removal performance.

[0006] Existing matching optimization methods mostly remain at the integer pixel level, and achieving sub-pixel accuracy relies on interpolation or local gradient fitting, which easily leads to getting trapped in local optima. At the same time, the parameters of each step in the entire matching process are independent of each other, lacking a closed-loop feedback mechanism to dynamically adjust the parameters of the upstream algorithm based on the optimization results.

[0007] Therefore, in order to solve the problems existing in the prior art, this application proposes an image matching method for linear array cameras based on particle swarm feedback. Summary of the Invention

[0008] In view of the shortcomings of the existing technology, the purpose of this invention is to provide an image matching method for a linear array camera based on particle swarm feedback.

[0009] To achieve the above objectives, the present invention provides the following technical solution: An image matching method for a linear scan camera based on particle swarm feedback includes: The image acquisition step involves acquiring real-time images of the target object using a line scan camera and obtaining the object's motion speed signal. The geometric correction step involves constructing a geometric correction model based on camera parameters, performing geometric error correction on the real-time image to obtain the corrected real-time image, and generating an image pair with the pre-stored template image. The motion compensation step involves constructing a motion compensation model, adjusting the scanning frequency of the linear scan camera based on the motion speed signal and the image quality feedback of the real-time image, and then re-acquiring the real-time image. The image matching step involves extracting and describing features from the corrected image pairs to obtain initial matching point pairs. The initial matching point pairs are then initially removed using a mismatch removal algorithm to obtain an inner point set and an initial transformation model. Based on the inner point set, the sub-pixel displacement is optimized using a particle swarm optimization algorithm to obtain the optimal sub-pixel displacement parameters, and the matching quality evaluation index is calculated. The compensation and adjustment step involves adjusting the judgment parameters of the error elimination algorithm according to the matching quality evaluation index, and then re-performing the mismatch elimination based on the adjusted parameters, thereby updating the eliminated interior point set and transformation model. In the image stitching step, the magnitude of the change in the transformation model is compared with a preset threshold. When the change is less than the preset threshold, the image is stitched using the final transformation model.

[0010] As a further improvement of the present invention, the motion compensation step includes: calculating the blur index of the real-time image, wherein the blur index is calculated based on the ratio of the gradient energy of the current image to the gradient energy of the standard clear image; converting the blur index into a frequency correction factor through a bounded nonlinear function; evaluating the motion speed signal to obtain an estimated speed; and calculating the scanning frequency at the next moment based on the estimated speed, the camera pixel spacing, and the frequency correction factor.

[0011] As a further improvement of the present invention, the geometric correction step includes constructing a geometric correction model based on the intrinsic and extrinsic parameters of the line scan camera. The intrinsic parameters include focal length, principal point, and distortion coefficient, and the extrinsic parameters include rotation matrix and translation vector. The geometric correction model is used to eliminate radial and tangential distortion of the acquired real-time image, and to compensate for scanning angle deviation and camera physical installation position deviation, and output a corrected distortion-free image.

[0012] As a further improvement of the present invention, the image matching step further includes constructing a dual-branch deep convolutional network, wherein the first branch is a spatial feature extraction branch, which encodes spatial features of the input image through multi-scale depth separable convolution, and the second branch is a temporal feature extraction branch, which splits the image into a row vector sequence by row, extracts inter-row correlation features through one-dimensional convolution along the scan line direction, and obtains a depth feature map by adaptive weighted fusion of the feature maps output by the branch.

[0013] As a further improvement of the present invention, the feature descriptor includes extracting the gradient direction histogram, gray-level distribution statistics and local binary pattern texture features of the local region of the image, and weighting and fusing the three feature vectors to obtain the feature descriptor.

[0014] As a further improvement of the present invention, the image matching step further includes calculating the distance between each feature point extracted from the template image and the feature descriptors of all feature points in the real-time image through a nearest neighbor matching strategy, and selecting the feature point with the smallest distance as the initial matching point pair.

[0015] As a further improvement of the present invention, the error elimination algorithm includes: calculating the Euclidean distance ratio and the connection angle deviation of each matching point pair; eliminating point pairs whose distance ratio exceeds a first preset threshold or whose angle deviation exceeds a second preset threshold; using the remaining point pairs as nodes; calculating the edge weights between nodes based on spatial distance and feature descriptor similarity; constructing a graph structure; dividing the nodes into several clusters using a clustering algorithm; eliminating clusters smaller than the smallest cluster size; and selecting the largest cluster as the inner point set for output.

[0016] As a further improvement of the present invention, the particle swarm algorithm includes: calculating an initial sub-pixel displacement estimate based on an initial transformation model; dividing the particle swarm into two parts; initializing the first part in the neighborhood of the initial sub-pixel displacement estimate according to a Gaussian distribution; and initializing the second part uniformly and randomly within a preset search range; setting the relative decrease in fitness between two adjacent generations as the fitness improvement rate; adjusting the inertia weight, individual learning factor, and social learning factor according to the fitness improvement rate; setting a stagnation counter; determining stagnation when the change in global optimal fitness for several consecutive iterations is less than a preset threshold; applying a Gaussian perturbation to the global optimal particle, with the perturbation amplitude increasing with the stagnation generation; reinitializing all particles to the current optimal neighborhood; initiating a secondary fine search; and outputting the optimal sub-pixel displacement parameters after the algorithm converges. Simultaneously, calculating and outputting a quality evaluation index, including the standard deviation of the matching residual and a convergence rate index, wherein the standard deviation of the matching residual is the standard deviation of the final projection error of all interior points, and the convergence rate index is the average fitness decrease rate of the last few generations.

[0017] As a further improvement of this invention, based on the standard deviation of the matching residuals and the convergence speed index output by the particle swarm optimization algorithm, and compared with three preset thresholds, the matching quality is divided into three levels: excellent, medium, and poor. A priority order for adjusting multiple adjustable parameters is set, with the cluster similarity threshold having the highest priority, followed by the distance ratio threshold, then the minimum cluster size, and the angle deviation threshold having the lowest priority. When the level is determined to be excellent, the cluster similarity threshold is adjusted; when it is determined to be medium, both the cluster similarity threshold and the distance ratio threshold are adjusted simultaneously, and error matching is re-performed; if the level is determined to be poor, all four parameters are adjusted, the minimum cluster size is reduced, and the angle deviation threshold is relaxed, triggering a re-execution of error matching removal, and some particle reset information is fed back to the particle swarm optimization algorithm to trigger sub-pixel optimization again.

[0018] The beneficial effects of this invention are as follows: This invention eliminates distortion and stretching / compression in image acquisition through geometric correction and motion compensation, enhances feature discrimination by utilizing a dual-branch deep network and combined feature descriptors, obtains a stable set of interior points by employing a two-stage mismatch elimination algorithm, achieves sub-pixel-level displacement optimization and outputs a quality evaluation index by improving the particle swarm optimization algorithm, and then dynamically adjusts the mismatch elimination parameters through hierarchical feedback closed loop, which significantly improves the matching accuracy and stitching quality of line scan cameras under complex conditions such as high-speed motion, shaking, and texture repetition, and realizes adaptive and highly robust image matching. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of the method according to Embodiment 2 of the present invention; Figure 3 This is a flowchart of the method in Embodiment 3 of the present invention. Detailed Implementation

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0021] Example 1: An image matching method based on particle swarm feedback for linear scan cameras, such as Figure 1 As shown, it includes: In the image acquisition step, the line scan camera scans the object under test line by line, generating a one-dimensional image signal for each line scanned. During the scanning process, the motion velocity signal output by the encoder connected to the camera is acquired simultaneously. This velocity signal reflects the actual movement rate of the object in the camera's field of view and serves as the basic input for subsequent motion compensation.

[0022] The geometric correction step utilizes pre-calibrated camera parameters, including principal point coordinates, focal length, and radial and tangential distortion coefficients in the intrinsic parameters, and rotation matrix and translation vector in the extrinsic parameters, to construct a geometric correction model. This model performs coordinate mapping transformation on the real-time acquired line-by-line images to eliminate radial and tangential distortion caused by the lens's own curvature. Simultaneously, the model compensates for scan line tilting or misalignment caused by camera mounting angle deviations and physical position offsets. The corrected real-time image, free of geometric distortion, is then paired with a pre-stored template image to form a matching image pair. The template image is a pre-acquired standard reference image that has also undergone the aforementioned geometric correction process, ensuring coordinate system consistency between the two.

[0023] The motion compensation step aims to match the line scan frequency of the linear scan camera with the speed of the moving object in real time, thereby avoiding image stretching or compression. This step employs an adaptive compensation model based on image content feedback.

[0024] First, the blur index of the real-time image is calculated. The blur index is defined as the ratio of the overall gradient energy of the current image to the gradient energy of a standard clear image. Gradient energy is obtained by calculating and summing the rate of change of brightness of each pixel in the horizontal and vertical directions. If the current image experiences motion blur due to velocity mismatch, its gradient energy will decrease, and the blur index will be greater than 1; if the image is over-compressed or stretched, resulting in overly sharp edges, the gradient energy may increase, and the blur index will be less than 1. This index directly reflects the degree of matching between the acquisition frequency and the object's velocity.

[0025] Next, the ambiguity index is input into a bounded nonlinear function for transformation. This function uses a hyperbolic tangent mapping and outputs a frequency correction factor. Since the output range of the hyperbolic tangent function is limited to a fixed interval, the correction factor will not excessively change the scanning frequency, thus ensuring system stability. When the ambiguity index deviates from 1, the correction factor increases or decreases accordingly, in the opposite direction to the mismatch trend.

[0026] Simultaneously, a time-series evaluation is performed on the motion velocity signal output by the encoder. Considering that the object's velocity may fluctuate instantaneously, a Kalman filter is used to smooth the velocity signal and perform a one-step prediction, resulting in a more accurate estimated velocity. The Kalman filter can fuse historical velocity observations with current measurements, effectively suppressing noise.

[0027] Finally, the scanning frequency for the next moment is calculated: the estimated velocity is divided by the physical pixel spacing corresponding to a single pixel in the camera to obtain the theoretical matching frequency; this frequency is then multiplied by a frequency correction factor, and a cumulative integral term of historical frequency error is added. The integral term is used to eliminate long-term steady-state deviations, allowing the scanning frequency to gradually converge to the optimal value. After the calculation is completed, the new scanning frequency is sent to the trigger controller of the linear scan camera. The camera uses this frequency for image acquisition in subsequent scanning cycles, eliminating stretching or compression distortion at its source.

[0028] The image matching step performs feature extraction and description on the corrected image pairs. Feature extraction employs a dual-branch deep convolutional network. The spatial branch uses multi-scale deep separable convolution to capture local texture and edges, while the temporal branch performs one-dimensional convolution along the image row direction to extract correlation features between scan lines. The outputs of the two branches are adaptively weighted and fused to form a depth feature map. Subsequently, a combined feature descriptor is constructed, fusing the gradient direction distribution, gray-level mean and variance, and local binary pattern texture of the local region. A nearest neighbor matching strategy is used: for each feature point in the template image, its descriptor distance to all feature points in the real-time image is calculated, and the point with the smallest distance is selected as the initial matching point pair.

[0029] The initial matching point pairs are removed using a mismatch removal algorithm. This algorithm first calculates the ratio of the Euclidean distance to the average distance of neighboring points for each matching point pair, as well as the angular deviation of the lines connecting the matching points; pairs with excessively large ratios or angular deviations are removed. The remaining point pairs are treated as graph nodes. Edge weights are constructed based on the spatial distance and feature similarity between nodes. A clustering algorithm is used to divide the nodes into clusters, removing clusters smaller than a minimum cluster threshold. The largest cluster is selected as the interior set, and the initial transformation model is calculated accordingly.

[0030] Based on the inlier set, an improved particle swarm optimization (PSO) algorithm is used for subpixel displacement optimization. The PSO algorithm uses the subpixel displacement as the optimization variable and the fitness function is the weighted sum of image grayscale differences and feature descriptor differences. During particle initialization, some particles are initialized with a Gaussian distribution in their neighborhood using displacement estimates obtained from the initial transformation model, while the remaining particles are uniformly and randomly distributed across the entire search range. During iteration, the inertia weight and learning factor are dynamically adjusted based on the fitness improvement rate. When the global optimal fitness stagnates for a long period, a Gaussian perturbation is applied to the optimal particle, and all particles are reset to their current optimal neighborhood, performing a second fine-tuning search. After convergence, the algorithm outputs the optimal subpixel displacement parameters and simultaneously calculates the matching quality evaluation index. This index contains two components: the standard deviation of the matching residuals, i.e., the statistical dispersion of the final projection error of all inliers; and the convergence speed index, i.e., the average fitness decline rate in the last few generations. The smaller the standard deviation of the residuals and the faster the convergence, the higher the matching quality.

[0031] The compensation adjustment step compares the standard deviation of the matching residuals and the convergence rate index output by the particle swarm optimization algorithm with three pre-set thresholds to classify the current matching quality into three levels: excellent, medium, and poor. The system internally maintains a parameter priority matrix: the cluster similarity threshold has the highest adjustment priority, followed by the distance ratio threshold, then the minimum cluster size, and the angle deviation threshold has the lowest priority.

[0032] In the optimal case, only a minor adjustment is made to the cluster similarity threshold, without re-executing mismatch removal. In the intermediate case, both the cluster similarity threshold and the distance ratio threshold are adjusted simultaneously, with the adjustment magnitude determined by the degree to which the standard deviation of the matching residual deviates from the ideal value, triggering a re-execution of mismatch removal and updating the interior point set and transformation model. In the poor case, all four parameters are adjusted, with the minimum cluster size further reduced and the angle deviation threshold relaxed. Simultaneously, some particle reset information is fed back to the particle swarm optimization algorithm, triggering a re-execution of sub-pixel optimization. This feedback mechanism forms a closed-loop adjustment, enabling the algorithm to adaptively improve the matching results.

[0033] The image stitching step compares the change between the current transformation model and the transformation model obtained in the previous iteration. The change can be measured by the rate of change of variance of each element of the transformation matrix or the rate of change of the size of the inlier set. When the change is less than a preset threshold, the closed-loop feedback is considered to have converged. The final transformation model is then used to perform geometric transformation and fusion on the image, completing the image stitching with sub-pixel accuracy.

[0034] Specifically, the motion compensation step includes: calculating the blur index of the real-time image, the blur index being calculated based on the ratio of the gradient energy of the current image to the gradient energy of the standard clear image; converting the blur index into a frequency correction factor through a bounded nonlinear function; evaluating the motion speed signal to obtain an estimated speed; and calculating the scanning frequency at the next moment based on the estimated speed, the camera pixel spacing, and the frequency correction factor.

[0035] The blur index of a real-time image is obtained by calculating the ratio of the gradient energy of the current image to that of a standard clear image. Gradient energy reflects the drastic change in brightness within the image: when the object's speed matches the scanning frequency, the image edges are sharp and the gradient energy is high; speed mismatch results in motion blur or stretching, and the gradient energy decreases accordingly. This ratio is input into a bounded nonlinear transformation function to obtain a frequency correction factor. This function uses a hyperbolic tangent form, and its output value is limited to preset upper and lower limits to avoid excessive correction amplitude leading to drastic fluctuations in the scanning frequency. The motion speed signal is evaluated using a Kalman filter. The Kalman filter combines historical speed observations with current measurements to output a smooth and predictive speed estimate. This estimate effectively suppresses instantaneous noise and random jitter in the encoder signal. Finally, the estimated speed is divided by the camera's physical pixel spacing to obtain the theoretical matching frequency, which is then multiplied by the frequency correction factor, and an integral term representing the historical frequency error is added to construct the scanning frequency for the next moment. The integral term eliminates long-standing steady-state biases, allowing the scanning frequency to gradually approach the optimal value. The calculated frequency is written to the trigger controller of the line scan camera in real time, and the camera operates at the new frequency in subsequent scanning cycles, thereby suppressing image stretching or compression at the source.

[0036] Specifically, the geometric correction step includes constructing a geometric correction model based on the intrinsic and extrinsic parameters of the line scan camera. The intrinsic parameters include focal length, principal point, and distortion coefficients, while the extrinsic parameters include rotation matrix and translation vector. The geometric correction model is used to eliminate radial and tangential distortion in the acquired real-time image and to compensate for scanning angle deviation and camera physical installation position deviation, outputting a corrected distortion-free image.

[0037] The intrinsic parameters of the line scan camera include focal length, principal point coordinates, and radial and tangential distortion coefficients. These parameters are obtained by least-squares fitting after taking multiple photos of the calibration board in a laboratory environment. The extrinsic parameters include rotation matrices and translation vectors, describing the pose relationship between the camera coordinate system and the world coordinate system. The geometric correction model is constructed as follows: For each frame of image acquired in real time, the normalized image plane position of each pixel is calculated based on its original coordinates. Then, the radial and tangential distortion correction terms are substituted to obtain the distortion offset. The distortion offset is then superimposed onto the original coordinates to eliminate lens distortion. The radial distortion correction term uses the distance from the pixel to the principal point as the independent variable and adopts a polynomial form with second-order coefficients. The tangential distortion correction term considers the offset introduced by the imperfect perpendicularity between the lens optical axis and the imaging plane during lens module assembly. After lens distortion correction, an affine transformation is performed on the entire image using a rotation matrix and a translation vector to compensate for minor rotation angles and physical position offsets during camera installation, especially the angular deviation between the scan line direction and the object's motion direction, as well as projection distortion caused by the height difference between the camera and the conveyor belt. After this processing, the output image is free from barrel or pincushion distortion, and has no scan line tilt or misalignment. Its coordinate system is identical to that of the standard template image, facilitating subsequent matching operations.

[0038] Specifically, the image matching step further includes constructing a dual-branch deep convolutional network specifically designed for the characteristics of line scan camera images. The first branch is a spatial feature extraction branch, employing a multi-scale depthwise separable convolutional structure. Depthwise separable convolution decomposes standard convolution into two steps: depthwise convolution and pointwise convolution, significantly reducing the number of parameters. The multi-scale design means setting convolutional kernels of various sizes or using dilated convolution in the same layer, enabling the network to simultaneously capture small-scale texture details and large-scale contour structures. The second branch is a temporal feature extraction branch. Given that line scan cameras perform line-by-line scanning imaging, there is a temporal or spatial continuity between adjacent scan lines. This branch splits the input image into several row vector sequences, each row vector representing the grayscale value or intermediate feature value of a row. Subsequently, a one-dimensional convolution operation is applied along the extension direction of the scan line, with the convolutional kernel size covering several adjacent rows, thereby extracting grayscale variation patterns, motion residuals, or local misalignment features between rows. The feature maps output by the two branches have the same spatial size and are merged through an adaptive weighted fusion module. The weighted fusion module incorporates learnable weight parameters that automatically adjust during training, allowing the network to focus on branch features that contribute significantly to the matching task during fusion. The fused deep feature map carries both local spatial and inter-row temporal information, exhibiting strong responsiveness to object edges, textures, and subtle deformations along the scanning direction, providing more discriminative feature representations for subsequent feature description and matching.

[0039] Example 2: Specifically, the feature descriptor includes extracting the gradient direction histogram, gray-level distribution statistics, and local binary pattern texture features of a local region of the image, and then weighting and fusing the three feature vectors to obtain the feature descriptor.

[0040] Feature descriptors are used to encode visual information of local regions of an image into a numerical vector for subsequent comparison of the similarity of feature points in different images. This embodiment uses three complementary feature types for weighted fusion.

[0041] The process of constructing a gradient orientation histogram is as follows: Within a local window surrounding a feature point, the gradient magnitude and direction of each pixel are calculated; the orientation angle is divided into several intervals, for example, every twenty degrees, and the cumulative sum of the gradient magnitudes within each interval is calculated to form a histogram vector. This vector describes the dominant orientation distribution of the edges within a local region and is robust to changes in illumination.

[0042] Gray-scale distribution statistics include the mean and variance of pixel gray-scale values ​​within a local window. The mean reflects the overall brightness level of the area, while the variance reflects the contrast strength. These two statistics are simple to calculate and can quickly distinguish regions with different gray-scale characteristics.

[0043] Local binary mode texture features generate binary codes by comparing the grayscale values ​​of each pixel within a window with its neighboring pixels. Specifically, the grayscale value of the center pixel is used as a threshold to binarize the surrounding pixels. These binarized pixels are then arranged into binary numbers and converted to decimal, serving as the texture code for that center point. The frequency of occurrence of various codes within the entire window is then calculated to obtain the texture histogram. Local binary mode is insensitive to monotonic lighting changes and has high computational efficiency.

[0044] The three feature vectors mentioned above are weighted and fused: the gradient direction histogram vector, gray-level statistics, and local binary pattern histogram vector are multiplied by their respective weight coefficients and then concatenated or spliced ​​to form a high-dimensional combined feature descriptor. The weight coefficients can be learned from training samples or set empirically to reduce redundant information while maintaining discriminative power in the fused descriptor.

[0045] Specifically, such as Figure 2 As shown, after obtaining the feature points and their feature descriptors of the template image and the real-time image, it is necessary to establish the correspondence between the feature points. This embodiment uses a nearest neighbor matching strategy to complete this initial pairing.

[0046] For each feature point extracted from the template image, the system calculates the distance between its descriptor vector and the descriptor vectors of all feature points in the real-time image. The distance metric used is Euclidean distance, which is calculated by taking the square root of the sum of the squares of the differences between the corresponding components of the two vectors. Euclidean distance measures the true difference between two descriptors in the feature space; a smaller value indicates greater similarity between the two local regions.

[0047] After obtaining all distance values, the real-time image feature point corresponding to the minimum value is selected, and this point is paired with the feature points of the template image to form an initial matching point pair. This process is equivalent to finding the most similar real-time feature point for each template feature point. Considering that there may be multiple similar regions in an image, relying solely on nearest neighbors may result in false matches. Therefore, in practice, a verification of the second-nearest distance ratio is introduced. That is, the ratio of the minimum distance to the second-smallest distance is calculated. If this ratio is less than a preset threshold, usually 0.8, the matching point pair is accepted; otherwise, it is rejected. This ratio verification effectively eliminates fuzzy matches in regions with repetitive textures.

[0048] After completing the nearest neighbor search for all template feature points, an initial set of matching point pairs is output for use in subsequent mismatch removal steps.

[0049] Specifically, the error elimination algorithm includes the following: the initial set of matching point pairs often contains a large number of erroneous matches, and this embodiment uses a two-stage adaptive algorithm for elimination.

[0050] The first stage involves pre-screening based on geometric constraints. For each matching point pair, the Euclidean distance between the two points is calculated, and the average distance of other matching point pairs in their neighborhood is also calculated. The ratio of the former to the latter is taken as the distance ratio. If the distance ratio exceeds a first preset threshold, it indicates that the scale or position of the matching point pair is inconsistent with the surrounding points, and it is likely a mismatch. Simultaneously, the angle of the line connecting the feature points in the template image to the corresponding points in the real-time image is calculated, and the median of the angles connecting all point pairs is calculated. If the absolute value of the angle of a point pair deviating from the median exceeds a second preset threshold, it is considered an angle anomaly. If either of these two anomalies is met, the point pair is removed.

[0051] The second stage uses graph clustering for local consistency filtering. The point pairs retained in the first stage are considered nodes in the graph structure. For any two nodes, their spatial distance in image space and the similarity between their feature descriptors are calculated. The spatial distance and feature similarity are weighted and converted into edge weights using an exponential function. Nodes with close spatial distance and high feature similarity have higher edge weights, indicating they tend to belong to the same geometrically consistent region. After constructing the complete graph, a density-based clustering algorithm, such as a noise-based density-based spatial clustering algorithm, is used to divide the nodes into several clusters. The clustering algorithm requires specifying a similarity threshold; nodes with edge weights exceeding this threshold are grouped into the same cluster. A minimum cluster size threshold is also set; any cluster with fewer nodes than this threshold is considered noise or isolated mismatches and is directly discarded. Finally, the cluster with the largest number of nodes is selected as the inlier set from the remaining clusters. Matching point pairs in this inlier set have highly consistent geometric transformation relationships, suitable for estimating the final transformation model. The inlier set is output to the subsequent sub-pixel optimization module, providing a reliable initial solution space for the particle swarm optimization algorithm.

[0052] Example 3: Specifically, such as Figure 3 As shown, the particle swarm optimization algorithm includes an initial subpixel displacement estimate derived from the initial transformation model output by the upstream mismatch elimination step. This transformation model describes the overall mapping relationship from the template image to the real-time image, and by applying it to the image center point, an approximate subpixel displacement can be calculated. The particle swarm is divided into two subgroups: in the first subgroup, particle positions are randomly generated according to a Gaussian distribution, with the mean of the Gaussian distribution set as the initial subpixel displacement estimate, and the standard deviation set to a small value, such as 0.5 pixels, so that most particles cluster around the most promising region; in the second subgroup, particle positions follow a uniform random distribution within a preset search range, such as a range of plus or minus five pixels, maintaining the ability to explore the entire solution space. The ratio of the number of particles in the two parts can be preset, for example, seven to three.

[0053] The fitness function evaluates the contribution of the sub-pixel displacement represented by each particle to the matching. In this embodiment, the fitness value is the weighted sum of the squared gray-level differences within the local window plus the Euclidean distance of the feature descriptors; a smaller fitness value indicates a more accurate match. In each iteration, the global optimal fitness value, i.e., the minimum value in the entire particle swarm's history, is recorded. The relative decrease in the global optimal fitness between two adjacent generations is defined as the fitness improvement rate, calculated by subtracting the current generation's optimal value from the previous generation's optimal value and then dividing by the previous generation's optimal value. This improvement rate reflects the speed of the optimization process in real time.

[0054] Three key parameters—inertia weight, individual learning factor, and social learning factor—are dynamically adjusted based on the fitness improvement rate. When the improvement rate is high, indicating the current search direction is effective, the inertia weight is increased to maintain the particle's original motion trend, while the two learning factors are decreased to prevent premature convergence to local maxima. When the improvement rate is low, meaning the search has entered a flat region, the inertia weight is decreased to enhance local fine-grained search capabilities, while the learning factors are increased to encourage particles to move closer to their historical optimal position and the global optimal position of the population. This adjustment mechanism ensures that the parameters no longer change linearly according to a predetermined algebra, but rather adaptively adjust based on the actual search progress.

[0055] Set up a stagnation counter. After each iteration, check the change in the global optimal fitness. If the change is less than a preset minimum threshold for multiple consecutive iterations, the search is considered stagnant. At this point, a perturbation operation is performed: a Gaussian perturbation is applied to the position of the global optimal particle, with the perturbation amplitude increasing with the number of stagnant generations (the longer the stagnation, the stronger the perturbation); simultaneously, the positions of all particles are reinitialized to be within the neighborhood of the current global optimal position, the search range of each coordinate axis is reduced to half its original value, and the particle velocities are reset to smaller random values. This operation is equivalent to initiating a secondary refined search, which helps to escape the local optimum trap.

[0056] The convergence criteria for the algorithm can be reaching the maximum number of iterations, or the change in the global optimal fitness across multiple consecutive generations being below a threshold and the stall counter not triggering a perturbation. Upon convergence, the optimal sub-pixel displacement parameter is output, representing the position of the globally optimal particle. Simultaneously, a matching quality evaluation index is calculated and output, comprising two components. The standard deviation of the matching residual is calculated by collecting the projection errors of all inliers after adjustment with the optimal sub-pixel displacement and calculating the standard deviation of these error values. A smaller standard deviation indicates better matching consistency among inliers. The convergence speed index is calculated by recording the global optimal fitness values ​​of the last few generations (e.g., the last ten generations) of the particle swarm optimization algorithm, calculating the relative decline rate for each generation, and then averaging these decline rates. A larger average value indicates faster convergence, reflecting a relatively simple matching problem or a relatively accurate initial estimate.

[0057] Specifically, the compensation adjustment step receives the standard deviation of the matching residuals and the convergence rate index output by the particle swarm optimization algorithm as the basis for evaluating the current matching quality. The system pre-stores three sets of thresholds, corresponding to two critical values ​​for the residual standard deviation and two critical values ​​for the convergence rate index. By comparing the actual values ​​with these thresholds, the matching quality is divided into three levels: excellent, medium, and poor. Excellent indicates that the matching residual standard deviation is less than the lower threshold and the convergence rate is fast, indicating that the current matching is very accurate; medium indicates that one of the indicators is between the two critical values; poor indicates that the residual standard deviation exceeds the higher threshold or the convergence rate is too slow, indicating a relatively serious matching bias.

[0058] For the four adjustable parameters involved in the mismatch removal algorithm, an adjustment priority order was set. The cluster similarity threshold determines whether two nodes belong to the same cluster in the clustering algorithm; its change has the most significant impact on the quality of the final inlier set, therefore it has the highest priority. The distance ratio threshold controls the strictness of the first-stage geometric pre-screening, with the next highest priority. The minimum cluster size threshold determines the minimum number of valid clusters allowed to be retained, with the next lowest priority. The angle deviation threshold mainly affects a very small number of abnormal matches, with the lowest priority.

[0059] For superior matching quality, only a minor adjustment is made to the cluster similarity threshold. Specifically, if the residual standard deviation is slightly higher than the ideal target value, the threshold is slightly lowered to relax the clustering conditions and retain more point pairs; conversely, the threshold is slightly increased to make the clustering more stringent. The adjustment is very small and does not trigger a re-execution of mismatch removal, avoiding unnecessary computational overhead.

[0060] For intermediate matching quality, both the cluster similarity threshold and the distance ratio threshold are adjusted. The adjustment magnitude is positively correlated with the degree to which the standard deviation of the matching residuals deviates from the ideal value: the greater the deviation, the larger the adjustment magnitude. The two parameters are adjusted in the same direction, i.e., the cluster similarity threshold is lowered while the distance ratio threshold is also lowered, appropriately relaxing the conditions in both the pre-screening and clustering stages to accommodate more potential inliers. After adjustment, the mismatch removal process is re-executed, and the inlier set and transformation model are recalculated using the new parameters.

[0061] For poor matching quality, enhancement adjustments are performed. All four parameters are adjusted: the cluster similarity threshold and distance ratio threshold are significantly reduced; the minimum cluster size threshold is reduced to half its original value, but an absolute lower limit, such as three, is set to avoid model instability due to retaining too few points; the angle deviation threshold is appropriately relaxed, as large matching residuals are often accompanied by angle estimation errors. Two additional operations are performed: triggering a re-execution of mismatch removal, and simultaneously feeding back reset information of some particles to the particle swarm optimization algorithm. The feedback information includes the position of the current optimal particle and an indication of a narrowed search range. Based on this, the particle swarm optimization algorithm reinitializes some of its particles and begins a new sub-pixel optimization iteration. This two-way closed-loop mechanism enables the entire matching system to self-correct even in poor matching conditions.

[0062] After the above-mentioned tiered adjustment, the system will check the new matching quality index again. If it still does not meet the requirements and the number of feedbacks has not reached the upper limit, the adjustment process can be repeated. When the change in the transformation model is less than the preset threshold or the maximum number of feedbacks is reached, the compensation adjustment ends.

[0063] The foregoing has illustrated and described the basic features, principles, and advantages of the present invention. It should be noted that the present invention is not limited to the above embodiments, but only to some embodiments. Any improvements and additions made without departing from the spirit and scope of the present invention are considered to be within the scope of protection of the present invention.

Claims

1. An image matching method for a linear scan camera based on particle swarm feedback, characterized in that, include: The image acquisition step involves acquiring real-time images of the target object using a line scan camera and obtaining the object's motion speed signal. The geometric correction step involves constructing a geometric correction model based on camera parameters, performing geometric error correction on the real-time image to obtain the corrected real-time image, and generating an image pair with the pre-stored template image. The motion compensation step involves constructing a motion compensation model, adjusting the scanning frequency of the linear scan camera based on the motion speed signal and the image quality feedback of the real-time image, and then re-acquiring the real-time image. The image matching step involves extracting and describing features from the corrected image pairs, constructing feature descriptors and obtaining initial matching point pairs, using a mismatch elimination algorithm to initially eliminate the initial matching point pairs to obtain an inner point set and an initial transformation model, using a particle swarm optimization algorithm based on the inner point set to obtain the optimal sub-pixel displacement parameters, and calculating the matching quality evaluation index. The compensation and adjustment step involves adjusting the judgment parameters of the error elimination algorithm according to the matching quality evaluation index, and then re-performing the mismatch elimination based on the adjusted parameters, thereby updating the eliminated interior point set and transformation model. In the image stitching step, the magnitude of the change in the transformation model is compared with a preset threshold. When the change is less than the preset threshold, the image is stitched using the final transformation model. The particle swarm optimization algorithm includes: calculating initial sub-pixel displacement estimates based on an initial transformation model; dividing the particle swarm into two parts; initializing the first part within the neighborhood of the initial sub-pixel displacement estimates using a Gaussian distribution; and initializing the second part uniformly and randomly within a preset search range; setting the relative decrease in fitness between two adjacent generations as the fitness improvement rate; adjusting inertia weights, individual learning factors, and social learning factors based on the fitness improvement rate; setting a stagnation counter; determining stagnation when the change in global optimal fitness over multiple consecutive iterations is less than a preset threshold; applying a Gaussian perturbation to the global optimal particle, with the perturbation amplitude increasing with the number of stagnation generations; reinitializing all particles to their current optimal neighborhood; initiating a secondary fine search; and outputting the optimal sub-pixel displacement parameters after the algorithm converges. Simultaneously, it calculates and outputs a quality evaluation index, including the standard deviation of the matching residual and a convergence rate index. The standard deviation of the matching residual is the standard deviation of the final projection error of all interior points, and the convergence rate index is the average fitness decrease rate of the last few generations.

2. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The motion compensation step includes calculating the blur index of the real-time image, which is calculated based on the ratio of the gradient energy of the current image to the gradient energy of the standard clear image. The blur index is then converted into a frequency correction factor using a bounded nonlinear function. The motion speed signal is evaluated to obtain an estimated speed. The scanning frequency at the next moment is calculated based on the estimated speed, the camera pixel spacing, and the frequency correction factor.

3. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The geometric correction step includes constructing a geometric correction model based on the intrinsic and extrinsic parameters of the line scan camera. The intrinsic parameters include focal length, principal point, and distortion coefficients, and the extrinsic parameters include rotation matrix and translation vector. The geometric correction model is used to eliminate radial and tangential distortion of the acquired real-time image, and to compensate for scanning angle deviation and camera physical installation position deviation, and output a corrected distortion-free image.

4. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The image matching step further includes constructing a dual-branch deep convolutional network. The first branch is a spatial feature extraction branch, which encodes spatial features of the input image through multi-scale deep separable convolution. The second branch is a temporal feature extraction branch, which splits the image into a row vector sequence by row and extracts inter-row correlation features through one-dimensional convolution along the scan line direction. The feature maps output by the branch are then fused through adaptive weighting to obtain a deep feature map.

5. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The feature descriptor includes extracting the gradient direction histogram, gray-level distribution statistics, and local binary pattern texture features of a local region of the image, and then weighting and fusing the three feature vectors to obtain the feature descriptor.

6. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The image matching step further includes calculating the distance between each feature point extracted from the template image and the feature descriptors of all feature points in the real-time image using a nearest neighbor matching strategy, and selecting the feature point with the smallest distance as the initial matching point pair.

7. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, The error elimination algorithm includes: calculating the Euclidean distance ratio and the angle deviation of the connection for each matching point pair; eliminating point pairs whose distance ratio exceeds a first preset threshold or whose angle deviation exceeds a second preset threshold; using the remaining point pairs as nodes; calculating the edge weights between nodes based on spatial distance and feature descriptor similarity; constructing a graph structure; dividing the nodes into several clusters using a clustering algorithm; eliminating clusters smaller than the smallest cluster size; and selecting the largest cluster as the inner point set for output.

8. The image matching method for a linear scan camera based on particle swarm feedback according to claim 1, characterized in that, Based on the standard deviation of the matching residuals and the convergence speed index output by the particle swarm optimization algorithm, the matching quality is compared with three preset thresholds to classify the matching quality into three levels: excellent, medium, and poor. The adjustment priority order of multiple adjustable parameters is set, with the cluster similarity threshold having the highest priority, followed by the distance ratio threshold, then the minimum cluster size, and the angle deviation threshold having the lowest priority. When the matching is determined to be excellent, the cluster similarity threshold is adjusted; when it is determined to be medium, both the cluster similarity threshold and the distance ratio threshold are adjusted, and error matching is re-executed; if it is determined to be poor, all four parameters are adjusted, the minimum cluster size is reduced, and the angle deviation threshold is relaxed, triggering a re-execution of error matching removal. Some particle reset information is also fed back to the particle swarm optimization algorithm to trigger sub-pixel optimization.

Citation Information

Patent Citations

  • Mechanical arm calibration and control device and method based on particle swarm optimization

    CN116476046A

  • Geometric correction method for images acquired by rail transit line-scan digital camera

    CN119359601A