Track defense area offset correction method based on image registration and optimal path selection, electronic equipment and storage medium

Through the method of image registration and optimal path selection, the Fourier-Merlin transform and phase correlation method are used to solve the accuracy of zone offset correction of the track monitoring system in bad weather, achieving efficient and stable zone correction effect, which is suitable for high-speed rail track monitoring in complex environments.

CN120278928APending Publication Date: 2025-07-08SHANDONG ZHIYANG HUITONG DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510327541.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing orbital monitoring system is difficult to accurately correct the zone offset caused by camera displacement in severe weather and complex environments. The traditional methods are inefficient and costly, and the existing image matching technology has limited generalization capabilities under extreme conditions.

Method used

The image matching quality is calculated by using a method based on image registration and optimal path selection through fast Fourier transform, coordinate system conversion and phase correlation methods, and the optimal transformation path is used to optimize the zone correction, including preprocessing, Fourier-Merlin transform, image matching under logarithmic polar coordinate system and optimal path selection.

Benefits of technology

Efficiently and accurately correct the zone offset caused by camera displacement in complex environments, improving the stability and robustness of the track monitoring system, reducing false alarms and missed reports, and is suitable for application scenarios in various complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278928A_ABST
    Figure CN120278928A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image offset correction, and particularly relates to a track defense area offset correction method based on image registration and optimal path selection, electronic equipment and a storage medium. The method comprises the following steps: firstly, calculating offset and a normalized signal power value of a target image and a source image by using Fourier-Mellin transform, and if the value is greater than a set threshold value, correcting a defense area coordinate in the source image according to the offset to obtain a defense area coordinate in the target image; if the value is smaller than a set threshold value, a series of middle frames between the target image and the source image are extracted, Fourier-Mellin transformation is carried out in sequence to search an optimal transformation path, and the defense area coordinates in the source image can be corrected step by step according to the optimal transformation path to obtain the defense area coordinates in the target image. The method can efficiently and accurately correct defence area deviation caused by camera displacement in a high-speed rail monitoring system in various environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image offset correction, and more specifically, relates to a method for offset correction of track defense areas based on image registration and optimal path selection, an electronic device, and a storage medium. Background Art

[0002] As an important part of the modern transportation system, the safety and reliability of high-speed railways are of crucial importance. To ensure the safe operation of high-speed trains, a large number of monitoring cameras are installed along the line to monitor the track status, train running conditions, and changes in the surrounding environment in real time. These monitoring data are of great significance for timely detecting potential threats (such as foreign object intrusion, track damage, etc.). However, in practical applications, due to the influence of various factors, such as wind force, temperature changes, and vibration, the monitoring cameras installed at fixed positions may undergo slight displacement or angular deflection, resulting in a deviation of the preset defense area in the captured image, that is, the so-called "defense area offset". This offset not only affects the accuracy of the monitoring system but also may lead to false alarms or missed reports of important events, posing a safety hazard to the operation of high-speed railways.

[0003] Current track monitoring systems usually rely on fixed monitoring points for continuous shooting and use image analysis technology to identify abnormal situations. However, during long-term use, it is inevitable for the monitoring cameras to undergo small displacements. Even a very small angular adjustment may cause a significant change in the content of the image, making the originally set defense area no longer accurately correspond to the actual track area. This situation is particularly obvious under complex weather conditions, such as strong winds, heavy rains, and snow disasters, because harsh weather will exacerbate the aging and wear of the equipment and increase the risk of camera displacement. In addition, the traditional method to solve such problems mainly relies on manual regular inspection and correction of the camera position. This method is not only inefficient but also costly and cannot meet the requirements of modern intelligent monitoring.

[0004] Currently, some automatic correction solutions using Fourier-Mellin transform mainly focus on image stabilization technology in static scenes, such as in patents like "Image Matching Method, Device, Computer Equipment, and Computer Readable Storage Medium (CN116563357B)" and "A Dynamic Visual SLAM Method Based on Frequency Domain and Semantics (CN116524026A)". These methods only perform registration on the target image and the source image through Fourier-Mellin transform and are applicable to relatively stable perspective change scenarios. However, for application scenarios such as high-speed railway track monitoring that require high-precision positioning and are often affected by extreme weather, the effects of these methods are often unsatisfactory.

[0005] Especially under adverse weather conditions and at night, how to effectively compensate for image differences caused by environmental factors has become an urgent problem to be solved. At the same time, although existing deep learning-based image matching technologies have made remarkable progress in many fields, their generalization ability is still limited due to the lack of sufficient training samples to cover all possible situations. Especially when faced with unseen adverse weather or night lighting conditions, the performance of the model may be greatly reduced, and it cannot provide reliable anti-region offset correction services.

[0006] In view of this, the present invention designs an anti-region offset correction method for rail tracks based on image registration and optimal path selection to solve the problem that the prior art cannot accurately achieve anti-region correction in the face of a dynamically changing complex environment. Summary of the Invention

[0007] The present invention aims to overcome at least one defect of the above prior art and provides an anti-region offset correction method for rail tracks based on image registration and optimal path selection. This method can efficiently and accurately correct the anti-region offset caused by camera displacement in the high-speed rail track monitoring system under various environments.

[0008] The present invention also discloses an electronic device for implementing the above method.

[0009] The present invention also discloses a computer-readable storage medium for implementing the above method.

[0010] The detailed technical solution of the present invention is as follows:

[0011] An anti-region offset correction method for rail tracks based on image registration and optimal path selection, the method comprising:

[0012] S1. Preprocess the target image and the source image;

[0013] S2. Apply the fast Fourier transform to convert the preprocessed grayscale image to a frequency-domain representation, and then use a high-pass filter to filter out high-frequency noise in the image while retaining the low-frequency components that reflect the main structure of the image;

[0014] S3. Convert the image processed in step S2 from the Cartesian coordinate system to the logarithmic polar coordinate system;

[0015] S4. Use the phase correlation method to perform the Mellin transform on the source image and the target image respectively in the logarithmic polar coordinate system, and calculate the translation amount and the normalized signal power value;

[0016] S5. Judge the image matching quality based on the normalized signal power value. If the matching quality is high, transform the anti-region in the source image to obtain the anti-region coordinates in the target image;

[0017] S6. If the matching quality is low, adjust the coordinates of the defense area in the target image by calculating the optimal transformation path.

[0018] Preferably according to the present invention, in step S1, the preprocessing refers to:

[0019] First, convert the color image into a grayscale image, and then apply a Hanning window to smooth the edges of the image.

[0020] Preferably according to the present invention, the conversion from the Cartesian coordinate system to the log-polar coordinate system in step S3 means mapping the frequency-domain image F(u, v) from the Cartesian coordinates (u, v) to the log-polar coordinates (ρ, θ), specifically as follows:

[0021] S31. Calculate the angular axis θ of the log-polar coordinates k :

[0022]

[0023] In formula (1), N θ represents the number of sampling points, with one sampling point per degree from 0 to 2π; θ k is the angular axis, which is the k-th angular value obtained by evenly dividing the entire circumference from 0 to 2π into N θ parts, in radians; k represents the discrete index of the angle, which is an integer with a value range of 0 to N θ -1.

[0024] S32. Calculate the minimum radial distance and the maximum radial distance:

[0025] The size of the image is M×N, and the center point (u0, v0) = (M / 2, N / 2). The minimum radial distance r min is 1, and the maximum radial distance r max is the distance from the center to the corner of the image:

[0026]

[0027] S33. Take the logarithm of the radial distance from r min to r max to obtain the range of the logarithmic radial axis ρ m :

[0028] ρ min = log(r min ), ρ max = log(r max ).

[0029] The number of its sampling points N ρ is min(M×N), and the value of each radial axis ρ m is:

[0030]

[0031] In Equation (3), ρ m represents the m-th discrete logarithm radial value, which is a sampling point on the radial axis and corresponds to the logarithmic radial coordinate in the log-polar coordinate system. It is the value obtained by taking the logarithm of the radial distance r and reflects the scale information of the image; N ρ represents the total number of sampling points on the logarithmic radial axis; m represents the discrete index of the logarithmic radial and is an integer with a value range from 0 to min(M×N)-1;

[0032] Through the above steps, an N θ ×N ρ grid can be obtained, and each point corresponds to a log-polar coordinate point (ρ m , θ k );

[0033] S34. Calculate the Cartesian coordinates (u m , v k ) in the original image corresponding to each log-polar coordinate point (ρ m,k , θ m,k ) through inverse mapping. (u m,k , v m,k ) indicates that the new coordinate point falls at a certain position in the original image. The specific method is as follows:

[0034] First, restore ρ m to the radial distance r m : r m = e ρm , and then calculate the corresponding Cartesian coordinates according to the angle θ k :

[0035] u m,k = u0 + r m ·cos(θ k ) (4)

[0037] v m,k = v0 + r m ·sin(θ k ) (5)

[0039] In Equations (4) and (5), u m,k and v m,k respectively represent the Cartesian horizontal and vertical coordinates corresponding to the log-polar coordinate (ρ m , θ k ); u0 and v0 represent the central horizontal and vertical coordinates of the frequency-domain image. For an image with a size of M×N, the center point (u0, v0) = (M / 2, N / 2), cos(θ k) represents the cosine value of the angle θ k and sin(θ k ) represents the sine value of the angle θ k ;

[0040] S35. If the obtained value of (u m,k v m,k ) is a decimal, then bilinear interpolation is required to estimate the pixel value at this position. The specific method is as follows:

[0041] First, find the four integer grid points (u1, v1), (u1, v2), (u2, v1), (u2, v2) closest to (u m,k v m,k );

[0042] Then, based on the values of these points in the original image, perform sum and distance weighted averaging to calculate the log-polar coordinate value F LP (ρ m , θ k ), and then fill this value back into (ρ m , θ k ):

[0043]

[0044] In Equation (6), w i,j represents the weight of bilinear interpolation and indicates the contribution of the four integer grid points to the target point;

[0045] S36. Repeat the above steps for all log-polar coordinate points (ρ m , θ k ), and finally obtain the log-polar coordinate image F LP (ρ, θ).

[0046] According to the preference of the present invention, the detailed steps of step S4 are as follows:

[0047] S41. Use the phase correlation method to find the peak position between the target image and the source image:

[0048] First, calculate the cross-power spectrum of the target image and the source image, which is defined as:

[0049]

[0050] In Equation (7), F1 represents the frequency domain representation of the source image in the log-polar coordinate system, F1 = F LP1 (ρ, θ); F2 represents the frequency domain representation of the target image in the log-polar coordinate system, F2 = F LP2 (ρ, θ); the superscript * represents the complex conjugate, and C(ρ, θ) represents the cross-power spectrum of the target image and the source image.

[0051] Then, the phase correlation function \(R(\rho,\theta)\) is calculated by performing an inverse Fourier transform on the cross-power spectrum:

[0052] \(R(\rho,\theta)=F\) -1 (C(\rho,\theta))(8)

[0053] In Equation (8), \(F\) -1 represents the inverse Fourier transform. After the inverse transform, the obtained phase correlation function \(R(\rho,\theta)\) is a function in the two-dimensional spatial domain, and its value represents the degree of coincidence between the target image and the source image in the spatial domain; actually, \(R(\rho,\theta)\) is a peak response function, and the peak position corresponds to the rotation information between the target image and the source image, and its peak position \((\Delta\rho,\Delta\theta)\) corresponds to the rotation angle of the target image relative to the source image;

[0054] S42. In the log-polar coordinate system, the scale difference is mainly reflected in the radial direction of the frequency. Due to the characteristics of the logarithmic coordinates, the scale difference is manifested as a linear change in the radius. Therefore, the scaling scale ratio of the image is deduced from the difference in the frequency domain radius: For two images with frequency domain radii \(r_1\) and \(r_2\) respectively, the scaling ratio Scale is obtained by calculating the ratio between them:

[0055]

[0056] S43. Use the obtained rotation angle and scaling scale ratio to complete geometric correction; after completing the geometric correction of rotation and scaling scale, use the phase correlation method again. First, calculate the cross-power spectrum in the Cartesian coordinate system, and then perform an inverse Fourier transform to obtain the phase correlation function \(R(\rho,\theta)\). The peak position \((\Delta\rho,\Delta\theta)\) of the phase correlation function in the Cartesian coordinate system indicates the translational offset between the two images;

[0057] S44. With the peak point as the center, select a \(5\times5\) neighborhood to calculate the normalized signal power value:

[0058] First, calculate the average power \(P\) raw of the neighborhood as follows:

[0059]

[0060] In Equation (10), \(N\) represents the number of neighborhood pixels; \((\rho,\theta)\) represents the coordinates of each point in the neighborhood in the log-polar coordinate system. Then, calculate the normalized signal power value \(P\) norm based on the average power of the neighborhood as follows:

[0061]

[0062] Output the normalized signal power value within a partial neighborhood centered on the peak point. This metric can be used as the basis for evaluating the registration quality and is subsequently used to optimize the selection of the optimal transformation link during multi-image sequence registration.

[0063] Preferably according to the present invention, in step S5, the result of the transformation in S4 is evaluated based on the normalized signal power value, and it is decided whether to accept this transformation. The detailed steps are as follows:

[0064] S51. Set an acceptance threshold and judge the magnitude relationship between the calculated normalized signal power and this threshold;

[0065] S52. If the normalized signal power value is greater than or equal to this threshold, it means that the matching quality between the current images is relatively high. Directly use the rotation, scaling, and translation offsets obtained in step S4 to correct the sector coordinates in the source image, thereby determining the correct sector position in the target image;

[0066] S53. If the normalized signal power value is lower than this threshold, it indicates that the matching quality is poor. At this time, enter step S6 to find the optimal transformation link to further optimize the registration result.

[0067] Preferably according to the present invention, the detailed steps of step S6 are as follows:

[0068] S61. Divide the time axis between the source image and the target image into one video frame per second.

[0069] S62. For each pair of adjacent frames i and j, calculate the normalized signal power value according to the Fourier-Mellin transform;

[0070] S63. Calculate the path score S(p) of each path p from the source image i ori to the target image i k . The path score S(p) is the product of the normalized signal power values between all adjacent frames on the path. A path containing frames i ori , i1, i2,..., i n , i k has the following path score:

[0071] S(p) = P(i ori , i1) × P(i1, i2) ×... × P(i n , i k )(12)

[0072] S64. Select the path with the highest score among all paths as the optimal path, that is:

[0073]

[0074] In formula (13), p(ik , i ori ) is the set of all paths from the source image i ori to the target image i k ;

[0075] S65. Correct the sector coordinates in the source image step by step to the target image according to the optimal path.

[0076] In another aspect of the present invention, an electronic device is further provided, including:

[0077] At least one processor; and

[0078] A memory storing instructions, when the instructions are executed by the at least one processor, causing the at least one processor to execute an orbit sector offset correction method based on image registration and optimal path selection as described above.

[0079] In another aspect of the present invention, a machine-readable storage medium is further provided, which stores executable instructions, and when the instructions are executed, the machine is caused to execute an orbit sector offset correction method based on image registration and optimal path selection as described above.

[0080] Compared with the prior art, the beneficial effects of the present invention are:

[0081] The present invention adopts a multi-step optimization strategy. First, the offset is calculated, and it is judged whether to accept the transformation according to the normalized signal power value. If the signal power value is lower than the set threshold, a series of intermediate frames are constructed as a bridge to find the optimal transformation path. Through this path optimization method, even in the case of poor initial matching quality or large differences, it can effectively compensate for the sector offset caused by camera displacement or angle change, ensuring the stability and accuracy of the orbit monitoring system. At the same time, it also improves the robustness and adaptability of the system, and is applicable to application scenarios under various complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 is a flowchart of the orbit sector offset correction method based on image registration and optimal path selection according to the present invention.

[0083] Figure 2 is an effect diagram of sector offset correction of a small-angle offset image using the Fourier-Mellin transform in an embodiment of the present invention.

[0084] Figure 3 is an effect diagram of sector offset correction of a large-angle offset image using the optimal transformation path in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.

[0086] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0087] Embodiment 1

[0088] Refer Figure 1 , this embodiment provides an automatic correction method for track zone offset based on image registration and optimal path selection. The method first processes the target image and the source image using Fourier-Mellin transform to calculate the translation amount and the normalized signal power value between them. Based on this value, it is judged whether to perform zone offset correction. If the signal power value is low, it indicates that the matching quality between the images is poor. In this case, in order to more accurately implement zone offset correction, the present invention proposes a strategy: by analyzing the consecutive frames between the target image and the source image, a series of intermediate frames are identified as bridges, thereby constructing an optimal transformation path from the source image to the target image. Once this optimal path is determined, the zone coordinates in the target image can be adjusted according to it. The selection of this optimal path is based on the principle of maximizing the Fourier-Mellin response value on the transformation path, ensuring that even in the case of low-quality matching, the most suitable transformation method can be found. Specifically as follows:

[0089] S1. Preprocess the target image and the source image;

[0090] In the preprocessing stage of step S1, in order to ensure the accuracy and efficiency of subsequent processing, a series of preprocessing operations need to be performed on the target image and the source image. First, convert the color image to a grayscale image because a color image usually has multiple channels (such as RGB). In many image processing tasks, a grayscale image is sufficient to capture the structural information of the image, and the color information often does not directly affect the subsequent processing results. After converting the color image to a grayscale image, the image will have only one intensity channel, so the computational amount of subsequent processing can be reduced; then apply the Hanning window. The Hanning window is a smoothing window function commonly used to reduce the edge effect of the signal during frequency domain conversion. When performing Fourier transform, mutations occur at the edges of the image, resulting in artifacts and unnatural high-frequency components in the frequency domain. By applying the Hanning window function, the edges of the image can be smoothed and these adverse effects can be reduced.

[0091] S2. Apply the fast Fourier transform;

[0092] The preprocessed image is transformed to the frequency domain representation using the Fast Fourier Transform (FFT) algorithm. FFT is an efficient method for computing the Discrete Fourier Transform (DFT). Through FFT, an image can be transformed from the spatial domain to the frequency domain representation, that is, from pixel values to frequency components. The low-frequency part in the frequency domain usually contains the overall structural information of the image, while the high-frequency part is distributed at the edges of the spectrum, corresponding to details, edges, or noise in the image. This step helps to extract the main structural features of the image and provides a basis for subsequent scale and translation calculations. In frequency domain analysis, high-frequency components are usually closely related to noise and details in the image, while low-frequency components carry the overall structural information of the image. By using a high-pass filter, high-frequency noise in the image can be effectively filtered out while retaining the low-frequency components that reflect the main structure of the image. This process helps to reduce the interference of noise and thus improve the accuracy of subsequent image processing steps.

[0093] S3. Coordinate system transformation;

[0094] To facilitate the calculation of scale and translation parameters, the image after FFT processing is transformed from the Cartesian coordinate system to the log-polar coordinate system. The frequency distribution in the Cartesian coordinate system (x, y) presents a rectangular region, while the log-polar coordinate system transforms this rectangular region into the form of a polar coordinate system, that is, represented by angle and radius. In the log-polar coordinate system, scale transformation is manifested as radial stretching, and translation transformation is manifested as angular translation. This transformation enables scaling and translation operations to be achieved through simple addition operations, thus simplifying the subsequent calculation process. The following is a detailed description of the transformation of the frequency domain image from Cartesian coordinates to log-polar coordinates:

[0095] After high-pass filtering, the low-frequency components of the frequency domain image F(u, v) are located at the center. The goal of the log-polar coordinate transformation is to map F(u, v) from Cartesian coordinates (u, v) to log-polar coordinates (ρ, θ). Assume the size of the image is M×N, where the center point is (u0, v0) = (M / 2, N / 2), and the minimum radial distance r min is set to 1, and the maximum radial distance r max is set to the distance from the center to the corner of the image:

[0096]

[0097] The angular axis θ k The number of sampling points N θ is set to one sampling point per degree from 0 to 2π, that is, 360. Then the calculation of the angular axis θ is as follows:

[0098]

[0099] where, N θIndicates the number of sampling points, with one sampling point per degree from 0 to 2π; θ k is the angle axis, which divides the entire circumference from 0 to 2π evenly into N θ parts, and the k-th angle value is obtained, with the unit of radians; k represents the discrete index of the angle, which is an integer and ranges from 0 to N θ -1.

[0100] Take the logarithm of the radial distance from r min to r max to obtain the range of the logarithmic radial axis ρ:

[0101] ρ min = log(r min ), ρ max = log(r max ). The number of sampling points N ρ is set to min(M×N), and the value of each radial axis ρ m is:

[0102]

[0103] where ρ m represents the m-th discrete logarithmic radial value, which is a sampling point on the radial axis, corresponding to the logarithmic radial coordinate in the logarithmic polar coordinate system. It is the value obtained by taking the logarithm of the radial distance r and reflects the scale information of the image; N ρ represents the total number of sampling points on the logarithmic radial axis; m represents the discrete index of the logarithmic radial, which is an integer and ranges from 0 to min(M×N)-1;

[0104] Through the above steps, a grid of N θ ×N ρ can be obtained, and each point corresponds to a (ρ m , θ k ). Since the image data is stored on an integer grid of (u, v), it is necessary to calculate the original coordinates (u m , v k ) corresponding to each logarithmic polar coordinate point (ρ m,k , θ m,k ) through inverse mapping. The specific method is:

[0105] First, restore ρ m to the radial distance r m : r m = e ρm

[0106] Then, calculate the corresponding Cartesian coordinates according to the angle θ k :

[0107] u m,k = u0 + r m·cos(θ k )

[0108] v m,k = v0 + r m ·sin(θ k )

[0109] These are usually (u m,k v m,k ) and are decimal numbers, indicating that the new coordinate point falls at a certain position in the original image. This value is not necessarily an integer.

[0110] Therefore, bilinear interpolation is used to estimate the pixel value at each (u m,k v m,k ) position from F(u, v) as follows:

[0111] Find the four integer grid points near (u m,k v m,k ), such as (u1, v2), (u1, v2), (u2, v1), (u2, v2).

[0112] Calculate the log-polar coordinate value of this point by weighted average of the values at these points in the original image according to the sum and distance

[0113] F LP (ρ m , θ k ), and then fill this value back into (ρ m , θ k ). Where w i,j represents the weight of bilinear interpolation, indicating the contribution of the four integer grid points to the target point.

[0114]

[0115] Repeat the above steps for all log-polar coordinate points (ρ m , θ k ), and finally obtain the log-polar coordinate image F LP (ρ, θ).

[0116] S4. Calculate the translation amount and the normalized signal power value for transformation evaluation;

[0117] Perform logarithmic polar coordinate decoupling geometric transformation in the logarithmic polar coordinate system, that is, the Mellin transform. Use the phase correlation technique of OpenCV to find the peak response position between two images, and thus calculate the rotation and scaling differences between them. The phase correlation technique is based on the phase information of the Fourier transform. The phase of an image contains information about the spatial position of the image, while the amplitude contains information about the intensity distribution of the image. When comparing two images, the amplitude information is often affected by noise and brightness changes, but the phase information is invariant to geometric transformations such as translation, rotation, and scaling. Determine the geometric transformation between two images by calculating the phase difference of the complex numbers in the frequency domain of the two images. In the logarithmic polar coordinate system, the distribution of frequency information changes, and the rotation transformation will be manifested as a translation in terms of angle. This transformation makes the rotation difference more obvious. Therefore, the rotation information of the image can be obtained by comparing the phase differences of the frequency domains of the two images.

[0118] Specifically, first calculate the cross-power spectrum of the target image and the source image, which is defined as:

[0119]

[0120] where F1 represents the frequency domain representation of the source image in the logarithmic polar coordinate system, F1 = F LP1 (ρ,θ); F2 represents the frequency domain representation of the target image in the logarithmic polar coordinate system, F2 = F LP2 (ρ,θ). The superscript * represents the complex conjugate, and C(ρ,θ) represents the cross-power spectrum of the target image and the source image.

[0121] Then calculate the phase correlation function R(ρ,θ) by performing the inverse Fourier transform on the cross-power spectrum:

[0122] R(ρ,θ) = F -1 (C(ρ,θ))

[0123] where F -1 represents the inverse Fourier transform. After the inverse transform, the obtained phase correlation function R(ρ,θ) is a function in the two-dimensional spatial domain, and its value represents the degree of coincidence of the target image and the source image in the spatial domain; actually, R(ρ,θ) is a peak response function, and the peak position corresponds to the rotation information between the target image and the source image, and its peak position (Δρ,Δθ) corresponds to the rotation angle of the target image relative to the source image;

[0124] Similar to rotation, the scaling transformation also affects the performance of the frequency-domain image. In the log-polar coordinate system, the scale difference is mainly reflected in the radial direction of the frequency. Due to the characteristics of the logarithmic coordinates, the scale difference is manifested as a linear change in the radius. Therefore, the scaling ratio of the image can be deduced from the difference in the frequency-domain radius. For example, if the frequency-domain radii of two images are r1 and r2 respectively, then the scaling ratio Scale can be obtained by calculating the ratio between them:

[0125]

[0126] After completing the geometric correction of rotation and scale, the cross-power spectrum is calculated again in the Cartesian coordinate system using the phase correlation method in the same way, and then the inverse Fourier transform is performed to obtain the phase correlation function R(ρ,θ), whose value is between [-1,1]. The peak position (Δρ,Δθ) of the phase correlation function in the Cartesian coordinate system indicates the translational offset between the two images. Taking the peak point as the center, the normalized signal power value is calculated for a 5×5 neighborhood. It is usually normalized to the range [0,1], which represents the confidence or quality of registration. 0 means that the two images are completely uncorrelated and the matching fails. 1 means that the two images are perfectly matched, with a clear peak and no noise interference. The average power of the neighborhood is calculated as follows:

[0127]

[0128] where N represents the number of neighborhood pixels; (ρ,θ) represents the coordinates of each point in the neighborhood in the log-polar coordinate system.

[0129] The normalized signal power value is as follows:

[0130]

[0131] Then, the normalized signal power value (in the range of 0-1) in a partial neighborhood centered on the peak point is output. This index can be used as the basis for evaluating the registration quality and is subsequently used to optimize the selection of the optimal transformation link when registering multiple image sequences.

[0132] S5. Judge the image matching quality based on the normalized signal power value; if the matching quality is high, transform the defense area in the source image to obtain the defense area coordinates in the target image;

[0133] Evaluate the result of the Fourier-Mellin transform based on the normalized signal power value and decide whether to accept the transform. Specifically, set an acceptance threshold of 0.65. If the calculated normalized signal power value is greater than this threshold, it is considered that the matching quality between the current images is high, and the rotation, scaling, and translation offsets obtained in step S4 can be directly used to correct the sector coordinates in the source image, thereby determining the correct sector position in the target image; if the normalized signal power value is below 0.65, it indicates that the matching quality is poor, and there may be large differences or noise interference. At this time, enter step S6 to find the optimal transformation link to further optimize the registration result.

[0134] Example results are as Figure 2 shown. After the above steps, the obtained normalized signal power value is 0.85, then Figure 2 the sector coordinates of the target image in

[0135] S6. If the matching quality is low, adjust the sector coordinates in the target image by calculating the optimal transformation path.

[0136] Divide the time axis between the source image and the target image into multiple small segments, one video frame per second. Assuming the total duration of the video is T seconds, there are a total of T pictures. Apply the Fourier-Mellin transform to the image frames and evaluate the transformation results of each frame to provide a score for path optimization. Select the path with the highest score among all possible paths as the optimal path, and gradually correct the sector coordinates in the source image to the target image.

[0137] Specifically, for each pair of adjacent frames i and j, calculate the normalized signal power value according to the Fourier-Mellin transform. This value is a numerical value between 0 and 1, indicating the reliability or quality of the transform. The higher the response value, the more accurate the transform. For each possible path p from the source image i ori to the target image i k , the score S(p) of the path is the product of the normalized signal power values between all adjacent frames on the path. If the path p contains frames i ori , i1, i2,..., i n , i k , then the score of this path is calculated as:

[0138] S(p) = P(i ori , i1) × P(i1, i2) ×... × P(i n , i k )

[0139] Then the optimal path P * is the path with the highest score among all possible paths. That is:

[0140]

[0141] Among them, p(i k , i ori ) is the set of all possible paths from the source image i ori to the target image i k .

[0142] Once the optimal transformation path is calculated, the coordinates of the defense area in the source image can be corrected step by step to obtain the coordinates of the defense area in the target image. Through this systematic path optimization method, even in the case of poor initial matching quality or large differences, it can effectively compensate for the defense area offset caused by camera displacement or angle change, ensuring the stability and accuracy of the track monitoring system. This method significantly reduces the situations of false alarms and missed alarms, providing a more reliable safety guarantee for high-speed rail operation. At the same time, it also improves the robustness and adaptability of the system, and is applicable to various application scenarios in complex environments.

[0143] Example results are as Figure 3 shown, demonstrating how to obtain the final defense area coordinates through the optimal transformation path. A total of three calculations are required for its optimal change path. First, the red defense area in the source image is sequentially transformed into the blue defense area in the intermediate transition image, and finally into the green defense area in the target image.

[0144] Embodiment 2

[0145] This embodiment also provides an electronic device, including:

[0146] At least one processor; and

[0147] A memory, the memory stores instructions, when the instructions are executed by the at least one processor, the at least one processor is caused to execute the method for correcting the track defense area offset based on image registration and optimal path selection as described above.

[0148] In this embodiment, the electronic device may include but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and so on.

[0149] Embodiment 3

[0150] This embodiment also provides a machine-readable storage medium, which stores executable instructions, and when the instructions are executed, the machine is caused to execute the method for correcting the track defense area offset based on image registration and optimal path selection as described above.

[0151] Specifically, a system or device equipped with a readable storage medium can be provided. On this readable storage medium, software program codes for implementing the functions of any one of the above embodiments are stored, and the computer or processor of the system or device is made to read and execute the instructions stored in the readable storage medium.

[0152] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0153] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0154] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An orbital defense zone offset correction method based on image registration and optimal path selection, characterized in that The method includes the following: S1. Preprocess the target image and the source image; S2. Apply the fast Fourier transform to convert the preprocessed grayscale image to a frequency-domain representation, and then use a high-pass filter to filter out the high-frequency noise in the image while retaining the low-frequency components that reflect the main structure of the image; S3. Convert the image processed in step S2 from the Cartesian coordinate system to the log-polar coordinate system; S4. Use the phase correlation method to perform the Mellin transform on the source image and the target image respectively in the log-polar coordinate system, and calculate the translation amount and the normalized signal power value; S5. Judge the image matching quality based on the normalized signal power value. If the matching quality is high, transform the defense area in the source image to obtain the defense area coordinates in the target image; S6. If the matching quality is low, adjust the defense area coordinates in the target image by calculating the optimal transformation path.

2. The orbit defense zone offset correction method based on image registration and optimal path selection according to claim 1, wherein In step S1, the preprocessing refers to: first convert the color image to a grayscale image, and then apply a Hanning window to smooth the edges of the image.

3. A method for correcting the offset of an orbital defense area based on image registration and optimal path selection according to claim 1, characterized in that The conversion from the Cartesian coordinate system to the log-polar coordinate system in step S3 means mapping the frequency-domain image F(u, v) from the Cartesian coordinates (u, v) to the log-polar coordinates (ρ, θ), specifically as follows: S31. Calculate the angular axis θ of the logarithmic polar coordinates k :[[-END]] In Equation (1), N θ represents the number of sampling points, with one sampling point per degree from 0 to 2π; θ k is the angle axis, which is obtained by evenly dividing the entire circumference from 0 to 2π into N θ parts, and is the k-th angle value, with the unit of radian; k represents the discrete index of the angle and is an integer, ranging from 0 to N θ - 1; S32. Calculate the minimum radial distance and the maximum radial distance: The size of the image is M×N, the center point (u0, v0) = (M / 2, N / 2), and the minimum radial distance r min is 1, and the maximum radial distance r max is the distance from the center to the corner of the image: S33. Take the logarithm of the radial distance from r min to r max to obtain the logarithmic radial axis ρ m in the range of: ρ min = log(r min ), ρ max = log(r max ); The number of sampling points N ρ is min(M×N), and the value of each radial axis ρ m is as follows: ρ m represents the m-th discrete logarithm radial value, which is a sampling point on the radial axis, corresponding to the logarithmic radial coordinate in the log-polar coordinate system. It is the value obtained by taking the logarithm of the radial distance r and reflects the scale information of the image; N ρ represents the total number of sampling points on the logarithmic radial axis; m represents the discrete index of the logarithmic radius and is an integer with a value range from 0 to min(M×N) - 1; An N can be obtained through the above steps θ ×N ρ grid, where each point corresponds to a log-polar coordinate point (ρ m , θ k ); S34. Calculate the Cartesian coordinates (u m , θ k ) in the original image corresponding to each log-polar coordinate point (ρ m,k , θ m,k ). (u m,k , v m,k ) indicates that the new coordinate point falls at a certain position in the original image. The specific method is as follows: First, reduce ρ m to the radial distance r m : Then, calculate the corresponding Cartesian coordinates according to the angle θ k : u m,k = u0 + r m · cos(θ k ) (4) v m,k = v0 + r m ·sin(θ k ) (5) In formulas (4) and (5), u m,k and v m,k respectively represent the Cartesian horizontal and vertical coordinates corresponding to the log-polar coordinates (ρ m , θ k ); u0 and v0 represent the central horizontal and vertical coordinates of the frequency-domain image. For an image with dimensions M×N, the center point (u0, v0) = (M / 2, N / 2), cos(θ k ) represents the cosine value of the angle θ k , and sin(θ k ) represents the sine value of the angle θ k . S35. If the obtained value of (u m,k v m,k ) is a decimal number, bilinear interpolation needs to be used to estimate the pixel value at this position. The specific method is as follows: First, find the four integer grid points (u1, v1), (u1, v2), (u2, v1), (u2, v2) that are closest to (u m,k v m,k ); Then, based on the values of these points in the original image, perform sum and distance weighted averaging to calculate the log-polar coordinate value F of this point LP (ρ m , θ k ), and then fill this value back into (ρ m , θ k ): In formula (6), w i,j represents the weight of bilinear interpolation, indicating the contribution of four integer grid points to the target point; S36. Repeat the above steps for all the log-polar points (ρ m , θ k ), and finally obtain the log-polar image F LP (ρ, θ).

4. A method for correcting the offset of an orbital defense zone based on image registration and optimal path selection according to claim 1, characterized in that, The detailed steps of step S4 are as follows: S41. Use the phase correlation method to find the peak position between the target image and the source image: First, calculate the cross-power spectrum of the target image and the source image, which is defined as: In Equation (7), F1 represents the frequency domain representation of the source image in the log-polar coordinate system, F1 = F LP1 (ρ, θ); F2 represents the frequency domain representation of the target image in the log-polar coordinate system, F2 = F LP2 (ρ, θ); the superscript * represents the complex conjugate, and C(ρ, θ) represents the cross-power spectrum of the target image and the source image; Then calculate the phase correlation function R(ρ, θ) by performing an inverse Fourier transform on the cross-power spectrum: R(ρ,θ) = F -1 (C(ρ,θ))(8) In Equation (8), F -1 represents the inverse Fourier transform; actually, R(ρ, θ) is a peak response function, and the peak position corresponds to the rotation information between the target image and the source image. Its peak position (Δρ, Δθ) corresponds to the rotation angle of the target image relative to the source image; S42. Deduce the scaling scale ratio of the image through the difference in the frequency-domain radius: For two images with frequency-domain radii r1 and r2 respectively, calculate the ratio between them to obtain the scaling ratio Scale: S43. Complete the geometric correction using the obtained rotation angle and scaling scale ratio; After completing the geometric correction of rotation and scaling scale, use the phase correlation method again. First, calculate the cross-power spectrum in the Cartesian coordinate system, and then perform an inverse Fourier transform to obtain the phase correlation function R(ρ, θ). The peak position (Δρ, Δθ) of the phase correlation function in the Cartesian coordinate system indicates the translation offset between the two images; S44. Select a 5×5 neighborhood centered on the peak point to calculate the normalized signal power value: First, calculate the average power P of the neighborhood raw , as follows: In formula (10), N represents the number of neighborhood pixels; (ρ, θ) represents the coordinates of each point in the neighborhood in the log-polar coordinate system; Then, calculate the normalized signal power value P based on the average power of the neighborhood norm , as follows: Output the normalized signal power value within a partial neighborhood centered on the peak point.

5. A method for correcting the offset of an orbital defense area based on image registration and optimal path selection according to claim 1, characterized in that Step S5 evaluates the result of the transformation in S4 based on the normalized signal power value and decides whether to accept the transformation. The detailed steps are as follows: S51. Set an acceptance threshold and judge the magnitude of the calculated normalized signal power and this threshold; S52. If the normalized signal power value is greater than or equal to this threshold, the matching quality between the current images is high. Directly use the rotation, scaling, and translation offset obtained in step S4 to correct the defense area coordinates in the source image, so as to determine the correct defense area position in the target image; S53. If the normalized signal power value is lower than this threshold, it indicates that the matching quality is poor. At this time, proceed to step S6 to find the optimal transformation link to further optimize the registration result.

6. A method for correcting the offset of an orbital security area based on image registration and optimal path selection according to claim 1, characterized in that The detailed steps of the said step S6 are as follows: S61. Divide the time axis between the source image and the target image into one video frame per second; S62. Calculate the normalized signal power value for each pair of adjacent frames i and j according to the Fourier-Mellin transform; S63. Calculate the path score S(p) for each path p from the source image i ori to the target image i k . The path score S(p) is the product of the normalized signal power values between all adjacent frames on the path. A path containing frames i ori , i1, i2,..., i n , i k has the following path score: S(p) = P(i ori , i1) × P(i1, i2) ×... × P(i n , i k )(12) S64. Select the path with the highest score among all paths as the optimal path, that is: In Equation (13), p(i k , i ori ) is the set of all paths from the source image i ori to the target image i k ; S65. Correct the coordinates of the defense area in the source image step by step to the target image according to the optimal path.

7. An electronic device, characterized in that, The said electronic device includes: At least one processor; and A memory that stores instructions, which when executed by the at least one processor, cause the at least one processor to execute an orbit defense area offset correction method based on image registration and optimal path selection as described in any one of claims 1 to 6.

8. A machine-readable storage medium, characterized in that, An executable instruction is stored on the said machine-readable storage medium, and when the instruction is executed, it causes the machine to execute an orbit defense area offset correction method based on image registration and optimal path selection as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Dynamic vision SLAM (Simultaneous Localization and Mapping) method based on frequency domain and semantics

    CN116524026A

  • Image matching method, apparatus, computer equipment and computer-readable storage medium

    CN116563357B