Panoramic radar image stitching method, device and computer equipment for fitting optical images
By mapping cross-modal feature point pairs and affine transformation matrices between optical and radar images, and combining this with a PID control algorithm, the error problem in radar image stitching was solved, achieving high-precision deformation detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-03-17
AI Technical Summary
In surface deformation detection, the observation range of a single-frame radar image is limited and cannot fully cover the surface area. Existing technologies suffer from nonlinear characteristics, mechanical errors, and error propagation effects during radar image stitching, resulting in insufficient accuracy and reliability of deformation detection.
By acquiring multiple sets of synchronously acquired optical and radar images, cross-modal feature point pairs are established, and affine transformation matrices are calculated to achieve the mapping and stitching of radar images to the optical coordinate system. Combined with PID control algorithms and feature point matching optimization, image stitching and turntable angle correction are performed.
It improves the accuracy and reliability of deformation detection, realizes cross-modal registration and seamless stitching, overcomes the error problems in traditional technologies, and ensures natural image stitching effect and high precision.
Smart Images

Figure CN120876219B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image stitching technology, and in particular to a panoramic radar image stitching method, apparatus, and computer equipment for fitting optical images. Background Technology
[0002] When detecting surface deformation, the observation range of a single radar image is limited and may not fully cover the surface area. For example, when detecting slope deformation, a single radar image cannot cover the entire slope area. Therefore, it is necessary to stitch together multiple radar images to expand the monitoring area and ensure comprehensive monitoring of the entire surface, thereby more accurately identifying potential deformation areas.
[0003] Traditional surface deformation detection radar systems face three core technical bottlenecks: First, the inherent nonlinear characteristics of radar imaging, such as the difference in range and azimuth resolution and slant range geometric projection distortion, cause non-uniform deformation of scan data when converting from polar coordinates to Cartesian coordinates. Even with ideal turntable angle encoder accuracy, artifacts such as edge burrs and abrupt intensity changes will still appear in the stitched radar panorama. Second, the system relies on a single data source, the turntable angle encoder, for image registration. Mechanical errors such as turntable axial jitter and gear backlash, along with environmental interference, will further amplify the stitching angle deviation between adjacent scan frames. Third, existing technologies lack a spatiotemporally synchronized external reference system. In large-scale, long-term monitoring, the frame-by-frame propagation effect of errors cannot be suppressed, causing millimeter-level deformation signals to be overwhelmed by centimeter-level stitching errors, ultimately leading to misjudgments of the surface safety status. Summary of the Invention
[0004] Therefore, it is necessary to provide a panoramic radar image stitching method, apparatus, and computer equipment that can improve the accuracy and reliability of deformation detection by fitting optical images to address the aforementioned technical problems.
[0005] A method for stitching panoramic radar images by fitting optical images, the method comprising:
[0006] Acquire multiple sets of simultaneously acquired optical and radar images;
[0007] The location information of the target object in the radar image is obtained, and the centroid coordinates of the target object in the optical image are obtained; the centroid coordinates and the location information constitute a cross-modal feature point pair;
[0008] Based on multiple sets of cross-modal feature point pairs, an affine transformation model from the radar coordinate system to the optical coordinate system is established. The unknown parameters in the affine transformation model are calculated, and the solution results of the unknown parameters are substituted into the affine transformation model to obtain the affine transformation matrix.
[0009] Based on the affine transformation matrix, each pixel in the radar image is mapped to the optical coordinate system to obtain the transformed radar image;
[0010] By stitching together multiple frames of the transformed radar images, a panoramic radar image is obtained.
[0011] The above-described panoramic radar image stitching method for fitting optical images acquires multiple sets of synchronously acquired optical and radar images. It obtains the positional information of the target object in the radar image and the centroid coordinates of the target object in the optical image. Based on multiple sets of cross-modal feature point pairs, it establishes an affine transformation model from the radar coordinate system to the optical coordinate system. It calculates the unknown parameters in the affine transformation model and substitutes the solution results into the affine transformation model to obtain the affine transformation matrix. Based on the affine transformation matrix, it maps each pixel in the radar image to the optical coordinate system to obtain the transformed radar image. By stitching together multiple frames of transformed radar images, a panoramic radar image is obtained. This method utilizes synchronously acquired optical data as the geometric "truth" and its spatiotemporal correlation with the radar image, overcoming the error limitations of a single sensor and achieving cross-modal registration and seamless stitching, thereby improving the accuracy and reliability of deformation detection.
[0012] In one embodiment, the camera that acquires the optical image and the radar device that acquires the radar image are coaxially mounted on the turntable;
[0013] The method further includes:
[0014] An optical panoramic image is obtained by stitching together optical images acquired at different acquisition times, and the optical reference point in the optical panoramic image and the radar mapping point in the radar panoramic image are determined; the optical reference point and the radar mapping point are pixels of the same object;
[0015] Calculate the vertical coordinate difference between the optical reference point and the radar mapping point, divide the vertical coordinate difference by the focal length of the camera to obtain a quotient, and convert the quotient into an angle difference using the arctangent function; the angle difference is the deviation between the actual rotation angle and the theoretical rotation angle of the turntable.
[0016] Calculate the horizontal coordinate difference between the optical reference point and the radar mapping point;
[0017] Based on the horizontal coordinate difference, the vertical coordinate difference, and the angle difference, the proportional term, integral term, and derivative term are calculated using a PID control algorithm.
[0018] The weighted sum of the proportional term, the integral term, and the differential term is superimposed on the encoded value corresponding to the actual rotation angle to obtain the corrected target encoded value, and the rotation of the turntable is controlled based on the target encoded value.
[0019] In this embodiment, an optical panoramic image is obtained by stitching together optical images acquired at different acquisition times. The optical reference point in the optical panoramic image and the radar mapping point in the radar panoramic image are determined. The vertical coordinate difference between the optical reference point and the radar mapping point is calculated, and the vertical coordinate difference is divided by the focal length of the camera to obtain the quotient. The quotient is converted into an angle difference using the arctangent function. The horizontal coordinate difference between the optical reference point and the radar mapping point is calculated. Based on the horizontal coordinate difference, vertical coordinate difference, and angle difference, the proportional term, integral term, and derivative term are calculated using a PID control algorithm. The weighted sum of the proportional term, integral term, and derivative term is superimposed on the encoded value corresponding to the actual rotation angle to obtain the corrected target encoded value. The rotation of the turntable is controlled based on the target encoded value. This overcomes the problem that the actual rotation angle deviates from the theoretical rotation angle due to factors such as temperature drift, installation error, and encoder nonlinearity, and realizes closed-loop feedback correction.
[0020] In one embodiment, acquiring the stitched optical panoramic image obtained from optical images acquired at different acquisition times includes:
[0021] Optical images acquired at different acquisition times are preprocessed to obtain multiple preprocessed images;
[0022] Feature points in each frame of the preprocessed image are detected, and a descriptor for each feature point is calculated based on the pixel values of the first neighboring region of each feature point.
[0023] Based on the descriptor, the Euclidean distance between each first feature point in the first image and each second feature point in the second image is calculated; the first image and the second image are two adjacent preprocessed images, the first feature point is the feature point in the first image, and the second feature point is the feature point in the second image;
[0024] Based on the Euclidean distance, the nearest neighbor and the second nearest neighbor of each first feature point are determined from each second feature point;
[0025] Taking any of the first feature points as the target point, when the ratio of the Euclidean distance between the target point and its nearest neighbor to the Euclidean distance between the target point and its second nearest neighbor is greater than a first threshold, the nearest neighbor and the target point are combined into an initial matching pair. This process is repeated for each of the first feature points to obtain multiple initial matching pairs.
[0026] The initial matching pairs are optimized to obtain optimized matching pairs;
[0027] The camera's pose parameters are calibrated based on the optimized matching pair, and perspective transformation is performed on each frame of the preprocessed image based on the calibrated pose parameters to obtain a transformed optical image. The transformed optical images are then stitched together to obtain an optical panoramic image.
[0028] By preprocessing optical images acquired at different acquisition times, the stability and accuracy of subsequent feature detection can be improved. Feature points in each preprocessed image are detected, and descriptors for each feature point are calculated based on the pixel values of its first neighbor region. Based on these descriptors, the Euclidean distance between each first feature point in the first image and each second feature point in the second image is calculated. Based on the Euclidean distance, the nearest and second nearest neighbors of each first feature point are determined from among the second feature points. Taking any first feature point as the target point, when the ratio of the Euclidean distance between the target point and its nearest neighbor to the Euclidean distance between the target point and its second nearest neighbor is greater than a first threshold, the nearest neighbor and the target point are combined into an initial matching pair. This process is repeated for each first feature point to obtain multiple initial matching pairs, effectively suppressing false matches and significantly reducing the false match rate. Optimizing the initial matching pairs yields optimized matching pairs, which eliminates any remaining erroneous matches and further improves matching reliability. By calibrating the camera's pose parameters based on optimized matching pairs, and performing perspective transformation on each frame of preprocessed images based on the calibrated pose parameters, transformed optical images are obtained. These transformed optical images are then stitched together to obtain an optical panoramic image. This ensures that there is no obvious distortion or misalignment of buildings, roads, and other features in the optical panoramic image, resulting in a natural visual effect.
[0029] In one embodiment, calculating the descriptor of each feature point based on the pixel values of the first neighboring region of each feature point includes:
[0030] Centered on the feature point, the first neighboring region of the feature point is divided into multiple sub-blocks;
[0031] In each sub-block, based on the pixel value of each pixel, the horizontal gradient and vertical gradient of each pixel are calculated, and based on the horizontal gradient and vertical gradient of each pixel, the gradient magnitude and gradient direction of each pixel are calculated.
[0032] The gradient direction corresponding to each sub-block is classified into the corresponding preset interval;
[0033] Based on the gradient direction and corresponding gradient magnitude in each preset interval, a descriptor for each feature point is generated.
[0034] In this embodiment, the first neighboring region of a feature point is divided into multiple sub-blocks centered on the feature point. In each sub-block, the horizontal and vertical gradients of each pixel are calculated based on the pixel value. Based on the horizontal and vertical gradients of each pixel, the gradient magnitude and gradient direction of each pixel are calculated. The gradient direction corresponding to each sub-block is classified into a corresponding preset interval. Based on the gradient direction and corresponding gradient magnitude in each preset interval, a descriptor for each feature point is generated. This can transform the disordered pixel information of each feature point into a structured description, ensuring that the descriptors of different feature points are unique, thereby avoiding obtaining incorrect initial matching pairs.
[0035] In one embodiment, optimizing the initial matching pair to obtain an optimized matching pair includes the following steps:
[0036] S11. Randomly select multiple initial matching pairs from the initial matching pairs as first matching pairs; the initial matching pairs that are not selected from the initial matching pairs are second matching pairs;
[0037] S12. Calculate the homography matrix based on the first matching pair;
[0038] S13. Based on the homography matrix, project the first feature point in the second matching pair onto the coordinate system of the second image to obtain the coordinates of the first projection point;
[0039] S14. Determine the coordinate difference between the first projection point and the first feature point associated with the first projection point as the first coordinate difference;
[0040] S15. Count the number of matching pairs in the second matching pair whose first coordinate difference is less than the second threshold, and repeat steps S1, S12, S13, S14 and S15 to obtain the number of matching pairs corresponding to each iteration. Determine the homography matrix corresponding to the maximum value among the number of matching pairs as the target homography matrix.
[0041] S16. Based on the target homography matrix, project the first feature point in the initial matching pair onto the coordinate system of the second image to obtain the coordinates of the second projection point;
[0042] S17. Determine the coordinate difference between the second projection point and the first feature point associated with the second projection point as the second coordinate difference, and remove the matching pairs in the initial matching pairs whose second coordinate difference is greater than the third threshold to obtain the optimized matching pairs.
[0043] In this embodiment, step S11 involves randomly selecting multiple initial matching pairs from the initial matching pairs as first matching pairs; the initial matching pairs that are not selected are designated as second matching pairs; step S12 involves calculating the homography matrix based on the first matching pairs; step S13 involves projecting the first feature point in the second matching pair onto the coordinate system of the second image based on the homography matrix to obtain the coordinates of the first projection point; step S4 involves determining the coordinate difference between the first projection point and the first feature point associated with the first projection point as the first coordinate difference; step S15 involves counting the number of matching pairs in the second matching pairs whose first coordinate difference is less than a second threshold, and repeating steps S11, S12, S13, S14, and S15 to obtain the number of matching pairs corresponding to each iteration, and then matching... S16. Based on the target homography matrix, the first feature point in the initial matching pair is projected onto the coordinate system of the second image to obtain the coordinates of the second projection point; S17. The coordinate difference between the second projection point and the first feature point associated with the second projection point is determined as the second coordinate difference, and matching pairs in the initial matching pair whose second coordinate difference is greater than the third threshold are eliminated to obtain optimized matching pairs. In this way, when the initial matching pair contains a large number of erroneous matches, these "outside points" can be automatically identified and eliminated, and "inside points" that conform to the global transformation model can be retained, i.e., optimized matching pairs. This provides reliable input for subsequent perspective transformation, image stitching, pose estimation, etc., and reduces visual defects such as stitching seams and distortion.
[0044] In one embodiment, calibrating the camera's pose parameters based on the optimized matching pair includes the following steps:
[0045] S21. Preset the three-dimensional coordinates and initial pose parameters of each pixel in the preprocessed image of each frame in the three-dimensional scene;
[0046] S22. Based on the initial pose parameters, determine the predicted position of each pixel's three-dimensional coordinates projected onto the preprocessed image;
[0047] S23. The difference between the predicted position of each pixel and the actual position in the preprocessed image is determined as the projection error;
[0048] S24. Based on the projection error, construct a nonlinear least squares optimization objective function;
[0049] S25. Calculate the minimum value of the nonlinear least squares optimization objective function using an iterative optimization algorithm, and update the initial pose parameters and three-dimensional coordinates in step S21 in reverse. Repeat steps S22 and S23 based on the updated pose parameters and three-dimensional coordinates until the projection error converges, and output the updated pose parameters.
[0050] In this embodiment, step S21 involves presetting the 3D coordinates and initial pose parameters of each pixel in the 3D scene of each preprocessed image frame; step S22 involves determining the predicted position of each pixel's 3D coordinates projected onto the preprocessed image based on the initial pose parameters; step S23 involves determining the projection error as the difference between the predicted position and the actual position of each pixel in the preprocessed image; step S24 involves constructing a nonlinear least squares optimization objective function based on the projection error; and step S25 involves calculating the minimum value of the nonlinear least squares optimization objective function using an iterative optimization algorithm and updating the initial pose parameters and 3D coordinates in step S21 in reverse. Steps S22 and S23 are repeated based on the updated pose parameters and 3D coordinates until the projection error converges, and the updated pose parameters are output. This ensures the consistency and optimality of pose parameter estimation among multiple preprocessed images.
[0051] In one embodiment, the radar image is acquired by means of:
[0052] The returned linear frequency modulated echo signal is acquired, and the linear frequency modulated echo signal is subjected to range pulse compression processing to obtain a pulse compressed signal; the range pulse compression processing includes deskewing, fast Fourier transform, and range migration effect correction;
[0053] The pulse compression signal is subjected to azimuth synthetic aperture radar imaging processing to obtain azimuth-compressed complex data;
[0054] The complex data is subjected to multi-view processing, the processed data is converted into an intensity image, and geocoding is performed to obtain a radar image.
[0055] In this embodiment, by acquiring the returned linear frequency modulated echo signal, range pulse compression processing is performed on the linear frequency modulated echo signal to obtain a pulse compressed signal. The pulse compressed signal is then subjected to azimuth synthetic aperture radar imaging processing to obtain azimuth compressed complex data. The complex data is then subjected to multi-view processing, and the multi-view processed data is converted into an intensity image and geocoded to obtain a radar image. This enables high-resolution radar imaging and the acquisition of high spatial resolution radar images.
[0056] In one embodiment, the method further includes:
[0057] Determine the second neighbor region of each pixel in the radar image, and calculate the average intensity and standard deviation of the second neighbor region;
[0058] Based on the average intensity and the standard deviation, the adaptive threshold corresponding to each pixel is obtained;
[0059] When the pixel value of a pixel is less than the corresponding adaptive threshold, the pixel is determined to be a candidate feature point in the radar image;
[0060] Clustering algorithms are used to filter the candidate feature points to obtain filtered feature points; the filtered feature points are used to determine the location information of the target object in the radar image.
[0061] In this embodiment, by determining the second neighboring region of each pixel in the radar image, the average intensity and standard deviation of the second neighboring region are calculated. Based on the average intensity and standard deviation, an adaptive threshold corresponding to each pixel is obtained. When the pixel value of a pixel is less than the corresponding adaptive threshold, the pixel is determined as a candidate feature point in the radar image. A clustering algorithm is used to filter the candidate feature points to obtain the filtered feature points. This can effectively overcome the problems of large speckle noise and uneven contrast in radar images, and achieve feature point filtering with high robustness and low false alarm rate.
[0062] A panoramic radar image stitching device for fitting optical images, the device comprising:
[0063] The image acquisition module is used to acquire multiple sets of synchronously acquired optical and radar images;
[0064] The information acquisition module is used to acquire the position information of the target object in the radar image and acquire the centroid coordinates of the target object in the optical image; the centroid coordinates and the position information constitute a cross-modal feature point pair;
[0065] The parameter calculation module is used to establish an affine transformation model from the radar coordinate system to the optical coordinate system based on multiple sets of cross-modal feature point pairs, calculate the unknown parameters in the affine transformation model, and substitute the solution results of the unknown parameters into the affine transformation model to obtain the affine transformation matrix.
[0066] The image conversion module is used to map each pixel in the radar image to the optical coordinate system based on the affine transformation matrix to obtain the converted radar image.
[0067] The image stitching module is used to stitch together multiple frames of the converted radar images to obtain a panoramic radar image.
[0068] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.
[0069] The aforementioned panoramic radar image stitching device and computer equipment for fitting optical images acquire multiple sets of synchronously acquired optical and radar images, obtain the position information of the target object in the radar image, and obtain the centroid coordinates of the target object in the optical image. Based on multiple sets of cross-modal feature point pairs, an affine transformation model from the radar coordinate system to the optical coordinate system is established. The unknown parameters in the affine transformation model are calculated, and the solution results of the unknown parameters are substituted into the affine transformation model to obtain the affine transformation matrix. Based on the affine transformation matrix, each pixel in the radar image is mapped to the optical coordinate system to obtain the transformed radar image. By stitching multiple frames of transformed radar images, a panoramic radar image is obtained. In this way, the synchronously acquired optical image can be used as the geometric "truth value" and its spatiotemporal correlation with the radar image, breaking through the error limitations of a single sensor and realizing cross-modal registration and seamless stitching, thereby improving the accuracy and reliability of deformation detection. Attached Figure Description
[0070] Figure 1 This is a flowchart illustrating a panoramic radar image stitching method for fitting optical images in one embodiment.
[0071] Figure 2 This is a schematic diagram of the overall process of a panoramic radar image stitching method for fitting optical images in one embodiment.
[0072] Figure 3 This is a structural block diagram of a panoramic radar image stitching device for fitting optical images in one embodiment;
[0073] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0075] In one embodiment, such as Figure 1 As shown, a panoramic radar image stitching method based on fitting optical images is provided, including the following steps:
[0076] S102, acquire multiple sets of synchronously acquired optical and radar images;
[0077] In this system, a camera for acquiring optical images and a radar device for acquiring radar images are coaxially mounted on a turntable. The radar device can be a rotating millimeter-wave radar with a frequency of 77 GHz, an azimuth resolution of 0.1°, an elevation scanning angle of ±30°, and a horizontal scanning angle of ±60°. The camera can be a global shutter industrial camera with a resolution of 20 megapixels and a frame rate of 30 fps. A beam splitter ensures that the camera's optical path is coaxial with the radar beam of the radar device, eliminating parallax. The beam splitter transmits millimeter waves (transmittance ≥95%) and reflects visible light (reflectivity ≥90%). After the turntable rotates to a designated position and stabilizes, the camera acquires optical images. The turntable is driven by a stepper motor with a step angle accuracy less than a first preset value. Each rotation of a preset angle triggers a radar beam transmission and reception. Specifically, for every 0.1° rotation of the turntable, the angle encoder in the turntable outputs a pulse signal to the FPGA (Field-Programmable Gate Array). Within 0.1μs of receiving the pulse, the FPGA sends a synchronization trigger signal. The rising edge of EN_RADAR (radar enable signal) triggers the radar transmission beam, lasting for 2ms. EN_CAM (camera enable signal) starts within 0.05ms after the rising edge of EN_RADAR, or starts synchronously, ensuring the camera exposure window completely covers the radar beam transmission period. The timestamp is based on the FPGA global clock count, incrementing every 0.1ms. The timestamp deviation between the radar image and the optical image is ≤0.1ms, ensuring spatiotemporal alignment between the radar and optical images. A turntable drives the radar image acquisition device to rotate and scan the surface to be deformed, emitting electromagnetic waves and receiving echoes to generate high-resolution radar images. Since the angle of a single radar scan is limited, the rotation angle recorded by the turntable angle encoder is used to stitch multiple radar images into a panoramic image. The angle encoder provides the absolute or relative measurement value of the turntable rotation angle, serving as the geometric reference for image stitching. The FPGA's synchronization logic uses a 100MHz crystal oscillator clock source, divided to 10MHz by a PLL (Phase Locked Loop) as the global synchronization clock. It receives pulse signals from the turntable angle encoder and generates three synchronization trigger signals. EN_RADAR triggers radar beam transmission and reception with a pulse width of 2ms. EN_CAM triggers camera exposure with a pulse width of 1ms, ensuring a timing overlap of ≥95% with the radar beam transmission.
[0078] Optical and radar images within the same group are acquired at the same time, while those between different groups are acquired at different times. Furthermore, all image data packets are appended with a unified timestamp in the format "hour:minute:second.millisecond", generated by the FPGA global clock. The alignment accuracy of synchronously acquired optical and radar images is 0.1ms. Specifically, each rotation of the turntable by a preset angle triggers data acquisition. The range-azimuth two-dimensional matrix radar echo data corresponding to the i-th rotation is associated with the optical image and the timestamp in the hour:minute:second.millisecond format to ensure that they represent the same spatial location. When the timestamp deviation is less than 0.1ms, it is determined that the radar and optical images are synchronously acquired and used as input data for cross-modal registration.
[0079] S104: Obtain the position information of the target object in the radar image and obtain the centroid coordinates of the target object in the optical image; the centroid coordinates and position information constitute a cross-modal feature point pair;
[0080] In radar images, the target object is a highly reflective feature point. Examples include rock edges, metal structures, and stable corner reflectors. Location information can also be understood as the centroid coordinates of the target object in the radar image. Because highly reflective feature points have strong energy, they are reflected as prominent colors and lines in optical images, and the same is true in radar images. Therefore, when determining the feature contours of a target object, the same target object shares the same feature descriptors, thus allowing the radar image and the optical image to correspond to the same target object.
[0081] The calculation process for the centroid coordinates of the target object involves extracting the feature contour of the target object, determining the circumscribed rectangle of the feature contour, and using the coordinates of the center point of the circumscribed rectangle as the centroid coordinates of the target object. Furthermore, before determining the feature contour of the target object in the optical image, the process includes filtering and denoising the optical image to determine the feature contour of the target object in the filtered and denoised optical image.
[0082] Methods for determining the feature contours of target objects in optical images include, but are not limited to, contour extraction based on edge detection, contour extraction based on image segmentation, and contour extraction based on deep learning. For example, the Canny edge detection algorithm can be used to extract the feature contours of objects with strong and distinct features in optical images.
[0083] S106. Based on multiple sets of cross-modal feature point pairs, an affine transformation model from the radar coordinate system to the optical coordinate system is established. The unknown parameters in the affine transformation model are calculated, and the solution results of the unknown parameters are substituted into the affine transformation model to obtain the affine transformation matrix.
[0084] The affine transformation model is: Where a, b, c, d, e and f are the unknown parameters to be solved in the affine transformation model, a, b, d and e are the rotation, scaling and shearing parameters, and c and f are the translation parameters; Let represent the location information of the target object in the i-th radar image acquired. This represents the location information of the target object in the i-th acquired optical image.
[0085] The unknown parameters in the affine transformation model can be solved using the least squares method. Specifically, the system of equations for the affine transformation model, based on multiple sets of cross-modal feature point pairs, is as follows: N represents the number of cross-modal feature point pairs. Let Ap = B, where p = [a, b, c, d, e, f]. T Perform singular value decomposition on matrix A Then the parameter solution is .
[0086] S108, based on the affine transformation matrix, maps each pixel in the radar image to the optical coordinate system to obtain the transformed radar image;
[0087] The coordinates of each pixel in the transformed radar image are obtained by multiplying the coordinates of each pixel in the radar image by the affine transformation matrix. The optical coordinate system is the panoramic coordinate system.
[0088] Furthermore, bilinear interpolation is performed on the transformed radar image to eliminate the jagged effect caused by coordinate discretization and achieve sub-pixel level alignment. Verification points are randomly selected in the overlapping area, such as rock fissures and structural edges, and the alignment error between the radar image and the optical image is calculated. If the error exceeds 1 pixel, the cross-modal feature point pairs are readjusted or the affine transformation matrix is optimized until the accuracy requirements are met.
[0089] S110 stitches together multiple frames of radar images to obtain a panoramic radar image.
[0090] While there may be overlapping areas in the various converted radar images, most areas are not identical.
[0091] Image stitching can also be understood as image fusion or image synthesis. The purpose of image stitching is to process overlapping areas and fill in gaps. Specifically, the pixel values of the overlapping areas are averaged, and the weights can be based on distance, signal-to-noise ratio, or Gaussian weights, etc., to ensure a smooth transition. Interpolation algorithms are then used to fill in the areas that were not captured.
[0092] The above-described panoramic radar image stitching method for fitting optical images acquires multiple sets of synchronously acquired optical and radar images. It obtains the positional information of the target object in the radar image and the centroid coordinates of the target object in the optical image. Based on multiple sets of cross-modal feature point pairs, it establishes an affine transformation model from the radar coordinate system to the optical coordinate system. It calculates the unknown parameters in the affine transformation model and substitutes the solution results into the affine transformation model to obtain the affine transformation matrix. Based on the affine transformation matrix, it maps each pixel in the radar image to the optical coordinate system to obtain the transformed radar image. By stitching together multiple frames of transformed radar images, a panoramic radar image is obtained. This method utilizes synchronously acquired optical data as the geometric "truth" and its spatiotemporal correlation with the radar image, overcoming the error limitations of a single sensor and achieving cross-modal registration and seamless stitching, thereby improving the accuracy and reliability of deformation detection.
[0093] In one embodiment, a camera for acquiring optical images and a radar device for acquiring radar images are coaxially mounted on a turntable, and the method further includes:
[0094] Obtain an optical panoramic image by stitching together optical images acquired at different acquisition times, and determine the optical reference point in the optical panoramic image and the radar mapping point in the radar panoramic image; the optical reference point and the radar mapping point are pixels of the same object;
[0095] Calculate the vertical coordinate difference between the optical reference point and the radar mapping point, divide the vertical coordinate difference by the camera's focal length to obtain the quotient, and convert the quotient into an angle difference using the arctangent function; the angle difference is the deviation between the actual rotation angle and the theoretical rotation angle of the turntable.
[0096] Calculate the difference in horizontal coordinates between the optical reference point and the radar mapping point;
[0097] Based on the differences in horizontal coordinates, vertical coordinates, and angles, the proportional, integral, and derivative terms are calculated using a PID control algorithm.
[0098] The weighted sum of the proportional, integral, and differential terms is superimposed onto the coded value corresponding to the actual rotation angle to obtain the corrected target coded value. The rotation of the turntable is then controlled based on the target coded value.
[0099] In this process, an optical image can be acquired at each acquisition moment. By mapping the pixels of each optical image to the same coordinate system, the mapped images of each optical image can be obtained. By stitching together the mapped images of each frame, an optical panoramic image can be obtained.
[0100] Optical reference points in optical panoramic images and radar mapping points in radar panoramic images can be detected using edge detection algorithms. Specifically, when detecting optical and radar panoramic images using edge detection algorithms, the same object in both types of images has the same feature descriptor, thus allowing the mapping of the same object between the two types of images to be determined. Further, the optical reference point can be any pixel of the object in the optical panoramic image, or the centroid of the object's outline in the optical panoramic image; similarly, the radar mapping point can be any pixel of the object in the radar panoramic image, or the centroid of the object's outline in the radar panoramic image.
[0101] The vertical coordinate difference refers to the difference in coordinates between the radar mapping point and the optical reference point in the vertical direction. The horizontal coordinate difference refers to the difference in coordinates between the radar mapping point and the optical reference point in the horizontal direction. Both the vertical and horizontal coordinate differences reflect the alignment error between the radar panoramic image and the optical panoramic image in the translational direction.
[0102] The proportional term responds quickly to the current angle difference, the integral term accumulates historical deviations to eliminate steady-state error, and the derivative term suppresses oscillations by predicting the trend of deviation changes. The weights of the proportional, integral, and derivative terms can be preset.
[0103] Furthermore, a preset duration is obtained, and the encoded value is corrected every preset duration to obtain the corrected target encoded value. For example, the encoded value is corrected every 10 milliseconds.
[0104] The corrected target encoding value is sent to the turntable's control module to drive the turntable's stepper motor to adjust the reference angle for subsequent scans. When the turntable rotates next time, it will rotate according to the new angle corresponding to the target encoding value, while continuously monitoring the deviation and iteratively correcting the encoding value, forming a closed-loop control chain of "detection-calculation-correction", ultimately suppressing the radar's stitching angle error within the second preset value.
[0105] In this embodiment, an optical panoramic image is obtained by stitching together optical images acquired at different acquisition times. The optical reference point in the optical panoramic image and the radar mapping point in the radar panoramic image are determined. The vertical coordinate difference between the optical reference point and the radar mapping point is calculated, and the vertical coordinate difference is divided by the focal length of the camera to obtain the quotient. The quotient is converted into an angle difference using the arctangent function. The horizontal coordinate difference between the optical reference point and the radar mapping point is calculated. Based on the horizontal coordinate difference, vertical coordinate difference, and angle difference, the proportional term, integral term, and derivative term are calculated using a PID control algorithm. The weighted sum of the proportional term, integral term, and derivative term is superimposed on the encoded value corresponding to the actual rotation angle to obtain the corrected target encoded value. The rotation of the turntable is controlled based on the target encoded value. This overcomes the problem that the actual rotation angle deviates from the theoretical rotation angle due to factors such as temperature drift, installation error, and encoder nonlinearity, and realizes closed-loop dynamic correction.
[0106] In one embodiment, acquiring an optical panoramic image stitched together from optical images acquired at different acquisition times includes:
[0107] Optical images acquired at different acquisition times are preprocessed to obtain multiple preprocessed images;
[0108] Feature points in each frame of the preprocessed image are detected, and descriptors for each feature point are calculated based on the pixel values of the first neighboring region of each feature point.
[0109] Based on the descriptor, calculate the Euclidean distance between each first feature point in the first image and each second feature point in the second image;
[0110] Based on Euclidean distance, the nearest neighbor and second nearest neighbor of each first feature point are determined from each second feature point;
[0111] Taking any first feature point as the target point, when the ratio of the Euclidean distance between the target point and its nearest neighbor to the Euclidean distance between the target point and its second nearest neighbor is greater than the first threshold, the nearest neighbor and the target point are combined into an initial matching pair. This process is repeated for each first feature point to obtain multiple initial matching pairs.
[0112] The initial matching pairs are optimized to obtain optimized matching pairs;
[0113] The pose parameters of the camera are calibrated by optimizing the matching pair, and perspective transformation is performed on each frame of preprocessed image based on the calibrated pose parameters to obtain the transformed optical image. The transformed optical images are then stitched together to obtain the optical panoramic image.
[0114] Image preprocessing includes, but is not limited to, denoising optical images, eliminating noise in optical images using Gaussian filtering, and enhancing the contrast of optical images.
[0115] Feature points are also pixels in the preprocessed image. Feature points can be obtained through feature point detection algorithms. Feature point detection algorithms include, but are not limited to, SIFT (Scale-Invariant Feature Transform) and ORB (Oriented Features from Accelerated Segment Test and Rotated Binary Robust Independent Elementary Features) algorithms.
[0116] The size of the first neighboring region can be determined by a first preset size. For example, if the first preset size is 16×16, then the 16×16 pixel area around the feature point is determined as the first neighboring region.
[0117] A descriptor is information used to describe the surrounding pixels of a feature point. The descriptor is calculated as follows: The first neighboring region of the feature point is divided into multiple sub-blocks centered on the feature point; within each sub-block, the horizontal and vertical gradients of each pixel are calculated based on its pixel value; the gradient magnitude and direction of each pixel are then calculated based on these gradients; the gradient directions corresponding to each sub-block are categorized into corresponding preset intervals; and a descriptor for each feature point is generated based on the gradient direction and corresponding gradient magnitude within each preset interval.
[0118] The first and second images are two adjacent preprocessed frames. The first feature point is a feature point in the first image, and the second feature point is a feature point in the second image. In a specific application, the descriptor of the first feature point A in the first image is a 128-dimensional descriptor D_A=[a1,a2,...,a...]. 128 The descriptor for the second feature point B in the second image is a 128-dimensional descriptor D_B = [b1,b2,...,b...]. 128 The Euclidean distance d(D_A,D_B) is calculated as follows: d(D_A,D_B)=[(a1-b1)²+(a2-b2)²+...+(a...b1)²+...b2 ... 128 -b 128 )²] 1 / 2 .
[0119] The nearest neighbor is the point in the second image that has the closest Euclidean distance to the first feature point in the first image among all the second feature points in the second image. The second nearest neighbor is the point in the second image that has the second closest Euclidean distance to the first feature point in the first image among all the second feature points in the second image. Each first feature point in the first image has one nearest neighbor and one second nearest neighbor.
[0120] The Euclidean distance between the target point and its nearest neighbor refers to the distance between the target point and its nearest neighbor. The Euclidean distance between the second nearest neighbor refers to the distance between the target point and its second nearest neighbor. A first threshold can be preset. For example, if the first threshold is 0.8, then when the ratio of the Euclidean distance between the target point and its nearest neighbor to the Euclidean distance between the target point and its second nearest neighbor is greater than 0.8, the target point's nearest neighbor and the target point are combined into an initial matching pair.
[0121] Traversing each first feature point can also be understood as taking each first feature point as the target point and judging whether the ratio corresponding to each target point is greater than the first threshold.
[0122] Optimization of the initial matching pair can be achieved through various methods, including but not limited to iterative optimization, graph optimization, and optimization using local geometric constraints.
[0123] Pose parameters include the rotation matrix and translation vector of each preprocessed image frame. The pose parameters of the calibrated camera based on optimized matching pairs can be achieved through bundle adjustment.
[0124] The perspective transformation of each preprocessed image frame is performed based on the calibrated pose parameters, which can also be understood as mapping the pixels in each preprocessed image frame to a unified panoramic coordinate system.
[0125] Furthermore, when stitching together the transformed optical images, a multi-band fusion technique is employed to eliminate seams in overlapping areas. The image is decomposed into Laplacian pyramid layers with different spatial frequencies. The high-frequency layer preserves edge details, while the low-frequency layer represents the overall structure. In overlapping areas, corresponding frequency bands are mixed with gradually varying weights. For example, a higher weight is given to the high-frequency layer in image edge areas to preserve details of rock fissures, while the low-frequency layer is dominant in smooth areas to maintain a natural transition. The result is a seamlessly stitched optical panoramic image with an alignment error of less than 0.1 pixels at the stitching point of any two frames, meeting the requirements for millimeter-level deformation monitoring.
[0126] In this embodiment, image preprocessing is performed on optical images acquired at different acquisition times to improve the stability and accuracy of subsequent feature detection. Feature points in each preprocessed image frame are detected, and descriptors for each feature point are calculated based on the pixel values of its first neighboring region. Based on these descriptors, the Euclidean distance between each first feature point in the first image and each second feature point in the second image is calculated. Based on the Euclidean distance, the nearest and second nearest neighbors of each first feature point are determined from among the second feature points. Taking any first feature point as the target point, when the ratio of the Euclidean distance between the target point and its nearest neighbor to the Euclidean distance between the target point and its second nearest neighbor is greater than a first threshold, the nearest neighbor and the target point are combined into an initial matching pair. This process is repeated for each first feature point to obtain multiple initial matching pairs, effectively suppressing false matches and significantly reducing the false match rate. Optimizing the initial matching pairs yields optimized matching pairs, which eliminates any remaining erroneous matches and further improves matching reliability. By calibrating the camera's pose parameters based on optimized matching pairs, and performing perspective transformation on each frame of preprocessed images based on the calibrated pose parameters, transformed optical images are obtained. These transformed optical images are then stitched together to obtain an optical panoramic image. This ensures that there is no obvious distortion or misalignment of buildings, roads, and other features in the optical panoramic image, resulting in a natural visual effect.
[0127] In one embodiment, a descriptor for each feature point is calculated based on the pixel values of its first neighboring region, including:
[0128] Centered on the feature point, the first neighboring region of the feature point is divided into multiple sub-blocks;
[0129] In each sub-block, based on the pixel value of each pixel, the horizontal gradient and vertical gradient of each pixel are calculated, and based on the horizontal gradient and vertical gradient of each pixel, the gradient magnitude and gradient direction of each pixel are calculated.
[0130] The gradient direction corresponding to each sub-block is classified into the corresponding preset interval;
[0131] Based on the gradient direction and corresponding gradient magnitude in each preset interval, a descriptor for each feature point is generated.
[0132] The formula for calculating the horizontal gradient is as follows: The formula for calculating the vertical gradient is: The formula for calculating the gradient magnitude is: The formula for calculating the gradient direction is: Let x be the x-coordinate of the pixel, y be the y-coordinate of the pixel, G(x) be the horizontal gradient of the pixel with x-coordinate, G(y) be the vertical gradient of the pixel with y-coordinate, and M(x,y) be the gradient magnitude of the pixel at (x,y). Let (x, y) be the gradient direction of the pixel. Let (x+1, y) be the pixel value of the pixel point. Let (x-1, y) be the pixel value of the pixel point. Let (x, y+1) be the pixel value of the pixel point. The pixel value of the pixel point (x, y-1).
[0133] Each sub-block contains multiple pixels, and each pixel corresponds to a gradient direction, so a sub-block corresponds to multiple gradient directions.
[0134] Preset intervals can be based on preset angle intervals. For example, if the preset angle interval is 45°, 360° can be divided into eight preset intervals: [0, 45°), [45°, 90°), [90°, 135°), [135°, 180°), [180°, 225°), [225°, 270°), [270°, 315°), [315°, 360°). When classifying gradient directions, they are assigned to the corresponding preset intervals based on their magnitude.
[0135] The gradient direction corresponding to each sub-block is categorized into a corresponding preset interval. This can also be understood as categorizing the gradient directions corresponding to each sub-block into different "bins". The gradient direction θ determines which preset interval a pixel belongs to. The gradient magnitude of the pixel is used as a "weight" and added to the corresponding preset interval. That is, the corresponding gradient magnitude is the gradient magnitude of the pixel to which each gradient direction belongs in the preset interval. For example, the gradient direction of the first pixel is 30°, which falls in the first preset interval [0°, 45°), so the gradient magnitude of the first pixel, 20, is added to the first preset interval. The gradient direction of the second pixel is 45°, which falls in the second preset interval [45°, 90°), so the gradient magnitude of the second pixel, 15, is added to the second preset interval.
[0136] In a specific application, with each feature point as the center, the surrounding 16×16 pixel area is divided into 4×4 sub-blocks. Within each sub-block, the horizontal and vertical gradients of each pixel are calculated first, and then the gradient magnitude and gradient direction of the pixel are calculated. Then, the gradient directions in each sub-block are classified into the corresponding intervals of the eight preset intervals, and an 8-dimensional orientation histogram is constructed based on the gradient magnitude. Finally, each of the 16 sub-blocks forms an 8-dimensional histogram, generating a 128-dimensional descriptor.
[0137] In this embodiment, the first neighboring region of a feature point is divided into multiple sub-blocks centered on the feature point. In each sub-block, the horizontal and vertical gradients of each pixel are calculated based on the pixel value. Based on the horizontal and vertical gradients of each pixel, the gradient magnitude and gradient direction of each pixel are calculated. The gradient direction corresponding to each sub-block is classified into a corresponding preset interval. Based on the gradient direction and corresponding gradient magnitude in each preset interval, a descriptor for each feature point is generated. This can transform the disordered pixel information of each feature point into a structured description, ensuring that the descriptors of different feature points are unique, thereby avoiding obtaining incorrect initial matching pairs.
[0138] In one embodiment, optimizing the initial matching pair to obtain an optimized matching pair includes the following steps:
[0139] S11. Randomly select multiple initial matching pairs from the initial matching pairs as the first matching pair; the initial matching pairs that are not selected from the initial matching pairs are the second matching pairs;
[0140] S12. Calculate the homography matrix based on the first matching pair;
[0141] S13. Based on the homography matrix, project the first feature point in the second matching pair onto the coordinate system of the second image to obtain the coordinates of the first projection point;
[0142] S4. Determine the coordinate difference between the first projection point and the first feature point associated with the first projection point as the first coordinate difference;
[0143] S15. Count the number of matching pairs in the second matching pair whose first coordinate difference is less than the second threshold, and repeat steps S11, S12, S13, S14 and S15 to obtain the number of matching pairs corresponding to each iteration. Determine the homography matrix corresponding to the maximum value among the number of matching pairs as the target homography matrix.
[0144] S16. Based on the target homography matrix, project the first feature point in the initial matching pair onto the coordinate system of the second image to obtain the coordinates of the second projection point;
[0145] S17. Determine the coordinate difference between the second projection point and the first feature point associated with the second projection point as the second coordinate difference, and remove the matching pairs in the initial matching pairs whose second coordinate difference is greater than the third threshold to obtain the optimized matching pairs.
[0146] Homography matrix is an important concept in computer vision. It describes the projective transformation relationship between point pairs on two images from different viewpoints. Mathematically, a homography matrix is a 3x3 matrix used to map a point on one plane to a corresponding point on another plane. Since the homography matrix has 8 degrees of freedom, at least 4 pairs of feature points need to be solved, so the number of first matching pairs is at least 4.
[0147] The process of calculating the homography matrix based on the first matching pair is as follows: construct a homogeneous linear equation system A·h=0 using the first matching pair, where h is a vector composed of 9 elements of the homography matrix H, and A is a matrix composed of the coordinate information of the feature points in the first matching pair; solve the homogeneous linear equation system by the singular value decomposition method to obtain the homography matrix H.
[0148] The coordinates of the first projection point can be obtained by multiplying the coordinates of the first feature point by the homography matrix. The coordinates of the second projection point can be obtained by multiplying the first feature point by the target homography matrix.
[0149] The first feature point associated with the first projection point refers to the feature point obtained by projection of the first projection point. For example, if the first feature point A is projected onto the coordinate system of the second image to obtain the first projection point B, then the first feature point associated with the first projection point B is the first feature point A.
[0150] The first coordinate difference includes the difference in the horizontal and vertical coordinates between the first projected point and the first feature point associated with it. The second and third thresholds can be preset.
[0151] In a specific application, the determination of optimized matching pairs can be achieved using the RANSAC (Random Sample Consensus) algorithm. Specifically, four first matching pairs are randomly selected from the initial matching pairs. The homography matrix is calculated, and the first feature points in the second matching pairs are projected onto the coordinate system of the second image using the homography matrix. The first coordinate difference is calculated, and inliers with an error of less than 1 pixel are retained. This random sampling and inlier counting process is repeated 1000 times. The homography matrix with the most inliers is selected as the target homography matrix. The first feature points in the initial matching pairs are projected onto the coordinate system of the second image using the target homography matrix to obtain the coordinates of the second projected points. The second coordinate difference is calculated, and matching pairs with a second coordinate difference greater than a third threshold are eliminated. Finally, the matching pair error rate is ensured to be less than 1%, resulting in optimized matching pairs.
[0152] In this embodiment, step S11 involves randomly selecting multiple initial matching pairs from the initial matching pairs as first matching pairs; the initial matching pairs that are not selected are designated as second matching pairs; step S12 involves calculating the homography matrix based on the first matching pairs; step S13 involves projecting the first feature point in the second matching pair onto the coordinate system of the second image based on the homography matrix to obtain the coordinates of the first projection point; step S4 involves determining the coordinate difference between the first projection point and the first feature point associated with the first projection point as the first coordinate difference; step S15 involves counting the number of matching pairs in the second matching pairs whose first coordinate difference is less than a second threshold, and repeating steps S11, S12, S13, S14, and S15 to obtain the number of matching pairs corresponding to each iteration, and then matching... S16. Based on the target homography matrix, the first feature point in the initial matching pair is projected onto the coordinate system of the second image to obtain the coordinates of the second projection point; S17. The coordinate difference between the second projection point and the first feature point associated with the second projection point is determined as the second coordinate difference, and matching pairs in the initial matching pair whose second coordinate difference is greater than the third threshold are eliminated to obtain optimized matching pairs. In this way, when the initial matching pair contains a large number of erroneous matches, these "outside points" can be automatically identified and eliminated, and "inside points" that conform to the global transformation model can be retained, i.e., optimized matching pairs. This provides reliable input for subsequent perspective transformation, image stitching, pose estimation, etc., and reduces visual defects such as stitching seams and distortion.
[0153] In one embodiment, calibrating the pose parameters of the camera based on optimized matching pairs includes the following steps:
[0154] S21. Preset the 3D coordinates and initial pose parameters of each pixel in the 3D scene of each frame of preprocessed image;
[0155] S22. Based on the initial pose parameters, determine the predicted position of each pixel's three-dimensional coordinates projected onto the preprocessed image;
[0156] S23. The difference between the predicted position of each pixel and the actual position in the preprocessed image is determined as the projection error;
[0157] S24. Based on the projection error, construct a nonlinear least squares optimization objective function;
[0158] S25. Calculate the minimum value of the nonlinear least squares optimization objective function using an iterative optimization algorithm, and update the initial pose parameters and three-dimensional coordinates in step S21 in reverse. Repeat steps S22 and S23 based on the updated pose parameters and three-dimensional coordinates until the projection error converges, and output the updated pose parameters.
[0159] The initial pose parameters include the rotation matrix and translation vector. The projection mapping formula for the predicted position is as follows: ,in, R is the projection function from 3D coordinates to 2D pixels, K is the intrinsic parameter matrix of the camera, and R is the projection function from 3D coordinates to 2D pixels. i It is the rotation matrix of the i-th preprocessed image, t i Let X be the translation vector of the preprocessed image in the i-th frame. j These are the three-dimensional coordinates of the j-th pixel. It is the predicted position of the three-dimensional coordinates of the j-th pixel in the i-th frame of the preprocessed image.
[0160] Projection error is the difference in coordinates between the predicted position of a pixel and its actual position in the preprocessed image. The formula for projection error is: x ij It represents the three-dimensional coordinates of the j-th pixel and its true location in the i-th preprocessed image. It is the predicted position of the j-th pixel in the i-th preprocessed image, e ij It is the projection error between the actual position and the predicted position of the j-th pixel in the i-th frame of the preprocessed image.
[0161] The nonlinear least squares optimization objective function is .
[0162] Iterative optimization algorithms can be damped least squares algorithms. Minimizing the nonlinear least squares optimization objective function involves the nonlinear coupling between the camera pose parameters and 3D coordinates. The damped least squares algorithm combines the fast convergence of the Gauss-Newton method with the robustness of gradient descent, introducing a damping factor on top of the linearized function. The formula for calculating the parameter update increment is: , where J is the Jacobian matrix of the optimization parameters, which include the three-dimensional coordinates and the initial pose parameters; Let be the parameter update amount; e be the error vector; λ be the damping factor; and I be the identity matrix. The algorithm iteratively adjusts the 3D coordinates and initial pose parameters until the change in projection error is less than 0.01 pixels, or until a maximum of 100 iterations are reached, thereby ensuring the consistency and optimality of pose estimation among multiple preprocessed images.
[0163] In this embodiment, step S21 involves presetting the 3D coordinates and initial pose parameters of each pixel in the 3D scene of each preprocessed image frame; step S22 involves determining the predicted position of each pixel's 3D coordinates projected onto the preprocessed image based on the initial pose parameters; step S23 involves determining the projection error as the difference between the predicted position and the actual position of each pixel in the preprocessed image; step S24 involves constructing a nonlinear least squares optimization objective function based on the projection error; and step S25 involves calculating the minimum value of the nonlinear least squares optimization objective function using an iterative optimization algorithm and updating the initial pose parameters and 3D coordinates in step S21 in reverse. Steps S22 and S23 are repeated based on the updated pose parameters and 3D coordinates until the projection error converges, and the updated pose parameters are output. This ensures the consistency and optimality of pose parameter estimation among multiple preprocessed images.
[0164] In one embodiment, the radar image is acquired by means of:
[0165] The returned linear frequency modulated echo signal is acquired, and range pulse compression processing is performed on the linear frequency modulated echo signal to obtain a pulse compressed signal; the range pulse compression processing includes deskewing, fast Fourier transform, and range migration effect correction;
[0166] The pulse compression signal is subjected to azimuth synthetic aperture radar imaging processing to obtain azimuth-compressed complex data.
[0167] Multi-view processing is performed on complex data, the processed data is converted into intensity images, and geocoding is performed to obtain radar images.
[0168] Among them, the linear frequency modulated (LFM) echo signal is the signal reflected back after the transmitted LFM signal encounters a ground target.
[0169] Range pulse compression processing of linear frequency modulated echo signals is the first step in SAR (Synthetic Aperture Radar) processing, with the goal of improving range resolution.
[0170] Range refers to the direction in which the radar beam points, that is, the straight line from the radar to the ground.
[0171] Pulse compression is the process of compressing a received long-duration, low-power linear frequency modulated echo signal into a narrow, high-amplitude pulse signal. Pulse compression greatly improves range resolution.
[0172] De-scoring involves mixing the received linear frequency modulated (LFM) echo signal with a "reference scoping signal," which can be the conjugate of the transmitted signal. This down-converts the LFM echo signal to a low-frequency or zero-frequency signal, reducing the data rate and computational load of subsequent processing.
[0173] The Fast Fourier Transform transforms the deskewed signal from the time domain to the frequency domain. Range migration correction corrects the echo trajectory of each target, "straightening" or "aligning" it to the same range gate.
[0174] Multi-view processing involves dividing the azimuth data into multiple "views," imaging each view independently, and then performing an incoherent average of these view images. Converting the multi-view processed data into an intensity image can be achieved by taking the square of the modulus of the complex data.
[0175] Geocoding refers to the process of converting intensity images from their inherent radar coordinate system to a standard geographic coordinate system.
[0176] In this embodiment, by acquiring the returned linear frequency modulated echo signal, range pulse compression processing is performed on the linear frequency modulated echo signal to obtain a pulse compressed signal. The pulse compressed signal is then subjected to azimuth synthetic aperture radar imaging processing to obtain azimuth compressed complex data. The complex data is then subjected to multi-view processing, and the multi-view processed data is converted into an intensity image and geocoded to obtain a radar image. This enables high-resolution radar imaging and the acquisition of high spatial resolution radar images.
[0177] In one embodiment, the method further includes:
[0178] Determine the second neighbor region of each pixel in the radar image, and calculate the average intensity and standard deviation of the second neighbor region;
[0179] Based on the average intensity and standard deviation, the adaptive threshold corresponding to each pixel is obtained;
[0180] When the pixel value of a pixel is less than the corresponding adaptive threshold, the pixel is determined as a candidate feature point in the radar image;
[0181] Clustering algorithms are used to filter candidate feature points to obtain filtered feature points; the filtered feature points are used to determine the location information of the same target object in the radar image.
[0182] The size of the second neighboring region can be determined based on a second preset size. For example, if the first preset size is 5×5, then the 5×5 pixel area surrounding the pixel is defined as the second neighboring region, centered on the pixel.
[0183] The average intensity is the average pixel value of all pixels in the second neighboring region. The standard deviation is the standard deviation of the pixel values of all pixels in the second neighboring region.
[0184] The formula for the adaptive threshold is T=μ+kσ, where μ is the average intensity, σ is the standard deviation, k is an empirical coefficient, and T is the adaptive threshold.
[0185] The corresponding adaptive threshold refers to the adaptive threshold calculated based on the pixel values of each pixel in the second neighboring region where the pixel is located.
[0186] Clustering algorithms include DBSCAN (Density-Based Spatial Clustering of Applications with Noise). Specifically, a maximum distance ε between two feature points is preset. If the distance between candidate feature points is less than or equal to the maximum distance ε, they are considered "reachable" from each other. A minimum number of points, MinPts, is preset to form a dense region. If there are at least MinPts candidate feature points within an ε-radius around a candidate feature point, that candidate feature point is considered a core point. The DBSCAN algorithm is then run, traversing all candidate feature points to find all core points and points within an ε-radius. All core points and candidate feature points within their respective ε-radius ranges are then considered feature points.
[0187] In this embodiment, by determining the second neighboring region of each pixel in the radar image, the average intensity and standard deviation of the second neighboring region are calculated. Based on the average intensity and standard deviation, an adaptive threshold corresponding to each pixel is obtained. When the pixel value of a pixel is less than the corresponding adaptive threshold, the pixel is determined as a candidate feature point in the radar image. A clustering algorithm is used to filter the candidate feature points to obtain the filtered feature points. This can effectively overcome the problems of large speckle noise and uneven contrast in radar images, and achieve feature point filtering with high robustness and low false alarm rate.
[0188] In some embodiments, a panoramic radar image stitching system for fitting optical images is provided, including a turntable, a camera for optical images and a radar device for acquiring radar images coaxially mounted on the turntable, and the camera and radar device achieve coaxial propagation of the optical path and the radar beam through a beam splitter prism.
[0189] The beam splitter is made of zinc selenide (ZnSe) material with a thickness of 10mm;
[0190] The surface of the beam splitter prism is coated with a composite coating layer, including a millimeter-wave transmission film and a visible light reflection film;
[0191] After precision assembly, the optical axis offset between the beam splitter and the central axis of the radar transmission beam is less than 0.01 mm; the angular deviation is less than 0.001°; the optical axis offset is calibrated using a laser interferometer, and the angular deviation is finely adjusted using a six-axis adjustment frame.
[0192] Furthermore, the turntable uses ceramic bearings as rotating support components. The coefficient of thermal expansion of ceramic bearings is ≤1ppm / ℃, in order to reduce mechanical deformation caused by changes in ambient temperature and ensure the directional stability of the system over a wide temperature range.
[0193] The panoramic radar image stitching system for fitting optical images also includes an FPGA-based synchronization control circuit, which is used to achieve high-precision time synchronization between radar acquisition and camera imaging. The FPGA's internal global clock network uses an H-tree structure for routing, ensuring that the clock deviation between each camera and radar device is less than 5 picoseconds; the FPGA trigger signal path adopts an equal-length design, with the path length error controlled within 1mm, corresponding to a transmission delay of less than 0.1 nanoseconds; the trigger signal is transmitted using LVDS (Low Voltage Differential Signaling), which has strong immunity to common-mode noise; the radar device, camera, and FPGA-based synchronization control circuit are powered by independent LDOs (Low Dropout Regulators), with power supply ripple less than 10mV, preventing power supply coupling interference.
[0194] This application also provides an application scenario in which the above-described panoramic radar image stitching method based on fitted optical images is applied. Specifically, the application of the panoramic radar image stitching method based on fitted optical images in this scenario is as follows:
[0195] The overall flowchart of the panoramic radar image stitching method for fitting optical images is as follows: Figure 2As shown. Specifically, the camera and radar equipment are coaxially mounted on a turntable. The linear frequency modulated echo signal is subjected to range pulse compression processing to obtain a pulse-compressed signal; the pulse-compressed signal is then subjected to azimuth synthetic aperture radar imaging processing to obtain azimuth-compressed complex data; the complex data undergoes multi-view processing, converting the multi-view processed data into an intensity image, and geocoding is performed to obtain a single-frame radar image. The position information of the target object in the radar image is obtained, and the centroid coordinates of the target object in the optical image are also obtained; the centroid coordinates and position information constitute cross-modal feature point pairs; based on multiple sets of cross-modal feature point pairs, an affine transformation model from the radar coordinate system to the optical coordinate system is established, the unknown parameters in the affine transformation model are calculated, and the solution results of the unknown parameters are substituted into the affine transformation model to obtain the affine transformation matrix; based on the affine transformation matrix, each pixel in the radar image is mapped to the optical coordinate system to obtain the transformed radar image; multiple frames of transformed radar images are stitched together to obtain a panoramic radar image. Optical images acquired at different acquisition times are preprocessed to obtain multiple preprocessed images. The SIFT algorithm is used to detect feature points in each preprocessed image, and descriptors for each feature point are calculated based on the pixel values of its first neighboring region. Based on the descriptors, the Euclidean distance between each first feature point in the first image and each second feature point in the second image is calculated. The first and second images are two adjacent preprocessed images; the first feature points are those in the first image, and the second feature points are those in the second image. Based on the Euclidean distance, each first feature point is determined from each second feature point. Nearest neighbor and second nearest neighbor; taking any first feature point as the target point, when the ratio of the Euclidean distance between the target point and the nearest neighbor to the Euclidean distance between the target point and the second nearest neighbor is greater than a first threshold, the nearest neighbor and the target point are combined into an initial matching pair. This process is repeated for each first feature point to obtain multiple initial matching pairs. The RANSAC algorithm is used to optimize the initial matching pairs to obtain optimized matching pairs. The camera pose parameters are calibrated based on the optimized matching pairs, and perspective transformation is performed on each frame of preprocessed image based on the calibrated pose parameters to obtain transformed optical images. The transformed optical images are then stitched together to obtain an optical panoramic image.The process involves determining the optical reference point in the optical panoramic image and the radar mapping point in the radar panoramic image; the optical reference point and the radar mapping point are pixels of the same object; calculating the vertical coordinate difference between the optical reference point and the radar mapping point, and dividing the vertical coordinate difference by the camera's focal length to obtain a quotient; converting the quotient into an angle difference using the arctangent function; the angle difference is the deviation between the actual rotation angle and the theoretical rotation angle of the turntable; calculating the horizontal coordinate difference between the optical reference point and the radar mapping point; based on the horizontal coordinate difference, vertical coordinate difference, and angle difference, calculating the proportional, integral, and derivative terms using a PID control algorithm; superimposing the weighted sum of the proportional, integral, and derivative terms onto the encoded value corresponding to the actual rotation angle to obtain the corrected target encoded value; and controlling the rotation of the turntable based on the target encoded value.
[0196] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0197] Based on the same inventive concept, this application also provides a panoramic radar image stitching device for implementing the panoramic radar image stitching method for fitting optical images as described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the panoramic radar image stitching device for fitting optical images provided below can be found in the limitations of the panoramic radar image stitching method for fitting optical images described above, and will not be repeated here.
[0198] In one embodiment, such as Figure 3 As shown, a panoramic radar image stitching device for fitting optical images is provided, comprising:
[0199] The image acquisition module is used to acquire multiple sets of synchronously acquired optical and radar images;
[0200] The information acquisition module is used to acquire the position information of the target object in the radar image and the centroid coordinates of the target object in the optical image; the centroid coordinates and position information constitute a cross-modal feature point pair;
[0201] The parameter calculation module is used to establish an affine transformation model from the radar coordinate system to the optical coordinate system based on multiple sets of cross-modal feature point pairs, calculate the unknown parameters in the affine transformation model, and substitute the solution results of the unknown parameters into the affine transformation model to obtain the affine transformation matrix.
[0202] The image conversion module is used to map each pixel in the radar image to the optical coordinate system based on the affine transformation matrix, so as to obtain the converted radar image.
[0203] The image stitching module is used to stitch together multiple frames of converted radar images to obtain a panoramic radar image.
[0204] The modules in the aforementioned panoramic radar image stitching device for fitting optical images can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0205] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores optical images, radar images, location information, centroid coordinates, affine transformation matrices, transformed radar images, and panoramic radar images. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a panoramic radar image stitching method that fits optical images.
[0206] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0207] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0208] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0209] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0210] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0211] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0212] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A panoramic radar image stitching method fitting an optical image, characterized by, The method comprises: acquiring multiple groups of synchronously collected optical images and radar images; acquiring position information of a target object in the radar images and acquiring a centroid coordinate of the target object in the optical images; the centroid coordinate and the position information constitute a cross-modal feature point pair; based on multiple groups of the cross-modal feature point pairs, an affine transformation model from a radar coordinate system to an optical coordinate system is established, unknown parameters in the affine transformation model are calculated, and a solution of the unknown parameters is substituted into the affine transformation model to obtain an affine transformation matrix; based on the affine transformation matrix, each pixel point in the radar image is mapped to the optical coordinate system to obtain a converted radar image; multiple frames of the converted radar images are spliced to obtain a radar panoramic image.
2. The method of claim 1, wherein, A camera for collecting the optical images and a radar device for collecting the radar images are coaxially installed on a turntable; The method further comprises: an optical panoramic image after splicing of optical images collected at different collection times is acquired, and an optical reference point in the optical panoramic image and a radar mapping point in the radar panoramic image are determined; the optical reference point and the radar mapping point are pixel points of the same object; a vertical coordinate difference between the optical reference point and the radar mapping point is calculated, and the vertical coordinate difference is divided by a focal length of the camera to obtain a quotient, and the quotient is converted into an angle difference through an inverse tangent function; the angle difference is a deviation between an actual rotation angle and a theoretical rotation angle of the turntable; a horizontal coordinate difference between the optical reference point and the radar mapping point is calculated; based on the horizontal coordinate difference, the vertical coordinate difference and the angle difference, a proportional term, an integral term and a differential term are calculated through a PID control algorithm; a weighted sum of the proportional term, the integral term and the differential term is added to an encoded value corresponding to the actual rotation angle to obtain a corrected target encoded value, and the rotation of the turntable is controlled based on the target encoded value.
3. The method of claim 2, wherein, The acquisition of the optical panoramic image after splicing of optical images collected at different collection times comprises: image preprocessing is performed on the optical images collected at different collection times to obtain multiple frames of preprocessed images; feature points in each frame of the preprocessed images are detected, and a descriptor of each feature point is calculated based on pixel values in a first adjacent region of the feature point; based on the descriptors, an Euclidean distance between each first feature point in a first image and each second feature point in a second image is calculated; the first image and the second image are two adjacent frames of preprocessed images, the first feature point is a feature point in the first image, and the second feature point is a feature point in the second image; based on the Euclidean distances, a nearest neighbor point and a second nearest neighbor point of each first feature point are determined from the second feature points; with any first feature point as a target point, when a ratio of an Euclidean distance between the target point and the nearest neighbor point to an Euclidean distance between the target point and the second nearest neighbor point is greater than a first threshold value, the nearest neighbor point and the target point are combined as an initial matching pair, and each first feature point is traversed to obtain multiple groups of initial matching pairs; optimizing the initial matching pairs to obtain optimized matching pairs; calibrating pose parameters of the camera based on the optimized matching pairs, performing perspective transformation on each of the preprocessed images based on the calibrated pose parameters to obtain transformed optical images, and stitching each of the transformed optical images to obtain an optical panorama.
4. The method of claim 3, wherein, The descriptor of each of the feature points is calculated based on pixel values of a first neighboring area of each of the feature points, including: dividing the first neighboring area of the feature point into a plurality of sub-blocks with the feature point as the center; in each of the sub-blocks, calculating a horizontal gradient and a vertical gradient of each pixel point based on pixel values of each pixel point, and calculating a gradient amplitude and a gradient direction of each pixel point based on the horizontal gradient and the vertical gradient of each pixel point; classifying the gradient direction of each of the sub-blocks into a corresponding preset interval; generating the descriptor of each of the feature points based on the gradient direction and the corresponding gradient amplitude in each of the preset intervals.
5. The method of claim 3, wherein, The optimization of the initial matching pairs to obtain optimized matching pairs includes the following steps: S11. Randomly selecting a plurality of initial matching pairs from the initial matching pairs as first matching pairs; the initial matching pairs that are not selected are second matching pairs; S12. Calculating a homography matrix based on the first matching pairs; S13. Projecting the first feature points in the second matching pairs into a coordinate system of the second image based on the homography matrix to obtain coordinates of first projection points; S4. Determining a first coordinate difference value between the first projection points and the first feature points associated with the first projection points; S15. Counting the number of matching pairs in the second matching pairs whose first coordinate difference values are less than a second threshold value, and repeating steps S11, S12, S13, S14 and S15 to obtain the number of matching pairs corresponding to each iteration, determining the maximum value of the number of matching pairs as a target homography matrix; S16. Projecting the first feature points in the initial matching pairs into the coordinate system of the second image based on the target homography matrix to obtain coordinates of second projection points; S17. Determining a second coordinate difference value between the second projection points and the first feature points associated with the second projection points, and eliminating the matching pairs in the initial matching pairs whose second coordinate difference values are greater than a third threshold value to obtain optimized matching pairs.
6. The method of claim 3, wherein, The calibration of the pose parameters of the camera based on the optimized matching pairs includes the following steps: S21. Pre-setting three-dimensional coordinates of each pixel point in each of the preprocessed images in a three-dimensional scene and initial pose parameters; S22. Determining a predicted position of the three-dimensional coordinates of each pixel point projected into a preprocessed image based on the initial pose parameters; S23. Determining a projection error between the predicted position of each pixel point and the actual position in the preprocessed image; S24. Constructing a nonlinear least squares optimization objective function based on the projection error; S25, calculating the minimum value of the nonlinear least squares optimization objective function by using an iterative optimization algorithm, and inversely updating the initial pose parameters and three-dimensional coordinates in step S21, repeating steps S22 and S23 based on the updated pose parameters and three-dimensional coordinates until the projection error converges, and outputting the updated pose parameters.
7. The method of claim 1, wherein, The radar image is obtained by: Obtaining the returned linear frequency modulation echo signal, performing distance direction pulse compression processing on the linear frequency modulation echo signal to obtain a pulse compression signal; the distance direction pulse compression processing includes slope correction processing, fast Fourier transform and range migration effect correction; Performing azimuth direction synthetic aperture radar imaging processing on the pulse compression signal to obtain azimuth compressed complex data; Performing multi-view processing on the complex data, converting the multi-view processed data into an intensity image, and performing geographic coding to obtain a radar image.
8. The method of claim 1, wherein, The method further comprises: Determining a second adjacent region of each pixel point in the radar image, calculating the average intensity and standard deviation of the second adjacent region; Based on the average intensity and the standard deviation, obtaining an adaptive threshold value corresponding to each pixel point respectively; When the pixel value of a pixel point is less than the corresponding adaptive threshold value, determining the pixel point as a candidate feature point in the radar image; Using a clustering algorithm to screen the candidate feature points to obtain screened feature points; the screened feature points are used to determine the position information of the target object in the radar image.
9. A panoramic radar image stitching apparatus that fits an optical image, characterized by, The device comprises: An image acquisition module for acquiring multiple groups of synchronously acquired optical images and radar images; An information acquisition module for acquiring the position information of a target object in the radar image and acquiring the centroid coordinates of the target object in the optical image; the centroid coordinates and the position information constitute a cross-modal feature point pair; A parameter calculation module for establishing an affine transformation model from a radar coordinate system to an optical coordinate system based on multiple groups of cross-modal feature point pairs, calculating unknown parameters in the affine transformation model, and substituting the solution of the unknown parameters into the affine transformation model to obtain an affine transformation matrix; An image conversion module for mapping each pixel point in the radar image to the optical coordinate system based on the affine transformation matrix to obtain a converted radar image; An image stitching module for stitching multiple frames of the converted radar image to obtain a radar panoramic image. 10.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-9. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 8.
Citation Information
Patent Citations
A vehicle-mounted monitoring all-round image splicing method based on ultrasonic radar
CN109785232A
Method of stitching optical images for 360 degree surround view using cameras and radars mounted around a vehicle
CN110855906A