A multi-lens large field of view high-resolution imaging method, device, equipment and medium
By combining a central lens and a fisheye lens, and adopting distortion correction and motion feature matching methods, the imaging problem caused by fisheye lens distortion is solved, and the robustness and quality of the multi-lens imaging system in complex scenes are improved.
Patent Information
- Application Number
- CN202411880753.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-19
AI Technical Summary
When existing multi-lens imaging systems deal with fisheye lens distortion, feature extraction and matching are easily affected by factors such as distortion and illumination changes, resulting in decreased stitching accuracy and robustness, especially poor imaging quality in scenes with weak textures, drastic illumination changes, and dynamic scenes.
A combination of a central lens and four fisheye lenses is used to improve imaging quality through simultaneous image acquisition, distortion correction, motion feature matching, and image fusion processing.
It effectively handles weak textures, drastic lighting changes, or dynamic scenes, improves robustness and imaging quality, and ensures image stitching accuracy and stability.
Smart Images

Figure CN119809995B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of lens technology, and in particular to a multi-lens large-field-of-view high-resolution imaging method, device, equipment and medium. Background Art
[0002] The so-called large field of view is the ability to provide a relatively wide field of view. Existing imaging systems use a combination of wide-angle lenses and fish-eye lenses to achieve large field of view imaging requirements. However, although wide-angle lenses can obtain a larger field of view, the edge resolution decreases significantly. Fish-eye lenses can obtain a larger field of view, but the image distortion is serious. Therefore, it is impossible to meet the requirements of high resolution while meeting the requirements of a large field of view.
[0003] A Chinese patent with publication number CN109191415B proposes a universal image fusion method that emphasizes image registration by calculating a homography matrix based on a reference field of view and a current field of view. This patent only proposes fusing two image acquisition systems to improve imaging quality, but does not take into account the distortion problem of fisheye lenses. When processing fisheye images using this patented method, feature extraction and matching are easily affected by factors such as distortion and illumination changes, resulting in reduced stitching accuracy and robustness.
[0004] Chinese patent publication number CN112258581B proposes an on-site calibration method for a panoramic camera with multiple fisheye lenses. This method can address the issues caused by fisheye lens distortion, but it is prone to mismatching during feature matching, resulting in poor image quality, especially in scenes with repetitive or weak textures. It also cannot effectively handle moving objects in dynamic scenes. Chinese patent publication number CN104835118A proposes a method for capturing panoramic images using two fisheye cameras. This method is mainly targeted at only two opposing fisheye lenses and cannot effectively handle moving objects in dynamic scenes.
[0005] In summary, the existing multi-lens imaging system including fisheye lens still has room for improvement in imaging, especially for scenes with weak texture, drastic lighting changes and dynamic scenes, where the existing methods are difficult to meet the actual imaging needs. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a multi-lens large field of view high-resolution imaging method, device, equipment and medium, aiming to improve the quality of multi-lens imaging.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a multi-lens large-field-of-view, high-resolution imaging method, wherein the multi-lens includes a central lens and four fisheye lenses, wherein the four fisheye lenses are geometrically distributed around the central lens, and the field of view angles of the central lens and the four fisheye lenses overlap, wherein the central lens is used to capture an image of a central area, and the optical resolution of the central lens is higher than that of the fisheye lenses, and the fisheye lenses are used to capture a large-field-of-view image of a peripheral area; the method comprises:
[0009] Synchronously acquire images collected by the central lens and each fisheye lens;
[0010] Performing distortion correction on the images captured by each fisheye lens based on the camera calibration parameters to obtain a corrected fisheye image;
[0011] Perform motion feature matching on the rectified fisheye image and the center lens image;
[0012] The corrected fisheye image is registered to the viewing angle of the central lens image based on the motion feature matching result;
[0013] Perform fusion processing on the registered images;
[0014] The fused image is post-processed to obtain the final image.
[0015] In a second aspect, the present invention further provides a multi-lens, large-field-of-view, high-resolution imaging device, the multi-lens comprising a central lens and four fisheye lenses, the four fisheye lenses being geometrically distributed around the central lens, the field of view angles of the central lens and the four fisheye lenses overlapping, the central lens being used to capture images of a central area, and having a higher optical resolution than the fisheye lenses, and the fisheye lenses being used to capture images of a large field of view of a peripheral area; the device comprising:
[0016] An image acquisition unit, used for synchronously acquiring images captured by the central lens and each fisheye lens;
[0017] A distortion correction unit, configured to perform distortion correction processing on the images captured by each fisheye lens based on camera calibration parameters to obtain a corrected fisheye image;
[0018] A motion feature matching unit, configured to perform motion feature matching on the corrected fisheye image and the center lens image;
[0019] an image registration unit, configured to register the corrected fisheye image to the viewing angle of the central lens image based on the motion feature matching result;
[0020] A fusion processing unit, used for performing fusion processing on the registered images;
[0021] The post-processing unit is used to perform post-processing on the fused image to obtain a final image.
[0022] In a third aspect, the present invention also provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a multi-lens, large-field-of-view, high-resolution imaging method as described above is implemented.
[0023] In a fourth aspect, the present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes a multi-lens, large-field-of-view, high-resolution imaging method as described above.
[0024] Compared with the prior art, the present invention has the following advantages: a multi-lens large field of view, high-resolution imaging method, wherein the multi-lens includes a central lens and four fisheye lenses, the four fisheye lenses are geometrically distributed around the central lens, and the field of view angles of the central lens and the four fisheye lenses overlap, the central lens is used to capture images of the central area, and the optical resolution of the central lens is higher than that of the fisheye lenses, and the fisheye lenses are used to capture large field of view images of the peripheral areas; the method includes: synchronously acquiring images captured by the central lens and each fisheye lens; performing distortion correction processing on the images captured by each fisheye lens based on camera calibration parameters to obtain a corrected fisheye image; performing motion feature matching on the corrected fisheye image and the central lens image; aligning the corrected fisheye image to the perspective of the central lens image based on the motion feature matching result; performing fusion processing on the aligned image; and post-processing the fused image to obtain the final image. By performing distortion correction processing on the fisheye lens and feature matching and registration on the central lens, the present invention can effectively cope with complex scenes such as weak textures, drastic lighting changes, or dynamics, thereby improving robustness and imaging quality.
[0025] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A flowchart of a multi-lens, large-field-of-view, high-resolution imaging method provided by a specific embodiment of the present invention; Figure 2A schematic block diagram of a multi-lens, large-field-of-view, high-resolution imaging device provided by a specific embodiment of the present invention; Figure 3 A schematic block diagram of a computer device provided in accordance with a specific embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions of the present invention in conjunction with specific embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] The present invention is applicable to large scenes, wide ranges and visual information collection scenarios that require clear details in the central area. For example, in the field of security monitoring, it is used for real-time monitoring of large areas in large public places such as airport waiting halls, railway station squares, large shopping mall atriums, etc. It is necessary to capture clear individual characteristics in crowded places (the central lens plays a role), and also take into account the dynamics of large fields of view such as surrounding passages, entrances and exits (the fisheye lens is responsible). For another example, in the field of intelligent transportation, it can be installed at complex intersections. The central lens focuses on traffic lights and vehicle license plates and driving status details of the main lanes, while the fisheye lens covers the traffic flow and pedestrian conditions of the entire intersection and surrounding branches. For another example, in the acquisition of virtual reality / augmented reality content, it can provide broad and detailed image materials for creating immersive scenes, meeting the needs of panoramic immersion and close-up of key elements. It can also be used for panoramic mapping of local areas in geographic surveying, such as recording the topography around city landmarks, taking into account the details of the building itself and the situation of the surrounding environment.
[0030] The multi-lens, large-field-of-view, high-resolution imaging method of the present invention can be integrated into an imaging device, which can be a professional-grade industrial camera integration device designed as an independent device with a protective shell and a stable mounting structure, suitable for long-term stable operation in harsh outdoor environments; it can also be a core imaging module embedded in smart terminal devices (such as high-end smartphones, professional-grade tablets for on-site survey recording), vehicle-mounted intelligent vision systems, drone aerial photography gimbal camera components, etc.
[0031] A multi-lens large-field-of-view high-resolution imaging method includes a central lens and four fisheye lenses. The four fisheye lenses are geometrically distributed around the central lens, and the field of view angles of the central lens and the four fisheye lenses overlap. The central lens is used to capture images of the central area, and the optical resolution of the central lens is higher than that of the fisheye lenses. The fisheye lenses are used to capture large-field-of-view images of the surrounding areas.
[0032] Specifically, the central lens is a high-resolution fixed-focus lens, such as the Canon EF 50mm f / 1.2LUSM. Its optical resolution enables clear capture of minute details in the central area. For example, when photographing cultural relic restoration scenes, it can accurately depict the texture and color changes on the artifact's surface, as well as the delicate manipulation of restoration tools. The four fisheye lenses are models suitable for capturing a wide field of view, such as the Nikon AF DX Fisheye-Nikkor 10.5mm f / 2.8G ED. They are evenly distributed in a circular geometry around the central lens, with the optical axes of adjacent fisheye lenses angled at 90 degrees, ensuring no blind spots in both horizontal and vertical directions and full coverage of the surrounding area.
[0033] A multi-lens large-field-of-view high-resolution imaging method includes the following steps: S10-S60.
[0034] S10, synchronously acquiring images captured by the central lens and each fisheye lens.
[0035] In this embodiment, synchronously acquiring images from the central lens and each fisheye lens is the foundation for subsequent high-quality imaging. By triggering all lenses (referring to the central lens and each fisheye lens) for simultaneous exposure, which can be achieved through hardware triggering or software synchronization, all images are captured at the same time, ensuring time synchronization. Image data is then read from all lenses.
[0036] S20 , performing distortion correction processing on the images captured by each fisheye lens based on camera calibration parameters to obtain a corrected fisheye image.
[0037] In one embodiment, step S20 specifically includes the following steps: S201 - S203 .
[0038] S201: Obtain camera calibration parameters of each fisheye lens.
[0039] In this embodiment, a calibration field is constructed, and a checkerboard calibration plate is selected. The side length accuracy of the checkerboard grid is controlled at ±0.05mm to ensure the accuracy of the calibration benchmark. The calibration field is equipped with a uniform, stable and constant color temperature (such as 5500K) lighting source to avoid uneven light causing shadows or reflections to interfere with camera acquisition and affect the calibration accuracy. The four fisheye lenses are sequentially mounted on a pan-tilt fixture with high-precision displacement adjustment function. The pan-tilt accuracy can reach ±0.1° rotation and ±0.1mm translation, which is convenient for adjusting the relative position of the lens and the calibration plate to ensure that clear and complete calibration images are obtained in different postures.
[0040] For each fisheye lens, we set up multiple different shooting positions and angle combinations. For example, we captured images at 0.5m, 1m, and 2m from the calibration plate, at horizontal angles of ±30° and ±60°, and vertical angles of ±20° and ±40°. We collected at least 10 images at each position and angle to ensure that all possible imaging states of the lens were covered. In total, we acquired approximately 60 calibration images for each lens.
[0041] The Zhang Zhengyou calibration algorithm is used in combination with the characteristics of the fisheye lens for optimization. For the calculation of internal parameters, the correspondence between the pixel coordinates of the checkerboard corner points in the image and the world coordinates is used to minimize the reprojection error through multiple iterations, and accurately solve the focal length (for example, the focal length of a fisheye lens is finally determined to be 8mm, with an error range of ±0.02mm), the coordinates of the principal point (the X coordinate of the principal point is accurate to ±0.5 pixels, and the Y coordinate accuracy is the same) and the radial distortion coefficient (generally the second-order and third-order radial distortion coefficients are accurate to ±0.001) and other parameters. In terms of external parameters, based on the obtained internal parameters, the imaging coordinate changes of the same world coordinate point under different lens postures are used to calculate the rotation and translation matrices using a method based on singular value decomposition (SVD). The accuracy of each element of the rotation matrix is controlled at ±0.005, and the accuracy of the translation matrix is ±0.1mm, ensuring that the spatial position and posture parameters of the lens relative to the calibration plate are accurate.
[0042] S202: Select a distortion correction algorithm to correct the collected fisheye image.
[0043] In this embodiment, a second-order polynomial model is used for radial distortion correction: d =x(1+k1r 2 +k2r 4 ), y d =y(1+k1r 2 +k2r 4 ), where (x,y) is the original image pixel coordinate, (x d ,y d ) is the corrected coordinate, k1 and k2 are radial distortion coefficients. Taking a fisheye lens image as an example, k1 = 0.25, k2 = -0.08, by traversing the original image pixel by pixel, the correction coordinates are calculated according to the above formula. For non-integer coordinates, bilinear interpolation is used to obtain the corrected pixel values. For example, the coordinates of a pixel point at the edge of the original image are (100.3, 200.7), and the calculated corrected coordinates are (98, 202.5). The bilinear interpolation is performed using the pixel values of its four neighborhoods to obtain the accurate pixel value after correction, ensuring the continuity of the image geometry. For tangential distortion, the formula is used:
[0044] x d =x+[2p1xy+p2(r 2 +2x2 )],y d =y+[p1(r 2 +2y 2 )+2p2xy], where p1 and p2 are the tangential distortion coefficients. Combined with the radial distortion correction step, this correction simultaneously corrects the effects of tangential distortion, ensuring that image details, especially textures at edges, are faithfully restored. For example, a circular object, which appears elliptical before correction due to tangential distortion, will be accurately restored after correction.
[0045] The fisheye image is divided into multiple concentric annular regions, and different correction strategies are adopted based on the degree of distortion in different regions. For example, a low-order polynomial model is used in the central region (radius 0-50 pixels). Because the distortion is relatively small, a simple model can be used for efficient correction and avoids image blur caused by excessive calculation. The middle annular region (radius 51-150 pixels) combines a high-order polynomial with a linear stretch transformation to correct the distortion while adjusting the uniformity of pixel distribution. The edge region (radius 151 pixels to the image boundary) uses a correction method based on a mapping lookup table (LUT). A large number of typical coordinate distortion mapping relationships are pre-calculated based on calibration parameters and stored in the LUT. During real-time correction, the corrected coordinates can be quickly indexed and retrieved, improving the speed and accuracy of edge correction. For example, in a city panorama, the edge building outlines are corrected using the segmented model, resulting in straight lines, clear details, and no obvious jagged edges or stretching.
[0046] S203: Generate a corrected fisheye image.
[0047] During the correction process, the image quality is monitored and optimized in real time. Image gradient information is used to assess the clarity of the corrected image. If local areas are blurred due to correction (e.g., the gradient amplitude falls below a set threshold), the interpolation algorithm parameters are automatically adjusted or the local smoothing filter strength is increased to restore image details while ensuring accurate geometric shapes. For example, in a natural scenery image captured with a fisheye lens, the edges of leaves may initially appear blurred after correction. By adaptively adjusting the filter parameters, the leaf edges are re-sharpened, resulting in a clear and natural texture.
[0048] Maintain image color consistency. Since distortion correction may affect color distribution, a color histogram matching technique is used to compare and match the color histogram of the corrected image with the color histogram of the original calibration plate image, and adjust the gains of the red, green, and blue channels to ensure that the image color is not distorted. For example, when photographing flowers, the color vividness of the flowers after correction is consistent with the real scene.
[0049] For steps S201-S203, the distortion in the fisheye lens image is effectively eliminated through precise calibration of parameters and adaptation algorithm. The fisheye image distortion is accurately corrected, the straight object imaging is restored to straightness, and the contour of the circular object is not deformed.
[0050] S30: Perform motion feature matching on the corrected fisheye image and the center lens image.
[0051] In one embodiment, step S30 specifically includes the following steps: S301 - S304 .
[0052] S301 : Perform feature point detection on the calibrated fisheye image and the center lens image using a feature detector.
[0053] In this embodiment, a Harris corner detector is selected for feature point detection for the center lens image, and a FAST detector is selected for feature point detection for the eye image.
[0054] S302. For each detected feature point, calculate feature descriptors of the pixels surrounding it, wherein the resolution-related information between pixels is enhanced when calculating the feature descriptor of the center lens image, and the field of view-related information is enhanced when calculating the feature descriptor of the fisheye image.
[0055] In this embodiment, the center lens image feature descriptor enhances resolution information by selecting a 3×3 neighborhood window centered on the feature point and subdividing the window into nine subregions. For each subregion, the gradient magnitude histograms in eight directions (0°-360° is divided into eight equal intervals) are calculated, forming a 9×8=72-dimensional feature vector. A high-precision Sobel operator is used to calculate the gradient magnitude and direction, and the operator kernel is fine-tuned based on the pixel sampling characteristics of the center lens to make it more suitable for extracting high-resolution image details. For example, when capturing microscopic circuit images of circuit boards, the high-resolution, sensitive HOG descriptor can accurately encode the direction and intensity differences of subtle grayscale changes between pixels for feature points generated by fine structures such as circuit intersections and solder joint edges, ensuring that detailed features of different circuits can be distinguished. Even for similar circuit layouts, the Euclidean distance difference between their feature descriptors can reach more than 10, effectively improving subsequent matching recognition.
[0056] The fisheye image feature descriptor emphasizes field-of-view characteristics by constructing a descriptor using annular partitioning statistics combined with radial gradient features. A circular area centered at the feature point and with a radius R (dynamically determined based on the fisheye image's field of view and the feature point's position, typically ranging from 10 to 50 pixels) is divided into three concentric annular bands. The radial grayscale gradient integral is calculated for each annular band. Statistics such as the grayscale mean and variance of the pixels within the annular band are also calculated. This combination forms a 3×4=12-dimensional basic feature vector. This vector is then expanded by combining information about the distance from the feature point to the image center and the relative angle of the annular band within which it resides, ultimately generating a 20-dimensional descriptor. In panoramic natural scenery photography, such as large field-of-view features captured by a fisheye lens, such as the curves of lake shores and mountain outlines, this descriptor effectively summarizes the distribution trends of features within the field of view and their relative position to the image center. Even under varying lighting conditions or with partial occlusion of the scene, the success rate of matching similar contour features remains above 80%, enhancing the robustness of fisheye image feature matching.
[0057] S303: Compare the feature descriptors of the fisheye image with the feature descriptors of the center lens image to find all feature point pairs that meet a similarity threshold.
[0058] In this embodiment, the normalized cross correlation (NCC) method is used to calculate the similarity of feature descriptors. For the feature descriptor pairs of the center lens and the fisheye lens images, the NCC value is calculated, and the initial similarity threshold is set to 0.7. This threshold is obtained by statistical analysis of matching experiments on a large number of combined images of different scenes (including indoor, outdoor, static, dynamic and other types), and minimizes false matches while ensuring a certain number of matches. For example, in the shooting of indoor home scenes, when matching furniture arrangements and decorative details, when the NCC value is greater than 0.7, it can be preliminarily determined as a potential matching feature point pair. After this initial screening, a set of possible matching point pairs can be quickly filtered out from the massive feature point combinations, accounting for about 20%-30% of the total number of combinations, greatly reducing the subsequent precise matching calculation amount while retaining most of the real matching point pairs. In the NCC calculation process, to improve computational efficiency, integral image technology is used to quickly obtain the pixel sum and square sum of the local area and reduce repeated calculations. For a 1000×1000 pixel image, the feature point matching calculation time is shortened by about 40% compared with the traditional pixel-by-pixel calculation method, meeting the needs of imaging systems with high real-time requirements.
[0059] S304: Eliminate incorrectly matched feature point pairs.
[0060] In one embodiment, step S304 specifically includes the following steps: S3041 - S3048 .
[0061] S3041. Randomly select a portion of feature points from all feature point pairs as a feature point pair subset.
[0062] In this embodiment, taking into account computational efficiency and model representativeness, a proportional random sampling rule is set. For a full set containing N feature point pairs, the extraction ratio is between 70% and 85% depending on the complexity of the scene and the amount of data. For example, when shooting a relatively simple indoor meeting scene with a relatively uniform feature distribution, N is 500 feature point pairs, and a subset of 75%, or 375 feature point pairs, is extracted; while for complex and changeable outdoor sports event scenes, the number of feature point pairs may be as high as 2,000. In this case, 80%, or 1,600 point pairs, are extracted as a subset to ensure that subsequent model estimation can cover sufficiently rich image feature information while avoiding inefficiency caused by excessive calculation.
[0063] During the sampling process, we introduce stratified sampling, stratifying feature points based on their location in the image (divided into central areas, intermediate rings, and edge areas) and their density (sparse areas versus dense areas). For example, in a drone-photographed city panorama, we sample more point pairs from the dense layer of feature points in densely populated downtown areas, while sampling fewer points from sparse layers in relatively open parks. This ensures that each layer has sufficient samples for the initial model construction, ensuring that the constructed transformation model takes into account the characteristics of all image components from the outset, improving the model's universality.
[0064] S3042. Construct a transformation model of the relative position parameters and distortion parameters of the central lens and the four fisheye lenses, and solve the model parameters using the data of the feature point pair subset.
[0065] In this embodiment, a model structure based on perspective transformation is selected to fully consider the complex relative positions and imaging distortion relationships between the central lens and the four fisheye lenses. The model parameters include rotation angles (three dimensions, rotation angles θx, θy, θz around the X, Y, and Z axes), translation amounts (Tx, Ty, Tz), scaling factors (Sx, Sy, Sz), and radial distortion coefficients related to lens distortion (k1 and k2 for fisheye lens correction compensation) and tangential distortion coefficients (p1 and p2), totaling 12 parameters. For example, in an in-vehicle multi-lens surround view system, the relative positions of the lenses will change slightly due to bumps during vehicle driving. These parameters can dynamically capture real-time geometric changes between lenses to ensure accurate image matching.
[0066] When solving model parameters using feature point subset data, an objective function is constructed based on the least squares principle. The coordinate correspondence between the feature point subset in the source image (fisheye image) and the target image (center lens image) is substituted into the function. An iterative optimization algorithm (such as the Levenberg-Marquardt algorithm) is then used to find the optimal solution for the parameters. Taking industrial inspection scenarios as an example, after multiple iterations (typically 10-15) of imaging matching of multiple feature points on a product surface, the convergence accuracy of each parameter reaches: rotation angle error ±0.1°, translation error ±0.5mm, scaling factor error ±0.02, and distortion coefficient error ±0.001. This lays a solid foundation for subsequent accurate reprojection error calculation.
[0067] S3043. Calculate the reprojection errors of all matched feature point pairs according to the transformation model.
[0068] S3044. According to the reprojection error, the matched feature point pairs are divided into inliers and outliers. Inliers refer to points whose reprojection error is less than a set threshold, and outliers refer to points whose reprojection error is greater than the set threshold.
[0069] For steps S3043-3044, in this embodiment, the reprojection error calculation process is as follows: Based on the solved transformation model, the coordinates of the feature points in the fisheye image are projected into the coordinate system of the central lens image through model transformation. The Euclidean distance between the projected coordinates and the actual coordinates of the corresponding matching feature points in the central lens image is calculated as the reprojection error. In the intelligent warehouse monitoring scenario, the reprojection error is calculated one by one for feature points such as shelves and cargo stacks. The calculation formula is:
[0070]
[0071] Where (x proj ,y proj ) is the projection coordinate, (x actual ,x actual ) are the actual coordinates.
[0072] When setting the reprojection error threshold, the lens resolution, scene scale, and imaging accuracy requirements are comprehensively considered. For high-resolution (e.g., 50 million pixels or more), small scenes (e.g., shooting with precision instruments in a laboratory), and where high-precision matching is required, the threshold is set at 1-2 pixels; for low-resolution, large scenes (e.g., panoramic surveillance of a city from high altitude), the threshold is relaxed to 3-5 pixels. For example, in museum cultural relic security monitoring, in order to accurately capture the displacement details of the cultural relics, the threshold is set at 1.5 pixels, effectively distinguishing between reliable matches and mismatched points, and ensuring the integrity of the cultural relics in subsequent image fusion.
[0073] After traversing all matching feature point pairs, those with reprojection errors less than a set threshold are marked as inliers and included in subsequent model optimization. Those with reprojection errors greater than the threshold are identified as outliers (mismatches) and temporarily stored for subsequent analysis (for statistical analysis of mismatch causes, etc.). For example, in outdoor construction monitoring, after reprojection error screening, approximately 60%-70% of the initial matching point pairs are confirmed as inliers. Outliers are mainly caused by feature extraction deviations caused by dust on the construction site, equipment occlusion, and light shadows. Subsequent targeted removal of these outliers can significantly improve the reliability of image registration.
[0074] S3045. Re-estimate the transformation model using the interior points.
[0075] S3046: Update the transformation model according to the iteration requirement to obtain multiple transformation models.
[0076] In steps S3045-S3046, in this embodiment, the transformation model parameters are re-estimated using the newly selected set of inliers, and the least squares method combined with an optimization algorithm is again used to solve the problem. At this point, because the inliers more accurately reflect the true correspondence between the images, the model parameter updates converge toward a more precise direction. Taking the robot visual navigation scenario as an example, the lens imaging is affected by the dynamic environment during the robot's movement. Each re-estimation of the model can adjust the parameters in real time. For example, if the perspective distortion of the fisheye lens is aggravated due to proximity to an obstacle, the radial distortion coefficient and translation parameters can be corrected in time through inlier feedback to ensure visual positioning accuracy. After iteration, the model parameter error is reduced by approximately 30%-40% compared to the previous round.
[0077] An upper limit is set for the number of iterations, generally 3-5 times, to prevent excessive iterations from falling into local optimality or affecting real-time performance due to waste of computing resources. At the same time, a dynamic termination condition is introduced to monitor the rate of change of model parameters between two adjacent iterations. When the rate of change of each parameter is less than the set minimum value (such as the rate of change of the rotation angle is less than 0.01° / time, and the rate of change of the translation amount is less than 0.1mm / time), the iteration is terminated in advance. In the real-time mapping scenario of UAVs, facing the rapid flight and image collection of complex terrain, this efficient iteration strategy can not only ensure model accuracy, but also complete model updates within seconds, adapting to real-time data processing requirements.
[0078] S3047. Select the transformation model with the largest number of inliers and the highest degree of matching with the layout and imaging characteristics of the central lens and the four fisheye lenses from multiple transformation models as the final transformation model, wherein the degree of matching with the layout and imaging characteristics is measured by calculating the difference between the model and the geometric optical model of the central lens and the fisheye lens.
[0079] S3048: Deem feature point pairs that do not conform to the final transformation model as mismatches and remove them.
[0080] For steps S3047-S3048, in this embodiment, the degree of matching is measured by the difference between the calculation model and the geometric optical model of the central lens and the fisheye lens. Starting from the principle of optical imaging, the perspective transformation model and the ideal lens imaging ray tracing model are compared in terms of light propagation path, imaging viewing angle, projection relationship, etc. For example, by calculating the mean deviation of the incident angle of light predicted by the calculation model and the incident angle of the actual lens optical model, as well as the sum of squares of the difference between the imaging size scaling ratio and the theoretical value, the model matching is comprehensively quantified. The lower the difference, the higher the matching. In the 3D modeling multi-lens acquisition scene, this indicator can accurately screen out the transformation model that best fits the physical laws of real lens imaging, ensuring the geometric consistency of the model in the reconstruction of complex 3D scenes.
[0081] Feature point pairs that don't match the final optimal transformation model are completely eliminated. In panoramic video stitching applications, this rigorous screening virtually eliminates ghosting and tearing caused by mismatches at the splicing points of adjacent frames, resulting in smooth, natural images and sub-pixel stitching accuracy. This significantly improves the visual quality of panoramic videos, providing a stable, high-quality image foundation for immersive virtual reality experiences and long-term security monitoring and retracing.
[0082] For steps S3041-S3048, through a sophisticated mismatch elimination process, the selected feature point pairs can accurately reflect the true geometric correspondence between the center lens and the fisheye lens image, so that the subsequent image registration error is controlled within a very small range (sub-pixel level), and seamless connection is achieved in the fields of panoramic stitching and multimodal image fusion. For example, the target trajectory in the 360° panoramic picture of security monitoring is continuous and smooth, without image jumps or misalignments caused by inaccurate registration. In addition, whether the light changes drastically (such as the transition between light and shadow at sunrise and sunset), the scene is subject to large dynamic interference (such as busy traffic intersections) or the lens's own state fluctuates (such as the vibration of drone flight), this solution can adaptively adjust the transformation model to eliminate mismatches caused by environmental factors and ensure imaging stability. In the intelligent traffic violation evidence collection multi-lens system, the characteristics of the violating vehicle can still be accurately locked under complex road conditions, with an error rate of less than 2%, thereby improving the accuracy of law enforcement evidence.
[0083] S40 , registering the corrected fisheye image to the viewing angle of the central lens image based on the motion feature matching result.
[0084] In one embodiment, step S40 specifically includes the following steps: S401 - S403 .
[0085] S401. Based on the matched feature point pairs, relative motion estimation is performed on the corrected fisheye image and the center lens image. During the motion estimation process, the stability of the installation structure of the center lens and the four fisheye lenses and the dynamic changes in the overlapping area of the field of view are combined, and a motion estimation model based on the optical flow method is used for estimation.
[0086] In this embodiment, the stability of the mounting structure of the central lens and the four fisheye lenses has a significant impact on relative motion estimation. For example, in a vehicle-mounted multi-lens imaging system, vibrations may occur during driving, but the lenses are fixed to the vehicle body via a sturdy anti-shake mounting bracket, and their overall relative position remains stable to a certain extent, with only slight low-frequency shaking. In the motion estimation model, the constraint information of this mounting structure is incorporated, and reasonable displacement and rotation range limit parameters are set. For example, based on the mechanical properties of the mounting bracket and past experimental data, the horizontal and vertical translation displacement fluctuation range is set within ±0.5mm, and the rotation angle variation range around each coordinate axis is set within ±0.2°. This is used as prior knowledge to filter out optical flow calculation results that do not conform to actual physical constraints, thereby improving the accuracy and reliability of motion estimation.
[0087] Because lenses have overlapping fields of view, the motion of objects in the overlapping area or changes in the lens's own posture can cause dynamic changes in the overlapping area in different shooting scenarios. For example, in a multi-lens system that monitors indoor human activity, when a person moves in the overlapping area between the center lens and the fisheye lens, the feature distribution in the overlapping area changes at different times. In the optical flow model, dynamic weights are set for the overlapping area, assigning higher weights to feature point pairs in the overlapping area (for example, setting a weight of 0.8 and a weight of 0.2 for feature point pairs in the non-overlapping area). This allows motion estimation to focus more on feature changes in the overlapping area, as the feature correspondence in these areas is more critical for subsequent image registration and can more accurately reflect the relative motion between images.
[0088] Consider a multi-camera scenario at a sporting event. The central lens focuses on the athletes' movements at the center of the field, while the fisheye lens covers the surrounding spectators and the overall field conditions. Among the matched feature point pairs, one corresponds to the corner points of a billboard at the edge of the field. Using feature-based optical flow, combined with the aforementioned lens optimization, the horizontal displacement of these feature point pairs in two adjacent frames is calculated to be 5 pixels (considering the relationship between image resolution and actual field distance, this corresponds to a horizontal displacement of approximately 0.5 meters) and a vertical displacement of 3 pixels (approximately 0.3 meters). Furthermore, based on the comprehensive calculation of multiple feature point pairs, a small rotation angle of approximately 0.1° around a coordinate axis is determined. This displacement and rotation information constitutes the initial results of relative motion estimation, providing key data for the subsequent construction of the geometric transformation matrix.
[0089] S402 : Constructing a geometric transformation matrix based on the motion estimation result and the relationship between the optical imaging center coordinates and the field angles of the central lens and the four fisheye lenses.
[0090] In this embodiment, the central lens and the four fisheye lenses each have different optical imaging center coordinates. The relative positional relationship of these coordinates in three-dimensional space determines the geometric transformation relationship between the images. Through a precise lens calibration process in advance, the coordinate values of the optical imaging center of each lens in a unified coordinate system can be obtained (for example, the imaging center coordinates of the central lens are (0,0,0), and the imaging center coordinates of a certain fisheye lens are (10mm,5mm,3mm). The coordinate units here are determined according to the actual calibration accuracy and are generally accurate to ±0.1mm). When constructing the geometric transformation matrix, these coordinate differences are incorporated into the calculation of the translation transformation. For example, when calculating the translation matrix elements of the perspective transformation from the fisheye image to the central lens image, the distance parameters required for translation in the X, Y, and Z directions are calculated based on the coordinate differences to ensure that the images can be correctly aligned in space.
[0091] Different lenses have different field of views, and the differences in their coverage and angular ranges also need to be reflected in the geometric transformation. For example, the field of view of the center lens is 60° (horizontally), and the field of view of a fisheye lens is 180° (horizontally). When aligning the fisheye image to the perspective of the center lens image, it is necessary to determine parameters such as the scaling factor and rotation angle based on the proportional relationship of the field of view and the corresponding relationship of the imaging area. Through trigonometric function relationships and imaging geometry principles, the horizontal and vertical scaling ratios of the fisheye image relative to the center lens image are calculated (assuming the horizontal scaling ratio is 0.5 and the vertical scaling ratio is 0.6) as well as the possible rotation angles (such as a 10° clockwise rotation). These scaling, rotation and other parameters are integrated with the previous translation parameters into a complete affine transformation matrix to construct a geometric transformation matrix that conforms to the actual imaging characteristics of the lens.
[0092] To simplify the explanation using a two-dimensional plane (actually a three-dimensional space transformation, the principle is similar), assume that after the above calculation, the translation parameters are (Tx = 10, Ty = 5) (unit is pixel), the scaling factor is (Sx = 0.8, Sy = 0.7), and the rotation angle is 15° (clockwise), then the constructed affine transformation matrix (expressed in homogeneous coordinate form) is as follows:
[0093] This matrix is the geometric transformation matrix required to align the fisheye image to the central lens image perspective, and is used for subsequent geometric transformation operations.
[0094] S403 , performing geometric transformation and image compensation on the corrected fisheye image according to the geometric transformation matrix, so that the corrected fisheye image is aligned with the central lens image in terms of viewing angle.
[0095] In one embodiment, step S403 specifically includes the following steps: S4031 - S4033 .
[0096] S4031. Perform coordinate transformation on each pixel point on the corrected fisheye image according to the transformation matrix.
[0097] In this embodiment, the coordinates of each pixel in the corrected fisheye image are transformed according to the constructed geometric transformation matrix. For example, the pixel with coordinates (x, y) in the fisheye image is multiplied by the above affine transformation matrix (calculated in homogeneous coordinates) to obtain the transformed coordinates (x', y'). The calculation formula is as follows (simplified in two dimensions):
[0098]
[0099] S4032. Calculate pixel values of the fisheye image after coordinate transformation using an interpolation method.
[0100] After coordinate transformation, the newly obtained pixel coordinates are often not integer coordinates, while the pixel values of the image are defined at integer coordinate positions. Therefore, bilinear interpolation can be used to obtain the accurate pixel values after transformation. For example, if the transformed coordinates are calculated to be (10.3, 20.7), the pixel value at the position (10.3, 20.7) is calculated using the pixel values of the four integer coordinate pixels around the pixel point ((10, 20), (10, 21), (11, 20), (11, 21)) according to the weights of bilinear interpolation. This ensures the continuity and rationality of the pixel values after the geometric transformation of the image, and avoids image aliasing and blurring.
[0101] S4033. Map the pixel values of the transformed fisheye image to the viewing angle of the central lens image according to the interpolation result.
[0102] After the previous coordinate transformation and interpolation calculations, the fisheye image has been adjusted accordingly in terms of geometry and pixel values. However, it still needs to be accurately "placed" within the perspective of the central lens image to achieve complete alignment of the two in terms of perspective. This is the purpose of pixel value mapping.
[0103] During implementation, the coordinate system of the center lens image is used as a reference, and the pixel values of the fisheye image after transformation and interpolation are filled into the corresponding area according to their corresponding coordinate positions. For example, the coordinate range of the center lens image is from (0,0) to (W center ,H center ), where W center Indicates the number of pixels in the width direction, H center Indicates the number of pixels in the height direction. After the pixel coordinates of the transformed fisheye image are appropriately scaled and offset, their pixel values are filled into the coordinate positions corresponding to the central lens image.
[0104] If parts of the fisheye image exceed the coordinate range of the central lens image during the mapping process, this can be addressed based on the specific application scenario. For example, in applications like panoramic stitching, the excess area can be cropped or image fusion boundary processing techniques can be used to ensure a smooth transition with the central lens image. In scenarios where full information must be preserved, the display range of the central lens image can be expanded or a scrolling view can be used to display the complete transformed fisheye image.
[0105] Taking the multi-lens imaging application of drones taking panoramic city photos as an example, the central lens focuses on the city's landmark buildings, while the fisheye lens captures a wide field of view, including surrounding streets and building complexes. After the aforementioned coordinate transformation, interpolation, and mapping operations, the image information captured by the fisheye lens, such as distant streets and building outlines, can be accurately integrated into the central lens's image perspective. For example, a curved street originally at the edge of the fisheye image, after transformation and mapping, naturally appears at the edge of the central lens's image with a geometric shape and pixel representation that conforms to the central lens's perspective, seamlessly connecting with the landmark building image captured by the central lens to form a complete, continuous, and uniformly perspectived city panoramic image, providing a good foundation for subsequent image fusion, analysis, and other operations.
[0106] Steps S4031-S4033, through detailed coordinate transformation, interpolation, and perspective mapping, ensure high-precision alignment of the corrected fisheye image with the center lens image. In applications such as panoramic image stitching and multi-lens surveillance systems, the transition between images is natural, with no noticeable misalignment, gaps, or distortion. This results in a more realistic and complete, high-resolution, large-field-of-view image, providing superior image data support for security monitoring, geographic mapping, virtual reality, and other fields.
[0107] S50: performing fusion processing on the registered images.
[0108] In one embodiment, step S50 specifically includes the following steps: S501 - S508 .
[0109] S501: Obtain the spectral response characteristic differences between the central lens and the four fisheye lenses.
[0110] The spectral response characteristics of the central lens and the four fisheye lenses can be measured in advance using a high-precision optical spectrum analyzer, such as the Ocean Optics USB4000 model. This analyzer covers the visible and near-infrared wavelengths (350-1000nm) and offers a spectral resolution of up to 1.5nm. The central lens and the four fisheye lenses are sequentially mounted on a fixture on the optical platform, ensuring that the lens optical axis is strictly parallel to the incident optical axis of the spectrum analyzer, with an error within ±0.1°, to ensure measurement accuracy.
[0111] The measurement environment is set as a darkroom to eliminate stray light interference, the internal temperature is constant at 25℃±1℃, and the humidity is maintained at 40%-60% to avoid environmental factors affecting the performance of the lens optical components and spectral measurement results.
[0112] For each lens, a standard D65 light source (color temperature 6500K, color rendering index greater than 95) is used as illumination. The light source is evenly introduced through an optical fiber and vertically illuminates a Spectralon standard whiteboard with a diffuse reflectivity close to 100%. The lens collects the reflected light from the whiteboard and focuses it onto the optical fiber probe of the spectrum analyzer. When measuring each lens, the light intensity values at multiple wavelengths (at least 500 data points evenly distributed in the measurement band) are collected. The measurement is repeated 5 times and the average value is recorded as the spectral response curve of the lens. For example, the center lens has a high response in the 550nm green light band, with a light intensity value of 800 count units, while a certain fisheye lens has a relatively low response in this band, with a only 600 count units. This accurately quantifies the differences in the spectral response characteristics of each lens and provides a key data basis for subsequent color correction.
[0113] S502 : performing color deviation analysis on the images collected by each fisheye lens and the central lens under a standard color chart to construct a color correction matrix.
[0114] S503 : Perform color correction on the registered fisheye image and the center lens image according to the color correction matrix.
[0115] For steps S502-S503, select an X-Rite ColorChecker 24-color standard color chart and place it in the center of the lens' field of view, ensuring that the chart fills approximately 30% of the image area to capture sufficient color information. Use the center lens and four fisheye lenses to capture images of the color chart, with an image resolution of at least 2000 × 3000 pixels. Ensure that the boundaries of the color chart blocks are clearly visible and that each color block has a sampling point of at least 100 pixels.
[0116] The captured image is converted to a color space, converting the RGB data into the CIE Lab uniform color space. This space better aligns with the human eye's visual perception and facilitates accurate quantification of color deviations. For example, in the center lens image, the red patch of the standard color chart (theoretical CIE Lab coordinates are L*=53.24, a*=80.09, b*=67.20) has measured coordinates of L*=52.80, a*=78.50, and b*=65.80. The difference between these values and the theoretical values is calculated. Similar operations are performed for each patch in each fisheye lens and center lens image to accumulate color deviation data.
[0117] Based on the collected color deviation data, a 3×3 color correction matrix M is constructed using the least squares method. Assuming the RGB values of the original image pixels are the vector [xyz] and the corrected RGB values are [x'y'z']. By solving the linear equation system M×[x yz]=[x'y'z'], the corrected image colors are made as close as possible to the theoretical colors of the standard color chart. For example, calculations reveal that certain elements of the correction matrix, such as M(1,1)=1.02, M(1,2)=-0.03, and M(1,3)=0.01, are obtained. Optimization operations such as singular value decomposition are performed on the matrix to ensure its stability and accuracy, minimizing noise amplification or color distortion during the color correction process. This matrix is then used to perform pixel-by-pixel color correction on the registered fisheye image and the central lens image, restoring true colors.
[0118] S504: Obtain the difference in illumination reflection between the central area and the peripheral area.
[0119] S505 : Calculate the average brightness of different regions based on the brightness adjustment algorithm of the region segmentation, and perform brightness adjustment according to a preset brightness balance rule.
[0120] For steps S504-S505, in the actual shooting scene, use a light sensor (such as the Hagner E4 all-sky irradiance sensor) to measure the light intensity and reflectivity of the central area and the surrounding areas respectively. Place the sensor at the center of the central lens field of view and at typical positions of the surrounding fisheye lens field of view (such as edges, corners, etc.), and collect light data at different time periods (such as early morning, noon, and evening) and in different weather conditions (sunny, cloudy, and cloudy). The measurement time interval for each position is 1 hour. Continuously collect data for 24 hours and take the statistical mean to accurately obtain the light reflection difference characteristics. For example, at noon on a sunny day, the light intensity in the central area reaches 1000 lux and the reflectivity is 0.3, while the light intensity in the surrounding area is only 600 lux and the reflectivity is 0.25, which clearly shows the brightness imbalance.
[0121] The image is divided into a central area, an intermediate transition area, and a peripheral area by using a region division method based on image grayscale histogram threshold segmentation combined with morphological processing. For example, a bimodal analysis is performed on the image grayscale histogram to determine the thresholds T1 = 120 and T2 = 200. Grayscale values less than T1 are classified as peripheral areas, those greater than T2 are classified as central areas, and those between the two are classified as transition areas. For different areas, the average brightness value is calculated (using the integral image fast algorithm to calculate the sum of the regional pixel grayscale and then average it). According to the preset brightness balance rules, such as maintaining the brightness ratio of the central area to the peripheral area within the range of 1.2-1.5, the brightness of the darker peripheral areas is enhanced through linear stretching, gamma correction, and other algorithms, and the gain of the overly bright central area is appropriately reduced to make the overall image brightness uniform and natural.
[0122] S506: Based on the theoretical model of the field of view overlap between the central lens and the four fisheye lenses and the actual imaging calibration results, the lens field of view model is analyzed to calculate the overlapping area.
[0123] Based on the optical design parameters (focal length, field of view, image plane size, etc.) of the central lens and the four fisheye lenses, a theoretical model of field of view overlap was constructed using the principles of geometric optics. In a three-dimensional coordinate system, with the intersection of the lens optical axes as the origin, trigonometric functions were used to calculate the field of view boundary equations for each lens, determining the theoretical extent of the overlap area. For example, given a 60° field of view for the central lens and a 180° field of view for the fisheye lenses, with focal lengths of 50mm and 8mm, respectively, the overlap boundary angle at a certain horizontal distance from the center was calculated, and a mathematical model was established to describe the geometry of the overlap area.
[0124] S507 , in the overlapping area, using a weight distribution strategy to fuse the images, where the weight distribution is dynamically adjusted based on the image quality evaluation indexes and the importance of the field of view of the central lens and the four fisheye lenses.
[0125] Comprehensively consider image quality factors such as clarity, contrast, and noise level. Clarity is measured using a gradient-based variance calculation method, calculating the variance of the gradient amplitude for each small area of the image (such as an 8×8 pixel block). A larger variance indicates higher clarity. Contrast is measured by calculating the dynamic range (the difference between the maximum and minimum values) of the image's grayscale histogram. Noise level is assessed using the local standard deviation statistic, with a smaller standard deviation indicating lower noise. For example, the clarity index value for the central area of a central lens image reaches 80 (normalized to a range of 0-100), the contrast is 40, and the noise level is 5. However, the corresponding index values for the peripheral areas of a fisheye lens are 60, 30, and 8, respectively. This quantifies the difference in image quality between these areas.
[0126] Determine the importance of field of view based on the shooting target and application scenario. In the scenario of security monitoring focusing on key areas, the importance of the field of view of the center lens is set to 0.7, and the total importance of the field of view of the surrounding fisheye lenses is 0.3; in the panoramic display scenario, the weight of the field of view of the fisheye lens is appropriately increased, such as 0.4 for the center lens and 0.15 for the average weight of the fisheye lens. Dynamically allocate fusion weights based on image quality indicators. If the clarity of a fisheye lens image suddenly increases in the overlapping area (such as due to good local light), increase its weight accordingly so that the fused image can use the best imaging part of each lens in real time to generate a high-quality fusion result. A weighted average fusion algorithm is used to calculate the fused RGB value of the pixels in the overlapping area according to the weights to achieve a smooth image transition.
[0127] S508: Smoothing the edges of the fusion area.
[0128] The Sobel operator, combined with morphological dilation and erosion, is used to detect the edges of the fusion region, resulting in a transition zone with a width of approximately 3-5 pixels. For example, at the junction of panoramic images, edge detection determines the location of the boundary pixels, and the transition zone is expanded by 2 pixels on both sides of the edge. This area is then smoothed to avoid abrupt fusion marks.
[0129] A smoothing method based on Gaussian weighted averaging is used. For each pixel point in the transition area, a 5×5 neighborhood window is set with it as the center. The neighborhood pixel weights are calculated according to the Gaussian distribution. The closer to the center pixel, the greater the weight. The neighborhood pixels are weighted averaged according to the weights to calculate the new pixel value, replacing the original pixel value. The algorithm is iterated multiple times (usually 3-5 times) to make the fusion edge brightness and color gradient natural, eliminate splicing gaps and color mutations, and improve the overall visual quality of the image.
[0130] In steps S501-S508, precise lens spectral response measurement and color chart calibration significantly reduce color deviations in images captured by different lenses. Compensation for differences in light reflection and brightness adjustment ensure uniform overall image brightness, eliminating noticeable contrast between bright and dark areas. This significantly improves visual comfort. For example, in nighttime scenes, this prevents loss of objects due to excessive brightness in the center and darkness around the periphery. For panoramic indoor displays, this creates a natural light distribution effect, making it easier for the human eye to capture image details and enhancing the viewing experience. Furthermore, overlapping area determination, appropriate weighting, and edge smoothing ensure seamless image stitching and fusion, virtually eliminating gaps, misalignment, and color and brightness jumps at the joints.
[0131] S60: Post-process the fused image to obtain a final image.
[0132] In one embodiment, step S60 specifically includes the following steps: S601 - S604 .
[0133] S601: Use an edge detection algorithm to identify gaps and discontinuous areas in the fused image.
[0134] In this embodiment, the Canny edge detection algorithm is selected as the basis and optimized for the characteristics of multi-lens fusion images. Since the fused image may have weak edge traces left over from previous processing such as registration and color correction, as well as the complex situation of mixed object edges in the actual scene, the high and low thresholds are adaptively adjusted based on the traditional Canny algorithm. First, the local grayscale gradient histogram of the image is calculated, and the high threshold is dynamically set to the 80th percentile of the gradient amplitude and the low threshold is set to the 20th percentile based on the histogram distribution. This can not only capture strong edges, but also connect weak edges through the low threshold to form a complete outline. For example, in panoramic city street scene images, the edges at the junction of buildings and the sky and the joints of different buildings can be accurately located after this optimization, and the edge positioning accuracy can reach the sub-pixel level, with the error controlled at ±0.3 pixels.
[0135] Integrating multi-scale analysis techniques, the system employs a Gaussian pyramid to perform multi-resolution image decomposition, detecting edges at different scales and fusing the results from each scale to enhance the detection of edges in gaps and discontinuous areas of varying sizes. For example, subtle defects such as gaps between leaves and discontinuities in architectural decorative lines can be clearly located at high resolution. Large discontinuous areas, such as gaps between large billboards and building walls, can be quickly outlined at low, coarse resolution, improving both comprehensiveness and accuracy of detection.
[0136] S602: Using a texture synthesis algorithm, generate a filling texture according to the texture features of the area surrounding the gap, and fill the gap and discontinuous area.
[0137] In this embodiment, the area around the gap is focused on, and the texture features are extracted using the local binary pattern (LBP) combined with the grayscale co-occurrence matrix (GLCM). LBP is used to describe the local micro-texture structure of the image, while GLCM quantifies the spatial correlation of pixel grayscale to obtain feature information such as texture direction, roughness, and contrast. For example, in a multi-lens fusion landscape image containing grass and a cobblestone road, the gap is located in the cobblestone road area. By analyzing the surrounding normal cobblestone texture, its grayscale periodic variation law (grayscale co-occurrence frequency in a specific direction in GLCM) and local texture pattern (LBP eigenvalue distribution) are extracted, and key features such as the main direction of the texture is determined to be approximately horizontal and the average roughness is 0.4 (normalized value).
[0138] Based on the extracted texture features, suitable sample blocks are selected around the gaps for filling. A search strategy based on feature similarity is used, with a threshold set such as regions with a texture feature Euclidean distance less than 0.5 (normalized distance) as candidate sample blocks. Sample blocks adjacent to the gap and with coherent textures are prioritized to ensure a smooth transition in the filling texture. For example, when filling gaps in a cobblestone pavement, multiple 8×8 pixel sample blocks are selected from adjacent intact slabs, whose texture features closely match the gap, to avoid introducing abrupt textures.
[0139] Using a block-based texture synthesis method, selected sample blocks are filled into the gap area one by one according to the principle of optimal matching. Using the PatchMatch algorithm, the filling position is first randomly initialized, and then an iterative search for better matching blocks is performed, continuously updating the filling results. In each iteration, the gradient difference and texture feature similarity between the filling block and the surrounding synthesized area are calculated, and the filling direction and position are optimized to ensure the continuous and smooth texture after filling. For example, when filling a gap in grass, as the iteration progresses, the texture of the filled grass blades gradually merges with the surrounding area. After 5-8 iterations, the growth direction and density of the grass blades in the gap are consistent with the surrounding area, making the filling traces visually difficult to distinguish, and the fusion degree between the filled area and the original image exceeds 90%.
[0140] S603: Analyze the color temperature of white object images captured by the central lens and the four fisheye lenses to construct a white balance adjustment matrix.
[0141] During the initial calibration and regular calibration of the multi-lens system, images of standard white objects (such as Spectralon whiteboards, white areas of Kodak gray cards, etc.) are specifically captured to ensure that the white objects are evenly illuminated, the light intensity is controlled at 500-1000 lux, the color temperature is close to the D65 (6500K) light source conditions, and the image resolution is not less than 2000×2000 pixels to avoid inaccurate color sampling due to uneven lighting or low resolution.
[0142] Grayscale world hypothesis correction preprocessing is performed on captured white object images to bring the image's average grayscale value closer to an ideal neutral gray (e.g., 128). This removes initial color shifts caused by factors like exposure bias and facilitates subsequent accurate color temperature analysis. For example, a white board image captured with a fisheye lens initially has an average grayscale of 110, which is corrected to 127 after linear stretching, improving basic color accuracy.
[0143] Using a color temperature estimation algorithm based on Planck's blackbody radiation law, we analyze the relative radiant energy distribution of white object images in different wavelength bands (e.g., the 400-700nm visible light band, sampled at 10nm intervals). We then estimate the lens color temperature by matching it with a standard color temperature spectrum curve. For example, the color temperature of a white object captured by a central lens is calculated to be 5800K, while that captured by a fisheye lens is 6200K, which deviates from the standard D65 color temperature.
[0144] A 3×3 white balance adjustment matrix is constructed based on the color temperature differences between lenses. Using the VonKries diagonal model, with the standard D65 color temperature as the benchmark, the gain adjustment coefficients for the red, green, and blue channels of each lens are calculated and populated into the matrix's diagonal elements. For example, for a fisheye lens with a higher color temperature, the blue channel gain is appropriately reduced (e.g., a coefficient of 0.9), the red channel gain is slightly increased (e.g., a coefficient of 1.05), and the green channel remains relatively stable (a coefficient of approximately 1.0), ensuring uniform white balance across different lens images.
[0145] S604: Perform color balance adjustment on the filled image according to the white balance adjustment matrix.
[0146] The padded fused image is multiplied by the white balance adjustment matrix on a pixel-by-pixel basis. Assuming the RGB value of an image pixel is the vector [xyz], and the adjusted RGB value is [x'y'z'], the color value is altered pixel by pixel using the formula [x'y'z'] = white balance adjustment matrix × [xy z]. For example, the original RGB value of a pixel in an image is [100120110], but after adjustment using the corresponding lens white balance matrix, it becomes [103120107]. This precisely compensates for color temperature differences, resulting in a truly neutral white in the white area. Other colors are also balanced and harmonious, reducing the color error rate (compared to standard colors) by approximately 80%.
[0147] During the adjustment process, pixel values are dynamically range-limited and normalized to avoid overflow or loss of precision due to matrix multiplication. This ensures that adjusted RGB values are within the legal range of 0-255. If values exceed this range, they are truncated or linearly scaled to maintain rich color depth, prevent color blocking or color discontinuity, and ensure high-quality color rendering in the final image.
[0148] By performing distortion correction on the fisheye lens and feature matching and registration with the central lens, the present invention can effectively cope with complex scenes such as weak textures, drastic lighting changes or dynamics, thereby improving robustness and imaging quality.
[0149] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0150] The embodiment of the present invention further provides a multi-lens large field of view high resolution imaging device, which is used to perform the steps of any embodiment of the multi-lens large field of view high resolution imaging method described above. Figure 2 , Figure 2FIG. 1 is a schematic block diagram of a multi-lens, large-field-of-view, high-resolution imaging device 100 provided in an embodiment of the present application. The multi-lens, large-field-of-view, high-resolution imaging device 100 specifically includes:
[0151] The image acquisition unit 110 is used to synchronously acquire the images captured by the central lens and each fisheye lens. The distortion correction unit 120 is used to perform distortion correction processing on the images captured by each fisheye lens based on the camera calibration parameters to obtain a corrected fisheye image. The motion feature matching unit 130 is used to perform motion feature matching on the corrected fisheye image and the central lens image. The image registration unit 140 is used to align the corrected fisheye image to the perspective of the central lens image based on the motion feature matching result. The fusion processing unit 150 is used to perform fusion processing on the registered image. The post-processing unit 160 is used to perform post-processing on the fused image to obtain the final image.
[0152] In one embodiment, the distortion correction unit 120 is specifically configured to: obtain camera calibration parameters of each fisheye lens; select a distortion correction algorithm to correct the captured fisheye image; and generate a corrected fisheye image.
[0153] In one embodiment, the motion feature matching unit 130 is specifically configured to: perform feature point detection on the corrected fisheye image and the center lens image using a feature detector; for each detected feature point, calculate feature descriptors of the pixels surrounding the feature point, wherein the resolution-related information between pixels is enhanced when calculating the feature descriptor of the center lens image, and the field of view-related information is enhanced when calculating the feature descriptor of the fisheye image; compare the feature descriptors of the fisheye image with the feature descriptors of the center lens image to find all feature point pairs that meet a similarity threshold; and eliminate incorrectly matched feature point pairs.
[0154] In one embodiment, the motion feature matching unit 130 is further specifically configured to: randomly select a portion of feature points from all feature point pairs as a feature point pair subset; construct a transformation model of the relative position parameters and distortion parameters of the central lens and the four fisheye lenses, and solve the model parameters using data from the feature point pair subset; calculate the reprojection errors of all matched feature point pairs according to the transformation model; divide the matched feature point pairs into inliers and outliers according to the reprojection errors, wherein the inliers refer to points whose reprojection errors are less than a set threshold, and the outliers refer to points whose reprojection errors are greater than the set threshold; use the inliers to re-estimate the transformation model; update the transformation model according to the iteration requirements to obtain multiple transformation models; select the transformation model with the largest number of inliers and the highest degree of matching with the layout and imaging characteristics of the central lens and the four fisheye lenses from the multiple transformation models as the final transformation model, wherein the degree of matching of the layout and imaging characteristics is measured by calculating the difference between the model and the geometric optical model of the central lens and the fisheye lens; and regard the feature point pairs that do not conform to the final transformation model as mismatches and are eliminated.
[0155] In one embodiment, the image registration unit 140 is specifically used to: perform relative motion estimation on the corrected fisheye image and the center lens image based on the matched feature point pairs, and during the motion estimation process, combine the stability of the installation structure of the center lens and the four fisheye lenses and the dynamic changes in the field of view overlapping area, and use a motion estimation model based on the optical flow method to perform estimation; construct a geometric transformation matrix based on the motion estimation results and the relationship between the optical imaging center coordinates and the field of view angle of the center lens and the four fisheye lenses; and perform geometric transformation and image compensation on the corrected fisheye image based on the geometric transformation matrix to align the corrected fisheye image with the center lens image in terms of perspective.
[0156] In one embodiment, the image registration unit 140 is further specifically configured to: perform coordinate transformation on each pixel point on the corrected fisheye image according to a transformation matrix; calculate pixel values of the fisheye image after coordinate transformation using an interpolation method; and map the pixel values of the transformed fisheye image to the viewing angle of the central lens image based on the interpolation result.
[0157] In one embodiment, the fusion processing unit 150 is specifically used to: obtain the difference in spectral response characteristics between the central lens and the four fisheye lenses; perform color deviation analysis on the images captured by each fisheye lens and the central lens under a standard color card to construct a color correction matrix; perform color correction on the aligned fisheye image and the central lens image according to the color correction matrix; obtain the difference in light reflection between the central area and the peripheral area; calculate the average brightness of different areas based on a brightness adjustment algorithm based on area segmentation, and adjust the brightness according to a preset brightness balance rule; analyze the lens field of view model based on the theoretical model of the field of view overlap of the central lens and the four fisheye lenses and the actual imaging calibration results to calculate the overlapping area; within the overlapping area, use a weight distribution strategy to fuse the images, and the weight distribution is dynamically adjusted according to the image quality evaluation index and field of view importance of the central lens and the four fisheye lenses; and smooth the edges of the fusion area.
[0158] In one embodiment, the post-processing unit 160 is further specifically used to: use an edge detection algorithm to identify gaps and discontinuous areas in the fused image; use a texture synthesis algorithm to generate a filling texture based on the texture features of the area surrounding the gap, and fill the gaps and discontinuous areas; construct a white balance adjustment matrix by performing color temperature analysis on the white object images captured by the central lens and the four fisheye lenses; and perform color balance adjustment on the filled image according to the white balance adjustment matrix.
[0159] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned multi-lens large field of view high-resolution imaging device 100 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of the description, it will not be repeated here.
[0160] The multi-lens large field of view high resolution imaging device can be implemented in the form of a computer program. The computer program can be used in Figure 3 Runs on the computer equipment shown.
[0161] See also Figure 3 , Figure 3 700 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 700 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.
[0162] like Figure 3 As shown, the computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-lens, large-field-of-view, high-resolution imaging method described above are implemented.
[0163] The computer device 700 includes a processor 720 , a memory, and a network interface 750 connected via a system bus 710 , wherein the memory may include a non-volatile storage medium 730 and an internal memory 740 .
[0164] The non-volatile storage medium 730 can store an operating system 731 and a computer program 732. When the computer program 732 is executed, the processor 720 can execute a multi-lens large field of view high-resolution imaging method.
[0165] The processor 720 is used to provide computing and control capabilities and support the operation of the entire computer device 700.
[0166] The internal memory 740 provides an environment for the operation of the computer program 732 in the non-volatile storage medium 730. When the computer program 732 is executed by the processor 720, the processor 720 can execute the multi-lens large field of view high-resolution imaging method.
[0167] The network interface 750 is used for network communication, such as sending assigned tasks, etc. It will be understood by those skilled in the art that Figure 3The structure shown in the figure is merely a block diagram of a portion of the structure related to the present invention and does not limit the computer device 700 to which the present invention is applied. The specific computer device 700 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. The processor 720 is used to execute program code stored in the memory to implement the multi-lens large field of view, high-resolution imaging method.
[0168] Those skilled in the art will understand that Figure 3 The embodiment of the computer device shown in the figure does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structure and function of the memory and processor are the same as those in the figure. Figure 3 The embodiments shown are consistent and will not be described again here.
[0169] It should be understood that in the embodiment of the present application, the processor 720 may be a central processing unit (CPU), and the processor 720 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0170] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multi-lens, large-field-of-view, high-resolution imaging method disclosed in an embodiment of the present invention.
[0171] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or units, or may be an electrical, mechanical or other form of connection.
[0172] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A multi-lens large field of view high-resolution imaging method, characterized in that: The multi-lens camera includes a central lens and four fisheye lenses, the four fisheye lenses are geometrically distributed around the central lens, and the field of view angles of the central lens and the four fisheye lenses overlap with each other. The central lens is used to capture an image of a central area, and the optical resolution of the central lens is higher than that of the fisheye lenses. The fisheye lenses are used to capture a large field of view image of a peripheral area. The method includes: Synchronously acquire images collected by the central lens and each fisheye lens; Performing distortion correction on the images captured by each fisheye lens based on the camera calibration parameters to obtain a corrected fisheye image; Perform motion feature matching on the rectified fisheye image and the center lens image; The corrected fisheye image is registered to the viewing angle of the central lens image based on the motion feature matching result; Perform fusion processing on the registered images; Post-processing the fused image to obtain the final image; The fusion processing of the registered images includes: Obtain the differences in spectral response characteristics between the central lens and the four fisheye lenses; Perform color deviation analysis on the images captured by each fisheye lens and the central lens under a standard color chart to construct a color correction matrix; Performing color correction on the registered fisheye image and the center lens image according to the color correction matrix; Obtain the difference in light reflection between the central area and the surrounding area; The brightness adjustment algorithm based on regional segmentation calculates the average brightness of different areas and adjusts the brightness according to the preset brightness balance rules; Based on the theoretical model of the field of view overlap between the central lens and the four fisheye lenses and the actual imaging calibration results, the lens field of view model is analyzed to calculate the overlapping area; In the overlapping area, the images are fused using a weight distribution strategy, and the weight distribution is dynamically adjusted according to the image quality evaluation indicators and field of view importance of the central lens and the four fisheye lenses; Smooths the edges of the blended area.
2. The multi-lens large field of view high-resolution imaging method according to claim 1, characterized in that: The performing feature matching on the corrected fisheye image and the center lens image includes: Perform feature point detection on the rectified fisheye image and the center lens image using a feature detector; For each detected feature point, the feature descriptor of the surrounding pixels is calculated. The resolution-related information between pixels is enhanced when calculating the feature descriptor of the center lens image, and the field of view-related information is enhanced when calculating the feature descriptor of the fisheye image. Compare the feature descriptors of the fisheye image with those of the center lens image and find all feature point pairs that meet the similarity threshold; Eliminate incorrectly matched feature point pairs.
3. The multi-lens large field of view high-resolution imaging method according to claim 2, characterized in that: The step of eliminating mismatched feature point pairs includes: Randomly select a part of feature points from all feature point pairs as feature point pair subsets; Construct a transformation model of the relative position parameters and distortion parameters of the central lens and the four fisheye lenses, and use the data of the feature point pair subset to solve the model parameters; According to the transformation model, calculate the reprojection error of all matched feature point pairs; According to the reprojection error, the matched feature points are divided into inliers and outliers. The inliers refer to the points whose reprojection error is less than the set threshold, and the outliers refer to the points whose reprojection error is greater than the set threshold. Re-estimate the transformation model using inliers; updating the transformation model according to iteration requirements to obtain multiple transformation models; The final transformation model is selected from multiple transformation models with the largest number of inliers and the highest degree of match with the layout and imaging characteristics of the central lens and the four fisheye lenses. The degree of match with the layout and imaging characteristics is measured by calculating the difference between the model and the geometric optical models of the central lens and the fisheye lenses. Feature point pairs that do not conform to the final transformation model are considered as mismatches and are removed.
4. The multi-lens large field of view high-resolution imaging method according to claim 1, characterized in that: The registering the corrected fisheye image to the viewing angle of the central lens image based on the motion feature matching result includes: Based on the matched feature point pairs, the relative motion of the corrected fisheye image and the center lens image is estimated. During the motion estimation process, the stability of the mounting structure of the center lens and the four fisheye lenses and the dynamic changes in the overlapping area of the field of view are combined, and a motion estimation model based on the optical flow method is used for estimation. According to the motion estimation results, a geometric transformation matrix is constructed by combining the coordinate relationship of the optical imaging center and the field angle relationship between the central lens and the four fisheye lenses; The corrected fisheye image is geometrically transformed and image compensated according to the geometric transformation matrix, so that the corrected fisheye image is aligned with the central lens image in terms of viewing angle.
5. The multi-lens large field of view high-resolution imaging method according to claim 4, characterized in that: The step of performing geometric transformation and image compensation on the corrected fisheye image according to the geometric transformation matrix so as to align the corrected fisheye image with the central lens image in terms of viewing angle includes: Perform coordinate transformation on each pixel point on the corrected fisheye image according to the transformation matrix; Use interpolation method to calculate the pixel value of the fisheye image after coordinate transformation; According to the interpolation result, the pixel values of the transformed fisheye image are mapped to the viewing angle of the central lens image.
6. The multi-lens large field of view high-resolution imaging method according to claim 1, characterized in that: The post-processing of the fused image to obtain a final image includes: Use edge detection algorithms to identify gaps and discontinuous areas in the fused image; Using texture synthesis algorithm, fill texture is generated according to the texture features of the area around the gap, and gaps and discontinuous areas are filled; By analyzing the color temperature of white object images captured by the central lens and four fisheye lenses, a white balance adjustment matrix is constructed; Performs color balance adjustments on the filled image according to the white balance adjustment matrix.
7. A multi-lens, large-field-of-view, high-resolution imaging device, which, when in operation, executes the multi-lens, large-field-of-view, high-resolution imaging method according to any one of claims 1 to 6, characterized in that: The multi-lens camera includes a central lens and four fisheye lenses. The four fisheye lenses are geometrically distributed around the central lens, and the field of view of the central lens and the four fisheye lenses overlap. The central lens is used to capture images of the central area, and the optical resolution of the central lens is higher than that of the fisheye lenses. The fisheye lenses are used to capture images of the surrounding area with a large field of view. The device includes: An image acquisition unit, used for synchronously acquiring images captured by the central lens and each fisheye lens; A distortion correction unit, configured to perform distortion correction processing on the images captured by each fisheye lens based on camera calibration parameters to obtain a corrected fisheye image; A motion feature matching unit, configured to perform motion feature matching on the corrected fisheye image and the center lens image; an image registration unit, configured to register the corrected fisheye image to the viewing angle of the central lens image based on the motion feature matching result; A fusion processing unit, used for performing fusion processing on the registered images; The post-processing unit is used to perform post-processing on the fused image to obtain a final image.
8. A computer device, characterized in that: The invention comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for multi-lens large field of view and high resolution imaging as claimed in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes a multi-lens, large-field-of-view, high-resolution imaging method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for acquiring panorama image by using two fish-eye camera lenses
CN104835118A
Image fusion methods, apparatus and electronic equipment
CN109191415B
A Field Calibration Method for Multi-Fisheye Lens Panoramic Cameras
CN112258581B
High-fidelity fisheye lens distortion correction method
CN108961155A
Method and device for calibrating dual fisheye lens panoramic camera, and storage medium and terminal thereof
US20190236805A1