Fisheye image multi-view fusion method and fusion system
By combining RGB images with depth images and using technologies such as feature point matching, distortion correction, nonlinear optimization and weighted fusion, the problem of traditional fisheye image stitching methods ignoring depth information is solved, and efficient and high-quality fisheye image fusion is achieved.
Patent Information
- Application Number
- CN202410975088.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-19
AI Technical Summary
The traditional fisheye image stitching method ignores the important role of depth information in the fusion process, resulting in the low quality of the generated panoramic images and low fusion efficiency.
By combining RGB images with depth images for splicing, feature point matching, distortion correction, nonlinear optimization and weighted fusion, we can achieve efficient and high-quality fusion of fisheye images.
It significantly improves the generation efficiency and quality of panoramic images, effectively eliminates stitching gaps, and improves the overall effect of image fusion.
Smart Images

Figure CN118967469B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a fisheye image multi-view fusion method and a fusion system. Background Art
[0002] In the field of image processing, the generation and fusion of panoramic images is an important research direction and is widely used in many fields such as virtual reality, autonomous driving, and security monitoring. Among them, fisheye images, due to their ultra-wide-angle characteristics, can capture a wide field of view within a limited number of shots, becoming an important data source for panoramic image stitching. However, images captured by fisheye lenses usually have severe distortion, and image information from different perspectives needs to undergo complex correction and fusion processing to generate high-quality panoramic images.
[0003] Traditional fisheye image stitching methods mostly focus on the correction and stitching of RGB images from a single perspective, ignoring the important role of depth information in the fusion process.
[0004] Therefore, it is necessary to provide a fisheye image multi-view fusion method and a fusion system to solve the above technical problems. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a fisheye image multi-view fusion method and a fusion system, which not only improves the generation efficiency of panoramic images, but also significantly improves the quality of image fusion by combining RGB images and depth images for splicing.
[0006] The present invention provides a method for fusion of multiple views of fisheye images, the method comprising the following steps:
[0007] Collecting multiple fisheye images to be fused from multiple perspectives, wherein each fisheye image to be fused includes an RGB image and a depth image;
[0008] Based on the RGB image, generating feature point pairs of adjacent RGB images by a feature point matching method;
[0009] Estimate a transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix;
[0010] Constructing an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solving the objective function using a nonlinear optimization algorithm;
[0011] After the solution, the RGB images are fused in a unified coordinate system using a weighted fusion strategy based on the depth information of the depth image, wherein the depth information is used to guide the allocation of weights to the RGB images during the fusion process.
[0012] Preferably, the collecting of multiple fisheye images to be fused from multiple viewing angles includes:
[0013] Use an RGB-D fisheye camera to capture the fisheye images to be fused from various preset viewing angles;
[0014] Separating the captured fisheye image to be fused into an RGB image and a depth image, wherein the separated RGB image is used to provide color information, and the depth image is used to provide spatial information;
[0015] Determine the overlapping areas of adjacent RGB images and adjacent depth images.
[0016] Preferably, the step of generating feature point pairs of adjacent RGB images by a feature point matching method based on the RGB image and the depth image includes:
[0017] For the overlapping areas of adjacent RGB images, Gaussian blur of different scales is used to construct a scale space and apply Gaussian blur and differential detection at multiple scales to identify key points and corresponding descriptors at different scales.
[0018] Initial matching pairs between adjacent RGB images are screened by calculating the Euclidean distance between the descriptors of key points and applying a ratio test;
[0019] The initial matching pairs are verified using the RANSAC method to generate feature point pairs.
[0020] Preferably, the estimating the transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs includes:
[0021] Apply the distortion correction model corresponding to the RGB-D fisheye camera and use the known distortion parameters to correct the distortion of the overlapping areas of adjacent RGB images and adjacent depth images.
[0022] Based on the verified feature point pairs, the transformation matrix including rotation and translation between the overlapping areas of adjacent RGB images is estimated by direct linear transformation using DLT.
[0023] A rotation matrix and a translation matrix are decomposed from the estimated transformation matrix, wherein the rotation matrix is used to indicate an angle change, and the translation matrix is used to indicate a displacement change.
[0024] Preferably, constructing an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solving the objective function using a nonlinear optimization algorithm includes:
[0025] Construct the objective function, the mathematical expression of the objective function is:
[0026]
[0027] Among them, E(T) is the objective function, T is the transformation matrix, ω i is the weight of the i-th feature point pair obtained based on depth information, is the error of the i-th feature point pair calculated based on the transformation matrix T, e i The mathematical expression of (T) is:
[0028]
[0029] in, and Respectively represent the homogeneous coordinates of the i-th feature point in adjacent RGB images, and Respectively represent the depth value of the i-th feature point in adjacent RGB images;
[0030] Through a nonlinear optimization algorithm, the gradient of the objective function is iteratively calculated and the transformation parameters are adjusted in the opposite direction, and the rotation matrix and the translation matrix are continuously updated until the objective function is minimized and the convergence condition is met.
[0031] Preferably, in a unified coordinate system, according to the depth information of the depth image, a weighted fusion strategy is used to fuse the RGB image, including:
[0032] The overlapping areas of the distortion-corrected adjacent RGB images and adjacent depth images are transformed into a common reference coordinate system;
[0033] Assigning a fusion weight to each pixel corresponding to the overlapping area of the adjacent RGB images based on the depth value of each pixel in the overlapping area of the adjacent depth images, wherein the fusion weight is inversely proportional to the depth value;
[0034] For each pixel in the overlapping area of adjacent RGB images, the weighted average RGB value is calculated according to the RGB values of the pixel at different viewing angles and the corresponding fusion weights;
[0035] The weighted average RGB value of each pixel in the overlapping area of adjacent RGB images is assigned to the pixel at the corresponding position in the fused image, and the fused image is synthesized by combining the non-overlapping areas of adjacent RGB images.
[0036] The present invention provides a fisheye image multi-view fusion system for a fisheye image multi-view fusion method, the fusion system comprising:
[0037] An image acquisition module, used for acquiring a plurality of fisheye images to be fused from a plurality of viewing angles, wherein each fisheye image to be fused includes an RGB image and a depth image;
[0038] A feature matching module, used to generate feature point pairs of adjacent RGB images by a feature point matching method based on the RGB image and the depth image;
[0039] A geometric transformation estimation module, used for estimating a transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix;
[0040] An optimization module, configured to construct an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solve the objective function using a nonlinear optimization algorithm;
[0041] The fusion module is used to fuse the RGB images in a unified coordinate system after solving the problem by adopting a weighted fusion strategy according to the depth information of the depth image, wherein the depth information is used to guide the allocation of weights to the RGB images during the fusion process.
[0042] Compared with the related art, the fisheye image multi-view fusion method and fusion system provided by the present invention have the following beneficial effects:
[0043] The present invention utilizes the information complementarity of RGB images and depth images, and realizes efficient and high-quality fusion of fisheye images through the steps of feature point matching, distortion correction, nonlinear optimization and weighted fusion. Specifically, firstly, fisheye image data including RGB images and depth images are collected from different viewing angles. Then, feature point matching is performed based on the RGB image, and the depth image information is used to assist in generating a more accurate transformation matrix. Then, the accuracy of the transformation matrix is further improved by constructing an objective function and solving it using a nonlinear optimization algorithm. Finally, in a unified coordinate system, a weighted fusion strategy is used to fuse the RGB images according to the depth information of the depth image, thereby effectively eliminating the splicing gaps and improving the quality of the panoramic image. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A flow chart of a fisheye image multi-view fusion method provided by the present invention;
[0045] Figure 2 A module structure diagram of a fisheye image multi-view fusion system provided by the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention, rather than all structures, are shown in the accompanying drawings. In addition, the embodiments of the present invention and the features in the embodiments may be combined with each other without conflict.
[0047] It should also be noted that, for ease of description, only the part relevant to the present invention but not all content is shown in the accompanying drawings. It should be mentioned before discussing exemplary embodiments in more detail that some exemplary embodiments are described as processing or methods depicted as flow charts. Although the flow chart describes each operation (or step) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of each operation can be rearranged. When its operation is completed, the processing can be terminated, but it can also have additional steps not included in the accompanying drawings. The processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0048] Embodiment 1
[0049] The present invention provides a method for fusion of multiple views of fisheye images, referring to Figure 1 As shown, the fusion method comprises the following steps:
[0050] S1: Collect multiple fisheye images to be fused from multiple perspectives, where each fisheye image to be fused includes an RGB image and a depth image.
[0051] In this embodiment, multiple fisheye cameras fixed at different positions are used, and these cameras cover different angles of the entire scene. The fisheye image to be fused output by each camera includes RGB information and corresponding depth information, and all cameras must be synchronized during acquisition to avoid spatial dislocation due to time difference; through multi-view acquisition, complete coverage of the scene can be obtained, and even if some areas are blocked from a single perspective, images from other perspectives can provide supplementary information. In addition, the depth image provides distance information, which is crucial for subsequent fusion.
[0052] S2: Based on the RGB image, generate feature point pairs of adjacent RGB images by a feature point matching method.
[0053] In this embodiment, a feature detection algorithm is used to extract key points on the RGB image, and descriptors are used to describe the local environment of these points. Then, for each pair of adjacent images, a feature matching algorithm is used to find the correspondence between the same feature points in the two images. At the same time, considering the depth information, key points with clear boundaries on the depth image can be preferentially selected to improve the matching accuracy.
[0054] S3: Estimate a transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix.
[0055] In this embodiment, first, the distortion correction model of the fisheye lens is applied to remove the radial distortion in the fisheye image to be fused, and then the RANSAC (Random Sampling Consensus) algorithm is used to calculate the rotation matrix and the translation matrix, that is, the relative posture between adjacent images, based on the matching feature point pairs.
[0056] Distortion correction ensures the geometric accuracy of the image, while the estimation of the transformation matrix establishes the geometric relationship between images.
[0057] S4: constructing an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solving the objective function using a nonlinear optimization algorithm.
[0058] In this embodiment, an objective function is constructed, which mainly considers the error of the transformation matrix and the matching error between the RGB image and the depth image. Then, a nonlinear optimization algorithm is used to minimize this function to obtain the optimal transformation parameters, and the transformation matrix can be further refined to reduce the registration error, so that the image fusion is more accurate, and the final fused image will be smoother and have no obvious seams.
[0059] S5: After the solution, in a unified coordinate system, according to the depth information of the depth image, a weighted fusion strategy is used to fuse the RGB image, wherein the depth information is used to guide the allocation of weights to the RGB image during the fusion process.
[0060] In this embodiment, a weight is assigned to each pixel according to the depth information in a unified coordinate system. Generally, the closer the pixel is to the camera, the higher the weight is. Then, the RGB values of the same pixel position in the RGB image are weighted averaged to obtain the final fused image; the depth-guided weighted fusion strategy can effectively handle the conflict of overlapping areas, give priority to displaying foreground objects, and reduce the blur of the background, producing high-quality panoramic images that are more natural and coherent visually.
[0061] Specifically, step S1 includes the following steps:
[0062] S101: Using an RGB-D fisheye camera to capture fisheye images to be fused from various preset viewing angles.
[0063] In this embodiment, multiple fisheye cameras with RGB-D (red, green, blue, and depth) capabilities are fixed at pre-planned positions, which are selected to ensure that they can collectively cover the entire scene of interest. Each camera captures an image containing RGB information and depth data, which is then used for multi-view image fusion. In order to maintain image synchronization, all cameras should trigger the shutter at the same time to avoid spatial inconsistency problems caused by small differences in time.
[0064] S102: Separate the captured fisheye image to be fused into an RGB image and a depth image, wherein the separated RGB image is used to provide color information, and the depth image is used to provide spatial information.
[0065] In this embodiment, in the captured RGB-D image, RGB and depth data are usually recorded in parallel, but the storage format may be different. Therefore, it is necessary to separate these two data types from the original image. The RGB image provides color information for reconstructing the color details of the scene, while the depth image provides the distance information of each pixel relative to the camera, which is the key to reconstructing the three-dimensional structure of the scene.
[0066] S103: Determine overlapping areas of adjacent RGB images and adjacent depth images.
[0067] In this embodiment, it is first necessary to identify which RGB images have visual intersections, which can be achieved by checking the matching of feature points between images. Similarly, the overlapping areas of the depth image can also be determined by analyzing the depth data, especially those areas where the depth information is continuous and has clear boundaries. Identifying overlapping areas helps to accurately perform image registration and fusion in subsequent steps. In particular, during the fusion process, pixels in the overlapping area will be used to evaluate the degree of match between different images, thereby ensuring that the final panoramic image has neither visual breaks nor obvious seams, and looks more coherent and natural as a whole.
[0068] Specifically, step S2 includes the following steps:
[0069] S201: For the overlapping areas of adjacent RGB images, Gaussian blurs of different scales are used to construct a scale space and apply Gaussian blur and differential detection at multiple scales to identify key points and corresponding descriptors at different scales.
[0070] In this embodiment, the scale-invariant feature transform (SIFT) is used to detect key points in different scale spaces, where the scale space is created by applying Gaussian blur filters of different scales to the RGB image, which helps detect features at different sizes. At each scale, extreme points are found as key point candidates, which are local maxima or minima in the scale space, which ensures the scale invariance of the selected key points. Next, the Gaussian difference (DoG) operator is applied to the neighborhood around each key point to further refine the position of the key point and calculate a descriptor for each key point. The descriptor is usually a vector that characterizes the local texture information of the image around the key point.
[0071] By detecting keypoints at multiple scales, we obtain feature points that are invariant to image scaling and rotation, which helps find consistent matches between images from different viewpoints, even when objects in the scene change. The use of descriptors ensures that keypoints can be reliably identified and matched even when lighting conditions change.
[0072] S202: Filter out initial matching pairs between adjacent RGB images by calculating the Euclidean distance between the descriptors of the key points and applying a ratio test.
[0073] In this embodiment, after obtaining the key points and their descriptors, the nearest neighbor search method is used to find the nearest neighbor and the next nearest neighbor of each descriptor in the key point descriptor set of another image. If the distance ratio of the nearest neighbor to the next nearest neighbor is less than a certain threshold (e.g., 0.8), it is considered that a matching pair has been found. This is because if a descriptor is very similar to another descriptor, it should be more similar than any other descriptor. This ratio test can exclude false matches.
[0074] S203: Using the RANSAC method to verify the initial matching pair and generate feature point pairs.
[0075] In this embodiment, RANSAC (RANdom SAmple Consensus) is an iterative algorithm for estimating the parameters of a mathematical model from a set of observed data while identifying and excluding outliers. In this step, RANSAC is used to verify and refine the initial matching pairs obtained from S202. RANSAC randomly selects the minimum number of matching pairs, calculates the model parameters, and then checks how many of the remaining matching pairs are consistent with the calculated model. This process is repeated multiple times, each time selecting a different subset of matching pairs, and finally selecting the model parameters containing the most consistent points as the best estimate.
[0076] Specifically, step S3 includes the following steps:
[0077] S301: Apply a distortion correction model corresponding to the RGB-D fisheye camera and use known distortion parameters to perform distortion correction on overlapping areas of adjacent RGB images and adjacent depth images.
[0078] In this embodiment, the fisheye lens will produce significant barrel distortion due to its ultra-wide-angle characteristics. By performing distortion correction corresponding to the distortion correction model of the RGB-D fisheye camera, the distortion of the image edge can be significantly reduced, making the matching between images more accurate, and also providing a good foundation for subsequent image registration and fusion.
[0079] S302: Based on the verified feature point pairs, estimate the transformation matrix including rotation and translation between the overlapping areas of adjacent RGB images through DLT direct linear transformation.
[0080] In this embodiment, direct linear transform (DLT) is a linear method for estimating the projection relationship between two images from corresponding points, i.e., the homography matrix. Given at least four pairs of verified feature point pairs, a linear system of equations can be set and solved to estimate the homography matrix H. This matrix contains rotation information and translation information, and can be used to project points in one image to corresponding positions in another image.
[0081] S303: Decomposing a rotation matrix and a translation matrix from the estimated transformation matrix, wherein the rotation matrix is used to indicate an angle change, and the translation matrix is used to indicate a displacement change.
[0082] In this embodiment, the transformation matrix can be decomposed into two parts: rotation and translation. The rotation matrix is obtained by normalizing the upper left 3x3 part of the transformation matrix, and the translation vector can be extracted from the last column of the transformation matrix. The rotation and translation information is clearly separated for easy understanding and application.
[0083] By decomposing the transformation matrix, we can clearly understand the changes in the relative rotation and translation between the two images, which is crucial for understanding the geometric relationship of the scene and performing accurate image registration and fusion operations.
[0084] Specifically, step S4 includes the following steps:
[0085] S401: construct the objective function, the mathematical expression of the objective function is:
[0086]
[0087] Among them, E(T) is the objective function, T is the transformation matrix, ω i is the weight of the i-th feature point pair obtained based on depth information, is the error of the i-th feature point pair calculated based on the transformation matrix T, ei The mathematical expression of (T) is:
[0088]
[0089] in, and Respectively represent the homogeneous coordinates of the i-th feature point in adjacent RGB images, and Respectively represent the depth value of the i-th feature point in the adjacent RGB images.
[0090] In this embodiment, by constructing the objective function, we are able to quantify the inconsistency between feature point pairs, which reflects the geometric differences between images. The use of weights ensures the importance of depth information in the error calculation, that is, feature point pairs closer to the camera have greater weights because they are more critical in the fusion process.
[0091] S402: Iteratively calculate the gradient of the objective function and adjust the transformation parameters in the opposite direction through a nonlinear optimization algorithm, and continuously update the rotation matrix and the translation matrix until the objective function is minimized and meets the convergence condition.
[0092] In this embodiment, the nonlinear optimization algorithm can effectively find the combination of transformation parameters that minimizes the objective function, thereby minimizing the reprojection error between feature point pairs. This step is crucial to improving the accuracy and robustness of image registration, because the optimization process takes into account the information of all feature point pairs and ensures the matching quality of the entire overlapping area. Ultimately, the obtained rotation matrix and translation matrix can accurately describe the relative position and orientation between adjacent RGB images, providing the necessary geometric information for image fusion.
[0093] Specifically, step S5 includes the following steps:
[0094] S501: Transforming overlapping areas of adjacent RGB images and adjacent depth images after distortion correction into a common reference coordinate system.
[0095] In this embodiment, the rotation and translation matrices estimated in step S3 are used to transform the overlapping areas of each image (RGB and depth) from their original coordinate systems to a unified reference coordinate system. This transformation is done by applying the above matrices to the coordinates of each pixel, ensuring that all image segments are aligned in the same coordinate frame, ready for the fusion process.
[0096] S502: assigning a fusion weight to each corresponding pixel in the overlapping area of adjacent RGB images based on the depth value of each pixel in the overlapping area of adjacent depth images, wherein the fusion weight is in inverse proportion to the depth value.
[0097] In this embodiment, in order to determine how the RGB values of each pixel should be mixed during the fusion process, a fusion weight is assigned to each pixel in the depth image according to its depth value. Specifically, pixels closer to the camera (i.e., with smaller depth values) receive higher weights, while pixels farther away from the camera (i.e., with larger depth values) receive lower weights. The purpose of this is to give priority to foreground objects because they are more visually eye-catching, while also reducing the impact of background objects.
[0098] S503: For each pixel in the overlapping area of adjacent RGB images, a weighted average RGB value is calculated according to the RGB values of the pixel at different viewing angles and the corresponding fusion weights.
[0099] In this embodiment, for each pixel in the overlapping area, the weighted average of the RGB values of the pixel position in all adjacent RGB images is calculated, where the weight is based on the fusion weight determined in step S502. By calculating the weighted average RGB value, information from different perspectives can be fused in the overlapping area while taking into account the influence of depth information, which helps to eliminate seams and inconsistencies and ensure the visual quality and coherence of the final fused image.
[0100] S504: assigning a weighted average RGB value of each pixel in the overlapping area of adjacent RGB images to the pixel at the corresponding position in the fused image, and combining the non-overlapping areas of the adjacent RGB images to synthesize the fused image.
[0101] In this embodiment, after the fusion of the overlapping areas is completed, the weighted average RGB value of each pixel is assigned to the pixel at the corresponding position in the fused image. At the same time, for the non-overlapping areas, the RGB value is directly copied from the original image without fusion. Finally, all the processed pixels are combined together to generate a complete, seamless fused image.
[0102] By applying the weighted average RGB value to the overlapping areas and combining the data from the non-overlapping areas, a high-quality panoramic image is created that not only covers the entire scene, but also appears visually smooth and consistent, with no obvious stitching marks. This approach ensures the efficiency and effectiveness of image fusion, and is very suitable for application scenarios that require wide viewing angles and high-resolution imaging.
[0103] The working principle of a fisheye image multi-view fusion method provided by the present invention is as follows: first, fisheye image data including RGB images and depth images are collected from different viewing angles; then, feature point matching is performed based on the RGB image, and the depth image information is used to assist in generating a more accurate transformation matrix; then, the accuracy of the transformation matrix is further improved by constructing an objective function and solving it using a nonlinear optimization algorithm; finally, in a unified coordinate system, a weighted fusion strategy is used to fuse the RGB images based on the depth information of the depth image, thereby effectively eliminating splicing gaps and improving the quality of the panoramic image.
[0104] Embodiment 2
[0105] The present invention provides a fisheye image multi-view fusion system for a fisheye image multi-view fusion method, referring to Figure 2 As shown, the fusion system comprises:
[0106] The image acquisition module 100 is used to acquire multiple fisheye images to be fused from multiple viewing angles, wherein each fisheye image to be fused includes an RGB image and a depth image.
[0107] The feature matching module 200 is used to generate feature point pairs of adjacent RGB images through a feature point matching method based on the RGB image and the depth image.
[0108] The geometric transformation estimation module 300 is used to estimate the transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix.
[0109] The optimization module 400 is used to construct an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solve the objective function using a nonlinear optimization algorithm.
[0110] The fusion module 500 is used to fuse the RGB images in a unified coordinate system after solving the problem, using a weighted fusion strategy based on the depth information of the depth image, wherein the depth information is used to guide the allocation of weights to the RGB images during the fusion process.
[0111] The present application is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0112] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, the storage medium including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0113] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
Claims
1. A fisheye image multi-view fusion method, characterized in that: The fusion method comprises the following steps: Collecting multiple fisheye images to be fused from multiple perspectives, wherein each fisheye image to be fused includes an RGB image and a depth image; Based on the RGB image, generating feature point pairs of adjacent RGB images by a feature point matching method; Estimate a transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix; Constructing an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solving the objective function using a nonlinear optimization algorithm; After the solution, in a unified coordinate system, the RGB images are fused using a weighted fusion strategy according to the depth information of the depth image, wherein the depth information is used to guide the allocation of weights to the RGB images during the fusion process; The fusion process specifically includes: The overlapping areas of the distortion-corrected adjacent RGB images and adjacent depth images are transformed into a common reference coordinate system; Assigning a fusion weight to each pixel corresponding to the overlapping area of the adjacent RGB images based on the depth value of each pixel in the overlapping area of the adjacent depth images, wherein the fusion weight is inversely proportional to the depth value; For each pixel in the overlapping area of adjacent RGB images, the weighted average RGB value is calculated according to the RGB values of the pixel at different viewing angles and the corresponding fusion weights; The weighted average RGB value of each pixel in the overlapping area of adjacent RGB images is assigned to the pixel at the corresponding position in the fused image, and the fused image is synthesized by combining the non-overlapping areas of adjacent RGB images.
2. The method for multi-view fusion of fisheye images according to claim 1, characterized in that: The collecting of multiple fisheye images to be fused from multiple viewing angles includes: Use an RGB-D fisheye camera to capture the fisheye images to be fused from various preset viewing angles; Separating the captured fisheye image to be fused into an RGB image and a depth image, wherein the separated RGB image is used to provide color information, and the depth image is used to provide spatial information; Determine the overlapping areas of adjacent RGB images and adjacent depth images.
3. The method for multi-view fusion of fisheye images according to claim 2, characterized in that: The method of generating feature point pairs of adjacent RGB images by a feature point matching method based on the RGB image and the depth image includes: For the overlapping areas of adjacent RGB images, Gaussian blur of different scales is used to construct a scale space and apply Gaussian blur and differential detection at multiple scales to identify key points and corresponding descriptors at different scales. Initial matching pairs between adjacent RGB images are screened by calculating the Euclidean distance between the descriptors of the key points and applying a ratio test; The initial matching pairs are verified using the RANSAC method to generate feature point pairs.
4. The method for multi-view fusion of fisheye images according to claim 3, characterized in that: The method of estimating the transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs includes: Apply the distortion correction model corresponding to the RGB-D fisheye camera and use the known distortion parameters to correct the distortion of the overlapping areas of adjacent RGB images and adjacent depth images. Based on the verified feature point pairs, the transformation matrix including rotation and translation between the overlapping areas of adjacent RGB images is estimated by direct linear transformation using DLT. A rotation matrix and a translation matrix are decomposed from the estimated transformation matrix, wherein the rotation matrix is used to indicate an angle change, and the translation matrix is used to indicate a displacement change.
5. The method for multi-view fusion of fisheye images according to claim 4, characterized in that: The constructing the objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solving the objective function using a nonlinear optimization algorithm, includes: Construct the objective function, the mathematical expression of the objective function is: in, is the objective function, is the transformation matrix, For the The weight of feature point pairs based on depth information, Based on the transformation matrix Calculated The error of feature point pairs, The mathematical expression is: in, and Respectively represent the first The homogeneous coordinates of the feature points, and Respectively represent the first The depth value of feature points; Through a nonlinear optimization algorithm, the gradient of the objective function is iteratively calculated and the transformation parameters are adjusted in the opposite direction, and the rotation matrix and the translation matrix are continuously updated until the objective function is minimized and the convergence condition is met.
6. A fisheye image multi-view fusion system, used to execute a fisheye image multi-view fusion method according to any one of claims 1 to 5, characterized in that: The fusion system comprises: An image acquisition module, used for acquiring a plurality of fisheye images to be fused from a plurality of viewing angles, wherein each fisheye image to be fused includes an RGB image and a depth image; A feature matching module, used to generate feature point pairs of adjacent RGB images by a feature point matching method based on the RGB image; A geometric transformation estimation module, used for estimating a transformation matrix between adjacent RGB images based on the distortion-corrected projection model and the matched feature point pairs, wherein the transformation matrix includes a rotation matrix and a translation matrix; An optimization module, configured to construct an objective function based on the transformation matrix and the matching error between the RGB image and the depth image, and solve the objective function using a nonlinear optimization algorithm; The fusion module is used to fuse the RGB images in a unified coordinate system after solving the problem by adopting a weighted fusion strategy according to the depth information of the depth image, wherein the depth information is used to guide the allocation of weights to the RGB images during the fusion process.
Citation Information
Patent Citations
Fisheye image splicing method and device
CN111461963A
Binocular image real-time splicing method, device and equipment based on weighted fusion strategy
CN117291804A