A dynamic surround view stitching method and system based on image overlap region feature perception

CN122453601APending Publication Date: 2026-07-24NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-24

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure QLYQS_1
    Figure QLYQS_1
Patent Text Reader

Abstract

The application discloses a kind of dynamic ring vision splicing method and system based on image overlap area feature perception, method includes based on the depth information and the feature information dynamic planning splicing path;According to the splicing path of planning, image splicing is executed, including the multi-band image fusion based on depth perception.The present application extracts local and global features of the overlapping area and calculates pixel-level depth information, constructs a geometric transformation model optimized by feature-depth combination, so that panoramic stitching can accurately identify the outline of close-range objects, and dynamically plan the optimal stitching line to bypass prominent objects;By searching the distance splicing template based on the real distance and automatically unifying the focal length of each camera, the field of view range is kept consistent;Through the multi-scale depth estimation network guided by features and joint reprojection error optimization, high-precision registration can still be maintained in scenes with few feature points or low overlap areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a dynamic surround view stitching method and system based on the perception of image overlap area features. Background Technology

[0002] With the advancement and development of technology, people's demand for information is increasing, and visual information, as the most effective means for humans to acquire information, accounts for the vast majority. However, due to the limited field of view of the human eye, a single camera cannot meet the requirements of a large field of view and high resolution. Therefore, video stitching technology has emerged, which combines video streams with overlapping areas captured by multiple cameras using certain registration and fusion techniques to form a panoramic video stream with a large field of view and high resolution. Video stitching technology has broad application prospects in many fields such as video surveillance, rail transit, video conferencing, virtual reality, and vehicle surround view.

[0003] Currently, in-vehicle surround view stitching systems are mainly designed for integrated vehicles. These systems simultaneously collect image information about the vehicle's surroundings, and after calibration, stitch them together to form a panoramic view, intuitively presenting the vehicle's location and surrounding environment. This significantly expands the driver's perception of their surroundings and effectively reduces traffic accidents. However, for trailer hitches with chain-like structures, since cameras are typically distributed between the front and body of the vehicle, the positional relationship between the cameras changes as the trailer hitch moves relative to the body. This necessitates real-time updates to stitching parameters, making traditional in-vehicle surround view stitching systems unsuitable.

[0004] However, existing video stitching methods still have some problems. First, most video stitching methods struggle to meet the increasing demands for real-time performance and visual quality. Second, when stitching 360-degree panoramic images from vehicles, the lack of unified control over the focal length of each camera before stitching results in inconsistent framing ranges for each camera, significantly increasing the difficulty of stitching. Furthermore, the traditional fixed stitching seams cause visual "tearing" or "invisibility" of close-up objects (such as shelves and people), affecting the clarity and visual realism of the panoramic image. These problems severely restrict the application and promotion of panoramic video stitching technology, necessitating a new technical solution to address these issues. Summary of the Invention

[0005] To address the aforementioned technical problems in related technologies, this invention proposes a dynamic surround-view stitching method and system based on image overlap area feature perception, which can overcome the above-mentioned shortcomings of the prior art.

[0006] To achieve the above-mentioned technical objectives, the technical solution of the present invention is implemented as follows: A dynamic lookaround stitching method based on image overlap region feature awareness includes the following steps: S1 acquires raw images from several cameras; S2 performs distortion correction on the original image; S3 extracts feature information of the overlapping region, the feature information including local features and global features; S4 calculates the depth information of the overlapping region; S5 dynamically plans the splicing path based on the depth information and the feature information; S6 performs image stitching according to the planned stitching path, including multi-band image fusion based on depth perception; S7 determines whether the splicing is complete. If not, return to step 1; otherwise, proceed to step 8. S8 outputs the stitched panoramic image.

[0007] Further, step 1 includes: The S101 uses multiple cameras to simultaneously capture video streams of the target scene; S102 decomposes the acquired video stream to obtain several independent image sequences; S103 performs color difference correction and exposure equalization on each individual image sequence.

[0008] Further, S2 includes: S201 acquires the internal and external parameters of the panoramic camera; S202 performs distortion modeling on the original image based on intrinsic and extrinsic parameters; S203 uses a distortion model to correct the distortion of the original image, resulting in a corrected image.

[0009] Further, S3 includes: S301 performs image feature extraction on the corrected image to obtain local and global features of the image; S302 constructs an image feature database to store feature descriptions and corresponding camera parameter indexes under different field of view angles; S303 calculates the feature similarity of overlapping regions to obtain preliminary feature matching point pairs.

[0010] Further, S4 includes: S401 uses a binocular ranging algorithm to calculate the true distance of objects in the overlapping area; S402 generates and retrieves distance stitching templates based on preset distance gradients, and extracts or calculates the initial stitching parameters at the current distance based on the real-time distance. The specific process is as follows: the relative pose and mapping table of the camera at different standard object distances are obtained in advance through calibration to form a template library. Based on the obtained real distance, the initial stitching parameters at the current distance are extracted or calculated from the library through linear interpolation or nearest neighbor matching algorithm. S403 employs a depth estimation network based on multi-scale feature fusion, which uses extracted local features as guiding information input to the network and outputs pixel-level depth maps of overlapping areas.

[0011] Further, S5 includes: S501 constructs a 3D point cloud overlapping region model based on the output pixel-level depth map; S502 introduces depth information as a weight into the geometric transformation model and calculates the dynamic geometric transformation parameters of the overlapping region by minimizing the reprojection error. S503 uses a dynamic programming algorithm to find the optimal stitching line with the minimum energy function in areas with high depth consistency, ensuring that the stitching seam bypasses nearby protruding objects.

[0012] Furthermore, the mathematical expression for the cost function of constructing the path search in S503 is: in, The color difference weights represent the pixel values ​​in the overlapping area. The geometric reprojection error representing the feature point pair. Represents the depth gradient cost. , , These are the preset weighting coefficients.

[0013] Further, S6 includes: S601 performs resampling and projection stitching on the overlapping area based on the planned dynamic path and the retrieved distance template parameters; S602 uses a deep learning-based image fusion network. Input the image pair to be stitched and its corresponding depth map, and the network will automatically adjust the fusion weights according to the distance of the objects. S603 performs temporal smoothing on the seams, using the path information from the previous frame to incrementally correct the current frame.

[0014] Further, S8 includes: S801 performs a quality assessment on the stitched panoramic image; If the quality assessment results do not meet the requirements, return to S1 and reassemble. If the quality assessment results meet the requirements, S803 will output the panoramic image to the display device.

[0015] A dynamic surround view stitching system based on image overlap region feature perception is provided, which uses the method described above to perform dynamic surround view stitching.

[0016] The beneficial effects of this invention are as follows: By extracting local and global features of overlapping areas and calculating pixel-level depth information, this invention constructs a geometric transformation model that jointly optimizes features and depth. This enables panoramic stitching to accurately identify the outlines of near-field objects and dynamically plan the optimal stitching line to bypass prominent objects, effectively solving the visual "tearing" or "invisibility" problems caused by traditional fixed stitching seams. By retrieving distance stitching templates based on real distances and automatically unifying the focal lengths of each camera, the framing range remains consistent, significantly reducing stitching difficulty and meeting real-time processing requirements. Through feature-guided multi-scale depth estimation networks and joint reprojection error optimization, high-precision registration can still be maintained in scenarios with few feature points or low overlap areas. By updating the relative pose parameters of the cameras in real time and performing temporal smoothing fusion, the invention automatically adapts to changes in camera position in dynamic motion scenarios such as chain trailers, eliminating parallax ghosting and distortion, and significantly improving the clarity, visual realism, and user experience of panoramic images. Detailed Implementation

[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the scope of protection of the present invention.

[0018] A dynamic lookaround stitching method based on image overlap region feature perception according to an embodiment of the present invention includes the following steps: Step 1: Acquire raw images from multiple cameras; Step 1 includes: Step 101: Simultaneously acquire video streams of the target scene using multiple cameras; Step 102: Decompose the acquired video stream to obtain multiple independent image sequences; Step 103: Perform color difference correction and exposure equalization processing on each independent image sequence.

[0019] Step 2: Perform distortion correction on the original image; Step 2 includes: Step 201: Obtain the internal and external parameters of the panoramic camera; Step 202: Determine the distortion model for the original image based on intrinsic and extrinsic parameters; Step 203: Use the distortion model to correct the distortion of the original image to obtain the corrected image.

[0020] Step 3: Extract feature information of overlapping regions; Step 3 includes: 301. Extract image features from the corrected image to obtain local features (such as ORB and SIFT operators) and global features (such as semantic segmentation maps or scene histograms). 302. Construct an image feature database to store feature descriptions and corresponding camera parameter indexes under different viewpoints; 303. Calculate the feature similarity of the overlapping regions to obtain preliminary feature matching point pairs.

[0021] Step 4: Calculate the depth information of the overlapping region; Step 4 includes: 401. Calculate the true distance of key objects in the overlapping area using a binocular ranging algorithm; 402. Generate and retrieve “distance stitching templates” based on preset distance gradients: Pre-obtain camera relative poses and mapping tables at different standard object distances (e.g., 2m, 5m, 10m) through calibration to form a template library; Based on the real-time distance obtained in step 401, extract or calculate the initial stitching parameters at the current distance from the library through linear interpolation or nearest neighbor matching algorithm. 403. Feature-guided refined depth estimation: A depth estimation network based on multi-scale feature fusion (such as an improved MonoDepth2 or PSMNet structure) is adopted. The local features extracted in step 301 are used as guiding information input into the network, and the pixel-level depth map of the overlapping area is output. The consistency of features is used to constrain the edge accuracy of depth estimation.

[0022] Step 5: Dynamically plan the stitching path based on depth and feature information; Step 5 includes: 501. Construct a 3D point cloud overlapping region model based on the depth map output in step 4; 502. Feature-Depth Joint Optimization Calculation: Depth information is introduced as a weight into the geometric transformation model. Higher feature registration weights are assigned to regions with drastic depth changes (edges of near-field objects). Dynamic geometric transformation parameters of overlapping regions are calculated by minimizing reprojection error. 503. Depth-based obstacle avoidance stitching path optimization: Using dynamic programming algorithm, the optimal stitching line with the minimum energy function is found in areas with high depth consistency (i.e. no significant protruding obstacles or flat scene depth), ensuring that the stitching seam bypasses nearby protruding scenes and avoids visual tearing.

[0023] Step 6: Perform image stitching according to the planned stitching path; Step 6 includes: 601. Based on the dynamic path planned in step 503 and the distance template parameters retrieved in step 402, perform resampling and projection stitching on the overlapping area; 602. Depth-aware multi-band image fusion: A deep learning-based image fusion network (such as an Encoder-Decoder structure fusion network) is adopted. The input is the image pair to be stitched and its corresponding depth map, which enables the network to automatically adjust the fusion weight according to the distance of the scene, enhance the details of the near scene, and smooth the transition of the far scene. 603. Perform temporal smoothing on the seams and use the path information of the previous frame to make incremental corrections to the current frame, reducing visual jumps in dynamic environments.

[0024] Step 7: Determine if the splicing is complete. If not, return to step 1; otherwise, proceed to step 8.

[0025] Step 8: Output the stitched panoramic image.

[0026] To facilitate understanding of the above technical solutions of the present invention, the following detailed description of the above technical solutions of the present invention will be provided through specific usage methods.

[0027] In practical application, the dynamic surround view stitching method based on image overlap area feature perception according to the present invention is implemented in the following specific steps: Step 1: Acquire raw images from multiple cameras; Step 101: Simultaneously acquire video streams of the target scene using multiple panoramic cameras, wherein each panoramic camera includes at least two lenses to acquire multiple images at different distances and angles. Step 102: Decompose the acquired video stream to obtain multiple independent image sequences; Step 103: Perform color difference correction and exposure equalization processing on each independent image sequence, using the histogram equalization method.

[0028] Step 2: Perform distortion correction on the original image; Step 201: Obtain the internal and external parameters of the panoramic camera, including lens focal length, image distance, field of view, and other parameters; Step 202: Determine the distortion model for the original image based on intrinsic and extrinsic parameters; Step 203: Use the distortion model to correct the distortion of the original image to obtain the corrected image.

[0029] Step 3: Extract feature information of overlapping regions; Step 301: Extract image features from the corrected image to obtain local and global features, including features such as edges, corners, and textures; Step 302: Construct an image feature database, including local features and global features; Step 303: Calculate the feature similarity of the overlapping regions and obtain the feature matching results.

[0030] Step 4: Calculate the depth information of the overlapping region; Step 401: Calculate the true distance of objects in the overlapping area using a binocular ranging algorithm. Use a binocular camera with a baseline length of 30mm and a ranging range of 0.5-15 meters. Step 402: Obtain the corresponding distance stitching template based on the actual distance, pre-generate distance stitching templates corresponding to n different distances, and pre-store the stitching parameters of the n different distances and the corresponding distance stitching templates in the memory of the panoramic camera, where n is an integer from 1 to 10; Step 403: Use a deep learning network to estimate the depth of the overlapping region. A depth estimation model based on a convolutional neural network is adopted, and the depth estimation accuracy is ±1 cm.

[0031] Step 5: Dynamically plan the stitching path based on depth and feature information; Step 501: Construct a spatiotemporally continuous three-dimensional overlapping region model: 5011. Point cloud fusion and denoising: Using the discrete depth data obtained in steps 401 and 403, combined with the camera's intrinsic and extrinsic parameters, the pixels in the overlapping area are mapped to a unified three-dimensional coordinate system to generate the initial point cloud. 5012. Dynamic Scene Reconstruction: A 3D model is constructed using point cloud representation. The spatiotemporal correlation between adjacent frames is used to smooth and filter the point cloud, eliminating outliers caused by sensor noise or the movement of dynamic objects, ensuring that the model can realistically reflect the geometric topological relationships of obstacles (such as shelves and cranes) in the overlapping area. Step 502: Calculation of geometric transformation parameters driven by joint feature-depth: 5021. Reprojection Error Modeling: When calculating geometric transformation parameters, not only the feature matching similarity of pixel planes should be considered, but depth constraints also need to be introduced. A reprojection error model based on the 3D coordinates of feature points is established. 5022. Parameter Solution and Transformation Model Selection: Based on the complexity of the overlapping area scenery, dynamically select the transformation model (such as homography transformation, affine transformation, or more complex Carthaginian / projection transformation); solve for the optimal geometric transformation matrix H by minimizing the joint loss function of "feature distance + depth difference"; 5023. Dynamically update pose parameters: For chain devices such as trailers, the relative pose parameters between cameras are corrected in real time using the transformation matrix obtained by solving, so as to provide a high-precision registration basis for path planning in step 503. Step 503: Optimize the stitching path using geometric transformation parameters and depth feature information, specifically including: 5031. Constructing the cost function for path search: Establish an objective function that comprehensively considers pixel differences, depth mutations, and feature alignment. Its mathematical expression is: in, The color difference weights represent the pixel values ​​in the overlapping area. The geometric reprojection error representing the feature point pair. Represents the depth gradient cost (used to identify and avoid the edges of nearby objects); , , These are preset weighting coefficients; 5032. Perform dynamic programming to find the optimal path: Using a dynamic programming algorithm, search for the path with the minimum cumulative cost in the energy map of the overlapping region as the optimal stitching line. This path will actively bypass regions with drastic depth changes or low feature matching, thereby avoiding the near-field "tearing" phenomenon; 5033. Perform smooth transition processing: With the optimal seam line as the center, set a transition width of 5-10 pixels, and use multi-band fusion or linear weighted fusion algorithm for incremental correction to eliminate visual abruptness.

[0032] Step 6: Perform image stitching; Step 601: Perform image stitching on the overlapping areas according to the stitching parameters of the distance stitching template. The stitching algorithm adopts an image fusion algorithm based on deep learning.

[0033] Step 602: Perform image fusion processing on the stitched images using a Doschler-based image fusion method.

[0034] Step 603: Apply special effects to the seams, using a median-based filtering method to reduce visual abruptness.

[0035] Step 7: Determine if the splicing is complete. If not, return to step 1; otherwise, proceed to step 8.

[0036] Step 8: Output the stitched panoramic image; Step 801: Evaluate the quality of the stitched panoramic image. Evaluation indicators include sharpness, detail, and color uniformity. Step 802: If the quality assessment result does not meet the requirements, return to step 1 to reassemble, and set the assessment threshold to 80 points. Step 803: If the quality assessment results meet the requirements, output the panoramic image to the display device.

[0037] In summary, by utilizing the technical solutions described above, this invention extracts local and global features of overlapping areas and calculates pixel-level depth information to construct a feature-depth joint optimization geometric transformation model. This enables panoramic stitching to accurately identify the outlines of nearby objects and dynamically plan the optimal stitching line to bypass prominent objects, effectively solving the visual "tearing" or "invisibility" problems caused by traditional fixed stitching seams. By retrieving distance stitching templates based on real distances and automatically unifying the focal lengths of each camera, the framing range remains consistent, significantly reducing stitching difficulty and meeting real-time processing requirements. Through feature-guided multi-scale depth estimation networks and joint reprojection error optimization, high-precision registration can still be maintained in scenarios with sparse feature points or low overlap areas. By updating the relative pose parameters of the cameras in real time and performing temporal smooth fusion, the system automatically adapts to changes in camera position in dynamic motion scenarios such as chain trailers, eliminating parallax ghosting and distortion, and significantly improving the clarity, visual realism, and user experience of panoramic images.

[0038] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A dynamic lookaround stitching method based on image overlap region feature perception, characterized in that, Includes the following steps: S1 acquires raw images from several cameras; S2 performs distortion correction on the original image; S3 extracts feature information of the overlapping region, the feature information including local features and global features; S4 calculates the depth information of the overlapping region; S5 dynamically plans the splicing path based on the depth information and the feature information; S6 performs image stitching according to the planned stitching path, including multi-band image fusion based on depth perception; S7 determines whether the splicing is complete. If not, return to step 1; otherwise, proceed to step 8. S8 outputs the stitched panoramic image.

2. The dynamic surround look stitching method based on image overlap region feature perception according to claim 1, characterized in that, Step 1 includes: The S101 uses multiple cameras to simultaneously capture video streams of the target scene; S102 decomposes the acquired video stream to obtain several independent image sequences; S103 performs color difference correction and exposure equalization on each individual image sequence.

3. The dynamic surround look stitching method based on image overlap region feature perception according to claim 1, characterized in that, S2 includes: S201 acquires the internal and external parameters of the panoramic camera; S202 performs distortion modeling on the original image based on intrinsic and extrinsic parameters; S203 uses a distortion model to correct the distortion of the original image, resulting in a corrected image.

4. The dynamic surround-look stitching method based on image overlap region feature perception according to claim 1, characterized in that, S3 includes: S301 performs image feature extraction on the corrected image to obtain local and global features of the image; S302 constructs an image feature database to store feature descriptions and corresponding camera parameter indexes under different field of view angles; S303 calculates the feature similarity of overlapping regions to obtain preliminary feature matching point pairs.

5. The dynamic surround look stitching method based on image overlap region feature perception according to claim 1, characterized in that, S4 includes: S401 uses a binocular ranging algorithm to calculate the true distance of objects in the overlapping area; S402 generates and retrieves distance stitching templates based on preset distance gradients, and extracts or calculates the initial stitching parameters at the current distance based on the real-time distance. The specific process is as follows: the relative pose and mapping table of the camera at different standard object distances are obtained in advance through calibration to form a template library. Based on the obtained real distance, the initial stitching parameters at the current distance are extracted or calculated from the library through linear interpolation or nearest neighbor matching algorithm. S403 employs a depth estimation network based on multi-scale feature fusion, which uses extracted local features as guiding information input to the network and outputs pixel-level depth maps of overlapping areas.

6. The dynamic surround-look stitching method based on image overlap region feature perception according to claim 5, characterized in that, S5 includes: S501 constructs a 3D point cloud overlapping region model based on the output pixel-level depth map; S502 introduces depth information as a weight into the geometric transformation model and calculates the dynamic geometric transformation parameters of the overlapping region by minimizing the reprojection error. S503 uses a dynamic programming algorithm to find the optimal stitching line with the minimum energy function in areas with high depth consistency, ensuring that the stitching seam bypasses nearby protruding objects.

7. The dynamic surround-look stitching method based on image overlap region feature perception according to claim 6, characterized in that, The mathematical expression for the cost function of path search in S503 is as follows: in, The color difference weights represent the pixel values ​​in the overlapping area. The geometric reprojection error representing the feature point pair. Represents the depth gradient cost. , , These are the preset weighting coefficients.

8. The dynamic surround-look stitching method based on image overlap region feature perception according to claim 5, characterized in that, S6 includes: S601 performs resampling and projection stitching on the overlapping area based on the planned dynamic path and the retrieved distance template parameters; S602 uses a deep learning-based image fusion network. Input the image pair to be stitched and its corresponding depth map, and the network will automatically adjust the fusion weights according to the distance of the objects. S603 performs temporal smoothing on the seams, using the path information from the previous frame to incrementally correct the current frame.

9. The dynamic surround look stitching method based on image overlap region feature perception according to claim 1, characterized in that, S8 includes: S801 performs a quality assessment on the stitched panoramic image; If the quality assessment results do not meet the requirements, return to S1 and reassemble. If the quality assessment results meet the requirements, S803 will output the panoramic image to the display device.

10. A dynamic surround-view stitching system based on image overlap region feature perception, characterized in that, Dynamic surround view stitching is performed using the method described in any one of claims 1 to 9.