Multi-lens picture automatic splicing fusion processing method in video communication

Through the video communication system with synchronous acquisition of multiple cameras and unified time synchronization, the feature recognition and terminal adaptation problems of traditional systems in complex backgrounds are solved, high-precision stitching and natural interaction of panoramic images are achieved, and the user experience is improved.

CN120614533APending Publication Date: 2025-09-09BEIJING ZHONGJI XINTUO TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510775594.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional video communication systems have difficulty accurately identifying key points in complex backgrounds, lack the ability to adapt to terminal devices, and have a single user interaction method, which cannot meet the needs of immersive experience and all-round visual perception.

Method used

It uses multiple cameras to synchronously capture video images, identify overlapping areas and key feature points, adopt unified time synchronization and standardized configuration, combine feature extraction algorithms and image matching, generate panoramic images, and support dynamic adaptation and natural interaction.

Benefits of technology

It achieves high-precision feature matching in complex scenarios, eliminates splicing traces, improves user interaction freedom and terminal adaptability, and ensures the continuity of the immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614533A_ABST
    Figure CN120614533A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-lens picture automatic splicing fusion processing method in video communication, which belongs to the technical field of multimedia communication, and comprises an acquisition module for synchronously acquiring video images of different visual angles through a plurality of cameras, and an analysis module for identifying overlapping areas and key feature points in each path of video, carrying out image matching, and sending the images to the video communication module. The fusion module is used for splicing and fusing the multiple paths of videos according to the image matching result to generate a panoramic picture, and the output module is used for outputting the spliced and fused panoramic picture to terminal equipment in real time. According to the invention, through a frame-level synchronization system with a unified time reference and standardized video parameter configuration, time sequence deviation and chromatic aberration distortion of multiple paths of videos are eliminated, the feature matching reliability in a complex scene is remarkably improved, the positioning precision of an overlapping region is ensured, splicing traces can be efficiently eliminated, and meanwhile, processing delay is reduced; and intelligent adaptation of the panoramic content on different terminal devices can be realized, and the degree of freedom of user control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimedia communication, and in particular relates to a method for automatically splicing and fusing multiple lens images in video communication. Background Art

[0002] With the rapid development of multimedia communication technology, the demand for multi-view video acquisition and display in application scenarios such as video conferencing, distance learning, and virtual live broadcast is increasing.

[0003] Traditional methods mostly use simple algorithms such as edge detection or template matching, which make it difficult to accurately identify key points in complex backgrounds. In addition, existing video communication systems generally lack the ability to adapt to terminal devices and cannot automatically adjust the picture ratio and clarity according to different screen sizes and resolutions. At the same time, user interaction methods are relatively simple and lack functions such as intelligent recommendation, voice focus, and gesture operation, which limits the user's viewing freedom and sense of participation, and cannot fully present scene information. As a result, it is difficult to meet users' needs for immersive experience and all-round visual perception. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for automatically splicing and fusing multi-lens images in video communication to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: comprising the following steps:

[0006] S1, acquisition module: synchronously acquire video images from different perspectives through multiple cameras;

[0007] S2, analysis module: identifies overlapping areas and key feature points in each video channel and performs image matching;

[0008] S3, fusion module: splices and fuses multiple videos based on image matching results to generate a panoramic image;

[0009] S4, output module: output the stitched and fused panoramic image to the terminal device in real time;

[0010] As a further preferred embodiment of the present technical solution: the S1 acquisition module: reasonably arranges multiple cameras in the target scene to ensure that there is a certain overlap of the field of view between the cameras to support subsequent stitching processing, and the cameras respectively cover different spatial angles and ranges to form a panoramic field of view;

[0011] As a further preferred embodiment of the present technical solution: in S1, all cameras are connected to a unified time synchronization mechanism to ensure that each video stream has accurate timestamp information when it is captured, achieve frame-level synchronization, unify video parameter configuration, and standardize settings for all cameras, which helps to reduce color difference and distortion problems in subsequent image processing. The underlying interface is called to simultaneously enable the video capture function of each camera, read each frame of image data in real time, and cache the captured image frames in timestamp order into an independent data buffer;

[0012] As a further preferred embodiment of the present technical solution: the S2 analysis module performs standardized preprocessing on the video frames, including grayscale processing, image enhancement and denoising operations to improve the accuracy of subsequent feature extraction, and uses the feature extraction algorithm SIFT to detect key points in each frame image and extract local invariant features. The formula is:

[0013]

[0014] In the formula, G(x,y,σ i ) represents the position (x, y) and scale σ i Gaussian filter response under i is a weighting coefficient used to enhance the stability of image edge and corner features. The detected key points generate corresponding feature descriptors to express the local image information around the key points. Between the images captured by adjacent cameras, the feature matching algorithm FLANN is used to compare the key points of the two frames to find the same feature points and establish a preliminary correspondence.

[0015] As a further preferred embodiment of the present technical solution: S2 uses the RANSAC algorithm to optimize the matching results, removes mismatched points caused by environmental interference and repeated textures, calculates the geometric transformation matrix between the images based on the successfully matched feature point pairs, uses the transformation matrix to perform a projective transformation on the images, estimates the boundary of the overlapping area between the two images, and forms an overlapping area mask map as an important reference for subsequent splicing and fusion;

[0016] As a further preferred embodiment of the present technical solution: the S3 fusion module obtains the feature matching results, overlapping area range and calculated geometric transformation matrix between each video channel from the analysis module, projects each video frame into a unified virtual canvas coordinate system according to the transformation matrix, and sequentially maps each projected image into a corresponding position on the panoramic canvas to form a multi-layer overlay structure;

[0017] As a further preferred embodiment of the present technical solution: in S3, for overlapping image areas, an image fusion algorithm is used to eliminate stitching gaps and color differences, the motion state of the current image content is analyzed in real time, and the optimal fusion method is intelligently selected. After stitching and fusion, the resulting panoramic image is encapsulated into a standard image format and transmitted to an output module for display;

[0018] As a further preferred embodiment of the present technical solution: the S4 output module receives the panoramic picture data stream, obtains a continuous panoramic image frame sequence from the fusion module, encodes and compresses the panoramic image frame sequence according to the universal video coding standard H.264, generates a transmittable standard video stream, and dynamically scales and crops the panoramic picture according to the screen resolution and aspect ratio of the target terminal device;

[0019] As a further preferred embodiment of the present technical solution: S4 supports multiple display mode options. Users can operate the panoramic image through touch and mouse. Combined with user behavior analysis and voice command recognition, it automatically recommends jumping to the current most concerned image area. If a certain image is missing or splicing fails, the system can add a prompt mark to the corresponding area, record key status information during the output process, and collect display effect evaluation through the user feedback channel for subsequent system optimization and algorithm upgrade;

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. In the present invention, a frame-level synchronization system with a unified time base and standardized video parameter configuration are used to eliminate the timing deviation and chromatic aberration distortion of multi-channel videos, laying the foundation for subsequent processing. A feature extraction algorithm with scale and rotation invariance is adopted, combined with an intelligent mismatch elimination mechanism, which significantly improves the reliability of feature matching in complex scenarios, ensures the positioning accuracy of overlapping areas, and can effectively eliminate splicing traces while reducing processing delays.

[0022] 2. In the present invention, dynamic image cropping and zooming technology is used to achieve intelligent adaptation of panoramic content on different terminal devices, and natural interactive operations such as gesture zooming and voice focusing are supported, thereby improving user control freedom. In addition, the image integrity and splicing status can be monitored in real time, and visual identification will be automatically provided in case of abnormalities to ensure a consistent experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The operating process of the method for automatically splicing and fusing multiple shots in video communication of the present invention is as follows Figure 1 ;

[0024] Figure 2 The operating process of the method for automatically splicing and fusing multiple shots in video communication of the present invention is as follows Figure 2 ;

[0025] Figure 3 The operating process of the method for automatically splicing and fusing multiple shots in video communication of the present invention is as follows Figure 3 . DETAILED DESCRIPTION

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0027] Example

[0028] See also Figure 1-Figure 3 As shown, the present invention provides a technical solution: comprising the following steps:

[0029] S1, acquisition module: synchronously acquire video images from different perspectives through multiple cameras;

[0030] S2, analysis module: identifies overlapping areas and key feature points in each video channel and performs image matching;

[0031] S3, fusion module: splices and fuses multiple videos based on image matching results to generate a panoramic image;

[0032] S4, output module: output the stitched and fused panoramic image to the terminal device in real time;

[0033] In this embodiment, specifically: the S1 acquisition module: reasonably arranges multiple cameras in the target scene to ensure that there is a certain overlap in the field of view between the cameras to support subsequent stitching processing. The cameras cover different spatial angles and ranges to form a panoramic field of view;

[0034] In this embodiment, specifically: in S1, all cameras are connected to a unified time synchronization mechanism to ensure that each video stream has accurate timestamp information when it is captured, achieving frame-level synchronization, calling the underlying interface and simultaneously enabling the video capture function of each camera, reading each frame of image data in real time, and caching the captured image frames in timestamp order into an independent data buffer;

[0035] In this embodiment, specifically: the S2 analysis module performs standardized preprocessing on the video frames, uses the feature extraction algorithm SIFT to detect key points in each frame image, and extracts local invariant features. The formula is:

[0036]

[0037] In the formula, G(x,y,σ i ) represents the position (x, y) and scale σ i Gaussian filter response under i is a weighting coefficient used to enhance the stability of image edge and corner features. The detected key points generate corresponding feature descriptors to express the local image information around the key points. Between the images captured by adjacent cameras, the feature matching algorithm FLANN is used to compare the key points of the two frames to find the same feature points and establish a preliminary correspondence.

[0038] In this embodiment, specifically: S2 uses the RANSAC algorithm to optimize the matching results, removes false matching points caused by environmental interference and repeated textures, calculates the geometric transformation matrix between the images based on the successfully matched feature point pairs, uses the transformation matrix to perform a projective transformation on the images, estimates the boundary of the overlapping area between the two images, and forms an overlapping area mask;

[0039] In this embodiment, specifically: the S3 fusion module obtains the feature matching results, overlapping area range, and calculated geometric transformation matrix between each video channel from the analysis module, projects each video frame into a unified virtual canvas coordinate system according to the transformation matrix, and sequentially maps each projected image to a corresponding position on the panoramic canvas to form a multi-layer overlay structure;

[0040] In this embodiment, specifically: S3, for overlapping image areas, uses an image fusion algorithm to eliminate stitching gaps and color differences, analyzes the motion state of the current image content in real time, intelligently selects the optimal fusion method, and after stitching and fusion, encapsulates the final panoramic image into a standard image format and transmits it to the output module for display;

[0041] In this embodiment, specifically: the S4 output module receives a panoramic image data stream, encodes and compresses the panoramic image frame sequence according to the universal video coding standard H.264 to generate a transmittable standard video stream, and dynamically scales and crops the panoramic image according to the screen resolution and aspect ratio of the target terminal device;

[0042] In this embodiment, specifically: S4 supports multiple display mode options. Users operate the panoramic picture through touch and mouse. Combined with user behavior analysis and voice command recognition, it automatically recommends jumping to the current most concerned picture area. If it is detected that a certain picture is missing or splicing fails, the system can add a prompt mark in the corresponding area, record key status information during the output process, and collect display effect evaluation through the user feedback channel.

[0043] Working principle or structural principle: multiple cameras are reasonably deployed in the target scene to ensure that there is a certain overlapping area of ​​vision between the cameras for subsequent image matching and stitching. All cameras are connected to a unified time synchronization mechanism to ensure that each video stream has accurate timestamp information when it is collected, so as to achieve strict synchronization at the frame level. At the same time, the cameras are configured in a standardized manner, including resolution, frame rate, color format, and image parameters such as white balance and exposure, to reduce color difference and distortion problems in subsequent image processing. The video capture function of each camera is enabled by calling the underlying interface, and each frame of image data is read in real time. It is cached in an independent data buffer in the order of timestamps, and its acquisition time, frame number and camera ID are recorded for synchronization analysis by the subsequent processing module. Subsequently, the system extracts the corresponding moment images from each video stream according to the timestamp to form a group of "synchronous frame groups" as input data for the subsequent analysis module.

[0044] After the analysis module receives the synchronized frame group from the acquisition module, it first performs standardized preprocessing on each video frame, including grayscale, histogram equalization, contrast enhancement and denoising operations, to improve image quality and lay the foundation for feature extraction. Subsequently, the SIFT algorithm is used to detect key points in the image. Key points are usually distributed at edges, corners or texture-rich areas and have strong repeatability. For each key point, a corresponding feature descriptor is generated to express the local image information around the point. Then, the FLANN algorithm is used to match key points between the images captured by adjacent cameras to find the same or similar feature point pairs and establish a preliminary correspondence. Then, the system further uses the RANSAC algorithm to eliminate mismatched points and retain reliable feature point pairs. The geometric transformation matrix between the images is calculated based on the successfully matched feature point pairs, and the matrix is ​​used to project one image into the coordinate system of another image, estimate the boundary of the overlapping area between the two images, and generate an overlapping area mask map as an important reference for subsequent splicing and fusion.

[0045] The fusion module receives the feature matching results, overlapping area range, and geometric transformation matrix from the analysis module and begins to perform image projection transformation and stitching operations. First, the video frames are projected into a unified virtual canvas coordinate system according to the transformation matrix, so that the images of different perspectives are spatially aligned. The images after projection transformation are sequentially mapped to the corresponding positions of the panoramic canvas to form a multi-layer overlay structure. For overlapping image areas, multiple image fusion algorithms are used to eliminate stitching gaps and color differences. At the same time, the system analyzes the motion state of the current image content in real time and dynamically selects the optimal fusion strategy. When the image is stable and rich in details, multi-band fusion is enabled to obtain high-quality effects. When the image moves violently, the weighted averaging method is used to reduce the impact of delay and artifacts.

[0046] The output module receives the panoramic image frame sequence from the fusion module and compresses and encodes it according to the universal video coding standard to generate a standard video stream suitable for network transmission. Subsequently, the system dynamically scales and crops the panoramic image according to the screen resolution and aspect ratio of the target terminal device to ensure its compatibility and display quality on various devices. The output module also supports multiple display mode options, including overall panoramic image display, split-screen viewing of the original camera image, and focus display of the area of ​​interest. Users can zoom, pan, and focus the image through touch, mouse, or remote control, improving viewing freedom and interactive experience. The system combines user behavior analysis and voice command recognition to automatically recommend jumping to the current most concerned area of ​​the image, such as the person speaking or the object being displayed. In the event of missing or splicing failure of a certain image channel, the system can add a prompt mark to the corresponding area and automatically update the display content after recovery to ensure the integrity of the image. Finally, the display effect evaluation is collected through the user feedback channel to provide data support for subsequent algorithm optimization and model upgrades.

[0047] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatically splicing and fusing multiple shots in video communication, characterized by: The following steps are involved: S1, acquisition module: synchronously acquire video images from different perspectives through multiple cameras; S2, analysis module: identifies overlapping areas and key feature points in each video channel and performs image matching; S3, fusion module: splices and fuses multiple videos based on image matching results to generate a panoramic image; S4, output module: output the stitched and fused panoramic image to the terminal device in real time.

2. The method for automatically splicing and fusing multiple shots in video communication according to claim 1, characterized in that: The S1 acquisition module: rationally arranges multiple cameras in the target scene to ensure that there is a certain overlap in the field of view between the cameras to support subsequent stitching processing. The cameras cover different spatial angles and ranges to form a panoramic field of view.

3. The method for automatically splicing and fusing multiple shots in video communication according to claim 2, characterized in that: In S1, all cameras are connected to a unified time synchronization mechanism to ensure that each video stream has accurate timestamp information when it is collected, achieve frame-level synchronization, call the underlying interface and simultaneously start the video capture function of each camera, read each frame of image data in real time, and cache the collected image frames in timestamp order into an independent data buffer.

4. The method for automatically splicing and fusing multiple shots in video communication according to claim 3, characterized in that: The S2 analysis module performs standardized preprocessing on the video frames, uses the feature extraction algorithm SIFT to detect key points in each frame, and extracts local invariant features. The formula is: In the formula, G(x,y,σ i ) represents the position (x, y) and scale σ i Gaussian filter response under i It is a weighting coefficient used to enhance the stability of image edge and corner features. The detected key points generate corresponding feature descriptors to express the local image information around the key points. Between the images captured by adjacent cameras, the feature matching algorithm FLANN is used to compare the key points of the two frames of images, find the same feature points, and establish a preliminary correspondence.

5. The method for automatically splicing and fusing multiple shots in video communication according to claim 4, characterized in that: The S2 uses the RANSAC algorithm to optimize the matching results, removes false matching points caused by environmental interference and repeated textures, calculates the geometric transformation matrix between the images based on the successfully matched feature point pairs, uses the transformation matrix to perform projective transformation on the images, estimates the boundary of the overlapping area between the two images, and forms an overlapping area mask map.

6. The method for automatically splicing and fusing multiple shots in video communication according to claim 5, characterized in that: The S3 fusion module obtains the feature matching results, overlapping area range and calculated geometric transformation matrix between each video from the analysis module, projects each video frame into a unified virtual canvas coordinate system according to the transformation matrix, and maps each projected image to the corresponding position of the panoramic canvas in turn to form a multi-layer overlay structure.

7. The method for automatically splicing and fusing multiple shots in video communication according to claim 6, characterized in that: The S3 uses an image fusion algorithm to eliminate stitching gaps and color differences in overlapping image areas, analyzes the motion state of the current image content in real time, and intelligently selects the optimal fusion method. After stitching and fusion, the final panoramic image is encapsulated into a standard image format and transmitted to the output module for display.

8. The method for automatically splicing and fusing multiple shots in video communication according to claim 7, characterized in that: The S4 output module receives the panoramic picture data stream, encodes and compresses the panoramic image frame sequence according to the universal video coding standard H.264, generates a transmittable standard video stream, and dynamically scales and crops the panoramic picture according to the screen resolution and aspect ratio of the target terminal device.

9. The method for automatically splicing and fusing multiple shots in video communication according to claim 8, characterized in that: The S4 supports multiple display mode options. Users can operate the panoramic screen through touch and mouse. Combined with user behavior analysis and voice command recognition, it automatically recommends jumping to the area of ​​the screen that is currently of greatest concern. If it detects that a certain image is missing or splicing fails, the system can add a prompt mark to the corresponding area, record key status information during the output process, and collect display effect evaluations through user feedback channels.

Citation Information

Cited By

  • Ultrahigh-definition video production system and method based on 5G and VR fusion

    CN122053937A

  • Ultra-high-definition video production system and method based on 5g and vr fusion

    CN122053937B