Video stitching and fusing method, system and terminal

Through the video stitching method of timestamp synchronization and multi-dimensional similarity calculation, the time logic chaos and splicing trace problems of video stitching and fusion in the prior art are solved, and efficient and natural video stitching effect is achieved.

CN120343314AActive Publication Date: 2025-07-18GOLDEN TIMES CULTURE COMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510827871.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing technology lacks multi-dimensional considerations when stitching and fusion of videos, resulting in confusion in time logic and obvious splicing traces of videos after stitching, affecting the video quality and viewing experience.

Method used

Through timestamp synchronization, keyframe feature vector extraction, time difference and place similarity calculation, combined with comprehensive similarity judgment, transitional segments are generated for video stitching, optimized homography matrix and brightness adjustment, and eliminated jitter and color differences.

Benefits of technology

It improves the accuracy and efficiency of video splicing, reduces splicing traces, ensures time continuity and logic, adapts to the video splicing needs in different scenarios, and achieves efficient automation and intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343314A_ABST
    Figure CN120343314A_ABST
Patent Text Reader

Abstract

The invention relates to a video stitching and fusion method, system and terminal, and belongs to the technical field of video processing, and the video stitching and fusion method comprises the steps: carrying out the timestamp synchronization of videos received from a master device end and a slave device end; acquiring an overlapping region of videos of the master device end and the slave device end, marking the overlapping region of the master device end as a fragment A, and marking the overlapping region of the slave device end as a fragment B; extracting a feature vector of the key frame A1 of the segment A and a feature vector of the key frame B1 of the segment B; calculating key frame similarity according to the feature vector of the key frame A1 and the feature vector of the key frame B1; obtaining the time difference and the place similarity of the key frame A1 and the key frame B1; calculating comprehensive similarity according to the key frame similarity, the time difference and the place similarity; judging whether the comprehensive similarity is greater than a similarity threshold value or not; and if yes, splicing and fusing the previous frames of the segment B and the segment A. According to the method and the device, the condition that obvious splicing traces appear after video splicing and fusion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and in particular, to a video splicing and fusion method, system, and terminal. Background Art

[0002] In today's video application fields, such as security monitoring, large event live broadcasts, virtual reality, etc., in order to obtain more comprehensive and rich video information, multiple devices are often used to collect videos simultaneously. These devices can be divided into a main device end and a slave device end, each collecting video images from different perspectives. However, effectively splicing and fusing the videos collected by these different devices to form a complete, coherent, and high-quality video image is a challenging task. This not only requires solving the synchronization problems of videos in time and space but also ensuring that the picture transition at the splicing is natural and avoiding obvious splicing marks.

[0003] Currently, when some existing technologies perform video splicing and fusion, they will consider the time synchronization problem of videos and calibrate the video timestamps of different devices through certain methods. There are also some technologies that will try to identify the overlapping areas of videos and perform splicing processing on the overlapping areas. In terms of judging the matching degree of video splicing, some technologies will extract the feature information of videos for similarity calculation. However, most existing technologies only focus on single-dimensional information, such as only paying attention to the feature similarity of videos or only considering the time synchronization factor, lacking comprehensive consideration of multiple important factors. Due to the lack of multi-dimensional comprehensive consideration, various problems are likely to occur when performing video splicing and fusion.

[0004] For example, relying solely on feature similarity for splicing may result in incorrect splicing because videos taken at different times and locations may have similar features, causing confusion in the time logic and space logic of the spliced video. Another example is that if only time synchronization is considered while ignoring the feature matching and location information of video content, it may cause the picture transition at the splicing to be unnatural, with obvious splicing marks, seriously affecting the quality and viewing experience of the spliced video. Summary of the Invention

[0005] In order to improve the situation of obvious splicing marks after video splicing and fusion, this application provides a video splicing and fusion method, system, and terminal.

[0006] In a first aspect, this application provides a video splicing and fusion method, adopting the following technical solution: A video splicing and fusion method includes: Synchronize the timestamps of the videos received from the main device end and the slave device end; Obtain the overlapping area of the videos of the master device end and the slave device end, mark the overlapping area of the master device end as segment A, and mark the overlapping area of the slave device end as segment B; Extract the feature vectors of the key frame A1 of segment A and the feature vectors of the key frame B1 of segment B; Calculate the key frame similarity according to the feature vectors of key frame A1 and the feature vectors of key frame B1; Obtain the time difference and location similarity between key frame A1 and key frame B1; Calculate the comprehensive similarity according to the key frame similarity, the time difference and the location similarity; Judge whether the comprehensive similarity is greater than the similarity threshold; If so, splice and fuse segment B with the previous frame of segment A; If not, generate a transition segment and splice and fuse the transition segment with segment A and segment B.

[0007] By adopting the above technical solution, the time stamps of the videos of the master device end and the slave device end are synchronized, which can ensure the consistency of the two videos in the time dimension. This avoids the problems of picture jumping or time logic confusion in the spliced video caused by time differences, making the spliced video show a continuous and smooth effect in time, and improving the accuracy and logic of the video content. By extracting the feature vectors of the key frames and calculating the similarity, and combining the time difference and location similarity to calculate the comprehensive similarity, the matching degree of the overlapping areas of the master device and slave device videos can be accurately judged. This multi-dimensional matching method greatly improves the accuracy of picture splicing, makes the picture transition natural at the splicing point, and reduces the splicing traces. Only the feature vectors of the key frames are extracted for similarity calculation instead of processing all frames, which greatly reduces the calculation amount and processing time. When dealing with large-scale video data, this can significantly improve the efficiency of splicing and fusion, reduce the demand for computing resources, and make the video splicing and fusion process more efficient and fast. Setting the similarity threshold and making a judgment can quickly decide whether to perform the splicing operation. This simple and effective decision-making mechanism avoids complex calculations and judgment processes, further improves the efficiency of video splicing, and makes the entire splicing and fusion process more automated and intelligent. In addition, this technical solution not only considers the feature similarity of the key frames, but also combines the time difference and location similarity for comprehensive judgment. The multi-dimensional consideration method makes the video splicing and fusion more flexible, can adapt to the video splicing requirements in different scenarios, and improves the situation of obvious splicing traces after video splicing and fusion.

[0008] Optionally, the step of synchronizing the time stamps of the videos received from the master device end and the slave device end specifically includes: The master device's clock source periodically broadcasts a time reference to the slave device, and the slave device's clock performs time calibration according to the deviation compensation formula; Generate a frame with a timestamp every 15 frames from the received videos of the master device and the slave device, and mark it as a time frame to achieve timestamp synchronization.

[0009] By adopting the above technical solution, the master device's clock source periodically broadcasts a time reference to the slave device, and the slave device's clock performs time calibration according to the deviation compensation formula. This operation can effectively eliminate the time deviation between the master and slave devices. Due to factors such as hardware differences and operating environments, the clocks of different devices may have slight errors, which will be amplified during the video stitching and fusion process, resulting in problems such as picture jumps and time logic confusion in the stitched video. By the master device broadcasting the time reference and the slave device performing deviation compensation, the times of the master and slave devices are made highly consistent, thus ensuring the continuity and accuracy of the stitched video in the time dimension. Generate a frame with a timestamp every 15 frames from the received videos of the master device and the slave device, and mark it as a time frame to achieve timestamp synchronization. This method generates time frames at a fixed frame interval, which can avoid the additional computational burden caused by frequent generation of time frames while ensuring the overall time synchronization of the video. Through this regularized time frame generation method, an ordered corresponding relationship in time is formed between the videos of the master and slave devices, further enhancing the accuracy of time synchronization.

[0010] Optionally, after the comprehensive similarity is greater than the similarity threshold and before stitching the overlapping region, the steps include: Obtain the current coordinates of the selected matching points and the transformed coordinates of the selected matching points after being transformed by the homography matrix H; Calculate the stitching reprojection error according to the current coordinates and the transformed coordinates; Determine whether the stitching reprojection error is greater than the error threshold; If so, optimize the homography matrix H.

[0011] By adopting the above technical solution, the current coordinates of the selected matching points and the transformed coordinates after being transformed by the homography matrix H are obtained, and the stitching reprojection error is calculated, which enables the matching degree of the video overlapping area between the master device and the slave device to be accurately measured before stitching. The reprojection error reflects the deviation between the matching points and their actual positions after transformation. By calculating and judging this error, the accuracy of stitching can be intuitively understood. When it is judged that the stitching reprojection error is greater than the error threshold, the homography matrix H is optimized, which is a key step to improve the stitching accuracy. The homography matrix H is used to describe the projective transformation relationship between two image planes, and its accuracy directly affects the stitching effect. Judging the stitching reprojection error and optimizing the matrix before stitching enables the system to better adapt to various complex shooting scenarios. Different shooting environments may cause changes in the characteristics of the matching points, thus affecting the accuracy of the homography matrix H.

[0012] Optionally, the steps after obtaining the overlapping area of the videos of the master device and the slave device include: Taking the A segment as the target stitching area, and judging whether the B segment shakes; If so, obtain the pixel motion vectors of two adjacent frames within the B segment; Determine the homography matrix H according to the pixel motion vectors; Eliminate the shake according to the inverse homography matrix.

[0013] By adopting the above technical solution, taking the A segment marked in the overlapping area of the master device as the target splicing area, and judging whether the B segment on the slave device jitters. This approach can promptly detect unstable factors that may affect the splicing effect. When it is detected that the B segment jitters, obtain the pixel motion vectors of two adjacent frames within the B segment, and determine the homography matrix H based on these pixel motion vectors. The pixel motion vectors reflect the displacement of pixels between adjacent frames. By analyzing these vectors, the characteristics and patterns of the video jitter can be accurately captured. The homography matrix H can describe the projective transformation relationship between two frames of images. Using the pixel motion vectors to determine the homography matrix H can precisely quantify the degree and direction of the video jitter. Eliminate the jitter according to the inverse homography matrix, which is a very effective video stabilization processing method. The inverse homography matrix is opposite to the determined homography matrix H. It can perform an inverse transformation on the jittery video, thereby canceling out the jitter effect of the video. By applying the inverse homography matrix, the jittery B segment can be corrected, enabling it to achieve a smooth transition with the target splicing area A segment during splicing, avoiding unnatural phenomena caused by jitter at the splicing location. Jitter will cause changes in the position and angle of the video. If not processed, it will result in inaccurate matching of the overlapping area during splicing, thereby generating splicing errors. By detecting and eliminating the jitter of the B segment, it can ensure that the pixels in the overlapping area are accurately aligned during splicing, reducing splicing errors.

[0014] Optionally, the steps before eliminating the jitter according to the inverse homography matrix include: Calculate the optical flow error term according to the pixel motion vectors; Obtain the rotation matrix of the current frame according to the rotation matrix of the previous frame; Calculate the rotation smooth term according to the rotation matrix of the previous frame and the rotation matrix of the current frame; Obtain the translational smooth term between the current frame and the previous frame; Calculate the total optimization objective according to the optical flow error term, the rotation smooth term, and the translational smooth term; Judge whether the total optimization objective is less than the optimization threshold; If so, eliminate the jitter according to the inverse homography matrix.

[0015] By adopting the above technical solution, the optical flow error term is calculated according to the pixel motion vector, and the pixel motion inconsistency caused by jitter between video frames can be quantified. The optical flow error term reflects the deviation between the actual pixel motion and the ideal pixel motion between adjacent frames. By calculating this error term, the area and degree of jitter can be accurately located. Calculating the rotation smoothing term and the translation smoothing term separately helps to further analyze the jitter from different dimensions. By calculating the rotation smoothing term through the rotation matrix of the previous frame and the current frame, the jitter degree of the video in the rotation direction can be measured; by obtaining the translation smoothing term of the current frame and the previous frame, the jitter of the video in the translation direction can be evaluated. The total optimization target is calculated according to the optical flow error term, the rotation smoothing term and the translation smoothing term, and multiple factors affecting the jitter are combined for evaluation. This comprehensive consideration of multiple factors avoids the limitations of single factor evaluation and can more accurately reflect the overall jitter condition of the video. By calculating the total optimization target, it can be determined whether the jitter of the current video has reached the level that needs to be processed, which provides a scientific basis for judging whether to perform jitter elimination operations in the future. Judging whether the total optimization target is less than the optimization threshold sets a clear standard for whether to use the inverse homography matrix to eliminate jitter. Only when the total optimization target is less than the optimization threshold is the jitter elimination operation performed, which can avoid unnecessary calculations and processing and improve the efficiency of the system. At the same time, it also ensures that the operation is only performed when it is really necessary to eliminate jitter, ensuring the effectiveness and pertinence of jitter elimination. Through the above series of steps, the jitter is accurately analyzed and effectively judged, and the jitter problem of the video on the slave device side can be handled before splicing. The stable video picture provides a good foundation for the subsequent splicing and fusion with the main device side video, making the picture transition at the splicing point more natural, reducing the splicing marks and picture incoherence caused by jitter.

[0016] Optionally, the video splicing and fusion method further includes: Calculate the RGB mean of the A segment and the B segment; Based on the color statistics of the overlapping area, a linear transformation matrix T is established; According to the original pixel value of the B segment, the linear transformation matrix T, the RGB mean value of the A segment and the RGB mean value of the B segment, the B segment is mapped to the color space of the A segment to obtain a corrected pixel value.

[0017] By adopting the above technical solution, during the video splicing process, due to differences in hardware parameters (such as sensor sensitivity, lens filters, etc.) and shooting environments (such as lighting conditions, angles, etc.) between the master device side and the slave device side, there may be obvious color differences between segment A and segment B. By calculating the RGB means of segment A and segment B, the color characteristics of the two segments can be quantified. Based on the color statistics in the overlapping area, a linear transformation matrix T is established, providing an accurate conversion rule for color mapping. According to this information, segment B is mapped to the color space of segment A, which can effectively eliminate the color differences between the two segments and make the spliced video more harmonious and consistent in color.

[0018] Optionally, the video splicing and fusion method further includes: Obtaining an optimized brightness value according to the target brightness value and a pre-constructed optimization objective function; Calculating the error value between the optimized brightness value and the target brightness value; Judging whether the error value is less than an error threshold; If not, optimizing the parameters of the optimization objective function.

[0019] By adopting the above technical solution, obtaining an optimized brightness value according to the target brightness value and a pre-constructed optimization objective function can adjust for the different brightness situations of the videos on the master device side and the slave device side. Due to differences in shooting environments, hardware parameters, etc. between different devices, the brightness of the videos is often inconsistent, and there will be a problem of abrupt picture brightness during splicing and fusion. By calculating the optimized brightness value and performing error calculation and judgment with the target brightness value, the system can continuously adjust the video brightness to make it approach the target brightness, thereby ensuring the consistency of the video picture brightness after splicing. Judge whether the error value is less than the error threshold, and if the condition is not met, optimize the parameters of the optimization objective function. This mechanism enables the video splicing and fusion system to have the ability of adaptive adjustment. When the brightness error is large, the system will automatically adjust the parameters of the optimization objective function to better adapt to different video brightness situations. When factors such as environmental light changes or device aging cause changes in the video brightness characteristics, the system can continuously optimize the brightness adjustment effect to ensure the high quality of the spliced video. Through brightness optimization and adaptive adjustment, this video splicing and fusion method can be applied to a variety of different application scenarios.

[0020] In a second aspect, the present application provides a video splicing and fusion system, adopting the following technical solution: A video splicing and fusion system, including: A time synchronization module for synchronizing the timestamps of the videos received from the master device side and the slave device side; An overlapping area determination module, configured to obtain the overlapping area of the videos of the master device and the slave device, and mark the overlapping area of the master device as segment A and the overlapping area of the slave device as segment B; A feature vector extraction module, configured to extract the feature vector of the key frame A1 of segment A and the feature vector of the key frame B1 of segment B; A calculation and processing module, configured to calculate the key frame similarity according to the feature vectors of key frame A1 and key frame B1; the calculation and processing module is further configured to obtain the time difference and location similarity between key frame A1 and key frame B1, and calculate the comprehensive similarity according to the key frame similarity, the time difference and the location similarity; A judgment module, configured to judge whether the comprehensive similarity is greater than the similarity threshold; A splicing and fusion module, configured to splice and fuse segment B with the previous frame of segment A when the judgment module judges it to be true.

[0021] In a third aspect, the present application provides a terminal, adopting the following technical solution: A terminal, comprising: A memory, storing a video splicing and fusion program; A processor, configured to execute the program stored on the memory to implement the steps of the above video splicing and fusion method.

[0022] In summary, the present application has at least the following beneficial effects: Synchronizing the timestamps of the videos of the master device and the slave device can ensure the consistency of the two videos in the time dimension, avoid problems such as picture jumps or time logic confusion in the spliced video caused by time differences, make the spliced video present a continuous and smooth effect in time, and improve the accuracy and logic of the video content. By extracting the feature vectors of key frames and calculating the similarity, and combining the time difference and location similarity to calculate the comprehensive similarity, the matching degree of the overlapping areas of the master device and the slave device videos can be accurately judged; the multi-dimensional matching method greatly improves the accuracy of picture splicing, makes the picture transition natural at the splicing point, and reduces the splicing traces. Setting the similarity threshold and making a judgment can quickly determine whether to perform the splicing operation, avoid complex calculation and judgment processes, further improve the efficiency of video splicing, and make the entire splicing and fusion process more automated and intelligent. In addition, this technical solution not only considers the feature similarity of key frames, but also combines the time difference and location similarity for comprehensive judgment. The multi-dimensional consideration method makes video splicing and fusion more flexible, can adapt to the video splicing requirements in different scenarios, and improves the situation where obvious splicing traces appear after video splicing and fusion. Description of the Drawings

[0023] Figure 1 It is a first flow chart of the method embodiment of the present application; Figure 2 is a second flow chart of the method embodiment of the present application; Figure 3 is a third flow chart of the method embodiment of the present application; Figure 4 is a fourth flow chart of the method embodiment of the present application; Figure 5 is a fifth flow chart of the method embodiment of the present application; Figure 6 This is the sixth flow chart of the method embodiment of the present application. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the appended drawings of the embodiments of the present invention. Figure 1 -Attached Figure 6 , the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The first embodiment of the present application discloses a video splicing and fusion method. Figure 1 As an implementation of the video splicing and fusion method, the video splicing and fusion method may include S110-S190: S110, synchronizing the timestamps of the videos received from the master device and the slave device; S120, obtaining the overlapping area of the videos on the master device and the slave device, marking the overlapping area on the master device as segment A, and marking the overlapping area on the slave device as segment B; S130, extracting a feature vector of a key frame A1 of segment A and a feature vector of a key frame B1 of segment B; S140, calculating key frame similarity based on the feature vector of key frame A1 and the feature vector of key frame B1; S150, obtaining the time difference and location similarity between key frame A1 and key frame B1; S160, calculating comprehensive similarity based on key frame similarity, time difference and location similarity; S170, determining whether the comprehensive similarity is greater than a similarity threshold; S180, if yes, then splice and merge the B segment with the previous frame of the A segment; S190. If not, generate a transition segment and splice and fuse the transition segment with segment A and segment B.

[0026] For S110, the steps of synchronizing the timestamps of the videos received from the master device side and the slave device side specifically include: Control the master device side clock source to broadcast the time reference to the slave device side periodically, and the slave device side clock performs time calibration according to the deviation compensation formula; Generate one frame with a timestamp every 15 frames from the received videos of the master device side and the slave device side, and mark it as a time frame to achieve timestamp synchronization.

[0027] Specifically, measure the transmission delay from the master clock of the master device side to each slave device side; then the master clock source can broadcast the time reference through the PTPv2 protocol, and the slave device side clock determines the adjustment deviation according to the deviation compensation formula, and then corrects the time according to the adjustment deviation. The deviation compensation formula is , average delay = ; is the timestamp of the master clock, is the timestamp of the slave clock, is the delay from the master device side to the slave device side, is the delay from the slave device side to the master device side. , is the corrected timestamp of the slave device side. The master device side and the slave device side can be understood as the main camera and the slave camera in this application.

[0028] One implementation manner for S120 - S150 is: Common image processing libraries (such as OpenCV) can be used to read the corresponding frames from the videos of the master device side and the slave device side; then perform the extraction operations of SIFT feature points and descriptors on the read video frames of the master and slave device sides respectively. The descriptor can represent the local feature information of the feature points. Then use FLANN (Fast Library for Approximate Nearest Neighbors) for feature matching, which can efficiently find the matching feature point pairs in the video frames of the master and slave device sides. Based on the matching feature point pairs, use the RANSAC (Random Sample Consensus) algorithm to calculate the homography matrix H. The homography matrix H can describe the projective transformation relationship between two planes, and through it, the overlapping area between the video frames can be obtained. Mark the overlapping area calculated in the video of the master device side as segment A, and mark the overlapping area of the video of the slave device side as segment B.

[0029] Input the pre - processed key frames into the pre - trained ResNet50 model, and extract the feature vectors before its fully - connected layer. The feature vectors can represent the main feature information of the key frames.

[0030] Key frame similarity , is the feature vector of key frame A1, is the feature vector of key frame B1.

[0031] One implementation for S160 - S180 is: Calculate the time difference corresponding to segment A and segment B through the video timestamp , if , is the maximum value of the acceptable time difference, then ; if , then .

[0032] Calculate the location distance d corresponding to segment A and segment B through the location information (such as latitude and longitude) carried by the video. If , is the maximum value of the acceptable location distance, then ; if , then . The comprehensive similarity .

[0033] When the comprehensive similarity is greater than the similarity threshold, to avoid time misalignment in the overlapping area, an image processing library (such as OpenCV) can be used to read the previous frame image of segment A from the main device - side video, denoted as , and then directly replace segment A with segment B and splice and fuse it with .

[0034] One implementation of generating a transition segment and splicing and fusing the transition segment with segment A and segment B can be: The feature vectors of segment A and segment B can be input into the generator of the GAN to generate video frames with transitional effects; then, the SIFT (Scale-Invariant Feature Transform) algorithm is used to extract the feature points and descriptors of the first frame of the transitional segment and the previous frame of segment A, and then FLANN (Fast Library for Approximate Nearest Neighbors) is used for feature matching to find the matching feature point pairs. According to the matching feature point pairs, the homography matrix is calculated using the RANSAC (Random Sample Consensus) algorithm to geometrically align it with segment A, and then the transformed first frame of the transitional segment and the previous frame of segment A are fused using the multi-band fusion algorithm. The multi-band fusion algorithm decomposes the image into sub-images of different frequency bands, fuses them separately and then reconstructs the image, which can effectively reduce the traces at the splicing location. Then, the ORB (Oriented FAST and Rotated BRIEF) algorithm is used to extract the feature points and descriptors of the last frame of the transitional segment and segment B, and then a brute-force matcher is used for feature matching to find the matching feature point pairs. According to the matching feature point pairs, the transformation matrix is calculated using the affine transformation algorithm to perform an affine transformation on the last frame of the transitional segment to geometrically align it with segment B, and a fade-in and fade-out fusion method is used to perform a gradual change process at the splicing location between the last frame of the transitional segment and segment B to make the transition more natural.

[0035] Referring to Figure 2 , after the comprehensive similarity is greater than the similarity threshold and before splicing the overlapping regions, the steps include S210 - S240: S210, obtain the current coordinates of the selected matching points and the transformed coordinates of the selected matching points after being transformed by the homography matrix H; S220, calculate the splicing reprojection error according to the current coordinates and the transformed coordinates; S230, determine whether the splicing reprojection error is greater than the error threshold; S240, if so, optimize the homography matrix H.

[0036] Specifically, one pair or a certain number (for example, 10 - 20) of matching points can be randomly selected from the matching point pairs obtained from the aforementioned feature matching as the selected matching points; a random number generator can be used to achieve random selection. For the selected matching points, record their coordinates in the video frame of the master device, and then use the homography matrix H to transform the current coordinates of the selected matching points.

[0037] Taking a pair of matching points as an example, the splicing reprojection error , is the transformed coordinates of the matching point after transformation, is the current coordinate. If multiple pairs of matching points are selected, the actual stitching reprojection error is the sum of squares or the mean square error of the stitching reprojection errors of multiple pairs of matching points. When the stitching reprojection error is greater than the error threshold, the homography matrix H is optimized.

[0038] Refer to Figure 3 , after obtaining the overlapping area of the main device end and slave device end videos, the subsequent steps include S310 - S340: S310, Take the A segment as the target stitching area and determine whether the B segment jitters; S320, If so, obtain the pixel motion vectors of two adjacent frames within the B segment; S330, Determine the homography matrix H according to the pixel motion vectors; S340, Eliminate the jitter according to the inverse homography matrix.

[0039] Specifically, use the Shi - Tomasi corner detection algorithm to detect feature points in the first frame of the B segment. The Shi - Tomasi corner detection is a commonly used feature point detection method that can detect corner points with obvious features in the image. Use the Lucas - Kanade optical flow algorithm to track the positions of these feature points in adjacent frames. The Lucas - Kanade optical flow algorithm is based on the assumption of invariant image grayness and calculates the optical flow by computing the displacement of feature points between adjacent frames. Calculate the average displacement of feature points between adjacent frames. If the average displacement is greater than a preset threshold (for example, set to 5 pixels according to experience), it is considered that the B segment has jittered.

[0040] Use the Farneback optical flow algorithm to calculate the optical flow between two adjacent frames of the B segment. The Farneback optical flow algorithm is a global optical flow calculation method based on polynomial expansion that can obtain a relatively dense optical flow field. Extract the motion vector of each pixel point from the optical flow field. The motion vector contains the displacement information of the pixel point in the horizontal and vertical directions. Select a sufficient number (generally at least 4 pairs) of matching points from the pixel motion vectors of two adjacent frames. The selection of matching points can be based on the stability and uniform distribution of feature points. Then use the RANSAC (Random Sample Consensus) algorithm combined with the selected matching points to calculate the homography matrix H. The RANSAC algorithm can effectively exclude incorrect matching points and improve the accuracy of the calculation of the homography matrix H. By calculating the inverse matrix of the homography matrix H, perform an inverse transformation on the frames with jitter to eliminate the jitter.

[0041] Refer to Figure 4 , before eliminating the jitter according to the inverse homography matrix, the subsequent steps include S410 - S470: S410, Calculate the optical flow error term according to the pixel motion vectors; S420. Obtain the rotation matrix of the current frame according to the rotation matrix of the previous frame; S430. Calculate the rotation smoothing term according to the rotation matrix of the previous frame and the rotation matrix of the current frame; S440. Obtain the translational smoothing term between the current frame and the previous frame; S450. Calculate the total optimization objective according to the optical flow error term, the rotation smoothing term, and the translational smoothing term; S460. Determine whether the total optimization objective is less than the optimization threshold; S470. If so, eliminate jitter according to the inverse homography matrix.

[0042] Specifically, the optical flow error term is , where V is the pixel motion vector between two adjacent frames, and is the pixel motion vector between two adjacent frames when there is no ideal motion.

[0043] Estimate the rotation matrices of the front and back frames and . Feature point matching and pose estimation algorithms such as the RANSAC algorithm can be used. The rotation smoothing term is used to constrain the rotation change between adjacent frames and avoid drastic rotational jitter. The smoothness of rotation is measured by calculating the difference between the rotation matrices of two adjacent frames; the difference between rotation matrices can be represented by the Frobenius norm of the matrix, and the rotation smoothing term is .

[0044] Convert the rotation matrix to a rotation vector, and then calculate the angle between the rotation vectors of two adjacent frames to measure the smoothness of rotation. Let the rotation vector of the previous frame be , and the rotation vector of the subsequent frame be , the rotation smoothing term .

[0045] Total optimization objective = . If the total optimization objective is less than the optimization threshold, perform an inverse transformation to eliminate jitter.

[0046] Referring to Figure 5 , as another implementation of the video stitching and fusion method, the video stitching and fusion method may include S510 - S530: S510. Calculate the RGB means of segment A and segment B; S520. Based on the color statistics of the overlapping region, establish a linear transformation matrix T; S530. According to the original pixel values of segment B, the linear transformation matrix T, the RGB mean of segment A, and the RGB mean of segment B, map segment B to the color space of segment A to obtain the corrected pixel values.

[0047] Specifically, for segment A and segment B, initialize three variables Rsum, Gsum, and Bsum respectively to accumulate the pixel values of each channel, and at the same time initialize the variable Pixel Count to record the total number of pixels. Use a double loop to traverse each pixel of segment A and segment B. For each pixel, obtain the values of its R, G, and B channels, and accumulate them into the corresponding Rsum, Gsum, and Bsum. Divide the accumulated Rsum, Gsum, and Bsum by the total number of pixels Pixel Count respectively, so as to obtain the RGB means of segment A and segment B. In the overlapping area, statistically analyze the color distributions of segment A and segment B respectively. The color space can be divided into several intervals, and the number of pixels in each interval is counted. For each color channel (R, G, B), find the corresponding relationship between the color distributions of segment A and segment B in the overlapping area. For example, the least squares method can be used to fit a straight line so that the straight line can describe the linear relationship between the two color distributions. According to the slope and intercept of the fitted straight line, construct a linear transformation matrix T. The linear transformation matrix T is a 3×3 diagonal matrix, and the elements on the diagonal correspond to the transformation coefficients of the R, G, and B channels respectively. Use a double loop to traverse each pixel of segment B. For the R, G, and B channel values of each pixel, first subtract the RGB mean of segment B, then multiply by the corresponding coefficient in the linear transformation matrix T, and finally add the RGB mean of segment A to obtain the corrected pixel value. Limit the corrected pixel value within the range of [0, 255] to avoid out-of-range situations. If the pixel value is less than 0, set it to 0; if the pixel value is greater than 255, set it to 255.

[0048] Referring to Figure 6 , as another implementation of the video splicing and fusion method, the video splicing and fusion method may further include S610 - S630: S610, obtain an optimized brightness value according to the target brightness value and the pre-constructed optimization objective function; S620, calculate the error value between the optimized brightness value and the target brightness value; S630, determine whether the error value is less than the error threshold; S640, if not, then optimize the parameters of the optimization objective function.

[0049] Specifically, the optimization objective function is ; assuming the target brightness value is [100, 105, 110], then if , expand the optimization objective function to , and then take the derivative of and set it to zero, and solve to get , , 。When the error value between the optimized brightness value and the target brightness value is greater than or equal to the error threshold, optimization can be performed. 。

[0050] The implementation principle of this embodiment is as follows: Synchronize the timestamps of the videos received from the master device and the slave device, then obtain the overlapping area of the videos of the master device and the slave device, mark the overlapping area of the master device as segment A, and mark the overlapping area of the slave device as segment B; take segment A as the target splicing area, and determine whether segment B jitters. If so, obtain the pixel motion vectors of two adjacent frames within segment B, and determine the homography matrix H according to the pixel motion vectors; calculate the optical flow error term according to the pixel motion vectors, obtain the rotation matrix of the current frame according to the rotation matrix of the previous frame, and calculate the rotation smooth term according to the rotation matrix of the previous frame and the rotation matrix of the current frame; obtain the translational smooth term between the current frame and the previous frame, calculate the total optimization target according to the optical flow error term, the rotation smooth term and the translational smooth term, and determine whether the total optimization target is less than the optimization threshold. If so, eliminate the jitter according to the inverse homography matrix; then extract the feature vectors of the key frame A1 of segment A and the feature vectors of the key frame B1 of segment B, and further calculate the key frame similarity according to the feature vectors of the key frame A1 and the feature vectors of the key frame B1; obtain the time difference and location similarity between the key frame A1 and the key frame B1, and calculate the comprehensive similarity according to the key frame similarity, the time difference and the location similarity. Determine whether the comprehensive similarity is greater than the similarity threshold. If so, obtain the current coordinates of the selected matching points and the transformed coordinates of the selected matching points after being transformed by the homography matrix H, and calculate the stitching reprojection error according to the current coordinates and the transformed coordinates. Determine whether the stitching reprojection error is greater than the error threshold. If so, optimize the homography matrix H until the stitching reprojection error is not greater than the error threshold, and then directly splice and fuse segment B with the previous frame of segment A; if not, generate a transition segment and splice and fuse the transition segment with segment A and segment B.

[0051] Based on the above method embodiments, the second embodiment of the present application discloses a video stitching and fusion system. The video stitching and fusion system of the embodiments of the present application can implement any of the above video stitching and fusion methods, and the specific working processes of each module in the video stitching and fusion system can refer to the corresponding processes in the above method embodiments.

[0052] For the sake of easy understanding, the following is an example: A video stitching and fusion system includes: A time synchronization module for synchronizing the timestamps of the videos received from the master device and the slave device; An overlapping area determination module for obtaining the overlapping area of the videos of the master device and the slave device, and marking the overlapping area of the master device as segment A and the overlapping area of the slave device as segment B; A feature vector extraction module, configured to extract the feature vectors of the key frame A1 of segment A and the feature vector of the key frame B1 of segment B; A calculation and processing module, configured to calculate the key frame similarity according to the feature vectors of the key frame A1 and the key frame B1; the calculation and processing module is further configured to obtain the time difference and location similarity between the key frame A1 and the key frame B1, and calculate the comprehensive similarity according to the key frame similarity, time difference and location similarity; A judgment module, configured to judge whether the comprehensive similarity is greater than the similarity threshold; A splicing and fusion module, configured to splice and fuse segment B with the previous frame of segment A when the judgment module judges yes.

[0053] The third embodiment of the present application provides a terminal. As an implementation manner of the terminal, the terminal may include: a memory and a processor; wherein, The memory is used to store a video splicing and fusion program; The processor is configured to execute the program stored on the memory to implement the steps of the above video splicing and fusion method.

[0054] Wherein, the memory may be communicatively connected to the processor through a communication bus, and the communication bus may be an address bus, a data bus, a control bus, etc.

[0055] In addition, the memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0056] And the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0057] The above are all preferred embodiments of the present application, and do not limit the protection scope of the present application in turn. Any feature disclosed in this specification (including the abstract and drawings), unless specifically stated, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically stated, each feature is only an example of a series of equivalent or similar features.

Claims

1. A video splicing and fusion method, characterized in that, Including: Synchronize the timestamps of the videos received from the master device and the slave device. Obtain the overlapping regions of the videos of the master device and the slave device, mark the overlapping region of the master device as segment A, and mark the overlapping region of the slave device as segment B. Extract the feature vectors of the key frame A1 of segment A and the feature vectors of the key frame B1 of segment B. Calculate the key frame similarity based on the feature vectors of key frame A1 and key frame B1. Obtain the time difference and location similarity between key frame A1 and key frame B1. Calculate the comprehensive similarity based on the key frame similarity, the time difference, and the location similarity. Determine whether the comprehensive similarity is greater than the similarity threshold. If so, splice and fuse segment B with the previous frame of segment A.

2. The video splicing and fusion method according to claim 1, wherein The step of synchronizing the timestamps of the videos received from the master device and the slave device specifically includes: Control the master device clock source to broadcast the time reference to the slave device periodically, and the slave device clock performs time calibration according to the deviation compensation formula. Generate one frame with a timestamp every 15 frames from the received videos of the master device and the slave device, and mark it as a time frame to achieve timestamp synchronization.

3. The video stitching and fusion method according to claim 2, characterized in that, After the comprehensive similarity is greater than the similarity threshold and before splicing the overlapping region, the steps include: Obtain the current coordinates of the selected matching points and the transformed coordinates of the selected matching points after being transformed by the homography matrix H. Calculate the splicing reprojection error based on the current coordinates and the transformed coordinates. Determine whether the splicing reprojection error is greater than the error threshold. If so, optimize the homography matrix H.

4. The video splicing and fusion method according to claim 2, characterized in that, After obtaining the overlapping regions of the videos of the master device and the slave device, the steps include: Take segment A as the target splicing region and determine whether segment B jitters. If so, obtain the pixel motion vectors of two adjacent frames within segment B. Determine the homography matrix H based on the pixel motion vectors. Eliminate the jitter according to the inverse homography matrix.

5. The video splicing and fusion method according to claim 4, wherein, Before eliminating the jitter according to the inverse homography matrix, the steps include: Calculate the optical flow error term based on the pixel motion vectors. Obtain the rotation matrix of the current frame based on the rotation matrix of the previous frame. Calculate the rotation smooth term based on the rotation matrix of the previous frame and the rotation matrix of the current frame. Obtain the translation smooth term between the current frame and the previous frame. Calculate the total optimization objective based on the optical flow error term, the rotation smooth term, and the translation smooth term. Determine whether the total optimization objective is less than the optimization threshold. If so, eliminate the jitter according to the inverse homography matrix.

6. The video splicing and fusion method according to claim 1, characterized in that The video splicing and fusion method further includes: Calculate the RGB means of segment A and segment B. Establish a linear transformation matrix T based on the color statistics of the overlapping region. Map segment B to the color space of segment A according to the original pixel values of segment B, the linear transformation matrix T, the RGB mean of segment A, and the RGB mean of segment B to obtain the corrected pixel values.

7. The video splicing and fusion method according to claim 1, wherein The video splicing and fusion method further includes: Obtain an optimized brightness value according to the target brightness value and a pre-constructed optimization objective function; Calculate the error value between the optimized brightness value and the target brightness value; Determine whether the error value is less than an error threshold; If not, optimize the parameters of the optimization objective function.

8. A video stitching and fusion system, characterized in that, Execute the video stitching and fusion method according to any one of claims 1-7. The video stitching and fusion system includes: A time synchronization module for synchronizing the timestamps of the videos received from the master device end and the slave device end; An overlapping area determination module for obtaining the overlapping areas of the videos of the master device end and the slave device end, marking the overlapping area of the master device end as segment A, and marking the overlapping area of the slave device end as segment B; A feature vector extraction module for extracting the feature vector of the key frame A1 of segment A and the feature vector of the key frame B1 of segment B; A calculation and processing module for calculating the key frame similarity according to the feature vector of key frame A1 and the feature vector of key frame B1; the calculation and processing module is further configured to obtain the time difference and location similarity between key frame A1 and key frame B1, and calculate the comprehensive similarity according to the key frame similarity, the time difference and the location similarity; A judgment module for judging whether the comprehensive similarity is greater than a similarity threshold; A stitching and fusion module for stitching and fusing segment B with the previous frame of segment A when the judgment module determines that it is; 9. A terminal, characterized in that, Including: A memory storing a video stitching and fusion program; A processor for executing the program stored on the memory to implement the steps of the video stitching and fusion method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image registration and multi-resolution fusion-based panoramic image splicing method

    CN108416732A

  • Synchronous key frame extraction video splicing method based on binocular camera

    CN110120012A

  • Video recording method and device

    CN114727147A

  • Video stitching method and device, computer equipment and storage medium

    CN114979758A

  • Video shooting method, device and system

    US20240015393A1