Video stitching method and device based on embedded platform
By using feature point matching and homography transformation matrix on the embedded platform, combining the best suture algorithm and in-depth in-depth method, high-quality and real-time video stitching is achieved, solving the problem of insufficient splicing quality and real-time in the existing technology.
Patent Information
- Application Number
- CN202510197894.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing video stitching methods cannot guarantee the real-time and quality requirements of stitching, especially when dealing with large field of view and high-resolution videos, it is difficult to meet the needs of users.
Using the video stitching method based on the embedded platform, the feature points are extracted from the first frame image data of the two video streams, match and purify, and obtain a homography transformation matrix. Then, the best suture algorithm and optimized histogram matching method and in-depth in-depth method are used for splicing and fusion, and the three-frame difference method is used to detect the moving area, and the suture and color correction functions are updated in real time.
High-quality and real-time video stitching is achieved, reducing the processing time of the stitching system, avoiding the problems of ghosting, blurring and splicing seams, and achieving good visual effects.
Smart Images

Figure CN120050373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video information processing, and in particular to a video stitching method and device based on an embedded platform. Background Art
[0002] The most effective way for humans to obtain information is through vision, and about 80% of the information is perceived through vision. Nowadays, multimedia information such as videos and images plays a crucial role in various fields such as surveillance and security, intelligent driving, medical diagnosis, and cultural heritage protection. With the continuous development of technology, people's demand for large field-of-view and high-resolution videos is increasing. However, most video shooting devices on the market have a small field of view, which is not as wide as the human eye's field of view, and it is difficult to meet this demand. To address this challenge, video stitching technology has emerged and provided a solution.
[0003] Video stitching technology is a technology that combines multiple video frames in chronological order into a continuous video. This technology is widely used in fields such as virtual reality, video surveillance, intelligent driving, and digital entertainment. However, most current video stitching methods cannot guarantee the real-time performance and quality requirements of stitching.
[0004] Traditional stitching line algorithms will cause obvious ghosting and stitching seams due to parallax, color, and brightness differences caused by different camera shooting angles. At the same time, the biggest difference between video stitching and image stitching is that there are moving objects in the video. If the stitching line is not updated when the moving object passes through the stitching line, it will cause ghosting and blurring problems. Summary of the Invention
[0005] A video stitching method, device, and storage medium based on an embedded platform proposed by the present invention can solve at least one of the technical problems in the background art.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A video stitching method based on an embedded platform includes the following steps:
[0008] Step 1: Extract feature points from the first-frame image data of two captured video streams, match and refine the feature points, and obtain a homography transformation matrix;
[0009] Step 2: In the first-frame images of the two video streams, select the left image to be stitched as the reference plane, perform perspective transformation on the right image to be stitched based on the homography matrix to align it with the reference plane, and obtain the overlapping area of the first-frame images of the two video streams;
[0010] Step 3: Use the optimal stitching line algorithm based on the dynamic programming method to calculate the optimal stitching line of the overlapping area for the first-frame images of the two video streams.
[0011] Step 4: Based on the found optimal stitching line, for the overlapping region of the first frame, use the optimized histogram matching method and the fade-in and fade-out method. First, determine the range of the fade-in and fade-out fusion region, and then calibrate the video frame images to be stitched based on the color changes of the video frame images within the range, and perform segmented fusion on the collected video frame images to obtain the fused panoramic image;
[0012] Step 5: Use the three-frame difference method to separate the moving region from the original image. When the moving region passes near the stitching line, set the current frame as the updated frame, update the stitching line and the color calibration function, and at the same time, to eliminate the cumulative error, update the homography matrix regularly;
[0013] Step 6: Generate a pixel mapping from the data of the first frame and the updated frame, save the color calibration function, and use the hardware acceleration method of the embedded platform to process the stitching and fusion process of other frames.
[0014] Furthermore, the specific method for registering the images in Step 1 is as follows:
[0015] Step 1-1: Use the ORB algorithm to extract feature points from the video stream first-frame image data and generate corresponding feature descriptors;
[0016] Step 1-2: For the generated feature points, use the nearest neighbor matching method based on the Hamming distance, and perform rough purification on the matched feature point pairs through the neighbor and second-nearest neighbor method, and use the point pairs smaller than the threshold as the matched point pairs;
[0017] Step 1-3: Use the RANSAC algorithm to perform secondary purification on the matched point pairs, and finally obtain the homography matrix of the image.
[0018] Furthermore, the optimal stitching line detection method in Step 3 is as follows:
[0019] Step 3-1: For the optimal stitching line algorithm based on dynamic programming, its goal is to find an optimal stitching line to make the visual effect of the stitching region reach the best. By performing pixel-level analysis on the overlapping region of the image, establish an energy function to measure the matching degree between each pixel. Then, use dynamic programming technology to start from one end of the image and gradually select the pixel points with the smallest energy function value to form a continuous stitching line.
[0020] First, define the energy function E(x,y) of the overlapping region to measure the color difference and geometric difference of the image overlapping region, where x and y are the image coordinates:
[0021] E(x,y) = rE c (x,y) + (1 - r)E g (x,y)
[0022] Among them, E c (x, y) measures the color difference, which is obtained by calculating the intensity difference of the coordinate pixel points. E g (x, y) measures the geometric gradient change difference, which is obtained by obtaining the gradient maps in the horizontal and vertical directions in the overlapping area of the image frame and then calculating their gradient differences. r is the weight between the color and geometric differences. Usually, the gradient change is more sensitive, so r is usually taken as 0.3;
[0023] Use the dynamic programming method to search for the stitching line. The specific method is as follows:
[0024] First, initialize the path weights and path indices of the first row. Then, perform dynamic programming calculations on each row of pixels. After considering boundary processing, calculate the cumulative intensity value of each pixel point. For the selection of the next pixel point, if the next pixel point is within the moving object range, skip this point. If it is not within the moving object range, determine the next pixel point by taking the minimum value method, and gradually update the path weights and path indices. Finally, obtain the optimal stitching line by backtracking the path with the minimum cumulative weight.
[0025] Furthermore, the fusion method for the first frame image of the collected video in step 4 is as follows:
[0026] Step 4-1: The fade-in and fade-out fusion method is a technique commonly used in video stitching and image fusion. The purpose is to achieve a smooth transition in the stitching area, thereby reducing visual discontinuity. By performing weighted averaging on the overlapping area, the fade-in and fade-out fusion method gradually adjusts the transparency of the image to make the stitched image transition naturally visually;
[0027] First, determine the range of the fade-in and fade-out fusion area, which is determined by the average brightness difference of the overlapping area:
[0028]
[0029] Among them, B is the range of the fade-in and fade-out fusion area, B min is the minimum range, which is preset according to the actual situation, B max is the maximum range, which is the shortest distance from the stitching line to the edge of the overlapping area:
[0030] B max = min(C(x, y) - O(x, y) L , O(x, y) R - C(x, y))
[0031] Among them, C(x, y) is the stitching line coordinate, O(x, y) L and O(x, y) R are the left and right boundary coordinates of the overlapping area respectively;
[0032] L diff is the average luminance difference of the image frame overlapping area, L avg is the average luminance, where L diff is calculated as follows:
[0033]
[0034] L avg is calculated as follows:
[0035]
[0036] where overlap is the overlapping area, w and h are the width and height of the overlapping area respectively, L point1 (x, y) is the pixel gray intensity within the overlapping range of the left image to be stitched, L point2 (x, y) is the pixel gray intensity within the overlapping area of the right image to be stitched;
[0037] Step 4-2: Histogram matching, also known as histogram specification, is a technique in image processing used to transform the histogram of an image into a histogram with a specific shape. This technique can be used to enhance the contrast of an image and can selectively enhance the contrast within a certain range of gray values.
[0038] After determining the range of the fade-in and fade-out fusion area, use the optimized histogram matching method to color-correct the video frame images to be stitched based on the color changes of the video frame images to be stitched within the fade-in and fade-out fusion area range;
[0039] Divide the fade-in and fade-out fusion area into several parts according to the set threshold. For each divided area, convert the color channels of the video frame images to be stitched to Lab, and calculate the histograms and normalized cumulative histograms of the two video frame images to be stitched in each color channel respectively. The calculation method of the histogram is as follows:
[0040]
[0041] where h s (i) is the histogram corresponding to each color channel of the reference video frame image, h t (i) is the histogram corresponding to each color channel of the target video frame image, W and H are the width and height of the divided area, I s (x, y) and I t (x, y) are the color channel values of the corresponding coordinates of the reference video frame image and the target video frame image respectively. The δ function is 1 when the color channel value is equal to i, otherwise it is 0;
[0042] The calculation method of the normalized cumulative histogram is as follows:
[0043]
[0044] Among them, C s (i) is the cumulative histogram corresponding to each color channel of the reference video frame image, C t (i) is the cumulative histogram corresponding to each color channel of the target video frame image, W and H are the width and height of the divided area. After the calculation is completed, the preliminary color mapping function M of each color channel is established:
[0045] M(j) = {i|C s (i) ≤ C t (j) ≤ C s (i + 1)}
[0046] Among them, C s (i) ≤ C t (j) ≤ C s (i + 1) represents the mapping condition, j is the color channel value corresponding to the target video frame image, and i is the color channel value corresponding to the reference video frame image;
[0047] After the preliminary color mapping function is established, considering the parallax effect existing in the fade-in / fade-out fusion area of the two images to be stitched, the calculated histograms are sorted in ascending order and the normalized cumulative histograms are recalculated respectively. Then, the numerical indexes less than the set threshold are selected in the normalized cumulative histograms, and their numerical positions in the preliminary color mapping function are inversely deduced and used as the noise areas to be removed. For the discontinuous parts in the processed preliminary color mapping function, smoothing processing is performed to obtain the corrected color mapping function, and the color correction of each color channel of the target video frame image is completed using the corrected color mapping function;
[0048] Repeat the above processing until the color correction of all divided areas is completed;
[0049] Step 4-3: The fusion of the first frame image of the video is based on the best stitching line algorithm. The fade-in / fade-out fusion area is fused according to the existing fade-in / fade-out method, and then the other areas are added to complete the final fusion:
[0050]
[0051] Among them, I 1 is the range from the reference image to the left boundary of the fade-in / fade-out fusion area, B is the range of the fade-in / fade-out fusion area, I 2 is the range from the right boundary of the fade-in / fade-out fusion area to the image to be stitched after projective transformation, r is the weight coefficient, and the calculation formula is as follows:
[0052]
[0053] Among them, x ris the abscissa of the right boundary of the fade-in and fade-out fusion region, x l is the abscissa of the left boundary of the fade-in and fade-out fusion region, x i is the abscissa of the current pixel point.
[0054] Further, the specific method for separating the moving image and detecting the updated frame in step 5 is as follows:
[0055] Step 5-1: Take three consecutive frames F n-1 、F n and F n+1 in the video stream, and calculate two difference images D1 and D2 through the following formula:
[0056] D1 = |F n+1 - F n |
[0057] D2 = |F n - F n-1 |
[0058] where D1 is the difference image between frames F n and F n+1 , D2 is the difference image between frames F n-1 and F n , and the final motion detection result M is obtained from the calculated difference images through the following formula:
[0059] M = D1 ∩ D2
[0060] Separate the detection result M as the motion region from the original image.
[0061] Judge whether there are pixel points in the current motion region that coincide with the pixel positions of the stitching line. If so, use the current frame as the updated frame to update the best stitching line and the color mapping function. Otherwise, continue to use the best stitching line and the color mapping function of the previous frame image.
[0062] Further, the specific implementation method for processing with other frames in step 6 is as follows:
[0063] Step 6-1: Convert the homography matrix obtained by registering the first frame image, the best stitching line obtained by the dynamic programming method, and the weight coefficients obtained by the optimized fade-in and fade-out method into the mapping coordinates corresponding to each pixel of the original image, save the color mapping function obtained by the optimized histogram matching method, and update the mapping coordinates and the old color mapping function in real time after the updated frame calculates the new homography matrix, best stitching line, weight coefficients, and color mapping function. Through the hardware interface function, realize fast coordinate mapping and color correction through hardware acceleration to achieve fast stitching and fusion of video frames.
[0064] On the other hand, the present invention also provides a multi-screen splicing device for display, which includes:
[0065] A main control module that uses RK3588 as the main control of the device and is configured to drive corresponding displays to display specific modules according to user settings; drive multiple splicing displays respectively based on multiple display data streams; ensure that each splicing display presents content based on its corresponding display data stream; process the video stream data to be spliced as a display data stream and output it to multiple splicing displays to display a complete picture corresponding to the image to be displayed;
[0066] A splicing logic processing module that is configured to establish a two-dimensional coordinate system according to user settings, convert the display order of each splicing display into two-dimensional coordinates, confirm the range of each display area after the first segmentation by traversing the two-dimensional coordinates of each splicing display, confirm the range of each display area after the second segmentation according to the current segmented areas and the display ranges of each splicing display after the first segmentation by the system, and output the data streams of each segmented area to multiple splicing displays after the second segmentation by a segmentation chip;
[0067] A display module that is configured to receive the segmented data stream after the second segmentation by the processing module and perform the multi-screen splicing display.
[0068] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0069] As can be seen from the above technical solutions, for the video splicing method and device based on an embedded platform of the present invention, first, video streams with overlapping regions are collected by two cameras, then a homography matrix is obtained by processing the first frame number on the embedded platform, and then the best stitching line algorithm is used to search for the stitching line in the overlapping region, and the optimized histogram matching method and fade-in and fade-out method are used to calibrate the color and splice and fuse the first frame data; the frame difference method is used to detect moving objects. If a moving object passes near the stitching line, the current frame is used as an update frame to update the stitching line and the color calibration function to ensure that there are no ghosting and blurring phenomena for the moving object near the stitching line. At the same time, to eliminate errors, the homography matrix is updated regularly; the data of the first frame and the update frame are used to generate a pixel mapping, and the color calibration function is saved. Other frames are directly subjected to coordinate mapping and color calibration, and the hardware acceleration is used to fuse the process. Finally, the spliced video is displayed in real time on a display device. This method improves the splicing quality while ensuring real-time performance, quickly obtains a high-quality spliced image and displays it in real time on a display device.
[0070] The advantages and positive effects of the present invention are as follows: After obtaining the homography matrix and the optimal stitching line through processing the first frame image, subsequent frames are directly stitched by means of hardware acceleration and the data of the first frame, ensuring the real-time performance of video stitching. During the image stitching process, a method based on frame difference is adopted to calculate the foreground area of moving objects in the video for updating the optimal stitching line. If the optimal stitching line calculated from the previous frame image passes through the foreground area of the current frame, it is necessary to recalculate the optimal stitching line for the current frame; otherwise, the stitching line of the previous frame is continued to be used. This method significantly reduces the processing time of the stitching system. During the image fusion process, an improved fade-in / fade-out algorithm is used to fuse the video frame images based on the optimal stitching line algorithm. The range of the fade-in / fade-out fusion area is confirmed by using the brightness difference between the overlapping area images, accelerating the processing speed. The video frame images to be stitched are divided into blocks based on the range of the fade-in / fade-out fusion area and the method of histogram matching correction is used to balance the colors while ensuring more delicate correction of color differences, adapting to the color changes in the vertical direction of the images. The fade-in / fade-out method is used to fuse the fade-in / fade-out fusion area, and then adding other areas can smoothly stitch the overlapping area of the images. This method effectively avoids the problems of ghosting, blurring, stitching seams and uneven transitions caused by the presence of moving objects and lighting differences, and obtains a seamless fusion and stitched image with good visual effects. Description of the Drawings
[0071] Figure 1 It is a flowchart of the video stitching method based on an embedded platform provided by an embodiment of the present invention;
[0072] Figure 2 It is a flowchart of the first frame image registration stage provided by an embodiment of the present invention;
[0073] Figure 3 It is a flowchart of searching for the stitching line for the first frame of the image to be stitched provided by an embodiment of the present invention;
[0074] Figure 4 It is a flowchart of fusing the first frame image provided by an embodiment of the present invention;
[0075] Figure 5 It is a flowchart of separating the motion area and detecting and updating the frame provided by an embodiment of the present invention;
[0076] Figure 6 It is a flowchart of processing other frames provided by an embodiment of the present invention;
[0077] Figure 7 It is a logic diagram of the display device provided by an embodiment of the present invention;
[0078] Figure 8The frame stitching effect provided by the embodiments of the present invention, where (a) is the left figure, (b) is the right figure, and (c) is the stitching result figure. Detailed implementation manners
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0080] As Figure 1 shown, the video stitching method based on an embedded platform described in this embodiment includes the following steps.
[0081] First, use feature points to register the first frame image, then use the optimal stitching line algorithm, optimized histogram matching method, and optimized fade-in and fade-out method to perform color correction and fusion on the first frame image. Use the three-frame difference method to separate the moving area, and determine whether to update the stitching line and fade-in and fade-out area by real-time monitoring whether the moving area passes through the stitching line. Use the hardware acceleration function of the embedded platform to complete the stitching and fusion of other frames, and finally send the fused video frame to the display device for output.
[0082] The specific steps of the video stitching based on the embedded platform described in the present invention are as follows:
[0083] Step 1: Extract feature points from the first frame image data of the two captured video streams, match and purify the feature points to obtain the homography transformation matrix.
[0084] Step 2: Select the left figure to be stitched as the reference plane, perform perspective transformation on the right figure to be stitched according to the homography matrix to align it with the reference plane, and obtain the overlapping area of the first frame images of the two video streams.
[0085] Step 3: Use the optimal stitching line algorithm based on the dynamic programming method to calculate the optimal stitching line of the overlapping area for the first frame images of the two video streams.
[0086] Step 4: Based on the found optimal stitching line, use the optimized histogram matching method and fade-in and fade-out method for the first frame overlapping area. First, determine the range of the fade-in and fade-out fusion area, then perform color correction on the video frame images to be stitched based on the color change of the video frame images within the range, and perform segmented fusion on the captured video frame images to obtain the fused panoramic image.
[0087] Step 5: Use the three-frame difference method to separate the moving area from the original image. When the moving area passes near the stitching line, set the current frame as the updated frame, update the stitching line and color correction function, and at the same time, to eliminate the cumulative error, regularly update the homography matrix.
[0088] Step 6: Generate a pixel map from the data of the first frame and the updated frame, save the color correction function, and use the hardware acceleration method of the embedded platform to process the stitching and fusion process of other frames.
[0089] As Figure 2 shown, the specific method for registering the first frame image in the present invention is as follows:
[0090] In step 1, feature points are extracted from the first frame image data of the two captured video streams using the ORB algorithm, and the feature points are matched and refined using the Hamming distance and the RANSAC algorithm to obtain a homography transformation matrix;
[0091] The specific method of step 1 is as follows:
[0092] Step 1-1: Use the ORB algorithm to extract feature points from the first frame image data of the video stream and generate corresponding feature descriptors;
[0093] In step 1-2, the nearest neighbor matching method based on the Hamming distance is used for the generated feature points, and the matching feature point pairs are roughly matched by the neighbor and the second nearest neighbor method. The point pairs smaller than the threshold are used as matching point pairs. According to a large number of experimental results, the threshold is generally taken as 0.6;
[0094] In step 1-3, the RANSAC algorithm is used to refine the matching point pairs, and finally the homography matrix of the image is obtained.
[0095] In step 2, the image to be stitched is subjected to perspective transformation to be in the same plane as the reference image. The specific method is as follows:
[0096] Step 2-1: Align the coordinates of the image to be stitched and the coordinates of the reference image according to the homography transformation formula:
[0097]
[0098] Among them, (x r , y r , z r ) is the transformed world coordinate. After the calculation, it is converted into the transformed two-dimensional coordinate (x, y), H is the homography matrix, and (x a , y a , 1) is the coordinate of the image to be stitched converted into the world coordinate;
[0099] In step 2-2, adjust the size of the canvas of the image to be stitched after transformation, and expand the width of the canvas size to the sum of the widths of the two images to facilitate subsequent stitching with the reference image;
[0100] In step 2-3, use the homography matrix to determine the overlapping area of the corresponding frames of the two video streams.
[0101] AsFigure 3 As shown in the figure, the specific method for stitching line search on the first frame of the image to be stitched in step 3 is as follows:
[0102] Step 3-1: The optimal stitching line algorithm based on dynamic programming finds an optimal stitching line on the basis of dynamic programming to optimize the visual effect of the stitching area. This method constructs an energy function to evaluate pixel matching degree by analyzing the pixels in the overlapping area in detail. Then, starting from one end of the image, dynamic programming technology is used to gradually select the pixels with the lowest energy function value to form a continuous stitching line to ensure the best stitching effect.
[0103] First, define the energy function E(x, y) of the overlapping area to measure the color difference and geometric difference of the image overlapping area, where x and y are image coordinates:
[0104] E(x,y) = rE c (x,y) + (1 - r)E g (x,y)
[0105] where, E c (x,y) measures the color difference, which is obtained by calculating the intensity difference of the coordinate pixel points. E g (x,y) measures the geometric gradient change difference, which is obtained by obtaining the gradient maps in the horizontal and vertical directions in the overlapping area of the image frames and then calculating their gradient differences. r is the weight between the color and geometric differences. Usually, the gradient change is more sensitive, so r is usually taken as 0.3;
[0106] Use the dynamic programming method to search for the stitching line. The specific method is as follows:
[0107] First, initialize the path weights and path indices of the first row. Then, perform dynamic programming calculations on each row of pixels. After considering boundary processing, calculate the cumulative intensity value of each pixel point. For the selection of the next pixel point, if the next pixel point is within the range of the moving object, skip this point. If it is not within the range of the moving object, determine the next pixel point by taking the minimum value, and gradually update the path weights and path indices. Finally, obtain the optimal stitching line by backtracking the path with the minimum cumulative weight.
[0108] As Figure 4 shown, the specific method for fusing the first frame image of the collected video in step 4 is as follows:
[0109] Step 4-1: The fade-in and fade-out fusion method is a technique commonly used in video stitching and image fusion. The purpose is to achieve a smooth transition in the stitching area, thereby reducing visual discontinuity. By performing weighted averaging on the overlapping area, the fade-in and fade-out fusion method gradually adjusts the transparency of the image to make the stitched image transition naturally visually;
[0110] The present invention uses an optimized fade-in and fade-out method. First, the range of the fade-in and fade-out fusion area is determined, which is determined by the average brightness difference of the overlapping area:
[0111]
[0112] where B is the calculated range of the fade-in and fade-out fusion area, B min is the minimum range, B max is the maximum range, which is the shortest distance from the suture to the edge of the overlapping area:
[0113] B max =min(C(x,y)-O(x,y) L ,O(x,y) R -C(x,y))
[0114] where C(x,y) is the suture coordinate, O(x,y) L and O(x,y) R are the left and right boundary coordinates of the overlapping area respectively;
[0115] L diff is the average brightness difference of the overlapping area of the image frame, L avg is the maximum brightness difference, where the calculation method of L diff is as follows:
[0116]
[0117] L avg The calculation method is as follows:
[0118]
[0119] where overlap is the overlapping area, w and h are the width and height of the overlapping area respectively, L point1 (x,y) is the pixel gray intensity within the overlapping range of the left image to be stitched, L point2 (x,y) is the pixel gray intensity within the overlapping area of the right image to be stitched;
[0120] Step 4-2: Histogram matching, also known as histogram specification, is a technique in image processing used to transform the histogram of an image into a histogram with a specific shape. This technique can be used to enhance the contrast of an image and can selectively enhance the contrast within a certain gray value range;
[0121] After determining the range of the fade-in and fade-out fusion area, use the optimized histogram matching method to color-correct the video frame images to be stitched based on the color changes of the video frame images to be stitched within the fade-in and fade-out fusion area;
[0122] The fade-in and fade-out fusion region is divided into several parts according to the set threshold. For each divided region, the color channels of the video frame images to be spliced are converted to Lab, and the histograms and normalized cumulative histograms of the two video frame images to be spliced in each color channel are calculated respectively. The calculation method of the histogram is as follows:
[0123]
[0124] where h s (i) is the histogram corresponding to each color channel of the reference video frame image, h t (i) is the histogram corresponding to each color channel of the target video frame image, W and H are the width and height of the divided region, I s (x,y) and I t (x,y) are the color channel values of the corresponding coordinates of the reference video frame image and the target video frame image respectively. The δ function is 1 when the color channel value is equal to i, and 0 otherwise;
[0125] The calculation method of the normalized cumulative histogram is as follows:
[0126]
[0127] where C s (i) is the cumulative histogram corresponding to each color channel of the reference video frame image, C t (i) is the cumulative histogram corresponding to each color channel of the target video frame image, W and H are the width and height of the divided region. After the calculation is completed, the preliminary color mapping function M of each color channel is established:
[0128] M(j) = {i|C s (i) ≤ C t (j) ≤ C s (i + 1)}
[0129] where C s (i) ≤ C t (j) ≤ c s (i + 1) represents the mapping condition, j is the color channel value of the target video frame image, and i is the color channel value of the reference video frame image;
[0130] After the initial color mapping function is established, considering the parallax effect existing in the fade-in / fade-out fusion area of the two images to be stitched, the calculated histograms are sorted in ascending order and the normalized cumulative histograms are recalculated respectively. Then, the numerical indices less than the set threshold are selected from the normalized cumulative histograms, and their numerical positions in the initial color mapping function are deduced inversely and used as the noise areas to be removed. For the discontinuous parts in the processed initial color mapping function, smoothing processing is performed to obtain the corrected color mapping function, and the color correction of each color channel of the target video frame image is completed using the corrected color mapping function;
[0131] Repeat the above processing until the color correction of all divided areas is completed;
[0132] Step 4-3: The fusion of the first frame image of the video is carried out by the optimal stitching line algorithm. The fade-in / fade-out fusion area is fused according to the existing fade-in / fade-out method, and then the final fusion is completed by adding other areas:
[0133]
[0134] Among them, I 1 is the range from the reference image to the left boundary of the fade-in / fade-out fusion area, B is the range of the fade-in / fade-out fusion area, I 2 is the range from the right boundary of the fade-in / fade-out fusion area to the image to be stitched after projective transformation, and r is the weight coefficient. The calculation formula is as follows:
[0135]
[0136] Among them, x r is the abscissa of the right boundary of the fade-in / fade-out fusion area, x k is the abscissa of the left boundary of the fade-in / fade-out fusion area, and x i is the abscissa of the current pixel point.
[0137] As Figure 5 shown, the specific method for separating the moving image and detecting the updated frame in step 5 is as follows:
[0138] Step 5-1: Take three consecutive frames F n-1 , F n and F n+1 in the video stream, and calculate two difference images D1 and D2 through the following formula:
[0139] D1 = |F n+1 - F n |
[0140] D2 = |F n - F n-1 |
[0141] Among them, D1 is the frame Fn and F n+1 The differential image of, D2 is the frame F n-1 and F n The differential image of, and the final motion detection result M is obtained from the calculated differential image through the following formula:
[0142] M = D1 ∩ D2
[0143] Separate the detection result M as the motion area from the original image;
[0144] Judge whether there are pixel points in the current motion area that coincide with the suture pixel positions. If so, use the current frame as the updated frame to update the optimal suture and color mapping function. Otherwise, continue to use the optimal suture and color mapping function of the previous frame image.
[0145] As Figure 6 shown, in step 6, the specific implementation method for processing other frames is as follows:
[0146] Step 6-1: Convert the homography matrix obtained by registering the first frame image, the optimal suture obtained by the dynamic programming method, and the weight coefficients obtained by the optimized fade-in and fade-out method into the mapping coordinates corresponding to each pixel of the original image. Save the color mapping function obtained by the optimized histogram matching method. After updating the frame to calculate the new homography matrix, optimal suture, weight coefficients, and color mapping function, also update the mapping coordinates and the old color mapping function in real time. Through the hardware interface function, realize fast coordinate mapping and color correction through hardware acceleration to achieve fast stitching and fusion of video frames.
[0147] As Figure 7 shown, the present invention also provides a multi-screen display device for verifying the stitching effect, which includes:
[0148] The main control module uses RK3588 as the device main control and is configured to drive the corresponding display to display a specific module according to user settings; drive multiple tiled displays based on multiple display data streams; ensure that each tiled display presents content based on its corresponding display data stream; process the video stream data to be tiled as the display data stream and output it to multiple tiled displays to display a complete picture corresponding to the image to be displayed;
[0149] The stitching logic processing module is configured to establish a two-dimensional coordinate system according to user settings, convert the display order of each tiled display into two-dimensional coordinates, confirm the range of each display area after one-time segmentation by traversing the two-dimensional coordinates of each tiled display, confirm the range of each display area after secondary segmentation according to the current segmentation areas and the display ranges of each tiled display after one-time segmentation by the system, and output the data streams of each segmentation area to multiple tiled displays after secondary segmentation by the segmentation chip;
[0150] A display module, configured to receive the segmented data stream after secondary segmentation by the processing module and perform the multi-screen splicing display.
[0151] Figure 8 For the input test video frames and the splicing effect, where Figure 8 (a) is the left image to be spliced, Figure 8 (b) is the right image to be spliced, Figure 8 (c) is the splicing result. It can be seen from the splicing result that through color calibration processing and dynamically adjusting the range B of the fade-in and fade-out fusion area in the present invention, the fusion range is expanded when the brightness difference is large, and the minimum width is maintained when the difference is small, avoiding excessive smoothing or residual seams. At the same time, the fusion width range limit B min ≤B≤B max ensures that the transition area guarantees the fade-in and fade-out effect without exceeding the overlapping boundary (the dotted box E in the figure), avoiding image content misalignment or distortion caused by excessive expansion. The splicing result shows that in areas with sudden changes in indoor and outdoor lighting and complex textures, this method can stably process various brightness and color differences and output a splicing result without visual discomfort under the constraint of B min ≤B≤B max and output a splicing result without visual discomfort.
[0152] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the above method.
[0153] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.
[0154] In another embodiment provided by the present application, a computer program product including instructions is further provided, which when running on a computer, causes the computer to execute any of the video splicing methods based on an embedded platform in the above embodiments.
[0155] It can be understood that the system, device, and storage medium provided by the embodiments of the present invention correspond to the method provided by the embodiments of the present invention. The explanations, examples, and beneficial effects of the relevant content can refer to the corresponding parts in the above method.
[0156] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0157] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0158] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0159] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video splicing method based on an embedded platform, characterized in that: The specific steps of video stitching based on embedded platform are as follows: Step 1: Extract feature points from the first frame images of the two collected video streams, match and purify the feature points, and obtain the homography transformation matrix; Step 2: In the first frame images of the two video streams, select the left image to be spliced as the reference plane, perform perspective transformation on the right image to be spliced based on the homography matrix and align it with the reference plane, and obtain the overlapping area of the first frame images of the two video streams; Step 3: Calculate the best stitching line of the overlapping area using the best stitching line algorithm based on dynamic programming for the first frame images of the two video streams; Step 4: Based on the best stitching line found, the optimized histogram matching method and fade-in and fade-out method are used for the overlapping area of the first frame. First, the range of the fade-in and fade-out fusion area is determined. Then, based on the color changes of the video frame images within the range, the color of the spliced video frame images is corrected, and the collected video frame images are segmented and fused to obtain the fused first frame image. Step 5: Use the three-frame difference method to separate the moving area from the original image. When the moving area passes near the stitching line, set the current frame as the update frame, update the stitching line and color correction function, and update the homography matrix regularly to eliminate the accumulated error. Step 6: Generate pixel maps from the data of the first frame and the updated frame, save the color correction function, and use the hardware acceleration method of the embedded platform to process the splicing and fusion process of other frames.
2. The video stitching method based on the embedded platform according to claim 1, characterized in that: The method for calculating the fade-in and fade-out blending area described in step 4 is as follows: The extent of the fade-in and fade-out blending area is determined by the average brightness difference in the overlapping areas: Where B is the range of the gradual in and out fusion area, B min is the minimum range, which is preset according to the actual situation. max is the maximum range, and is the shortest distance from the seam line to the edge of the overlapping area: B max =min(C(x,y)-O(x,y) L ,O(x,y) R -C(x,y)) Where C(x,y) is the coordinate of the suture line, O(x,y) L and O(x,y) R They are the left and right boundary coordinates of the overlapping area respectively; L diff is the average brightness difference in the overlapping area of the image frames, L avg is the average brightness, where L diff The calculation method is as follows: L avg The calculation method is as follows: Where overlap is the overlapping area, w and h are the width and height of the overlapping area respectively, L point1 (x, y) is the grayscale intensity of the pixel in the overlapping range of the left image to be spliced, L point2 (x, y) is the grayscale intensity of the pixel in the overlapping area of the right image to be stitched.
3. The video stitching method based on the embedded platform according to claim 1, characterized in that: The method for color correction of the first frame image described in step 4 is as follows: After determining the range of the fade-in and fade-out fusion area, the optimized histogram matching method is used to perform color correction on the video frame image to be spliced based on the color change of the video frame image to be spliced within the range of the fade-in and fade-out fusion area; The fade-in and fade-out fusion area is divided into several parts according to the set threshold. For each divided area, the color channel of the video frame image to be stitched is converted to Lab, and the histogram and normalized cumulative histogram of the two video frame images to be stitched in each color channel are calculated respectively. The histogram calculation method is as follows: where h s (i) is the histogram of the reference video frame image in each color channel, h t (i) is the histogram of the target video frame image in each color channel, W and H are the width and height of the divided area, I s (x,y) and I t (x, y) are the color channel values of the corresponding coordinates of the reference video frame image and the target video frame image, respectively. The delta function is 1 when the color channel value is equal to i, otherwise it is 0; The normalized cumulative histogram is calculated as follows: Among them C s (i) is the cumulative histogram corresponding to each color channel of the reference video frame image, C t (i) is the cumulative histogram corresponding to each color channel of the target video frame image, W and H are the width and height of the divided area, and after the calculation is completed, the preliminary color mapping function M of each color channel is established: M(j)={i|C s (i)≤C t (j)≤C s (i+1)} Among them C s (i)≤C t (j)≤C s (i+1) represents the mapping condition, j is the value of each color channel corresponding to the target video frame image, and i is the value of each color channel corresponding to the reference video frame image; After the preliminary color mapping function is established, the calculated histograms are sorted in ascending order and the normalized cumulative histogram is recalculated, taking into account the parallax effect of the two images to be spliced in the fade-in and fade-out fusion area. The numerical indexes less than the set threshold are selected in the normalized cumulative histograms, and their numerical positions in the preliminary color mapping function are reversed and removed as noise areas. The non-continuous parts in the processed preliminary color mapping function are smoothed to obtain the corrected color mapping function, and the corrected color mapping function is used to complete the color correction of each color channel of the target video frame image. The above process is repeated until the color correction of all divided areas is completed.
4. The video stitching method based on the embedded platform according to claim 1, characterized in that: The method for fusing the first frame of the video described in step 4 is as follows: The first frame of the video is fused based on the best stitching algorithm. The fade-in and fade-out fusion area is fused according to the existing fade-in and fade-out method, and the other areas are added to complete the final fusion: Among them, I1 is the range from the reference image to the left edge of the fade-in and fade-out fusion area, B is the range of the fade-in and fade-out fusion area, I2 is the range from the right edge of the fade-in and fade-out fusion area to the image to be spliced after projection transformation, r is the weight coefficient, and the calculation formula is as follows: Among them, x r is the horizontal coordinate of the right edge of the fading-in and fading-out fusion area, x l is the horizontal coordinate of the left edge of the fade-in and fade-out blending area, x i The horizontal coordinate of the current pixel.
5. The video stitching method based on the embedded platform according to claim 1, characterized in that: The specific method for registering the first frame image in step 1 is as follows: Step 1-1: Use the ORB algorithm to extract feature points from the first frame of the video stream and generate corresponding feature descriptors; Step 1-2: The generated feature points are matched using the nearest neighbor matching method based on the Hamming distance, and the matching feature point pairs are roughly purified using the neighbor and next nearest neighbor method, and the point pairs that are less than the threshold are taken as matching point pairs; Step 1-3: Use the RANSAC algorithm to perform secondary purification on the matching point pairs and finally obtain the homography matrix of the image.
6. The video stitching method based on the embedded platform according to claim 1, characterized in that: The best suture detection method described in step 3 is as follows: Define the energy function E(x,y) of the overlapping area to measure the color difference and geometric difference of the overlapping area of the image, where x and y are the image coordinates: E(x,y)=rE c (x,y)+(1-r)E g (x,y) Among them, E c (x, y) measures the color difference, which is obtained by calculating the intensity difference of the coordinate pixel point, E g (x, y) measures the difference in geometric gradient changes. It obtains the horizontal and vertical gradient maps in the overlapping area of the image frame, and then calculates their gradient differences. r is the weight between color and geometric differences.
7. The video stitching method based on the embedded platform according to claim 1 is characterized in that: Step 3 uses dynamic programming to search for suture lines. The specific method is: First, the path weight and path index of the first row are initialized. Then, dynamic programming calculation is performed on each row of pixels. After considering the boundary processing, the cumulative intensity value of each pixel is calculated. For the selection of the next pixel, if the next pixel is within the range of the moving object, it is skipped. If it is not within the range of the moving object, the next pixel is determined by taking the minimum value, and the path weight and path index are gradually updated. Finally, the optimal stitching line is obtained by backtracing the path with the minimum cumulative weight.
8. The video stitching method based on the embedded platform according to claim 1, characterized in that: The specific method of separating the moving image and detecting the update frame in step 5 is as follows: Step 5-1: Take three consecutive frames F from the video stream n-1 、F n and F n+1 , two differential images D1 and D2 are calculated by the following formula: D1=|F n+1 -F n | D2=|F n -F n-1 | Among them, F1 is frame D n and F n+1 The difference image of frame F n-1 and F n The calculated differential image is used to obtain the final motion detection result M through the following formula: M=D1∩D2 The detection result M is separated from the original image as the motion area. Determine whether there are pixels in the current motion area that coincide with the stitching line pixel position. If so, use the current frame as the update frame to update the best stitching line and color mapping function. Otherwise, continue to use the best stitching line and color mapping function of the previous frame image.
9. The video stitching method based on the embedded platform according to claim 1, characterized in that: The specific implementation method of using hardware acceleration to process other frames in step 6 is as follows: The homography matrix obtained by registration of the first frame image, the optimal stitching line obtained by dynamic programming and the weight coefficient obtained by optimized fade-in and fade-out method are converted into mapping coordinates corresponding to each pixel of the original image, and the color mapping function obtained by optimizing the histogram matching method is saved. After the new homography matrix, optimal stitching line, weight coefficient and color mapping function are calculated for the update frame, the mapping coordinates and the old color mapping function are also updated in real time. Through the hardware interface function, fast coordinate mapping and color correction are realized through hardware acceleration to achieve fast splicing and fusion of video frames.
10. A multi-screen splicing device for display, used to implement the video splicing method based on an embedded platform according to any one of claims 1 to 9, characterized in that: include, The main control module uses RK3588 as the device main control and is configured to drive the corresponding display to display a specific module according to user settings; Based on multiple display data streams, drive multiple spliced displays respectively; ensure that each spliced display presents content based on its corresponding display data stream; process the video stream data to be spliced as a display data stream and output it to multiple spliced displays to display a complete picture corresponding to the image to be displayed; The splicing logic processing module is configured to establish a two-dimensional coordinate system according to user settings, convert the display order of each splicing display into two-dimensional coordinates, confirm the range of each display area to be divided once by traversing the two-dimensional coordinates of each splicing display, confirm the range of each display area to be divided twice according to the current divided areas and the display range of each splicing display after the system performs the first division, and output the data stream of each divided area to multiple splicing displays after the second division by the division chip; The display module is configured to receive the segmented data stream after secondary segmentation by the processing module and perform the multi-screen splicing display.
Citation Information
Patent Citations
Multichannel real-time video splicing processing system
CN103856727A
Image splicing method based on optimal suture line self-selection area gradual-in gradual-out algorithm
CN112365518A
Array camera color correction method based on optimized histogram matching
CN113079275A
Video fusion algorithm based on dynamic optimal suture line and improved fade-in and fade-out method
CN113221665A
Image splicing method, device and equipment and readable storage medium
CN115880154A