Video stitching method, system, device, medium and product
By tracing filtering, combined feature fusion and resolution enhancement processing of the video images to be stitched, combined with dynamic weighted image fusion and stitching, the problem of excessive feature fusion and low efficiency of high-resolution video processing in video stitching is solved, and efficient and accurate video stitching effect is achieved.
Patent Information
- Application Number
- CN202510546516.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video stitching technologies are prone to excessive fusion when feature fusion, resulting in feature redundancy, affecting matching accuracy, and increasing computational burden when processing high-resolution videos, affecting the real-time and processing efficiency of the system.
By performing guide filtering on the video images to be stitched, noise and interference are removed, and then combined feature fusion processing and resolution enhancement processing are carried out to obtain feature fusion video images and resolution enhancement video images, and finally dynamic weighted image fusion stitching is performed.
The generation of repeated feature points is reduced, the efficiency and accuracy of feature fusion is improved, the computational burden is increased during high-resolution video processing is avoided, and the quality and processing efficiency of video stitching are improved.
Smart Images

Figure CN120075375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a video stitching method, system, device, medium and product. Background Art
[0002] Video stitching is a technology that synthesizes multiple frames of video or multiple camera video sequences into a panoramic image, and is widely used in fields such as security monitoring, virtual reality, and video conferencing. In recent years, video stitching technology has developed rapidly, and deep learning and optimization algorithms have been introduced to improve the stitching quality and processing efficiency. Due to its great potential shown in academic research and practical applications, video stitching has received extensive attention in recent years.
[0003] Currently, when the video stitching technology performs feature fusion, it is prone to over-fusion, resulting in feature redundancy, affecting the matching accuracy, and requiring more computing resources for multi-feature fusion. Especially when processing high-resolution videos, it may lead to an increase in the computational burden, affecting the real-time performance and processing efficiency of the system. Summary of the Invention
[0004] In view of this, in order to solve the above-mentioned technical problems, the present invention provides a video stitching method, system, device, medium and product.
[0005] The first aspect of the present invention provides a video stitching method, including:
[0006] Obtain two video images to be stitched, and perform guided filtering on the two video images to be stitched to obtain two filtered video images;
[0007] Perform joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively to obtain a feature fusion video image and a resolution enhancement video image;
[0008] Perform dynamic weighted image fusion stitching on the feature fusion video image and the resolution enhancement video image to obtain a stitched video image.
[0009] Preferably, this method further includes:
[0010] Perform preprocessing on the two video images to be stitched; wherein, the preprocessing includes one or more of Gaussian filtering, histogram equalization enhancement, and guided filtering for defogging.
[0011] Preferably, performing guided filtering on the two video images to be stitched to obtain two filtered video images includes:
[0012] Take each of the video images to be stitched as an input image, and construct a loss function based on the least squares method with the goal of minimizing the difference between the input image and the output image after guided filtering processing;
[0013] Determine the linear transformation coefficients in the linear transformation function through the loss function; wherein, the linear transformation function is constructed according to the local linear transformation mapping relationship between the guidance image and the output image after the guided filtering processing;
[0014] Update the linear transformation function through the linear transformation coefficients, and use the updated linear transformation function to output the output image after the guided filtering processing corresponding to the input image, and use the output image after the guided filtering processing as the filtered video image.
[0015] Preferably, perform joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively, including: the process of performing joint feature fusion processing on the two filtered video images;
[0016] The process of performing joint feature fusion processing on the two filtered video images includes:
[0017] Extract multiple joint feature operators of the two filtered video images;
[0018] Perform normalization processing on the multiple joint feature operators;
[0019] Perform similar feature aggregation according to the preset reachable neighborhood range of each normalized joint feature operator, and obtain multiple joint features;
[0020] Perform vector fusion on the multiple joint features to obtain a feature fusion video image.
[0021] Preferably, the method further includes:
[0022] Perform PCA dimensionality reduction processing on the feature fusion video image.
[0023] Preferably, perform joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively, including: the process of performing resolution enhancement processing on the two filtered video images;
[0024] The process of performing resolution enhancement processing on the two filtered video images includes:
[0025] Adjust the two filtered video images to the same scale according to a preset ratio under the constraint of keeping the aspect ratio of the image unchanged, and obtain two first video images;
[0026] Project any one of the first video images into the coordinate system of the reference image through a homography matrix to obtain a second video image;
[0027] Perform stitching edge processing on the first video image and the second video image to obtain a stitched edge video image;
[0028] Reset the resolution of the stitched edge video image to a required preset resolution to obtain a third video image;
[0029] Crop the invalid edge area of the third video image and normalize the size of the cropped third video image to obtain the resolution-enhanced video image.
[0030] Preferably, the dynamically weighted image fusion stitching of the feature fusion video image and the resolution-enhanced video image to obtain the stitched video image includes:
[0031] Extract the color space features in the feature fusion video image and the resolution-enhanced video image; the color space features include saturation, contrast, and color richness;
[0032] Determine the weight value of image fusion stitching according to the color space features;
[0033] Perform weighted fusion stitching on the feature fusion video image and the resolution-enhanced video image according to the weight value to obtain the stitched video image.
[0034] In a second aspect, the present invention also provides a video stitching system, including:
[0035] A guided filtering module for obtaining two video images to be stitched and performing guided filtering processing on the two video images to be stitched to obtain two filtered video images;
[0036] An image processing module for respectively performing joint feature fusion processing and resolution enhancement processing on the two filtered video images to obtain a feature fusion video image and a resolution-enhanced video image;
[0037] An image stitching module for performing dynamically weighted image fusion stitching on the feature fusion video image and the resolution-enhanced video image to obtain the stitched video image.
[0038] In a third aspect, the present invention also provides an electronic device, which includes a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor executes the steps of the video stitching method as described in the first aspect.
[0039] Fourthly, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the steps of the video splicing method described in the first aspect are realized.
[0040] Fifthly, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is made to execute the steps of the video splicing method described in the first aspect.
[0041] As can be seen from the above technical solutions, the present invention performs guided filtering on the video images to be spliced to remove noise and interference in the images. Then, joint feature fusion processing and resolution enhancement processing are performed on the filtered video images to obtain a feature fusion video image and a resolution enhanced video image. Thus, by jointly fusing multiple features, the generation of duplicate feature points is reduced, the efficiency and accuracy of feature fusion are improved, the increase in computational burden when processing high-resolution videos is avoided, and dynamic weighted image fusion splicing is performed on these two images, thereby dynamically weighted splicing by combining the high efficiency of low-resolution image splicing and the rich details of high-resolution images, improving the quality and processing efficiency of video splicing. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is an application environment diagram of a video splicing method provided by an embodiment of the present invention;
[0044] Figure 2 It is a flowchart of a video splicing method provided by an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of a guided filtering operation provided by an embodiment of the present invention;
[0046] Figure 4 It is a flowchart of PCA dimensionality reduction provided by an embodiment of the present invention;
[0047] Figure 5 It is a general schematic diagram of a video splicing method provided by an embodiment of the present invention;
[0048] Figure 6Schematic structural diagram of a video splicing system provided by an embodiment of the present invention;
[0049] Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] The video splicing method provided by the embodiments of the present application can be applied to, for example Figure 1 the application environment shown in the figure. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed in the cloud or other network servers. The terminal 101 or the server 102 obtains two video images to be spliced, and performs guided filtering on the two video images to be spliced to obtain two filtered video images; performs joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively to obtain a feature fusion video image and a resolution enhancement video image; performs dynamic weighted image fusion splicing on the feature fusion video image and the resolution enhancement video image to obtain a spliced video image.
[0052] The terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc.
[0053] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0054] As Figure 2 shown, the embodiments of the present application provide a video splicing method. Taking the method applied to Figure 1 the terminal 101 or the server 102 in the figure as an example, it includes the following steps S1 to S3. Among them:
[0055] Step S1: Obtain two video images to be spliced, and perform guided filtering on the two video images to be spliced to obtain two filtered video images.
[0056] Among them, the video images are collected by the camera as the video images to be spliced.
[0057] Since mean filtering has a high degree of blindness, it affects the denoising effect and causes edge blurring in the image. The bilateral filter has a very high computational complexity. Due to the mathematical model, the bilateral filter sometimes undergoes gradient reversal, resulting in image loss.
[0058] In the embodiments of the present application, guided filtering is used to perform guided filtering on two video images to be stitched.
[0059] Among them, guided filtering is an edge-preserving filtering method. By considering the structural information in the image, it can preserve the edge details of the image while removing noise. The guided filtering process is guided by guiding the image. Through guided filtering, two filtered video images can be obtained. These images remove noise while preserving the important edges and structural information in the image, and a guided filter with linear computational complexity is added to improve the computational efficiency. The phenomenon of gradient reversal is avoided, thereby improving the accuracy of image processing.
[0060] Step S2: Perform joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively to obtain a feature fusion video image and a resolution-enhanced video image.
[0061] Among them, in the two filtered video images, a part of the filtered video images are subjected to joint feature fusion processing. After feature point matching, low-resolution matching images of the two images are obtained. At the same time, the distribution range of the matching points in the image is statistically analyzed, and the maximum and minimum values of the horizontal and vertical coordinates of the matching points in the two images are found. The smallest rectangular area containing all the matching points is used as an estimate of the overlapping area. Another part of the filtered video images are subjected to resolution adjustment to obtain high-resolution images. According to the estimated overlapping area, stitching edge processing is performed to obtain a high-resolution matching image of direct stitching of the two images, thereby combining the high efficiency of low-resolution image stitching and the rich details of high-resolution images, and improving the quality and efficiency of video stitching.
[0062] Among them, by performing joint feature fusion processing on the filtered video images, the key features in the image can be extracted and effectively fused, thereby improving the accuracy and stability of stitching. At the same time, by performing resolution enhancement processing on the filtered video images, the resolution of the image can be improved while maintaining the image details, making the stitched video image clearer and more delicate.
[0063] Step S3: Perform dynamic weighted image fusion stitching on the feature fusion video image and the resolution-enhanced video image to obtain the stitched video image.
[0064] Among them, by adopting dynamic weighted image fusion stitching, the weight value can be dynamically adjusted based on multiple factors such as image quality, stitching accuracy, and processing efficiency, so as to obtain the optimal stitching effect. During the stitching process, by comprehensively considering the advantages of the feature fusion video image and the resolution enhancement video image, and using the dynamic weighted strategy for fusion, the quality of the stitched video image can be further improved.
[0065] It should be noted that the present invention performs guided filtering processing on the video images to be stitched to remove noise and interference in the images. Then, the filtered video images are subjected to joint feature fusion processing and resolution enhancement processing to obtain a feature fusion video image and a resolution enhancement video image. Thus, by jointly fusing multiple features, the generation of duplicate feature points is reduced, the efficiency and accuracy of feature fusion are improved, the computational burden during the processing of high-resolution videos is avoided, and by performing dynamic weighted image fusion stitching on these two images, dynamic weighted stitching is carried out by combining the high efficiency of low-resolution image stitching and the rich details of high-resolution images, improving the quality and processing efficiency of video stitching.
[0066] In some embodiments, the method further includes:
[0067] Preprocessing two video images to be stitched; among them, the preprocessing includes one or more of Gaussian filtering, histogram equalization enhancement, and guided filtering for defogging.
[0068] Among them, Gaussian filtering is a linear smoothing filter, suitable for eliminating Gaussian noise and smoothing the image. Histogram equalization enhancement is to adjust the histogram of the image to make its distribution more uniform, thereby enhancing the contrast of the image. Guided filtering for defogging is to use a guiding image to filter the haze image to remove the haze in the image and improve the clarity and visibility of the image. By preprocessing the two video images to be stitched, the effects of subsequent guided filtering processing, joint feature fusion processing, and resolution enhancement processing can be further improved, thereby improving the quality and stability of video stitching.
[0069] In some embodiments, performing guided filtering processing on two video images to be stitched to obtain two filtered video images includes:
[0070] Step S101: Taking each video image to be stitched as an input image, constructing a loss function with the goal of minimizing the difference between the input image and the output image after guided filtering processing;
[0071] Step S102: Determining the linear transformation coefficients in the linear transformation function through the loss function; wherein, the linear transformation function is constructed based on the local linear transformation mapping relationship between the guiding image and the output image after guided filtering processing;
[0072] Step S103, updating the linear transformation function through the linear transformation coefficient, and using the updated linear transformation function to output the output image after the guided filtering processing corresponding to the input image, and using the output image after the guided filtering processing as the filtered video image.
[0073] Among them, guided filtering uses a guide image to guide the filtering process. The guide image can be either the image itself or a different image. Its filtering effect is similar to bilateral filtering. It is used as an edge-preserving smoothing operator and has better performance near the edge.
[0074] In guided filtering, the output value of each pixel is a linear transformation of the pixel values in its neighborhood. The coefficients of this linear transformation are determined by minimizing the difference between the input image and the output image after guided filtering. This difference is usually quantified by the least squares method. The coefficients in the linear transformation function are locally determined, which means that they vary according to different areas of the image, making it possible to remove noise while maintaining the edge and structural information of the image. This feature of guided filtering makes it widely used in the field of image processing, especially in scenes where edge details need to be maintained.
[0075] like Figure 3 As shown, the output is considered through the local guide image, and as the guide image of the guided filtering operation, the filtered output image is set It is a guide image middle Pixel-centered window The local linear transformation in , that is, the linear transformation function is:
[0076]
[0077] In the formula, , Respectively Centered window The linear transformation coefficients of .
[0078] For a square window with a radius (The square window refers to the window of local linear transformation, and the neighborhood radius is usually set to 9). , where Q is the output image, a is the linear transformation coefficient, and I is the guide image. Only in the boot image There is an edge only when there is an edge. In order to determine the size of the linear transformation coefficient, set the constraint condition ,in, is the filtered input image, and the output As input The model and remove the noise or texture components Among them, the texture components can be extracted through the gray-level co-occurrence matrix or the texture energy distribution can be analyzed by converting the image to the frequency domain using the Fourier transform. The texture features reflect the structural features such as the repeatability, directionality, and roughness of the image.
[0079] While maintaining the local linearity of the linear transformation function, seek a solution to minimize the difference between the filtered output and the input At this time, based on the least squares method, a loss function is constructed with the goal of minimizing the difference between the input image and the output image after guided filtering:
[0080]
[0081] In the formula, , , is the corresponding regularization parameter, , and are the number of pixels, mean, and variance of the window respectively.
[0082] When the loss function ensures that the output image is smooth enough, it can fit the structural information of the input image as much as possible. It is used to solve the linear transformation coefficients.
[0083] The final output image is:
[0084]
[0085] In the formula, , , thus establishing a mapping from each pixel point from to , and through this mapping, the input image is processed by guided filtering to obtain the output image, and the output image after guided filtering is used as the filtered video image.
[0086] In some embodiments, joint feature fusion processing and resolution enhancement processing are respectively performed on two filtered video images, including: the process of performing joint feature fusion processing on two filtered video images;
[0087] The process of performing joint feature fusion processing on two filtered video images includes:
[0088] Step S201, extract multiple joint feature operators of the two filtered video images.
[0089] Among them, the joint feature operator is used for the registration of video images. The joint feature operator includes Speed Up Robust Features (SURF), Histogram of Oriented Gradient (HOG) features, optical flow, and color histogram.
[0090] Among them, SURF is an accelerated version of Scale Invariant Feature Transform (SIFT), and it is good at processing images with blur and rotation.
[0091] The Histogram of Oriented Gradient feature is a feature descriptor. Through this feature, objects in computer vision and image processing can be intuitively detected. The formation structure of its features mainly calculates and statistically analyzes the gradient direction histogram of the local area of the image.
[0092] Optical flow describes the motion change of image pixels over time and is particularly suitable for motion detection of targets in videos.
[0093] The color histogram is used to describe the statistical information of the color distribution in an image or video frame. A common approach is to decompose the image into different color channels (such as the Red Green Blue (RGB) channel or Hue, Saturation, Value (HSV)), and then count the frequencies of different color values in each channel. The joint feature operator can comprehensively consider the above information, making it easier to capture the common information between different videos, thereby improving the video stitching effect.
[0094] Step S202: Perform normalization processing on multiple joint feature operators.
[0095] Among them, the magnitudes of different feature vectors are different, and they need to be normalized before fusion. Suppose we have joint feature operators of different dimensions , and its normalization formula is:
[0096]
[0097] Among them, is the normalized joint feature operator, and are the mean and standard deviation of the features respectively. After normalization, each feature value will be concentrated in the same range, which is beneficial for subsequent fusion.
[0098] Step S203: Aggregate similar features according to the preset reachable neighborhood range of each normalized joint feature operator, and obtain multiple joint features.
[0099] Exemplarily, in the embodiments of the present application, at each feature point, the percentage of the short side of the image size is used as the radius, and a circle is drawn with each feature point as the origin to serve as the reachable neighborhood range, marking the influence range of the feature point. Descriptors are extracted for all feature points within the circle, the Euclidean distance or Hamming distance between two of them is calculated, and a threshold is set. If the distance is less than the threshold, they are considered feature-similar. Non-maximum suppression is used to retain the feature point with the smallest descriptor distance and suppress the remaining redundant points.
[0100] By identifying similar or duplicate features within the circle and aggregating them into a main feature. By integrating the features within the circle, such as SURF, HOG, optical flow, and color histograms, a comprehensive combined feature can be formed.
[0101] Among them, different colors or shapes can be used to mark the aggregated main feature to ensure its clear visibility in the image. And based on the initial candidate points of the feature detection algorithm, the central feature point needs to meet high distinctiveness, stability, and uniform spatial distribution. The center point of the circle is marked as the center of the main feature, and all aggregated features are displayed within the circle.
[0102] Step S204: Perform vector fusion on multiple combined features to obtain a feature-fused video image.
[0103] Among them, the SURF, HOG, optical flow, and color histogram features included in the combined feature can be directly fused through vector splicing. Assume that the respective feature vectors are 、 、 and , then the fused feature vector is:
[0104] .
[0105] It can be understood that through vector splicing, feature information of different dimensions is integrated together. The dimension of the spliced feature vector is the sum of the feature dimensions, forming a more comprehensive feature description. This feature fusion method can make full use of the advantages of various feature operators, capture more image information, and thus improve the accuracy and robustness of video splicing.
[0106] Among them, the dimension of the spliced feature vector may be very high, and the computational cost is large. To reduce the computational complexity, it further includes: performing PCA dimensionality reduction processing on the feature-fused video image.
[0107] Among them, PCA dimensionality reduction refers to the dimensionality reduction technology of Principal Component Analysis (PCA). This algorithm uses the covariance matrix to map the dataset samples from a high-dimensional space to a low-dimensional space, maximizing the preservation of the original data information characteristics, effectively solving the curse of dimensionality problem, and at the same time trying to retain the variance information of the data.
[0108] As Figure 4 shown, the process of PCA dimensionality reduction includes:
[0109] (1) Standardization of data samples: Assume that the dataset matrix is a matrix with r columns of data and c sample values in each column . First, perform standardization processing on the matrix, subtract the mean value of each column from each element in the matrix, and generate a new data matrix with zero mean as:
[0110]
[0111] Calculate the covariance matrix of the dataset matrix. Covariance can calculate the variance between the average data and other elements in the row.
[0112] In the dimensionality reduction process, it is expected that the values of the original data after mapping can be maximally dispersed. In mathematical terms, this degree of dispersion can be represented by variance, and the formula is as shown:
[0113]
[0114] Find the first basis vector, which is the direction in which the variance value is the largest after the projection of all data on this vector. After finding the direction with the largest variance after the first projection, then find the second projection direction. If the one with the largest variance after projection is still selected, then it will basically coincide with the first direction. Therefore, it is hoped that these two vectors are independent and orthogonal to each other. Then the covariance can be used to represent the correlation between the two vectors, and the expression is as follows:
[0115]
[0116] When the value is equal to 0, it means that these two fields are completely independent.
[0117] Perform eigenvalue decomposition on matrix C to solve the eigenvalues and the eigenvectors corresponding to the eigenvalues.
[0118] Arrange the solved eigenvalues in descending order, and then sequentially obtain the eigenvectors corresponding to the top k eigenvalues from largest to smallest, and combine these vectors into matrix . is the matrix after dimensionality reduction to k dimensions.
[0119] In the embodiments of the present application, in order to improve the resolution of video stitching, resolution enhancement processing is performed on the feature fusion video image. The resolution enhancement processing can adopt super-resolution reconstruction technology, and by using the prior information of the image and the interpolation algorithm, the low-resolution image is reconstructed into a high-resolution image.
[0120] In some embodiments, joint feature fusion processing and resolution enhancement processing are respectively performed on two filtered video images, including: the process of performing resolution enhancement processing on the two filtered video images;
[0121] The process of performing resolution enhancement processing on the two filtered video images includes:
[0122] Step S211: Under the constraint of keeping the aspect ratio of the image unchanged, the two filtered video images are adjusted to the same scale according to a preset ratio to obtain two first video images.
[0123] Among them, by keeping the aspect ratio of the image, it is adjusted to the same height or width according to the ratio. If strict size uniformity is required, it can be achieved by cropping the redundant area or edge filling (such as black, white or mirror pixels).
[0124] Step S212: Project any one of the first video images into the coordinate system of the reference image through a homography matrix to obtain a second video image.
[0125] The target video image is transformed through a homography matrix to align it with the reference video image in the same coordinate system. The homography matrix is a geometric transformation matrix used to describe the planar projection relationship between two images, and it can map the points in one image to the corresponding points in another image. By aligning the two video images, the parallax between them can be eliminated, providing more accurate image information for subsequent video stitching. Among them, the common transformation types of the homography matrix include affine transformation or perspective transformation.
[0126] Step S213: Perform stitching edge processing on the first video image and the second video image to obtain a stitched edge video image.
[0127] Among them, the stitching edge processing is to eliminate the gaps or artifacts at the stitching edge and improve the naturalness and smoothness of the stitched image. The stitching edge processing can include steps such as edge fusion, fade processing or feathering processing. Edge fusion calculates the pixel value differences on both sides of the stitching edge and performs smooth transition, making the stitching edge less noticeable. Fade processing gradually changes the pixel values at the stitching edge to achieve natural transition. Feathering processing blurs the stitching edge, making the edge area softer and fusing more naturally with the surrounding image.
[0128] Find the optimal stitching path for the overlapping area through graph cut or dynamic programming to avoid moving objects or misaligned parts.
[0129] Step S214: Reset the resolution of the stitched edge video image to the required preset resolution to obtain a third video image.
[0130] Among them, if the size of the stitched image is too small, use an interpolation algorithm (such as bilinear, Lanczos algorithm) to enlarge it to the required resolution.
[0131] Step S215: Crop the invalid edge area of the third video image and normalize the size of the cropped third video image to obtain a resolution-enhanced video image.
[0132] Among them, cropping the invalid edge area is to remove the edge black edges or blurred areas generated due to image enlargement to ensure the neatness and clarity of the video image. For example, remove the invalid edge area generated by the transformation (such as the black filled part) and retain the valid content.
[0133] Size normalization is to adjust the cropped video image to a unified size. For example, force it to be adjusted to a standard size (such as 1920×1080) according to the application scenario, and content-aware scaling may be required.
[0134] In some embodiments, perform dynamic weighted image fusion stitching on the feature fusion video image and the resolution-enhanced video image to obtain the stitched video image, including:
[0135] Step S301: Extract the color space features in the feature fusion video image and the resolution-enhanced video image; the color space features include saturation, contrast, and color richness.
[0136] Among them, saturation is a measure of the vividness of a color, which describes the purity or intensity of the color. High saturation means the color is more vivid and pure, while low saturation makes the color appear darker and less saturated.
[0137] Contrast refers to the brightness difference between the brightest and darkest areas in an image. High contrast can enhance the clarity and layering of the image, making the image look sharper and more vivid. While low contrast may cause the details of the image to be blurred and the overall image to appear rather flat.
[0138] Color richness describes the types and quantities of colors in an image. An image with high color richness contains more color information, can show more details and color variations, and makes the image look more colorful.
[0139] In the embodiments of the present application, saturation can be calculated through color space conversion, and the formula is:
[0140]
[0141] Among them, S is the saturation, and are respectively the maximum and minimum values of the color channels in the image.
[0142] The contrast of the image is represented by the following formula:
[0143]
[0144] Among them, C is the contrast, is the maximum luminance value in the image, is the minimum luminance value in the image.
[0145] The color richness can be measured using the color distribution of the image or the diversity of the histogram:
[0146]
[0147] Among them, R is the color richness, is the probability of each color in the color histogram, and R is normalized.
[0148] Step S302: Determine the weight value for image fusion and stitching according to the color space characteristics;
[0149] During the fusion process of video images in different scenarios, due to the differences in the content of high- and low-resolution stitched images, the weight value for image fusion and stitching needs to be dynamically adjusted. The present invention dynamically adjusts the weight value for image fusion and stitching based on the saturation, contrast, and color richness values of each frame of the image to be stitched in the video, so that the stitched video image is more natural and coordinated in terms of color, brightness, and details. For example, when the saturation of the feature fusion video image is relatively high, its weight in the stitching result can be appropriately increased to maintain the vividness of the color; while when the contrast of the resolution-enhanced video image is higher, its weight can be increased to enhance the clarity and sense of hierarchy of the image.
[0150] Among them, the weight value for image fusion and stitching is calculated as:
[0151]
[0152] In the formula, is the weight value for image fusion and stitching.
[0153] Step S303: Perform weighted fusion and stitching on the feature fusion video image and the resolution-enhanced video image according to the weight value to obtain the stitched video image.
[0154] Specifically, the feature fusion video image and the resolution enhancement video image are weighted according to the calculated weight values, that is, the value of each pixel point is weighted and averaged according to the corresponding weight, so as to obtain the spliced video image. This dynamic weighted image fusion and splicing method can dynamically adjust the weight according to the actual features of the image, making the spliced video image more natural and coordinated in terms of color, brightness and details, and improving the splicing effect. At the same time, this method can also effectively reduce the splicing traces generated due to image splicing, making the spliced video image smoother and more coherent.
[0155] The formula for weighted fusion and splicing is as follows:
[0156] +(1 -
[0157] In the formula, is the pixel point of the spliced video image, is the weight value, is the pixel point of the feature fusion video image, is the pixel point of the resolution enhancement video image.
[0158] In summary, the video splicing method provided by the embodiment of the present application, as Figure 5 shown, through preprocessing two video images to be spliced, inputting the images into the guided filtering module, after filtering the two video images to be spliced by the guided filtering module, where a part of the images are subjected to feature extraction, joint feature and PCA dimensionality reduction, and another part of the images are subjected to resolution enhancement processing, and then the feature fusion video image and the resolution enhancement video image are subjected to steps such as dynamic weighted image fusion and splicing, realizing the precise splicing of multiple video images. This method can be widely applied to fields such as video surveillance, virtual reality, panoramic shooting, etc., and has important practical value.
[0159] Based on the same inventive concept, the embodiment of the present application also provides a video splicing system for implementing the above-mentioned video splicing method.
[0160] The implementation solution provided by this system to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the video splicing system provided below can refer to the limitations on the video splicing method in the above text, and will not be repeated here.
[0161] As Figure 6 shown, the embodiment of the present application provides a video splicing system, including:
[0162] A guided filtering module 100, configured to obtain two video images to be spliced, and perform guided filtering processing on the two video images to be spliced to obtain two filtered video images;
[0163] An image processing module 200, configured to perform joint feature fusion processing and resolution enhancement processing on two filtered video images respectively, to obtain a feature fusion video image and a resolution enhanced video image;
[0164] An image stitching module 300, configured to perform dynamic weighted image fusion stitching on the feature fusion video image and the resolution enhanced video image, to obtain a stitched video image.
[0165] In some embodiments, the system further includes: a preprocessing module, configured to:
[0166] Preprocess two video images to be stitched; wherein, the preprocessing includes one or more of Gaussian filtering, histogram equalization enhancement, and guided filter defogging.
[0167] In some embodiments, a guided filter module 100, configured to:
[0168] Take each video image to be stitched as an input image, and construct a loss function based on the least squares method with the goal of minimizing the difference between the input image and the output image after guided filter processing;
[0169] Determine the linear transformation coefficients in the linear transformation function through the loss function; wherein, the linear transformation function is constructed according to the local linear transformation mapping relationship between the guidance image and the output image after guided filter processing;
[0170] Update the linear transformation function through the linear transformation coefficients, and use the updated linear transformation function to output the output image after guided filter processing corresponding to the input image, and take the output image after guided filter processing as the filtered video image.
[0171] In some embodiments, the image processing module 200, configured to:
[0172] Extract multiple joint feature operators of two filtered video images;
[0173] Perform normalization processing on multiple joint feature operators;
[0174] Perform similar feature aggregation according to the preset reachable neighborhood range of each normalized joint feature operator, and obtain multiple joint features;
[0175] Perform vector fusion on multiple joint features to obtain a feature fusion video image.
[0176] In some embodiments, the system further includes: a dimensionality reduction module, configured to:
[0177] Perform PCA dimensionality reduction processing on the feature fusion video image.
[0178] In some embodiments, the image processing module 200 is configured to:
[0179] Adjust two filtered video images to the same scale according to a preset ratio under the constraint of keeping the aspect ratio of the images unchanged, to obtain two first video images;
[0180] Project any one of the first video images into the coordinate system of the reference image through a homography matrix to obtain a second video image;
[0181] Perform stitching edge processing on the first video image and the second video image to obtain a stitched edge video image;
[0182] Reset the resolution of the stitched edge video image to a required preset resolution to obtain a third video image;
[0183] Crop the invalid edge regions of the third video image, and perform size normalization on the cropped third video image to obtain a resolution-enhanced video image.
[0184] In some embodiments, the image stitching module 300 is configured to:
[0185] Extract the color space features in the feature fusion video image and the resolution-enhanced video image; the color space features include saturation, contrast, and color richness;
[0186] Determine the weight values for image fusion stitching according to the color space features;
[0187] Perform weighted fusion stitching on the feature fusion video image and the resolution-enhanced video image according to the weight values to obtain a stitched video image.
[0188] As Figure 7 shown, an embodiment of the present application provides an electronic device. The electronic device 10 includes a memory 20 and a processor 30. When the computer program stored in the memory 20 is executed by the processor 30, the processor 30 is caused to execute the steps of the video stitching method in the above embodiments.
[0189] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the video stitching method in the above embodiments are implemented.
[0190] An embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the steps of the video stitching method described in the above embodiments.
[0191] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, electronic devices, computer storage media, and computer program products described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0192] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0193] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0194] In several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media, computer program products, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0195] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0196] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0197] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks or optical discs and other various media that can store program codes.
[0198] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A video splicing method, characterized in that: include: Acquire two video images to be stitched, and perform guided filtering on the two video images to be stitched to obtain two filtered video images; Performing joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively to obtain a feature fusion video image and a resolution enhanced video image; Dynamic weighted image fusion and splicing are performed on the feature fused video image and the resolution enhanced video image to obtain a spliced video image.
2. The video stitching method according to claim 1, characterized in that: Also includes: Preprocessing is performed on the two video images to be stitched; wherein the preprocessing includes one or more of Gaussian filtering, histogram equalization enhancement, and guided filtering defogging.
3. The video stitching method according to claim 1, characterized in that: Performing guided filtering processing on the two video images to be spliced to obtain two filtered video images, including: Taking each of the video images to be stitched as an input image, constructing a loss function based on the least squares method with the goal of minimizing the difference between the input image and the output image after guided filtering processing; Determining linear transformation coefficients in a linear transformation function by means of the loss function; wherein the linear transformation function is constructed according to a local linear transformation mapping relationship between a guide image and an output image after the guide filtering process; The linear transformation function is updated by the linear transformation coefficient, and the output image after the guided filtering processing corresponding to the input image is outputted by using the updated linear transformation function, and the output image after the guided filtering processing is used as the filtered video image.
4. The video stitching method according to claim 1, characterized in that: The two filtered video images are respectively subjected to joint feature fusion processing and resolution enhancement processing, including: a process of performing joint feature fusion processing on the two filtered video images; The process of performing joint feature fusion processing on the two filtered video images includes: Extracting a plurality of joint feature operators of the two filtered video images; Normalizing the plurality of joint feature operators; Similar features are aggregated according to the preset reachable neighborhood range of each normalized joint feature operator to obtain multiple joint features; Vector fusion is performed on a plurality of the joint features to obtain a feature fused video image.
5. The video stitching method according to claim 1 or 4, characterized in that: Also includes: The feature fusion video image is subjected to PCA dimensionality reduction processing.
6. The video stitching method according to claim 1, characterized in that: The two filtered video images are respectively subjected to joint feature fusion processing and resolution enhancement processing, including: a process of performing resolution enhancement processing on the two filtered video images; The process of performing resolution enhancement processing on the two filtered video images includes: The two filtered video images are adjusted to the same scale according to a preset ratio under the constraint of keeping the aspect ratio of the images unchanged, so as to obtain two first video images; Projecting any one of the first video images into the coordinate system of the reference image through a homography matrix to obtain a second video image; Performing splicing edge processing on the first video image and the second video image to obtain a splicing edge video image; Resetting the resolution of the spliced edge video image to a desired preset resolution to obtain a third video image; The invalid edge region of the third video image is cropped, and the size of the cropped third video image is normalized to obtain the resolution-enhanced video image.
7. The video stitching method according to claim 1, characterized in that: The step of performing dynamic weighted image fusion and splicing on the feature fused video image and the resolution enhanced video image to obtain a spliced video image includes: Extracting color space features from the feature-fused video image and the resolution-enhanced video image; the color space features include saturation, contrast, and color richness; Determining a weight value for image fusion and splicing according to the color space characteristics; The feature fused video image and the resolution enhanced video image are weightedly fused and spliced according to the weight value to obtain the spliced video image.
8. A video splicing system, characterized in that: include: A guided filtering module is used to obtain two video images to be spliced, and perform guided filtering on the two video images to be spliced to obtain two filtered video images; An image processing module, used for performing joint feature fusion processing and resolution enhancement processing on the two filtered video images respectively, to obtain a feature fusion video image and a resolution enhanced video image; The image stitching module is used to perform dynamic weighted image fusion stitching on the feature fused video image and the resolution enhanced video image to obtain a stitched video image.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the video stitching method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the video stitching method according to any one of claims 1 to 7 are implemented.