A real-time video stitching method based on optical flow calculation
By adopting multi-stage optical flow calculation method in VR technology, the problem of difficult to achieve high-quality and low-latency real-time video splicing in the prior art is solved, and a more refined and realistic splicing effect is achieved.
Patent Information
- Application Number
- CN201980099652.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-08-22
AI Technical Summary
In VR technology, existing video stitching methods are difficult to achieve high-quality and low-latency real-time video stitching, affecting users' immersive experience.
The real-time video stitching method based on optical flow calculation is adopted, and the target optical flow information is obtained through multi-level optical flow calculation, and the overlapping areas of the original video are spliced based on this information.
It achieves shorter computing time and higher computing efficiency, and the resulting spliced video picture is more refined, the splicing effect is more realistic and the delay is lower.
Smart Images

Figure CN114287126B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing, and in particular, to a real-time video stitching method based on optical flow calculation. Background Art
[0002] VR (Virtual Reality) technology can bring panoramic simulation images and immersive experiences to users and has been applied in many fields such as games, movies, medicine, and design. Among them, ODSV video is an important factor for VR to enable users to have a sense of immersion. When the ODSV video is used in combination with an HMD (Head Mounted Display), the display can track the head movement of the user and present video content corresponding to the user's field of view.
[0003] In order to provide users with a more comfortable and realistic immersive experience, higher quality is required in video correction (such as distortion calibration, brightness consistency adjustment), stitching, stereoscopic information reconstruction, and subsequent encoding, and more details in the original video need to be retained during the processing. Also, in order to improve the real-time performance of the picture and enhance the interactivity of VR, the end-to-end system from the camera lens to the HMD display screen also needs to have lower latency. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a real-time video stitching method based on optical flow calculation to achieve the above effects.
[0005] To achieve the above purpose, a first aspect of the present disclosure provides a video processing method, including: obtaining a plurality of original videos, where the plurality of original videos are videos obtained by a plurality of dynamic image recorders arranged at preset positions;
[0006] Determining an overlapping area of each adjacent original video according to a preset rule corresponding to the preset position, where the adjacent original videos are the original videos obtained by the adjacent dynamic image recorders arranged; performing multi-level optical flow calculation on the original videos within each overlapping area to obtain a plurality of target optical flow information; and stitching the overlapping areas of each adjacent original video based on the target optical flow information to obtain a target video.
[0007] Optionally, after obtaining the multiple original videos, the method further includes: performing time synchronization on the multiple original videos; calibrating the synchronized multiple original videos; determining the overlapping regions of each adjacent original video according to the preset rule corresponding to the preset position, including: determining the overlapping regions of each adjacent calibrated original video according to the preset rule corresponding to the preset position; splicing the overlapping regions of each adjacent original video based on the target optical flow information, including: splicing the overlapping regions of each adjacent calibrated original video based on the target optical flow information.
[0008] Optionally, performing multi-level optical flow calculation on the original videos within each overlapping region to obtain multiple target optical flow information includes: performing the following operations on multiple target frame images of two original videos within the overlapping region: respectively downsampling the corresponding frame images of the two original videos to obtain two image sequences arranged from low to high in resolution; determining the first initial optical flow information of the first image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the image sequence; based on the first initial optical flow information, determining the optical flow information of the first target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and using the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence as the first initial optical flow information, where the first target image is an image other than the first image in the first image sequence; repeating the step of determining the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence based on the first initial optical flow information in the order from low to high of the resolution of the first target image until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and using the final optical flow information as the target optical flow information corresponding to the target frame image.
[0009] Optionally, the method further includes: determining candidate optical flow information using the gradient descent method based on each of the initial optical flow information; determining the optical flow information of the target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence based on the initial optical flow information includes: determining the optical flow information of the target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence based on the initial optical flow information and / or the candidate optical flow information.
[0010] Optionally, before downsampling the target frame images of the two original videos respectively, the method further includes: determining the sobel features in the target frame images; and determining an error function for optical flow calculation based on the sobel features.
[0011] Optionally, the method further includes: after obtaining the target optical flow information of the target frame image, determining the type of the adjacent frame image according to the relevance between the target frame image and the adjacent frame image; when the adjacent frame image is a P-frame image, determining the optical flow information of a preset image in the first image sequence of the adjacent frame image relative to the image with the corresponding resolution in the second image sequence based on the target optical flow information of the target frame image, and using the optical flow information of the preset image relative to the image with the corresponding resolution in the second image sequence as the second initial optical flow information; repeating, in the order from low to high of the resolution of the second target image, the step of determining the optical flow information of the second target image relative to the image with the corresponding resolution in the second image sequence based on the second initial optical flow information until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and using the final optical flow information as the target optical flow information corresponding to the adjacent frame image, where the second target image is an image in the first image sequence of the adjacent frame image with a resolution higher than that of the preset image.
[0012] Optionally, the step of stitching the overlapping regions of each adjacent original video based on the target optical flow information to obtain a target video includes: stitching the overlapping regions according to the distance between the dynamic image recorders, the pupil distance information, and the target optical flow information to obtain two target videos respectively for the left eye and the right eye.
[0013] Optionally, the method further includes: partitioning the background blocks and object blocks in the target video based on the target optical flow information; and performing ODSV video encoding on the target video according to the partitioning result.
[0014] In a second aspect of the present disclosure, there is provided a video processing apparatus, including: a video acquisition module configured to acquire a plurality of original videos, where the plurality of original videos are videos acquired by a plurality of dynamic image recorders arranged at preset positions within the same time period; an overlapping module configured to determine an overlapping region of each adjacent original video according to a preset rule corresponding to the preset position, where the adjacent original videos are the original videos acquired by the dynamically arranged adjacent dynamic image recorders; an optical flow calculation module configured to perform multi-level optical flow calculation on the original videos within each overlapping region to obtain a plurality of target optical flow information; and a stitching module configured to stitch the overlapping regions of each adjacent original video based on the target optical flow information to obtain a target video.
[0015] Optionally, the device further includes a time synchronization module and a calibration module. The time synchronization module is configured to perform time synchronization on the multiple original videos; the calibration module is configured to calibrate the synchronized multiple original videos; the overlapping module is configured to determine an overlapping area of each adjacent original video after calibration according to a preset rule corresponding to the preset position; the splicing module is configured to splice the overlapping areas of each adjacent original video after calibration based on the target optical flow information.
[0016] Optionally, the optical flow calculation module includes multiple optical flow calculation sub-modules. Each optical flow calculation sub-module is configured to perform the following operations on multiple target frame images of two original videos within the overlapping area: perform downsampling on corresponding frame images of the two original videos respectively to obtain two image sequences arranged from low to high in resolution; determine first initial optical flow information of a first image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the image sequence; based on the first initial optical flow information, determine optical flow information of a first target image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, and use the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence as the first initial optical flow information, where the first target image is an image in the first image sequence other than the first image; repeat the step of determining the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence based on the first initial optical flow information in the order from low to high of the resolution of the first target image until determining final optical flow information of an image with the same resolution as the original video in the first image sequence relative to an image with a corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the target frame image.
[0017] Optionally, the optical flow calculation module further includes a candidate calculation sub-module configured to determine candidate optical flow information using the gradient descent method based on each of the initial optical flow information; the optical flow calculation sub-module is further configured to determine optical flow information of a target image in the first image sequence relative to an image with a corresponding resolution in the second image sequence based on the initial optical flow information and / or the candidate optical flow information.
[0018] Optionally, the optical flow calculation module further includes a function determination sub-module configured to determine sobel features in the target frame image; based on the sobel features, determine an error function for optical flow calculation.
[0019] Optionally, the optical flow calculation module further includes a frame type processing sub-module, configured to determine the type of the adjacent frame based on the correlation between the target frame and the adjacent frame after obtaining the target optical flow information of the target frame; when the adjacent frame is a P-frame, based on the target optical flow information of the target frame, determine the optical flow information of a preset image in the first image sequence of the adjacent frame with respect to an image with a corresponding resolution in the second image sequence, and use the optical flow information of the preset image with respect to the image with the corresponding resolution in the second image sequence as the second initial optical flow information; repeat the step of determining the optical flow information of a second target image with respect to an image with a corresponding resolution in the second image sequence based on the second initial optical flow information in ascending order of the resolution of the second target image until determining the final optical flow information of an image with the same resolution as the original video in the first image sequence with respect to an image with the corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the adjacent frame, where the second target image is an image in the first image sequence of the adjacent frame with a resolution higher than that of the preset image.
[0020] Optionally, the stitching module is further configured to stitch the overlapping area according to the distance of the dynamic video recorder, the pupil distance information, and the target optical flow information to obtain two target videos for the left eye and the right eye respectively.
[0021] Optionally, the apparatus further includes an encoding module, configured to partition the background block and the object block in the target video based on the target optical flow information; and perform ODSV video encoding on the target video according to the partitioning result.
[0022] In a third aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0023] In a fourth aspect of the present disclosure, there is provided an electronic device, including a memory and a processor, where a computer program is stored on the memory, and the processor is configured to execute the computer program in the memory to implement the steps of the method according to any one of the first aspect of the present disclosure.
[0024] Through the above technical solutions, the optical flow of the overlapping area of the acquired original video is calculated by means of multi-level optical flow calculation, and the overlapping area of the original video is stitched based on the obtained target optical flow information, so that a stitched video with finer pictures, more realistic stitching effects, and lower latency can be obtained with shorter calculation time and higher calculation efficiency.
[0025] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. Brief Description of the Drawings
[0026] The drawings are used to provide a further understanding of the present disclosure and form a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:
[0027] Figure 1 is a flowchart of a video processing method shown in an exemplary disclosed embodiment.
[0028] Figure 2 is a schematic diagram of the arrangement of a dynamic video recorder shown in an exemplary disclosed embodiment.
[0029] Figure 3 is a schematic diagram of an overlapping area shown in an exemplary disclosed embodiment.
[0030] Figure 4 is an example diagram of an overlapping area of an ODSV shooting system shown in an exemplary disclosed embodiment.
[0031] Figure 5 is another example diagram of an overlapping area of an ODSV shooting system shown in an exemplary disclosed embodiment.
[0032] Figure 6 is a block diagram of a video processing device shown in an exemplary disclosed embodiment.
[0033] Figure 7 is a block diagram of an electronic device shown in an exemplary disclosed embodiment.
[0034] Figure 8 is a block diagram of an electronic device shown in an exemplary disclosed embodiment. Detailed Description of the Embodiments
[0035] The following details the specific embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and understanding the present disclosure, and are not used to limit the present disclosure.
[0036] Figure 1 is a flowchart of a video processing method shown in an exemplary disclosed embodiment. As Figure 1 shown, the method includes the following steps.
[0037] S11. Obtain a plurality of original videos, where the plurality of original videos are videos obtained by a plurality of dynamic video recorders arranged at preset positions.
[0038] The dynamic image recorder can be a camera, a camera-equipped camera, a mobile terminal with a camera function (such as a mobile phone, a tablet computer, etc.). In order to achieve better image effects, in the following embodiments of this application, the dynamic image recorder is taken as an example of a camera for illustration.
[0039] In order to achieve a panoramic photography effect, multiple cameras (4, 6, 8, etc.) can be arranged around an axis or an axis. A smaller number of cameras may cause serious distortion around the lens or a photographic blind area, while an excessive number of cameras may lead to a larger processing volume and a slower processing speed. Preferably, as Figure 2 shown is a schematic diagram of a possible arrangement of dynamic image recorders, with 6 cameras arranged equidistantly around an axis.
[0040] It should be noted that the original video described in this disclosure can be a video file that has been shot and needs to be processed, or it can be a real-time transmitted image. For the former, this disclosure can process it into a panoramic video while retaining more details of the original video and maintaining a higher image quality. For the latter, this disclosure can, on the basis of retaining more details and maintaining a higher image quality, through the high-efficiency video processing method provided by this disclosure, process it into a panoramic image with low-latency real-time output.
[0041] Optionally, after obtaining multiple original videos, the multiple original videos can also be time-synchronized, calibrated, and the calibrated original videos can be used as the original videos for subsequent use in this disclosure.
[0042] Use the camera parameters to correct the original video obtained by this camera, where the camera parameters include the horizontal FOV and vertical FOV, focal length, and radial distortion coefficient. Imaging distortion is an optical aberration that represents the magnification of an image within the field of view at a fixed imaging distance, and the radial distortion coefficient is used to describe the relationship between the theoretical and actual image heights. If a wide-angle or fish-eye lens is used during imaging, additional high barrel distortion will be introduced, and the image magnification decreases as the distance from the image increases. Therefore, the objects around the imaging boundary are greatly compressed and need to be stretched.
[0043] S12. Determine the overlapping area of each adjacent original video according to the preset rule corresponding to the preset position.
[0044] Among them, the adjacent original videos are the original videos obtained by the adjacent-arranged dynamic image recorders.
[0045] Different arrangement positions correspond to different rules. For example, as Figure 3As shown in the figure, the area enclosed by the lines extending from both ends of the camera is the image that the camera can capture. Among them, the area enclosed by the included angle formed by the dotted lines represents the area that can be captured by two cameras simultaneously, and this area is the overlapping area described in the present disclosure. The video images within an overlapping area will be captured by two adjacent cameras.
[0046] S13. Perform multi-level optical flow calculation on the original videos within each of the overlapping areas to obtain multiple target optical flow information.
[0047] Optical flow is used to represent the pixel movement direction between two images. When calculating the displacement of each pixel, dense optical flow provides a finer-grained and more intuitive representation of the correspondence between images. For ODSV videos, the optical flow information between each frame of the videos captured by adjacent cameras can provide a reference for the splicing position and splicing method of this frame, and can also represent the information of the parallax generated by different placement positions of the cameras.
[0048] The existing fastest optical flow algorithm, FlowNet2-s, takes about 80 ms to generate a stitched frame image with a size of 4096×2048 pixels, and the maximum transmission frame rate it can reach is 12 fps. For an optical flow-based high-definition ODSV system, to achieve a low-latency real-time effect presentation requires at least 25 fps, and this method far fails to meet such requirements. In addition, the optical flow accuracy of this method is much lower than the multi-level optical flow calculation method proposed in the present disclosure.
[0049] Specifically, the multi-level optical flow algorithm means that for the same frame at the same moment of two original videos within an overlapping area, they are respectively downsampled to form two image sequences with resolutions from low to high. Starting from the images with lower resolutions, calculate the optical flow information between two images pairwise, and use the optical flow information calculated from the images with lower resolutions as a reference quantity to be added to the optical flow calculation of the next level until the optical flow information of the images with higher resolutions is calculated. Because there is not much available pixel information in the initially low-resolution images, the calculation efficiency is very high. And through multiple layers of refined calculations, it provides available references for the calculations of subsequent high-resolution images. Therefore, the processing time required for the calculations of high-resolution images is shorter compared with traditional optical flow calculation methods. Therefore, the multi-level optical flow algorithm can improve the calculation speed and calculation accuracy of optical flow.
[0050] In a possible implementation manner, the following method can be used to perform multi-level optical flow calculation on two original videos within an overlapping area:
[0051] For each frame of the two original videos within the overlapping area, perform the following operations:
[0052] Downsample the corresponding frame images of two original videos respectively to obtain two image sequences arranged from low to high in resolution; for example, if the resolution of the original video is 1125, downsample a certain frame of it to generate an image sequence with resolutions of 10, 18, 36, 90, 276, and 1125 respectively, where the resolution of the last layer of images is the same as that of the original video.
[0053] Determine the first initial optical flow information of the first image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the image sequence;
[0054] Based on the first initial optical flow information, determine the optical flow information of the first target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and use the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence as the first initial optical flow information. The first target image is an image in the first image sequence other than the first image;
[0055] In the order from low to high of the resolution of the first target image, repeat the step of determining the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence based on the first initial optical flow information until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the target frame image.
[0056] That is to say, perform multi-level optical flow calculation on each frame of the original video, so that accurate target optical flow information can be obtained.
[0057] In another possible embodiment, the target optical flow information of the original video in the overlapping area can also be calculated by the following method:
[0058] After obtaining the target optical flow information of the target frame image, determine the type of the adjacent frame image according to the correlation between the target frame image and the adjacent frame image; among them, the types of adjacent frame images are divided into I frames and P frames. A P frame is an image with a correlation with an I frame and has similar frame features to the I frame, and can be obtained by data change based on the I frame, while the correlation between an I frame and its previous frame is relatively low, and it cannot be obtained by simple data change based on its previous frame.
[0059] When the adjacent frame image is an I frame image, downsample it and perform multi-level optical flow calculation step by step using the method in the previous embodiment.
[0060] When the adjacent frame is a P-frame, based on the target optical flow information of the target frame, determine the optical flow information of a preset image in the first image sequence of the adjacent frame relative to an image with a corresponding resolution in the second image sequence, and use the optical flow information of the preset image relative to the image with the corresponding resolution in the second image sequence as the second initial optical flow information;
[0061] Repeat, in the order from the lowest to the highest resolution of the second target image, the step of determining the optical flow information of the second target image relative to an image with a corresponding resolution in the second image sequence based on the second initial optical flow information, until determining the final optical flow information of an image with the same resolution as the original video in the first image sequence relative to an image with the corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the adjacent frame, where the second target image is an image in the first image sequence of the adjacent frame with a resolution higher than that of the preset image.
[0062] At this time, the preset image in the first image sequence is not necessarily the image with the lowest resolution. On the basis of already having the exact optical flow information of the previous frame as a reference, the calculation of several layers of images with lower resolutions can be ignored, and only the multi-level optical flow calculation is performed for the subsequent several layers of images with higher resolutions, which can improve the efficiency of optical flow calculation and save calculation time.
[0063] It should be noted that before performing optical flow calculation, the frame can be preprocessed to exclude pixel value differences caused by different lens exposures, lens aperture settings, or video coding effects.
[0064] In a possible implementation, determine the sobel feature in the target frame, and based on the sobel feature, determine the error function during optical flow calculation.
[0065] Specifically, the sobel feature matrix can be used as the feature on the normalized luminance component of the output image of the optical flow calculation, and the error function can be calculated through the following formula.
[0066]
[0067]
[0068]
[0069] Among them, is the x component of the sobel feature of the pixel point (x, y), is the y component of the sobel feature of the pixel point (x, y), is the pixel point is the x component of the sobel feature, is a pixel point for the y component of the sobel feature of is the optical flow of the pixel (x, y). c is the smoothing factor for coherent constraint optimization, which can be the reciprocal of the overlap width, is the spatial average value of the optical flow.
[0070] Moreover, when performing hierarchical optical flow calculation, the hexagonal iteration method can be used for each pixel point to find its optical flow. Each time of iteration, the side length of the hexagon is expanded, and the best optical flow is searched within the hexagonal region each time. Among them, the side length of the hexagon used in the iteration can avoid taking smaller values to reduce the error that the iterative search falls into the local minimum. Also, since there is not much information in the images with lower resolutions in the previous several layers, the number of iterations can be set to be very small. And the images with higher resolutions in the subsequent several layers have already preset the optical flow information obtained in the previous layers, so it is not necessary to perform too many iterations to find the optimal optical flow. Specifically, the number of iterations can be inversely proportional to the resolution and at most does not exceed 10 times, which can not only ensure the efficiency of the iteration but also ensure the quality of the optical flow obtained by the iteration.
[0071] Moreover, in a possible implementation manner, based on the initial optical flow information obtained in each layer, the gradient descent method can be used to determine the candidate optical flow information, and the optical flow information of the target image can be determined based on the initial optical flow information and / or the candidate optical flow information. This can also avoid the error that the iterative search generates the local minimum.
[0072] Specifically, the candidate optical flow information can be obtained through the following formula.
[0073]
[0074]
[0075] where,( ) is the candidate optical flow of the nth layer,( ) is the target optical flow of the nth layer, is the step size of the gradient descent.
[0076] In addition, since the pixels in some regions of the picture tend to be stationary between frames and hardly change, therefore, these regions can be identified by calculating the regional energy of the regions, and the pixel points in these regions can be ignored during the optical flow calculation to reduce the computational amount of the optical flow calculation.
[0077] S14. Based on the target optical flow information, splice each of the overlapping regions of the adjacent original videos to obtain the target video.
[0078] The corresponding relationship between the picture pixels of the images within the overlapping region can be determined through the target optical flow information, thereby determining the stitching relationship between the pictures, achieving panoramic stitching, and reducing stitching traces.
[0079] Optionally, considering the application of panoramic video in ODSV video, when performing stitching and obtaining the target video, the overlapping region can also be stitched according to the distance of the dynamic video recorder, the pupil distance information, and the target optical flow information to obtain two target videos respectively for the left eye and the right eye.
[0080] Since stereoscopic vision is caused by the depth of field of an object and the corresponding binocular parallax, given the camera distance, pupil distance, and the parallax (i.e., optical flow) of each pixel, stereoscopic vision can be re-established in ODSV.
[0081] Figure 4 is an example diagram of the overlapping region of an ODSV shooting system, which consists of four cameras evenly located on the equator. A and B are the positions of the left and right cameras, while L and R are the left eye and the right eye respectively. The line segments AB and LR represent the camera distance and the general pupil distance respectively, and the point P is the position of the object. Figure 5 is another example diagram of the overlapping region of an ODSV shooting system. By using ERP, the quarter-spherical space is flattened into a rectangular image. The ERP overlap has a constant spacing coordinate system, where the vertical axis is latitude and the horizontal axis is longitude. Figure 5 The four sides of the rectangle in Figure 4 can be found through the one-to-one correspondence in Figure 5 The left side of the rectangle in Figure 4 is the optical axis of camera A, that is, the Figure 5 y-axis in Figure 4 The right side of the rectangle in Figure 5 is the optical axis of camera B, that is, the x-axis in Figure 5 The top and bottom sides of the rectangle in are the extended North Pole and South Pole of the quarter-spherical space. Figure 4 in and are the distances between the imaging points and the corresponding optical axes, and they are in a direct proportional relationship with , , , where W represents the width of the overlapping region. For monocular applications, the viewing point is the Figure 4 origin (O) in , and the imaging position of any item is the same for both eyes. Therefore, the theoretical imaging angle of P can be calculated using the following formula
[0082]
[0083]
[0084]
[0085] Where D is the distance between O and P, and OA = OB is the radius of the camera lens.
[0086] Meanwhile, binocular observation introduces two separate observation points L and R, from which the imaging angle of point P can be calculated from the following equations 1 and 1.
[0087]
[0088]
[0089] Where OL is the distance between O and L. Without considering special cases, the above calculations apply to a single camera, and in the ODSV system, two calculations can be used to calculate two pairs of imaging angles respectively.
[0090] For non-occluded regions, the above method can be used, where the object can be captured by two adjacent cameras simultaneously. However, for occluded parts, the details of the images captured by a single camera will be lost. In these regions, the common positioning directions from the two cameras have different depths of field (one foreground and one background). Without the reference of other adjacent cameras, using the and calculated values are completely different.
[0091] The first step in solving occluded regions is to identify them. For non-occluded regions, the left-to-right and right-to-left optical flows should be roughly opposite, i.e., (x, y) + (x, y) ≈ 0, and this relationship no longer holds for occluded regions. Therefore, we use the following rating rule based on vector length to find occluded regions:
[0092]
[0093] Where |...| is the Euclidean length of the vector. The rating rule based on vector length is used instead of the rating rule based on dot product because the optical flow of the background is usually small and may be contaminated by noise in the calculation. Whether the current pixel should be filled with foreground content can be checked by the synthesis method described in the formula for calculating the imaging angle above to determine the background region and the foreground region. In addition, the results of occlusion detection in the previous frame are also used to reduce the noise of occlusion detection in the current frame.
[0094] Once the occluded area is determined, we can generate the foreground object of the virtual viewpoint using only the information from one camera (e.g., using the right part of the left camera). Correspondingly, the right occluded part is synthesized using the image of the right camera, and the left occluded part is synthesized using the image of the left camera. Since the rendering method is different for the three regions (the left occluded part, the right occluded part, and the non-occluded part), this may lead to discontinuity at the boundaries. To solve this problem, we apply a Gaussian filter to the boundaries to ensure a smooth transition between different regions and reduce binocular rivalry.
[0095] It should be noted that after obtaining the target video, the target video can be encoded with ODSV video encoding so that the video can be presented in the VR system.
[0096] Optionally, when encoding, the background blocks and object blocks in the target video can be partitioned based on the target optical flow information; according to the result of the partitioning, the target video is encoded with ODSV video encoding to improve the encoding efficiency.
[0097] Among them, the spatial distribution of the optical flow can be used to segment objects from the background. Based on the optical flow information, the intra-frame variance and inter-frame variance can be obtained, and the intra-frame variance and inter-frame variance are used to represent the spatial and temporal variances of the optical flow of the current partition.
[0098] The intra-frame variance represents the temporal change of the optical flow of the current block. If the current block is part of a static background or a static object, the intra-frame variance will be very small. In this case, the encoding methods using the skip mode and the merge mode are usually more efficient because the motion vectors of these regions will remain unchanged.
[0099] Since the changing objects usually have a certain depth, the pixels within the object should contain highly similar optical flow values (i.e., similar parallax degrees and small inter-frame variance values). In this case, we do not need to partition the partition into very small segments. On the contrary, a large inter-frame variance indicates that the details of this region are more and the motion performance is very fine, so the necessity of partitioning exists.
[0100] Through the above technical methods, at least the following technical effects can be achieved.
[0101] By calculating the optical flow of the overlapping area of the obtained original video through multi-level optical flow calculation and splicing the overlapping area of the original video based on the calculated target optical flow information, a spliced video with finer pictures, more realistic splicing effects, and lower latency can be obtained with shorter calculation time and higher calculation efficiency.
[0102] Figure 6is a block diagram of a video processing device shown according to an exemplary disclosed embodiment. As Figure 6 shown, the device 600 includes:
[0103] A video acquisition module 601, configured to acquire a plurality of original videos, where the plurality of original videos are videos acquired by a plurality of dynamic image recorders arranged at preset positions within the same time period.
[0104] An overlapping module 602, configured to determine an overlapping region of each adjacent original video according to a preset rule corresponding to the preset position, where the adjacent original videos are the original videos acquired by the dynamically arranged image recorders arranged adjacent to each other.
[0105] An optical flow calculation module 603, configured to perform multi-level optical flow calculation on the original videos within each overlapping region to obtain a plurality of target optical flow information.
[0106] A splicing module 604, configured to splice the overlapping regions of each adjacent original video based on the target optical flow information to obtain a target video.
[0107] Optionally, the device further includes a time synchronization module and a calibration module. The time synchronization module is configured to perform time synchronization on the plurality of original videos; the calibration module is configured to calibrate the plurality of synchronized original videos; the overlapping module is configured to determine an overlapping region of each adjacent original video after calibration according to a preset rule corresponding to the preset position; the splicing module is configured to splice the overlapping regions of each adjacent original video after calibration based on the target optical flow information.
[0108] Optionally, the optical flow calculation module includes a plurality of optical flow calculation sub-modules. Each optical flow calculation sub-module is configured to perform the following operations on two target frame images of the two original videos within the overlapping region: downsample the corresponding frame images of the two original videos respectively to obtain two image sequences arranged from low to high in resolution; determine the first initial optical flow information of the first image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the image sequence; based on the first initial optical flow information, determine the optical flow information of the first target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and use the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence as the first initial optical flow information, where the first target image is an image other than the first image in the first image sequence; repeat the step of determining the optical flow information of the first target image relative to the image with the corresponding resolution in the second image sequence based on the first initial optical flow information in the order from low to high of the resolution of the first target image until the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with the corresponding resolution in the second image sequence is determined, and use the final optical flow information as the target optical flow information corresponding to the target frame image.
[0109] Optionally, the optical flow calculation module further includes a candidate calculation sub-module, configured to determine candidate optical flow information using the gradient descent method based on each of the initial optical flow information; the optical flow calculation sub-module is further configured to determine the optical flow information of the target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence based on the initial optical flow information and / or the candidate optical flow information.
[0110] Optionally, the optical flow calculation module further includes a function determination sub-module, configured to determine the sobel feature in the target frame image; based on the sobel feature, determine the error function during optical flow calculation.
[0111] Optionally, the optical flow calculation module further includes a frame type processing sub-module, configured to determine the type of the adjacent frame according to the correlation between the target frame and the adjacent frame after obtaining the target optical flow information of the target frame; when the adjacent frame is a P-frame, based on the target optical flow information of the target frame, determine the optical flow information of a preset image in the first image sequence of the adjacent frame relative to the image with the corresponding resolution in the second image sequence, and use the optical flow information of the preset image relative to the image with the corresponding resolution in the second image sequence as the second initial optical flow information; in the order from low to high of the resolution of the second target image, repeatedly execute the step of determining the optical flow information of the second target image relative to the image with the corresponding resolution in the second image sequence based on the second initial optical flow information until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with the corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the adjacent frame, where the second target image is an image in the first image sequence of the adjacent frame with a resolution higher than that of the preset image.
[0112] Optionally, the stitching module is further configured to stitch the overlapping area according to the distance of the dynamic video recorder, the pupil distance information, and the target optical flow information, to obtain two target videos respectively for the left eye and the right eye.
[0113] Optionally, the device further includes an encoding module, configured to partition the background block and the object block in the target video based on the target optical flow information; and perform ODSV video encoding on the target video according to the result of the partitioning.
[0114] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0115] Through the above technical solution, the optical flow of the overlapping area of the acquired original video is calculated by means of multi-level optical flow calculation, and the overlapping area of the original video is stitched based on the obtained target optical flow information, so that a stitched video with finer pictures, more realistic stitching effect and lower latency can be obtained with shorter calculation time and higher calculation efficiency.
[0116] Figure 7 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. As Figure 7 shown, the electronic device 700 may include: a processor 701, a memory 702. The electronic device 700 may further include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.
[0117] Among them, the processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in the above video processing method. The memory 702 is used to store various types of data to support the operation of the electronic device 700. These data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 703 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory 702 or sent through the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.
[0118] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the above video processing method.
[0119] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above video processing method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 702 including program instructions, and the above program instructions can be executed by the processor 701 of the electronic device 700 to complete the above video processing method.
[0120] Figure 8 is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 can be provided as a server. Referring to Figure 8 , the electronic device 800 includes a processor 822, the number of which can be one or more, and a memory 832 for storing computer programs executable by the processor 822. The computer programs stored in the memory 832 can include one or more modules each corresponding to a set of instructions. In addition, the processor 822 can be configured to execute the computer program to execute the above video processing method.
[0121] In addition, the electronic device 800 can further include a power supply component 826 and a communication component 850. The power supply component 826 can be configured to perform power management of the electronic device 800, and the communication component 850 can be configured to implement communication of the electronic device 800, for example, wired or wireless communication. In addition, the electronic device 800 can further include an input / output (I / O) interface 858. The electronic device 800 can operate based on an operating system stored in the memory 832, such as Windows ServerTM, Mac OSXTM, UnixTM, LinuxTM, and so on.
[0122] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions, and when the program instructions are executed by a processor, the steps of the above video processing method are implemented. For example, the computer-readable storage medium may be the memory 832 including the program instructions as described above, and the above program instructions may be executed by the processor 822 of the electronic device 800 to complete the above video processing method.
[0123] In another exemplary embodiment, there is also provided a computer program product, which includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for executing the above video processing method when executed by the programmable device.
[0124] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0125] In addition, it should be noted that, in the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination manners.
[0126] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
Claims
1. A real-time video stitching method based on optical flow calculation, characterized in that, the method includes: Obtaining a plurality of original videos, where the plurality of original videos are videos obtained by a plurality of dynamic image recorders arranged at preset positions; Determining an overlapping area of each adjacent pair of original videos according to a preset rule corresponding to the preset position, where the adjacent original videos are the original videos obtained by the dynamically image recorders arranged adjacent to each other; Performing multi-level optical flow calculation on the original videos within each overlapping area to obtain a plurality of target optical flow information; Based on the target optical flow information, stitching the overlapping areas of each adjacent pair of original videos to obtain a target video; The performing multi-level optical flow calculation on the original videos within each overlapping area to obtain a plurality of target optical flow information includes: Performing the following operations on a plurality of target frame images of two original videos within the overlapping area: Performing downsampling on the corresponding frame images of the two original videos respectively to obtain two image sequences arranged from low to high in resolution; Determining first initial optical flow information of a first image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the first image sequence; Based on the first initial optical flow information, determining optical flow information of a first target image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, and using the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence as the first initial optical flow information, where the first target image is an image in the first image sequence other than the first image; Repeating the step of determining the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence based on the first initial optical flow information in the order from low to high of the resolution of the first target image until determining the final optical flow information of an image with the same resolution as the original video in the first image sequence relative to the image with a corresponding resolution in the second image sequence, and using the final optical flow information as the target optical flow information corresponding to the target frame image; After obtaining the target optical flow information of the target frame image, determining the type of the adjacent frame image according to the correlation between the target frame image and the adjacent frame image; When the adjacent frame image is a P-frame image, determining the optical flow information of a preset image in the first image sequence of the adjacent frame image relative to an image with a corresponding resolution in the second image sequence of the adjacent frame image based on the target optical flow information of the target frame image, and using the optical flow information of the preset image relative to the image with a corresponding resolution in the second image sequence of the adjacent frame image as the second initial optical flow information; where the preset image is not the image with the lowest resolution in the first image sequence of the adjacent frame image; Repeat the step of determining the optical flow information of the image with the corresponding resolution in the second image sequence of the second target image relative to the adjacent frame image based on the second initial optical flow information in ascending order of the resolution of the second target image until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence of the adjacent frame image relative to the image with the corresponding resolution in the second image sequence of the adjacent frame image, and use the final optical flow information as the target optical flow information corresponding to the adjacent frame image, where the second target image is an image in the first image sequence of the adjacent frame image with a resolution higher than the preset image.
2. The method according to claim 1, wherein, after obtaining the multiple original videos, the method further includes: performing time synchronization on the multiple original videos; calibrating the synchronized multiple original videos; The determining the overlapping region of each adjacent original video according to the preset rule corresponding to the preset position includes: determining the overlapping region of each adjacent calibrated original video according to the preset rule corresponding to the preset position; The splicing the overlapping region of each adjacent original video based on the target optical flow information includes: splicing the overlapping region of each adjacent calibrated original video according to the target optical flow information.
3. The method according to claim 1, wherein, the method further includes: determining candidate optical flow information using the gradient descent method based on each of the initial optical flow information; The determining the optical flow information of the target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence based on the initial optical flow information includes: determining the optical flow information of the target image in the first image sequence relative to the image with the corresponding resolution in the second image sequence based on the initial optical flow information and / or the candidate optical flow information.
4. The method according to claim 1, wherein, before respectively downsampling the target frame images of the two original videos, the method further includes: determining the sobel feature in the target frame image; determining the error function during optical flow calculation based on the sobel feature.
5. The method according to claim 1, wherein, The splicing the overlapping region of each adjacent original video based on the target optical flow information to obtain the target video includes: splicing the overlapping region according to the distance of the dynamic video recorder, the pupil distance information and the target optical flow information to obtain two target videos respectively for the left eye and the right eye.
6. The method according to any one of claims 1-5, wherein, the method further includes: partitioning the background block and the object block in the target video based on the target optical flow information; performing ODSV video encoding on the target video according to the result of the partitioning.
7. A real-time video splicing device based on optical flow calculation, wherein, the device includes: A video acquisition module, configured to acquire a plurality of original videos, where the plurality of original videos are videos acquired by a plurality of dynamic image recorders arranged at preset positions within the same time period; An overlapping module, configured to determine an overlapping region of each adjacent original video according to a preset rule corresponding to the preset position, where the adjacent original videos are the original videos acquired by the dynamically arranged adjacent image recorders; An optical flow calculation module, configured to perform multi-level optical flow calculation on the original videos within each overlapping region to obtain a plurality of target optical flow information; A splicing module, configured to splice the overlapping regions of each adjacent original video based on the target optical flow information to obtain a target video; Specifically, the optical flow calculation module is configured to perform the following operations on a plurality of target frame images of two original videos within the overlapping region: Perform downsampling on the corresponding frame images of the two original videos respectively to obtain two image sequences arranged from low to high in resolution; Determine first initial optical flow information of a first image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, where the first image is the image with the lowest resolution in the first image sequence; Based on the first initial optical flow information, determine optical flow information of a first target image in the first image sequence relative to an image with a corresponding resolution in the second image sequence, and use the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence as the first initial optical flow information, where the first target image is an image other than the first image in the first image sequence; Repeat the step of determining the optical flow information of the first target image relative to the image with a corresponding resolution in the second image sequence based on the first initial optical flow information in the order from low to high of the resolution of the first target image until determining the final optical flow information of the image with the same resolution as the original video in the first image sequence relative to the image with a corresponding resolution in the second image sequence, and use the final optical flow information as the target optical flow information corresponding to the target frame image; Specifically, after obtaining the target optical flow information of the target frame image, the optical flow calculation module is configured to determine the type of the adjacent frame image according to the correlation between the target frame image and the adjacent frame image; When the adjacent frame image is a P-frame image, determine the optical flow information of a preset image in the first image sequence of the target frame image relative to an image with a corresponding resolution in the second image sequence of the adjacent frame image based on the target optical flow information of the target frame image, and use the optical flow information of the preset image relative to the image with a corresponding resolution in the second image sequence of the adjacent frame image as the second initial optical flow information; where the preset image is not the image with the lowest resolution in the first image sequence of the adjacent frame image; Repeat the step of determining the optical flow information of the image with the corresponding resolution in the second image sequence of the second target image relative to the adjacent frame picture based on the second initial optical flow information in ascending order of the resolution of the second target image until the final optical flow information of the image with the same resolution as the original video in the first image sequence of the adjacent frame picture relative to the image with the corresponding resolution in the second image sequence of the adjacent frame picture is determined, and use the final optical flow information as the target optical flow information corresponding to the adjacent frame picture, where the second target image is an image in the first image sequence of the adjacent frame picture with a resolution higher than the preset image.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the steps of the method for real-time video stitching based on optical flow calculation according to any one of claims 1-6.
9. An electronic device, characterized in that, comprising: a memory, on which a computer program is stored; a processor for executing the computer program in the memory to implement the steps of the method for real-time video stitching based on optical flow calculation according to any one of claims 1-6.
Citation Information
Patent Citations
Video generation method, device and equipment for virtual viewpoint
CN109379577A
Stitching frames into a panoramic frame
US20180063513A1