A real-time arbitrary-view, free-view video generation method and system

By using dense optical flow field estimation and optical flow embedding techniques in a free-view video system, embedded optical flow and mask matrices are generated, solving the problem of high computational resource and transmission bandwidth requirements in high-concurrency scenarios. This enables the generation of high-quality arbitrary-view virtual views under low-resource conditions and supports personalized services for multiple users.

CN119011871BActive Publication Date: 2025-10-17SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411010941.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-10-17
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

Existing free-viewpoint video systems have high demands for computing resources and transmission bandwidth in high-concurrency scenarios, making it difficult to achieve high-quality generation of arbitrary viewpoints in real time, especially performing poorly on low-end user terminal devices.

Method used

By acquiring multiple frames of color images of the same calibration object from different perspectives, a dense optical flow field estimation network is used to determine the bidirectional optical flow. An embedded optical flow and mask matrix are generated in the optical flow embedder. Combined with forward warping and fusion processing, a virtual view of any perspective is interpolated in real time.

Benefits of technology

It enables the efficient generation of high-quality virtual views with arbitrary perspectives under conditions of low computing resources and transmission bandwidth, and is suitable for edge servers and clients, supporting multi-user personalized free-view video services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119011871B_ABST
    Figure CN119011871B_ABST
Patent Text Reader

Abstract

The present disclosure provides a real-time arbitrary view video generation method and system, wherein the real-time arbitrary view generation method comprises: acquiring multiple frames of color images of different views of the same calibration object; inputting the multiple frames of color images into a preset dense optical flow field estimation network to determine bidirectional optical flow; inputting the multiple frames of color images and the bidirectional optical flow into an optical flow embedder to determine embedded optical flow and a mask matrix; determining first optical flow and second optical flow according to a third view input by a user and the embedded optical flow; and performing forward warping processing on the multiple frames of color images according to the first optical flow and the second optical flow to determine a complete virtual view at the third view. Through the above technical solution, real-time arbitrary view generation is realized, which is lightweight, efficient, reduces the consumption of computing resources, and improves user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of immersive self-media, in particular to a real-time arbitrary view and free-view video generation method and system. BACKGROUND

[0002] Free-view video is a new type of immersive media form with strong interactive characteristics, which has attracted more attention than typical application scenarios such as cloud gaming and remote virtual reality, and is expected to change the way we consume visual content.

[0003] Generally, a free-view system includes a multi-view acquisition system, free-view content production, and encoding transmission and a client. The multi-view acquisition system aims to provide multi-angle and multi-directional video source information for the free-view system. However, due to the limitations of hardware costs and data volume, the acquisition system can only be limited to sparse and limited number of cameras for shooting. Virtual view synthesis technology aims to obtain other uncollected view information from limited view information. The DIBR (Depth Image-Based Rendering) method is the most commonly used view synthesis method in free-view systems. However, due to the introduction of occlusion and black holes in three-dimensional image distortion, the synthesis results are often unsatisfactory. In addition, the acquisition of accurate depth maps also faces great challenges.

[0004] Free-view systems can generally be divided into two models: central and distributed.

[0005] In the central model, the required viewpoints of different users are synthesized at the server end. Some existing real-time view synthesis methods require sufficient computing resources, so a server can only serve a limited number of user terminals. As the number of user access increases, the number of servers also needs to increase accordingly. This model is difficult to cope with high concurrency scenarios and can cause additional response delays during the interaction process.

[0006] The distributed model can serve multiple users at the same time because it performs the view synthesis process at the client end. However, the Multiview-VideoPlus-Depth (MVD) representation required for view synthesis needs to be transmitted to the user, which can result in high transmission bandwidth. In addition, the view synthesis method requires a large amount of processing power, which is not friendly to some low-end user terminals. SUMMARY

[0007] In view of the defects in the prior art, the purpose of the present disclosure is to provide a real-time arbitrary view generation method and system, and a free view video method and system.

[0008] In order to achieve the above-mentioned purpose, according to a first aspect of the present disclosure, a real-time arbitrary view generation method is provided, comprising:

[0009] Obtaining multiple frames of color images of different views of the same calibration object, wherein the multiple frames of color images of different views of the same calibration object include a first color texture image of a first view and a second color texture image of a second view;

[0010] Inputting the first color texture image and the second color texture image into a preset dense optical flow field estimation network to determine a bidirectional optical flow;

[0011] Inputting the first color texture image, the second color texture image and the bidirectional optical flow into an optical flow embedder to determine a preset number of embedded optical flows and a mask matrix;

[0012] According to a third view input by a user and the preset number of embedded optical flows, determining a first preset number of first optical flows of the first color texture image to a virtual view of the third view and a first preset number of second optical flows of the second color texture image to the virtual view of the third view;

[0013] According to the first preset number of first optical flows and the first preset number of second optical flows, performing forward warping processing on the first color texture image and the second color texture image respectively, and performing fusion processing on the images after the forward warping processing according to the mask matrix to determine a complete virtual view at the third view.

[0014] Optionally, the step of inputting the first color texture image, the second color texture image and the bidirectional optical flow into the optical flow embedder to determine a preset number of embedded optical flows and a mask matrix comprises:

[0015] Performing feature extraction processing on the first color texture image and the second color texture image to determine image features of the color images of the first view and the second view at different resolution scales;

[0016] Performing first warping processing on the image features of the color images of the first view and the second view according to an initial optical flow to determine image features after the first warping processing;

[0017] Sequentially performing alignment processing and splicing processing on the image features after the first warping processing and original image features of a warping end point to determine spliced image features;

[0018] The stitching image features are sequentially encoded and decoded to determine a second preset number of optical flow residuals, wherein the second preset number has the same value as the preset logarithm;

[0019] The initial optical flow is optimized according to the second preset number of optical flow residuals to determine a preset logarithm of embedded optical flow and a mask matrix.

[0020] Optionally, the first preset number of first optical flows of the first color texture image to a virtual view of the third view and the first preset number of second optical flows of the second color texture image to the virtual view of the third view are determined according to the third view input by the user and the preset logarithm of embedded optical flow, and the method comprises:

[0021]

[0022] wherein, represents the first optical flow, and p represents the third view, represents the embedded optical flow from the first view to the second view, represents the second optical flow, represents the embedded optical flow from the second view to the first view.

[0023] Optionally, the fusion processing comprises first fusion processing and second fusion processing.

[0024] The first color texture image and the second color texture image are subjected to forward warping processing by using the first preset number of first optical flows and the first preset number of second optical flows, and the images subjected to the forward warping processing are subjected to fusion processing by using the mask matrix to determine a complete virtual view at the third view, and the method comprises:

[0025] The first color texture image is subjected to forward warping to the third view according to the first preset number of first optical flows by using a forward warping operator to determine a first preset number of first-side virtual third color texture images;

[0026] The second color texture image is subjected to forward warping to the third view according to the first preset number of second optical flows by using the forward warping operator to determine a first preset number of second-side virtual third color texture images;

[0027] The first preset number of first-side virtual third color texture images and the first preset number of second-side virtual third color texture images are subjected to the first fusion processing, respectively, to determine first-side fusion images and second-side fusion images subjected to the first fusion processing;

[0028] The first side fusion image and the second side fusion image are subjected to the second fusion processing according to the mask matrix, and a complete virtual view at the third view angle is determined.

[0029] Optionally, the second fusion processing of the first side fusion image and the second side fusion image according to the mask matrix to determine the complete virtual view at the third view angle comprises:

[0030] I p = M ¢ I left→p + (1-M) ¢ I right→p

[0031]

[0032]

[0033] wherein I p represents the complete virtual image at the third view angle, M represents the mask matrix, I left→p represents the first side fusion image, I right→p represents the second side fusion image, represents a first side virtual third color texture image, represents a second side virtual third color texture image, and k represents the first preset number.

[0034] According to a second aspect of the present disclosure, there is provided a real-time arbitrary view generation system, comprising:

[0035] An acquisition module is configured to acquire a plurality of color images of a same calibration object at different view angles, wherein the plurality of color images of the same calibration object at different view angles comprise a first color texture image at a first view angle and a second color texture image at a second view angle.

[0036] A first determination module is configured to input the first color texture image and the second color texture image into a preset dense optical flow field estimation network to determine a bidirectional optical flow.

[0037] A second determination module is configured to input the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder to determine a preset number of embedded optical flows and a mask matrix.

[0038] A third determination module is configured to determine, according to a third view angle input by a user and the preset number of embedded optical flows, a first preset number of first optical flows of the first color texture image to a virtual view at the third view angle and a first preset number of second optical flows of the second color texture image to the virtual view at the third view angle.

[0039] a fourth determination module, configured to perform forward warping on the first color texture image and the second color texture image according to the first preset number of first optical flows and the first preset number of second optical flows, and to fuse the images after the forward warping according to the mask matrix to determine a complete virtual view at the third perspective.

[0040] According to a third aspect of the present disclosure, a method for generating a free-viewpoint video is provided, comprising:

[0041] Acquire multiple frames of color images of the same calibration object from different perspectives;

[0042] Combining the multiple color image frames in pairs;

[0043] Obtaining a preset logarithmic embedded optical flow and mask matrix in real time from each of the two-by-two combined multi-frame color images according to the method provided by the first aspect of the present disclosure;

[0044] Adaptively splicing the pairwise combination of the multi-frame color images and the preset logarithmic embedded optical flows and mask matrices to form a plurality of composite images in the spatial domain;

[0045] Encoding the multiple composite images, dividing them into video segments in the time domain, and transmitting them through the HLS protocol;

[0046] Different clients interactively select target synthetic images and download them, and synthesize a complete virtual view at the user's viewing angle in real time according to the method provided in the first aspect of the present disclosure.

[0047] Optionally, adaptively splicing each of the two-by-two combination of the multi-frame color images and the preset logarithmic embedded optical flows and mask matrices to form a plurality of composite images in the spatial domain includes:

[0048] Performing channel-dimensional synthesis processing on the preset logarithmic embedded optical flows and mask matrices corresponding to each of the two-by-two combinations of the multi-frame color images to determine a three-channel matrix corresponding to each of the two-by-two combinations of the multi-frame color images;

[0049] Performing spatial domain synthesis processing on each of the two-by-two combinations of the multi-frame color images and the three-channel matrix corresponding to each of the two-by-two combinations of the multi-frame color images to determine a synthesized image;

[0050] The multiple frames of color images at different viewing angles and all pairwise combinations of embedded optical flows and mask matrices are subjected to resolution reduction processing to form a total composite image.

[0051] Optionally, the different clients interactively select target synthesized images and download, and complete virtual views at the viewing angle of the user are synthesized in real time according to the method provided in the first aspect of the present disclosure, comprising:

[0052] A global viewing angle index is adopted to download the synthesized image corresponding to the viewing angle of the user and the total synthesized image;

[0053] In one video slice time, the viewing angle of the user falls within the downloaded synthesized image, and the complete virtual view at the viewing angle of the user is synthesized in real time according to the method provided in the first aspect of the present disclosure by using the synthesized image;

[0054] In one video slice time, the viewing angle of the user falls outside the downloaded synthesized image, and the complete virtual view at the viewing angle of the user is synthesized in real time according to the method provided in the first aspect of the present disclosure by using the total synthesized image;

[0055] After one video slice time, the user downloads the complete virtual video at the viewing angle and the total synthesized image.

[0056] According to the fourth aspect of the present disclosure, a free-viewpoint video generation system is provided, comprising:

[0057] A collection server module is configured to acquire multiple frames of color images at different viewing angles of the same calibration object;

[0058] A cloud processing and content distribution network module is configured to combine the multiple frames of color images two by two and distribute them to an edge server module by a content distribution network;

[0059] An edge server module is configured to acquire a preset number of embedded optical flows and mask matrices in real time for each of the two-by-two combined multiple frames of color images; adaptively stitch each of the two-by-two combined multiple frames of color images and the preset number of embedded optical flows and mask matrices to form multiple synthesized images in a spatial domain; encode the multiple synthesized images, divide them into video segments in a time domain, and transmit them through an HLS protocol;

[0060] A client module is configured to enable different clients to interactively select target synthesized images and download, and complete virtual views at the viewing angle of the user are synthesized in real time.

[0061] Compared with the prior art, the embodiments of the present disclosure have at least one of the following beneficial effects:

[0062] By the above technical solution, after acquiring color images of different perspectives facing the same calibration object, the two-by-two combined images are based on the collected multiple frames, and a virtual intermediate perspective view of any perspective between the two-by-two combined images is interpolated in real time, which is lightweight, high and efficient, can use less computing resources to interpolate high-quality virtual views of any perspective in real time, and is convenient to deploy in edge server or client, which is beneficial to the development of free perspective video generation system.

[0063] The free perspective video generation method provided by the present disclosure is a free perspective video generation method for multiple users. The free perspective video generation system built based on this method decouples the number of access users and the load of the edge server, so that a single edge server can provide personalized free perspective video services for multiple users. BRIEF DESCRIPTION OF DRAWINGS

[0064] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments, made with reference to the accompanying drawings:

[0065] Figure 1 is a flowchart of a real-time arbitrary perspective generation method according to an exemplary embodiment.

[0066] Figure 2 is a flowchart of a method for determining a preset logarithmic embedded optical flow and mask matrix according to an exemplary embodiment.

[0067] Figure 3 is a process diagram of real-time arbitrary perspective generation according to an exemplary embodiment.

[0068] Figure 4 is a flowchart of a method for determining a complete virtual view at a third perspective according to an exemplary embodiment.

[0069] Figure 5 is a flowchart of a free perspective video generation method according to an exemplary embodiment.

[0070] Figure 6 is a schematic diagram of the arrangement of an image collector according to an exemplary embodiment.

[0071] Figure 7 is a block diagram of a real-time arbitrary perspective generation system according to an exemplary embodiment.

[0072] Figure 8 is a schematic diagram of the architecture of a free perspective video generation system according to an exemplary embodiment. DETAILED DESCRIPTION

[0073] The application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be pointed out that, for those skilled in the art, without departing from the concept of the application, a number of modifications and improvements can be made. These all belong to the protection scope of the application.

[0074] Figure 1 A flow chart of a real-time arbitrary view generation method according to an exemplary embodiment is shown.

[0075] As shown in Figure 1 , a real-time arbitrary view generation method includes S11 to S15.

[0076] S11, a plurality of color images of different views of the same calibration object are acquired.

[0077] The plurality of color images of different angles of the same calibration object include a first color texture image of a first view and a second color texture image of a second view.

[0078] The image collectors are arranged around the same calibration object at a preset angle and a preset radius, and can adopt devices such as cameras or video cameras. The baseline between every two image collectors, i.e. the distance between the two image collectors, is set to an appropriate length, and the length of the baseline should not be too long.

[0079] The two image collectors adjacent in space are respectively referred to as a first image collector and a second image collector. In the present disclosure, the image collected by the first image collector is the first color texture image of the first view, and the image collected by the second image collector is the second color texture image of the second view, and the first color texture image and the second color texture image are collected synchronously.

[0080] In the present disclosure, the first view can also be regarded as a left view, and the second view can also be regarded as a right view.

[0081] In a possible embodiment, the image collector can adopt an industrial video camera, the baseline between every two image collectors is set to 40 cm, and the first color texture image and the second color texture image are both in RGB format.

[0082] S12, the first color texture image and the second color texture image are input into a preset dense optical flow field estimation network to determine a bidirectional optical flow.

[0083] The preset dense optical flow field estimation network is an open source network, the first color texture image and the second color texture image are input into the preset dense optical flow field estimation network, and the preset dense optical flow field estimation network such as a RAFT network is used to simultaneously calculate the bidirectional pixel-by-pixel offset of the first color texture image to the second color texture image and the second color texture image to the first color texture image, that is, bidirectional optical flow, including optical flow F1 and optical flow F2, wherein the optical flow F1 and the optical flow F2 are both two-dimensional vector matrix images.

[0084] The downsampling processing before the first color texture image and the second color texture image are input and the upsampling processing after the output of the preset dense optical flow field estimation network can be determined according to factors such as computing resources.

[0085] In the present disclosure, the first color texture image and the second color texture image are first subjected to downsampling processing, the first color texture image and the second color texture image are downsampled to 1 / 2 of the size of the original image, and then a low-resolution optical flow is calculated through a RAFT network (Recurrent All-Pairs Field Transforms for Optical Flow), and finally the low-resolution optical flow and a visibility mask are subjected to upsampling processing to the size of the original image to obtain the optical flow F1 and the optical flow F2.

[0086] Through the above technical solution, the preset dense optical flow field estimation network is used, on the one hand, the effect is good, and on the other hand, the performance and speed can be balanced by adjusting the parameters to ensure the real-time performance of the network.

[0087] S13, input the first color texture image, the second color texture image and the bidirectional optical flow into the optical flow embedder, and determine the preset number of embedded optical flows and mask matrices.

[0088] The preset number of values is k, and each pair of embedded optical flows includes an embedded optical flow from the first view to the second view and an embedded optical flow from the second view to the first view. The value of k is determined according to the deployment of the application. Specifically, the obtained preset number of embedded optical flows and mask matrices need to be encoded and transmitted through a network in a free view video system. The larger the value of k, the better the quality of the virtual view image generated, but at the same time, the amount of encoded data is larger, and the network transmission bandwidth occupied is also higher.

[0089] The mask matrix is used to describe the visibility of the first view and the second view.

[0090] S14, determining a first preset number of first optical flows from the first color texture image to the virtual view of the third view angle and a first preset number of second optical flows from the second color texture image to the virtual view of the third view angle according to the third view angle input by the user and the preset logarithmic embedded optical flow.

[0091] The preset logarithmic embedded optical flow has the same value as the first preset number.

[0092] The optical flow represents the pixel offset from one image to another image, and is represented by a two-dimensional vector, which can represent the horizontal offset or the vertical offset in the image matrix.

[0093] S15, performing forward warping processing on the first color texture image and the second color texture image according to the first preset number of first optical flows and the first preset number of second optical flows respectively, and performing fusion processing on the images after the forward warping processing according to the mask matrix to determine the complete virtual view at the third view angle.

[0094] The warping process is a process of adding the pixel offset to the image.

[0095] As an example, there is a whole right pixel offset from the first color texture image of the first view to the second color texture image of the second view, and the warping process is equivalent to moving the first color texture image of the first view to the right by an offset.

[0096]

[0097] Wherein, x1 represents the horizontal coordinate of the pixel of the first color texture image, y1 represents the vertical coordinate of the pixel of the first color texture image, u1 represents the horizontal coordinate of the pixel of the second color texture image, v1 represents the vertical coordinate of the pixel of the second color texture image, and F1 represents the optical flow from the first color texture image to the second color texture image.

[0098] As another example, there is a whole left pixel offset from the second color texture image of the second view to the first color texture image of the first view, and the warping process is equivalent to moving the first color texture image of the second view to the left by an offset.

[0099]

[0100] Wherein, u2 represents the horizontal coordinate of the pixel of the first color texture image, v2 represents the vertical coordinate of the pixel of the first color texture image, x2 represents the horizontal coordinate of the pixel of the second color texture image, y2 represents the vertical coordinate of the pixel of the second color texture image, and F2 represents the optical flow from the second color texture image to the first color texture image.

[0101] In the present disclosure, when performing pixel mapping according to the first optical flow and the second optical flow, i.e., the warping process, no inverse warping process is adopted, but a forward warping operator σ is adopted to perform forward warping processing.

[0102] By adopting the forward warping operator to perform forward warping processing, optical flow reversal can be avoided, the first color texture image and the first optical flow are directly used to obtain an image at the third perspective, the efficiency of the virtual view is improved, and artifacts are reduced.

[0103] As an example, the first color texture image at the first perspective is forward warped to the target position p, i.e., at the third perspective: wherein I left represents the first color texture image, F left→p represents a first optical flow, and σ represents a forward warping operator.

[0104] wherein the process of forward warping the first color texture image at the first perspective to the target position p includes:

[0105] mapping the first color texture image pixel q to the third perspective according to the first optical flow, and bilinearly sampling the pixel p, i.e.,

[0106] p = bilinear(q + F left→p [q])

[0107] wherein bilinear represents a bilinear sampling operator.

[0108] Bilinear weights:

[0109] u = p - (q + F left→p [q])

[0110] wherein u represents bilinear weights.

[0111] Bilinear kernel:

[0112] b(u) = max(0, 1 - |u x |) · max(0, 1 - |u y |)

[0113] wherein b(u) represents a bilinear kernel, u x represents an x component of the bilinear weights u, and u y represents a y component of the bilinear weights u.

[0114] The current pixel value is determined by bilinear weighting:

[0115]

[0116] wherein I p [p] represents the current pixel value.

[0117] As another example, the second color texture image of the second view is forward warped to the target position p, i.e. the third view, as follows: where I right represents the second color texture image, F right→p represents a second optical flow.

[0118] where the process of forward warping the second color texture image of the second view to the target position p is the same as the method of forward warping the first color texture image of the first view to the target position p, which will not be described here.

[0119] Thus, the first view can obtain k groups of images mapped to the third view, and the second view can obtain k groups of images mapped to the third view.

[0120] The images after the forward warping process are fused by using the mask matrix to determine the complete virtual view at the third view.

[0121] Through the above technical solution, after obtaining color images of different views facing the same calibration object, the two-by-two combined images are obtained based on the collected multiple images, and a virtual intermediate view of any view between the two-by-two combined images is interpolated in real time. It is lightweight and efficient, can use less computing resources to interpolate high-quality virtual views of any view in real time, and is convenient to deploy in edge server or client, which is conducive to the development of free view video generation system.

[0122] Figure 2 is a flowchart of a method for determining a preset number of embedded optical flows and mask matrices according to an example embodiment. Figure 3 is a process diagram of real-time arbitrary view generation according to an example embodiment.

[0123] As Figure 2 , Figure 3 shown, in some possible embodiments, the first color texture image, the second color texture image and the bidirectional optical flow are input into the optical flow embedder to determine a preset number of embedded optical flows and mask matrices, including S21 to S25.

[0124] S21, the first color texture image and the second color texture image are subjected to feature extraction processing to determine the image features of the color images of the first view and the second view at different resolution scales.

[0125] Wherein, the feature extractor can be used to extract the image features of the first color texture image and the second color texture image at different resolution scales.

[0126] S22 , performing a first distortion process on the image features of the color images of the first viewing angle and the second viewing angle according to the initial optical flow, to determine the image features after the first distortion process.

[0127] S23 , sequentially aligning and splicing the features of the image after the distortion process and the features of the original image at the distortion end point to determine spliced ​​image features.

[0128] Among them, the first distortion processing is the distortion process mentioned above, which distorts the image features of one perspective to another perspective through the distortion process and aligns them with the image features of the other perspective, and splices the aligned image features to determine the image features after splicing.

[0129] S24 , performing encoding and decoding processing on the spliced ​​image features in sequence to determine a second preset number of optical flow residuals.

[0130] The value of the second preset number is the same as the value of the preset logarithm.

[0131] like Figure 3 As shown, in a possible embodiment, the spliced ​​image features are input into the Encoder layer for optical flow encoding, and then passed through the Decoder layer for optical flow decoding, wherein the Decoder layer and the Encoder layer share features through jump connections and output a set of residuals for optimizing the initial optical flow. The number of the residuals in the set is k, which has the same meaning and value as k in step S13 of the present disclosure.

[0132] S25 , optimizing the initial optical flow according to a second preset number of optical flow residuals to determine a preset number of embedded optical flows and a mask matrix.

[0133] Continuing with the above example, we can use k optical flow residuals to optimize the initial optical flow and determine a set of k pairs of embedded optical flows and mask matrices for describing the first and second perspectives, where the k pairs of embedded optical flows are:

[0134]

[0135] in, represents the embedded optical flow from the first view to the second view, Represents the embedded optical flow from the second view to the first view.

[0136] In some possible embodiments, determining, based on the third perspective input by the user and a preset logarithm of embedded optical flows, a first preset number of first optical flows from the first color texture image to the virtual view of the third perspective and a first preset number of second optical flows from the second color texture image to the virtual view of the third perspective includes:

[0137]

[0138] wherein, represents the first optical flow, p represents the third view angle, represents the embedding optical flow from the first view angle to the second view angle, represents the second optical flow, represents the embedding optical flow from the second view angle to the first view angle.

[0139] Figure 4 is a method flow chart for determining a complete virtual view at a third view angle according to an exemplary embodiment.

[0140] As Figure 3 , Figure 4 shown in some possible embodiments, a first preset number of first optical flows and a first preset number of second optical flows are used to perform forward warping processing on a first color texture image and a second color texture image, and a mask matrix is used to perform fusion processing on the images after the forward warping processing, and the complete virtual view at the third view angle is determined, including S31 to S34.

[0141] In the present disclosure, the fusion processing includes first fusion processing and second fusion processing.

[0142] S31, using a forward warping operator to forward warp the first color texture image to the third view angle according to the first preset number of first optical flows, to determine a first preset number of first-side virtual third color texture images.

[0143] S32, using a forward warping operator to forward warp the second color texture image to the third view angle according to the first preset number of second optical flows, to determine a first preset number of second-side virtual third color texture images.

[0144] According to the above example, in S31 to S32, according to k pairs of optical flows composed of the first optical flow and the second optical flow, the first view angle and the second view angle can respectively obtain k groups of images mapped to the target position, i.e. k groups of first-side virtual third color texture images and k groups of second-side virtual third color texture images, wherein the first-side virtual third color texture image and the second-side virtual third color texture image are both images of the third view angle:

[0145]

[0146] wherein, represents the first-side virtual third color texture image, represents the first-side second-side virtual third color texture image, I left represents the first color texture image, I right represents the second color texture image, represents the first optical flow, represents the second optical flow.

[0147] S33, performing first fusion processing on the first preset number of first-side virtual third color texture images and the first preset number of second-side virtual third color texture images respectively to determine a first-side fusion image and a second-side fusion image after first fusion processing.

[0148] wherein the first-side fusion image:

[0149]

[0150] wherein I left→p represents the first-side fusion image, represents the first-side virtual third color texture image, and k represents the first preset number.

[0151] the second-side fusion image:

[0152]

[0153] wherein I right→p represents the second-side fusion image, represents the second-side virtual third color texture image.

[0154] S34, performing second fusion processing on the first-side fusion image and the second-side fusion image according to the mask matrix to determine a complete virtual view at the third view angle.

[0155] Based on the above example,

[0156] I p = M O I left→p + (1-M) O I right→p

[0157] wherein I p represents the complete virtual image at the third view angle.

[0158] In the present disclosure, the mask matrix is used as a weight for weighted sum processing.

[0159] In one possible embodiment, the leftmost region at the first view angle is invisible in the third view angle, and after the warping processing, the leftmost region of the first-side virtual third color texture image is an invalid region, at this time, M = 0, and the mask matrix M is learned by the convolutional neural network.

[0160] Based on the same inventive concept, the present disclosure also provides a free-view video generation method.

[0161] Figure 5 is a flowchart of a free-view video generation method according to an exemplary embodiment.Figure 6 is a schematic diagram of an arrangement of an image collector according to an example embodiment.

[0162] As shown in Figure 5 the disclosure also provides a free-view video generation method, which includes S101-S106.

[0163] S101, acquiring multiple frames of color images of different views of the same calibration object.

[0164] As shown in Figure 6 as an example, the image collector can be an industrial camera that supports 4K / 120FPS scene shooting.

[0165] In the disclosure, the number of industrial cameras can be 12, arranged in a fixed arc, with a field of view angle of 70 degrees, and 12 industrial cameras are used to shoot the scene. The image information collected by the 12 industrial cameras is transmitted to the cloud for processing.

[0166] S102, combining the multiple frames of color images two by two.

[0167] S103, acquiring a preset number of embedded optical flow and mask matrices for each two-by-two combined multiple frames of color images according to a real-time arbitrary view generation method.

[0168] In the above example, the multiple frames of color images, i.e. videos, collected by the 12 industrial cameras are processed in the cloud, which can include encoding and compression processing, and then distributed to the edge media server. The edge media server of the disclosure deploys a real-time arbitrary view generation network for real-time virtual view interpolation.

[0169] The edge media server combines two adjacent videos received from the 12 videos two by two, a total of 11 pairs of left and right views, and then acquires the corresponding preset number of embedded optical flow and mask matrices for each two-by-two combination.

[0170] S104, adaptively splicing the multiple frames of color images and their preset number of embedded optical flow and mask matrices for each two-by-two combination to form multiple composite images in the spatial domain.

[0171] In some possible embodiments, S104 specifically includes S201-S203.

[0172] S201, performing channel dimension synthesis processing on the corresponding preset number of embedded optical flow and mask matrices of each two-by-two combined multiple frames of color images to determine a three-channel matrix corresponding to each two-by-two combined multiple frames of color images.

[0173] Wherein, the embedded optical flow is 2 channels, the mask matrix is 1 channel, after channel dimension synthesis processing, a three-channel matrix is formed, the resolution of which is consistent with the resolution of the two two-combined multi-frame color images.

[0174] S202, spatial domain synthesis processing is performed on each two two-combined multi-frame color image and the three-channel matrix corresponding to each two two-combined multi-frame color image to determine a synthesis image.

[0175] Wherein, the two two-combined multi-frame color images are respectively placed at the upper left and the upper right, and the three-channel matrix is placed at the lower left. The lower right part of the synthesis image can be optionally added with a texture image of other view angle other than the two two-combined ones and the embedded optical flow and mask matrix.

[0176] S203, the multi-frame color images of different view angles and the embedded optical flow and mask matrix of all two two-combined ones are all subjected to resolution reduction processing to form a total synthesis image.

[0177] S105, the multiple synthesis images are encoded, divided into video segments in time domain and transmitted through HLS protocol.

[0178] In some possible embodiments, S105 specifically includes S301 to S303.

[0179] S301, each synthesis image is encoded and then time-sliced in time sequence;

[0180] S302, each slice is divided into fixed time size to form multiple video segments distributed in time domain and space domain;

[0181] S303, the video segments are transmitted by using a standard streaming media transmission protocol (such as HLS, etc.).

[0182] As an example, a spatio-temporal segmentation method is used to cope with the multi-user scenario. The multiple synthesis image streams are divided into a series of video slices in time domain and space domain. Specifically, 12 adaptively organized synthesis images form a slice in space domain, and then each slice is divided in time sequence. This embodiment uses an edge media server as a media resource repository, wherein the server load is independent of the number of clients, and the operation of the client will be described separately. As for content transmission, the HTTP Adaptive Streaming transmission protocol divides the video content into video slices with the same length for transmission, the number of frames of each video slice is an integer multiple of the encoding GOP size, and the first frame is an all-intra reference frame (I frame), so that each video slice can be independently decoded. The HLS protocol is one of them, which is very suitable for transmission of time-space video slices.

[0183] S106, different clients select and download target composite images through interaction, and generate complete virtual views at the viewing angle of the user in real time according to the real-time arbitrary viewing angle generation method.

[0184] In a possible embodiment, S106 specifically comprises S401 to S404.

[0185] S401, a global viewing angle index is adopted to download a composite image corresponding to the viewing angle of the user and a total composite image;

[0186] S402, within a video slice time, the viewing angle of the user falls within the downloaded composite image, a complete virtual view at the viewing angle of the user is generated in real time according to the real-time arbitrary viewing angle generation method by using the composite image;

[0187] S403, within a video slice time, the viewing angle of the user falls outside the downloaded composite image, a complete virtual view at the viewing angle of the user is generated in real time according to the real-time arbitrary viewing angle generation method by using the total composite image;

[0188] S404, after a video slice time, the user downloads a complete virtual video at the viewing angle and the total composite image.

[0189] As an example, users of different clients select and download viewing angle composite images to be viewed through interaction, and generate complete virtual views at the viewing angle in real time. In this embodiment, when a user of a different client requests a corresponding view point in an interactive manner, that is, a complete virtual view at the viewing angle, an operation of searching for a corresponding view index from a global lookup table (corresponding to the global viewing angle index mentioned above, containing the position information of different viewing angles in a composite image, and the corresponding viewing angle tile can be found through the lookup table) will be performed. If the required view point is in the current composite image, the viewing angle extraction unit extracts the corresponding viewing angle tile bit stream, and decodes it through the video slice decoding unit, and if it is not in the current composite image, the corresponding viewing angle tile bit stream will be extracted in the total composite image. The interaction of the user directly determines the next downloaded composite image video segment, that is, the composite image closest to the required viewing angle. The video segment decoding unit is responsible for decoding the extracted bit stream into YUV format, and then playing by the video player based on OpenGL.

[0190] Based on the same concept, the present disclosure also provides a real-time arbitrary viewing angle generation system, Figure 7 is a block diagram of a real-time arbitrary viewing angle generation system according to an exemplary embodiment. Referring to Figure 7The real-time arbitrary perspective generation system 100 includes: an acquisition module 110 , a first determination module 120 , a second determination module 130 , a third determination module 140 , and a fourth determination module 150 .

[0191] An acquisition module 110 is configured to acquire multiple color images of the same calibration object from different perspectives, wherein the multiple color images of the same calibration object from different perspectives include a first color texture image from a first perspective and a second color texture image from a second perspective;

[0192] A first determining module 120 is configured to input the first color texture image and the second color texture image into a preset dense optical flow field estimation network to determine a bidirectional optical flow;

[0193] A second determining module 130 is configured to input the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder to determine an embedded optical flow and a mask matrix of a preset logarithm;

[0194] a third determining module 140, configured to determine, based on a third perspective input by a user and the preset logarithm of embedded optical flows, a first preset number of first optical flows from the first color texture image to the virtual view of the third perspective, and a first preset number of second optical flows from the second color texture image to the virtual view of the third perspective;

[0195] a fourth determination module, configured to perform forward warping on the first color texture image and the second color texture image according to the first preset number of first optical flows and the first preset number of second optical flows, and to fuse the images after the forward warping according to the mask matrix to determine a complete virtual view at the third perspective.

[0196] Regarding the embodiment of the above-mentioned real-time arbitrary perspective generation system, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0197] Figure 8 The figure is a schematic diagram showing the architecture of a free-viewpoint video generation system according to an exemplary embodiment.

[0198] Based on the same concept, Figure 8 As shown, the present disclosure also provides a free-viewpoint video generation system, which mainly includes an acquisition module, a cloud processing and content distribution network module, an edge server module and a client module.

[0199] The acquisition server module is used to obtain multiple frames of color images of the same calibration object from different perspectives;

[0200] a cloud processing and content distribution network module for combining the multiple frames of color images two by two and distributing them to the edge server module by a content distribution network;

[0201] an edge server module for obtaining a preset number of embedded optical flow and mask matrices for each of the two-by-two combined multiple frames of color images in real time; adaptively stitching each of the two-by-two combined multiple frames of color images and the preset number of embedded optical flow and mask matrices thereof to form a plurality of composite images in a spatial domain; and encoding the plurality of composite images, dividing them into video segments in a time domain, and transmitting them through an HLS protocol;

[0202] a client module for different clients to select target composite images through interaction and download, and to synthesize complete virtual views at viewing angles of users in real time.

[0203] The functions of each module can be found in the above-mentioned embodiments of the free-view video method, and will not be described here again.

[0204] In actual applications, when building the free-view video generation system of the embodiment, a pipeline design mode can be used, and a first-in-first-out queue data structure can be used to realize data transmission and thread separation, so that the units in the system are highly parallel executed, so that the overall delay bottleneck of the system is only limited to the module with the longest time consumption, thereby achieving good real-time performance, and in addition, the performance of the entire system can be greatly improved through heterogeneous computing, including CPU and GPU.

[0205] Table 1 Performance test results of the free-view video system

[0206]

[0207] Table 1 is a performance test result table of the pipeline of the free-view video generation system of the embodiment of the disclosure. The delay of each unit of the free-view video generation system and the average delay of 1000 frames are tested, as shown in Table 1, the free-view video generation system of the embodiment of the disclosure achieves good real-time performance. It should be noted that only two clients, PC and mobile phone, are used to test the client of the embodiment, but the system can support theoretically any number of clients to access.

[0208] Through the above technical solution, the free-view video generation system built based on this method decouples the number of accessed users and the load of the edge server, so that a single edge server can provide personalized free-view video services for multiple users.

[0209] Those skilled in the art will appreciate that embodiments of the disclosure can be provided as a method, or as a computer program product. Accordingly, the disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the disclosure can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0210] The disclosure is described in reference to the flow diagrams and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the disclosure. It should be understood that each flow and / or block in the flow diagrams and / or block diagrams, and combinations of flows and / or blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified by the flow or flows and / or block or blocks.

[0211] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow diagrams and / or block diagrams flow or flows and / or block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified by the flow or flows and / or block or blocks.

[0212] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow diagrams and / or block diagrams flow or flows and / or block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified by the flow or flows and / or block or blocks.

[0213] The specific embodiments of the present disclosure have been described. It is to be understood that the disclosure is not limited to the specific embodiments described above and various modifications, alterations or permutations can be made to the disclosure by those skilled in the art without departing from the scope of the present disclosure. The above preferred features can be combined in any manner possible, provided that they are not mutually exclusive.

Claims

1. A real-time arbitrary perspective generation method, characterized in that: include: Acquire multiple frames of color images of the same calibration object from different perspectives, wherein the multiple frames of color images of the same calibration object from different perspectives include a first color texture image from a first perspective and a second color texture image from a second perspective; Inputting the first color texture image and the second color texture image into a preset dense optical flow field estimation network to determine a bidirectional optical flow; Inputting the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder to determine an embedded optical flow and a mask matrix of a preset logarithm; Determining, based on a third perspective input by a user and the preset logarithm of embedded optical flows, a first preset number of first optical flows from the first color texture image to a virtual view of the third perspective, and a first preset number of second optical flows from the second color texture image to the virtual view of the third perspective; performing forward warping on the first color texture image and the second color texture image respectively using the first preset number of first optical flows and the first preset number of second optical flows, and fusing the images after the forward warping using the mask matrix to determine a complete virtual view at the third perspective; The step of inputting the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder to determine an embedded optical flow and a mask matrix of a preset logarithm includes: performing feature extraction processing on the first color texture image and the second color texture image to determine image features of the color images at different resolution scales of the first and second perspectives; Performing a first warping process on the image features of the color images of the first viewing angle and the second viewing angle using an initial optical flow to determine image features after the first warping process; Aligning and splicing the image features after the first warping process with the original image features at the warping end point in sequence to determine spliced ​​image features; performing encoding and decoding on the stitched image features in sequence to determine a second preset number of optical flow residuals, wherein the value of the second preset number is the same as the value of the preset logarithm; The initial optical flow is optimized according to the second preset number of optical flow residuals to determine a preset number of embedded optical flows and a mask matrix.

2. The method according to claim 1, characterized in that The determining, based on the third perspective input by the user and the preset logarithm of embedded optical flows, a first preset number of first optical flows from the first color texture image to the virtual view of the third perspective and a first preset number of second optical flows from the second color texture image to the virtual view of the third perspective includes: in, represents the first optical flow, p represents the third viewing angle, represents the embedded optical flow from the first view to the second view, represents the second optical flow, represents the embedded optical flow from the second view to the first view.

3. The method according to claim 1, characterized in that The fusion process includes a first fusion process and a second fusion process; The method further comprises: performing forward warping on the first color texture image and the second color texture image respectively using the first preset number of first optical flows and the first preset number of second optical flows, and fusing the images after the forward warping using the mask matrix to determine a complete virtual view at the third perspective, including: Using a forward warping operator, forward warping the first color texture image to the third perspective according to the first preset number of first optical flows, to determine a first preset number of first-side virtual third color texture images; Using the forward warping operator, forward warping the second color texture image to the third perspective according to the first preset number of second optical flows, to determine a first preset number of second-side virtual third color texture images; performing the first fusion process on the first preset number of first-side virtual third color texture images and the first preset number of second-side virtual third color texture images, respectively, to determine a first-side fusion image and a second-side fusion image after the first fusion process; The first side fusion image and the second side fusion image are subjected to the second fusion process according to the mask matrix to determine a complete virtual view at the third viewing angle.

4. The method according to claim 3, characterized in that performing the second fusion process on the first side fusion image and the second side fusion image according to the mask matrix to determine a complete virtual view at the third perspective, include: I p =M⊙I left→p +(1-M)⊙I right→p Among them, I p represents the complete virtual image at the third viewing angle, M represents the mask matrix, I left→p represents the first side fusion image, I right→p represents the second side fused image, represents the first side virtual third color texture image, represents the second side virtual third color texture image, and k represents the first preset number.

5. A real-time arbitrary perspective generation system, characterized in that: include: an acquisition module, configured to acquire multiple color images of the same calibration object from different perspectives, wherein the multiple color images of the same calibration object from different perspectives include a first color texture image from a first perspective and a second color texture image from a second perspective; a first determining module, configured to input the first color texture image and the second color texture image into a preset dense optical flow field estimation network to determine a bidirectional optical flow; a second determining module, configured to input the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder, and determine an embedded optical flow and a mask matrix of a preset logarithm; a third determining module, configured to determine, based on a third perspective input by a user and the preset logarithm of embedded optical flows, a first preset number of first optical flows from the first color texture image to the virtual view of the third perspective, and a first preset number of second optical flows from the second color texture image to the virtual view of the third perspective; a fourth determining module, configured to perform a forward warping process on the first color texture image and the second color texture image according to the first preset number of first optical flows and the first preset number of second optical flows, and to fuse the images after the forward warping process according to the mask matrix to determine a complete virtual view at the third perspective; The second determining module inputs the first color texture image, the second color texture image, and the bidirectional optical flow into an optical flow embedder to determine an embedded optical flow and a mask matrix of a preset logarithm, including: performing feature extraction processing on the first color texture image and the second color texture image to determine image features of the color images at different resolution scales of the first and second perspectives; Performing a first warping process on the image features of the color images of the first viewing angle and the second viewing angle using an initial optical flow to determine image features after the first warping process; Aligning and splicing the image features after the first warping process with the original image features at the warping end point in sequence to determine spliced ​​image features; performing encoding and decoding on the stitched image features in sequence to determine a second preset number of optical flow residuals, wherein the value of the second preset number is the same as the value of the preset logarithm; The initial optical flow is optimized according to the second preset number of optical flow residuals to determine a preset number of embedded optical flows and a mask matrix.

6. A method for generating a free-viewpoint video, characterized in that: include: Acquire multiple frames of color images of the same calibration object from different perspectives; Combining the multiple color image frames in pairs; Each of the two-by-two combination of multiple frames of color images The method according to claim 1 obtains the preset logarithmic embedded optical flow and mask matrix in real time; Adaptively splicing the pairwise combination of the multi-frame color images and the preset logarithmic embedded optical flows and mask matrices to form a plurality of composite images in the spatial domain; Encoding the multiple composite images, dividing them into video segments in the time domain, and transmitting them through the HLS protocol; Different clients interactively select target synthetic images and download them, and synthesize a complete virtual view at the user's viewing angle in real time according to the method described in any one of claims 2-3.

7. The method according to claim 6, characterized in that Adaptively splicing each of the two-by-two combined multi-frame color images and their preset logarithmic embedded optical flows and mask matrices to form multiple synthetic images in the spatial domain includes: Performing channel-dimensional synthesis processing on the preset logarithmic embedded optical flows and mask matrices corresponding to each of the two-by-two combinations of the multi-frame color images to determine a three-channel matrix corresponding to each of the two-by-two combinations of the multi-frame color images; Performing spatial domain synthesis processing on each of the two-by-two combinations of the multi-frame color images and the three-channel matrix corresponding to each of the two-by-two combinations of the multi-frame color images to determine a synthesized image; The multiple frames of color images at different viewing angles and all pairwise combinations of embedded optical flows and mask matrices are subjected to resolution reduction processing to form a total composite image.

8. The method according to claim 7, characterized in that The different clients interactively select and download target synthetic images, and synthesize a complete virtual view at the user's viewing angle in real time according to the method according to any one of claims 2 to 3, comprising: Using a global viewpoint index, download the composite image and the total composite image corresponding to the viewpoint viewed by the user; Within a video slice time, the viewing angle of the user falls within the downloaded synthetic image, and the synthetic image is used to synthesize a complete virtual view at the viewing angle of the user in real time according to any one of claims 3-4; During a video slice time, the viewing angle of the user falls outside the downloaded synthetic image, and the total synthetic image is used to synthesize a complete virtual view at the viewing angle of the user in real time according to any one of claims 3-4; After a video slicing time, the user downloads the complete virtual video and the total composite image at the viewing angle.

9. A free viewpoint video generation system, characterized in that: include: The acquisition server module is used to obtain multiple frames of color images of the same calibration object from different perspectives; A cloud processing and content distribution network module, configured to combine the multiple frames of color images in pairs and distribute the images to the edge server module via a content distribution network; The edge server module is used to transmit each of the two-by-two combination of multiple frames of color images The method according to claim 1 obtains a preset logarithmic embedded optical flow and mask matrix in real time; adaptively splices each of the two-by-two combinations of multiple color frames and their preset logarithmic embedded optical flow and mask matrix to form multiple composite images in the spatial domain; encodes the multiple composite images, divides them into video segments in the temporal domain, and transmits them via the HLS protocol; The client module is used for different clients to interactively select target synthetic images and download them, and to synthesize a complete virtual view at the user's viewing angle in real time according to the method described in any one of claims 2-3.

Citation Information

Patent Citations

  • Multi-user free view angle video method and system based on real-time virtual view angle interpolation

    CN114897681A

  • Method and device for generating video intermediate frame

    CN115065796A