Video generation method, device, equipment and storage medium
By performing depth estimation and edge recognition on the original image, generating foreground and background expansion pixels, and performing three-dimensional modeling, the problems of large amount of calculation and low accuracy in the prior art are solved, and efficient and widely applicable video generation is achieved.
Patent Information
- Application Number
- CN202311245123.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-09-25
AI Technical Summary
The prior art video generation method under image guidance is large in calculation, slow in processing, and the accuracy of the semantic segmentation model is not high in untrained scenarios.
By performing depth estimation of the original image, identifying foreground and background edge pixels, and expanding out foreground and background expansion pixels, performing three-dimensional modeling, and finally generating video.
It reduces the amount of calculation, improves the efficiency of video generation, has a wider scope of application, avoids the accuracy of semantic segmentation models in untrained scenarios, and the generated video is closer to the real scene.
Smart Images

Figure CN117354480B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, virtual reality, and deep learning, and can be applied to scenarios such as artificial intelligence-based content generation and the metaverse. In particular, it relates to a video generation method, apparatus, device, and storage medium. Background Art
[0002] In an image-guided video generation method, a dynamic video can be generated based on a static image. The static images used for video generation can cover a wide range of real-scene images and some stylized generated images.
[0003] Such a video generation method can be used in a text-to-video system. For example, based on text information, an image can be generated through an image generation model, and then a video corresponding to the image can be produced through methods such as three-dimensional camera movement. Summary of the Invention
[0004] The present disclosure provides a video generation method, apparatus, device, and storage medium.
[0005] According to one aspect of the present disclosure, a video generation method is provided, including:
[0006] Performing depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image;
[0007] Performing edge recognition based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels;
[0008] Generating foreground extended pixels obtained by expanding from the foreground edge pixels to the periphery according to the original image, and generating background extended pixels obtained by expanding from the background edge pixels to the periphery;
[0009] Performing three-dimensional modeling based on the pixels in the original image, the foreground extended pixels, and the background extended pixels;
[0010] Generating a video based on the three-dimensional model obtained by modeling.
[0011] According to another aspect of the present disclosure, a video generation apparatus is provided, including:
[0012] An estimation module for performing depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image;
[0013] An identification module for performing edge recognition based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels;
[0014] An expansion module for generating foreground expansion pixels obtained by expanding outward from the foreground edge pixels and generating background expansion pixels obtained by expanding outward from the background edge pixels based on the original image;
[0015] A modeling module for performing three-dimensional modeling based on the pixels in the original image, the foreground expansion pixels, and the background expansion pixels;
[0016] A generation module for generating a video based on the three-dimensional model obtained by modeling.
[0017] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect embodiment of the present disclosure.
[0018] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect embodiment of the present disclosure.
[0019] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the first aspect embodiment of the present disclosure when executed by a processor.
[0020] The video generation method, apparatus, device, and storage medium provided by the present disclosure, after obtaining the depth of each pixel in the original image by performing depth estimation on each pixel in the original image, perform edge recognition based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels. Based on the original image, generate foreground expansion pixels obtained by expanding outward from the foreground edge pixels, and generate background expansion pixels obtained by expanding outward from the background edge pixels. Then perform three-dimensional modeling based on the pixels in the original image, the foreground expansion pixels, and the background expansion pixels, and generate a video based on the three-dimensional model obtained by modeling.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a schematic flowchart of a video generation method provided by an embodiment of the present disclosure;
[0024] Figure 2 Schematic flowchart of another video generation method provided by an embodiment of the present disclosure;
[0025] Figure 3 Schematic diagram of foreground and background edge division;
[0026] Figure 4 Schematic flowchart of a possible video generation method;
[0027] Figure 5 Schematic structural diagram of a video generation device 500 provided by an embodiment of the present disclosure;
[0028] Figure 6 Schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure is shown. Detailed implementation manners
[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] In related technologies, a hierarchical representation method of depth maps (LDI, layered depth image) is used to estimate monocular depth information of an input image, and each pixel is clustered and divided into different levels according to its depth value. For the missing pixels in each level image, the surrounding pixels are diffused to the missing pixels through a heuristic method. Thus, a three-dimensional scene is constructed, and a video is rendered based on the camera's moving lens trajectory.
[0031] However, this method in related technologies requires dividing multiple levels and predicting missing pixels for each level, resulting in a slow processing speed and a large amount of computation.
[0032] To address the problem of low efficiency in related technologies, in the present disclosure, a single real scene or a stylized original image generated by AI is input. First, monocular depth estimation is performed on it, and then edge pixels, including foreground edge pixels and background edge pixels, are extracted based on the estimated depth. Then, corresponding regions, namely the region of foreground extended pixels and the region of background extended pixels, are diffused according to the foreground edge pixels and background edge pixels. The RGB values of the extended regions are drawn, and then a three-dimensional model is generated based on this, and a video is generated by adopting a dynamic moving lens method.
[0033] In the present disclosure, since the original image is only divided into a foreground layer and a background layer, and only the pixels extended from the foreground-background edge are drawn with RGB values, the overall process only requires two model predictions for drawing and depth estimation, greatly reducing the processing speed and the amount of computation compared to all existing methods.
[0034] Moreover, in the present disclosure, foreground edge pixels and background edge pixels are identified based on depth. Compared with the method of identifying edge pixels based on semantic segmentation, the applicable range is wider, avoiding the problem of low precision and recall rate of the semantic segmentation model in un-trained scenarios.
[0035] Figure 1 The flowchart of a video generation method provided by an embodiment of the present disclosure. The method provided by this embodiment can be executed by a video generation device, such as Figure 1 as shown, including:
[0036] Step 101, perform depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image.
[0037] Among them, the original image can be machine-generated or manually drawn, and this embodiment does not make any limitations in this regard.
[0038] For example: applied to a text-to-video system, taking paragraphs in lyrics or other constraint conditions as input, generating an image through an image generation model, using the generated image as the original image, and then outputting a video through the video generation method in this embodiment. Multiple videos generated from images can be further spliced into a long video such as an MV.
[0039] Another example: applied to a cover story, for the pictures recommended daily in the user's photo album, using the picture as the original image, and executing the video generation method in this embodiment to output the corresponding video for dynamic display on the album home page.
[0040] Another example: applied to the generation of e-commerce live posters, using a single-frame poster as the original image, and the camera coordinates of the last frame can be designed to be the same as those of the first frame, so that a loopable video can be output and used as a background for loop playback during e-commerce live broadcasts.
[0041] Optionally, the object presented in the original image includes a foreground and a background, and there are some depth differences between the object as the foreground and the object as the background.
[0042] In the case where each pixel point in the original image does not carry a depth value, the original image can be depth-estimated by means of monocular vision.
[0043] When each pixel in the original image carries a depth value, the depth value carried by each pixel can be directly adopted. Alternatively, the carried depth value is fused with the depth value obtained by estimating the depth of the original image in a monocular vision manner. In this embodiment, no specific limitation is imposed on the specific depth estimation method for obtaining the depth of each pixel. It should be noted that the pixels mentioned in this embodiment and subsequent embodiments may be a single pixel or a pixel unit composed of several pixels combined together.
[0044] Step 102: Perform edge recognition based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels.
[0045] Among them, the foreground edge pixels refer to the pixels belonging to the edge of the foreground object.
[0046] Similarly, the background edge pixels refer to the pixels belonging to the edge of the background object.
[0047] Optionally, for the original image, edge recognition is performed according to the change of pixel depth. This is because the foreground and background objects are usually located on different depth layers and there will be a depth difference between them. Based on this, the foreground edge pixels and background edge pixels can be recognized. The foreground edge pixels have a depth difference from the adjacent background edge pixels and do not have a large depth difference from the adjacent foreground pixels. Similarly, the background edge pixels do not have a large depth difference from the adjacent background pixels and have a depth difference from the adjacent foreground edge pixels.
[0048] Step 103: Generate foreground extended pixels obtained by expanding from the foreground edge pixels to the periphery according to the original image, and generate background extended pixels expanded from the background edge pixels to the periphery.
[0049] Among them, the foreground extended pixels are the pixels located outside the foreground that are invisible due to the perspective in the original image. The background extended pixels are the pixels located outside the background that are invisible due to occlusion under the perspective of the original image.
[0050] Optionally, according to the original image, predict the pixel values of the pixels that are invisible due to the perspective in the original image, which can specifically be RGB values. And according to the original image, predict the background pixels that are invisible due to foreground occlusion.
[0051] Step 104: Perform 3D modeling based on the pixels, the foreground extended pixels, and the background extended pixels in the original image.
[0052] Optionally, the model for performing 3D modeling is a 3D mesh model using triangular patches, or it can be other models. No limitation is imposed on this in this embodiment.
[0053] Step 105, generate a video based on the three-dimensional model obtained by modeling.
[0054] Optionally, in the three-dimensional space where the three-dimensional model is located, based on the camera movement trajectory, locate a number of camera positions; determine the imaging diagrams of the three-dimensional model from the perspectives of each of the camera positions; and arrange the corresponding imaging diagrams according to the moments when the camera is at each of the camera positions in the camera movement trajectory to obtain the generated video.
[0055] In this embodiment, after obtaining the depths of the pixels in the original image by performing depth estimation on each pixel in the original image, edge recognition is performed according to the depths of the pixels in the original image to obtain foreground edge pixels and background edge pixels. According to the original image, foreground extended pixels obtained by expanding outward from the foreground edge pixels and background extended pixels obtained by expanding outward from the background edge pixels are generated. Then, three-dimensional modeling is performed according to the pixels, foreground extended pixels, and background extended pixels in the original image, and a video is generated based on the three-dimensional model obtained by modeling. Since only the depths of the pixels in the original image need to be edge-recognized and divided into two categories, namely foreground edge pixels and background edge pixels, the number of divided levels is reduced compared with the related LDI technology. Moreover, in the present disclosure, only the foreground extended pixels obtained by expanding outward from the foreground edge pixels and the background extended pixels obtained by expanding outward from the background edge pixels need to be complemented according to the original image, and the amount of calculation is reduced compared with the method in the related technology that needs to complement the edges of the existing pixels and missing pixels and the missing pixels in different levels, thereby improving the efficiency of video generation.
[0056] As a video generation method guided by pictures, 3D camera movement can cover a wide range of real-scene images and some stylized generated images, and transform a static image into a dynamic camera movement video of any length. To clearly illustrate the video generation process in the scenario of 3D camera movement, another video generation method is provided. Figure 2 It is a schematic flowchart of another video generation method provided by the embodiments of the present disclosure. The method provided by this embodiment can be executed by a video generation device, such as Figure 2 shown, and includes:
[0057] Step 201, input the original image.
[0058] Step 202, perform depth estimation on each pixel in the original image to obtain the depths of the pixels in the original image.
[0059] Optionally, a depth model can be used for depth estimation, and the depth model can estimate the depth of pixels based on monocular depth estimation technology.
[0060] For the input original image I, a depth model is used to generate a single-channel depth map D for depth estimation. The value of each pixel in the depth map D indicates the depth, that is, the depth value, representing the value of the Z-axis of the object it shows in the world coordinate system. The larger the depth value, the more forward the pixel is and the more likely it is to be a foreground pixel; conversely, it is a background pixel.
[0061] Step 203: Based on the depth, perform depth edge extraction to extract foreground edge pixels and background edge pixels from the original image.
[0062] Optionally, according to the depth of each pixel in the original image, an edge detection algorithm with an adaptive threshold is used to determine an edge map. The pixel value of each pixel in the edge map is used to indicate whether the corresponding pixel in the original image is an edge pixel. For adjacent edge pixels in the edge map, according to the depth difference of the corresponding pixels in the original image, the pixels in the original image are divided into foreground edge pixels and background edge pixels. Only the foreground and background need to be divided for the edge part, which improves the efficiency of foreground and background recognition to a certain extent.
[0063] As a possible implementation, perform Canny edge detection with an adaptive threshold on the depth map D to achieve depth edge extraction and obtain an edge map C. The edge map C is a binary map. The pixel value greater than zero in the edge map C indicates that there is a depth mutation at the pixel position, which may be the edge between the foreground and the background, so as to extract foreground edge pixels and background edge pixels from the original image accordingly. The pixel value equal to zero in the edge map C indicates that the pixel position does not belong to the edge.
[0064] Furthermore, before and after extracting foreground edge pixels and background edge pixels in step 203, several possible correction processes can be performed:
[0065] As the first possible correction method, after determining the edge map by using an edge detection algorithm with an adaptive threshold according to the depth of each pixel in the original image and before dividing the pixels in the original image into foreground edge pixels and background edge pixels, several connected regions can be determined in the edge map according to the pixel values of the pixels in the edge map. The pixel value of each pixel in the connected region indicates that it belongs to an edge pixel. Remove the connected regions with the number of pixel points less than the number threshold from the several connected regions to form as large connected regions as possible, avoiding interference caused by misrecognition and affecting the visual effect.
[0066] For example: for the edge map C, identify the 8-neighborhood connected regions, count the number of pixels in each connected region, and obtain the connected regions with the number of pixels less than the number threshold a (for example, the value range of a can be configured between 8 and 12, and the typical value is 10). In the edge map C, set the pixels in the connected regions with the number of pixels less than the number threshold a to 0.
[0067] As a second possible correction method, after determining the edge map using an edge detection algorithm with an adaptive threshold based on the depth of each pixel in the original image, before dividing the pixels in the original image into foreground edge pixels and background edge pixels, several connected regions can also be determined in the edge map according to the pixel values of the pixels in the edge map, where the pixel values of the pixels in the connected regions indicate belonging to edge pixels; merge two connected regions in the several connected regions with a distance less than the distance threshold to form a connected region as large as possible, avoiding interference caused by misidentification and affecting the visual effect.
[0068] For example: count the set of pixels {p_e} with the number of pixels greater than 1 and equal to 1 in all 8-neighborhoods in the edge map C. These pixel points are the end pixels of all edge lines in the edge map C. If the distance between two end pixels in {p_e} is less than the distance threshold b (for example, the value range of b can be configured between 15 and 25 pixel distances, and the typical value is 20 pixel distances), and these two end pixels do not belong to the same connected region, then draw a straight line between these two points in the edge map C, that is, connect and merge them into the same connected region, that is, connect adjacent lines into a straight line. This can effectively alleviate the problem of discontinuous depth edges detected by edge detection.
[0069] It should be noted that the first possible correction method and / or the second possible correction method described above can be executed for the edge map C. Take the edge map after correction as the edge map C1. Similar to the edge map C, the edge map C1 is also a binary map.
[0070] Divide the foreground edge pixels and background edge pixels based on the edge map C1.
[0071] Optionally, for each pixel p with a pixel value greater than 0 in the edge map C1, Figure 3 For the schematic diagram of foreground and background edge division, such as Figure 3If the coordinates of the pixel p shown are (0, 0), traverse its 4 adjacent pixels q ∈ {p_(-1, 0), p_(0, +1), p_(+1, 0), p_(0, -1)} respectively, as well as the pixel q2 adjacent to q on the extension line pointed by p, where q2 = q + Δi, and the pixel p2 adjacent to p on the reverse extension line, where p2 = p - Δi, where Δi is the offset relative to the pixel q or p. If (p + p2) > (q + q2), it indicates that p is a foreground edge and q is a background edge. Add p to the foreground edge pixel set N and add q to the background edge pixel set F. It should be noted that Figure 3 Take Δi as one pixel as an example for illustration.
[0072] As a third possible correction method, identify the first target pixels that are the endpoints of the foreground edges from the set of foreground edge pixels; pass through the pixels in the set of foreground edge pixels to determine the paths connecting any two first target pixels; delete the pixels in the set of foreground edge pixels that have not passed through any path as redundant pixels.
[0073] For example: after determining the foreground edge pixel set N and the background edge pixel set F, for each pixel point p in the set N, count the number of pixels in the 8-neighborhood of the edge map C1 whose pixel values are greater than 1 and equal to 1. These points are the endpoint pixels of all the lines in N. Starting from the pixels in NE, traverse all the pixels in N with a 4-neighborhood breadth-first search to obtain all possible paths {P} from each pixel in NE to other pixels in NE. Only keep the points in the N set that are on the path {P}, and delete the points not on the path in N, thereby removing redundant foreground edge pixels.
[0074] As a fourth possible correction method, for any one of the background edge pixels, query multiple second target pixels in the neighborhood; determine the second target pixel with the minimum depth from the multiple second target pixels; in the case where the second target pixel with the minimum depth does not belong to the foreground edge pixels, add the second target pixel with the minimum depth to the set of background edge pixels to correct the interference caused by mis-identification.
[0075] For example: for the pixels in N, calculate the pixels in its 4-neighborhood, and determine the pixel with the minimum depth (i.e., the depth value recorded in the depth map D) that is not in the N set, and add it to the F set.
[0076] As a fifth possible correction method, for any one of the background edge pixels, query whether there are adjacent background edge pixels; in the case where there are no adjacent background edge pixels, delete the background edge pixel to correct the isolated pixels caused by mis-identification. For example: filter the pixels in the F set that are not adjacent to any other pixels in F.
[0077] As the sixth possible correction method, contour pixels on the periphery of edge pixels are determined in the edge map; the depth corresponding to the contour pixels is used to replace the depth of the enclosed edge pixels, ensuring an obvious depth mutation between foreground edge pixels and background edge pixels. Optionally, in order to determine the contour pixels on the periphery of edge pixels in the edge map, the edge pixels in the edge map can be dilated to obtain a first dilated map; the edge pixels in the first dilated map are dilated again to obtain a second dilated map; the edge pixels overlapping with the first dilated map are removed from the second dilated map, and the remaining edge pixels are used as the contour pixels.
[0078] For example: the edge map C1 is dilated three times to obtain Cd, then Cd is dilated once again to obtain Cd2, and Cd2 - C1 is used to obtain Cc, that is, the binary map of the edge pixels in Cd. For all pixels in Cc with values greater than 0, their adjacent pixels are traversed in a 4-neighborhood depth-first manner. For the pixels belonging to the N or F set encountered for the first time during the traversal, the depth value at the position of the starting pixel (the pixel point in Cc) of the traversal is used to replace the depth value at the current pixel position, thereby updating the depth of the edge pixels.
[0079] As the seventh possible correction method, the depth of the foreground edge pixels is increased according to the depth of the neighborhood pixels of the foreground edge pixels; the depth of the background edge pixels is decreased according to the depth of the neighborhood pixels of the background edge pixels. Ensure an obvious depth mutation between the foreground edge pixels and the background edge pixels.
[0080] For example: for each pixel p in N, the depth value of pixel p is replaced with the maximum depth value among the 9 neighborhood pixels including itself.
[0081] For another example: for each pixel p in F, the depth value of pixel p is replaced with the minimum depth value among the 9 neighborhood pixels including itself.
[0082] Step 204, based on the foreground edge pixels and the background edge pixels, the depth of the foreground extended pixels on the periphery of the foreground edge pixel points is extended, and the depth of the background extended pixels on the periphery of the background edge pixels is extended according to the depth of the background edge pixels.
[0083] Optionally, the foreground edge pixels and the background edge pixels can be added as elements to a queue, where the attribute information of the elements includes: the first coordinate of the pixel to which the element belongs, the foreground / background flag indicating whether the pixel belongs to the foreground or the background, the diffusion step number, and the second coordinate of the starting pixel of diffusion. Take out elements from the queue one by one. Whenever a target element is taken out, based on the foreground / background flag of the target element, the pixel at the corresponding first coordinate position in the foreground image or the background image is configured with a set non-zero value; and, if the diffusion step number of the target element is greater than zero, query the neighboring pixels of the target element in the original image, and use the neighboring pixels of the target element as new elements and add them to the queue. Wherein, the diffusion step number of the new element is the diffusion step number of the target element minus one, the second coordinate of the new element is the first coordinate of the target element, and the foreground / background flag of the new element is the same as that of the target element.
[0084] As a possible implementation manner, the diffusion step number of the background edge pixel to which the element belongs is the first step number; the diffusion step number of the foreground edge pixel to which the element belongs is the second step number; wherein, the second step number is greater than the first step number. Thus, the diffusion of the foreground edge pixels is more than that of the background edge pixels, and the background content is more presented by the foreground diffusion pixels after completion.
[0085] For example: A new foreground mask Mn and a background mask Mf can be initialized first. Both Mf and Mn are binary images, where all pixel values are 0; initialize a depth map Dm with all values being 0. The values in the depth map Dm are of floating-point type and represent the depth information of the pixels to be inpainted. And initialize a weight map Wf with all values being 0. The values in the weight map Wf are of floating-point type and represent the weights for weighted fusion of the pixels in the completed background area and the original background pixels.
[0086] Construct a queue Q to make the expansion process more orderly. Each element in the queue Q contains (p, is_near, step, root), where p is the pixel coordinate, is_near is used to mark whether it is a foreground pixel or a background pixel, step is the diffusion step number, and root is the starting pixel coordinate of traversal. For the pixels in N, the step is initially set to Sn (for example, it can take a value of 200), and for the pixels in F, the step is initially set to Sf (for example, it can take a value of 10).
[0087] Put all the pixels in N and F into the queue Q in order, and then loop to take elements from the queue. For each element q = (q[p], q[is_near], q[step], q[root]) taken out, set the depth map Dm[q[p]] to D[q[root]]. When q[is_near] is true, set Mn[q] to 1, otherwise set Mf[q] to 1, and at the same time set Wf[q] = q[step] / Sf. If q[step] > 0, visit the adjacent pixel positions of its 4-neighborhood q[p].neighbor: if q[p].neighbor[i] is not in Q and has not been visited, then deposit (q[p].neighbor[i], q[is_near], q[step] - 1, q[root]) into Q. Loop like this until Q is empty, and a foreground mask map diffused for Sn steps, a background mask map diffused for Sf steps, and the diffused depth map Dm of all traversed pixels can be obtained. Due to some pixel missing in the original image due to perspective occlusion, by extending with reference to the depth pixels of the pixel points in the original image, it is convenient for more complete and realistic scene modeling in subsequent 3D modeling.
[0088] Step 205, perform pixel value completion according to the pixel values of each pixel in the original image to determine the pixel values of foreground extended pixels and background extended pixels.
[0089] Optionally, synthesize a target mask map according to the foreground extended pixels and the background extended pixels. Input the original image and the mask map into a drawing model to determine the pixel values of each extended pixel in the target mask map according to the depth and pixel values of each pixel in the original image, and according to the depth of each extended pixel in the target mask map; wherein, the extended pixels include the foreground extended pixels and the background extended pixels. In this way, the invisible pixels in the original image can be complemented to a certain extent by referring to the original image, which is convenient for more complete and realistic scene modeling in subsequent 3D modeling.
[0090] As a possible implementation, the foreground extended pixels can be used as the foreground mask map, the background extended pixels can be used as the background mask map, and the foreground mask map and the background mask map are summed to obtain the target mask map. At the same time, both the foreground and the background are predicted, so that both the foreground and the background are complemented, avoiding the problem of a rather rigid connection between the foreground and the background. For example: sum Mf and Mn to get M, then use M as the mask map and the original image I as the input, and use the LAMA algorithm to complete the RGB channels of the non-zero regions in M to obtain the completed image R. The pixel values of this image and the original image are fused according to the weight map, so as to achieve a gradually transitional picture effect.
[0091] Step 206: Perform 3D modeling based on the pixels, foreground extended pixels, and background extended pixels in the original image to obtain a 3D model of the mesh.
[0092] Optionally, establish connection edges between the pixels in the original image and neighboring pixels; in the case where the two pixels connected by the connection edge are the foreground edge pixel and the background edge pixel respectively, delete the connection edge; establish connection edges between the foreground extended pixels and the background edge pixels with a depth matching that in the original image according to the depth of the foreground extended pixels; combine the connection edges to obtain a plurality of triangular patches in 3D space; render each triangular patch according to the pixel values of the pixels in the original image and according to the pixel values of the foreground edge pixels and the background edge pixels to obtain a 3D model that is more in line with the actual scene.
[0093] As a possible implementation, construct a graph structure model G, where each node G[w][h][c] in G represents a pixel, where w and h are the coordinate positions of the pixel in images or mask graphs such as I and M, and c represents the pixel point at the c-th layer of this pixel position. The levels of the pixel points should ensure that the depth values are arranged from large to small, that is, the pixels in the upper layer are more inclined to the foreground.
[0094] G[w][h][c] has several attributes {v, d, is_near, is_far, edgelist}, where v is the RGB value of this pixel point, d is the depth value of this pixel point, is_near is a flag indicating whether this pixel point is a foreground edge pixel, is_far is a flag indicating whether this pixel point is a background edge pixel, and edgelist is a list of the edges between this pixel and other pixels, restricting that the edges can only exist between adjacent pixels with a spatial coordinate of 4-neighborhood, and the layer number is not restricted.
[0095] First, establish connection edges between the pixels in the original image and neighboring pixels; in the case where the two pixels connected by the connection edge are the foreground edge pixel and the background edge pixel respectively, delete the connection edge. Initialize G for all pixels {p} in the original image I, p[v] = I[p], p[d] = D[p], if p is in N, then p[is_near] = true (True), if p is in F, then p[is_far] = True; for any 4-neighborhood pixel q of p, if p[is_near] = True and q[is_far] = True, or p[is_far] = True and q[is_near] = True, then there is no connection edge between p and q, otherwise generate a connection edge epq and add it to the connection edge set p[edgelist].
[0096] Furthermore, according to the depth of the foreground expanded pixels, connection edges are established between the background edge pixels that match the depth in the original image. For all non-zero pixel sets {p}n in Mn, they are inserted into G[w][h][0] according to their coordinates, that is, the first layer at the coordinate position, and then p[v]=R[p] and p[d]=Dm[p] are set. Then, traverse the 4-neighborhood pixel positions q. If q is in {p}n or q[is_far]=True, a connection edge epq is constructed between p and q and added to p[edgelist]. For all non-zero pixel sets {p}f in Mf, they are inserted into G[w][h][-1] according to their coordinates, that is, the last layer at the coordinate position, and then p[v]=R[p]*Wf[p]+I[P]*(1-Wf[p]) and p[d]=Dm[p] are set. Traverse the 4-neighborhood pixel positions q. If q is in {p}f or q[is_far]=True, a connection edge epq is constructed between p and q and added to p[edgelist].
[0097] Step 207, obtain the camera movement trajectory.
[0098] Optionally, if a video of f frames is to be generated, f camera coordinates pose need to be constructed, and the camera movement trajectory F={pose} is composed of f camera coordinates.
[0099] As a possible implementation, the camera movement trajectory is zoom in / zoom out. For example: pose={[1,0,0,0],[0,1,0,0],……,[0,0,cos((i / f)π / 4 - π / 4)*Z + X,0],[0,0,0,0]}, where i is the frame number, and Z and X are constants. It can present a variable-speed zoom in / zoom out movement.
[0100] As another possible implementation, the camera movement trajectory is horizontal rotation. For example: pose={[cos(-θ / 180π),0,-sin(-θ / 180π),-sin(σ*π / 4)*X],[0,1,0,-sin(σ*π / 4)*Y],[sin(-θ / 180π),0,cos(-θ / 180π),-sin(σ*π / 4)*Z],[0,0,0,0]}, where θ=-Z + 2Z*i / f, X, Y, and Z are constants.
[0101] As yet another possible implementation, the camera movement trajectory is a planar circle. For example: pose={[1,0,0,-cos(σ*π)*X],[0,1,0,-sin(σ*π)*X],[0,0,1,0],[0,0,0,0]}, where X, Y are constants.
[0102] It should be noted that those skilled in the art can also adjust to obtain new trajectories or design other trajectories based on this, and this embodiment does not limit this.
[0103] Step 208: Generate a video based on the camera movement trajectory and the three-dimensional model.
[0104] Optionally, in the three-dimensional space where the three-dimensional model is located, based on the camera movement trajectory, locate a number of camera positions; determine the imaging diagrams of the three-dimensional model from the perspectives of each camera position; and arrange the corresponding imaging diagrams according to the moments when the camera is at each camera position in the camera movement trajectory to obtain the generated video. Since multiple optional trajectories can be configured, the way of generating the video is more flexible and variable, meeting the needs of different users.
[0105] It can be seen that the video generation process can be divided into input images as Figure 4 shown, and then perform depth estimation on the input images, perform depth edge extraction based on the depth obtained from the depth estimation, expand the foreground edge pixels and background edge pixels extracted, and perform pixel completion. Based on the pixels in the original image and the completed pixels, construct a three-dimensional mesh model, and perform frame-by-frame rendering on the three-dimensional mesh model in combination with the camera trajectory to obtain the output video.
[0106] In this embodiment, after performing depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image, perform edge recognition according to the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels. According to the original image, generate foreground extended pixels extended from the foreground edge pixels to the periphery, and generate background extended pixels extended from the background edge pixels to the periphery. Then, perform three-dimensional modeling according to the pixels, foreground extended pixels, and background extended pixels in the original image, and generate a video based on the three-dimensional model obtained from the modeling. Since only the edges of each pixel in the original image need to be recognized and divided into two categories: foreground edge pixels and background edge pixels, the number of divided levels is reduced compared with the related LDI technology. Moreover, in the present disclosure, only the foreground extended pixels extended from the foreground edge pixels to the periphery and the background extended pixels extended from the background edge pixels to the periphery need to be completed according to the original image, and the calculation amount is reduced compared with the method in the related technology that needs to complete the edges of the existing pixels and missing pixels and the missing pixels at different levels, improving the efficiency of video generation.
[0107] Figure 5 The following is a schematic structural diagram of a video generation device 500 provided by an embodiment of the present disclosure. As Figure 5 shown, it includes: an estimation module 501, a recognition module 502, an expansion module 503, a modeling module 504, and a generation module 505.
[0108] An estimation module 501 for performing depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image.
[0109] An identification module 502 for performing edge identification based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels.
[0110] An expansion module 503 for generating foreground expansion pixels obtained by expanding from the foreground edge pixels to the periphery based on the original image, and generating background expansion pixels expanded from the background edge pixels to the periphery.
[0111] A modeling module 504 for performing three-dimensional modeling based on the pixels in the original image, the foreground expansion pixels, and the background expansion pixels.
[0112] A generation module 505 for generating a video based on the three-dimensional model obtained by modeling.
[0113] In some possible embodiments, the expansion module 503 includes:
[0114] An expansion unit for expanding to obtain the depth of the foreground expansion pixels outside the foreground edge pixel points according to the depth of the foreground edge pixels, and expanding to obtain the depth of the background expansion pixels outside the background edge pixels according to the depth of the background edge pixels;
[0115] A completion unit for determining the pixel values of the foreground expansion pixels and the pixel values of the background expansion pixels according to the pixel values of each pixel in the original image.
[0116] Optionally, the completion unit is used for:
[0117] Synthesizing a target mask map according to the foreground expansion pixels and the background expansion pixels;
[0118] Inputting the original image and the mask map into a drawing model to determine the pixel values of each expansion pixel in the target mask map according to the depth and pixel values of each pixel in the original image, and according to the depth of each expansion pixel in the target mask map; wherein, the expansion pixels include the foreground expansion pixels and the background expansion pixels.
[0119] Wherein, the completion unit synthesizing the target mask map according to the foreground expansion pixels and the background expansion pixels includes: using the foreground expansion pixels as a foreground mask map, using the background expansion pixels as a background mask map, and summing the foreground mask map and the background mask map to obtain the target mask map.
[0120] Optionally, the expansion unit is used for:
[0121] Add the foreground edge pixels and the background edge pixels as elements to the queue, where the attribute information of the elements includes: the first coordinate of the pixel to which the element belongs, the foreground / background flag indicating whether the pixel belongs to the foreground or the background, the diffusion step number, and the second coordinate of the starting pixel of diffusion;
[0122] Take elements from the queue one by one. Whenever a target element is taken out, configure the pixel at the corresponding first coordinate position in the foreground image or the background image as a set non-zero value based on the foreground / background flag of the target element; and,
[0123] If the diffusion step number of the target element is greater than zero, query the neighboring pixels of the target element in the original image, and add the neighboring pixels of the target element as new elements to the queue;
[0124] wherein, the diffusion step number of the new element is the diffusion step number of the target element minus one, the second coordinate of the new element is the first coordinate of the target element, and the foreground / background flag of the new element is the same as that of the target element.
[0125] Optionally, the diffusion step number of the pixel to which the element belongs being the background edge pixel is the first step number; the diffusion step number of the pixel to which the element belongs being the foreground edge pixel is the second step number; wherein, the second step number is greater than the first step number.
[0126] In some possible embodiments, the recognition module 502 includes:
[0127] An edge recognition unit, configured to determine an edge map by using an edge detection algorithm with an adaptive threshold according to the depth of each pixel in the original image, where the pixel value of each pixel in the edge map is used to indicate whether the corresponding pixel in the original image is an edge pixel;
[0128] A partitioning unit, configured to partition the pixels in the original image into foreground edge pixels and background edge pixels according to the depth difference of the corresponding pixels in the original image for adjacent edge pixels in the edge map.
[0129] Optionally, the recognition module 502 further includes:
[0130] A first correction unit, configured to determine a plurality of connected regions in the edge map according to the pixel values of each pixel in the edge map, where the pixel values of each pixel in the connected regions indicate belonging to edge pixels; remove the connected regions with the number of pixel points less than the number threshold from the plurality of connected regions.
[0131] Optionally, the recognition module 502 further includes:
[0132] A second correction unit, configured to determine a plurality of connected regions in the edge map according to the pixel values of the pixels in the edge map, where the pixel values of the pixels in the connected regions indicate belonging to edge pixels; and merge two connected regions in the plurality of connected regions whose distance is less than a distance threshold.
[0133] Optionally, the recognition module 502 further includes:
[0134] A third correction unit, configured to identify first target pixels that are foreground edge endpoints from the set of foreground edge pixels; determine a path connecting any two first target pixels through the pixels in the set of foreground edge pixels; and delete the pixels in the set of foreground edge pixels that are not passed through any path as redundant pixels.
[0135] Optionally, the recognition module 502 further includes:
[0136] A fourth correction unit, configured to query a plurality of second target pixels in the neighborhood of any one of the background edge pixels; determine the second target pixel with the minimum depth from the plurality of second target pixels; and add the second target pixel with the minimum depth to the set of background edge pixels when the second target pixel with the minimum depth does not belong to the foreground edge pixels.
[0137] Optionally, the recognition module 502 further includes:
[0138] A fifth correction unit, configured to query whether there are adjacent background edge pixels for any one of the background edge pixels; and delete the background edge pixel when there are no adjacent background edge pixels.
[0139] Optionally, the recognition module 502 further includes:
[0140] A sixth correction unit, configured to determine contour pixels outside the edge pixels in the edge map; and replace the depth of the edge pixels surrounded by the depth corresponding to the contour pixels.
[0141] Wherein, optionally, the sixth correction unit determines the contour pixels outside the edge pixels in the edge map, including: dilating the edge pixels in the edge map to obtain a first dilated map; dilating the edge pixels in the first dilated map again to obtain a second dilated map;
[0142] Removing the edge pixels overlapping with the first dilated map from the second dilated map to use the remaining edge pixels as the contour pixels.
[0143] In some possible embodiments, the depth of the foreground edge pixels is greater than the depth of the background edge pixels; based on this, the recognition module 502 further includes:
[0144] A seventh correction unit is configured to increase the depth of the foreground edge pixels according to the depth of the neighboring pixels of the foreground edge pixels, and decrease the depth of the background edge pixels according to the depth of the neighboring pixels of the background edge pixels.
[0145] In some possible embodiments, the modeling module 504 is configured to:
[0146] Establish connection edges between the pixels in the original image and their neighboring pixels;
[0147] Delete the connection edges when the two pixels connected by the connection edges are the foreground edge pixels and the background edge pixels respectively;
[0148] Establish connection edges between the foreground extended pixels and the background edge pixels with matching depths in the original image according to the depth of the foreground extended pixels;
[0149] Combine the connection edges to obtain a plurality of triangular patches in three-dimensional space;
[0150] Render each triangular patch according to the pixel values of the pixels in the original image, and according to the pixel values of the foreground edge pixels and the background edge pixels, to obtain a three-dimensional model.
[0151] In some possible embodiments, the generation module 505 includes:
[0152] Locate a number of camera positions in the three-dimensional space where the three-dimensional model is located based on the camera movement trajectory;
[0153] Determine the imaging diagrams of the three-dimensional model from the perspectives of each camera position;
[0154] Arrange the corresponding imaging diagrams according to the moments when the camera is at each camera position in the camera movement trajectory to obtain the generated video.
[0155] In this embodiment, after estimating the depth of each pixel in the original image to obtain the depth of each pixel in the original image, edge recognition is performed based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels. According to the original image, foreground extended pixels obtained by extending outward from the foreground edge pixels and background extended pixels obtained by extending outward from the background edge pixels are generated. Subsequently, 3D modeling is performed based on the pixels, foreground extended pixels, and background extended pixels in the original image, and video generation is performed based on the 3D model obtained by the modeling. Since it is only necessary to classify the edges of each pixel in the original image into two categories: foreground edge pixels and background edge pixels through edge recognition, the number of classification levels is reduced compared with the related LDI technology. Moreover, in the present disclosure, it is only necessary to complement the foreground extended pixels obtained by extending outward from the foreground edge pixels and the background extended pixels obtained by extending outward from the background edge pixels according to the original image. Compared with the related technology in which the edges of existing pixels and missing pixels in different levels, as well as the missing pixels, all need to be complemented, the amount of calculation is reduced, and the efficiency of video generation is improved.
[0156] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0157] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0158] As Figure 6 shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 801, the ROM 802, and the RAM 603 are connected to each other through a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.
[0159] Multiple components in device 600 are connected to I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0160] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the video generation method in any other suitable manner (e.g., by means of firmware).
[0161] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0163] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0165] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0166] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system or a server combined with a blockchain.
[0167] Among them, it should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0168] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0169] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video generation method, comprising: Performing depth estimation on each pixel in the original image to obtain the depth of each pixel in the original image; Performing edge recognition according to the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels; Generating foreground extended pixels obtained by extending outward from the foreground edge pixels and generating background extended pixels extending outward from the background edge pixels according to the original image; Performing 3D modeling according to the pixels in the original image, the foreground extended pixels, and the background extended pixels; Generating a video based on the 3D model obtained by modeling; The generating foreground extended pixels obtained by extending outward from the foreground edge pixels and generating background extended pixels extending outward from the background edge pixels according to the original image includes: Extending to obtain the depth of foreground extended pixels outside the foreground edge pixel points according to the depth of the foreground edge pixels, and extending to obtain the depth of background extended pixels outside the background edge pixels according to the depth of the background edge pixels; Taking the foreground extended pixels as a foreground mask map, taking the background extended pixels as a background mask map, and summing the foreground mask map and the background mask map to obtain a target mask map; Inputting the original image and the mask map into a drawing model to complete the RGB channel of non-zero regions in the target mask map according to the depth and pixel values of each pixel in the original image, and determining the pixel values of each extended pixel in the target mask map; wherein, the extended pixels include the foreground extended pixels and the background extended pixels.
2. The method according to claim 1, wherein, The extending to obtain the depth of foreground extended pixels outside the foreground edge pixel points according to the depth of the foreground edge pixels and extending to obtain the depth of background extended pixels outside the background edge pixels according to the depth of the background edge pixels includes: Adding the foreground edge pixels and the background edge pixels as elements to a queue, wherein the attribute information of the elements includes: the first coordinate of the pixel to which the element belongs, a foreground / background flag indicating whether the pixel belongs to the foreground or the background, the number of diffusion steps, and the second coordinate of the starting pixel of diffusion; Taking elements one by one from the queue. Whenever a target element is taken out, configuring the pixel at the corresponding first coordinate position in the foreground map or the background map as a set non-zero value based on the foreground / background flag of the target element; and If the number of diffusion steps of the target element is greater than zero, querying the neighboring pixels of the target element in the original image, and adding the neighboring pixels of the target element as new elements to the queue; wherein the number of diffusion steps of the new element is the number of diffusion steps of the target element minus one, the second coordinate of the new element is the first coordinate of the target element, and the foreground / background flag of the new element is the same as that of the target element.
3. The method according to claim 2, wherein, The number of diffusion steps of the pixel to which the element belongs being the background edge pixel is the first number of steps; the number of diffusion steps of the pixel to which the element belongs being the foreground edge pixel is the second number of steps; wherein the second number of steps is greater than the first number of steps.
4. The method according to any one of claims 1 to 3, wherein, Performing edge recognition based on the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels, including: Based on the depth of each pixel in the original image, determining an edge map using an edge detection algorithm with an adaptive threshold, where the pixel value of each pixel in the edge map is used to indicate whether the corresponding pixel in the original image is an edge pixel; For adjacent edge pixels in the edge map, dividing the pixels in the original image into foreground edge pixels and background edge pixels according to the depth difference of the corresponding pixels in the original image.
5. The method according to claim 4, wherein After determining the edge map using the edge detection algorithm with an adaptive threshold based on the depth of each pixel in the original image, further including: According to the pixel values of the pixels in the edge map, determining a plurality of connected regions in the edge map, where the pixel values of the pixels in the connected regions indicate belonging to edge pixels; Removing the connected regions with the number of pixel points less than the number threshold from the plurality of connected regions.
6. The method according to claim 4, wherein, After determining the edge map using the edge detection algorithm with an adaptive threshold based on the depth of each pixel in the original image, further including: According to the pixel values of the pixels in the edge map, determining a plurality of connected regions in the edge map, where the pixel values of the pixels in the connected regions indicate belonging to edge pixels; Merging two connected regions with a distance less than the distance threshold among the plurality of connected regions.
7. The method according to claim 4, wherein, After dividing the pixels in the original image into foreground edge pixels and background edge pixels for adjacent edge pixels in the edge map according to the depth difference of the corresponding pixels in the original image, further including: Identifying first target pixels serving as foreground edge endpoints from the set of foreground edge pixels; Determining a path connecting any two first target pixels passing through the pixels in the set of foreground edge pixels; Deleting the pixels in the set of foreground edge pixels that have not passed through any path as redundant pixels.
8. The method according to claim 4, wherein After dividing the pixels in the original image into foreground edge pixels and background edge pixels for adjacent edge pixels in the edge map according to the depth difference of the corresponding pixels in the original image, further including: Querying a plurality of second target pixels in the neighborhood for any one of the background edge pixels; Determining the second target pixel with the minimum depth from the plurality of second target pixels; In the case where the second target pixel with the minimum depth does not belong to the foreground edge pixels, adding the second target pixel with the minimum depth to the set of background edge pixels as a background edge pixel.
9. The method according to claim 4, wherein After dividing the pixels in the original image into foreground edge pixels and background edge pixels for adjacent edge pixels in the edge map according to the depth difference of the corresponding pixels in the original image, further including: Querying whether there are adjacent background edge pixels for any one of the background edge pixels; Deleting the background edge pixel in the case where there are no adjacent background edge pixels.
10. The method according to claim 4, wherein After determining the edge map using the edge detection algorithm with an adaptive threshold based on the depth of each pixel in the original image, further including: Determine the contour pixels outside the peripheral of the edge pixels in the edge map; Replace the depth of the enclosed edge pixels with the depth corresponding to the contour pixels.
11. The method according to claim 10, wherein, The determining the contour pixels outside the peripheral of the edge pixels in the edge map includes: Dilate the edge pixels in the edge map to obtain a first dilated map; Dilate the edge pixels in the first dilated map again to obtain a second dilated map; Remove the edge pixels overlapping with the first dilated map from the second dilated map, and use the remaining edge pixels as the contour pixels.
12. The method according to claim 4, wherein The depth of the foreground edge pixels is greater than the depth of the background edge pixels; After dividing the pixels in the original image into foreground edge pixels and background edge pixels according to the depth difference of the corresponding pixels in the original image for adjacent edge pixels in the edge map, it further includes: Adjust the depth of the foreground edge pixels according to the depth of the neighborhood pixels of the foreground edge pixels; Adjust the depth of the background edge pixels according to the depth of the neighborhood pixels of the background edge pixels.
13. The method according to claim 1, wherein The three-dimensional modeling according to the pixels in the original image, the foreground extended pixels, and the background extended pixels includes: Establish connection edges between the pixels in the original image and their neighborhood pixels; Delete the connection edges when the two pixels connected by the connection edges are the foreground edge pixels and the background edge pixels respectively; Establish connection edges between the foreground extended pixels and the background edge pixels with matching depths in the original image according to the depth of the foreground extended pixels; Combine the connection edges to obtain a plurality of triangular patches in three-dimensional space; Render each triangular patch according to the pixel values of the pixels in the original image, and according to the pixel values of the foreground edge pixels and the background edge pixels, to obtain a three-dimensional model.
14. The method according to any one of claims 1-3, wherein, The video generation based on the three-dimensional model obtained by modeling includes: In the three-dimensional space where the three-dimensional model is located, locate a number of camera positions based on the camera movement trajectory; Determine the imaging maps of the three-dimensional model from the perspectives of each camera position; Arrange the corresponding imaging maps according to the moments when the camera is at each camera position in the camera movement trajectory to obtain the generated video.
15. A video generation device, including: An estimation module for estimating the depth of each pixel in the original image to obtain the depth of each pixel in the original image; An identification module for performing edge identification according to the depth of each pixel in the original image to obtain foreground edge pixels and background edge pixels; An extension module for generating foreground extended pixels extended from the foreground edge pixels to the periphery according to the original image, and generating background extended pixels extended from the background edge pixels to the periphery; A modeling module for performing three-dimensional modeling according to the pixels in the original image, the foreground extended pixels, and the background extended pixels; A generation module for generating a video based on the three-dimensional model obtained by modeling; The extension module includes: An expansion unit for expanding to obtain the depth of foreground expansion pixels outside the perimeter of the foreground edge pixel according to the depth of the foreground edge pixel, and expanding to obtain the depth of background expansion pixels outside the perimeter of the background edge pixel according to the depth of the background edge pixel; A completion unit for determining the pixel values of the foreground expansion pixels and the pixel values of the background expansion pixels according to the pixel values of the pixels in the original image; The completion unit is configured to: Use the foreground expansion pixels as a foreground mask map, use the background expansion pixels as a background mask map, sum the foreground mask map and the background mask map to obtain a target mask map; Input the original image and the mask map into a drawing model to complete the RGB channels of the non-zero regions in the target mask map according to the depth and pixel values of the pixels in the original image, and determine the pixel values of the expansion pixels in the target mask map; wherein, the expansion pixels include the foreground expansion pixels and the background expansion pixels.
16. The device according to claim 15, wherein The expansion unit is configured to: Add the foreground edge pixels and the background edge pixels as elements to a queue, wherein the attribute information of the elements includes: the first coordinate of the pixel to which the element belongs, a foreground / background flag for indicating whether the pixel belongs to the foreground or the background, the number of diffusion steps, and the second coordinate of the starting pixel of the diffusion; Take elements one by one from the queue. Whenever a target element is taken out, configure the pixel at the corresponding first coordinate position in the foreground map or the background map as a set non-zero value based on the foreground / background flag of the target element; and, If the number of expansion steps of the target element is greater than zero, query the neighboring pixels of the target element in the original image, and add the neighboring pixels of the target element as new elements to the queue; wherein, the number of expansion steps of the new element is the number of diffusion steps of the target element minus one, the second coordinate of the new element is the first coordinate of the target element, and the foreground / background flag of the new element is the same as that of the target element.
17. The apparatus according to claim 16, wherein, The number of diffusion steps of the pixel to which the element belongs being the background edge pixel is the first number of steps; the number of diffusion steps of the pixel to which the element belongs being the foreground edge pixel is the second number of steps; wherein, the second number of steps is greater than the first number of steps.
18. The apparatus according to any one of claims 15 - 17, wherein, The recognition module includes: An edge recognition unit for determining an edge map using an edge detection algorithm with an adaptive threshold according to the depth of the pixels in the original image, wherein the pixel values of the pixels in the edge map are used to indicate whether the corresponding pixels in the original image are edge pixels; A partitioning unit for partitioning the pixels in the original image into foreground edge pixels and background edge pixels according to the depth difference of the corresponding pixels in the original image for adjacent edge pixels in the edge map.
19. The apparatus according to claim 18, wherein, The recognition module further includes: A first correction unit for determining a number of connected regions in the edge map according to the pixel values of the pixels in the edge map, wherein the pixel values of the pixels in the connected regions indicate belonging to edge pixels; removing the connected regions with the number of pixel points less than a number threshold from the number of connected regions.
20. The apparatus according to claim 18, wherein, The recognition module further includes: A second correction unit, configured to determine a plurality of connected regions in the edge map according to the pixel values of the pixels in the edge map, where the pixel values of the pixels in the connected regions indicate belonging to edge pixels; and merge two connected regions in the plurality of connected regions whose distance is less than a distance threshold.
21. The apparatus according to claim 18, wherein, The recognition module further includes: A third correction unit, configured to identify first target pixels serving as foreground edge endpoints from the set of foreground edge pixels; determine a path connecting any two first target pixels through the pixels in the set of foreground edge pixels; and delete the pixels in the set of foreground edge pixels that are not passed through any path as redundant pixels.
22. The apparatus according to claim 18, wherein The recognition module further includes: A fourth correction unit, configured to query a plurality of second target pixels in the neighborhood for any one of the background edge pixels; determine the second target pixel with the minimum depth from the plurality of second target pixels; and add the second target pixel with the minimum depth to the set of background edge pixels as a background edge pixel in the case where the second target pixel with the minimum depth does not belong to the foreground edge pixels.
23. The apparatus according to claim 18, wherein, The recognition module further includes: A fifth correction unit, configured to query whether there are adjacent background edge pixels for any one of the background edge pixels; and delete the background edge pixel in the case where there are no adjacent background edge pixels.
24. The device according to claim 18, wherein The recognition module further includes: A sixth correction unit, configured to determine contour pixels outside the edge pixels in the edge map; and replace the depth of the enclosed edge pixels with the depth corresponding to the contour pixels.
25. The apparatus according to claim 24, wherein, The sixth correction unit is further configured to: Dilate the edge pixels in the edge map to obtain a first dilated map; Dilate the edge pixels in the first dilated map again to obtain a second dilated map; Remove the edge pixels overlapping with the first dilated map from the second dilated map, so as to use the remaining edge pixels as the contour pixels.
26. The apparatus according to claim 18, wherein The depth of the foreground edge pixels is greater than the depth of the background edge pixels; The recognition module further includes: A seventh correction unit, configured to increase the depth of the foreground edge pixels according to the depth of the neighborhood pixels of the foreground edge pixels; Reduce the depth of the background edge pixels according to the depth of the neighborhood pixels of the background edge pixels.
27. The device according to claim 15, wherein The modeling module is configured to: Establish connection edges between the pixels in the original image and the neighborhood pixels; Delete the connection edges in the case where the two pixels connected by the connection edges are the foreground edge pixels and the background edge pixels respectively; Establish connection edges between the foreground extended pixels and the background edge pixels with matching depth in the original image according to the depth of the foreground extended pixels; Combine the connection edges to obtain a plurality of triangular patches in three-dimensional space; Render each of the triangular patches according to the pixel values of the pixels in the original image and according to the pixel values of the foreground edge pixels and the background edge pixels to obtain a three-dimensional model.
28. The device according to any one of claims 15 - 17, wherein The generation module includes: Locate a plurality of camera positions in the three-dimensional space where the three-dimensional model is located based on the camera movement trajectory; Determine the imaging diagrams of the three-dimensional model from the perspectives of each of the camera positions; Arrange the corresponding imaging diagrams according to the moments when the camera is at each of the camera positions in the camera movement trajectory to obtain the generated video.
29. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-14.
30. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.
31. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-14.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN114677426A