Image generation method and device and electronic equipment

By performing background filling and area extraction and filling on the second viewing angle image during the video conversion process, the problem of inaccurate filling of the foreground object is solved, and the visual quality of the video is improved.

CN120014110APending Publication Date: 2025-05-16VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510103008.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When converting monocular video to binocular video, the front-back occlusion relationship changes due to changes in view angles, resulting in inaccurate filling of missing areas of foreground objects, affecting the visual quality of the video.

Method used

By generating a second viewing angle image corresponding to the first viewing angle image, including the missing area, and then background filling of the foreground area and the missing area in the second viewing angle image, a pure background image is obtained, and the area image corresponding to the missing area is extracted, and it is filled to the missing area to ensure that the filled area is consistent with the background.

Benefits of technology

It effectively avoids the impact of foreground information on the filling of missing areas, improves the accuracy and consistency of filling, and improves the visual quality of the converted video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014110A_ABST
    Figure CN120014110A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device and electronic equipment, and belongs to the technical field of electronics. The image generation method comprises the steps that according to a first visual angle image, a second visual angle image corresponding to the first visual angle image is generated, and the second visual angle image comprises a first missing area; filling a foreground region and a first missing region in the second view angle image to obtain a second view angle background image; extracting a region image corresponding to the first missing region from the second visual angle background image; and filling the first missing region in the second view angle image with the region image to obtain a processed second view angle image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of electronic technology, and specifically relates to an image generation method, device and electronic equipment. Background Art

[0002] The booming development of virtual reality (VR) technology has placed higher requirements on the processing of video content. Converting monocular video to binocular video is extremely important for achieving an immersive experience. Usually, monocular video can be converted to binocular video based on depth information or parallax information, but the change in perspective may cause the front and back occlusion relationship of objects in the scene to change. For example, if the invisible background object that was originally blocked by the foreground object becomes visible, the area originally occupied by the foreground object will be missing in the converted video, and this area needs to be filled.

[0003] In the related art, the missing area can usually be filled based on the average color, such as calculating the average color value of the background area around the missing area, and then using the average color value to fill the missing area.

[0004] However, since the foreground color may interfere with the calculation of the above average color value, the calculated average color value may differ greatly from the color value of the background area, so that the color of the filled area is inconsistent with the background area, resulting in poor visual quality of the converted video. Summary of the invention

[0005] The purpose of the embodiments of the present application is to provide an image generation method, device and electronic device, which can improve the visual quality of the converted video.

[0006] In a first aspect, an embodiment of the present application provides an image generation method, the method comprising: generating a second perspective image corresponding to the first perspective image based on a first perspective image, the second perspective image including a first missing area; filling the foreground area and the first missing area in the second perspective image to obtain a second perspective background image; extracting a regional image corresponding to the first missing area from the second perspective background image; and filling the regional image into the first missing area in the second perspective image to obtain a processed second perspective image.

[0007] In a second aspect, an embodiment of the present application provides an image generation device, which includes a processing module and an acquisition module. The processing module is used to generate a second perspective image corresponding to the first perspective image based on the first perspective image, wherein the second perspective image includes a first missing area; and fill the foreground area and the first missing area in the second perspective image to obtain a second perspective background image. The acquisition module is used to extract a regional image corresponding to the first missing area from the second perspective background image obtained by the processing module. The processing module is also used to fill the regional image extracted by the acquisition module into the first missing area in the second perspective image to obtain a processed second perspective image.

[0008] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the image generation method described in the first aspect are implemented.

[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the image generation method described in the first aspect are implemented.

[0010] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the image generation method as described in the first aspect.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the image generation method as described in the first aspect.

[0012] In the embodiment of the present application, a second perspective image corresponding to the first perspective image can be generated based on the first perspective image, and the second perspective image includes a first missing area; then the foreground area and the first missing area in the second perspective image are filled to obtain a second perspective background image; then the regional image corresponding to the first missing area is extracted from the second perspective background image, and the regional image is filled into the first missing area in the second perspective image to obtain a processed second perspective image. In this way, since the second perspective background image is obtained by filling the foreground area and the first missing area in the second perspective image as the background, the second perspective background image is a pure background image, and the regional image extracted from the second perspective background image is also a pure background image. When the regional image is used to fill the first missing area in the second perspective image, the influence of the foreground image information on the filling of the first missing area can be avoided, the accuracy of the filling can be improved, the consistency and coordination between the filled first missing area and the background area of ​​the second perspective image can be ensured, and the visual quality of the processed second perspective image can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is one of the flowcharts of the image generation method provided in the embodiment of the present application;

[0014] Figure 2 This is the second flow chart of the image generation method provided in the embodiment of the present application;

[0015] Figure 3 is a schematic diagram of an example of a perspective image and a background image provided in an embodiment of the present application;

[0016] Figure 4 This is the third flow chart of the image generation method provided in the embodiment of the present application;

[0017] Figure 5 This is the fourth flowchart of the image generation method provided in the embodiment of the present application;

[0018] Figure 6 is a schematic diagram of an example of a background image, a mask image, and a region image provided in an embodiment of the present application;

[0019] Figure 7 is a schematic diagram of an example of a regional image and a viewing angle image provided in an embodiment of the present application;

[0020] Figure 8 This is the fifth flowchart of the image generation method provided in the embodiment of the present application;

[0021] Fig. 9 This is the sixth flowchart of the image generation method provided in the embodiment of the present application;

[0022] Fig.10is a schematic diagram of an example of a perspective image, a mask image, and a background image provided in an embodiment of the present application;

[0023] Fig.11 This is the seventh flow chart of the image generation method provided in the embodiment of the present application;

[0024] Fig.12 is a schematic diagram of an image generating device provided in an embodiment of the present application;

[0025] Fig.13 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0026] Fig.14 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0028] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0029] The terms "at least one (item)", "at least one of" and the like in the specification and claims of the present application refer to any one, any two or a combination of more than two of the objects included therein. For example, at least one (item) of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c" and "a, b and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" refers to two or more, and its meaning is similar to that of "at least one (item)".

[0030] The image generation method, device and electronic device provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0031] The image generation method provided in the embodiment of the present application can be applied to three-dimensional (3D) video playback scenarios.

[0032] For example, in a VR usage scenario, a binocular video generated from a monocular video is displayed in a VR device to present different images to the user's two eyes. These images are processed by the user's brain to produce a sense of stereo and depth, thereby enhancing the realism and immersion of the virtual environment.

[0033] As another example, in a 3D viewing scene, a binocular video generated from a monocular video is played in a viewing device, and the binocular video presents different effects to the user's two eyes, giving the user an immersive viewing experience.

[0034] The embodiment of the present application provides an image generation method, which can first generate a second perspective image corresponding to the first perspective image according to the first perspective image, wherein the second perspective image includes a first missing area; then fill the foreground area and the first missing area in the second perspective image to obtain a second perspective background image; then extract the regional image corresponding to the first missing area from the second perspective background image, and fill the regional image into the first missing area in the second perspective image to obtain a processed second perspective image. In this way, since the second perspective background image is obtained by filling the foreground area and the first missing area in the second perspective image as the background, the second perspective background image is a pure background image, and the regional image extracted from the second perspective background image is also a pure background image. When the regional image is used to fill the first missing area in the second perspective image, the influence of the foreground image information on the filling of the first missing area can be avoided, the filling accuracy can be improved, the consistency and coordination between the filled first missing area and the background area of ​​the second perspective image can be ensured, and the visual quality of the processed second perspective image can be improved.

[0035] The execution subject of the image generation method provided in the embodiment of the present application may be an image generation device. Exemplarily, the image generation device may be an electronic device, or a functional component or functional entity in the electronic device. The image generation method provided in the embodiment of the present application will be described exemplarily below by taking the execution subject as an electronic device as an example.

[0036] Figure 1 is a flow chart of the image generation method provided in the embodiment of the present application, such as Figure 1 As shown, the image generation method provided in the embodiment of the present application may include the following steps 101 to 104.

[0037] Step 101: The electronic device generates a second perspective image corresponding to the first perspective image according to the first perspective image.

[0038] In the embodiment of the present application, the first viewing angle image is an image of the first viewing angle, which can be understood as an image seen at the first viewing angle. The second viewing angle image is an image of the second viewing angle, which can be understood as an image seen at the second viewing angle.

[0039] It can be understood that the first viewing angle and the second viewing angle may be different viewing angles, that is, the first viewing angle image and the second viewing angle image are images of different viewing angles.

[0040] Exemplarily, the first perspective image may be a left perspective image of a human eye, and the second perspective image may be a right perspective image of a human eye. Alternatively, the first perspective image may be a right perspective image of a human eye, and the second perspective image may be a left perspective image of a human eye.

[0041] Optionally, in the embodiment of the present application, the first-perspective image may be an image stored in the electronic device or an image extracted by the electronic device from the first-perspective video. The specific embodiment of the present application is not limited.

[0042] The image stored in the electronic device may be an image acquired by the electronic device from other devices or an image collected by the electronic device. The first-perspective video may be a first-perspective video.

[0043] Optionally, in the embodiment of the present application, the electronic device may periodically extract the first-perspective image from the first-perspective video. And, for each extracted first-perspective image, the electronic device may generate a processed second-perspective image corresponding to the first-perspective image according to the method provided in the embodiment of the present application.

[0044] In some embodiments of the present application, the electronic device may periodically extract the first-perspective image from the first-perspective video at a preset frame rate.

[0045] Optionally, in the embodiment of the present application, the preset frame rate may be the frame rate of the 3D video required by the user, and the preset frame rate is less than or equal to the frame rate of the first-view video. For example, assuming that the frame rate of the 3D video required by the user is 30 frames per second (fps), the preset frame rate may be 30fps.

[0046] It should be noted that the preset frame rate can be set or adjusted according to actual needs, and the present embodiment of the application does not limit this. For example, the preset frame rate can be 25fps, 30fps, 50fps or 60fps.

[0047] In some embodiments of the present application, the electronic device may periodically extract the first-perspective image from the first-perspective video according to a preset time period.

[0048] It should be noted that the above-mentioned preset time period can be set or adjusted according to actual needs, and the embodiment of the present application does not limit this. For example, the preset time period can be 10ms, 30ms or 20ms.

[0049] In the embodiment of the present application, the second viewing angle image may include a first missing area, which may also be referred to as a hole area.

[0050] In the embodiment of the present application, the first missing area may be an area that is invisible in the first viewing angle image but visible in the second viewing angle image.

[0051] It should be noted that since the change in viewing angle will cause the change in the front and back occlusion relationship of objects in the scene, the background area that is invisible in the first viewing angle image may become visible in the second viewing angle image. However, since the first viewing angle image does not include the image information of the background area, when the second viewing angle image is generated based on the first viewing angle image, the image of the background area in the second viewing angle image cannot be generated. Therefore, the image of the background area in the second viewing angle image is missing, that is, the second viewing angle image includes the first missing area.

[0052] For example, in the first perspective, area A is blocked by the foreground object, but after changing the perspective, in the second perspective, area A is not blocked by the foreground object. However, since the image information of area A is not included in the first perspective image, the image of area A is missing in the second perspective image. In this case, area A is the first missing area in the second perspective image.

[0053] Optionally, in the embodiment of the present application, the above step 101 can be specifically implemented by the following step 1011.

[0054] Step 1011: The electronic device uses a geometric transformation algorithm to perform pixel transformation on the first-perspective image according to the parallax image to obtain a second-perspective image.

[0055] In the embodiment of the present application, the above-mentioned parallax image can be used to represent the position difference of the pixel points in the image between the first viewing angle and the second viewing angle.

[0056] Optionally, in the embodiment of the present application, the position of each pixel in the above parallax image is the same as the position of each pixel in the above first-view image. For example, if the position of a pixel in the parallax image is (1, 1), then the pixel in the first-view image is also (1, 1).

[0057] Optionally, in the embodiment of the present application, the pixel value of each pixel in the above-mentioned parallax image can be used to represent the offset direction and offset amount of the position of the pixel in the second perspective image relative to the position of the pixel in the first perspective image. In other words, for a pixel in the first perspective image, after the pixel is moved along the offset direction by the offset amount, the position of the pixel is the position of the pixel in the second perspective image.

[0058] Exemplarily, the above-mentioned offset direction can be determined according to the perspective relationship between the first perspective and the second perspective. For example, assuming that the first perspective is a left perspective and the second perspective is a right perspective, the offset direction can be horizontal to the left. Assuming that the first perspective is a right perspective and the second perspective is a left perspective, the offset direction can be horizontal to the right. Assuming that the first perspective is upward relative to the second perspective, the offset direction can be vertically upward, and assuming that the first perspective is downward relative to the second perspective, the offset direction can be vertically downward.

[0059] Exemplarily, the offset direction is left or right offset, and a positive number indicates a left offset, and a negative number indicates a right offset. Assume that the value of pixel A in the disparity image is 2, indicating that the position of pixel A in the second-perspective image is horizontally offset by 2 pixels to the left relative to the position of pixel A in the first-perspective image. Then, after moving pixel A in the first-perspective image horizontally by 2 pixels to the left, the position of pixel A is the position of pixel A in the second-perspective image. For example, assuming that the position of pixel A in the disparity image is (10, 1), the position of pixel A in the first-perspective image is also (10, 1). After moving pixel A horizontally by 2 pixels to the left, the position of pixel A in the second-perspective image is (8, 1).

[0060] Exemplarily, the offset direction is left or right offset, and a positive number indicates a left offset, and a negative number indicates a right offset. Assume that the value of pixel point B in the disparity image is -1, indicating that the position of pixel point B in the second-perspective image is horizontally offset by 1 pixel to the right relative to the position of pixel point B in the first-perspective image. Then, after moving pixel point B in the first-perspective image to the right by 1 pixel in the horizontal direction, the position of pixel point B is the position of pixel point B in the second-perspective image. For example, assuming that the position of pixel point B in the disparity image is (3, 1), the position of pixel point B in the first-perspective image is also (3, 1). After moving pixel point B to the right by 1 pixel in the horizontal direction, the position of pixel point B in the second-perspective image is (4, 1).

[0061] Optionally, in an embodiment of the present application, the electronic device may acquire a first image and a second image, generate an initial disparity image based on the first image and the second image through a stereo matching algorithm, and then multiply the initial disparity image by a preset coefficient to obtain a disparity image.

[0062] Optionally, in an embodiment of the present application, the first image is an image of a first perspective, the second image is an image of a second perspective, the first image and the second image are images respectively acquired by two acquisition devices at the same acquisition time point, and the baseline length between the two acquisition devices is less than the baseline length of a human eye. The baseline length between the two acquisition devices is the distance between the optical centers of the two acquisition devices, and the baseline length of a human eye is the distance between a person's left eye and right eye.

[0063] Optionally, in an embodiment of the present application, the above-mentioned preset coefficient can be determined according to the baseline length between the two acquisition devices, the baseline length of the human eye, and the focal lengths of the two acquisition devices.

[0064] Exemplarily, taking the electronic device as a mobile phone, the mobile phone includes a main camera and a sub-camera, the distance between the optical centers of the main camera and the sub-camera is the baseline length, and the baseline length is smaller than the baseline length of the human eye, then the electronic device can obtain the image taken by the main camera (i.e., the first image mentioned above) and the image taken by the sub-camera (i.e., the second image mentioned above) at the same time point, and generate an initial disparity image based on the image taken by the main camera and the image taken by the sub-camera, and then determine the preset coefficient according to the baseline length between the two acquisition devices, the baseline length of the human eye and the focal length of the two acquisition devices, and multiply the pixel value of each pixel in the initial disparity image by the initial coefficient to obtain a disparity image, and the pixel value of each pixel in the disparity image is the pixel value after multiplying the initial coefficient.

[0065] Optionally, in an embodiment of the present application, the electronic device may first preprocess the first image and the second image so that the first image and the second image are on the same plane and the epipolar lines are in the horizontal direction, or so that the first image and the second image are on the same plane and the epipolar lines are in the vertical direction. Then, the matching cost of the first image and the second image is calculated to generate a disparity space image (DSI), which is a three-dimensional image, in which each disparity value corresponds to a cost map, and then each cost map is filtered and denoised. Then, for each pixel point, the electronic device finds the cost map with the minimum matching cost from the multiple filtered cost maps corresponding to the pixel point, and uses the disparity value corresponding to the cost map as the pixel value of the pixel point. By analogy, the pixel value of each pixel point can be determined, and then the pixel value of each pixel point is multiplied by the initial coefficient to obtain a disparity image.

[0066] It should be noted that the disparity image obtained by the above stereo matching algorithm is used to represent the position difference between the first and second perspectives of the pixels in the image in the horizontal direction, that is, the pixel value of each pixel in the disparity image is used to represent the position of the pixel in the second perspective image, relative to the position of the pixel in the first perspective image, and the amount of the horizontal shift to the left or right. Alternatively, the disparity image obtained by the above stereo matching algorithm is used to represent the position difference between the first and second perspectives of the pixels in the image in the vertical direction, that is, the pixel value of each pixel in the disparity image is used to represent the position of the pixel in the second perspective image, relative to the position of the pixel in the first perspective image, and the amount of the vertical shift upward or downward.

[0067] Optionally, in an embodiment of the present application, the electronic device may input the first perspective image into a depth estimation model, output a first depth image through the depth estimation model, and generate a disparity image according to the first depth image.

[0068] It should be noted that the above-mentioned depth estimation model can be obtained by training a deep learning model in a supervised training manner through a large number of single-view images and depth images.

[0069] Optionally, in an embodiment of the present application, the first depth image may be used to represent the depth of pixels in the first perspective image, that is, the pixel value of each pixel in the first depth image represents the distance between the pixel and the shooting source in the first perspective image.

[0070] Optionally, in an embodiment of the present application, the electronic device can calculate a disparity image based on the first depth image, the intrinsic parameters and extrinsic parameters of a capture device that captures the first perspective image, the focal length of the capture device, the baseline length between two capture devices in the electronic device, and the baseline length of the human eye.

[0071] For example, taking a pixel in the first depth image as an example, the electronic device can calculate a new pixel value of the pixel based on the pixel value of the pixel, the internal and external parameters of the acquisition device, the focal length of the acquisition device, the baseline length between the two acquisition devices, and the baseline length of the human eye. Similarly, the new pixel value of each pixel can be obtained, and then a disparity image can be obtained.

[0072] Optionally, in an embodiment of the present application, for each pixel in the first perspective image, the electronic device can move the position of the pixel according to the pixel value of the pixel in the parallax image. The position of the pixel after the movement is the position of the pixel in the second perspective image, and so on. After moving the position of each pixel in the first perspective image, the second perspective image can be obtained.

[0073] Exemplarily, for the first pixel in the first perspective image, the electronic device determines that the moving direction is horizontally to the right, and the moving distance is 2 pixels based on the pixel value of the first pixel in the parallax image. Then the electronic device can move the first pixel to the right by 2 pixels in the horizontal direction. At this time, the position of the first pixel is the position of the first pixel in the second perspective image. Similarly, each pixel can be moved to the corresponding position in the second perspective image to obtain the second perspective image.

[0074] For example, taking the first perspective image as the left perspective image and the second perspective image as the right perspective image, for each pixel in the left perspective image, the electronic device can move each pixel horizontally to the left by a corresponding offset according to the pixel value of each pixel in the disparity map, and then determine the position of each pixel in the right perspective image to finally obtain the right perspective image.

[0075] For example, taking the first perspective image as the right perspective image and the second perspective image as the left perspective image, for each pixel in the right perspective image, the electronic device can move each pixel to the right horizontally by a corresponding offset according to the pixel value of each pixel in the disparity map, and then determine the position of each pixel in the left perspective image to finally obtain the left perspective image.

[0076] It should be noted that the above only takes the translation of pixels as an example to illustrate the process of generating the second perspective image. In other examples, the electronic device can also rotate, project, or perform other transformations on the pixels in the first perspective image according to the parallax image to obtain the second perspective image, which is not limited in the embodiments of the present application.

[0077] In this way, the electronic device uses a geometric transformation algorithm to deform the first-view image based on the parallax image to obtain the second-view image, providing a basis for subsequent processing.

[0078] Optionally, in an embodiment of the present application, the first-perspective image may be a frame of image in a first-perspective video, and the first-perspective video may include multiple frames of image; the step 101 may be specifically implemented through the following step 1012.

[0079] Step 1012: The electronic device generates a second perspective image corresponding to the first perspective image according to the first perspective image and N frames of images.

[0080] In the embodiment of the present application, the N frames of images may be N frames of images adjacent to the first viewing angle image in the multiple frames of images, where N is a positive integer, and the N frames of images are smaller than the multiple frames of images.

[0081] Optionally, in an embodiment of the present application, the above-mentioned first-perspective image can be any frame image in the first-perspective video, that is, each frame image in the first-perspective video can be used as a first-perspective image, and for each frame image in the first-perspective video, a second-perspective image corresponding to the frame image can be generated.

[0082] Optionally, in an embodiment of the present application, the electronic device may combine an optical flow method, a parallax estimation algorithm, and a geometric transformation algorithm to generate a second perspective image corresponding to the first perspective image based on the first perspective image and N frames of images.

[0083] Optionally, in an embodiment of the present application, the electronic device may use an optical flow method to calculate a motion vector field between a first-perspective image and N-frame images, and the motion vector field may include motion information of each pixel in the image, including motion speed and motion direction. In addition, the electronic device may input the first-perspective image and N-frame images into a depth estimation model, perform disparity estimation through the depth estimation model, output a second depth image, and generate a first disparity image based on the second depth image. Then, the electronic device may fuse the disparity image with the motion vector field to obtain a second disparity image. Next, the electronic device uses a geometric transformation algorithm to perform pixel transformation on the first-perspective image based on the second disparity image to obtain a second-perspective image.

[0084] In this way, the electronic device combines the optical flow method with the disparity estimation algorithm to make full use of motion information and depth information to improve the accuracy of the generated disparity map, and then generates a second disparity image based on the disparity map using a geometric transformation algorithm, which can improve the accuracy of the second disparity image. In addition, each time the second perspective image is generated, the motion information is combined to improve the coherence between adjacent second perspective images, thereby ensuring the smoothness of the final generated second perspective video.

[0085] Optionally, in an embodiment of the present application, the first perspective image may be one of the M frames of images extracted by the electronic device from the first perspective video. Exemplarily, the electronic device may extract the M frames of images from the first perspective video and sort the M frames of images according to the acquisition time point. Then, the electronic device may generate a second perspective image corresponding to the first perspective image based on the first perspective image and the N frames of images adjacent to the first perspective image in the M frames of images.

[0086] It should be noted that the specific implementation of the electronic device generating the second perspective image corresponding to the first perspective image based on the first perspective image and N frame images adjacent to the first perspective image in M ​​frame images can be found in the relevant description of the above embodiment. To avoid repetition, this embodiment will not be repeated here.

[0087] Step 102: The electronic device fills the foreground area and the first missing area in the second viewing angle image to obtain a second viewing angle background image.

[0088] In the embodiment of the present application, the foreground area may be an area where the foreground object is located in the second viewing angle image. The first missing area may be an area that is invisible in the first viewing angle image but visible in the second viewing angle image.

[0089] Optionally, in an embodiment of the present application, the electronic device can fill the foreground area and the first missing area as the background image based on the background image excluding the foreground area and the first filling area in the second perspective image, to obtain a pure background image of the second perspective, i.e., the above-mentioned second perspective background image.

[0090] Optionally, in the embodiment of the present application, combined with Figure 1 ,like Figure 2 As shown, the above step 102 can be specifically implemented through the following step 1021.

[0091] Step 1021: The electronic device fills the foreground area and the first missing area in the second perspective image based on the pixel values ​​of the background area in the second perspective image to obtain a second perspective background image.

[0092] In the embodiment of the present application, the above-mentioned background area may be the area where the background objects are located in the second perspective image.

[0093] Optionally, in an embodiment of the present application, the electronic device may use a preset background filling algorithm to fill both the foreground area and the first missing area in the second perspective image as the background to obtain a second perspective background image.

[0094] Exemplarily, the above-mentioned preset background filling algorithm may include an algorithm based on texture synthesis or an image restoration algorithm based on an image filling model. The image filling model may be obtained by training a deep learning model with a large number of images to be filled and background images, and the image filling model may be used to fill the area to be filled in the input image to obtain a background image.

[0095] Optionally, in an embodiment of the present application, the electronic device may adopt a texture synthesis-based method to analyze the pixel values ​​of the background area in the second-perspective image, and generate filling pixel values ​​(also referred to as filling textures) similar to the pixel values ​​of the background area based on the analysis results, and then use the filling pixel values ​​to fill the foreground area and the first missing area in the second-perspective image to obtain a second-perspective background image.

[0096] Exemplarily, the electronic device may generate a first filling pixel value according to the pixel values ​​of the background area surrounding the foreground area in the second perspective image, and then fill the foreground area in the second perspective image with the first filling pixel value. Furthermore, the electronic device may generate a second filling pixel value according to the pixel values ​​of the background area surrounding the first missing area in the second perspective image, and then fill the first missing area in the second perspective image with the second filling pixel value.

[0097] Optionally, in an embodiment of the present application, the electronic device can mark the foreground area and the first missing area in the first-perspective image, and input the marked second-perspective image into an image filling model, extract features (such as pixel values) of the background area in the second-perspective image through the image filling model, and fill the foreground area and the first missing area in the second-perspective image to obtain a second-perspective background image.

[0098] For example, Figure 3 (a) is a second perspective image 31, which includes a foreground area 32 and a first missing area 33. The electronic device uses a preset background filling algorithm to fill the foreground area 32 and the first missing area 33 in the second perspective image 31, so that the foreground area 31 and the first missing area 33 are converted into the background, and then a second perspective background image 34 is obtained, as shown in FIG. Figure 3 As shown in (b) in .

[0099] Optionally, in the embodiment of the present application, the electronic device may further smooth the second perspective background image to eliminate stitching marks and improve the image quality and visual effect of the filled second perspective background image.

[0100] In this way, the electronic device fills the foreground area and the first missing area in the second perspective image as the background image by background filling, and obtains the second perspective background image including only the background, which can avoid the influence of the foreground on subsequent processing.

[0101] Step 103: The electronic device extracts a region image corresponding to the first missing region from the second viewing angle background image.

[0102] Optionally, in the embodiment of the present application, combined with Figure 1 ,like Figure 4 As shown, the above step 103 can be specifically implemented through the following steps 1031 and 1032.

[0103] Step 1031: The electronic device determines, based on the missing area mask image, a first area corresponding to the first missing area in the second viewing angle background image.

[0104] In the embodiment of the present application, the missing area mask image can be used to represent an area that is invisible in the first perspective image but visible in the second perspective image.

[0105] Optionally, in the embodiment of the present application, the missing area mask image may be a missing area mask image of the second viewing angle, which has the same size as the second viewing angle image, and the color of the missing area in the missing area mask image is different from the color of other areas except the missing area.

[0106] Optionally, in the embodiment of the present application, combined with Figure 4 ,like Figure 5 As shown, before the above step 1031, the image generation method provided in the embodiment of the present application may further include the following step 1030.

[0107] Step 1030: The electronic device generates a missing area mask image based on the first missing area in the second viewing angle image.

[0108] Optionally, in an embodiment of the present application, the electronic device may adopt a preset missing area extraction algorithm to determine the first missing area in the second perspective image, and then set each pixel in the first missing area to white (or set the pixel value of each pixel to 1), and set other pixel points in the second perspective image except the first missing area to black (or set the pixel values ​​of other pixel points to 0), or, set each pixel in the first missing area to black, and set other pixel points in the second perspective image except the first missing area to white, so as to generate a missing area mask image.

[0109] Exemplarily, the above-mentioned preset missing region extraction algorithm may include a threshold method, an edge detection algorithm, or an algorithm based on a missing region detection model.

[0110] The missing area detection model can be obtained by training a deep learning model with a large number of sample images with marked missing areas.

[0111] Optionally, in an embodiment of the present application, the electronic device may convert the second perspective image into a grayscale image, and set a second threshold, and determine the area where pixels in the second perspective image with pixel values ​​lower than the second threshold are located as the first missing area.

[0112] It should be noted that the second threshold can be set based on experience or experimental data, and the embodiment of the present application does not limit this.

[0113] Optionally, in an embodiment of the present application, the electronic device may use an edge detection algorithm (such as the Canny algorithm) to detect edges in the second-perspective image, and further process and analyze the edge detection results to determine the first missing area in the second-perspective image.

[0114] Optionally, in an embodiment of the present application, the electronic device may input the second perspective image into a missing area detection model, identify the missing area in the second perspective image through the missing area detection model, and output the position information of the first missing area.

[0115] In this way, the electronic device generates a missing area mask image based on the first missing area in the second perspective image, so as to quickly extract the image to be filled into the first missing area from the second perspective background image according to the missing area mask image, thereby improving the efficiency of image filling.

[0116] Optionally, in an embodiment of the present application, the electronic device may determine, based on the position information of the first missing area in the missing area mask image, an area indicated by the position information in the second perspective background image as the first area corresponding to the first missing area.

[0117] For example, taking the case where each pixel in the first missing area in the missing area mask image is white, the electronic device can first determine the position information of the white area in the missing area mask image, and then determine the area indicated by the position information in the second perspective background image, and determine the area as the first area corresponding to the first missing area.

[0118] Step 1032: The electronic device extracts the image of the first area from the second viewing angle background image to obtain a regional image.

[0119] For example, Figure 6 (a) is the second viewing angle background image 34, Figure 6 (b) is the missing region mask image 35. According to the missing region mask image 35, the first region corresponding to the first missing region is determined in the second perspective background image 34, and then the image of the first region is extracted from the second perspective background image 34 to obtain the region image 36, as shown in FIG. Figure 6 As shown in (c) in .

[0120] In this way, the electronic device extracts the regional image corresponding to the first missing area from the second-perspective background image based on the missing area mask image. Since the second-perspective background image is a pure background image, the regional image extracted from the second-perspective background image is also a pure background image and will not be affected by the foreground image. Then, the missing area of ​​the second-perspective image is filled according to the regional image. The missing area of ​​the second-perspective image can be filled as the background, and the background is coordinated and unified with the original background in the second-perspective image, thereby improving the quality of the filled image.

[0121] Step 104: The electronic device fills the first missing area in the second viewing angle image with the regional image to obtain a processed second viewing angle image.

[0122] Optionally, in the embodiment of the present application, the electronic device may overwrite the pixel values ​​of the regional image to the first missing region in the second-perspective image to obtain a processed second-perspective image.

[0123] For example, Figure 7 (a) in FIG. 3 is a region image 36, Figure 7 (b) is a second perspective image 31, which includes a first missing region 33. The region image 36 is filled into the first missing region 33 of the second perspective image 31 to obtain a processed second perspective image 37, as shown in FIG. Figure 7 As shown in (c) in .

[0124] In an image display method provided by an embodiment of the present application, in the process of filling the first missing area of ​​the second perspective image, the background is filled in the foreground area and the first missing area in the second perspective image to obtain the second perspective background image, and then the area image is extracted from the second perspective background image according to the missing area mask image, and then the first missing area in the second perspective image is filled with the area image, thereby avoiding filling errors caused by improper foreground and background processing in traditional methods, avoiding the influence of foreground image information on the filling of the first missing area, improving the accuracy of background filling, reducing image defects caused by confusion between foreground and background, ensuring the consistency and coordination between the filled first missing area and the background area of ​​the second perspective image, and improving the visual quality of the processed second perspective image.

[0125] In an embodiment of the present application, when the above-mentioned first-perspective image is an image extracted by the electronic device from a first-perspective video, the electronic device can generate a processed second-perspective image corresponding to each first-perspective image in the first-perspective video, and these processed second-perspective images can constitute a second-perspective video. Then, the electronic device fuses the first-perspective video and the second-perspective video to obtain a binocular video, and displays the binocular video to the user, that is, presents different images to the user's two eyes, and these images are processed by the user's brain to produce a sense of stereo and depth, thereby enhancing the user's sense of reality and immersion in the virtual environment.

[0126] Optionally, in the embodiment of the present application, combined with Figure 1 ,like Figure 8 As shown, before "the electronic device fills the foreground area in the second perspective image" in the above step 102, the image generation method provided in the embodiment of the present application may also include the following steps 201 to 203.

[0127] Step 201: The electronic device generates a first-perspective foreground mask image according to the first-perspective image.

[0128] Optionally, in the embodiment of the present application, the foreground mask image of the first perspective may be a foreground mask image of the first perspective, which may be used to represent the foreground image in the first perspective image.

[0129] Optionally, in the embodiment of the present application, the first-perspective foreground mask image has the same size as the first-perspective image, and the foreground area in the first-perspective foreground mask image has different colors from other areas except the foreground area.

[0130] Optionally, in an embodiment of the present application, the electronic device may adopt a preset foreground extraction algorithm to determine a first-perspective foreground image in the first-perspective image, that is, a foreground image in the first-perspective image, and then set each pixel in the first-perspective foreground image to white (or set the pixel value of each pixel to 1), and set other pixel points in the first-perspective image except the first-perspective foreground image to black (or set the pixel values ​​of other pixel points to 0), or, set each pixel in the first-perspective foreground image to black, and set other pixel points in the first-perspective image except the first-perspective foreground image to white, so that a first-perspective foreground mask image can be generated.

[0131] Exemplarily, the above-mentioned preset foreground extraction algorithm may include a threshold method, an edge detection algorithm, or a foreground detection algorithm based on a target detection model.

[0132] Among them, the above-mentioned target detection model can be obtained by training a deep learning model through a large number of sample images and foreground images of sample images.

[0133] Optionally, in an embodiment of the present application, the electronic device may mark pixel points having pixel values ​​greater than a first threshold as foregrounds based on the pixel value of each pixel point in the disparity image, and determine the image of the area where the multiple pixel points marked as foregrounds are located as the first-perspective foreground image.

[0134] It should be noted that the first threshold may be determined by analyzing the pixel values ​​of the foreground area and the background area in the disparity image, or may be set based on experience or experimental data, which is not limited in the embodiments of the present application.

[0135] Optionally, in an embodiment of the present application, the electronic device may convert the first-perspective image into a grayscale image, and then identify the edges in the first-perspective image through an edge detection algorithm, and refine the detected edges to reduce the width of the edges, and then use a contour tracking algorithm to extract the contour of the object from the refined edges, and then use morphological operations (such as dilation) to fill the area inside the contour to form a foreground area, separate the foreground area from the background area, and obtain a first-perspective foreground image that only includes the foreground object.

[0136] Optionally, in an embodiment of the present application, the electronic device may input the first-perspective image into a target detection model, identify the foreground object in the first-perspective image through the target detection model, and output the first-perspective foreground image.

[0137] In this way, the foreground image in the first-perspective image can be accurately obtained through the above method, and then the first-perspective foreground mask image can be generated for subsequent processing.

[0138] Step 202: The electronic device generates a second-viewing angle foreground mask image according to the first-viewing angle foreground mask image.

[0139] Optionally, in the embodiment of the present application, the second-perspective foreground mask image is a foreground mask image of the second perspective, which can be used to represent the foreground image in the second-perspective image.

[0140] Optionally, in the embodiment of the present application, combined with Figure 8 ,like Fig. 9 As shown, the above step 202 can be specifically implemented through the following step 2021.

[0141] Step 2021: The electronic device uses a geometric transformation algorithm to perform pixel transformation on the first-view foreground mask image according to the disparity image to obtain a second-view foreground mask image.

[0142] In the embodiment of the present application, the above-mentioned parallax image can be used to represent the position difference of the pixel points in the image between the first viewing angle and the second viewing angle.

[0143] It should be noted that, for the explanation of the parallax image, reference may be made to the relevant description in the above step 1011 , and to avoid repetition, this embodiment will not be described again here.

[0144] Optionally, in an embodiment of the present application, for each pixel point in the first-perspective foreground mask image, the electronic device can move the position of the pixel point according to the pixel value of the pixel point in the parallax image, and the position of the pixel point after the movement is the position of the pixel point in the second-perspective foreground mask image, and so on. After moving the position of each pixel point in the first-perspective foreground mask image, the second-perspective foreground mask image can be obtained.

[0145] It should be noted that the implementation principle of the above step 2021 is the same as the implementation principle of the above step 1011. Therefore, the specific implementation of the above step 2021 can refer to the relevant description of the above step 1011. To avoid repetition, this embodiment will not be repeated here.

[0146] In this way, the electronic device uses a geometric transformation algorithm to map each pixel in the first-perspective foreground mask image based on the disparity image to obtain the second-perspective foreground mask image, which can ensure the accurate transmission and synchronous update of the foreground image during the perspective conversion process, thereby improving the accuracy of the second-perspective foreground mask image.

[0147] Step 203: The electronic device determines the foreground area in the second-viewing angle image according to the second-viewing angle foreground mask image.

[0148] Optionally, in an embodiment of the present application, the electronic device may determine, based on the position information of the foreground area in the second-perspective foreground mask image, the area indicated by the position information in the second-perspective image as the foreground area.

[0149] In this way, since the second-perspective foreground mask image is obtained by mapping the first-perspective foreground mask image, the accuracy of the foreground image during the perspective conversion process is ensured. Therefore, determining the foreground area based on the second-perspective foreground mask image can improve the accuracy of the determined foreground area. Therefore, when the foreground area is filled subsequently, the accuracy of the filling can be improved.

[0150] Optionally, in an embodiment of the present application, before the above-mentioned step 102, the electronic device may generate a second-perspective foreground mask image and a missing area mask image by the manner of the above-mentioned embodiment, and then, the specific implementation of the above-mentioned step 102 may include: the electronic device synthesizes the second-perspective foreground mask image and the missing area mask image to obtain a mask image of the area to be filled, and determines the area to be filled in the second-perspective image according to the mask image of the area to be filled, and the area to be filled includes the foreground area and the first missing area, and then fills the area to be filled in the second-perspective image to obtain a second-perspective background image.

[0151] For example, Fig.10 (a) is a second perspective image 31, and the second perspective image 31 includes a foreground area 32 and a first missing area 33. Fig.10 (b) is a mask image 38 of the area to be filled, which is composed of a second perspective foreground mask image and a missing area mask image. The electronic device determines the area to be filled in the second perspective image according to the mask image 38 of the area to be filled, and the area to be filled includes the foreground area 32 and the first missing area 33, and then fills the area to be filled in the second perspective image 31 to obtain a second perspective background image 34, as shown in FIG. Fig.10 As shown in (c) in .

[0152] It should be noted that the above embodiment is to illustrate the image generation method provided by the embodiment of the present application by taking the second perspective image generated from the first perspective image as an example. In some embodiments of the present application, the electronic device can generate images of various perspectives (i.e., perspectives different from the first perspective) after processing based on the first perspective image. Next, the electronic device is to generate a third perspective image processed based on the first perspective image as an example for explanation.

[0153] Optionally, the image generation method provided in the embodiment of the present application may further include the following step 105.

[0154] Step 105: The electronic device generates a third-perspective image corresponding to the first-perspective image based on the first-perspective image.

[0155] In the embodiment of the present application, the third-perspective image may include a second missing area, and the third-perspective image and the second-perspective image may be images of different perspectives.

[0156] Optionally, in the embodiment of the present application, the second missing area may be used to represent an area that is invisible in the first-perspective image but visible in the third-perspective image.

[0157] It should be noted that the specific implementation of the above step 105 can refer to the relevant description of the above step 101. To avoid repetition, this embodiment will not be repeated here.

[0158] Furthermore, after the electronic device generates a third perspective image corresponding to the first perspective image, it can also fill the foreground area and the second missing area in the second perspective image to obtain a third perspective background image, and then extract the area image corresponding to the second missing area from the third perspective background image, and fill the area image into the second missing area in the third perspective image to obtain a processed third perspective image.

[0159] It should be noted that the processing process after the above electronic device generates the third-viewing angle image can refer to the relevant description of the above embodiment. To avoid repetition, this embodiment will not be described again.

[0160] In this way, the image generation method provided in the embodiment of the present application can not only convert a single-perspective video (monocular video) into a dual-perspective video (binocular video), but also convert a single-perspective video into a multi-perspective (such as three-perspective, four-perspective, etc.) video, thereby generating high-quality videos for more complex spatial VR scenes or other multi-perspective application scenarios.

[0161] Fig.11 30 is a flowchart of an image generation method provided in an embodiment of the present application. Taking a binocular video playback scenario as an example, the method may include the following steps 301 to 308.

[0162] Step 301: The electronic device obtains a left-view image and a parallax image.

[0163] Exemplarily, the electronic device extracts a single frame image from a monocular video source at a preset frame rate or a preset time period as a left-view image, where the left-view image includes monocular visual information of a scene.

[0164] Exemplarily, the electronic device uses a stereo matching algorithm to generate a disparity image based on the first image and the second image in the binocular video with a shorter baseline. Alternatively, the electronic device inputs the left-view image into a depth estimation model to obtain a depth image, and generates a disparity image based on the depth image.

[0165] It should be noted that the left-view image may be the first-view image in the above embodiment.

[0166] Step 302: The electronic device generates a left-view foreground mask image according to the left-view image.

[0167] It should be noted that the left-view foreground mask image may be the first-view foreground mask image in the above embodiment.

[0168] Step 303: The electronic device generates a right-view image according to the parallax image and the left-view image.

[0169] Exemplarily, the electronic device may use the disparity map and a geometric transformation algorithm to perform pixel transformation on the left-view image to obtain a right-view image, where the right-view image includes the missing area.

[0170] It should be noted that the above-mentioned right-view image may be the second-view image in the above-mentioned embodiment, and the above-mentioned missing area may be the first missing area in the above-mentioned embodiment.

[0171] Step 304: The electronic device generates a right-view foreground mask image according to the left-view foreground mask image and the parallax image.

[0172] Exemplarily, the electronic device may utilize the disparity map and adopt a geometric transformation algorithm to perform pixel transformation on the left-view foreground mask image to obtain the right-view foreground mask image.

[0173] It should be noted that the above-mentioned right-view foreground mask image may be the second-view foreground mask image in the above-mentioned embodiment.

[0174] Step 305: The electronic device generates a missing area mask image based on the right viewing angle image.

[0175] Step 306: The electronic device fills the foreground area and the missing area in the right-view image according to the right-view foreground mask image and the missing area mask image to obtain a right-view background image.

[0176] Exemplarily, the electronic device determines the foreground area and the missing area as the filling area, and fills the filling area using a preset background filling method, so that the foreground area and the missing area are converted into the background during the filling, and a pure background image of the right perspective is obtained, that is, the right perspective background image.

[0177] For example, Fig.10 (a) is a right-view image 31, which includes a foreground area 32 and a missing area 33. Fig.10 (b) is a mask image 38 of the area to be filled, which is composed of a right-view foreground mask image and a missing area mask image. The electronic device determines the area to be filled in the right-view image 31 according to the mask image 38 of the area to be filled, and the area to be filled includes the foreground area 32 and the missing area 33, and then fills the area to be filled in the right-view image 31 to obtain the right-view background image 34, as shown in FIG. Fig.10 As shown in (c) in .

[0178] Step 307: The electronic device extracts a region image corresponding to the missing region in the right-view image from the right-view background image.

[0179] For example, Figure 6 (a) is the right-view background image 34, Figure 6 (b) in FIG. 3 is a missing region mask image 35. According to the missing region mask image 35, a region image 36 is extracted from the right perspective background image 34. Figure 6 As shown in (c) in .

[0180] Step 308: The electronic device fills the missing area in the right-view image with the regional image to obtain a processed right-view image.

[0181] For example, Figure 7 (a) in FIG. 3 is a region image 36, Figure 7 (b) is a right perspective image 31, which includes a missing area 33. The region image 36 is filled into the missing area 33 of the right perspective image 31 to obtain a processed right perspective image 37, as shown in FIG. Figure 7 As shown in (c) in .

[0182] Furthermore, the electronic device generates a plurality of processed right-view images, generates a right-view video according to the plurality of processed right-view images, and synthesizes the left-view video and the right-view video to obtain a binocular video and display it.

[0183] It should be noted that the specific implementation of the above steps 301 to 308 can refer to the relevant description of the above method embodiment. To avoid repetition, this embodiment will not be repeated here.

[0184] The image generation method provided by the embodiment of the present application is a pure background image because the right-perspective background image is obtained by filling the foreground area and the missing area in the right-perspective image as the background. Therefore, the right-perspective background image is a pure background image. Then, the regional image extracted from the right-perspective background image is also a pure background image. When the missing area in the right-perspective image is filled with the regional image, the influence of the foreground image information on the filling of the missing area can be avoided, the filling accuracy can be improved, the consistency and coordination between the filled missing area and the background area of ​​the right-perspective image can be ensured, and the visual quality of the processed right-perspective image can be improved.

[0185] It should be noted that the above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or any two or more can be executed in combination with each other. The specific implementation method can be determined according to actual usage requirements, and the embodiments of the present application do not impose any restrictions on this.

[0186] The image generation method provided in the embodiment of the present application can be executed by an image generation device. In the embodiment of the present application, the image generation device provided in the embodiment of the present application is described by taking the image generation method executed by the image generation device as an example.

[0187] Fig.12 11 is a schematic diagram of the structure of an image generating device provided in an embodiment of the present application. The image generating device 1100 includes: a processing module 1101 and an acquisition module 1102.

[0188] Among them, the above-mentioned processing module 1101 is used to generate a second perspective image corresponding to the first perspective image based on the first perspective image, and the second perspective image includes the first missing area; and fill the foreground area and the first missing area in the second perspective image to obtain a second perspective background image.

[0189] The acquisition module 1102 is used to extract a region image corresponding to the first missing region from the second viewing angle background image obtained by the processing module 1101 .

[0190] The processing module 1101 is further configured to fill the first missing region in the second viewing angle image with the regional image extracted by the acquisition module 1102 to obtain a processed second viewing angle image.

[0191] Optionally, in an embodiment of the present application, the processing module 1101 is specifically used to fill the foreground area and the first missing area in the second perspective image based on the pixel value of the background area in the second perspective image to obtain the second perspective background image.

[0192] Optionally, in an embodiment of the present application, the above-mentioned processing module 1101 is also used to generate a first-perspective foreground mask image based on the first-perspective foreground mask image; and to generate a second-perspective foreground mask image based on the first-perspective foreground mask image; and to determine the foreground area in the second-perspective image based on the second-perspective foreground mask image.

[0193] Optionally, in an embodiment of the present application, the above-mentioned processing module 1101 is specifically used to perform pixel transformation on the first-perspective foreground mask image according to the disparity image and adopt a geometric transformation algorithm to obtain the second-perspective foreground mask image; the disparity image is used to represent the position difference of the pixel points in the image between the first perspective and the second perspective.

[0194] Optionally, in an embodiment of the present application, the above-mentioned acquisition module 1102 is specifically used to determine a first area corresponding to a first missing area in the second perspective background image based on a missing area mask image, and the missing area mask image is used to represent an area that is invisible in the first perspective image and visible in the second perspective image; and, extract the image of the first area from the second perspective background image to obtain a regional image.

[0195] Optionally, in the embodiment of the present application, the processing module 1101 is further configured to generate a missing area mask image based on the first missing area in the second viewing angle image.

[0196] Optionally, in an embodiment of the present application, the above-mentioned processing module 1101 is specifically used to perform pixel transformation on the first perspective image according to the disparity image and adopt a geometric transformation algorithm to obtain a second perspective image; the disparity image is used to represent the position difference of pixel points in the image between the first perspective and the second perspective.

[0197] Optionally, in an embodiment of the present application, the first-perspective image is a frame image in a first-perspective video, and the first-perspective video includes multiple frame images; the processing module 1101 is specifically used to generate a second-perspective image corresponding to the first-perspective image based on the first-perspective image and N frame images, and the N frame images are N frame images adjacent to the first perspective image in the multiple frame images, and N is a positive integer.

[0198] Optionally, in an embodiment of the present application, the processing module 1101 is further used to generate a third perspective image corresponding to the first perspective image based on the first perspective image, wherein the third perspective image includes a second missing area, and the third perspective image and the second perspective image are images of different perspectives.

[0199] The image generating device provided by the embodiment of the present application is a pure background image because the second-perspective background image is obtained by filling the foreground area and the first missing area in the second-perspective image as the background. Therefore, the regional image extracted from the second-perspective background image is also a pure background image. When the first missing area in the second-perspective image is filled with the regional image, the influence of the foreground image information on the filling of the first missing area can be avoided, the filling accuracy can be improved, the consistency and coordination between the filled first missing area and the background area of ​​the second-perspective image can be ensured, and the visual quality of the processed second-perspective image can be improved.

[0200] The image generating device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or a device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device, an augmented reality / virtual reality device, a robot, a wearable device, a super mobile personal computer, a netbook or a personal digital assistant, etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0201] The image generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0202] The image generating device provided in the embodiment of the present application can implement each process implemented by each embodiment of the above-mentioned image generating method, and will not be described again here to avoid repetition.

[0203] Alternatively, if Fig.13 As shown, an embodiment of the present application further provides an electronic device 900, including a processor 901 and a memory 902, wherein the memory 902 stores programs or instructions that can be executed on the processor 901, and when the program or instructions are executed by the processor 901, the various steps of the above-mentioned image generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described again here.

[0204] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0205] Fig.14A schematic diagram of the hardware structure of an electronic device to implement an embodiment of the present application.

[0206] The electronic device 1000 includes but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010 and other components.

[0207] Those skilled in the art will appreciate that the electronic device 1000 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 1010 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Fig.14 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.

[0208] Among them, the processor 1010 is used to generate a second perspective image corresponding to the first perspective image based on the first perspective image, and the second perspective image includes a first missing area; and fill the foreground area and the first missing area in the second perspective image to obtain a second perspective background image; and extract the area image corresponding to the first missing area from the second perspective background image; and fill the area image into the first missing area in the second perspective image to obtain a processed second perspective image.

[0209] Optionally, in an embodiment of the present application, the processor 1010 is specifically used to fill the foreground area and the first missing area in the second perspective image based on the pixel value of the background area in the second perspective image to obtain the second perspective background image.

[0210] Optionally, in an embodiment of the present application, the above-mentioned processor 1010 is also used to generate a first-perspective foreground mask image based on the first-perspective foreground mask image; and to generate a second-perspective foreground mask image based on the first-perspective foreground mask image; and to determine the foreground area in the second-perspective image based on the second-perspective foreground mask image.

[0211] Optionally, in an embodiment of the present application, the above-mentioned processor 1010 is specifically used to perform pixel transformation on the first-perspective foreground mask image according to the disparity image and adopt a geometric transformation algorithm to obtain a second-perspective foreground mask image; the disparity image is used to represent the position difference of pixel points in the image between the first perspective and the second perspective.

[0212] Optionally, in an embodiment of the present application, the above-mentioned processor 1010 is specifically used to determine a first area corresponding to a first missing area in the second perspective background image based on a missing area mask image, and the missing area mask image is used to represent an area that is invisible in the first perspective image and visible in the second perspective image; and, extract the image of the first area from the second perspective background image to obtain a regional image.

[0213] Optionally, in the embodiment of the present application, the processor 1010 is further configured to generate a missing area mask image based on the first missing area in the second perspective image.

[0214] Optionally, in an embodiment of the present application, the processor 1010 is specifically used to perform pixel transformation on the first perspective image according to the disparity image by using a geometric transformation algorithm to obtain a second perspective image; the disparity image is used to represent the position difference of pixel points in the image between the first perspective and the second perspective.

[0215] Optionally, in an embodiment of the present application, the first-perspective image is a frame image in a first-perspective video, and the first-perspective video includes multiple frame images; the processor 1010 is specifically used to generate a second-perspective image corresponding to the first-perspective image based on the first-perspective image and N frame images, and the N frame images are N frame images adjacent to the first perspective image in the multiple frame images, and N is a positive integer.

[0216] Optionally, in an embodiment of the present application, the processor 1010 is further used to generate a third perspective image corresponding to the first perspective image based on the first perspective image, wherein the third perspective image includes a second missing area, and the third perspective image and the second perspective image are images of different perspectives.

[0217] The electronic device provided in the embodiment of the present application fills the foreground area and the first missing area in the second perspective image as the background to obtain the second perspective background image. The second perspective background image is a pure background image, and the regional image extracted from the second perspective background image is also a pure background image. When the regional image is used to fill the first missing area in the second perspective image, the influence of the foreground image information on the filling of the first missing area can be avoided, the filling accuracy can be improved, the consistency and coordination between the filled first missing area and the background area of ​​the second perspective image can be ensured, and the visual quality of the processed second perspective image can be improved.

[0218] It should be understood that in the embodiment of the present application, the input unit 1004 may include a graphics processor (Graphics Processing Unit, GPU) 10041 and a microphone 10042, and the graphics processor 10041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0219] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0220] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 1010.

[0221] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned image generation method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0222] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0223] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned image generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0224] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0225] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned image generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0226] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0227] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0228] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. An image generation method, characterized in that: The method comprises: Generate a second perspective image corresponding to the first perspective image according to the first perspective image, wherein the second perspective image includes the first missing area; Filling the foreground area and the first missing area in the second viewing angle image to obtain a second viewing angle background image; Extracting a region image corresponding to the first missing region from the second viewing angle background image; The first missing area in the second viewing angle image is filled with the regional image to obtain a processed second viewing angle image.

2. The method according to claim 1, characterized in that The step of filling the foreground area and the first missing area in the second viewing angle image to obtain the second viewing angle background image includes: Based on the pixel values ​​of the background area in the second viewing angle image, the foreground area and the first missing area in the second viewing angle image are filled to obtain a second viewing angle background image.

3. The method according to claim 1 or 2, characterized in that: Before filling the foreground area in the second viewing angle image, the method further includes: generating a first-perspective foreground mask image according to the first-perspective image; Generating a second-viewing angle foreground mask image according to the first-viewing angle foreground mask image; According to the second-view foreground mask image, a foreground area in the second-view image is determined.

4. The method according to claim 3, characterized in that The step of generating a second viewing angle foreground mask image according to the first viewing angle foreground mask image comprises: According to the disparity image, a geometric transformation algorithm is used to perform pixel transformation on the first perspective foreground mask image to obtain the second perspective foreground mask image; the disparity image is used to represent the position difference of pixel points in the image between the first perspective and the second perspective.

5. The method according to claim 1, characterized in that The step of extracting a region image corresponding to the first missing region from the second viewing angle background image includes: Determining a first area corresponding to the first missing area in the second perspective background image based on a missing area mask image, wherein the missing area mask image is used to represent an area that is invisible in the first perspective image and visible in the second perspective image; The image of the first area is extracted from the second viewing angle background image to obtain the area image.

6. The method according to claim 5, characterized in that The method further comprises: Based on the first missing area in the second viewing angle image, the missing area mask image is generated.

7. The method according to claim 1, characterized in that The step of generating a second perspective image corresponding to the first perspective image according to the first perspective image includes: According to the disparity image, a geometric transformation algorithm is used to perform pixel transformation on the first perspective image to obtain the second perspective image; the disparity image is used to represent the position difference of pixel points in the image between the first perspective and the second perspective.

8. The method according to claim 1, characterized in that The first-perspective image is a frame of image in a first-perspective video, and the first-perspective video includes multiple frames of image; The step of generating a second perspective image corresponding to the first perspective image according to the first perspective image includes: A second perspective image corresponding to the first perspective image is generated according to the first perspective image and N frame images, wherein the N frame images are N frame images adjacent to the first perspective image in the multiple frame images, and N is a positive integer.

9. The method according to claim 1, characterized in that: The method further comprises: A third perspective image corresponding to the first perspective image is generated according to the first perspective image, wherein the third perspective image includes a second missing area, and the third perspective image and the second perspective image are images of different perspectives.

10. An image generating device, characterized in that: The device comprises: a processing module and an acquisition module; The processing module is used to generate a second perspective image corresponding to the first perspective image according to the first perspective image, wherein the second perspective image includes the first missing area; and fill the foreground area and the first missing area in the second perspective image to obtain a second perspective background image; The acquisition module is used to extract a region image corresponding to the first missing region from the second viewing angle background image obtained by the processing module; The processing module is further used to fill the area image extracted by the acquisition module into the first missing area in the second viewing angle image to obtain a processed second viewing angle image.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the image generating method according to any one of claims 1 to 9 are implemented.