Image processing method, device and electronic device
By incorporating depth and optical flow images from multiple frames, the image processing method enhances supersampling efficiency and accuracy, addressing the limitations of existing methods by leveraging more image prior information.
Patent Information
- Application Number
- JP2025525325
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-20
- Filing Date
- 2024-02-08
- Publication Date
- 2025-12-16
AI Technical Summary
Existing supersampling methods for converting low-resolution images to high-resolution images suffer from low accuracy due to the lack of prior information about the low-resolution images, leading to inefficient and inaccurate supersampling processes.
An image processing method that involves acquiring a first image and its corresponding depth image, along with second images and their depth images and optical flow images from multiple frames, to perform splicing and optical flow reconstruction, thereby enhancing the supersampling process with more image prior information.
This approach improves supersampling efficiency and accuracy by utilizing more image information, reducing the number of pixels processed and enhancing the resolution of the supersampled images.
Smart Images

Figure 2025540591000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority from Chinese Patent Application No. 202310181542.2, filed on February 20, 2023, the entire contents of which are hereby incorporated by reference.
[0002] The disclosed embodiments relate to an image processing method, apparatus, and electronic device. [Background technology]
[0003] Supersampling allows low-resolution images in a video to be converted into high-resolution images, resulting in even higher resolution video.
[0004] Currently, a low-resolution image is input to a pre-trained supersampling model, and a supersampling process is performed on the low-resolution image based on the supersampling model to obtain a supersampled high-resolution image. However, in the above method, the supersampling model cannot obtain the prior information of the low-resolution image, and the accuracy of the supersampled image obtained by the supersampling model is low. Therefore, how to improve the accuracy of the supersampled image has become an urgent problem to be solved. Summary of the Invention [Means for solving the problem]
[0005] The present disclosure provides an image processing method, apparatus and electronic device, which can be used to solve the technical problem of low accuracy of supersampling.
[0006] According to a first aspect, the present disclosure provides an image processing method, the method comprising: acquiring a first image of a video and a first depth image corresponding to the first image; acquiring second images for the first N frames of the first image, second depth images and optical flow images corresponding to the second images for each frame, where N is an integer greater than 0; and determining a supersampled image corresponding to the first image based on the first image, the first depth image, the second images of the first N frames, N second depth images, and N optical flow images.
[0007] According to a second aspect, the present disclosure provides an image processing device including a first acquisition module, a second acquisition module, and a determination module; the first acquisition module is adapted to acquire a first image of a video and a first depth image corresponding to the first image; the second acquisition module is used to acquire second images of the first N frames of the first image, second depth images and optical flow images corresponding to the second images of each frame, where N is an integer greater than 0; The determination module is used to determine a super-sampled image corresponding to the first image based on the first image, the first depth image, the second image of the first N frames, N second depth images, and N optical flow images.
[0008] According to a third aspect, embodiments of the present disclosure provide an electronic device comprising a processor and a memory; The memory stores computer-executable instructions; The processor executes computer-executable instructions stored in the memory, causing the at least one processor to perform the first aspect and various of the image processing methods that may be involved in the first aspect.
[0009] According to a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions, when executed by a processor, realizing the first aspect and various image processing methods that may be involved in the first aspect.
[0010] According to a fifth aspect, an embodiment of the present disclosure provides a computer program product including a computer program that, when executed by a processor, implements the first aspect and various image processing methods that may be involved in the first aspect.
[0011] The present disclosure provides an image processing method, apparatus, and electronic device, in which the electronic device acquires a first image of a video and a first depth image corresponding to the first image, acquires second images of the first N frames of the first image, second depth images and optical flow images corresponding to the second images of each frame, where N is an integer greater than 0, and determines a supersampled image corresponding to the first image based on the first image, the first depth image, the second images of the first N frames, the N second depth images, and the N optical flow images. Based on the above method, the electronic device directly performs supersampling convolution on the low-resolution image, which reduces the number of pixels in the convolution process and thus improves supersampling efficiency, and the electronic device can perform supersampling processing on the first image based on image information of the first image and image information of the first N frames of images, which allows the electronic device to obtain more image prior information and thus improves the accuracy of image supersampling.
[0012] In order to more clearly explain the embodiments of the present disclosure, the following will briefly introduce the drawings that need to be used in the embodiments. It is clear that the drawings in the following description are some embodiments of the present disclosure, and those skilled in the art can further obtain other drawings based on these drawings without any creative work. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure. [Figure 3]FIG. 3 is a schematic diagram of a second image according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram of an optical flow image corresponding to the second image according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a schematic diagram of a target optical flow image according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a schematic diagram of a process for determining a third spliced image according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a structural schematic diagram of a super-sampling network according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a process diagram of an image processing method according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a schematic diagram of a method for training a supersampling network according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic diagram of the training process of a supersampling network according to an embodiment of the present disclosure. [Figure 11] FIG. 11 is a structural schematic diagram of an image processing device according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a structural schematic diagram of another image processing device according to an embodiment of the present disclosure. [Figure 13] FIG. 13 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] Reference will now be made in detail to illustrative embodiments, examples of which are illustrated in the drawings. When the following description refers to the drawings, like numbers in different drawings refer to the same or similar elements unless otherwise indicated. The embodiments described in the following illustrative examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as set forth in the appended claims.
[0015] For ease of understanding, the concepts involved in the disclosed embodiments are explained below.
[0016] Electronic device: A device with wireless transmission and reception capabilities. The electronic device can be located on land, indoors or outdoors, handheld, wearable, or vehicle-mounted. The electronic device may be a mobile phone, a tablet computer (Pad), a computer with wireless transmission and reception capabilities, virtual reality (VR) electronic device, augmented reality (AR) electronic device, a wireless terminal for industrial control, an in-vehicle electronic device, a wireless terminal for self-driving, a wireless electronic device for remote medical care, a wireless electronic device for a smart grid, a wireless electronic device for transportation safety, a wireless electronic device for a smart city, a wireless electronic device for a smart home, a wearable electronic device, or the like. The electronic device according to the disclosed embodiments may also be referred to as a terminal, user equipment (UE), access electronic device, in-vehicle terminal, industrial control terminal, UE unit, UE station, mobile station, mobile base, remote station, remote electronic device, mobile device, UE electronic device, wireless communication device, UE agent, or UE device. The electronic device may be stationary or mobile.
[0017] Optical flow reconstruction processing: Optical flow reconstruction processing is an image processing method for reconstructing an image based on the optical flow information of the image. The optical flow information between any two images can indicate the movement information of pixel positions between the two images. For example, for adjacent images A and B, an electronic device can obtain an optical flow image (an image containing optical flow information) between images A and B. The electronic device can reconstruct image B based on the optical flow image and image A, and the electronic device can also reconstruct image A based on the optical flow image and image B. For example, an electronic device can obtain an optical flow image related to image A and image B. If each value of the optical flow image is +2, it means that image B can be reconstructed by moving each pixel of image A to the right by two pixels.
[0018] As should be explained, the reconstructed image B may be the same as or different from the original image B, and the reconstructed image A may be the same as or different from the original image A, which is mainly determined by the accuracy of the optical flow information contained in the optical flow image.
[0019] In related art, a low-resolution image in a video can be converted into a high-resolution image through supersampling, thereby obtaining a higher-resolution video. However, before performing supersampling convolution on the low-resolution image, the electronic device must perform feature extraction on the low-resolution image to obtain a low-resolution feature image, then perform upsampling on the low-resolution feature image to obtain a high-resolution feature image, and finally perform supersampling convolution on the high-resolution feature image to obtain a high-resolution image. In this way, the supersampling convolution of the high-resolution feature image takes a long time, reducing the supersampling efficiency. Furthermore, the electronic device only performs supersampling on the low-resolution image of the current frame, which has little prior information about the low-resolution image obtained by the electronic device, making supersampling more difficult and further reducing the accuracy of the supersampling image.
[0020] In order to solve the technical problems in the related art, an embodiment of the present disclosure provides an image processing method, in which an electronic device acquires a first image of a video and a first depth image corresponding to the first image, acquires a second image of the first N frames of the first image, a second depth image and an optical flow image corresponding to the second image of each frame, performs a splicing process on the first image and the first depth image to obtain a first spliced image, performs a splicing process on the second image of each frame and the second depth image corresponding to the second image to obtain N second spliced images, and determines a supersampled image corresponding to the first image based on the first spliced image, the N second spliced images, and the N optical flow images. In this way, the electronic device performs a supersampling convolution process on the low-resolution image, thereby shortening the time required for image supersampling and improving supersampling efficiency. Furthermore, because the first spliced image and the N second spliced images contain a lot of image prior information about the first image, the electronic device can accurately perform the supersampling process on the first image, further improving the accuracy of image supersampling.
[0021] Hereinafter, an application scenario of the disclosed embodiment will be described with reference to FIG.
[0022] FIG. 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure. Referring to FIG. 1, the scenario includes an electronic device, a rendered image of a previous first frame, and a rendered image of a current frame. A stitching process is performed on the rendered image of the previous first frame and a depth image B corresponding to the rendered image of the previous first frame, and optical flow processing is performed on the stitched image to obtain a stitched image of the rendered image of the current frame reconstructed through optical flow and depth image a. Depth image a is the depth image of the current frame for which optical flow reconstruction is performed based on depth image B. A stitching process is performed on the rendered image of the current frame and a depth image A corresponding to the rendered image of the current frame to obtain a stitched image of the rendered image of the current frame and depth image A.
[0023] 1 , the electronic device receives the optical flow reconstructed joint image of the current frame's rendered image and depth image a, and the joint image of the current frame's rendered image and depth image A, and the electronic device can output a supersampled image of the current frame's rendered image, where the resolution of the supersampled image of the current frame is greater than the resolution of the current frame's rendered image. In this way, the electronic device can perform supersampling on the current frame's rendered image based on more image information, further improving the accuracy of the supersampling, and the electronic device can perform supersampling on a low-resolution image, so that fewer pixels are processed by the electronic device, further improving the supersampling efficiency.
[0024] However, FIG. 1 only shows an example of an application scene of the disclosed embodiment, and does not limit the application scene of the disclosed embodiment.
[0025] The following specific examples will be used to describe in detail the technical solutions of the present disclosure and how they can be used to solve the above technical problems. Some of the following specific examples can be combined with each other, and the same or similar concepts or processes may not be described again in a specific example. The following examples will be described with reference to the drawings.
[0026] 2 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure. Referring to FIG. 2, the method may include steps S201 to S203.
[0027] S201, obtain a first image of a video and a first depth image corresponding to the first image.
[0028] The implementation body of the embodiments of the present disclosure may be an electronic device or an image processing device provided in the electronic device. The image processing device may be realized based on software, or the image processing device may be realized based on a combination of software and hardware, and the embodiments of the present disclosure are not limited thereto.
[0029] The video may be a rendered video. For example, the video may be a virtual video, a virtual animation, or the like rendered by a rendering machine. The video may also be other videos, and the embodiments of the present disclosure are not limited thereto. Optionally, the first image may be an image of any frame of the video. For example, the first image may be an image of the second frame, the third frame, or the fourth frame of the video.
[0030] Optionally, the first depth image may be a depth image of the first image. The first depth image may indicate a distance between a pixel in the first image and a camera. For example, the first image includes an object, and if the distance between the object and the camera is small, the color of the object in the first depth image corresponding to the first image is dark; if the distance between the object and the camera is large, the color of the object in the first depth image corresponding to the first image is dark. The distance between each pixel in the first image and the camera can be determined based on the first depth image.
[0031] Alternatively, the electronic device may receive video output by a rendering machine. For example, the rendering machine may render a virtual video, and the electronic device may be connected to the rendering machine, so that after the rendering machine generates the virtual video, the rendering machine can transmit the virtual video to the electronic device. Alternatively, the electronic device may receive video transmitted by a server. For example, the electronic device may be connected to the server, so that after the server collects the rendered video, the server can transmit the rendered video to the electronic device.
[0032] It should be noted that the electronic device may acquire the video based on other possible implementations, and the embodiments of the present disclosure are not limited thereto.
[0033] Optionally, the electronic device can process the first image based on an image identification algorithm and further obtain a first depth image corresponding to the first image. For example, the electronic device can process the first image using a pre-trained depth image acquisition model, and the depth image acquisition model can output the first depth image corresponding to the first image. For example, when a rendering machine renders an image of each frame in a video, it can render a depth image corresponding to the image of each frame and further transmit the depth image corresponding to the image of each frame to the electronic device.
[0034] As should be explained, the electronic device can also obtain the depth image of the first image based on other possible implementation methods (for example, after the server collects the video, it can extract the depth image corresponding to the image of each frame in the video and send the first image and the first depth image to the electronic device), and the embodiments of the present disclosure are not limited thereto.
[0035] S202, obtain second images of the first N frames of the first image, second depth images and optical flow images corresponding to the second images of each frame.
[0036] The second image may be an image of the first N frames of the first image in the video, where N is an integer greater than 0. Optionally, the electronic device can acquire the second image in the video based on the first image and the value of N. For example, if N is 1, the second image may be an image of the first frame before the first image; if N is 2, the second image may be an image of the first frame and an image of the second frame before the first image; if N is 3, the second image may be an image of the first frame, an image of the second frame, and an image of the third frame before the first image; if N is 4, the second image may be an image of the first frame, an image of the second frame, an image of the third frame, and an image of the fourth frame before the first image.
[0037] The second image will be described below with reference to FIG.
[0038] 3 is a schematic diagram of a second image according to an embodiment of the present disclosure. Referring to FIG. 3, a video is included. The video includes image 1, image 2, image 3, image 4, and image 5, and the images of five frames in the video are arranged in the playback order of the video. If N is 2 and the first image is image 3, the second image may be the image of two frames before image 3, and the second image may include image 1 and image 2.
[0039] The second depth image may be a depth image corresponding to the second image. For example, if the number of second images is one, the number of second depth images is one, and if the number of second images is two, the number of second depth images is two. For example, if the electronic device acquires second image A and second image B, the electronic device can acquire a second depth image corresponding to second image A and a second depth image corresponding to second image B.
[0040] As is necessary to explain, the method for the electronic device to acquire the second depth image can refer to step S201, and the disclosed embodiment will not describe it in detail.
[0041] The optical flow image corresponding to the second image can indicate pixel changes between the image of the frame following the second image and the second image. For example, if the second image contains 100 pixels, the image of the frame following the second image also contains 100 pixels, and the optical flow image corresponding to the second image also contains 100 values. For example, if each value in the optical flow image corresponding to the second image is +2, it means that each pixel in the second image is shifted two pixels to the right to obtain the image of the frame following the second image. For example, if the value of position 1 in the optical flow image corresponding to the second image is -2, it means that the pixel at position 1 in the second image is shifted two pixels to the left in the image of the frame following the second image, and if the value of position 2 in the optical flow image corresponding to the second image is +10, it means that the pixel at position 2 in the second image is shifted ten pixels to the right in the image of the frame following the second image.
[0042] Hereinafter, an optical flow image corresponding to the second image will be described with reference to FIG.
[0043] 4 is a schematic diagram of an optical flow image corresponding to a second image according to an embodiment of the present disclosure. Referring to FIG. 4, the image includes a second image and an optical flow image corresponding to the second image. Taking the pixel at point A in the second image as an example, the corresponding value in the optical flow image at the position of point A is +2. Therefore, after processing the second image using the optical flow image, the pixel at point A is shifted two pixels to the right to obtain the next frame of the second image. In the next frame of the second image, the pixel at point A in the second image is shifted to the position of point a in the next frame of the image.
[0044] Optionally, the number of second images is the same as the number of optical flow images. For example, when the electronic device acquires one second image, the electronic device can acquire an optical flow image corresponding to the second image, and perform an optical flow reconstruction process on the second image using the optical flow image to obtain an image of a frame next to the second image. When the electronic device acquires two second images, the electronic device can acquire two optical flow images corresponding to the two second images, and perform an optical flow reconstruction process on the second image of the previous first frame based on the optical flow image corresponding to the second image of the previous first frame to obtain an optical flow reconstruction image of the current frame. Based on the optical flow image corresponding to the second image of the previous second frame, the electronic device can perform an optical flow reconstruction process on the image of the previous second frame to obtain an optical flow reconstruction image of the previous first frame. Finally, based on the optical flow image of the previous first frame, the electronic device can perform an optical flow reconstruction process on the optical flow reconstruction image of the previous first frame to obtain an optical flow reconstruction image of the current frame.
[0045] Optionally, the electronic device can obtain an optical flow image corresponding to the second image based on an optical flow algorithm. For example, the electronic device can process the video based on an optical flow algorithm and further obtain an optical flow image (the optical flow image is an optical flow image corresponding to the previous frame image) between every two frame images (the images may be any two frame images, and the embodiments of the present disclosure are not limited thereto), the electronic device can store the optical flow image in a cache memory, and when the electronic device obtains the second image, the electronic device can determine the optical flow image corresponding to the second image in the cache memory.
[0046] As needs to be explained, the electronic device can also acquire the optical flow image corresponding to the second image based on other possible implementation methods (for example, when rendering a video based on a rendering machine, the rendering machine can render the optical flow image between every two frame images and store it in a cache memory, and when the electronic device requests to acquire the optical flow image corresponding to the second image, the rendering machine can acquire the optical flow image in the cache memory and send the optical flow image to the electronic device), but the embodiments of the present disclosure are not limited thereto.
[0047] S203: Determine a super-sampled image corresponding to the first image based on the first image, the first depth image, the second images of the first N frames, the N second depth images and the N optical flow images.
[0048] The supersampled image is an image obtained after a supersampling process is performed on the first image. Optionally, the image resolution of the supersampled image is greater than the image resolution of the first image. For example, if the first image includes 100 pixels, after the first image is supersampling-performed, the supersampled image corresponding to the first image may include 200 pixels. Thus, the resolution of the supersampled image corresponding to the first image is greater than the resolution of the first image.
[0049] Specifically, the electronic device can determine a super-sampled image corresponding to a first image according to the following possible implementation: perform a splicing process on the first image and the first depth image to obtain a first spliced image, perform a splicing process on the second image of each frame and a second depth image corresponding to the second image to obtain N second spliced images, and determine a super-sampled image corresponding to the first image according to the first spliced image, the N second spliced images, and the N optical flow images.
[0050] The first joined image may be an image obtained by joining the first image and the first depth image. For example, the electronic device may join the first image and the first depth image vertically to obtain the first depth image, or the electronic device may join the first image and the first depth image horizontally to obtain the first depth image. The electronic device may also join the first image and the first depth image based on other possible implementation methods, and the embodiments of the present disclosure are not limited thereto.
[0051] The second joined image may be an image obtained by joining the second image and the second depth image corresponding to the second image. For example, when the electronic device acquires the second image A, the second image B, the second depth image 1, and the second depth image 2, and the second depth image 1 is the depth image of the second image A and the second depth image 2 is the depth image of the second image B, the electronic device can perform a joining process on the second image A and the second depth image 1 to obtain one second joined image, and can join the second image B and the second depth image 2 to obtain another second joined image.
[0052] As is necessary to explain, the electronic device may stitch the second image and the second depth image vertically, or the electronic device may stitch the second image and the second depth image horizontally, or the electronic device may stitch the second image and the second depth image based on other possible implementation manners, and the embodiments of the present disclosure are not limited thereto. In this way, the first stitched image and the second stitched image can fuse the depth information of the image and improve the accuracy of image supersampling.
[0053] Specifically, when the electronic device determines a supersampled image corresponding to the first image based on the first joined image, N second joined images, and N optical flow images, it determines at least one target optical flow image associated with each second joined image in the N optical flow images, and determines the supersampled image based on the N second joined images, at least one target optical flow image associated with each second joined image, and the first joined image.
[0054] The target optical flow image is used to perform optical flow reconstruction processing on the second joined image to reconstruct the second image in the second joined image into the first image, and to reconstruct the second depth image into the first depth image.
[0055] The target optical flow image will be described below with reference to FIG.
[0056] 5 is a schematic diagram of a target optical flow image according to an embodiment of the present disclosure. Referring to FIG. 5, the target optical flow image includes image 1, image 2, image 3, image 4, optical flow image A between image 1 and image 2, optical flow image B between image 2 and image 3, and optical flow image C between image 3 and image 4. If N is 3 and the first image is 4, it is determined that the second image includes image 1, image 2, and image 3.
[0057] 5, the target optical flow image a corresponding to the second spliced image of image 1 includes optical flow image A, optical flow image B, and optical flow image C. The target optical flow image b corresponding to the second spliced image of image 2 includes optical flow image B and optical flow image C. The target optical flow image c corresponding to the second spliced image of image 3 includes optical flow image C.
[0058] Optionally, the electronic device determines the supersampled image based on the N second joined images, at least one target optical flow image associated with each second joined image, and the first joined image, specifically by performing an optical flow reconstruction process on the second joined images based on at least one target optical flow image associated with each second joined image to obtain N third joined images, and determining the supersampled image based on the N third joined images and the first joined image.
[0059] The third spliced image is an image obtained after the second spliced image has undergone an optical flow reconstruction process based on at least one target optical flow image. For example, when N is 1, the second image is an image of a frame previous to the first image (current frame), and the electronic device performs an optical flow reconstruction process on the second spliced image corresponding to the image of the previous frame based on the second spliced image corresponding to the image of the previous frame and one optical flow image corresponding to the image of the previous frame to obtain a third spliced image. When N is 2, the second image is an image of a first frame previous to the first image and an image of a second frame previous to the first image, and the electronic device performs an optical flow reconstruction process on the second spliced image corresponding to the image of the previous first frame based on the second spliced image corresponding to the image of the previous first frame and the optical flow image corresponding to the image of the previous first frame to obtain one third spliced image. When N is 2, the second image is an image of a first frame previous to the first image and an image of a second frame previous to the first image, and the electronic device performs an optical flow reconstruction process on the second spliced image corresponding to the image of the previous second frame based on the second spliced image corresponding to the image of the previous second frame, the optical flow image corresponding to the image of the previous second frame, and the optical flow image corresponding to the image of the previous first frame to obtain another third spliced image.
[0060] Hereinafter, the process of determining the third joined image will be described with reference to FIG.
[0061] 6 is a schematic diagram of a process for determining a third joined image according to an embodiment of the present disclosure. Referring to FIG. 6, an electronic device (not shown in FIG. 6) can perform optical flow processing on the joined image between the image of the previous second frame and a second depth image A corresponding to the image of the previous second frame based on an optical flow image corresponding to the image of the previous second frame, to obtain an optical flow-reconstructed joined image between the image of the previous first frame and a second depth image B, where the second depth image B is a depth image corresponding to the image of the first frame before optical flow reconstruction is performed based on the second depth image A and the optical flow image corresponding to the image of the previous second frame.
[0062] 6, the electronic device performs optical flow processing on a joint image between the optically flow-reconstructed image of the previous first frame and a second depth image B based on an optical flow image corresponding to the image of the previous first frame, thereby obtaining a joint image between the optically flow-reconstructed image of the current frame and a second depth image C, where the second depth image C is a depth image (third joint image) corresponding to the image of the current frame for which optical flow reconstruction is performed based on the second depth image B and the optical flow image corresponding to the image of the previous first frame. In this way, the third joint image includes image information of the current frame optically flow-reconstructed using the images of the first N frames. Therefore, the third joint image includes more image prior information of the current frame, which can further improve the accuracy of supersampling.
[0063] Optionally, the electronic device can determine the super-sampled image according to the following possible implementation: input N third combined images and the first combined image into a super-sampling network to obtain a super-sampled image corresponding to the first image, and the super-sampling network is used to perform super-sampling processing on the image; the last layer network of the super-sampling network is a pixel reordering network, and the pixel reordering network is used to perform up-sampling processing on the feature image;
[0064] The structure of the supersampling network will be described below with reference to FIG.
[0065] FIG. 7 is a schematic diagram of the structure of a supersampling network according to an embodiment of the present disclosure. Referring to FIG. 7, the supersampling network includes a U-net structure consisting of 10 convolutional layers and two up / downsampling layers. For example, the first convolutional layer has 64 channels, and the second convolutional layer has 32 channels. Specifically, the upsampling layer output by the supersampling network is a pixel rearrangement layer implemented based on pixel shuffle. The upsampling layer receives the first and third low-resolution joined images and outputs a high-resolution supersampled image. In this way, the resolution of the feature image processed by the intermediate layer of the supersampling network is reduced, which further shortens the processing time of the supersampling network and improves the supersampling efficiency. Furthermore, the supersampling network can obtain more image information (depth information, pixel information, etc.), thereby improving the accuracy of supersampling.
[0066] As should be noted, a supersampling network can also be another lightweight encoder-decoder network with a similar structure, by removing the batch normalization layer in the lightweight encoder-decoder network and replacing the output layer with a pixel shuffle layer.
[0067] The disclosed embodiments provide an image processing method, in which an electronic device can acquire a first image of a video and a first depth image corresponding to the first image, the electronic device can acquire a second image of the first image for the first N frames, a second depth image and an optical flow image corresponding to the second image for each frame, perform a splicing process on the first image and the first depth image to obtain a first spliced image, perform a splicing process on the second image of each frame and a second depth image corresponding to the second image to obtain N second spliced images, and determine a supersampled image corresponding to the first image based on the first spliced image, the N second spliced images and the N optical flow images. In this way, the supersampling network of the electronic device can directly process low-resolution images, and the resolution of the feature images processed by the intermediate layers of the supersampling network is low, further improving supersampling efficiency, and the first spliced image and the N second spliced images contain more depth information and pixel information of the first image, so that the electronic device can accurately perform supersampling on the first image, further improving the accuracy of the supersampled image.
[0068] Based on the embodiment shown in FIG. 2, the steps of the image processing method will be described below with reference to FIG.
[0069] FIG. 8 is a schematic diagram of an image processing method according to an embodiment of the present disclosure. Referring to FIG. 8, the system includes a rendering machine. The rendering machine can render a low-resolution RGB map (having a size of H×W×3) of a current frame and a first depth image corresponding to the low-resolution RGB map of the current frame. The rendering machine can render a low-resolution RGB map (having a size of H×W×3) of a previous frame, a second depth image corresponding to the low-resolution RGB map of the previous frame, and an optical flow image (having a size of H×W×2). The optical flow image can indicate the displacement of corresponding pixels between the current frame and the previous frame, where a first channel represents the offset direction and magnitude of the image in the x direction, and a second channel represents the offset and magnitude of the image in the y direction.
[0070] 8, a splicing process is performed on the low-resolution RGB image of the current frame and the first depth image to obtain a first spliced image, which includes the low-resolution RGB image of the current frame and the first depth image, and a splicing process is performed on the low-resolution RGB image of the previous frame and the second depth image to obtain a second spliced image, which includes the low-resolution RGB image of the previous frame and the second depth image.
[0071] Referring to FIG. 8, optical flow processing is performed on the second joint image based on the optical flow image to obtain a third joint image. The third joint image includes an RGB map of the current frame reconstructed by optical flow and a second depth image of the current frame reconstructed by optical flow. The first joint image and the second joint image are input to a supersampling network, which can output a high-resolution RGB map (2H×2W×3) of the current frame. In this way, the resolution of the feature image processed by the intermediate layer of the supersampling network is low, further improving supersampling efficiency. Furthermore, the first joint image and the third joint image contain more depth information and pixel information than the first image. Therefore, the electronic device can accurately perform supersampling processing on the first image, further improving the accuracy of the supersampling image.
[0072] According to any one of the above embodiments, the above image processing method further includes a method for training a supersampling network, which will be described below with reference to FIG.
[0073] 9 is a schematic diagram of a method for training a supersampling network according to an embodiment of the present disclosure. Referring to FIG. 9, the method process includes steps S901 to S903.
[0074] S901, obtain multi-frame sample images in a sample video, and obtain a sample depth image and a sample optical flow image corresponding to each frame of the sample image.
[0075] Alternatively, the sample video may be a rendered video. For example, the sample video may be a game animation video, a virtual character video, etc. The electronic device may receive the sample video sent from another device. For example, the electronic device may be connected to a rendering machine, so that the rendering machine can render the sample video and then send the sample video to the electronic device.
[0076] As is necessary to explain, during the training phase, the electronic device can acquire a sample image of each frame of the sample video, a sample depth image and a sample optical flow image corresponding to the sample image of each frame, and the acquisition method of the sample image, the sample depth image and the sample optical flow image corresponding to the sample image can refer to the embodiment shown in Figure 2, and the disclosed embodiment will not be described in detail here.
[0077] S902, obtain a sample super-sampled image corresponding to the sample image of each frame.
[0078] A sampled supersampled image is a supersampled image corresponding to a sampled image. The resolution of the sampled image is smaller than that of the sampled supersampled image corresponding to the sampled image. For example, an electronic device can capture two images with the same image content, one with a lower resolution and the other with a higher resolution, and the image with the lower resolution can be the sampled image, and the image with the higher resolution can be the sampled supersampled image corresponding to the sampled image. For example, in actual application, the image of each frame of a standard-resolution sample video can be the sample image, and the image of each frame of a super-resolution sample video corresponding to the standard-resolution sample video can be the sampled supersampled image.
[0079] As should be understood, the electronic device can also obtain sample super-sampled images corresponding to the sample images of each frame based on other possible implementation methods, and the embodiments of the present disclosure are not limited thereto.
[0080] S903, based on the multi-frame sample images, the first M frame images of the sample images of each frame, the sample depth images and sample super-sampled images of the sample images of each frame, and the sample depth images and sample optical flow images corresponding to the first M frame images of the sample images of each frame, a super-sampling network is trained.
[0081] M is an integer greater than or equal to N. Specifically, for a sample image of a current frame, the electronic device can train a supersampling network according to the following possible implementation: perform splicing processing on the sample image and the sample depth image of the sample image to obtain a sample first spliced image, perform splicing processing on the first M frame images of the sample image and the sample depth image corresponding to the first M frame images of the sample image to obtain a sample second spliced image of M frames, perform optical flow processing on the sample second spliced image of M frames according to the sample optical flow image of M frames corresponding to the first M frame images of the sample image to obtain a sample third spliced image of M frames, and train a supersampling network according to the sample first spliced image, the sample third spliced image of M frames, and the sample supersampling image corresponding to the sample image.
[0082] The sample second spliced image includes the first M frame images of the sample image of one frame and the sample depth image corresponding to the first M frame images of the sample image of the frame. As is necessary to explain, the splicing process for the sample image and the sample depth image of the sample image, and the splicing process for the first M frame images of the sample image and the sample depth image corresponding to the first M frame images of the sample image can be referred to the embodiment shown in FIG. 2, and the embodiment of the present disclosure will not be described in detail.
[0083] The sample third joint image is an image after optical flow reconstruction is performed on the sample second joint image. As is necessary to explain, the method of the electronic device acquiring the sample third joint image can refer to the embodiment shown in FIG. 2, and the embodiment of the present disclosure will not describe it in detail.
[0084] Optionally, the electronic device may train a supersampling network based on the sample first spliced image, the sample third spliced image of the M frame, and the sample supersampling image corresponding to the sample image, specifically by processing the sample first spliced image and the sample third spliced image of the M frame based on the supersampling network to obtain a predicted supersampling image, and training the supersampling network based on the loss between the predicted supersampling image and the sample supersampling image. For example, the electronic device may construct a loss function based on the predicted supersampling image predicted by the supersampling network and the actual sample supersampling image, and further update network parameters in the supersampling network through the loss function.
[0085] Optionally, the electronic device can further assist in training the supersampling network based on the judgment device. For example, the judgment device can be configured with four convolutional layers, and the judgment device can judge whether the predicted supersampling image output from the supersampling network is true or false, and further train the supersampling network based on the judgment result. For example, in actual application, the true / false judgment device can determine that the image output from the supersampling network is a predicted image, and during training, the image output from the supersampling network approaches the actual image (i.e., the judgment device cannot determine whether the image output from the supersampling network is a predicted image or an actual image).
[0086] Optionally, after the electronics acquires the multi-frame sample images in the sample video, the electronics can further pre-process the multi-frame sample images, where the pre-processing may include at least one of blur processing, noise processing, compression distortion processing, and ringing distortion processing.
[0087] The blurring process may include Gaussian blurring process and motion blurring process, the noise process may include Gaussian noise processing, Poisson noise processing, color noise processing and gray noise processing, the compression distortion process may be JPEG (Joint Photographic Experts Group) compression processing, and the ringing distortion process may be one or more of ringing distortion (sinc filter). In this way, processing the sample image with the above preprocessing method can simulate the original low-resolution image and further improve the robustness of the supersampling network.
[0088] As needs to be explained, in the training stage, SMAA anti-aliasing processing can be performed on the sample super-sampled images acquired by the electronic device, and thus the anti-aliasing effect of the super-sampling network can be enhanced.
[0089] As should be noted, in actual application, the electronic device can randomly pre-process the multi-frame sample images. For example, after the electronic device acquires the multi-frame sample images, the electronic device pre-processes 30% of the sample images and does not pre-process the remaining 70% of the sample images. When pre-processing the sample images, the pre-processing method for each sample image can be a random selection method (for example, randomly select one or more of the above pre-processing methods to process the sample image). The electronic device can also pre-process the sample images according to other methods, and the embodiments of the present disclosure are not limited thereto.
[0090] The training process of the supersampling network will now be described with reference to FIG.
[0091] 10 is a schematic diagram of a training process of a supersampling network according to an embodiment of the present disclosure. Referring to FIG. 10, a system includes a rendering machine. The rendering machine can render a low-resolution sample image A of a current frame and a sample first depth image corresponding to sample image A. The rendering machine can render a low-resolution sample image B of a previous frame, a sample second depth image corresponding to sample image B, and a sample optical flow image.
[0092] 10, sample image A is pre-processed, and the pre-processed sample image A and the first depth image are subjected to a joining process to obtain a first joined image, which includes the pre-processed low-resolution sample image A and the sample first depth image. Sample image B is pre-processed, and the pre-processed sample image B and the sample second depth image are subjected to a joining process to obtain a second joined image, which includes the pre-processed low-resolution sample image B and the sample second depth image.
[0093] 10, optical flow processing is performed on the second joint image based on the optical flow image to obtain a third joint image. The third joint image includes a sample image B of the current frame reconstructed by optical flow and a sample second depth image of the current frame reconstructed by optical flow. The first joint image and the second joint image are input to a supersampling network, which can output a predicted high-resolution image.
[0094] 10, an electronic device (not shown in FIG. 10) can update parameters in the supersampling network based on the loss between the predicted high-resolution image and the sample supersampling image corresponding to the low-resolution sample image A. The electronic device can process the predicted high-resolution image based on the determination device, and update parameters in the supersampling network based on the output result by the determination device.
[0095] The disclosed embodiments provide a method for training a supersampling network, which involves obtaining multiple sample images of frames in a sample video, sample depth images and sample optical flow images corresponding to the sample images of each frame, obtaining sample super-sampled images corresponding to the sample images of each frame, and training the supersampling network based on the sample images of multiple frames, the first M frame images of each sample image of each frame, the sample depth images and sample super-sampled images of each sample image of each frame, and the sample depth images and sample optical flow images corresponding to the first M frame images of each sample image of each frame. In this way, the supersampling network can directly process low-resolution images, so the resolution of the feature images processed by the intermediate layer of the supersampling network is low, which further improves the training efficiency of the supersampling network, and pre-processing the sample images can simulate the original low-resolution images, thereby improving the robustness of the supersampling network.
[0096] 11 is a structural schematic diagram of an image processing device according to an embodiment of the present disclosure. Referring to FIG. 11, the image processing device 110 includes a first acquisition module 111, a second acquisition module 112, and a determination module 113; The first acquisition module 111 is used to acquire a first image of a video and a first depth image corresponding to the first image; the second acquisition module 112 is used to acquire second images of the first N frames of the first image, second depth images and optical flow images corresponding to the second images of each frame, where N is an integer greater than 0; The determination module 113 is used to determine a super-sampled image corresponding to the first image based on the first image, the first depth image, the second image of the first N frames, N second depth images, and N optical flow images.
[0097] According to one or more embodiments of the present disclosure, the determination module 113 specifically: performing a joining process on the first image and the first depth image to obtain a first joined image; performing a joining process on the second image of each frame and a second depth image corresponding to the second image to obtain N second joined images; determining a super-sampled image corresponding to the first image based on the first joined image, the N second joined images, and the N optical flow images.
[0098] According to one or more embodiments of the present disclosure, the determination module 113 specifically: determining a target optical flow image of at least one frame associated with each second joined image among the N optical flow images, the target optical flow image being used to reconstruct a second image in the second joined image into the first image and reconstruct the second depth image into the first depth image by performing an optical flow reconstruction process on the second joined image; determining the supersampled image based on the N second joined images, at least one target optical flow image associated with each of the second joined images, and the first joined image.
[0099] According to one or more embodiments of the present disclosure, the determination module 113 specifically: performing an optical flow reconstruction process on the second joined images based on at least one target optical flow image associated with each of the second joined images to obtain N third joined images; determining the super-sampled image based on the N third joined images and the first joined image.
[0100] According to one or more embodiments of the present disclosure, the determination module 113 specifically: A supersampling network is used to input the N third joined images and the first joined image to obtain a supersampling image corresponding to the first image; The supersampling network is used to perform supersampling processing on an image.
[0101] According to one or more embodiments of the present disclosure, the network at the last layer of the supersampling network is a pixel reordering network, which is used to perform upsampling processing on the feature image.
[0102] The image processing device provided by the disclosed embodiment can be used to implement the technical solutions of the above method embodiment, and the realization principles and technical effects are similar, so this embodiment will not be described in detail here.
[0103] 12 is a structural schematic diagram of another image processing device according to an embodiment of the present disclosure. Based on the embodiment shown in FIG. 11, referring to FIG. 12, the image processing device 110 further includes a training module 114, which includes: Obtaining sample images of multiple frames in a sample video, a sample depth image and a sample optical flow image corresponding to the sample images of each frame; obtaining a sample supersampled image corresponding to the sample image of each frame; The supersampling network is trained based on the multi-frame sample images, the first M frame images of the sample images of each frame, the sample depth images and sample supersampling images of the sample images of each frame, and the sample depth images and sample optical flow images corresponding to the first M frame images of the sample images of each frame, where M is an integer greater than or equal to N.
[0104] According to one or more embodiments of the present disclosure, the training module 114 further comprises: used to preprocess the multi-frame sample images; The pre-processing includes at least one of blurring processing, noise processing, compression distortion processing, and ringing distortion processing.
[0105] The image processing device provided by the disclosed embodiment can be used to implement the technical solutions of the above method embodiment, and the realization principles and technical effects are similar, so this embodiment will not be described in detail here.
[0106] FIG. 13 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. FIG. 13 shows a structural schematic diagram of an electronic device 1300 for implementing an embodiment of the present disclosure. The electronic device includes, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (abbreviated as PDAs), tablet computers (Portable Android Devices (abbreviated as PADs), portable multimedia players (abbreviated as PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and disk-type computers. The electronic device shown in FIG. 13 is merely an example and does not limit the functionality and scope of use of the embodiment of the present disclosure.
[0107] 13, electronic device 1300 may include a processing unit (e.g., a central processor, a graphics processor, etc.) 1301, which can perform various appropriate operations and processes based on programs stored in read only memory (abbreviated as ROM) 1302 or programs loaded from storage device 1308 into random access memory (abbreviated as RAM) 1303. RAM 1303 further stores various programs and data necessary for the operation of electronic device 1300. Processing unit 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.
[0108] Typically, input devices 1306, including, for example, a touch screen, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1307, including, for example, a liquid crystal display (LCD), speaker, oscillator, etc.; storage devices 1308, including, for example, a tape, hard disk, etc.; and communication devices 1309 may be connected to the I / O interface 1305. The communication devices 1309 may allow the electronic device 1300 to communicate wirelessly or via wires with other devices to exchange data. While FIG. 13 illustrates electronic device 1300 having various devices, it should be understood that it need not embody or include all of the devices shown. Instead, it may embody or include more or fewer devices.
[0109] In particular, according to embodiments of the present disclosure, the processes described with reference to the flowcharts are implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program embodied on a computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such embodiments, the computer program is downloaded and installed from a network via the communication device 1309, or installed from the storage device 1308, or installed from the ROM 1302. When executed by the processing device 1301, the computer program performs the functions defined in the methods of the embodiments of the present disclosure.
[0110] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media include, but are not limited to, an electrical connection with one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program for use in or in connection with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the computer-readable program code is included. Such propagated data signals may take several forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may transmit, propagate, or carry a program used by or in connection with an instruction execution system, apparatus, or device. Program code contained in a computer-readable medium may be transmitted over any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0111] The computer-readable medium may be included in the electronic device, or may exist separately from the electronic device.
[0112] The computer-readable medium includes one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods illustrated in the above embodiments.
[0113] The disclosed embodiments provide a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, realize various of the image processing methods that may be involved in the embodiments.
[0114] An embodiment of the present disclosure provides a computer program product including a computer program that, when executed by a processor, implements the various image processing methods that may be involved in the above embodiments.
[0115] Computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user computer, partially on the user computer, as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0116] The flowcharts and block diagrams in the figures illustrate possible organizational structures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions shown in the blocks may be executed in a different order than that shown in the figures. For example, two blocks shown in sequence may actually be executed substantially in parallel, or may be executed in the reverse order depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0117] The units described in the disclosed embodiments may be implemented in a software manner or a hardware manner, and the names of the units do not limit the units in some circumstances, for example, the first obtaining unit may be further described as "a unit for obtaining at least two Internet Protocol addresses."
[0118] The functionality described herein above may be performed, at least in part, by one or more hardware logic elements. For example, without limitation, typical types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0119] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program used in or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include one or more wire-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] It should be noted that the modifications "one" and "multiple" referred to in this disclosure are exemplary and not limiting. Those skilled in the art will understand that "one or more" should be understood unless the context clearly dictates otherwise.
[0121] The names of messages or information exchanged between devices in the disclosed embodiments are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0122] It is to be understood that, before using the technical solutions disclosed in each embodiment of the present disclosure, the types, scope of use, and usage scenarios of personal information related to the present disclosure should be properly informed to users in accordance with relevant laws and regulations, and their consent should be obtained.
[0123] For example, in response to receiving a user's active request, prompt information may be sent to the user to clearly prompt the user that the requested operation requires the acquisition and use of the user's personal information. This allows the user to freely choose whether to provide personal information to software or hardware, such as an electronic device, application program, server, or storage medium, that performs the operation of the technical solution of the present disclosure, based on the prompt information. As an optional, non-limiting implementation method, the method of sending prompt information to the user in response to receiving a user's active request may be, for example, a pop-up window, and the prompt information may be presented in text form in the pop-up window. The pop-up window may also include a selection control for the user to select "Agree" or "Do not agree" to provide personal information to the electronic device. It should be understood that the above notification and user consent acquisition process is merely exemplary and does not limit the implementation method of the present disclosure. Other methods that comply with relevant laws and regulations may also be applicable to the implementation method of the present disclosure.
[0124] It should be understood that the data related to the present technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of corresponding laws, regulations and related provisions. The data may include information, parameters, messages, etc., such as flow switching instruction information.
[0125] The above description merely describes the preferred embodiments and applied technical principles of the present disclosure. Those skilled in the art will understand that the scope of disclosure contained in the present disclosure is not limited to the technical solution consisting of a specific combination of the above technical features, but also covers other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the idea of the above disclosure, such as, for example, a technical solution formed by replacing the above features with technical features having similar functions disclosed in the present disclosure (but not limited to these).
[0126] Additionally, although operations are described in a particular order, this should not be understood as requiring that the operations be performed in the particular order or sequence shown. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although the above description includes several specific implementation details, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.
[0127] Although the present subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. 1. An image processing method, comprising: acquiring a first image of a video and a first depth image corresponding to the first image; acquiring second images for the first N frames of the first image, second depth images and optical flow images corresponding to the second images for each frame, where N is an integer greater than 0; and determining a supersampled image corresponding to the first image based on the first image, the first depth image, the second images of the first N frames, N second depth images, and N optical flow images.
2. The step of determining a super-sampled image corresponding to the first image based on the first image, the first depth image, the second image of the first N frames, N second depth images, and N optical flow images includes: performing a joining process on the first image and the first depth image to obtain a first joined image; performing a stitching process on the second image of each frame and a second depth image corresponding to the second image to obtain N second stitched images; and determining a supersampled image corresponding to the first image based on the first joined image, the N second joined images, and the N optical flow images.
3. The step of determining a super-sampled image corresponding to the first image based on the first joint image, the N second joint images, and the N optical flow images includes: determining a target optical flow image of at least one frame associated with each second joined image in the N optical flow images, the target optical flow image being used to reconstruct a second image in the second joined image into the first image and to reconstruct the second depth image into the first depth image by performing an optical flow reconstruction process on the second joined image; and determining the supersampled image based on the N second joined images, at least one target optical flow image associated with each of the second joined images, and the first joined image.
4. The step of determining the supersampled image based on the N second joined images, at least one target optical flow image associated with each of the second joined images, and the first joined image includes: performing an optical flow reconstruction process on the second joined images based on at least one target optical flow image associated with each of the second joined images to obtain N third joined images; and determining the supersampled image based on the N third spliced images and the first spliced image.
5. The step of determining the super-sampled image based on the N third joined images and the first joined image includes: inputting the N third spliced images and the first spliced image into a supersampling network to obtain a supersampled image corresponding to the first image; The method of claim 4 , wherein the supersampling network is used to perform supersampling operations on an image.
6. The method of claim 5 , wherein the last layer network of the supersampling network is a pixel reordering network, and the pixel reordering network is used to perform upsampling processing on the feature image.
7. The supersampling network is obtained by training it in the following manner: Obtain multi-frame sample images in the sample video, sample depth images and sample optical flow images corresponding to the sample images of each frame; Obtain a sample supersampled image corresponding to the sample image of each frame, The method of claim 5 or 6, wherein the supersampling network is trained based on the multi-frame sample images, the first M frame images of the sample images of each frame, sample depth images and sample supersampling images of the sample images of each frame, and sample depth images and sample optical flow images corresponding to the first M frame images of the sample images of each frame, wherein M is an integer greater than or equal to N.
8. After obtaining multi-frame sample images in the sample video, the method includes: further comprising pre-processing the multi-frame sample images; The method of claim 7 , wherein the pre-processing includes at least one of blurring, noise, compression distortion, and ringing distortion.
9. An image processing device including a first acquisition module, a second acquisition module, and a determination module, the first capture module is configured to capture a first image of a video and a first depth image corresponding to the first image; the second acquisition module is configured to acquire second images of the first N frames of the first image, second depth images and optical flow images corresponding to the second images of each frame, where N is an integer greater than 0; The determination module is an image processing device configured to determine a supersampled image corresponding to the first image based on the first image, the first depth image, the second image of the first N frames, N second depth images, and N optical flow images.
10. An electronic device comprising a processor and a memory, The memory stores computer-executable instructions; The electronic device causes the processor to execute the image processing method according to any one of claims 1 to 8 by executing computer executable instructions stored in the memory.
11. A computer-readable storage medium having computer-executable instructions stored therein, the computer-readable storage medium realizing the image processing method of any one of claims 1 to 8 when a processor executes the computer-executable instructions.
Citation Information
Patent Citations
Video high-temporal-spatial-resolution signal processing method combining optical flow method and deep network
CN110634105A
Depth map super-resolution optimization method and device, processing equipment and storage medium
CN113284081A
Video processing method and device, computer equipment and medium
CN114612841A
Video rendering processing method and device, equipment and storage medium
CN115035230A
High-resolution image generation apparatus, high-resolution image generation method, and high-resolution image generation program
US20160086311A1