Image generation method and device, electronic equipment, storage medium and chip

By acquiring the image motion vector and optical flow data for distortion processing, and combining with the neural network to generate new images, the time and calculation cost problems of extrapolated frame insertion on mobile devices are solved, and high-quality image generation is achieved.

CN120378566APending Publication Date: 2025-07-25艾酷软件技术(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510590668.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When the prior art performs extrapolation of frame insertion on mobile devices, the time cost, calculation cost and frame insertion quality cannot be taken into account, resulting in soaring power consumption and heating problems.

Method used

By obtaining the motion vector data and optical flow data between the current image and its depth image and historical image, the motion vector image and optical flow vector image are determined, and the distortion process is performed based on the depth image to generate a new image, avoiding the use of G-Buffer of the frame to be generated, and weighted processing is combined with the neural network.

Benefits of technology

Improves the accuracy and fluency of image generation, reduces display delay and calculation costs, and improves the quality and coherence of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378566A_ABST
    Figure CN120378566A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, electronic equipment, a storage medium and a chip, and belongs to the technical field of image processing. The image generation method comprises the following steps: acquiring a first image, a depth image of the first image, and motion vector data and optical flow data between the first image and at least two frames of historical images; respectively determining a motion vector diagram and an optical flow vector diagram according to the motion vector data and the optical flow data; based on the depth image, performing distortion processing on the first image according to the motion vector diagram and the optical flow vector diagram to obtain a first distortion diagram and a second distortion diagram; and generating a second image according to the first image, the first distorted graph and the second distorted graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and specifically relates to an image generation method and apparatus, an electronic device, a storage medium, and a chip. Background Art

[0002] With the progress of real-time rendering technology, electronic devices can create more realistic interactive environments, such as real-time path tracing effects in games. Games or VR (Virtual Reality) applications often require photo-realistic quality.

[0003] Currently, for mobile devices, with the display technology iterating to the 2K resolution (3200×1440 pixels) and 120Hz adaptive refresh rate standard, the electronic devices of each manufacturer already have the hardware foundation to present delicate picture quality and smooth dynamic pictures. However, compared with the traditional 1080P resolution, the 2K resolution needs to process about 180% more pixel amounts. At this time, computationally intensive technologies such as ray tracing and particle effects will further occupy the resources of the GPU (Graphics Processing Unit) or CPU (Central Processing Unit), resulting in a soaring power consumption and heat generation problem. Therefore, advanced rendering acceleration technology is particularly important. Among them, the frame generation technology is a technology that can be used to increase the frame rate to obtain a smoother and jitter-free rendering effect. It can bring a smooth visual effect on the premise of ensuring high-quality rendering quality. The core of the frame generation technology is time sampling and reconstruction. By obtaining a higher sampling rate in time to accelerate the generation of the target frame, the frame generation technology mainly includes interpolated frames and extrapolated frames.

[0004] Among them, the interpolated frame generates a new frame between two rendered frames. It is the mainstream method for video frame interpolation and desktop game frame interpolation. This method relies on the motion estimation of the previous and next frames and can generate relatively accurate composite frames. However, since the generated frame depends on the previous frame and the next frame as inputs, this frame interpolation method needs to wait until the next frame is completely rendered before starting the frame interpolation work. As Figure 2 shown, for the interpolated frame method, the original frame I1 needs to wait until the target frame I 0.5 is calculated and displayed before it can be displayed, which increases the display delay in the rendering process. The extrapolated frame, on the other hand, generates a new frame based on the rendered frame at the previous moment and does not bring additional delay. As Figure 2 shown, for the extrapolated frame method, the original frame I1 can be immediately displayed after rendering. However, in related technologies, as Figure 2 shown, most of the extrapolated frames use the G-Buffer (Geometry Buffer) of the target frame I 0.5 to guide the target frame I0.5 The generation of which requires waiting for the G-Buffer at the 0.5 moment and modifying the deferred rendering pipeline, bringing additional computational costs and engine integration costs. For the extrapolation frame interpolation method that does not apply G-Buffer guidance, due to the lack of information about future frames, usually poor results will be produced. In this way, when performing frame interpolation, it is impossible to balance the time cost, computational cost, and frame interpolation quality of frame interpolation. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an image generation method, device, electronic device, storage medium, and chip, which can balance the time cost, computational cost, and frame interpolation quality of frame interpolation.

[0006] In a first aspect, the embodiments of the present application provide an image generation method, which includes: obtaining a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; respectively determining a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data; based on the depth image, respectively performing warping processing on the first image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image; generating a second image according to the first image, the first warped image, and the second warped image.

[0007] In a second aspect, the embodiments of the present application provide an image generation device, which includes: an obtaining unit, configured to obtain a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; a processing unit, configured to respectively determine a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data; the processing unit is further configured to, based on the depth image, respectively perform warping processing on the first image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image; the processing unit is further configured to generate a second image according to the first image, the first warped image, and the second warped image.

[0008] In a third aspect, the embodiments of the present application provide an electronic device, which includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the image generation method as described in the first aspect are implemented.

[0009] In a fourth aspect, the embodiments of the present application provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the image generation method as described in the first aspect are implemented.

[0010] In a fifth aspect, the embodiments of the present application provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the image generation method as described in the first aspect.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the image generation method as in the first aspect.

[0012] In the image generation method provided by the embodiment of the present application, a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images are obtained; according to the motion vector data and the optical flow data, a motion vector map and an optical flow vector map are respectively determined; based on the depth image, the first image is respectively distorted according to the motion vector map and the optical flow vector map to obtain a first distorted image and a second distorted image; a second image is generated according to the first image, the first distorted image and the second distorted image. Through the above image generation method, the first image of the current frame and its depth image, the motion vector data and the optical flow data between the first image and at least two historical images are obtained, and then according to the motion vector data and the optical flow data, a motion vector map and an optical flow vector map are respectively determined, which can describe the motion changes between images, accurately locate the motion trajectory and spatial position relationship of objects in the image, improve the accuracy of subsequent image generation, and do not need to use the G-Buffer of the frame to be generated, avoiding additional calculation costs and engine integration costs. On this basis, based on the depth image, the first image is respectively distorted according to the motion vector map and the optical flow vector map to obtain a first distorted image and a second distorted image, and then a second image is generated according to the first image, the first distorted image and the second distorted image, which can fuse the original information and the distorted information adjusted by motion and spatial information, make the generated image transition naturally with the previous and subsequent frames in the time and space dimensions, reduce the visual jump feeling, and improve the smoothness and coherence of the video or image sequence. In this way, on the one hand, generating a new image based on the current frame and historical frames can reduce the display delay in the frame interpolation rendering process. On the other hand, constructing a two-dimensional motion relationship from the current frame to the frame to be generated based on the motion vector data and the optical flow data to generate an image can avoid additional calculation costs and engine integration costs, improve the accuracy of image generation, and moreover, fusing the original information and the distorted information adjusted by motion and spatial information to generate an image can improve the smoothness and coherence of the generated image, improve the quality of the generated image, and take into account the time cost, calculation cost and interpolation quality of frame interpolation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is one of the flow schematic diagrams of the image generation method provided by the embodiment of the present application;

[0014] Figure 2 is a principle comparison diagram of interpolation frames and extrapolation frames in related technologies;

[0015] Figure 3The second flowchart of the image generation method provided by the embodiments of the present application;

[0016] Figure 4 The first schematic diagram of the principle of the image generation method provided by the embodiments of the present application;

[0017] Figure 5 The second schematic diagram of the principle of the image generation method provided by the embodiments of the present application;

[0018] Figure 6 The first principle framework diagram of the image generation method provided by the embodiments of the present application;

[0019] Figure 7 The second principle framework diagram of the image generation method provided by the embodiments of the present application;

[0020] Figure 8 The third principle framework diagram of the image generation method provided by the embodiments of the present application;

[0021] Figure 9 The structural block diagram of the image generation device provided by the embodiments of the present application;

[0022] Figure 10 The structural block diagram of the electronic device provided by the embodiments of the present application;

[0023] Figure 11 The schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.

[0025] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0026] Next, the image generation method provided by the embodiments of the present application will be described in detail in conjunction with the accompanying drawings, through specific embodiments and their application scenarios.

[0027] As Figure 1 shown, an embodiment of the present application provides an image generation method, which may include the following S102 to S108:

[0028] S102: Obtain a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two frames of historical images.

[0029] The image generation method proposed in the embodiment of the present application is executed by an electronic device, which may specifically be a smart electronic device such as a smart phone, a tablet computer, a notebook computer, and a smart watch, etc., and no specific limitation is made here.

[0030] Wherein, the current moment is defined as moment 0, the first image I0 is an image calculated by the rendering engine through the rendering pipeline at moment 0, and the historical image is an image calculated by the rendering engine through the rendering pipeline at a historical moment.

[0031] Further, at least two frames of historical images include the first two frames of the first image I0, namely the first historical image I -1 and the first historical image I -2 .

[0032] Further, the above-mentioned motion vector data includes the first rendering motion vector M -1 between the first image I0 and the first historical image I 0→-1 , and the second rendering motion vector M -1 between the first historical image I -2 and the second historical image I -1→-2 .

[0033] Further, the above-mentioned optical flow data includes the first optical flow F -1 between the first image I0 and the first historical image I 0→-1 , and the second optical flow F -1 between the first historical image I -2 and the second historical image I -2→-1 .

[0034] Wherein, the depth image D0 and the first rendering motion vector M 0→-1 are G-Buffer data rendered at moment 0, the second rendering motion vector M -1→-2 and the second optical flow F -2→-1 are data determined according to the G-Buffer and calculation cache of the first historical image I -1 and the second historical image I -2 , and the first optical flow F 0→-1 is estimated according to the first image I0 and the first historical image I -1 .

[0035] Specifically, when the operating load of the image generation method provided in the embodiments of the present application is the GPU of the SoC (System on a Chip), a traditional algorithm based on energy or block matching is used to estimate the first optical flow F 0→-1 , while when the operating load of the image generation method provided in the embodiments of the present application is the NPU of the SoC, an algorithm based on a neural network can be used to estimate the first optical flow F 0→-1 .

[0036] Specifically, in the image generation method provided in the embodiments of the present application, the first image I0, the first historical image I -1 , the first historical image I -2 , the depth image D0, the first rendering motion vector M 0→-1 , the second rendering motion vector M -1→-2 , the second optical flow F -2→-1 are obtained from the rendering engine, and the first optical flow F -1 is estimated based on the first image I0 and the first historical image I 0→-1 . In this way, the spatial and temporal information of the image can be comprehensively captured. The depth image can reflect the three-dimensional structure of the scene, and the motion vector data and optical flow data can describe the motion changes between images. Through multi-dimensional data fusion, a rich and accurate information basis can be provided for extrapolation and interpolation, and the motion trajectory and spatial position relationship of objects in the image can be accurately located, thereby improving the subsequent interpolation accuracy.

[0037] S104: Determine a motion vector map and an optical flow vector map respectively according to the motion vector data and the optical flow data.

[0038] Among them, the motion vector map is used to describe the rendering motion changes between images, and the optical flow vector map is used to describe the optical flow motion changes between images.

[0039] Specifically, in the image generation method provided in the embodiments of the present application, after obtaining the motion vector data and the optical flow data between the first image I0 and at least two frames of historical images, according to the motion vector data, namely the first rendering motion vector M 0→-1 and the second rendering motion vector M -1→-2 , the motion vector map M 0→t from time 0 to the target time t is determined, and according to the optical flow data, namely the first optical flow F 0→-1 and the second optical flow F -2→-1 , the optical flow vector map F 0→tIn this way, the motion changes between images can be described, the motion trajectories and spatial position relationships of objects in the images can be accurately located, the accuracy of subsequent image generation can be improved, and the G-Buffer of the frame to be generated does not need to be used, avoiding additional computational costs and engine integration costs.

[0040] S106: Based on the depth image, respectively according to the motion vector map and the optical flow vector map, perform warping processing on the first image to obtain a first warped image and a second warped image.

[0041] Among them, image warping processing is an operation of deforming and transforming an image. By changing the positions and arrangements of pixel points in the image, the image presents a warped and deformed effect.

[0042] Specifically, in the image generation method provided in the embodiments of the present application, after obtaining the motion vector map M 0→t and the optical flow vector map F 0→t , under the guidance of the depth image D0, according to the motion vector map M 0→t perform forward warping processing on the first image I0 to obtain a first warped image and under the guidance of the depth image D0, according to the optical flow vector map F 0→t perform forward warping processing on the first image I0 to obtain a second warped image In this way, the corresponding vector maps are determined based on the motion vector data and the optical flow data, and then the corresponding warped images are obtained through warping processing based on the depth image. The vector maps can accurately depict the motion states of objects in the images, making the warping processing of the images more in line with the actual motion situation, and being able to more reasonably predict the content of the intermediate frame images during frame interpolation, reducing frame interpolation errors caused by inaccurate motion estimation.

[0043] Among them, in the actual application process, the forward warping processing of the first image I0 can be specifically implemented through atomic operations in the GPU of the SoC.

[0044] S108: Generate a second image according to the first image, the first warped image, and the second warped image.

[0045] Among them, the second image is the image to be rendered at the target time t.

[0046] Specifically, in the image generation method provided in the embodiments of the present application, after obtaining the first warped image and the second warped image , then generate a second image according to the first image I0, the first warped image and the second warped image to obtain a second image In this way, by using the warped image to transform the original image, different postures and positions of an object during movement can be simulated. When the generated image is used as the interpolation result, the generated image has a more natural visual transition with the previous and subsequent frames, can effectively reduce the jitter and incoherence of the moving picture, and improve the smoothness of video playback.

[0047] The above image generation method provided by the embodiments of the present application includes obtaining a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; respectively determining a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data; based on the depth image, respectively performing warping processing on the first image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image; generating a second image according to the first image, the first warped image and the second warped image. Through the above image generation method, the first image of the current frame, its depth image, the motion vector data and the optical flow data between the first image and at least two historical images are obtained, and then the motion vector map and the optical flow vector map are respectively determined according to the motion vector data and the optical flow data, which can describe the motion changes between images, accurately locate the motion trajectory and spatial position relationship of the object in the image, improve the accuracy of subsequent image generation, and do not need to use the G-Buffer of the frame to be generated, avoiding additional calculation costs and engine integration costs. On this basis, based on the depth image, the first image is respectively warped according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image, and then a second image is generated according to the first image, the first warped image and the second warped image, which can fuse the original information and the warped information adjusted by motion and spatial information, make the generated image transition naturally with the previous and subsequent frames in the time and space dimensions, reduce the visual jump feeling, and improve the smoothness and coherence of the video or image sequence. In this way, on the one hand, generating a new image based on the images of the current frame and the historical frames can reduce the display delay in the interpolation rendering process. On the other hand, constructing a two-dimensional motion relationship from the current frame to the frame to be generated based on the motion vector data and the optical flow data to generate an image can avoid additional calculation costs and engine integration costs, improve the accuracy of image generation, and moreover, fusing the original information and the warped information adjusted by motion and spatial information to generate an image can improve the smoothness and coherence of the generated image, improve the quality of the generated image, and take into account the time cost, calculation cost and interpolation quality of interpolation.

[0048] In the embodiments of the present application, the above S104 may specifically include the following S104a and S104b:

[0049] S104a: Based on the principle of uniformly accelerated motion, respectively determine first vector data and second vector data according to the motion vector data and the optical flow data.

[0050] Among them, the above-mentioned constant acceleration motion may specifically include uniform motion and uniformly accelerated motion.

[0051] Specifically, in the image generation method provided in the embodiments of the present application, assuming that the rendering motion and the optical flow motion between images are constant acceleration motions, based on the principle of constant acceleration motion, a two-dimensional motion relationship from the rendering frame at time 0 to the rendering frame at the target time t is constructed. Furthermore, according to the motion vector data, the first vector data is determined, and according to the optical flow data, the second vector data is determined. In this way, the vector data is determined based on the principle of constant acceleration motion, making the simulation of the object motion in the image more in line with the real physical laws, being able to more accurately restore the object motion trajectory and state, and improving the representation accuracy of the motion vector map and the optical flow vector map for the real motion.

[0052] In the actual application process, specifically, the first vector data and the second vector data can be calculated according to the following formula (1):

[0053]

[0054] where f 0→t represents the first vector data or the second vector data, t represents the time interval between the rendering times of the second image and the first image, v0 represents the initial velocity of the rendering motion and the optical flow motion between images in the physical assumption of constant acceleration motion, and α τ represents the acceleration of the rendering motion and the optical flow motion between images in the physical assumption of constant acceleration motion, and both τ and k represent time.

[0055] S104b: According to the first vector data and the second vector data, perform coordinate mapping on the historical images respectively to obtain a motion vector map and an optical flow vector map.

[0056] Specifically, in the image generation method provided in the embodiments of the present application, after obtaining the first vector data and the second vector data, then according to the first vector data, perform coordinate mapping on the historical images to obtain the motion vector map M 0→t , and according to the second vector data, perform coordinate mapping on the historical images to obtain the optical flow vector map F 0→t , so as to map the motion data at time 0 to the target time t. In this way, using the determined vector data to perform coordinate mapping on the historical images to generate the vector map can reasonably adjust the positions of the pixel points in the historical images, making the transition of the images in the time dimension more natural and coherent, in line with the position changes caused by the actual motion of the object, and avoiding image distortion or unnatural motion phenomena caused by coordinate mapping deviation.

[0057] Among them, the above-mentioned process of coordinate mapping involves the problem of many-to-one mapping, which can be specifically implemented through atomic operations in the GPU of the SoC.

[0058] That is, in the image generation method provided in the embodiments of the present application, as Figure 4 shown, the generation process of the motion vector map M 0→t and the optical flow vector map F 0→t both include a vector estimation stage based on a motion hypothesis and a starting coordinate warping stage. According to the optical flow F -1→0 and F -2→-1 generate F 0→t , according to the rendered motion vector M 0→-1 and M -1→-2 generate M 0→t , where F -1→0 is the optical flow data with the same value as F 0→-1 but in the opposite direction. In the vector estimation stage based on the motion hypothesis, a constant acceleration motion hypothesis is made for the two-dimensional plane. Through the above formula (1), the motion data from the initial velocity and acceleration at time 0 to the target time t is obtained by integration, that is, the first vector data and the second vector data. In the starting coordinate warping stage, as Figure 5 shown, the starting coordinates of the motion data are mapped, and the motion data at time 0 is mapped to the target time t.

[0059] Among them, in the vector estimation stage based on the motion hypothesis, the direction of the optical flow and the direction of the rendered motion vector can be opposite or the same, and no specific limitation is made here.

[0060] It can be understood that common extrapolation interpolation techniques use neural network-based generation methods or projection methods based on the camera VP (View-Project) matrix. The images obtained in this way have problems such as poor domain adaptation ability or foreground projection errors. Based on the physical hypothesis of constant acceleration motion, the present application constructs a two-dimensional motion relationship from the current frame to the target frame, and then performs coordinate mapping. The coordinate mapping process does not require the G-Buffer of the target frame, and can completely remove the computational cost and engine integration cost of the target frame deferred rendering.

[0061] In the above embodiments provided by the present application, based on the constant acceleration motion principle, according to the motion vector data and the optical flow data, the first vector data and the second vector data are respectively determined; according to the first vector data and the second vector data, the historical images are respectively subjected to coordinate mapping to obtain the motion vector map and the optical flow vector map. In this way, the motion trajectory and state of the object can be restored more accurately, the representation accuracy of the motion vector map and the optical flow vector map for real motion can be improved, and it is ensured that the motion vector map and the optical flow vector map can accurately reflect the object motion in various motion situations, providing reliable support for subsequent image warping and image generation.

[0062] In the embodiments of the present application, the above S104a may specifically include the following S110 and S112:

[0063] S110: When the frame rate of the first image is greater than a preset threshold, based on the principle of uniform motion, determine the first vector data and the second vector data.

[0064] For the specific value of the preset threshold, those skilled in the art can set it according to the actual situation. For example, the preset threshold is set to 30fps, and no specific limitation is made here.

[0065] Specifically, in the image generation method provided in the embodiments of the present application, when the frame rate of the first image is greater than the preset threshold, it can be specifically assumed that both the rendering motion and the optical flow motion between the images are uniform motions. Based on the principle of uniform motion, a two-dimensional motion relationship from the rendering frame at time 0 to the rendering frame at the target time t is constructed. Then, according to the motion vector data, the first vector data is determined, and according to the optical flow data, the second vector data is determined. In this way, when the frame rate is high, a uniform motion model is adopted, which simplifies the calculation and conforms to the actual situation that the object motion is approximately uniform at high frame rates.

[0066] In the actual application process, when the frame rate of the first image is greater than the preset threshold, the first vector data and the second vector data can be specifically calculated according to the following formula (2):

[0067] f 0→t = v0t = f -1→0 t, (2)

[0068] Where, f 0→t represents the first vector data or the second vector data, v0 represents the initial velocity of the rendering motion and the optical flow motion between the images under the physical assumption of uniform motion, t represents the time interval between the rendering times of the second image and the first image, f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image.

[0069] S112: When the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determine the first vector data and the second vector data.

[0070] Specifically, in the image generation method provided in the embodiments of the present application, when the frame rate of the first image is less than or equal to a preset threshold, it can be specifically assumed that both the rendering motion and the optical flow motion between the images are uniformly accelerated motions. Based on the principle of uniformly accelerated motion, a two-dimensional motion relationship from the rendering frame at time 0 to the rendering frame at the target time t is constructed. Then, according to the motion vector data, the first vector data is determined, and according to the optical flow data, the second vector data is determined. In this way, when the frame rate is low, using the uniformly accelerated motion model can more accurately depict complex motion states such as acceleration and deceleration that may exist at low frame rates, making the motion modeling more in line with the actual scenario and improving the accuracy of vector data determination.

[0071] In the actual application process, when the frame rate of the first image is less than or equal to a preset threshold, the first vector data and the second vector data can be specifically calculated according to the following formula (3):

[0072]

[0073] where f 0→t represents the first vector data or the second vector data, v0 represents the initial velocity of the rendering motion and the optical flow motion between the images in the physical assumption of uniformly accelerated motion, a0 represents the acceleration of the rendering motion and the optical flow motion between the images in the physical assumption of uniformly accelerated motion, t represents the time interval between the rendering times of the second image and the first image, and f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image.

[0074] In the above embodiments provided by the present application, when the frame rate of the first image is greater than the preset threshold, based on the principle of uniform motion, the first vector data and the second vector data are determined according to the following formula; f 0→t = f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, the first vector data and the second vector data are determined according to the following formula; where f 0→t represents the first vector data or the second vector data, t represents the time interval between the rendering times of the second image and the first image, and f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1Represents the motion vector data or optical flow data between the second frame of historical image and the first frame of historical image. In this way, it ensures that high-quality motion vector maps and optical flow vector maps can be generated under various frame rate conditions, providing a reliable basis for subsequent image warping and image generation, enhancing the adaptability of image generation, and implementing a balance between computational accuracy and resource utilization according to the strategy of dynamically adjusting the calculation method according to the frame rate, improving the practicality of image generation in different scenarios.

[0075] In the embodiments of the present application, the above S108 may specifically include the following S108a and S108b:

[0076] S108a: Using a neural network, generate a third image and a fourth image according to the first image, the first warped image, and the second warped image.

[0077] Among them, the third image is a weight map. A weight map is an image with weight information, and the weight value represents the importance degree of a specific attribute or feature, which is represented by a gray value or color coding. For example, in image fusion, the weight map determines the contribution ratio of each image.

[0078] That is, the third image is used to indicate the proportion of the first warped image and the second warped image in the second image.

[0079] Furthermore, the fourth image is a repair map. Some hole regions without pixel information or missing pixel information will be generated during the forward warping of the image. The repair map can reason based on the context information of the image, so as to fill the hole regions in the image.

[0080] That is, the fourth image is used to supplement the missing pixel information in the second image.

[0081] Specifically, in the image generation method provided in the embodiments of the present application, after obtaining the first warped image and the second warped image, input the first image, the first warped image, and the second warped image into the neural network, so as to use the neural network to generate and output a third image and a fourth image according to the input data.

[0082] Among them, the above neural network may adopt a symmetric encoding and decoding structure similar to U-Net (a variant of a deep learning model based on the U-Net architecture, and U-Net adopts an encoder-decoder structure), which is not specifically limited here.

[0083] S108b: Perform weighted processing on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image.

[0084] Specifically, in the image generation method provided in the embodiments of the present application, after the neural network outputs the third image and the fourth image, the first warped image and the second warped image are weighted according to the third image and the fourth image to obtain the second image. In this way, by using the output results of the neural network to weight the first warped image and the second warped image, the neural network can adaptively select the rendering color at the target time t. Moreover, the neural network can also reason and fill the hole regions generated during the forward warping process according to the context, balance the contributions of different warped images to the final image, make the generated image more in line with expectations in terms of texture, structure, etc., and improve the overall visual quality.

[0085] It can be understood that the extrapolation interpolation method usually performs poorly in the shadow and particle special effect regions because it does not consider the light and shadow problems or uses the neural network to implicitly optimize the problem regions. The image generation method provided in the embodiments of the present application explicitly generates the light and shadow problem regions, based on the physical assumption of constant acceleration motion, uses the optical flow data and the motion vector data to generate a new image, and adaptively weights the two warped images by using the neural network to obtain a new image. This not only enables the neural network to learn a smooth and natural weight map, but also solves the problem of poor generalization ability of the generative network through an explicit construction method.

[0086] In the above embodiments provided by the present application, the neural network is used to generate the third image and the fourth image according to the first image, the first warped image, and the second warped image. The third image is used to indicate the proportion of the first warped image and the second warped image in the second image, and the fourth image is used to supplement the missing pixel information in the second image; the first warped image and the second warped image are weighted according to the third image and the fourth image to obtain the second image. In this way, the neural network is used to mix and repair the warped images to obtain a new image, balance the contributions of different warped images to the final image, make the generated image more in line with expectations in terms of texture, structure, etc., and improve the quality of the generated image.

[0087] In the embodiments of the present application, the above S108b may specifically include the following S114:

[0088] S114: According to the formula The first warped image and the second warped image are weighted to obtain the second image.

[0089] Specifically, in the image generation method provided in the embodiments of the present application, the warped images may be mixed and repaired according to the following formula (4) to obtain the second image:

[0090]

[0091] where represents the second image, represents the third image, Represents the fourth image, Represents the first warped image, Represents the second warped image.

[0092] In the above embodiments provided by the present application, according to the following formula, the first warped image and the second warped image are weighted to obtain the second image; Wherein, Represents the second image, Represents the third image, Represents the fourth image, Represents the first warped image, Represents the second warped image. In this way, the fusion degree of different image components can be accurately controlled, so that the generated image can better take into account the advantageous features of each warped image, avoid the over - presentation or under - presentation of a certain image feature, realize the fine fusion of image elements, make the generated image more in line with expectations in terms of texture, structure, etc., and improve the quality of the generated image.

[0093] In summary, as Figure 3 shown, the image generation method provided by the embodiments of the present application may specifically include the following S202 to S210:

[0094] S202: Obtain the first image and its depth image, the first historical image, the first rendering motion vector, the second rendering motion vector, and the second optical flow.

[0095] S204: Estimate the first optical flow according to the first image and the first historical image.

[0096] S206: Determine the coordinate mapping relationship from the first image to the second image: According to the motion vector data and the optical flow data, determine the motion vector map and the optical flow vector map respectively.

[0097] S208: Based on the depth image, warp the first image according to the motion vector map and the optical flow vector map respectively to obtain the first warped image and the second warped image.

[0098] S210: Use a neural network to mix and repair the first warped image and the second warped image to obtain the second image.

[0099] That is, the embodiments of the present application provide an AI (Artificial Intelligence) extrapolation and interpolation method that can be applied to mobile devices and takes into account both performance and effects. During the rendering interpolation process, according to appropriate physical motion assumptions, the motion of the current frame and the historical frame is mapped to a future moment. At the same time, based on the implicit learning characteristics of the neural network, the correct choice is adaptively made between geometric motion and light and shadow motion for different coordinates. In this way, users can obtain high - frame - rate and high - quality rendering effects on mobile devices.

[0100] It can be understood that, from a hardware perspective, currently electronic devices usually implement advanced CV (Computer Vision) and CG (Computer Graphics) tasks, such as large model local inference and millisecond-level image processing tasks, through heterogeneous computing architecture integration solutions and external AI acceleration chip solutions. The image generation method provided in the embodiments of the present application can also implement low-latency rendering frame interpolation functions on mobile devices through the cooperation of multiple components of the SoC or external independent chips.

[0101] Specifically, the operating carrier of the image generation method provided in the embodiments of the present application can be an SoC. Among them, as Figure 6 shown, optical flow calculation can be performed in the NPU (Neural Network Processing Unit) of the SoC, or, as Figure 7 shown, optical flow calculation can also be implemented in the GPU of the SoC. In this way, without increasing the rendering cost and pipeline integration cost, the problem of extrapolation frame interpolation on mobile devices can be solved with high quality through the application software solution on the general SoC.

[0102] Furthermore, as Figure 8 shown, the operating carrier of the image generation method provided in the embodiments of the present application can also be an external independent chip. At this time, compared with the operating carrier SoC, although the cost is higher, the latency and power consumption are lower.

[0103] In the actual application process, the image generation method provided in the embodiments of the present application can be specifically applied to scenarios such as game rendering, mobile animation rendering, and visualization of mobile rendering engine development, which are not specifically limited herein.

[0104] For the image generation method provided in the embodiments of the present application, the execution subject can be an image generation device. In the embodiments of the present application, taking the image generation device executing the above image generation method as an example, the image generation device provided in the embodiments of the present application is described.

[0105] As Figure 9 shown, the embodiments of the present application provide an image generation device 500, which may include the following acquisition unit 502 and processing unit 504.

[0106] The acquisition unit 502 is configured to acquire a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images;

[0107] The processing unit 504 is configured to determine a motion vector map and an optical flow vector map respectively according to the motion vector data and the optical flow data;

[0108] The processing unit 504 is further configured to perform warping processing on the first image based on the depth image, respectively according to the motion vector map and the optical flow vector map, to obtain a first warped image and a second warped image;

[0109] The processing unit 504 is further configured to generate a second image according to the first image, the first warped image, and the second warped image.

[0110] The image generation device 500 provided by the embodiment of the present application obtains a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; determines a motion vector map and an optical flow vector map respectively according to the motion vector data and the optical flow data; performs warping processing on the first image based on the depth image, respectively according to the motion vector map and the optical flow vector map, to obtain a first warped image and a second warped image; generates a second image according to the first image, the first warped image, and the second warped image. Through the above image generation device 500, the first image of the current frame, its depth image, the motion vector data and the optical flow data between the first image and at least two historical images are obtained, and then the motion vector map and the optical flow vector map are respectively determined according to the motion vector data and the optical flow data, which can describe the motion changes between images, accurately locate the motion trajectory and spatial position relationship of the objects in the image, improve the accuracy of subsequent image generation, and do not need to use the G-Buffer of the frame to be generated, avoiding additional calculation costs and engine integration costs. On this basis, further perform warping processing on the first image based on the depth image, respectively according to the motion vector map and the optical flow vector map, to obtain a first warped image and a second warped image, and then generate a second image according to the first image, the first warped image, and the second warped image, which can fuse the original information and the warped information adjusted by the motion and spatial information, make the generated image transition naturally with the previous and subsequent frames in the time and space dimensions, reduce the visual jump feeling, and improve the smoothness and coherence of the video or image sequence. In this way, on the one hand, generating a new image based on the images of the current frame and the historical frames can reduce the display delay in the frame interpolation rendering process. On the other hand, constructing a two-dimensional motion relationship from the current frame to the frame to be generated based on the motion vector data and the optical flow data to generate an image can avoid additional calculation costs and engine integration costs, improve the accuracy of image generation, and moreover, fusing the original information and the warped information adjusted by the motion and spatial information to generate an image can improve the smoothness and coherence of the generated image, improve the quality of the generated image, and take into account the time cost, calculation cost and interpolation quality of frame interpolation.

[0111] In the embodiment of the present application, the processing unit 504 is specifically configured to: respectively determine first vector data and second vector data according to the motion vector data and the optical flow data based on the principle of constant acceleration motion; perform coordinate mapping on the historical images respectively according to the first vector data and the second vector data to obtain a motion vector map and an optical flow vector map.

[0112] Based on the principle of constant acceleration motion, the above embodiments provided by the present application respectively determine the first vector data and the second vector data according to the motion vector data and the optical flow data; according to the first vector data and the second vector data, the historical images are respectively subjected to coordinate mapping to obtain a motion vector map and an optical flow vector map. In this way, the motion trajectory and state of the object can be restored more accurately, the representation accuracy of the motion vector map and the optical flow vector map for the real motion can be improved, and it is ensured that the motion vector map and the optical flow vector map can accurately reflect the object motion in various motion situations, providing reliable support for subsequent image distortion and image generation.

[0113] In the embodiment of the present application, the processing unit 504 is specifically configured to: when the frame rate of the first image is greater than a preset threshold, based on the principle of uniform motion, determine the first vector data and the second vector data according to the following formula; f 0→t =f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determine the first vector data and the second vector data according to the following formula; where, f 0→t represents the first vector data or the second vector data, t represents the interval duration between the rendering times of the second image and the first image, f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image.

[0114] Based on the principle of uniform motion, the above embodiments provided by the present application determine the first vector data and the second vector data according to the following formula; f 0→t =f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determine the first vector data and the second vector data according to the following formula; where, f 0→t represents the first vector data or the second vector data, t represents the interval duration between the rendering times of the second image and the first image, f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image. In this way, it is ensured that high-quality motion vector maps and optical flow vector maps can be generated under various frame rate conditions, providing a reliable basis for subsequent image distortion and image generation, enhancing the adaptability of image generation, and the strategy of dynamically adjusting the calculation method according to the frame rate realizes the balance between calculation accuracy and resource utilization, improving the practicality of image generation in different scenarios.

[0115] In an embodiment of the present application, the processing unit 504 is specifically configured to: use a neural network to generate a third image and a fourth image according to a first image, a first warped image, and a second warped image, where the third image is used to indicate the proportion of the first warped image and the second warped image in a second image, and the fourth image is used to supplement missing pixel information in the second image; and perform a weighted processing on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image.

[0116] In the above embodiment provided by the present application, a neural network is used to generate a third image and a fourth image according to a first image, a first warped image, and a second warped image, where the third image is used to indicate the proportion of the first warped image and the second warped image in a second image, and the fourth image is used to supplement missing pixel information in the second image; and perform a weighted processing on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image. In this way, the neural network is used to mix and repair the warped images to obtain a new image, balance the contributions of different warped images to the final image, make the generated image more in line with expectations in terms of texture, structure, etc., and improve the quality of the generated image.

[0117] In an embodiment of the present application, the processing unit 504 is specifically configured to: perform a weighted processing on the first warped image and the second warped image according to the following formula to obtain the second image; where, represents the second image, represents the third image, represents the fourth image, represents the first warped image, represents the second warped image.

[0118] In the above embodiment provided by the present application, a weighted processing is performed on the first warped image and the second warped image according to the following formula to obtain the second image; where, represents the second image, represents the third image, represents the fourth image, represents the first warped image, represents the second warped image. In this way, the fusion degree of different image components can be accurately controlled, so that the generated image can better take into account the advantageous features of each warped image, avoid excessive or insufficient presentation of a certain image feature, realize the fine fusion of image elements, make the generated image more in line with expectations in terms of texture, structure, etc., and improve the quality of the generated image.

[0119] The image generation device 500 in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0120] The image generation device 500 in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0121] The image generation device 500 provided in the embodiments of the present application can implement Figure 1 and Figure 3 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0122] Optionally, as Figure 10 shown, the embodiments of the present application further provide an electronic device 600, including a processor 602 and a memory 604. A program or instruction that can run on the processor 602 is stored on the memory 604. When the program or instruction is executed by the processor 602, it implements each step of the above image generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0123] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0124] Figure 11 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0125] The electronic device 700 includes, but is not limited to, components such as a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710.

[0126] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 710 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than those shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0127] Among them, the processor 710 is used to obtain a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two frames of historical images.

[0128] The processor 710 is further used to respectively determine a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data.

[0129] The processor 710 is further used to respectively perform warping processing on the first image based on the depth image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image.

[0130] The processor 710 is further used to generate a second image according to the first image, the first warped image, and the second warped image.

[0131] In an embodiment of the present application, a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images are obtained; according to the motion vector data and the optical flow data, a motion vector map and an optical flow vector map are respectively determined; based on the depth image, the first image is respectively distorted according to the motion vector map and the optical flow vector map to obtain a first distorted image and a second distorted image; a second image is generated according to the first image, the first distorted image and the second distorted image. In an embodiment of the present application, the first image of the current frame and its depth image, the motion vector data and the optical flow data between the first image and at least two historical images are obtained, and then according to the motion vector data and the optical flow data, a motion vector map and an optical flow vector map are respectively determined, which can describe the motion changes between images, accurately locate the motion trajectory and spatial position relationship of objects in the images, improve the accuracy of subsequent image generation, and do not need to use the G-Buffer of the frame to be generated, avoiding additional calculation costs and engine integration costs. On this basis, based on the depth image, the first image is respectively distorted according to the motion vector map and the optical flow vector map to obtain a first distorted image and a second distorted image, and then a second image is generated according to the first image, the first distorted image and the second distorted image, which can fuse the original information and the distorted information adjusted by the motion and spatial information, make the generated image transition naturally with the previous and subsequent frames in the time and space dimensions, reduce the visual jump feeling, and improve the smoothness and coherence of the video or image sequence. In this way, on the one hand, generating a new image based on the current frame and the historical frames can reduce the display delay in the frame interpolation rendering process. On the other hand, constructing a two-dimensional motion relationship from the current frame to the frame to be generated based on the motion vector data and the optical flow data to generate an image can avoid additional calculation costs and engine integration costs, improve the accuracy of image generation, and moreover, fusing the original information and the distorted information adjusted by the motion and spatial information to generate an image can improve the smoothness and coherence of the generated image, improve the quality of the generated image, and take into account the time cost, calculation cost and interpolation quality of the frame interpolation.

[0132] Optionally, the processor 710 is specifically configured to: based on the principle of constant acceleration motion, respectively determine first vector data and second vector data according to the motion vector data and the optical flow data; according to the first vector data and the second vector data, respectively perform coordinate mapping on the historical images to obtain a motion vector map and an optical flow vector map.

[0133] Based on the principle of constant acceleration motion, the above embodiments provided by the present application determine the first vector data and the second vector data respectively according to the motion vector data and the optical flow data; and respectively perform coordinate mapping on the historical images according to the first vector data and the second vector data to obtain the motion vector map and the optical flow vector map. In this way, the motion trajectory and state of the object can be restored more accurately, the representation accuracy of the motion vector map and the optical flow vector map for the real motion can be improved, and it is ensured that the motion vector map and the optical flow vector map can accurately reflect the object motion in various motion situations, providing reliable support for subsequent image distortion and image generation.

[0134] Optionally, the processor 710 is specifically configured to: when the frame rate of the first image is greater than a preset threshold, based on the principle of uniform motion, determine the first vector data and the second vector data according to the following formula; f 0→t = f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determine the first vector data and the second vector data according to the following formula; where, f 0→t represents the first vector data or the second vector data, t represents the time interval between the rendering times of the second image and the first image, f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image.

[0135] Based on the principle of uniform motion, the above embodiments provided by the present application determine the first vector data and the second vector data according to the following formula; f 0→t = f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determine the first vector data and the second vector data according to the following formula; where, f 0→t represents the first vector data or the second vector data, t represents the time interval between the rendering times of the second image and the first image, f -1→0 represents the motion vector data or the optical flow data between the first frame of historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second frame of historical image and the first frame of historical image. In this way, it is ensured that high-quality motion vector maps and optical flow vector maps can be generated under various frame rate conditions, providing a reliable basis for subsequent image distortion and image generation, enhancing the adaptability of image generation, and the strategy of dynamically adjusting the calculation method according to the frame rate realizes the balance between calculation accuracy and resource utilization, improving the practicality of image generation in different scenarios.

[0136] Optionally, the processor 710 is specifically configured to: use a neural network to generate a third image and a fourth image according to the first image, the first warped image, and the second warped image, where the third image is used to indicate the proportion of the first warped image and the second warped image in the second image, and the fourth image is used to supplement the missing pixel information in the second image; and perform a weighted process on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image.

[0137] In the above embodiments provided by the present application, a neural network is used to generate a third image and a fourth image according to the first image, the first warped image, and the second warped image, where the third image is used to indicate the proportion of the first warped image and the second warped image in the second image, and the fourth image is used to supplement the missing pixel information in the second image; and a weighted process is performed on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image. In this way, the warped images are mixed and repaired using a neural network to obtain a new image, balancing the contributions of different warped images to the final image, making the generated image more in line with expectations in terms of texture, structure, etc., and improving the quality of the generated image.

[0138] Optionally, the processor 710 is specifically configured to: perform a weighted process on the first warped image and the second warped image according to the following formula to obtain the second image; where represents the second image, represents the third image, represents the fourth image, represents the first warped image, represents the second warped image.

[0139] In the above embodiments provided by the present application, a weighted process is performed on the first warped image and the second warped image according to the following formula to obtain the second image; where represents the second image, represents the third image, represents the fourth image, represents the first warped image, represents the second warped image. In this way, the fusion degree of different image components can be accurately controlled, enabling the generated image to better take into account the advantageous features of each warped image, avoiding the over - presentation or under - presentation of a certain image feature, achieving the fine fusion of image elements, making the generated image more in line with expectations in terms of texture, structure, etc., and improving the quality of the generated image.

[0140] It should be understood that in the embodiments of the present application, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also referred to as a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0141] The memory 709 can be used to store software programs and various data. The memory 709 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 can include a volatile memory or a non-volatile memory, or the memory 709 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0142] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 710 either.

[0143] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the image generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0144] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.

[0145] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement each process of the above-mentioned embodiment of the image generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0146] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0147] The embodiments of the present application provide a computer program product, which is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the image generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0148] It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus including a series of elements not only includes those elements but also includes other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0150] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. An image generation method, characterized in that, including: obtaining a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; respectively determining a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data; based on the depth image, respectively performing warping processing on the first image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image; generating a second image according to the first image, the first warped image and the second warped image.

2. The image generation method according to claim 1, wherein The respectively determining a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data includes: based on the principle of constant acceleration motion, respectively determining first vector data and second vector data according to the motion vector data and the optical flow data; respectively performing coordinate mapping on the historical images according to the first vector data and the second vector data to obtain the motion vector map and the optical flow vector map.

3. The image generation method according to claim 2, characterized in that, The based on the principle of constant acceleration motion, respectively determining first vector data and second vector data according to the motion vector data and the optical flow data includes: when the frame rate of the first image is greater than a preset threshold, based on the principle of uniform motion, determining the first vector data and the second vector data according to the following formula; f 0→t = f -1→0 t; when the frame rate of the first image is less than or equal to the preset threshold, based on the principle of uniformly accelerated motion, determining the first vector data and the second vector data according to the following formula; Among them, f 0→t represents the first vector data or the second vector data, t represents the time interval between the rendering times of the second image and the first image, and f -1→0 represents the motion vector data or the optical flow data between the first-frame historical image and the first image, and f -2→-1 represents the motion vector data or the optical flow data between the second-frame historical image and the first-frame historical image.

4. The image generation method according to claim 1, wherein The generating a second image according to the first image, the first warped image and the second warped image includes: using a neural network to generate a third image and a fourth image according to the first image, the first warped image and the second warped image, where the third image is used to indicate the proportion of the first warped image and the second warped image in the second image, and the fourth image is used to supplement missing pixel information in the second image; performing weighted processing on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image.

5. The image generation method according to claim 4, wherein The performing weighted processing on the first warped image and the second warped image according to the third image and the fourth image to obtain the second image includes: performing weighted processing on the first warped image and the second warped image according to the following formula to obtain the second image; Among them, represents the second image, represents the third image, represents the fourth image, represents the first distorted image, represents the second distorted image.

6. An image generation device, characterized in that, including: an obtaining unit, configured to obtain a first image, a depth image of the first image, motion vector data and optical flow data between the first image and at least two historical images; a processing unit, configured to respectively determine a motion vector map and an optical flow vector map according to the motion vector data and the optical flow data; the processing unit is further configured to, based on the depth image, respectively perform warping processing on the first image according to the motion vector map and the optical flow vector map to obtain a first warped image and a second warped image; the processing unit is further configured to generate a second image according to the first image, the first warped image and the second warped image.

7. The image generation device according to claim 6, wherein The processing unit is specifically configured to: Based on the principle of constant acceleration motion, the first vector data and the second vector data are respectively determined according to the motion vector data and the optical flow data; According to the first vector data and the second vector data, coordinate mapping is respectively performed on the historical image to obtain the motion vector map and the optical flow vector map.

8. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the image generation method described in any one of claims 1 to 5 are implemented.

9. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by the processor, the steps of the image generation method described in any one of claims 1 to 5 are implemented.

10. A chip, characterized in that, It includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the image generation method described in any one of claims 1 to 5.