Image processing method and device, electronic equipment, chip and storage medium
By obtaining the rendering information of two adjacent frames in the video frame sequence, combined with deep learning technology, predicting and generating high-quality next frames, the problems of inaccurate motion vector prediction and high computational complexity in the existing technology are solved, and image quality improvement and equipment battery life are achieved.
Patent Information
- Application Number
- CN202510437215.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
AI Technical Summary
When generating the next frame of the video prediction technology, the data during the rendering process is not fully utilized, resulting in inaccurate prediction of motion vectors, deterioration of picture quality, high computational complexity and long delay, and cannot be effectively applied to mobile terminals.
By obtaining the adjacent two frames of the video frame sequence and their rendering information, including motion vectors, depth information and albedo information, combined with deep learning technology, the motion vectors of the next frame are predicted, and transformed and fusion processing is performed to generate a high-quality next frame.
It improves the picture quality of the next frame of the picture, reduces the computational complexity, reduces the power consumption of the device, extends the battery life of the device, and improves the user experience.
Smart Images

Figure CN120264015A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to an image processing method, apparatus, electronic device, chip, and storage medium. Background Art
[0002] With the rapid development of digital technologies, the generation and consumption of video data have increased exponentially. From short videos on social media platforms to high-definition movies, security surveillance videos, and real-time video streams of autonomous vehicles, video data is everywhere. This vast amount of video data not only poses higher requirements for storage and transmission but also brings new challenges and opportunities to video processing technologies.
[0003] In this context, video prediction, as a cutting-edge video processing technology, has become particularly important. Among them, video prediction aims to analyze the existing video frames in a video frame sequence to infer and generate future video frames. The core of this technology lies in capturing the spatio-temporal patterns and dynamic changes in the video to achieve accurate prediction of future frames. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems in the related art to some extent.
[0005] To this end, this application proposes an image processing method, apparatus, electronic device, chip, and storage medium to implement video prediction by integrating two adjacent frames in a video frame sequence and the rendering information corresponding to the two adjacent frames to generate the next frame after the two adjacent frames. This can not only improve the image quality of the predicted next frame but also reduce the need for complex calculations as video prediction only utilizes the relevant information of two adjacent frames, thereby reducing the power consumption of the device, extending the battery life of the device, and enhancing the user experience.
[0006] An embodiment of one aspect of this application proposes an image processing method, including:
[0007] Obtain two adjacent frames in a video frame sequence and the rendering information corresponding to the frames;
[0008] Generate the next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames.
[0009] An embodiment of another aspect of this application proposes an image processing apparatus, including:
[0010] An obtaining module, configured to obtain two adjacent frames in a video frame sequence and the rendering information corresponding to the frames;
[0011] A generation module, configured to generate a next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames.
[0012] Another embodiment of this application provides a chip, which includes:
[0013] A graphics processing unit (GPU), configured to obtain two adjacent frames in a video frame sequence and the rendering information corresponding to the frames, and send the rendering information corresponding to the frames to a neural network processing unit (NPU);
[0014] The NPU is configured to predict a motion vector of a next frame after the two adjacent frames according to the rendering information corresponding to the frames, and send the motion vector to the GPU;
[0015] The GPU is further configured to generate the next frame according to the two adjacent frames and the motion vector.
[0016] Another embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image processing method as described in the foregoing aspect.
[0017] Another embodiment of this application provides another chip, which includes an interface circuit and a processing circuit that are coupled to each other. The interface circuit is configured to input or output signals, and the processing circuit is configured to execute the image processing method as described in the foregoing aspect.
[0018] Another embodiment of this application provides a non-transitory computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, they implement the image processing method as described in the foregoing aspect.
[0019] Another embodiment of this application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method as described in any of the foregoing aspects.
[0020] The image processing method, device, electronic device, chip, and storage medium provided by this application perform video prediction by integrating two adjacent frames in a video frame sequence and the rendering information corresponding to the two adjacent frames, and generate a next frame after the two adjacent frames. This can not only improve the image quality of the predicted next frame, but also, since video prediction only requires using the relevant information of two adjacent frames, it can reduce the need for complex calculations, thereby reducing the power consumption of the device, extending the battery life of the device, and improving the user experience.
[0021] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0022] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:
[0023] Figure 1 is a schematic flowchart of the first image processing method provided by the embodiment of the present application;
[0024] Figure 2 is a schematic flowchart of the second image processing method provided by the embodiment of the present application;
[0025] Figure 3 is a schematic flowchart of the third image processing method provided by the embodiment of the present application;
[0026] Figure 4 is a schematic flowchart of the fourth image processing method provided by the embodiment of the present application;
[0027] Figure 5 is a schematic flowchart of the fifth image processing method provided by the embodiment of the present application;
[0028] Figure 6 is a schematic structural diagram of the first chip provided by the embodiment of the present application;
[0029] Figure 7 is a schematic diagram of the video prediction principle provided by the embodiment of the present application;
[0030] Figure 8 is a schematic structural diagram of an image processing device provided by the embodiment of the present application;
[0031] Figure 9 is a schematic structural diagram of an electronic device provided by the embodiment of the present application;
[0032] Figure 10 is a schematic structural diagram of the second chip proposed by the embodiment of the present application. Detailed Embodiments
[0033] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0034] Video prediction mainly includes the following two methods: model-based prediction method and learning-based prediction method.
[0035] Model-based prediction method: For some scenarios with clear physical laws, such as the motion scenario of an object, etc. For the motion scenario of a rigid body, factors such as its rotation and translation can also be considered. For example, for a rotating gyroscope, by analyzing physical parameters such as its initial angular velocity, angular acceleration, and the translation velocity of the center of mass, its posture and position in subsequent video frames can be predicted.
[0036] Learning-based prediction method: Starting from the perspective of probability, analyze the statistical laws of pixel changes in the video frame sequence. For example, in a traffic video, count the appearance frequency and driving direction of vehicles in different lanes and at different time periods, and by establishing a probability model, such as a Markov chain, predict the position and state of the vehicle at the next moment. This method can also be used to predict texture changes in the video. For example, in a video of wind blowing grass, count the probability distribution of the direction and amplitude of grass blade swinging, and then predict the shape of grass blades in future video frames.
[0037] In any one of the embodiments of the present application, the application scenarios of video prediction include but are not limited to the following scenarios:
[0038] (1) Video stream transmission scenario: Predicting the picture content of video frames in advance can help the system optimize bandwidth allocation. For example, for parts of the video with high predictability (such as relatively static parts in a landscape video), the transmission bit rate can be appropriately reduced, while more bandwidth is allocated to parts of the video that are difficult to predict (such as suddenly appearing animals or vehicles). In video coding standards, exploration is also being carried out on how to use video prediction to more efficiently perform inter-frame prediction to reduce the bit rate of the video and improve coding efficiency.
[0039] (2) Autonomous driving scenario: The perception system of the vehicle can use video prediction to better understand the surrounding environment. For example, predict the next actions of pedestrians and other vehicles, and assist autonomous driving vehicles in making more reasonable decisions, such as decelerating in advance and avoiding; or, for complex road condition changes, such as suddenly appearing obstacles on the road or lane-changing behaviors of other vehicles, video prediction can help the autonomous driving system plan the driving path of the vehicle in advance to ensure driving safety.
[0040] In the related art, the following several technical solutions are mainly adopted to implement video prediction:
[0041] The first solution: Calculate the motion vectors of the current frame and the previous frame of the video stream, and based on this motion vector, predict the motion vector of the next frame, and apply the motion vector of the next frame to the current frame to obtain the picture content of the next frame (i.e., the predicted frame).
[0042] The second solution: Based on the game state of the current frame and the historical sequence of user inputs, predict the game state of the next frame, and generate one or more predicted game screens accordingly.
[0043] The third solution: Separate the dynamic and static regions in the video, specifically perform frame interpolation on the dynamic region, reduce frame interpolation anomalies, and save frame interpolation power consumption.
[0044] The fourth solution: During the frame interpolation of video frames, generate the motion vector of the next frame by calculating the motion vectors of at least three image frames, and perform frame interpolation based on this to obtain the content of the next frame (i.e., the predicted frame).
[0045] However, the above-mentioned first solution directly predicts the motion vectors of adjacent frames based on the image content, without considering other data in the rendering process (such as depth, motion vector (or called motion vector, Motion Vector, abbreviated as MV), user interface (User Interface, abbreviated as UI) information, etc.), lacking physical information in the rendering process, resulting in inaccurate prediction of motion vectors and deteriorated image quality.
[0046] The above-mentioned second solution requires a large amount of historical information for state estimation and decision-making, which is applicable to personal computers (Personal Computer, abbreviated as PC) and cloud game scenarios. However, caching and decision-making on mobile devices will bring huge computational consumption and latency, resulting in the inability to implement the solution.
[0047] The above-mentioned third solution is similar to the first solution, directly using images to predict the motion vectors of two adjacent frames, without using other data in the rendering process (such as depth, MV, UI information, etc.), lacking physical information in the rendering process, resulting in inaccurate prediction of motion vectors and deteriorated image quality. Moreover, in game scenarios, there are rarely completely static regions (the background has clouds and grass shaking). If pixel-level static region division is performed, it will cause line breakage in the frame interpolation result, reducing the image quality and user experience of the generated image.
[0048] The above-mentioned fourth solution requires at least three frames of image information to predict the fourth frame, needs to cache and process more data, has higher latency, and this solution does not consider factors such as depth that affect motion, limiting the image quality of the generated image.
[0049] Therefore, in view of at least one of the problems existing in the above-mentioned related technologies, the present application proposes an image processing method, apparatus, electronic device, chip, and storage medium.
[0050] The following describes the image processing method, apparatus, electronic device, chip, and storage medium of the embodiments of the present application with reference to the accompanying drawings.
[0051] Figure 1 Schematic flowchart of the first image processing method provided by the embodiments of the present application.
[0052] It should be noted that the image processing method of the embodiments of the present application can be applied to an image processing device. In some possible embodiments, the image processing device can be configured in an electronic device or a chip so that the electronic device or the chip can perform image processing functions. Additionally, in some possible embodiments, the image processing device can also be software in the electronic device, etc.
[0053] In any one of the embodiments of the present application, the chip can be integrated into an electronic device. The chip includes, among others, a Graphics Processing Unit (GPU), a Neural Processing Unit (NPU), a Central Processing Unit (CPU), an Image Signal Processing (ISP), an Application-Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Field-Programmable Gate Array (FPGA), a System On Chip (SOC), a Reduced Instruction Set Computer (RISC), etc., which are not listed one by one here.
[0054] Among them, the electronic device includes but is not limited to: terminals, vehicles, PCs, etc. Among them, a terminal is an entity on the user side for receiving or transmitting signals, such as a mobile phone. A terminal can also be called a terminal device (terminal), user equipment (UE for short), mobile station (MS for short), mobile terminal (MT for short), etc. A terminal can be a car with communication functions, a smart car, a mobile phone, a wearable device, a tablet computer (Pad), a computer with wireless transceiver functions, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, and so on. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal.
[0055] As Figure 1 shown, the image processing method may include the following steps S101 to S102:
[0056] Step S101, obtain two adjacent frames of the video frame sequence and the rendering information corresponding to the frames.
[0057] Among them, each video frame (referred to as a frame in this application) in the video frame sequence has corresponding rendering information, where the rendering information includes relevant parameters used in the rendering process of the corresponding frame, and is used to improve the effect and efficiency of video processing and graphics rendering.
[0058] Exemplarily, the rendering information of each frame includes but is not limited to at least one of the following: the first item, the motion vector (MV) of the frame, denoted as the first motion vector in this application; the second item, the depth information of the frame, denoted as the first depth information in this application; the third item, the albedo information of the frame, denoted as the first albedo information in this application; the fourth item, the bone point information of the character in the frame, where the bone point information can be used to indicate the positions of the respective bone points of the character, and the character includes but is not limited to human characters and animal characters.
[0059] Among them, the first motion vector can be a two-channel tensor, which is used to indicate the motion components of each pixel point in the corresponding picture. Among them, the motion components of each pixel point include the motion information of the pixel point in two channels, that is, the motion information in the X direction and the motion information in the Y direction.
[0060] Among them, in the case where the picture is a non-first-frame picture in the video frame sequence, for example, mark the picture as the nth (n is a positive integer greater than 1) frame picture in the video frame sequence. The first motion vector of the nth frame picture is used to indicate the motion information of the nth frame picture distorted towards the (n - 1)th frame picture. In the case where the picture is the first-frame picture in the video frame sequence, the first motion vector of this picture can be a set vector, such as a 0 vector.
[0061] Among them, the first depth information is used to indicate the depth of each pixel point in the corresponding picture. Exemplarily, the first depth information can be presented in the form of a grayscale image, and each pixel value in the image is used to indicate the depth at the position of the pixel point.
[0062] Among them, the first albedo information is used to indicate the albedo of each pixel point in the corresponding picture. Among them, the albedo is used to indicate the texture and color of the corresponding pixel point. Among them, the color of each pixel point refers to the color of the object to which the pixel point belongs itself, that is, the original color before image post-processing (such as lighting processing, adding shadows and special effects, etc.).
[0063] In the embodiments of the present application, two adjacent pictures in the video frame sequence can be obtained, and the rendering information of the two adjacent pictures can be obtained.
[0064] Step S102, generate the next picture after the two adjacent pictures according to the two adjacent pictures and the rendering information corresponding to the pictures.
[0065] In the embodiments of the present application, the two adjacent pictures and the rendering information corresponding to the two adjacent pictures can be comprehensively used for video prediction to generate the next picture after the two adjacent pictures.
[0066] Exemplarily, mark the two adjacent pictures as including the (i - 1)th frame picture and the ith frame picture in the video frame sequence, then the generated next picture can be the (i + 1)th frame picture.
[0067] As an application scenario, the image processing method provided in this application can be applied to a video stream transmission scenario: the video frame sequence may include each video frame in the video stream sent from the client to the server, and two adjacent frames may include two adjacent video frames in the video stream. In this application, the execution entity (such as a terminal) can respond to the upload operation of the client sending the video stream to the server, and predict the (i + 1)-th video frame (i.e., the next frame) based on the currently uploaded video frame (such as the i-th frame) and the previous frame (the (i - 1)-th frame) in the video stream. Thus, bandwidth allocation can be performed based on the predicted (i + 1)-th video frame, and the video stream can be encoded according to the allocated bandwidth and then uploaded to the server.
[0068] For example, for a video part with high predictability in the video stream, such as a relatively static video part in a landscape video, the transmission bit rate can be appropriately reduced, and more bandwidth can be allocated to the video part that is difficult to predict in the video stream, such as a suddenly appearing animal or vehicle.
[0069] Thus, during the video stream transmission process, predicting the content of subsequent video frames in advance can optimize bandwidth allocation and reduce transmission delay.
[0070] As another application scenario, the image processing method provided in this application can be applied to an autonomous driving scenario: the video frame sequence may include multiple captured frames continuously captured by an in-vehicle camera of a vehicle, and two adjacent frames may include two captured frames captured by the in-vehicle camera adjacent times. In this application, the execution entity (such as a vehicle) can respond to the power-on operation of the vehicle, and predict the (i + 1)-th captured frame (i.e., the next frame) based on the currently captured frame (such as the i-th frame) and the previously captured frame (the (i - 1)-th frame) by the in-vehicle camera. Thus, action prediction can be performed on obstacles in the environment where the vehicle is located based on the (i + 1)-th captured frame to obtain an action prediction result, which is denoted as the first prediction result in this application, so as to be used for controlling the driving of the vehicle based on the first prediction result.
[0071] Among them, action prediction includes predicting various behaviors of obstacles (such as pedestrians, other vehicles, etc.), such as moving direction, speed change, and whether to stop. The purpose of action prediction is to "control the driving of the vehicle", that is, based on the action prediction result, to adjust the driving state of the vehicle, such as accelerating, decelerating, steering, avoiding, etc.
[0072] Alternatively, the execution entity can predict the position information of obstacles in the environment where the vehicle is located and / or the lane-changing behavior of the obstacles based on the (i + 1)-th captured frame to obtain a second prediction result, so as to perform path planning for the vehicle based on the second prediction result to ensure driving safety.
[0073] Alternatively, the execution entity can perform path planning and driving control on the vehicle according to the first prediction result and the second prediction result.
[0074] Thus, during the autonomous driving of a vehicle, predicting the subsequent captured images of an in-vehicle camera in advance enables the vehicle's perception system to better understand the surrounding environment using video prediction, thereby ensuring the driving safety of the vehicle.
[0075] As another application scenario, the image processing method provided in this application can be applied to a game scenario: the video frame sequence may include multiple rendered game images in the game, and two adjacent frames of images may include game images obtained by two adjacent renderings in the game. In this application, an execution entity (such as a terminal) can, in response to a trigger operation for the game, predict the game image of the (i + 1)-th frame (i.e., the next image) based on the currently rendered game image (such as the i-th frame) and the previously rendered game image (such as the (i - 1)-th frame) in the game, so as to perform post-image processing based on the game image of the (i + 1)-th frame and render it for display.
[0076] Thus, in the game field, predicting the subsequent game images in the game in advance helps to achieve a smoother user operation response and improve the realism and immediacy of the interaction.
[0077] It should be noted that the above application scenarios are only for illustrative purposes, but this application is not limited thereto and can also be applied to other fields, and the video frame sequence and two adjacent frames of images in the video frame sequence can be obtained by other means.
[0078] The image processing method according to the embodiment of this application comprehensively performs video prediction by using two adjacent frames of images in the video frame sequence and the rendering information corresponding to the two adjacent frames of images to generate the next image after the two adjacent frames of images. This can not only improve the image quality of the predicted next image, but also, since video prediction only needs to use the relevant information of two adjacent frames of images, it can reduce the need for complex calculations, thereby reducing the power consumption of the device, extending the battery life of the device, and improving the user experience.
[0079] The embodiment of this application provides another image processing method. Figure 2 It is a schematic flowchart of the second image processing method provided by the embodiment of this application.
[0080] It should be noted that this image processing method can be executed alone, or can also be executed in combination with any one of the embodiments in this application or possible implementation manners in the embodiments, or can also be executed in combination with any one of the technical solutions in the related art. The embodiment of this application does not limit this.
[0081] As Figure 2 shown, this image processing method may include the following steps S201 to S204:
[0082] Step S201: Obtain two adjacent frames in the video frame sequence and the rendering information corresponding to the frames; wherein, the rendering information includes the first motion vector and the first depth information of the corresponding frame.
[0083] Wherein, in the case that a certain frame is a non-first frame in the video frame sequence, the first motion vector of this frame is used to indicate the motion information of this frame distorted towards the previous frame, and in the case that this frame is the first frame in the video frame sequence, the first motion vector of this frame is a set vector.
[0084] It should be noted that the explanatory description of step S201 can be referred to the relevant description in any embodiment of the present application, and will not be elaborated here.
[0085] Step S202: Predict the second motion vector of the next frame after the two adjacent frames according to the first motion vector and the first depth information corresponding to the two adjacent frames.
[0086] Wherein, the two adjacent frames may include the (i - 1)-th frame and the i-th frame in the video frame sequence, and the next frame may include the (i + 1)-th frame, where i is a positive integer greater than 1.
[0087] Wherein, the second motion vector is used to indicate the motion information of the (i + 1)-th frame distorted towards the i-th frame, and the second motion vector is used to indicate the motion component of each pixel point in the (i + 1)-th frame. Each pixel point's motion component includes the motion information of this pixel point in two channels, namely the motion information in the X direction and the motion information in the Y direction. Exemplarily, the second motion vector can be denoted as
[0088] As an example, the deep learning technology in the field of artificial intelligence can be adopted to predict the motion vector of the (i + 1)-th frame according to the first motion information of the (i - 1)-th frame, the first motion information of the i-th frame, the first depth information of the (i - 1)-th frame, and the first depth information of the i-th frame. In the present application, it is denoted as the second motion vector.
[0089] As another example, the rendering information may further include the first albedo information of the corresponding frame. In the present application, the deep learning technology can be adopted to predict the second motion vector of the (i + 1)-th frame according to the first motion information of the (i - 1)-th frame, the first motion information of the i-th frame, the first depth information of the (i - 1)-th frame, the first depth information of the i-th frame, the first albedo information of the (i - 1)-th frame, and the first albedo information of the i-th frame.
[0090] Step S203: Based on the second motion vector and the first motion vectors corresponding to the two adjacent frames, transform the two adjacent frames to obtain two transformed frames.
[0091] In an embodiment of the present application, the (i - 1)-th frame of the picture can be transformed according to the second motion vector of the (i + 1)-th frame of the picture and the first motion vector of the (i - 1)-th frame of the picture to obtain a transformed frame of the picture, which is denoted as the (i - 1)-th transformed frame of the picture in the present application, and the i-th frame of the picture is transformed according to the second motion vector of the (i + 1)-th frame of the picture and the first motion vector of the i-th frame of the picture to obtain another transformed frame of the picture, which is denoted as the i-th transformed frame of the picture in the present application.
[0092] In any one of the embodiments of the present application, the final motion vector from the (i - 1)-th frame of the picture to the (i + 1)-th frame of the picture can be generated according to the second motion vector of the (i + 1)-th frame of the picture and the first motion vector of the (i - 1)-th frame of the picture, which is denoted as the fifth motion vector in the present application, and based on this fifth motion vector, the (i - 1)-th frame of the picture is transformed to obtain the (i - 1)-th transformed frame of the picture.
[0093] Exemplarily, the (i - 1)-th frame of the picture is marked as I i-1 , the first motion vector of the (i - 1)-th frame of the picture is MV i-1 , the fifth motion vector is MV (i-1)to(i+1) , then the (i - 1)-th transformed frame of the picture can be: warp(I i-1 , MV (i-1)to(i+1) ).
[0094] In any one of the embodiments of the present application, the final motion vector from the i-th frame of the picture to the (i + 1)-th frame of the picture can be generated according to the second motion vector of the (i + 1)-th frame of the picture and the first motion vector of the i-th frame of the picture, which is denoted as the sixth motion vector in the present application, and based on this sixth motion vector, the i-th frame of the picture is transformed to obtain the i-th transformed frame of the picture.
[0095] Exemplarily, the i-th frame of the picture is marked as I i , the first motion vector of the i-th frame of the picture is MV i , the sixth motion vector is MV (i)to(i+1) , then the i-th transformed frame of the picture can be: warp(I i , MV (i)to(i+1) ).
[0096] In summary, targeted transformation processing can be performed on two adjacent frames of pictures, improving the accuracy of subsequent video prediction, that is, improving the generation quality of the (i + 1)-th frame of the picture.
[0097] Step S204: Fuse the two transformed frames of the picture to obtain the next frame of the picture.
[0098] In an embodiment of the present application, the (i - 1)-th transformed frame of the picture and the i-th transformed frame of the picture can be fused based on an image fusion technology to obtain the (i + 1)-th frame of the picture.
[0099] As an example, the following formula can be used to fuse the transformed frame of the (i-1)-th frame and the transformed frame of the i-th frame:
[0100]
[0101] where M is the weight, denotes the generated frame of the (i + 1)-th frame.
[0102] The image processing method according to the embodiment of the present application combines the first motion vector and the first depth information in the rendering information of two adjacent frames, predicts the second motion vector of the next frame after two adjacent frames, can improve the prediction accuracy of the second motion vector, and then based on the accurate second motion vector, transform and fuse the two adjacent frames to obtain the next frame, which can improve the generation quality of the next frame.
[0103] The embodiment of the present application provides another image processing method, Figure 3 which is a schematic flowchart of the third image processing method provided by the embodiment of the present application.
[0104] It should be noted that this image processing method can be executed alone, or can be executed together with any one of the embodiments or possible implementation manners in the present application, or can also be executed together with any one of the technical solutions in the related art. The embodiment of the present application does not limit this.
[0105] As Figure 3 shown, this image processing method may include the following steps S301 to S306:
[0106] Step S301, obtain two adjacent frames in the video frame sequence and the rendering information corresponding to the frames; wherein, the rendering information includes the first motion vector and the first depth information of the corresponding frame.
[0107] It should be noted that the explanation of step S301 can refer to the relevant description in any embodiment of the present application, and will not be elaborated here.
[0108] Step S302, generate an initial motion vector of the next frame after two adjacent frames according to the first motion vector and the first depth information corresponding to the two adjacent frames.
[0109] Among them, the two adjacent frames may include the frame of the (i-1)-th frame and the frame of the i-th frame in the video frame sequence, and the next frame may include the frame of the (i + 1)-th frame, where i is a positive integer greater than 1.
[0110] In the embodiment of the present application, the first motion vector MV of the (i-1)-th frame i-1 and the first motion vector MV of the i-th frame i, the first depth information D of the (i - 1)-th frame image i-1 , the first depth information D of the i-th frame image i , to generate an initial motion vector of the (i + 1)-th frame image
[0111] As a possible implementation, first, according to the first motion vector MV of the (i - 1)-th frame image i-1 and the first motion vector MV of the i-th frame image i , calculate the motion difference of the pixel pair; wherein, the pixel pair includes any first pixel point of the (i - 1)-th frame image and the second pixel point corresponding to the first pixel point in the i-th frame image. That is, according to MV i-1 , determine the motion component of the first pixel point, and according to MV i , determine the motion component of the second pixel point, and take the difference between the motion component of the first pixel point and the motion component of the second pixel point as the motion difference of the pixel pair.
[0112] After that, it can be determined whether the motion difference of the pixel pair is less than a set difference threshold. If the motion difference of the pixel pair is less than the difference threshold, then take the motion component of the second pixel point in the first motion information of the i-th frame image as the motion component of the third pixel point corresponding to the second pixel point in the (i + 1)-th frame image.
[0113] If the motion difference of the pixel pair is greater than or equal to the difference threshold, then according to the first depth information D of the (i - 1)-th frame image i-1 and the first depth information D of the i-th frame image i , calculate the depth difference of the pixel pair, that is, according to D i-1 , determine the depth of the first pixel point, and according to D i , determine the depth of the second pixel point, and take the difference between the depth of the first pixel point and the depth of the second pixel point as the depth difference of the pixel pair. Then, according to the motion component of the second pixel point and the depth difference of the pixel pair, calculate the motion component of the third pixel point corresponding to the second pixel point in the (i + 1)-th frame image.
[0114] Thus, in this application, according to the motion components of each third pixel point in the (i + 1)-th frame image, an initial motion vector of the (i + 1)-th frame image can be generated
[0115] Exemplarily, mark the coordinate positions of the first pixel point, the second pixel point, and the third pixel point as (x, y), and the difference threshold as T1. The following formula can be used to calculate and obtain
[0116]
[0117] Among them, α is a preset parameter.
[0118] In summary, by comprehensively considering the motion difference and depth difference between two adjacent frames of images to calculate the initial motion vector of the next frame after the two adjacent frames of images, the calculated initial motion vector can be made smoother and more stable, thereby improving the quality of subsequent video prediction.
[0119] Step S303: Generate a third motion vector from two adjacent frames of images pointing to the next frame according to the first motion vector corresponding to the two adjacent frames of images and the initial motion vector of the next frame.
[0120] In the embodiment of the present application, the initial motion vector from the (i - 1)-th frame of image to the (i + 1)-th frame of image can be calculated according to the first motion vector of the (i - 1)-th frame of image and the initial motion vector of the (i + 1)-th frame of image. In the present application, it is denoted as the third motion vector MV'. (i-1)to(i+1) 。
[0121] Similarly, the third motion vector MV' from the i-th frame of image to the (i + 1)-th frame of image can be calculated according to the first motion vector of the i-th frame of image and the initial motion vector of the (i + 1)-th frame of image. (i)to(i+1) 。
[0122] Step S304: Predict the second motion vector of the next frame based on the third motion vector corresponding to the two adjacent frames of images and the first depth information.
[0123] Exemplarily, based on deep learning technology, according to the third motion vector MV' corresponding to the (i - 1)-th frame of image (i-1)to(i+1) , the third motion vector MV' corresponding to the i-th frame of image (i)to(i+1) , the first depth information D of the (i - 1)-th frame of image i-1 , the first depth information D of the i-th frame of image i , predict the second motion vector of the i-th frame of image
[0124] Step S305: Transform the two adjacent frames of images based on the second motion vector and the first motion vector corresponding to the two adjacent frames of images to obtain two transformed frames of images.
[0125] Step S306: Fuse the two transformed frames of images to obtain the next frame.
[0126] It should be noted that the explanations of steps S305 to S306 can refer to the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0127] The image processing method according to the embodiment of the present application first generates an initial motion vector corresponding to the next frame according to the first motion vector and the first depth information corresponding to two adjacent frames, and then generates a third motion vector from the two adjacent frames to the next frame according to the first motion vector corresponding to the two adjacent frames and the initial motion vector of the next frame. Furthermore, by comprehensively considering the third motion vector corresponding to the two adjacent frames and the first depth information, the second motion vector of the next frame is predicted, which can improve the prediction accuracy of the second motion vector.
[0128] Another image processing method is provided in the embodiment of the present application. Figure 4 It is a schematic flowchart of the fourth image processing method provided in the embodiment of the present application.
[0129] It should be noted that this image processing method can be executed alone, or can be executed in combination with any one of the embodiments or possible implementation manners in the present application, or can also be executed in combination with any one of the technical solutions in the related art. The embodiments of the present application do not limit this.
[0130] As Figure 4 shown, this image processing method may include the following steps S401 to S408:
[0131] Step S401: Obtain two adjacent frames in the video frame sequence and the rendering information corresponding to the frames; wherein, the rendering information includes the first motion vector, the first depth information, and the first albedo information corresponding to the corresponding frame.
[0132] Step S402: Generate an initial motion vector of the next frame after the two adjacent frames according to the first motion vector and the first depth information corresponding to the two adjacent frames.
[0133] Step S403: Generate a third motion vector from the two adjacent frames to the next frame according to the first motion vector corresponding to the two adjacent frames and the initial motion vector of the next frame.
[0134] It should be noted that the explanations of steps S401 to S403 can be found in the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0135] Step S404: Based on the third motion vector corresponding to the two adjacent frames, perform a transformation process on the first albedo information of the two adjacent frames to obtain the second albedo information corresponding to the two adjacent frames.
[0136] Among them, the two adjacent frames may include the (i - 1)-th frame and the i-th frame in the video frame sequence, and the next frame may include the (i + 1)-th frame, where i is a positive integer greater than 1.
[0137] In an embodiment of the present application, the third motion vector MV′ from the (i - 1)-th frame to the (i + 1)-th frame can be used (i-1)to(i+1) , to perform a transformation process on the first albedo information A of the (i - 1)-th frame i-1 , to obtain the transformed A i-1 , which is denoted as the second albedo information A corresponding to the (i - 1)-th frame in the present application (i-1)to(i+1) .
[0138] Similarly, the third motion vector MV′ from the i-th frame to the (i + 1)-th frame can be used (i)to(i+1) , to perform a transformation process on the first albedo information A of the i-th frame i , to obtain the transformed A i , which is denoted as the second albedo information A corresponding to the i-th frame in the present application (i)to(i+1) .
[0139] Step S405: Based on the third motion vectors corresponding to two adjacent frames, perform a transformation process on the first depth information of the two adjacent frames to obtain the second depth information corresponding to the two adjacent frames.
[0140] In an embodiment of the present application, the third motion vector MV′ from the (i - 1)-th frame to the (i + 1)-th frame can be used (i-1)to(i+1) , to perform a transformation process on the first depth information D of the (i - 1)-th frame i-1 , to obtain the transformed D i-1 , which is denoted as the second depth information D corresponding to the (i - 1)-th frame in the present application (i-1)to(i+1) .
[0141] Similarly, the third motion vector MV′ from the i-th frame to the (i + 1)-th frame can be used (i)to(i+1) , to perform a transformation process on the first depth information D of the i-th frame i , to obtain the transformed D i , which is denoted as the second depth information D corresponding to the i-th frame in the present application (i)to(i+1) .
[0142] Step S406: Predict the second motion vector of the next frame according to the second albedo information and the second depth information corresponding to two adjacent frames, and according to the first motion vector and the initial motion vector of the latter frame in the two adjacent frames.
[0143] As an example, the second albedo information A (i-1)to(i+1) and the second depth information D (i-1)to(i+1) corresponding to the (i - 1)-th frame, the first motion vector MV i , the second albedo information A (i)to(i+1) and the second depth information D (i)to(i+1), and the initial motion vector of the (i + 1)-th frame Input it into a deep learning model or neural network, so that the model predicts the final motion vector corresponding to the i-th frame, which is denoted as the second motion vector in this application
[0144] Step S407: Based on the second motion vector and the first motion vectors corresponding to two adjacent frames, perform transformation on the two adjacent frames to obtain two transformed frames
[0145] Step S408: Fuse the two transformed frames to obtain the next frame after the two adjacent frames
[0146] It should be noted that the explanations of steps S407 to S408 can refer to the relevant descriptions in any embodiment of this application, and will not be elaborated here
[0147] In any embodiment of this application, the second albedo information A corresponding to the (i - 1)-th frame (i-1)to(i+1) and the second depth information D (i-1)to(i+1) , the first motion vector MV corresponding to the i-th frame i , the second albedo information A (i)to(i+1) and the second depth information D (i)to(i+1) , and the initial motion vector of the (i + 1)-th frame Predict the mask weight (Mask, abbreviated as M). In this application, the two transformed frames can be weighted based on the mask weight to obtain the (i + 1)-th frame. The implementation principle can refer to formula (1), and will not be elaborated here
[0148] As an example, the deep learning model or neural network can also output the mask weight M. In this application, based on the mask weight M output by the model, the (i - 1)-th transformed frame and the i-th transformed frame can be weighted using the above formula (1) to obtain the (i + 1)-th frame
[0149] Exemplarily, the Unet model can be used to predict the second motion vector and the mask weight M of the i-th frame, where is a 2-channel tensor representing the motion information in the XY direction; M is a single-channel mask weight, which is activated by sigmoid (activation function) to ensure that its value range is between 0 and 1
[0150] The image processing method according to the embodiment of the present application synthesizes the first motion vector, the first depth information, and the first albedo information in the rendering information of two adjacent frames of images to predict the second motion vector of the next frame of image after the two adjacent frames of images, which can further improve the prediction accuracy of the second motion vector. Furthermore, based on the accurate second motion vector, the two adjacent frames of images are transformed and fused to obtain the next frame of image, which can further improve the generation quality of the next frame of image.
[0151] The embodiment of the present application provides another image processing method. Figure 5 It is a schematic flowchart of the fifth image processing method provided by the embodiment of the present application.
[0152] It should be noted that this image processing method can be executed alone, or it can be executed in combination with any one of the embodiments in the present application or the possible implementation manners in the embodiments, or it can also be executed in combination with any one of the technical solutions in the related technologies. The embodiments of the present application do not limit this.
[0153] As Figure 5 shown, this image processing method may include the following steps S501 to S509:
[0154] Step S501, obtain two adjacent frames of images in the video frame sequence and the rendering information corresponding to the images; wherein, the rendering information includes the first motion vector and the first depth information of the corresponding image, and the bone point information of the character in the corresponding image.
[0155] Step S502, generate an initial motion vector of the next frame of image after the two adjacent frames of images according to the first motion vector and the first depth information corresponding to the two adjacent frames of images.
[0156] It should be noted that the explanations of steps S501 to S502 can refer to the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0157] Step S503, obtain the bone point information of the character in the next frame of image; wherein, the bone point information of the character in the next frame of image is predicted according to the bone point information of the character in the two adjacent frames of images.
[0158] Among them, the bone point information of any frame of image is used to indicate the positions of the respective bone points of the character in the image. Among them, the character includes but is not limited to a human character and an animal character.
[0159] Among them, the two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1.
[0160] In the embodiment of the present application, the bone point information B of the character in the (i - 1)-th frame of image can be used.i-1 and the skeletal point information B of the character in the i-th frame of the picture i , predict the skeletal point information of the (i + 1)-th frame of the picture
[0161] Exemplarily, B i-1 and B i can be input into a neural network (such as a fully connected neural network) for predicting the skeletal point positions, to obtain the skeletal point information of the (i + 1)-th frame of the picture
[0162] It should be noted that since the dimension of the skeletal point information is very low and the number of layers of the fully connected neural network is small, the computing resources required for this step are low, which can improve the prediction speed.
[0163] Step S504, based on the skeletal point information of the next picture and the skeletal point information of the latter frame in two adjacent frames of pictures, determine the fourth motion vector of the skeletal points of the character in the next picture.
[0164] In the embodiment of the present application, according to the skeletal point information of the (i + 1)-th frame of the picture and the skeletal point information B of the i-th frame of the picture i , the motion vector of the skeletal point positions of the character in the (i + 1)-th frame of the picture can be calculated, which is denoted as the fourth motion vector in the present application. That is, the skeletal points of B i and are in one-to-one correspondence, and the motion relationship of the corresponding skeletal points from the i-th frame of the picture to the (i + 1)-th frame of the picture can be calculated to obtain the fourth motion vector.
[0165] Exemplarily, the fourth motion vector of the skeletal point positions of the character in the (i + 1)-th frame of the picture can be marked as
[0166] Step S505, based on the fourth motion vector, adjust the initial motion vector of the next picture.
[0167] Considering that only the skeletal point positions in the fourth motion vector have values, and the rest are all 0 (that is, the motion vector of the skeletal points cannot judge the motion information of the surrounding area and can only judge the single-point motion), in the present application, the fourth motion vector can be subjected to Gaussian blur to spread or blur the motion information of to the surrounding pixel points, so that the surrounding pixel points also have motion information. Thus, in the present application, based on the Gaussian-blurred fourth motion vector, the initial motion vector of the (i + 1)-th frame of the picture can be adjusted.
[0168] Exemplarily, the following formula can be used to adjust the initial motion vector of the (i + 1)-th frame of the picture Make adjustments:
[0169]
[0170] Among them, Gaussian(·) represents Gaussian blur, and β is a preset parameter.
[0171] Step S506: Generate a third motion vector from two adjacent frames of pictures to the next picture according to the adjusted initial motion vector of the next picture and the first motion vector of two adjacent frames of pictures.
[0172] In the embodiment of the present application, the third motion vector MV' from the (i - 1)-th frame of picture to the (i + 1)-th frame of picture can be calculated according to the first motion vector MV i-1 of the (i - 1)-th frame of picture and the adjusted initial motion vector of the (i + 1)-th frame of picture . (i-1)to(i+1) .
[0173] Similarly, the third motion vector MV' from the i-th frame of picture to the (i + 1)-th frame of picture can be calculated according to the first motion vector MV i of the i-th frame of picture and the adjusted initial motion vector of the (i + 1)-th frame of picture . (i)to(i+1) .
[0174] Step S507: Predict the second motion vector of the next picture based on the third motion vector corresponding to two adjacent frames of pictures and the first depth information.
[0175] Step S508: Transform two adjacent frames of pictures based on the second motion vector and the first motion vectors corresponding to two adjacent frames of pictures to obtain two transformed frames of pictures.
[0176] Step S509: Fuse two transformed frames of pictures to obtain the next picture.
[0177] It should be noted that the explanations of steps S507 to S509 can be referred to the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0178] The image processing method of the embodiment of the present application comprehensively calculates the motion vector of the next picture by combining the motion vectors, depth information, and skeletal point information of the character in two adjacent frames of pictures, which can improve the rationality and accuracy of the calculation results.
[0179] To implement the above embodiments, the present application also proposes a chip.
[0180] Figure 6 It is a schematic structural diagram of the first chip provided by the embodiment of the present application.
[0181] As Figure 6As shown, the chip 600 may include a GPU 610 and an NPU 620 that communicate with each other. Among them,
[0182] The GPU 610 is configured to obtain two adjacent frames in the video frame sequence and the rendering information corresponding to the frames, and send the rendering information corresponding to the frames to the NPU 620;
[0183] The NPU 620 is configured to predict the motion vector of the next frame after two adjacent frames according to the rendering information corresponding to the frames, and send the motion vector to the GPU 610;
[0184] Among them, the motion vector of the next frame may be the second motion vector in the foregoing embodiments.
[0185] The GPU 610 is further configured to generate the next frame according to two adjacent frames and the motion vector.
[0186] In any one of the embodiments of the present application, two adjacent frames in the video frame sequence include: two adjacent video frames in the video stream sent from the client to the server; the chip 600 is further configured to: perform bandwidth allocation according to the next frame, and encode the video stream according to the allocated bandwidth and upload it to the server. In any one of the embodiments of the present application, two adjacent frames in the video frame sequence include the shooting pictures captured by the in-vehicle camera of the vehicle twice in succession; the chip 600 is further configured to perform at least one of the following:
[0187] Based on the next frame, perform action prediction on the obstacles in the environment where the vehicle is located to obtain a first prediction result; wherein, the first prediction result is used for controlling the driving of the vehicle;
[0188] Based on the next frame, predict the position information of the obstacles in the environment where the vehicle is located and / or the lane-changing behavior of the obstacles to obtain a second prediction result; wherein, the second prediction result is used for path planning of the vehicle.
[0189] In any one of the embodiments of the present application, two adjacent frames in the video frame sequence include the game pictures rendered twice in succession in the game; the chip 600 is further configured to: perform post-image processing based on the next frame and render for display.
[0190] In any one of the embodiments of the present application, the rendering information includes at least one of the following:
[0191] The first motion vector corresponding to the frame; wherein, in response to the corresponding frame being a non-first frame in the video frame sequence, the first motion vector is used to indicate the motion information of the corresponding frame distorted towards the previous frame, and in response to the corresponding frame being the first frame in the video frame sequence, the first motion vector is a set vector;
[0192] The first depth information of the corresponding image;
[0193] The first albedo information of the corresponding image;
[0194] The bone point information of the character in the corresponding image; wherein, the bone point information is used to indicate the positions of each bone point of the character.
[0195] In any embodiment of the present application, the NPU 620 is configured to: predict the second motion vector of the next image according to the first motion vector and the first depth information corresponding to two adjacent frames of images.
[0196] In any embodiment of the present application, the GPU 610 is configured to: transform two adjacent frames of images based on the second motion vector and the first motion vector corresponding to two adjacent frames of images to obtain two transformed frames of images; fuse the two transformed frames of images to obtain the next image.
[0197] In any embodiment of the present application, the NPU 620 is configured to: generate an initial motion vector of the next image according to the first motion vector and the first depth information corresponding to two adjacent frames of images; generate a third motion vector from two adjacent frames of images to the next image according to the first motion vector and the initial motion vector of the next image corresponding to two adjacent frames of images; predict the second motion vector of the next image based on the third motion vector and the first depth information corresponding to two adjacent frames of images.
[0198] In any embodiment of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the NPU 620 is configured to: obtain the motion difference of the pixel pair according to the first motion vector corresponding to two adjacent frames of images; wherein, the pixel pair includes any first pixel point of the (i - 1)-th frame of image and the second pixel point corresponding to the first pixel point in the i-th frame of image; in response to the motion difference being less than the difference threshold, use the motion component of the second pixel point in the first motion information of the i-th frame of image as the motion component of the third pixel point corresponding to the second pixel point in the (i + 1)-th frame of image; or, in response to the motion difference being greater than or equal to the difference threshold, determine the depth difference of the pixel pair according to the first depth information corresponding to two adjacent frames of images, and determine the motion component of the third pixel point according to the motion component of the second pixel point and the depth difference; generate the initial motion vector of the (i + 1)-th frame of image according to the motion components of each third pixel point in the (i + 1)-th frame of image.
[0199] In any one of the embodiments of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the NPU 620 is configured to: perform transformation processing on the first albedo information of the two adjacent frames of images based on the third motion vector corresponding to the two adjacent frames of images to obtain the second albedo information corresponding to the two adjacent frames of images; perform transformation processing on the first depth information of the two adjacent frames of images based on the third motion vector corresponding to the two adjacent frames of images to obtain the second depth information corresponding to the two adjacent frames of images; predict the second motion vector of the (i + 1)-th frame of image according to the second albedo information and the second depth information corresponding to the two adjacent frames of images, and according to the first motion vector and the initial motion vector of the i-th frame of image.
[0200] In any one of the embodiments of the present application, the GPU 610 is configured to: obtain the mask weight sent by the NPU; where the mask weight is predicted by the NPU according to the second albedo information and the second depth information corresponding to the two adjacent frames of images, and according to the first motion vector and the initial motion vector of the i-th frame of image; perform weighting on the two transformed frames of images based on the mask weight to obtain the (i + 1)-th frame of image.
[0201] In any one of the embodiments of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the NPU 620 is configured to: obtain the bone point information of the character in the (i + 1)-th frame of image; where the bone point information of the character in the (i + 1)-th frame of image is predicted according to the bone point information of the character in the two adjacent frames of images; determine the fourth motion vector of the bone points of the character in the (i + 1)-th frame of image based on the bone point information of the (i + 1)-th frame of image and the bone point information of the i-th frame of image; adjust the initial motion vector of the (i + 1)-th frame of image based on the fourth motion vector; generate the third motion vector pointing from the two adjacent frames of images to the (i + 1)-th frame of image according to the adjusted initial motion vector and according to the first motion vectors of the two adjacent frames of images.
[0202] In any one of the embodiments of the present application, the chip 600 may further include: a CPU, where the CPU is configured to predict the bone point information of the character in the (i + 1)-th frame of image according to the bone point information of the character in the two adjacent frames of images, and send the bone point information of the character in the (i + 1)-th frame of image to the NPU 620.
[0203] In any embodiment of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the GPU 610 is configured to: generate a fifth motion vector from the (i - 1)-th frame of image to the (i + 1)-th frame of image according to the second motion vector and the first motion vector of the (i - 1)-th frame of image; perform transformation on the (i - 1)-th frame of image based on the fifth motion vector to obtain the transformed (i - 1)-th frame of image; generate a sixth motion vector from the i-th frame of image to the (i + 1)-th frame of image according to the second motion vector and the first motion vector of the i-th frame of image; perform transformation on the i-th frame of image based on the sixth motion vector to obtain the transformed i-th frame of image.
[0204] It should be noted that the above explanation of the embodiments of the image processing method also applies to the chip of this embodiment, and will not be elaborated here.
[0205] In the chip of the embodiment of the present application, by using the NPU to comprehensively predict the motion vector of the next frame of image after two adjacent frames of images in the video frame sequence, the prediction efficiency and prediction accuracy can be improved. Furthermore, the GPU comprehensively uses the accurate motion vector and two adjacent frames of images for video prediction to generate the next frame of image, which can not only improve the image quality of the predicted next frame of image, but also reduce the demand for complex calculations by only using the relevant information of two adjacent frames of images, thereby reducing the power consumption of the device, prolonging the battery life of the device, and enhancing the user experience.
[0206] In any embodiment of the present application, taking the scenario where the solution provided by the present application is applied to a game scene as an example, the modeling information (such as the bone point information of a character, etc.) and physical information (such as depth, albedo, motion vector, etc.) used in the game rendering process, as well as the (i - 1)-th game frame and the i-th game frame, can be used to generate the (i + 1)-th game frame. During the process of generating the game frame, the motion vector (or motion vector) is calculated more accurately, and the generated image quality is better. Moreover, only the previous and next two game frames are used for image processing to obtain the (i + 1)-th game frame, which can significantly reduce the rendering delay, rendering power consumption, and calculation overhead. While prolonging the battery life, it can effectively reduce the latency of the game and improve the user's game experience, enabling the solution to be implemented on a mobile terminal.
[0207] As an example, video prediction can be achieved through the interaction between the NPU, GPU, and CPU in the chip. The video prediction principle is as Figure 7 shown, and mainly includes the following steps:
[0208] The GPU side executes the following steps a to b:
[0209] Step a: Access the game rendering pipeline and read the motion vector MV i-1 of the (i - 1)-th frame (such as the previous frame) game screen I generated during the rendering process i-1 , depth information D i-1 , albedo information A i-1 , and read the motion vector MV i of the i-th frame (such as the current frame) game screen I i , depth information D i , albedo information A i .
[0210] Step b: Transmit the above information to the NPU.
[0211] The CPU side executes the following steps c to d:
[0212] Step c: (Optional) Read the bone point information B i-1 of the character in the (i - 1)-th frame game screen during the modeling process i and the bone point information B
[0213] of the character in the i-th frame game screen. Step d: (Optional) Use a neural network (such as a fully connected neural network) to predict the bone point information i-1 of the (i + 1)-th frame game screen based on the bone point information B i of two adjacent frame game screens. Since the dimension of the bone point information is very low and the number of layers of the fully connected neural network is small, this calculation step can be executed on the CPU.
[0214] The NPU side executes the following steps e to h:
[0215] Step e: After the NPU receives the information transmitted by the GPU, based on the motion vector MV i-1 of the (i - 1)-th frame game screen I i-1 , depth information D i-1 , and based on the motion vector MV i of the i-th frame game screen I i , depth information D i , generate the initial motion vector of the (i + 1)-th frame game screen
[0216] First, the motion difference (such as the difference value) between MV i-1 and MV i can be calculated. If the difference of the pixel point at the coordinate (x, y) is less than the difference threshold T1, then at this pixel point:
[0217]
[0218] Otherwise, at this pixel point, there is:
[0219]
[0220] where α is a preset parameter.
[0221] Step f: (Optional) Read the skeletal point information of the (i + 1)-th frame of the game screen from the CPU side and accordingly correct or fine-tune the initial motion vector in Step 2
[0222] First, based on the skeletal point information B of the character in the i-th frame of the game screen i and the skeletal point information of the (i + 1)-th frame of the game screen calculate the motion vector of the skeletal point position of the character in the (i + 1)-th frame of the game screen
[0223] After that, based on perform fine-tuning on Exemplarily, the fine-tuning formula can be as follows:
[0224]
[0225] where Gaussian(·) represents Gaussian blur and β is a preset parameter.
[0226] Step g: Use the motion vector MV of the (i - 1)-th frame of the game screen i-1 and the initial motion vector of the (i + 1)-th frame of the game screen to generate the initial motion vector MV′ from the (i - 1)-th frame of the game screen to the (i + 1)-th frame of the game screen (i-1)to(i+1) , and accordingly transform the albedo information A i-1 and depth information D i-1 of the (i + 1)-th frame of the game screen to obtain the transformed albedo information A (i-1)to(i+1) and the transformed depth information D (i-1)to(i+1) ; Similarly, use the motion vector MV of the i-th frame of the game screen i and the initial motion vector of the (i + 1)-th frame of the game screen to generate the initial motion vector MV′ from the i-th frame of the game screen to the (i + 1)-th frame of the game screen (i)to(i+1) , and accordingly transform the albedo information A i and depth information D i of the i-th frame of the game screen to obtain the transformed albedo information A (i)to(i+1) and the transformed depth information D (i)to(i+1) .
[0227] Step h: Use the transformed albedo information A (i-1)to(i+1)and A (i)to(i+1) and the transformed D (i-1)to(i+1) and D (i)to(i+1) the motion vector MV of the i-th frame of the game screen i the initial motion vector of the (i + 1)-th frame of the game screen Input into the neural network to obtain the final motion vector of the (i + 1)-th frame of the game screen and the mask weight Mask, and transmit them to the GPU
[0228] Exemplarily, a UNet network can be used. The output of the UNet network can be a 2-channel tensor, representing the motion information in the XY direction; Mask is a single-channel mask weight, activated by sigmoid (activation function) to ensure that its value range is between 0 and 1
[0229] The GPU side performs the following step j
[0230] Step j: Use the final motion vector of the (i + 1)-th frame of the game screen the motion vector MV of the (i - 1)-th frame of the game screen I i-1 and the motion vector MV of the i-th frame of the game screen I i-1 to perform transformation processing on the (i - 1)-th frame of the game screen I i and the i-th frame of the game screen I i , and weight them using the mask weight to obtain the (i + 1)-th frame of the game screen i-1 and the i-th frame of the game screen I i (i.e., the predicted image), and return it to the GPU. The specific formula is as follows (i.e., the predicted image), return it to the GPU. The specific formula is as follows
[0231]
[0232] where MV (i-1)to(i+1) refers to the motion vector from the (i - 1)-th frame of the game screen to the (i + 1)-th frame of the game screen determined based on the final motion vector of the (i + 1)-th frame of the game screen and the motion vector MV of the (i - 1)-th frame of the game screen i-1 , warp(I i-1 , MV (i-1)to(i+1) ) refers to the image obtained by transforming the (i - 1)-th frame of the game screen I using MV (i-1)to(i+1) (denoted as the transformed image in this application); MV i-1 refers to the motion vector from the i-th frame of the game screen to the (i + 1)-th frame of the game screen determined based on the final motion vector of the (i + 1)-th frame of the game screen (i)to(i+1) and the motion vector MV of the i-th frame of the game screen and the motion vector MV of the i-th frame of the game screen i , warp(I i , MV(i)to(i+1) ) refers to using MV (i)to(i+1) to transform the i-th frame of the game screen I i into the transformed screen.
[0233] After predicting the (i + 1)-th frame of the game screen UI texture mapping, other post-processing, and display can be performed to render and display the game screen in real time.
[0234] In summary, by using the rendering information such as the motion vectors, depth, and albedo of the previous and next frames, and the scene modeling information such as the skeletal point information of the character, the next frame is comprehensively rendered. In this way, the motion vectors are calculated more accurately, the latency is lower, and the generated image quality is better.
[0235] To implement the above embodiments, an embodiment of the present application also proposes an image processing device.
[0236] Figure 8 FIG. is a schematic structural diagram of an image processing device provided by an embodiment of the present application.
[0237] As Figure 8 shown, the image processing device 800 may include: an acquisition module 810 and a generation module 820.
[0238] Among them, the acquisition module 810 is configured to acquire two adjacent frames of the video frame sequence and the rendering information corresponding to the frames;
[0239] The generation module 820 is configured to generate the next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames.
[0240] Further, in an implementation manner of the embodiment of the present application, two adjacent frames of the video frame sequence include: two adjacent video frames in the video stream sent from the client to the server; the image processing device 800 may further include:
[0241] An allocation module, configured to perform bandwidth allocation according to the next frame, and encode the video stream according to the allocated bandwidth and upload it to the server.
[0242] In an implementation manner of the embodiment of the present application, two adjacent frames of the video frame sequence include the shooting pictures taken by the in-vehicle camera of the vehicle twice in succession; the image processing device 800 may further include:
[0243] A prediction module, configured to perform at least one of the following:
[0244] Based on the next frame, perform action prediction on the obstacles in the environment where the vehicle is located to obtain a first prediction result; wherein, the first prediction result is used to control the driving of the vehicle;
[0245] Based on the next frame, predict the position information of obstacles in the environment where the vehicle is located and / or the lane-changing behavior of the obstacles, and obtain a second prediction result; wherein, the second prediction result is used for path planning of the vehicle.
[0246] In an implementation manner of the embodiments of the present application, two adjacent frames in the video frame sequence include game frames obtained by adjacent two renderings in the game; the image processing apparatus 800 may further include:
[0247] A processing module, configured to perform post-image processing based on the next frame and render for display.
[0248] In an implementation manner of the embodiments of the present application, the rendering information includes at least one of the following:
[0249] The first motion vector corresponding to the frame; wherein, in response to the corresponding frame being a non-first frame in the video frame sequence, the first motion vector is used to indicate the motion information of the corresponding frame distorted towards the previous frame, and in response to the corresponding frame being the first frame in the video frame sequence, the first motion vector is a set vector;
[0250] The first depth information corresponding to the frame;
[0251] The first albedo information corresponding to the frame;
[0252] The bone point information of the character in the corresponding frame; wherein, the bone point information is used to indicate the positions of the respective bone points of the character.
[0253] In an implementation manner of the embodiments of the present application, the generation module 820 is configured to: predict the second motion vector of the next frame according to the first motion vector and the first depth information corresponding to two adjacent frames; perform transformation on the two adjacent frames based on the second motion vector and the first motion vector corresponding to the two adjacent frames to obtain two transformed frames; and fuse the two transformed frames to obtain the next frame.
[0254] In an implementation manner of the embodiments of the present application, the generation module 820 is configured to: generate an initial motion vector of the next frame according to the first motion vector and the first depth information corresponding to two adjacent frames; generate a third motion vector from the two adjacent frames to the next frame according to the first motion vector corresponding to the two adjacent frames and the initial motion vector of the next frame; and predict the second motion vector of the next frame based on the third motion vector corresponding to the two adjacent frames and the first depth information.
[0255] In an implementation manner of the embodiment of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the generation module 820 is configured to: obtain the motion difference of pixel pairs according to the first motion vectors corresponding to two adjacent frames of images; where the pixel pairs include any first pixel point in the (i - 1)-th frame of image and the second pixel point corresponding to the first pixel point in the i-th frame of image; in response to the motion difference being less than the difference threshold, use the motion component of the second pixel point in the first motion information of the i-th frame of image as the motion component of the third pixel point corresponding to the second pixel point in the (i + 1)-th frame of image; or, in response to the motion difference being greater than or equal to the difference threshold, determine the depth difference of the pixel pairs according to the first depth information corresponding to two adjacent frames of images, and determine the motion component of the third pixel point according to the motion component of the second pixel point and the depth difference; generate the initial motion vector of the (i + 1)-th frame of image according to the motion components of the respective third pixel points in the (i + 1)-th frame of image.
[0256] In an implementation manner of the embodiment of the present application, two adjacent frames of images include the (i - 1)-th frame of image and the i-th frame of image, and the next frame of image includes the (i + 1)-th frame of image, where i is a positive integer greater than 1; the generation module 820 is configured to: perform transformation processing on the first albedo information of two adjacent frames of images based on the third motion vectors corresponding to the two adjacent frames of images to obtain the second albedo information corresponding to the two adjacent frames of images; perform transformation processing on the first depth information of two adjacent frames of images based on the third motion vectors corresponding to the two adjacent frames of images to obtain the second depth information corresponding to the two adjacent frames of images; predict the second motion vector of the (i + 1)-th frame of image according to the second albedo information and the second depth information corresponding to the two adjacent frames of images, and according to the first motion vector and the initial motion vector of the i-th frame of image.
[0257] In an implementation manner of the embodiment of the present application, the generation module 820 is configured to: obtain a mask weight; where the mask weight is predicted according to the second albedo information and the second depth information corresponding to two adjacent frames of images, and according to the first motion vector and the initial motion vector of the i-th frame of image; perform weighting on the two transformed frames of images based on the mask weight to obtain the (i + 1)-th frame of image.
[0258] In an implementation manner of the embodiment of the present application, two adjacent frames of pictures include the (i - 1)-th frame of picture and the i-th frame of picture, and the next picture includes the (i + 1)-th frame of picture, where i is a positive integer greater than 1; a generating module 820 is configured to: obtain the skeletal point information of the character in the (i + 1)-th frame of picture; wherein, the skeletal point information of the character in the (i + 1)-th frame of picture is predicted based on the skeletal point information of the character in two adjacent frames of pictures; determine a fourth motion vector of the skeletal points of the character in the (i + 1)-th frame of picture based on the skeletal point information of the (i + 1)-th frame of picture and the skeletal point information of the i-th frame of picture; adjust the initial motion vector of the (i + 1)-th frame of picture based on the fourth motion vector; and generate a third motion vector pointing from the two adjacent frames of pictures to the (i + 1)-th frame of picture according to the adjusted initial motion vector and according to the first motion vector of the two adjacent frames of pictures.
[0259] In an implementation manner of the embodiment of the present application, two adjacent frames of pictures include the (i - 1)-th frame of picture and the i-th frame of picture, and the next picture includes the (i + 1)-th frame of picture, where i is a positive integer greater than 1; a generating module 820 is configured to: generate a fifth motion vector pointing from the (i - 1)-th frame of picture to the (i + 1)-th frame of picture according to the second motion vector and the first motion vector of the (i - 1)-th frame of picture; transform the (i - 1)-th frame of picture based on the fifth motion vector to obtain a transformed (i - 1)-th frame of picture; generate a sixth motion vector pointing from the i-th frame of picture to the (i + 1)-th frame of picture according to the second motion vector and the first motion vector of the i-th frame of picture; and transform the i-th frame of picture based on the sixth motion vector to obtain a transformed i-th frame of picture.
[0260] It should be noted that the foregoing explanations of the embodiments of the image processing method also apply to the image processing apparatus of this embodiment, and will not be elaborated here.
[0261] In the image processing apparatus of the embodiment of the present application, video prediction is performed by integrating two adjacent frames of pictures in a video frame sequence and the rendering information corresponding to the two adjacent frames of pictures to generate the next picture after the two adjacent frames of pictures. This can not only improve the image quality of the predicted next picture, but also, since video prediction only needs to use the relevant information of two adjacent frames of pictures, it can reduce the demand for complex calculations, thereby reducing the power consumption of the device, extending the battery life of the device, and enhancing the user experience.
[0262] To implement the above embodiments, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the image processing method described in any of the foregoing embodiments is implemented.
[0263] Figure 9A schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 900 may be a vehicle, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0264] Referring to Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.
[0265] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.
[0266] The memory 904 is configured to store various types of data to support the operation of the electronic device 900. Examples of these data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0267] The power component 906 provides power to various components of the electronic device 900. The power component 906 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 900.
[0268] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0269] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.
[0270] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0271] The sensor assembly 914 includes one or more sensors for providing a status assessment of various aspects of the electronic device 900. For example, the sensor assembly 914 can detect the on / off state of the electronic device 900, the relative positioning of components, such as the display and keypad of the electronic device 900. The sensor assembly 914 can also detect a change in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and a change in the temperature of the electronic device 900. The sensor assembly 914 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 can also include a light sensor, such as a Complementary Metal-Oxide-Semiconductor (CMOS) or Charge-Coupled Device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0272] The communication component 916 is configured to facilitate communication between the electronic device 900 and other devices in a wired or wireless manner. The electronic device 900 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0273] In an exemplary embodiment, the electronic device 900 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above method.
[0274] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions. The above instructions can be executed by a processor 920 of the electronic device 900 to complete the above method. For example, the non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0275] To implement the above embodiments, the present application also proposes a chip. The chip includes an interface circuit and a processing circuit that are coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to execute the image processing method provided in any of the foregoing embodiments.
[0276] Figure 10 It is a schematic structural diagram of the second chip proposed in the embodiments of the present application. Reference can be made to Figure 10 the schematic structural diagram of the chip 1000 shown, but not limited thereto.
[0277] The chip 1000 includes a processing circuit 1001, and the processing circuit 1001 is configured to execute any of the above image processing methods.
[0278] In some embodiments, the chip 1000 further includes one or more interface circuits 1002. Optionally, the interface circuit 1002 is connected to the memory 1003. The interface circuit 1002 can be used to receive signals from the memory 1003 or other devices, and the interface circuit 1002 can be used to send signals to the memory 1003 or other devices. For example, the interface circuit 1002 can read the instructions stored in the memory 1003 and send the instructions to the processing circuit 1001.
[0279] In some embodiments, the interface circuit 1002 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1001 performs other steps.
[0280] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be replaced with each other.
[0281] In some embodiments, the chip 1000 further includes one or more memories 1003 for storing instructions. Optionally, all or part of the memory 1003 can be outside the chip 1000.
[0282] To implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method as described in any of the foregoing method embodiments.
[0283] To implement the above embodiments, the present application also proposes a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it implements the image processing method as described in any of the foregoing method embodiments.
[0284] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0285] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0286] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0287] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered a sequenced list of executable instructions for implementing a logical function and may be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (Random Access Memory, abbreviated as RAM), a read-only memory (Read-Only Memory, abbreviated as ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (Compact Disc Read-Only Memory, abbreviated as CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0288] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0289] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0290] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0291] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An image processing method, characterized in that, including: obtaining two adjacent frames in a video frame sequence and rendering information corresponding to the frames; generating a next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames.
2. The method according to claim 1, wherein The two adjacent frames in the video frame sequence include: two adjacent video frames in a video stream sent from a client to a server; The method further includes: performing bandwidth allocation according to the next frame, and encoding the video stream according to the allocated bandwidth and uploading it to the server.
3. The method according to claim 1, wherein The two adjacent frames in the video frame sequence include captured images obtained by an in-vehicle camera of a vehicle in two adjacent captures; The method further includes at least one of the following: performing motion prediction on an obstacle in the environment where the vehicle is located based on the next frame to obtain a first prediction result; wherein, the first prediction result is used for controlling the driving of the vehicle; predicting position information of an obstacle in the environment where the vehicle is located and / or a lane-changing behavior of the obstacle based on the next frame to obtain a second prediction result; wherein, the second prediction result is used for path planning of the vehicle.
4. The method according to claim 1, characterized in that, The two adjacent frames in the video frame sequence include game screens obtained by rendering in a game in two adjacent times; The method further includes: performing post-image processing based on the next frame and rendering for display.
5. The method according to any one of claims 1-4, characterized in that, The rendering information includes at least one of the following: a first motion vector corresponding to a frame; wherein, in response to the corresponding frame being a non-first frame in the video frame sequence, the first motion vector is used to indicate motion information of the corresponding frame distorted towards the previous frame, and in response to the corresponding frame being the first frame in the video frame sequence, the first motion vector is a set vector; a first depth information corresponding to a frame; a first albedo information corresponding to a frame; skeleton point information of a character in a corresponding frame; wherein, the skeleton point information is used to indicate positions of respective skeleton points of the character.
6. The method according to claim 5, wherein The generating a next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames includes: predicting a second motion vector of the next frame according to the first motion vector and the first depth information corresponding to the two adjacent frames; transforming the two adjacent frames based on the second motion vector and the first motion vector corresponding to the two adjacent frames to obtain two transformed frames; fusing the two transformed frames to obtain the next frame.
7. The method according to claim 6, characterized in that, The predicting a second motion vector of the next frame according to the first motion vector and the first depth information corresponding to the two adjacent frames includes: generating an initial motion vector of the next frame according to the first motion vector and the first depth information corresponding to the two adjacent frames; generating a third motion vector pointing from the two adjacent frames to the next frame according to the first motion vector corresponding to the two adjacent frames and the initial motion vector of the next frame; predicting the second motion vector of the next frame based on the third motion vector corresponding to the two adjacent frames and the first depth information.
8. The method according to claim 7, characterized in that, The two adjacent frames include the (i - 1)-th frame and the i-th frame, and the next frame includes the (i + 1)-th frame, where i is a positive integer greater than 1; generating the initial motion vector of the next frame according to the first motion vector and the first depth information corresponding to the two adjacent frames includes: Obtaining the motion difference of a pixel pair according to the first motion vector corresponding to the two adjacent frames; wherein, the pixel pair includes any first pixel point of the (i - 1)-th frame and a second pixel point corresponding to the first pixel point in the i-th frame; In response to the motion difference being less than the difference threshold, using the motion component of the second pixel point in the first motion information of the i-th frame as the motion component of the third pixel point corresponding to the second pixel point in the (i + 1)-th frame; or, In response to the motion difference being greater than or equal to the difference threshold, determining the depth difference of the pixel pair according to the first depth information corresponding to the two adjacent frames, and determining the motion component of the third pixel point according to the motion component of the second pixel point and the depth difference; Generating the initial motion vector of the (i + 1)-th frame according to the motion components of the third pixel points in the (i + 1)-th frame.
9. The method according to claim 7, wherein The two adjacent frames include the (i - 1)-th frame and the i-th frame, and the next frame includes the (i + 1)-th frame, where i is a positive integer greater than 1; predicting the second motion vector of the next frame based on the third motion vector and the first depth information corresponding to the two adjacent frames includes: Performing a transformation process on the first albedo information of the two adjacent frames based on the third motion vector corresponding to the two adjacent frames to obtain the second albedo information corresponding to the two adjacent frames; Performing a transformation process on the first depth information of the two adjacent frames based on the third motion vector corresponding to the two adjacent frames to obtain the second depth information corresponding to the two adjacent frames; Predicting the second motion vector of the (i + 1)-th frame according to the second albedo information and the second depth information corresponding to the two adjacent frames, and according to the first motion vector and the initial motion vector of the i-th frame.
10. The method according to claim 9, wherein Fusing the two transformed frames to obtain the next frame includes: Obtaining a mask weight; wherein, the mask weight is predicted according to the second albedo information and the second depth information corresponding to the two adjacent frames, and according to the first motion vector and the initial motion vector of the i-th frame; Weighting the two transformed frames based on the mask weight to obtain the (i + 1)-th frame.
11. The method according to claim 7, characterized in that, The two adjacent frames include the (i - 1)-th frame and the i-th frame, and the next frame includes the (i + 1)-th frame, where i is a positive integer greater than 1; Generating the third motion vector from the two adjacent frames to the next frame according to the first motion vector corresponding to the two adjacent frames and the initial motion vector of the next frame includes: Obtain the bone point information of the character in the (i + 1)-th frame; wherein, the bone point information of the character in the (i + 1)-th frame is predicted based on the bone point information of the character in two adjacent frames. Based on the bone point information of the (i + 1)-th frame and the bone point information of the i-th frame, determine the fourth motion vector of the bone points of the character in the (i + 1)-th frame. Based on the fourth motion vector, adjust the initial motion vector of the (i + 1)-th frame. According to the adjusted initial motion vector and the first motion vectors of the two adjacent frames, generate the third motion vector from the two adjacent frames to the (i + 1)-th frame.
12. The method according to claim 6, wherein The two adjacent frames include the (i - 1)-th frame and the i-th frame, and the next frame includes the (i + 1)-th frame, where i is a positive integer greater than 1. The transformation of the two adjacent frames based on the second motion vector and the first motion vectors corresponding to the two adjacent frames to obtain two transformed frames includes: According to the second motion vector and the first motion vector of the (i - 1)-th frame, generate the fifth motion vector from the (i - 1)-th frame to the (i + 1)-th frame. Based on the fifth motion vector, transform the (i - 1)-th frame to obtain the (i - 1)-th transformed frame. According to the second motion vector and the first motion vector of the i-th frame, generate the sixth motion vector from the i-th frame to the (i + 1)-th frame. Based on the sixth motion vector, transform the i-th frame to obtain the i-th transformed frame.
13. An image processing apparatus, characterized in that, It includes: An acquisition module, configured to acquire two adjacent frames in a video frame sequence and the rendering information corresponding to the frames. A generation module, configured to generate the next frame after the two adjacent frames according to the two adjacent frames and the rendering information corresponding to the frames.
14. A chip, characterized in that, The chip includes: A graphics processing unit (GPU), configured to acquire two adjacent frames in a video frame sequence and the rendering information corresponding to the frames, and send the rendering information corresponding to the frames to a neural network processing unit (NPU). The NPU is configured to predict the motion vector of the next frame after the two adjacent frames according to the rendering information corresponding to the frames, and send the motion vector to the GPU. The GPU is further configured to generate the next frame according to the two adjacent frames and the motion vector.
15. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A chip, characterized in that, The chip includes an interface circuit and a processing circuit that are mutually coupled. The interface circuit is used to input or output signals, and the processing circuit is used to implement the method according to any one of claims 1 to 12.
17. A non-transitory computer-readable storage medium storing computer program instructions thereon, characterized in that, When the program instruction is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
18. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.