Video processing method and device, equipment, storage medium and program product

By collecting user input operations to generate event description sequences and predict motion information, the problem of insufficient quality and efficiency in video frame interpolation under rapid rotation or sudden movements is solved, achieving higher quality and more efficient video frame interpolation generation.

CN121397174APending Publication Date: 2026-01-23VASTAI TECH (SHANGHAI) INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511934862.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing video frame interpolation technology is insufficient in terms of quality and efficiency in situations of rapid rotation or sudden movements, resulting in blurred vision or ghosting for users.

Method used

The system collects user input, generates an event description sequence, and uses this sequence and the current video frame to generate predicted motion information for interpolation to generate the next video frame.

Benefits of technology

It improves the quality and efficiency of video frame interpolation, ensuring smooth video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397174A_ABST
    Figure CN121397174A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a video processing method and device, equipment, a storage medium and a program product. The method provided by the invention comprises the following steps: collecting an input operation of a user in a video playing process; generating an event description sequence of the input operation based on the operation information of the input operation; generating predicted motion information for a first video frame of the video based on the event description sequence and the first video frame, the predicted motion information indicating inter-frame motion of pixels in the first video frame; generating a second video frame of the video at least based on the first video frame and the predicted motion information; and playing the second video frame of the video. In this way, according to the embodiment of the invention, the generation quality and the generation efficiency of the video frame insertion can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device, storage medium and program product for video processing. BACKGROUND

[0002] In the field of video processing, video fluency is an important factor to guarantee user experience and platform competitiveness. In some scenarios, people will use video frame interpolation technology to ensure video fluency. Therefore, the generation quality and generation efficiency of video frame interpolation become the focus of attention. SUMMARY

[0003] In a first aspect of the present disclosure, a method for video processing is provided. The method comprises: collecting an input operation of a user in a process of playing a video; generating an event description sequence of the input operation based on operation information of the input operation, the event description sequence comprising event description information of at least one event corresponding to the input operation, the event description information comprising time information, type information and operation parameters of the corresponding event, wherein the at least one event comprises a plurality of events, and the order of the event description information corresponding to the plurality of events in the event description sequence is determined based on the time order of the plurality of events; generating prediction motion information for a first video frame of the video based on the event description sequence and the first video frame, the prediction motion information indicating inter-frame motion of pixels in the first video frame; generating a second video frame of the video based on at least the first video frame and the prediction motion information; and playing the second video frame of the video.

[0004] In a second aspect of the present disclosure, an apparatus for video processing is provided. The apparatus comprises: an operation collection module configured to collect an input operation of a user in a process of playing a video; a first generation module configured to generate an event description sequence of the input operation based on operation information of the input operation, the event description sequence comprising event description information of at least one event corresponding to the input operation, the event description information comprising time information, type information and operation parameters of the corresponding event, wherein the at least one event comprises a plurality of events, and the order of the event description information corresponding to the plurality of events in the event description sequence is determined based on the time order of the plurality of events; a second generation module configured to generate prediction motion information for a first video frame of the video based on the event description sequence and the first video frame, the prediction motion information indicating inter-frame motion of pixels in the first video frame; a third generation module configured to generate a second video frame of the video based on at least the first video frame and the prediction motion information; and a content playing module configured to play the second video frame of the video.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] In this way, embodiments of the present disclosure are able to collect input operations of a user in a video playing process, and generate predicted motion information of a current video frame (e.g., a first video frame) based on the input operations and the current video frame, so that the generation quality and generation efficiency of a video frame (e.g., a second video frame) can be improved.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above-described and other features and advantages of various embodiments of the present disclosure will become more apparent by reference to the following detailed description and accompanying drawings. In the drawings, like reference numerals indicate like elements, and: Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown; Figure 2 A flowchart showing an example process of video processing according to some embodiments of the present disclosure is shown; Figure 3 A schematic diagram showing an example preset time period according to some embodiments of the present disclosure is shown; Figure 4 A schematic structural block diagram of an example apparatus for video processing according to some embodiments of the present disclosure is shown; and Figure 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0011] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0012] It should be noted that the titles of any sections / sub-sections provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / sub-section. Furthermore, embodiments described in any section / sub-section can be combined with any other embodiments described in the same section / sub-section and / or different section / sub-section in any manner.

[0013] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicitly and implicitly recited definitions can also be found below. The terms "first," "second," and the like can refer to different or identical objects. Other explicit and implicit definitions can also be found below.

[0014] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the present disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all collection, acquisition, processing, processing, forwarding, use, etc. of data are performed on the premise that users are aware of and confirm. Accordingly, when implementing embodiments of the present disclosure, the type of data or information that can be involved, the range of use, the scenario of use, etc. should be informed to users and authorized by users in a proper manner according to relevant laws and regulations. The specific informing and / or authorization manner can vary according to actual situations and application scenarios, and the scope of the present disclosure is not limited in this respect.

[0015] In the present specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. Users refuse to process personal information other than the necessary information required for basic functions, which will not affect the user's use of basic functions.

[0016] Traditional methods rely on previous and subsequent video frames or optical flow information to generate interpolated frames. This requires waiting for the next video frame to arrive before motion information can be calculated, thus affecting the efficiency of interpolation. Furthermore, rapid rotation of objects or camera angles within the video, or other sudden movements, can affect the quality of the interpolated frames, resulting in blurred vision or ghosting for the user.

[0017] This disclosure proposes a video processing scheme. The scheme includes: during video playback, acquiring user input operations; generating an event description sequence based on operation information of the input operations, the event description sequence including event description information of at least one event corresponding to the input operation, the event description information including time information, type information, and operation parameters of the corresponding event, wherein at least one event includes multiple events, and the order of the event description information corresponding to the multiple events in the event description sequence is determined based on the time order of the multiple events; generating predicted motion information for the first video frame based on the event description sequence and a first video frame of the video, the predicted motion information indicating inter-frame motion of pixels in the first video frame; generating a second video frame of the video based at least on the first video frame and the predicted motion information; and playing the second video frame of the video.

[0018] In this way, embodiments of the present disclosure can collect user input operations during video playback and generate predicted motion information for the current video frame based on the input operations and the current video frame (e.g., the first video frame), thereby improving the generation quality and efficiency of video interpolation (e.g., the second video frame).

[0019] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0020] Example Environment Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include terminal device 110.

[0021] In this example environment 100, terminal device 110 may run an application 120 that supports user interface interaction. Application 120 may be any suitable type of application for user interface interaction, and examples may include, but are not limited to, game applications, video applications, live streaming applications, social applications, or other suitable applications. User 140 may interact with application 120 via terminal device 110 and / or its attached devices.

[0022] exist Figure 1If the application 120 is active, the terminal device 110 can present, through the application 120, an interface 150 for supporting interface interaction in the environment 100.

[0023] In some embodiments, the terminal device 110 communicates with the server 130 to implement the provisioning of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable game console, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).

[0024] The server 130 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and big data and artificial intelligence platform. The server 130 may, for example, include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 can provide background services for the application 120 in the terminal device 110 that supports content presentation.

[0025] A communication connection can be established between the server 130 and the terminal device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the terminal device 110 can implement signaling interaction through the communication connection therebetween.

[0026] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0027] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0028] Example process Figure 2 A flowchart of an example video processing procedure 200 according to some embodiments of the present disclosure is shown. Procedure 200 can be implemented at terminal device 110. Reference will be made below. Figure 1 Describe the process 200.

[0029] like Figure 2 As shown in box 210, the terminal device 110 collects user input operations during video playback.

[0030] In some embodiments, the video played by the terminal device 110 may be associated with a virtual scene, such virtual scene including but not limited to: game virtual scene (e.g., single-player game, cloud game), live streaming virtual scene, and simulation virtual scene. For ease of description, the following description uses a game virtual scene as an example to illustrate this disclosure.

[0031] In some embodiments, user input may include physical operations performed by the user on a physical device within a preset time period.

[0032] In some embodiments, during video playback, adjacent video frames on the terminal device 110 have a certain time interval on the playback timeline. Figure 3 For example, Figure 3 The diagram illustrates an example preset time period according to some embodiments of the present disclosure. Assuming that the current video frame 301 of the video is at the current time, for example, time T0, and the next video frame 302 of the current video frame is at the next time after the current time, for example, time T1, then the terminal device 110 can collect the user's physical operations on the physical device between time T0 and time T1.

[0033] In some embodiments, such a physical device can be used to control a virtual scene. As an example, the physical device can be used to control the adjustment of the camera view in the virtual scene, control the movement, jumping, aiming, shooting, throwing objects, or triggering new skills of a virtual character in the virtual scene, etc. Exemplarily, such a physical device may include keyboard keys, a mouse, a touchpad, a game controller, and a display screen, etc. It is understood that this disclosure is not intended to limit the specific type of physical device.

[0034] In some embodiments, a user's physical operations on a physical device may include, but are not limited to, moving a mouse, clicking keys on a keyboard, and sliding a touchpad. It is understood that this disclosure is not intended to limit the specific physical operations performed.

[0035] As an example, the terminal device 110 can collect, in real time, a physical operation of the user on the physical device in a preset time period as an input operation of the user in the process of rendering the game virtual scene.

[0036] Returning to Figure 2 At block 220, the terminal device 110 generates an event description sequence of the input operation based on the operation information of the input operation, the event description sequence including event description information of at least one event corresponding to the input operation, the event description information including time information, type information, and operation parameters of the corresponding event.

[0037] In some embodiments, the terminal device 110 can determine a plurality of candidate events corresponding to the input operation based on the operation information of the input operation.

[0038] Taking the game virtual scene as an example, the terminal device 110 can collect the input operation of the user in the game virtual scene, and further, the terminal device 110 can determine a plurality of candidate events corresponding to the input operation based on the operation information of the input operation, for example, candidate event 1 (the user moves the mouse to the right, the virtual camera view in the virtual scene is adjusted by ×× degrees), candidate event 2 (the user clicks the keyboard “W key”, the virtual character in the virtual scene moves forward), candidate event 3 (the user clicks the right mouse button, the virtual character in the virtual scene opens the camera), and candidate event 4 (the user clicks the keyboard “space bar”, the virtual character in the virtual scene jumps), and the like.

[0039] Additionally or alternatively, the terminal device 110 can also determine type information, time information, operation parameters (such as intensity information and duration) and the like corresponding to the plurality of candidate events based on the operation information of the input operation. Specifically, the type information can indicate the type corresponding to the candidate event, the time information can indicate the time when the candidate event occurs, for example, the operation parameters can indicate the intensity parameter of the event, the duration of the event, and the like.

[0040] In some embodiments, the terminal device 110 can determine at least one event satisfying a preset condition from the plurality of candidate events based on the type of the plurality of candidate events.

[0041] Specifically, the terminal device 110 can be preconfigured with an event type table, which can indicate, for example, the event types satisfying the preset condition, such as the event type of controlling the adjustment of the virtual camera view in the virtual scene, the event type of controlling the movement of the virtual character in the virtual scene, and the like. Additionally or alternatively, the event type table can also include, for example, the event types of controlling the virtual character in the virtual scene to jump, aim, shoot, throw an object, or trigger a new skill.

[0042] It can be understood that the event types in the above event type table are only exemplary descriptions, and the disclosure is not intended to limit the specific event types in the event type table. For ease of description, the following is introduced taking the event types in the event type table including the event type of controlling the adjustment of the virtual camera view angle in the virtual scene and the event type of controlling the movement of the virtual character in the virtual scene as examples.

[0043] In some embodiments, the terminal device 110 can determine at least one event satisfying the preset condition from the plurality of candidate events based on the event type table. Specifically, the terminal device 110 can determine the at least one event from the plurality of candidate events based on whether the type of the plurality of candidate events matches the event type in the event type table.

[0044] Taking the above-mentioned candidate event 1 to candidate event 4 as an example, the terminal device 110 can determine the candidate event 1 as the event 1 satisfying the preset condition in response to the type of the candidate event 1 in the plurality of candidate events matching the event type in the event type table; the terminal device 110 can determine the candidate event 2 as the event 2 satisfying the preset condition in response to the type of the candidate event 2 in the plurality of candidate events matching the event type in the event type table; the terminal device 110 can determine the candidate event 3 as an event not satisfying the preset condition in response to the type of the candidate event 3 in the plurality of candidate events not matching the event type in the event type table; the terminal device 110 can determine the candidate event 4 as an event not satisfying the preset condition in response to the type of the candidate event 4 in the plurality of candidate events not matching the event type in the event type table, and the like.

[0045] In this way, the terminal device 110 can determine at least one event (e.g., the event 1 and the event 2) satisfying the preset condition from the plurality of candidate events based on the type of the plurality of candidate events.

[0046] Further, the terminal device 110 can generate an event description sequence corresponding to the at least one event based on the operation information. Specifically, the terminal device 110 can construct the event description sequence based on the time order of the plurality of events in response to the at least one event including the plurality of events, wherein the order of the event description information corresponding to the plurality of events in the event description sequence is determined based on the time order.

[0047] As an example, the terminal device 110 can construct the event description sequence based on the time order in which the plurality of events occur in response to the at least one event satisfying the preset condition from the plurality of candidate events including the plurality of events (e.g., the event 1 and the event 2). For example, the terminal device 110 can place the event description information corresponding to the event 1 before the event description information corresponding to the event 2 in response to the time at which the event 1 occurs being earlier than the time at which the event 2 occurs.

[0048] In some embodiments, the event description sequence can further comprise one event and its corresponding event description information.

[0049] In this way, the terminal device 110 can obtain the event description sequence corresponding to the at least one event satisfying the preset condition.

[0050] At block 230, the terminal device 110 generates prediction motion information for the first video frame based on the event description sequence and the first video frame, the prediction motion information indicating inter-frame motion of pixels in the first video frame.

[0051] In some embodiments, the terminal device 110 can determine a first feature representation of the event description sequence and a second feature representation of the first video frame. Specifically, the terminal device 110 can encode the event description sequence by using an encoding technique and / or an encoding tool to obtain the corresponding first feature representation. As an example, the terminal device 110 can encode the event type in the event description sequence based on an encoding technique such as one-hot or embedding representation, and in addition, the terminal device 110 can encode the time information, operation parameters in the event description sequence based on a sine-cosine encoder. In some embodiments, if the event description sequence comprises multiple events, such first feature representation can be represented as, the feature representation corresponding to event 1, the feature representation corresponding to event 2, … event 1 and event 2 can be sorted based on the time sequence of event occurrence. It can be understood that the above-mentioned encoding technique and / or encoding tool are only exemplary descriptions, and the disclosure does not intend to limit the specific encoding manner of the event description sequence.

[0052] In some embodiments, the terminal device 110 can encode the first video frame of the video by using a video frame encoding tool to obtain the corresponding second feature representation. In some embodiments, the first video frame of the video can be, for example, the current video frame 301 as shown in the following figure. Figure 3 It can be understood that the video frame encoding tool can be determined by those skilled in the art according to the needs, and the disclosure does not limit this.

[0053] In this way, the terminal device 110 can determine the first feature representation of the event description sequence and the second feature representation of the first video frame.

[0054] In some embodiments, the prediction motion information can indicate the inter-frame motion of pixels in the first video frame, for example, the motion vector of the pixels. As an example, assuming that the prediction motion information is (0, +1), the prediction motion information can indicate that the pixel A in the first video frame will move one unit distance to the right, for example, the prediction motion information can indicate that the pixel A in the first video frame moves from the position (0, 0) to the position (0, 1).

[0055] In some embodiments, the terminal device 110 or the server 130 can pre-train a motion prediction model, which can be trained based on the following procedure, for example. First, a large number of event description sequences of users and corresponding screen recording data can be collected. When constructing the training sample, the event description sequence, the current video frame and at least one historical video frame are jointly used as the input of the model. The supervision label of the model can be set as the real motion information (e.g., the real optical flow field) calculated from the continuous video frames, for example.

[0056] The model can employ an end-to-end multi-modal joint training paradigm, for example. The encoder is used to extract the feature representation of the event description sequence, the current video frame and at least one historical video frame, respectively. Then, the fusion module (e.g., based on the attention mechanism) is used to learn the dynamic association among the three. Finally, the decoder is used to output the predicted motion information (e.g., the predicted motion field). The training target is to minimize the error between the predicted motion information and the real motion information through the loss function, and finally obtain the pre-trained model. It can be understood that the above model training process is only exemplary, and the present disclosure does not intend to limit the training process of the model.

[0057] Further, the terminal device 110 can provide the pre-trained model with the first feature representation and the second feature representation to generate the predicted motion information for the first video frame. Specifically, the terminal device 110 can determine at least one historical video frame of the first video frame. Further, the terminal device 110 can provide the model with the first feature representation, the second feature representation and a third feature representation of the at least one historical video frame to generate the predicted motion information.

[0058] In some embodiments, the at least one historical video frame of the first video frame can be a preceding video frame of the first video frame, for example. Such at least one historical video frame can be one or more, and such at least one historical video frame can include, for example, the last video frame adjacent to the first video frame in the video. It can be understood that the number of at least one historical video frame can be determined by those skilled in the art according to the needs, and the present disclosure does not limit it.

[0059] In some embodiments, the model processing the above three feature representations can be a neural network model or any appropriate type of generative model, for example. It can be understood that the present disclosure does not intend to limit the specific type of the model.

[0060] In this way, the embodiments of the present disclosure can generate the predicted motion information of the first video frame based on the event description sequence, the first video frame and the feature representation corresponding to at least one historical video frame of the first video frame, respectively.

[0061] In some embodiments, the terminal device 110 can obtain pixel motion information of the first video frame. In some embodiments, such pixel motion information may, for example, include motion vector information or optical flow information. As an example, the terminal device 110 may, for example, obtain at least one historical video frame of the first video frame, and further, the terminal device 110 may, based on the at least one historical video frame and the first video frame, determine the pixel motion information of the first video frame by means of optical flow extrapolation, linear interpolation or non-linear interpolation, etc.

[0062] In some embodiments, the terminal device 110 can generate a correction signal for the pixel motion information based on the event description sequence. Specifically, the terminal device 110 may, based on at least one event (e.g., event 1 "user moves mouse to the right" and event 2 "user clicks keyboard 'W key'") in the event description sequence, generate a correction signal for the pixel motion information.

[0063] Taking event 1 as an example, the terminal device 110 may, for example, generate a correction signal 1 of "adjusting the virtual camera perspective in the virtual scene by ×× degrees"; taking event 2 as an example, the terminal device 110 may, for example, generate a correction signal 2 of "moving the virtual character in the virtual scene forward".

[0064] Further, the terminal device 110 may, based on the correction signal (e.g., the correction signal 1, the correction signal 2), adjust the pixel motion information to generate predicted motion information for the first video frame. As an example, the terminal device 110 may, based on the correction signal, adjust the direction information and the amplitude information of the pixel motion information of the first video frame to generate the predicted motion information for the first video frame.

[0065] In this way, the embodiments of the present disclosure can generate a correction signal for the pixel motion information based on the event description sequence, and further, adjust the pixel motion information based on the correction signal to generate predicted motion information for the first video frame.

[0066] In some embodiments, the video may, for example, be associated with a virtual scene, such as a game virtual scene. In some embodiments, the terminal device 110 may, based on the event description sequence, generate a control parameter for a virtual camera of the virtual scene. Specifically, in response to the event description sequence including an event related to controlling the virtual camera in the virtual scene, the terminal device 110 may, based on the event, generate a corresponding control parameter. As an example, in response to the event description sequence including event 1 "user moves mouse to the right", the terminal device 110 may, for example, generate a control parameter for the virtual camera of the virtual scene, such as adjusting the angle of the virtual camera to the right by ×× degrees.

[0067] Additionally or alternatively, the terminal device 110 can generate the corresponding control parameter based on the event in response to the event description sequence including an event related to the virtual role in the control virtual scene. As an example, the terminal device 110 can generate the control parameter for the virtual role of the virtual scene, e.g., control the forward movement of the virtual role, in response to the event description sequence including the event 2 “user clicks the keyboard ‘W key’”.

[0068] In some embodiments, the terminal device 110 can determine the initial motion information for the first video frame based on the control parameter. Further, the terminal device 110 can adjust the initial motion information based on the pixel motion information of the first video frame to generate the predicted motion information, the pixel motion information including the motion vector information or the optical flow information. As an example, the terminal device 110 can adjust the local pixel information of the initial motion information based on the pixel motion information of the first video frame to generate the predicted motion information for the first video frame.

[0069] In this way, the embodiments of the present disclosure can generate the control parameter for the virtual camera of the virtual scene based on the event description sequence, and determine the initial motion information for the first video frame based on the control parameter, further, can adjust the initial motion information based on the pixel motion information of the first video frame to generate the predicted motion information.

[0070] Returning to Figure 2 In block 240, the terminal device 110 generates a second video frame of the video based on at least the first video frame and the predicted motion information. In some embodiments, the first video frame can be, for example, a previous video frame of the second video frame, to Figure 3 For example, such a second video frame can be, for example, the next video frame 302 of the current video frame 301.

[0071] In some embodiments, the terminal device 110 can perform an interpolation operation based on the pixel information of the first video frame and the predicted motion information to generate the second video frame of the video. As an example, the terminal device 110 can determine the motion direction and displacement amount of each pixel or pixel block in the first video frame at the next time according to the predicted motion information. Subsequently, based on the pixel information (such as color, brightness value) of the first video frame, the pixel is displaced and interpolated along the motion vector direction, thereby generating the second video frame of the video.

[0072] Specifically, the terminal device 110 can implement the above process by, for example, motion-compensated interpolation operation: first, the terminal device 110 can project the pixels of the first video frame forward to the target time position according to the motion vector field, forming an initial morphed frame (or motion-compensated frame); second, for the problem of multiple pixels mapping to the same position (collision) or un-mapped pixels (holes), pixel fusion and hole filling can be performed by weighted average, median filtering or reliability-based fusion strategy. In addition, for the motion occlusion area, inference can be made in combination with motion boundary information and spatial interpolation. Based on this process, the terminal device 110 can generate a visually continuous and high-quality second video frame by performing smoothing filtering and post-processing on the reconstructed pixel grid. It can be understood that the above implementation process is only an example, and the present disclosure is not intended to limit the process.

[0073] In some embodiments, the terminal device 110 or the server 130 can pre-train an image generation model. The image generation model can be trained based on, for example, the following process: first, training data can be obtained, which can include a plurality of continuous video frame sequences, each video frame sequence can include two consecutive video frames (e.g., a reference frame and a target frame to be predicted) and the real motion information between them; second, noise can be added to the target frame step by step to generate a series of noisy video frames; then, the pixel information of the reference frame, the real motion information and the current noise step number are input as conditions, and the noisy video frame is input as a target, which are jointly input into the neural network model to be trained; finally, the neural network model is trained to learn to predict the added noise on the target frame, and to minimize the difference between the predicted noise and the real added noise. In this way, by using the conditional injection mechanism in the training process, the features of the reference frame and the features of the real motion information can be fused into the intermediate representation of the model, so that the model establishes a mapping relationship from the combination of “reference frame-motion information” to “target frame”, and finally obtains the pre-trained image generation model. It can be understood that the above training process is only an example, and the present disclosure is not intended to limit the specific training process of the image generation model.

[0074] Further, the terminal device 110 can provide the pixel information of the first video frame and the predicted motion information to this pre-trained image generation model to generate the second video frame. As an example, the terminal device 110 can provide the pixel information of the first video frame and the predicted motion information as a condition input to this pre-trained image generation model. The pre-trained image generation model can iteratively denoise the initial noise signal based on the mapping relationship established in the training stage to generate a second video frame that is continuous with the first video frame and conforms to the motion trajectory indicated by the predicted motion information.

[0075] In some embodiments, the terminal device 110 can generate the second video frame of the video based on the pixel information of the first video frame and the predicted motion information.

[0076] It can be understood that the above-mentioned manner of generating the second video frame of the video is only exemplary, and the terminal device 110 may, for example, also generate the second video frame of the video by using any appropriate combination of the above-mentioned manners, and the present disclosure is not intended to limit the specific method of generating the second video frame.

[0077] Returning to Figure 2 At block 250, the terminal device 110 plays the second video frame of the video.

[0078] In order to realize the video picture synchronized with the input operation of the user, in some embodiments, the terminal device 110 can directly play the second video frame of the video in response to the generation of the second video frame. As an example, the terminal device 110 can directly present the game picture corresponding to the second video frame in response to the generation of the second video frame in the game virtual scene.

[0079] In this way, the embodiments of the present disclosure can collect the input operation of the user in the video playing process, and generate the predicted motion information of the current video frame (e.g., the first video frame) based on the input operation and the current video frame, so that the generation quality and the generation efficiency of the video interpolation (e.g., the second video frame) can be improved.

[0080] Example apparatus and device Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above-mentioned method or process. Figure 4 An exemplary structural block diagram of an example apparatus 400 for video processing according to certain embodiments of the present disclosure is shown. The apparatus 400 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0081] As Figure 4As shown, the apparatus 400 includes an operation collecting module 410 configured to collect an input operation of a user in a process of playing a video; a first generating module 420 configured to generate an event description sequence of the input operation based on operation information of the input operation, the event description sequence including event description information of at least one event corresponding to the input operation, the event description information including time information, type information and operation parameters of the corresponding event, wherein the at least one event includes a plurality of events, and an order of the event description information corresponding to the plurality of events in the event description sequence is determined based on a time order of the plurality of events; a second generating module 430 configured to generate predicted motion information for a first video frame of the video based on the event description sequence and the first video frame, the predicted motion information indicating inter-frame motion of pixels in the first video frame; a third generating module 440 configured to generate a second video frame of the video based on at least the first video frame and the predicted motion information; and a content playing module 450 configured to play the second video frame of the video.

[0082] In some embodiments, the first generating module 420 is further configured to determine a plurality of candidate events corresponding to the input operation based on the operation information of the input operation; and determine at least one event satisfying a preset condition from the plurality of candidate events based on types of the plurality of candidate events; and generate the event description sequence corresponding to the at least one event based on the operation information.

[0083] In some embodiments, the second generating module 430 is further configured to determine a first feature representation of the event description sequence and a second feature representation of the first video frame; and provide the first feature representation and the second feature representation to a model to generate the predicted motion information for the first video frame.

[0084] In some embodiments, the second generating module 430 is further configured to determine at least one historical video frame of the first video frame; provide the first feature representation, the second feature representation and a third feature representation of the at least one historical video frame to the model to generate the predicted motion information.

[0085] In some embodiments, the second generating module 430 is further configured to obtain pixel motion information of the first video frame, the pixel motion information including motion vector information or optical flow information; generate a correction signal for the pixel motion information based on the event description sequence; and adjust the pixel motion information based on the correction signal to generate the predicted motion information for the first video frame.

[0086] In some embodiments, the video is associated with a virtual scene, and the second generation module 430 is further configured to generate control parameters for a virtual camera for the virtual scene based on an event description sequence; determine initial motion information for a first video frame based on the control parameters; and adjust the initial motion information based on the pixel motion information of the first video frame to generate predicted motion information, wherein the pixel motion information includes motion vector information or optical flow information.

[0087] In some embodiments, the third generation module 440 is further configured to perform interpolation operations based on the pixel information and predicted motion information of the first video frame to generate a second video frame of the video; or to provide pixel information and predicted motion information to an image generation model to generate a second video frame.

[0088] In some embodiments, the video is associated with a virtual scene, and the input operations include physical operations performed by the user on a physical device within a preset time period, the physical device being used to control the virtual scene.

[0089] The modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the modules in device 400 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0090] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Terminal equipment 110.

[0091] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and nonvolatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically-erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable media, and can include machine- readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 500.

[0092] The electronic device 500 can further include additional removable / non-removable, volatile / nonvolatile storage media. Although not shown in the electronic device 500, a disk drive and a disk controller are typically provided to interface with a removable, non-removable media disk (e.g., a "floppy disk"), and a disk drive and a disk controller are typically provided to interface with a removable, non-removable media disk (e.g., a "floppy disk"). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure. Figure 5

[0093] The communication unit 540 enables communication with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0094] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can further communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 540, as needed, with one or more devices that enable a user to interact with the electronic device 500, or with any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0095] ​According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0096] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0097] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0098] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0099] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk drive, or any other suitable non-transitory computer readable medium can store the computer program product.

[0100] Various implementations of the disclosure have been described in detail above. The foregoing description is exemplary and explanatory only, and is not intended to be exhaustive or to limit various implementations of the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings without departing from the scope and spirit of the disclosure. It is intended that the scope of the disclosure be limited only by the claims and the equivalents thereof. The use of the terms "including," "containing," "comprising," "having," "in involving," "portions," "elements," "components," "steps," "phases," "processes," "operations," "steps," "stages," "procedures," "methods," "mechanisms," "devices," "systems," "apparatuses," "units," "means," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "

Claims

1. A method of video processing, the method comprising: The method comprises: acquiring an input operation of a user during playing of a video; generating an event description sequence of the input operation based on operation information of the input operation, the event description sequence comprising event description information of at least one event corresponding to the input operation, the event description information comprising time information, type information and operation parameters of a corresponding event, wherein the at least one event comprises a plurality of events, and an order of the event description information corresponding to the plurality of events in the event description sequence is determined based on a time order of the plurality of events; generating predicted motion information for a first video frame of the video based on the event description sequence and the first video frame, the predicted motion information indicating inter-frame motion of pixels in the first video frame; generating a second video frame of the video based at least on the first video frame and the predicted motion information; and playing the second video frame of the video.

2. The method of claim 1, wherein, The generating of the event description sequence of the input operation based on the operation information of the input operation comprises: determining a plurality of candidate events corresponding to the input operation based on the operation information of the input operation; and determining the at least one event satisfying a preset condition from the plurality of candidate events based on types of the plurality of candidate events; and generating the event description sequence corresponding to the at least one event based on the operation information.

3. The method of claim 1, wherein, The generating of the predicted motion information for the first video frame based on the event description sequence and the first video frame of the video comprises: determining a first feature representation of the event description sequence and a second feature representation of the first video frame; and providing the first feature representation and the second feature representation to a model to generate the predicted motion information for the first video frame.

4. The method of claim 3, wherein, The providing of the first feature representation and the second feature representation to a model to generate the predicted motion information for the first video frame comprises: determining at least one historical video frame of the first video frame; providing the first feature representation, the second feature representation and a third feature representation of the at least one historical video frame to the model to generate the predicted motion information.

5. The method of claim 1, wherein, The generating of the predicted motion information for the first video frame based on the event description sequence and the first video frame of the video comprises: acquiring pixel motion information of the first video frame, the pixel motion information comprising motion vector information or optical flow information; generating a correction signal for the pixel motion information based on the event description sequence; and adjusting the pixel motion information based on the correction signal to generate the predicted motion information for the first video frame.

6. The method of claim 1, wherein, The video is associated with a virtual scene, and the generating of the predicted motion information for the first video frame based on the event description sequence and the first video frame of the video comprises: generating a control parameter of a virtual camera of the virtual scene based on the event description sequence; determining initial motion information for the first video frame based on the control parameter; and adjust the initial motion information based on pixel motion information of the first video frame to generate the predicted motion information, the pixel motion information comprising motion vector information or optical flow information.

7. The method of claim 1, wherein, generating a second video frame of the video based on at least the first video frame and the predicted motion information comprises: performing an interpolation operation based on pixel information of the first video frame and the predicted motion information to generate the second video frame of the video; or providing the pixel information and the predicted motion information to an image generation model to generate the second video frame.

8. The method of claim 1, wherein, the video is associated with a virtual scene, and the input operation comprises a physical operation of a user on a physical device within a preset time period, the physical device being used to control the virtual scene.

9. An apparatus for video processing, the apparatus comprising: The apparatus comprises: an operation collection module configured to collect an input operation of a user during playing of a video; a first generation module configured to generate an event description sequence of the input operation based on operation information of the input operation, the event description sequence comprising event description information of at least one event corresponding to the input operation, the event description information comprising time information, type information and operation parameters of a corresponding event, wherein the at least one event comprises a plurality of events, and an order of the event description information corresponding to the plurality of events in the event description sequence is determined based on a time order of the plurality of events; a second generation module configured to generate predicted motion information for a first video frame of the video based on the event description sequence and the first video frame; a third generation module configured to generate a second video frame of the video based on at least the first video frame and the predicted motion information, the predicted motion information indicating inter-frame motion of pixels in the first video frame; and a content playing module configured to play the second video frame of the video. 10.An electronic device, comprising at least one processor; and at least one memory, characterized in that, The at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having stored thereon computer- executable instructions, wherein, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 8.

12. A computer program product, the computer program product being tangibly stored in a computer storage medium and comprising computer-executable instructions, the computer program product being characterized in that, The computer-executable instructions, when executed by the device, cause the device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cloud game acceleration method, device, readable storage medium and computer equipment

    CN110227260A

  • Method and apparatus for improving video quality

    CN115917585A

  • Method for allowing streaming of video content between server and electronic device, and server and electronic device for streaming video content

    CN118339838A

  • Video rendering method, video rendering device and terminal device

    CN120692416A

  • Data processing apparatus and method

    US20250170486A1