Method, apparatus and storage medium for processing video data
By acquiring the color and motion information of video frames, determining the motion vectors and color residuals of pixel blocks, and constructing encoding information, the problem of large video frame data volume is solved, and storage space is reduced while transmission efficiency is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECH SHANGHAI
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, video frames contain a large amount of data, and how to reduce the size of video frames is an urgent problem to be solved.
By acquiring color and motion information from video frames, the motion vectors and color residual information of pixel blocks are determined, and the encoding information of the video data is constructed.
While ensuring the integrity of video frames, the space required to store video data has been reduced, and storage and transmission efficiency has been improved.
Smart Images

Figure CN122120443A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia technology, and in particular to a method, apparatus, device and storage medium for processing video data. Background Technology
[0002] With the rapid development of computer technology, the demand for transmitting information via video is increasing.
[0003] In related technologies, video frames of video data can carry a rich variety of image information, conveying a wealth of meaning through elements such as color, outline, structure, and the spatial relationships of different objects in the image.
[0004] However, storing or transmitting video frames requires a large amount of data, and how to reduce the size of video frames is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for processing video data, the technical solution of which is as follows:
[0006] According to one aspect of this application, a method for processing video data is provided, the method being executed by a computer device, the method comprising:
[0007] Obtain the color information of the i-th video frame and the (i+1)-th video frame, and obtain the motion information of at least one pixel in the (i+1)-th video frame. The i-th video frame and the (i+1)-th video frame are used to display the virtual environment, where i is a positive integer.
[0008] Based on the motion information of the at least one pixel, the motion vector of at least one pixel block in the (i+1)th video frame is determined. The at least one pixel includes a second pixel. The motion information of the second pixel is used to indicate the difference between the second position of the second pixel in the (i+1)th video frame and the first position of the first pixel in the i-th video frame. The second pixel and the first pixel correspond to the same position in the virtual environment.
[0009] Based on the motion vector of the at least one pixel block, color residual information corresponding to the at least one pixel block is determined in the color information of the i-th video frame and the (i+1)-th video frame, respectively.
[0010] The encoding information of the video data is constructed based on the color information of the i-th video frame, the motion vector of the at least one pixel block, and the color residual information.
[0011] According to another aspect of this application, a video data processing apparatus is provided, the apparatus comprising:
[0012] The acquisition module is used to acquire the color information of the i-th video frame and the (i+1)-th video frame, and to acquire the motion information of at least one pixel in the (i+1)-th video frame. The i-th video frame and the (i+1)-th video frame are used to display the virtual environment, where i is a positive integer.
[0013] The processing module is configured to determine the motion vector of at least one pixel block in the (i+1)th video frame based on the motion information of the at least one pixel, wherein the at least one pixel includes a second pixel, and the motion information of the second pixel is used to indicate the difference between the second position of the second pixel in the (i+1)th video frame and the first position of the first pixel in the i-th video frame, wherein the second pixel and the first pixel correspond to the same position in the virtual environment;
[0014] The processing module is further configured to determine, based on the motion vector of the at least one pixel block, the color residual information corresponding to the at least one pixel block in the color information of the i-th video frame and the (i+1)-th video frame, respectively;
[0015] A construction module is used to construct the encoding information of the video data based on the color information of the i-th video frame, the motion vector of the at least one pixel block, and the color residual information.
[0016] In an optional design of this application, the first set of pixels is a subset of the at least one pixel, and each pixel in the first set of pixels is a pixel in the first pixel block;
[0017] The processing module is also used for:
[0018] The motion vector of the first pixel block is determined based on the motion information of the pixels in the first pixel set.
[0019] In an optional design of this application, the processing module is further configured to:
[0020] The average value of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block;
[0021] Alternatively, the median of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block;
[0022] Alternatively, the modulo value of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block.
[0023] In an optional design of this application, the processing module is further configured to:
[0024] Based on the motion information of the pixels in the first pixel set, the reference vector of the first pixel block is determined;
[0025] Based on the first region position of the first pixel block in the (i+1)th video frame and the reference vector, the color sub-information of the reference pixel block is determined in the color information of the i-th video frame;
[0026] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed a color threshold, the reference vector is determined as the motion vector of the first pixel block.
[0027] In an optional design of this application, the color threshold is the minimum of at least one reference difference degree; the reference difference degree is used to indicate the difference between the color sub-information of the neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
[0028] In an optional design of this application, the processing module is further configured to:
[0029] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds a color threshold, a correction pixel block is determined.
[0030] The motion vector of the first pixel block is determined based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th video frame.
[0031] Wherein, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
[0032] In an optional design of this application, the motion information of the second pixel includes distance sub-information in the first direction and the second direction;
[0033] Wherein, the first direction and the second direction are the directions of the two vertical sides of the i-th video frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first direction and the second direction.
[0034] In an optional design of this application, the motion information of the second pixel includes direction sub-information and distance sub-information;
[0035] The direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position.
[0036] In an optional design of this application, the processing module is further configured to:
[0037] Perform data normalization and / or coordinate axis transformation on the motion information of the at least one pixel to obtain the corrected motion information of the at least one pixel;
[0038] The motion vector of the at least one pixel block is determined based on the corrected motion information of the at least one pixel.
[0039] In an optional design of this application, the acquisition module is further configured to:
[0040] Acquire control information in the virtual environment, the control information being used to control virtual objects and / or virtual cameras in the virtual environment;
[0041] Based on the control information, the motion information of at least one pixel in the (i+1)th video frame is determined.
[0042] In an optional design of this application, the processing module is further configured to:
[0043] Based on the motion vector of the first pixel block, the color information of the associated pixel block is determined in the color information of the i-th video frame, wherein the at least one pixel block includes the first pixel block;
[0044] The color residual information of the first pixel block is determined based on the difference between the color information of the associated pixel block and the color information of the first pixel block.
[0045] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the video data processing method as described above.
[0046] According to another aspect of this application, a computer-readable storage medium is provided, wherein a computer program and a video stream are stored in the computer program, the computer program being executed by a processor to implement the above-described video encoding method to generate the video stream.
[0047] According to another aspect of this application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium, wherein a processor reads from the computer-readable storage medium and executes the computer instructions to implement the video data processing method described above.
[0048] This application also provides a method for generating a video stream, wherein the video stream is obtained by encoding video data using the above-described video data processing method.
[0049] The beneficial effects of the technical solution provided in this application include at least the following:
[0050] The motion vectors determined by motion information fully utilize the virtual spatial positions of pixels on the video frame in the virtual environment. Encoding information is constructed based on the motion vectors, color residual information, and color information of the i-th video frame. Compared to directly storing the color information of the (i+1)-th video frame, this reduces the space required to store video data. By using the motion vectors and color residual information of pixel blocks, the color information of the (i+1)-th video frame can be obtained based on the color information of the i-th video frame. While reducing the storage space of video data, the integrity of the color information of the (i+1)-th video frame is ensured, thus guaranteeing the integrity of the image of the (i+1)-th video frame. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a schematic diagram of a computer system provided in an exemplary embodiment of this application;
[0053] Figure 2 This is a schematic diagram of a video data processing method provided in an exemplary embodiment of this application;
[0054] Figure 3 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0055] Figure 4 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0056] Figure 5 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0057] Figure 6 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0058] Figure 7 This is a schematic diagram illustrating the processing of video data provided in an exemplary embodiment of this application;
[0059] Figure 8 This is a schematic diagram illustrating the processing of video data provided in an exemplary embodiment of this application;
[0060] Figure 9 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0061] Figure 10 This is a flowchart of a video data processing method provided in an exemplary embodiment of this application;
[0062] Figure 11 This is a structural block diagram of a video data processing apparatus provided in an exemplary embodiment of this application;
[0063] Figure 12 This is a structural block diagram of a server provided in an exemplary embodiment of this application.
[0064] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0066] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0067] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0068] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the video frames, motion information and other information involved in this application were obtained with full authorization.
[0069] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0070] Before describing the embodiments of this application, a brief introduction to the video encoding process in related technologies will be given first.
[0071] A video signal is a sequence of images consisting of multiple frames. A frame represents the spatial information of a video signal. Taking YUV mode as an example, a frame includes a luminance sample matrix (Y) and two chrominance sample matrices (Cb and Cr). From the perspective of how video signals are acquired, they can be divided into two types: those captured by a camera and those generated by a computer. Due to differences in statistical characteristics, the corresponding compression coding methods may also differ.
[0072] In some mainstream video coding technologies, such as H.265 / HEVC (High Efficient Video Coding), H.266 / VVC (Versatile Video Coding), and AVS (Audio Video Coding Standard) (e.g., AVS3), a hybrid coding framework is used to perform a series of operations and processes on the input raw video signal:
[0073] 1. Block Partition Structure: The input image is divided into several non-overlapping processing units, each of which undergoes a similar compression operation. This processing unit is called a CTU (Coding Tree Unit) or LCU (Large Coding Unit). Further subdivisions can be made below the CTU to obtain one or more basic coding units, called CUs (Coding Units). Each CU is the most basic element in a coding process. The following describes the various coding methods that can be used for each CU.
[0074] 2. Predictive Coding: This includes intra-frame prediction and inter-frame prediction. The original video signal is predicted from a selected reconstructed video signal to obtain a residual video signal. The encoder needs to determine the most suitable predictive coding mode from among many possible modes for the current CU and inform the decoder. Intra-frame prediction refers to predicting a signal from a region within the same image that has already been encoded and reconstructed. Inter-frame prediction refers to predicting a signal from another encoded image that is different from the current image (i.e., the current video frame).
[0075] 3. Transform Coding and Quantization: The residual video signal undergoes transform operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform) to convert the signal into the transform domain, where the coefficients are called transform coefficients. In the transform domain, the signal undergoes further lossy quantization, losing some information to make the quantized signal more suitable for compression. Some video coding standards may offer more than one transform option; therefore, the encoder needs to select one transform for the current CU and inform the decoder. The fineness of quantization is usually determined by the quantization parameter. A larger QP (Quantization Parameter) value means that coefficients with a wider range of values will be quantized into the same output, which usually results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that coefficients with a smaller range of values will be quantized into the same output, which usually results in less distortion and a higher bitrate.
[0076] 4. Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of each value, ultimately outputting a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected mode and motion vectors, also requires entropy coding to reduce the bit rate. Statistical coding is a lossless coding method that effectively reduces the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) or Content-Adaptive Binary Arithmetic Coding (CABAC).
[0077] 5. Loop Filtering: An encoded image undergoes inverse quantization, inverse transform, and prediction compensation (the reverse operations of steps 2-4 above) to obtain a reconstructed decoded image. Compared to the original image, the reconstructed image differs in some information due to quantization, resulting in distortion. Filtering the reconstructed image, such as deblocking, SAO (Sample Adaptive Offset), or ALF (Adaptive Lattice Filter), can effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will serve as a reference for subsequent encoded images to predict future signals, the above filtering operations are also called loop filtering, or filtering operations within the encoding loop.
[0078] As can be seen from the above encoding process, at the decoding end, for each CU, after obtaining the compressed bitstream, the decoder first performs entropy decoding to obtain various mode information and quantized transform coefficients. Each coefficient undergoes inverse quantization and inverse transform to obtain the residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to that CU can be obtained. After adding the two, the reconstructed signal is obtained. Finally, the reconstructed value of the decoded image needs to undergo a loop filtering operation to generate the final output signal.
[0079] Many mainstream video coding standards, such as HEVC, VVC, and AVS3, employ block-based hybrid coding frameworks. They divide the raw video data into a series of coded blocks and combine prediction, transform, and entropy coding methods to achieve video data compression. Motion compensation is a commonly used prediction method in video coding. Based on the redundancy characteristics of video content in the temporal or spatial domains, motion compensation derives the predicted value of the current coded block from the already coded regions. These prediction methods include inter-frame prediction, intra-frame block copy prediction, and intra-frame string copy prediction. In specific coding implementations, these prediction methods may be used individually or in combination. For coded blocks using these prediction methods, one or more two-dimensional displacement vectors are typically explicitly or implicitly encoded in the bitstream to indicate the displacement of the current coded block (or its sibling block) relative to one or more reference blocks.
[0080] Figure 1 A schematic diagram of a computer system provided in one embodiment of this application is shown. This computer system can implement a system architecture for a video data processing method. The computer system may include: a terminal 100 and a server 200.
[0081] Terminal 100 can be an electronic device such as a mobile phone, tablet computer, or PC (Personal Computer). A client application for the target application can be installed and run on terminal 100. This target application can be a video data processing application, a model building application, or other applications that provide data processing functions; this application does not limit the specific application. Furthermore, this application does not limit the form of the target application, including but not limited to apps, mini-programs, etc., installed on terminal 100, and it can also be a webpage.
[0082] Server 200 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Server 200 can be the backend server for the aforementioned target application, used to provide backend services to the clients of the target application.
[0083] The video data processing method provided in this application embodiment can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Figure 1Taking the illustrated implementation environment as an example, the video data processing method can be executed by terminal 100 (e.g., by a client of the target application installed and running on terminal 100), by server 200, or by interaction and cooperation between terminal 100 and server 200. This application does not limit the scope of the method. It should be noted that the technical solutions provided in this application can be applied to H.266 / VVC standards, H.265 / HEVC standards, AVS (such as AVS3), or next-generation video encoding / decoding standards. This application does not limit the scope of the method.
[0084] Furthermore, the technical solution of this application can be combined with blockchain technology. For example, in the video data processing method disclosed in this application, some data involved (such as color information of video frames, encoding information of video data, etc.) can be stored on the blockchain. The terminal 100 and the server 200 can communicate through a network, such as a wired or wireless network.
[0085] Next, the video data processing method in this application will be introduced:
[0086] Figure 2 A schematic diagram of a video data processing method provided in an exemplary embodiment of this application is shown.
[0087] For example, color information of the first video frame 302 and the second video frame 304 is obtained. The first video frame 302 and the second video frame 304 are two adjacent video frames in the video data. The first video frame 302 is the video frame displayed before the second video frame 304.
[0088] The second video frame 304 comprises a pixel block arranged in three rows and five columns, with each pixel block containing a total of 64 pixels arranged in eight rows and eight columns. It is understood that in different embodiments, the second video frame 304 may be divided into more or fewer pixel blocks, and the number and arrangement of pixels within each pixel block may differ from this embodiment.
[0089] Taking the first pixel block 304a in the second video frame 304 as an example, the first pixel block 304a is a pixel block located in the second row and third column; the first pixel block 304a includes 64 pixels, and the motion information 306 of each pixel in the first pixel block 304a is obtained.
[0090] Motion information 306 is used to indicate the positional difference of the same virtual spatial location in the virtual environment between the first video frame 302 and the second video frame 304. For example, a virtual ramp exists in the first video frame 302, located at the left frame position 3020 of the first video frame 302; while in the second video frame 304, the virtual ramp is located at the right frame position 3040 of the second video frame 304; motion information 306 is used to indicate the positional difference between the left frame position 3020 and the right frame position 3040.
[0091] Motion information 306 for each pixel in the first pixel block 304a, including the virtual spatial positions corresponding to the 64 pixels in the first pixel block 304a, and the positional differences in the first video frame 302 and the second video frame 304. Figure 2 The direction and distance of positional differences are indicated by directional arrows.
[0092] The average value of the motion information 306 of each pixel in the first pixel block 304a is determined as the motion vector of the first pixel block 304a. The motion vector is the overall representation information of the first pixel block 304a, which is statistically obtained from each pixel in the first pixel block 304a.
[0093] For example, based on the position of the first pixel block 304a in the second video frame 304, the second pixel block 302a in the first video frame 302 is obtained by translating it in the direction and distance indicated by the motion vector.
[0094] The color information of the second pixel block 302a is obtained from the color information of the first video frame 302 according to the position of the second pixel block 302a; the color information of the first pixel block 304a is obtained from the color information of the second video frame 304 according to the position of the first pixel block 304a.
[0095] Color residual information 310 is calculated based on the difference between the color information of the second pixel block 302a and the color information of the first pixel block 304a. For example, the color residual information 310 and the motion vector of the first pixel block 304a can be used to obtain the color information of the first pixel block 304a based on the color information of the first video frame 302. Repeating the above process yields the motion vectors and color residual information of each pixel block in the second video frame 304; this allows the color information of the second video frame 304 to be obtained based on the color information of the first video frame 302, using the motion vectors and color residual information of multiple pixel blocks. Constructing video data encoding information based on the color information of the first video frame 302 and the motion vectors and color residual information of the pixel blocks in the second video frame 304 reduces the space required to store the video data compared to directly constructing the encoding information based on the color information of the first video frame 302 and the second video frame 304.
[0096] The following examples will illustrate the video data processing method.
[0097] Figure 3 A flowchart illustrating a video data processing method provided in an exemplary embodiment of this application is shown. The method can be executed by a computer device. The method includes:
[0098] Step 510: Obtain the color information of the i-th video frame and the (i+1)-th video frame, and obtain the motion information of at least one pixel in the (i+1)-th video frame;
[0099] For example, the i-th video frame and the i+1-th video frame are two adjacent video frames in a video information. The color information can be directly indicated through information in the form of red-green-blue (RGB) channels, or indirectly indicated through information in the form of hue-saturation-value (HSV) channels. This application does not limit the data format of the color information. For example, the i-th video frame and the i+1-th video frame are used to display a virtual environment; for example, the i-th video frame and the i+1-th video frame are two video frames in chronological order from front to back, with the i-th video frame being the first in the chronological order and the i+1-th video frame being the second in the chronological order; for example, the i-th video frame and the i+1-th video frame are two adjacent video frames.
[0100] For example, motion information is used to indicate the positional change of a pixel in the (i+1)th video frame relative to the pixel in the ith video frame. At least one pixel can be a pixel at any position in the (i+1)th video frame, a pixel at the center of the (i+1)th video frame, and / or a pixel at the center of a region after the (i+1)th video frame is divided into multiple regions (or pixel blocks) according to a preset size. This application does not limit the position of the pixel. For example, the motion information of at least one pixel is a set of information, such as a set of motion information for each pixel within the at least one pixel.
[0101] Step 520: Determine the motion vector of at least one pixel block in the (i+1)th video frame based on the motion information of at least one pixel.
[0102] For example, for motion information of at least one pixel, the at least one pixel includes a second pixel; the motion information of the second pixel is used to indicate the difference between a second position of the second pixel in the (i+1)th video frame and a first position of the first pixel in the ith video frame. For example, the motion information of the second pixel is used to indicate the positional difference between two pixels on two video frames, and the exemplary positional difference is a vector information including direction and distance on the plane where the video frames are located.
[0103] For example, the virtual environment can be obtained by using a virtual camera to observe a virtual environment. This virtual environment can be provided by an application installed on a computer device, or by a program installed on another device. Understandably, the virtual environment can be provided by a game application or a simulation application. For example, the virtual environment can be any of a two-dimensional, 2.5-dimensional, or three-dimensional virtual environment. For example, the second pixel and the first pixel correspond to the same position in the virtual environment; the motion information of the second pixel can indicate the changes in the observed view of the virtual environment between two video frames. For example, the pixel data of the second pixel is the color information of the first position in the virtual environment displayed in the (i+1)th video frame, and the pixel data of the first pixel is the color information of the first position in the virtual environment displayed in the ith video frame; the visual information presented by the pixel data of the two pixels is at the same position in the virtual environment; taking a three-dimensional spatial environment as an example, the second pixel and the first pixel correspond to the same position in the three-dimensional spatial coordinate system describing the three-dimensional spatial environment. Understandably, a virtual environment can be implemented as at least one of two-dimensional virtual environments, 2.5-dimensional virtual environments, etc., and the same position is the same coordinate position in a coordinate system with the virtual environment as the reference frame.
[0104] For example, a virtual environment may contain virtual characters, virtual items, and other virtual objects, and the pixel data of the second pixel and the pixel data of the first pixel may correspond to the same virtual object in the virtual environment.
[0105] For example, due to differences in the observation position and / or observation angle of the virtual environment in different video frames, the second position of the second pixel in the (i+1)th video frame differs from the first position of the first pixel in the ith video frame. The motion information, describing this positional difference, as described above, includes vector information in the plane of the video frame, encompassing both direction and distance. For example, both the first and second positions are positions in the two-dimensional coordinate system of the video frame. For example, the difference between the first and second positions is a relative positional difference in the two-dimensional coordinate system.
[0106] For example, a pixel block includes at least two pixels in the (i+1)th video frame, and the motion vector of the pixel block is obtained based on the motion information of the pixels; for example, the motion information of the pixels is used to characterize the motion vector of the pixel block.
[0107] For example, the motion vector of a pixel block can be determined based on the motion information of one or more pixels. For example, it is not excluded that the motion information of pixels can be used to determine the motion vector of multiple pixel blocks. For example, when determining the motion vector of a pixel block based on the motion information of at least two pixels, the motion vector of the at least two pixels can be determined by at least one of the following methods: statistical average, median, mode, etc. This application does not limit the method of determining the motion vector.
[0108] Step 530: Based on the motion vector of at least one pixel block, determine the color residual information corresponding to at least one pixel block in the color information of the i-th video frame and the (i+1)-th video frame, respectively.
[0109] For example, the motion vector of a pixel block is used to indicate the positional difference between two pixel blocks on two video frames; for example, the color residual information is the color difference between a pixel block in the (i+1)th video frame and its corresponding pixel block in the ith video frame. For example, the color residual information indicates the color difference at each pixel point within a pixel block.
[0110] Step 540: Construct the encoding information of the video data based on the color information of the i-th video frame, and the motion vector and color residual information of at least one pixel block;
[0111] For example, for the (i+1)th video frame, the motion vector of at least one pixel block indicates the difference from the i-th video frame in the pixel block dimension, representing a difference information over a large field of view. Color residual information, based on the motion vector, describes a difference information over a small field of view in the dimension of each pixel. By describing the differences between the i-th and i+1th video frames from different dimensions, the color information of the (i+1)th video frame can be indicated based on the color information of the i-th video frame, ensuring the integrity of the color information of the (i+1)th video frame. The encoded information of the constructed video data does not need to directly carry the color information of the (i+1)th video frame, reducing the volume of the encoded video data and improving the efficiency of video data storage and / or transmission. For example, the encoded video data can be transmitted between different computer devices in the form of one or more encoded video streams.
[0112] In summary, the method provided in this embodiment fully utilizes the virtual spatial positions of pixels on a video frame in a virtual environment by determining motion vectors through motion information. Encoding information is constructed based on motion vectors, color residual information, and the color information of the i-th video frame. Compared to directly storing the color information of the (i+1)-th video frame, this reduces the space required to store video data. By using the motion vectors and color residual information of pixel blocks, the color information of the (i+1)-th video frame can be obtained based on the color information of the i-th video frame. While reducing the storage space of video data, the integrity of the color information of the (i+1)-th video frame is ensured, guaranteeing the completeness of the image of the (i+1)-th video frame.
[0113] Figure 4 A flowchart illustrating a video data processing method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 520 can be implemented as step 522:
[0114] Step 522: Determine the motion vector of the first pixel block based on the motion information of the pixels in the first pixel set;
[0115] In this embodiment, the motion vector of the first pixel block in the (i+1)th video frame is used as an example for description.
[0116] For example, in step 510, motion information of at least one pixel is obtained. A subset of the at least one pixel is a first pixel set, which includes one or at least two pixels. Each pixel in the first pixel set is a pixel in a first pixel block. The first pixel block is any one of one or more pixel blocks obtained by dividing the (i+1)th video frame according to a preset size.
[0117] For example, the motion vector of the first pixel block is determined based on the motion information of the pixels within the first pixel block. For example, the motion vector can be determined using one or at least two pixels from the first pixel set.
[0118] For example, when the motion vector is determined based on the motion information of a pixel in the first pixel set, the motion information of a pixel is determined as the motion vector of the first pixel block.
[0119] For example, when the motion vector is determined based on the motion information of at least two pixels in the first pixel set, statistical calculations are performed on the motion information of the at least two pixels to obtain the motion vector of the first pixel block. Further, performing statistical calculations on the motion information of the at least two pixels includes, but is not limited to, at least one of the following methods: statistical average, median, mode, etc. Further, step 522 can be implemented as any one of the following:
[0120] • The average value of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block;
[0121] For example, the average value of motion information is used to indicate the average positional difference of at least two pixels in two videos; the motion vector determined based on the average value can make the sum of the color residual information of each pixel in the pixel block approach the minimum value, thereby reducing the storage space required for color residual information from the dimension of the sum of the color residual information of each pixel.
[0122] For example, the motion vector of the pixel block is:
[0123]
[0124] Among them, MV x It is the motion vector of the x-th pixel block, where n is the number of pixels in the pixel block, mv ix It is the motion information of the i-th pixel in the x-th pixel block; similarly, MV y mv is the motion vector of the y-th pixel block, where n is the number of pixels in the pixel block. iy It represents the motion information of the i-th pixel in the y-th pixel block.
[0125] • The median of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block;
[0126] For example, the median of motion information indicates the motion information of a pixel segment where the positional difference between at least two pixels in two videos is at an intermediate level. It uses the motion information of a real pixel directly as the motion vector of the pixel block. It takes into account that the sum of the color residual information of each pixel approaches the minimum value, and directly selects the motion vector from the motion information. It can preliminarily verify whether the stored motion vector is correct by whether there is a motion vector in the motion information.
[0127] For example, the motion vector of the pixel block is:
[0128]
[0129] Among them, MV med,xIt is the motion vector determined by the median in the x-th pixel block, where n is the number of pixels in the pixel block. It is the xth pixel in the xth pixel block. Motion information of each pixel It is the xth pixel in the xth pixel block. Motion information of individual pixels; similarly, MV med,y is the motion vector of the y-th pixel block, and n is the number of pixels in the pixel block. It is the y-th pixel in the y-th pixel block. Motion information of each pixel It is the y-th pixel in the y-th pixel block. Motion information of each pixel.
[0130] • The motion vector of the first pixel block is determined by the modulo value of the motion information of at least two pixels in the first pixel set;
[0131] For example, the median of motion information is the motion information of the maximum number of pixels whose positions differ equally in two videos; it is the motion information of a real pixel directly used as the motion vector of the pixel block; and it is possible to directly verify whether the stored motion vector is correct by checking whether the motion vector is the maximum number of identical information in the motion information.
[0132] In summary, the method provided in this embodiment, through motion vectors determined by motion information, fully utilizes the corresponding virtual spatial positions of pixels on video frames in the virtual environment. The motion vector of the first pixel block is determined based on the motion information of the pixels within the first pixel block, ensuring that the motion vector of the first pixel block can characterize the motion of the pixels within the first pixel block. Encoding information is constructed based on the motion vector, color residual information, and color information of the i-th video frame. Compared to directly storing the color information of the (i+1)-th video frame, this reduces the space required to store video data. Through the motion vector and color residual information of the pixel block, the color information of the (i+1)-th video frame can be obtained based on the color information of the i-th video frame. While reducing the storage space of video data, the integrity of the color information of the (i+1)-th video frame is ensured, guaranteeing the integrity of the image of the (i+1)-th video frame.
[0133] Figure 5 A flowchart illustrating a video data processing method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 4 In the illustrated embodiment, step 522 can be implemented as steps 522a, 522b, and 522c:
[0134] Step 522a: Determine the reference vector of the first pixel block based on the motion information of the pixels in the first pixel set;
[0135] For example, the reference vector is determined based on motion information. For example, the reference vector is an intermediate parameter in the process of determining the motion vector; referring to the description in step 522, the reference information can be determined directly or statistically based on the motion information of one or more pixels within the first pixel block. For the method of determining the reference vector, please refer to... Figure 4 The method for determining the motion vector in the corresponding embodiments will not be repeated here.
[0136] Step 522b: Based on the position of the first region of the first pixel block in the (i+1)th video frame and the reference vector, determine the color sub-information of the reference pixel block in the color information of the i-th video frame;
[0137] For example, the position of the second region where the reference pixel block is located is obtained by moving the reference vector from the first region position of the first pixel block. The reference vector indicates the direction and distance of movement relative to the first region position. The color sub-information of the reference pixel block is the information located in the second region position from the color information of the i-th video frame.
[0138] Step 522c: If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed the color threshold, the reference vector is determined as the motion vector of the first pixel block;
[0139] For example, if the difference does not exceed the color threshold, the similarity between the first pixel block and the reference pixel block meets the color threshold requirement, and the reference vector is directly determined as the motion vector of the first pixel block. For example, the color threshold constrains the similarity between the reference pixel block indicated by the reference vector and the first pixel block. This avoids the problem of increased video data encoding volume due to the complex information required to be recorded in the color residual information caused by excessive differences between the first pixel block and the reference pixel block.
[0140] For example, the color threshold can be a pre-set empirical value or determined based on the color information of the i-th video frame; this application does not limit the method of determining the color threshold. In an alternative implementation, the color threshold is the minimum value among at least one reference difference degree.
[0141] For example, the reference difference is used to indicate the difference between the color sub-information of neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
[0142] Neighboring pixel blocks have the same size as the reference pixel block and are located in the peripheral region of the reference pixel block. The size of the peripheral region is larger than that of the reference pixel block; for example, the center point of the peripheral region is the center point of the reference pixel block. There are usually overlapping pixel blocks between neighboring and reference pixel blocks, but it is not impossible for pixels in a neighboring pixel block to be completely different from pixels in the reference pixel block.
[0143] In summary, the method provided in this embodiment fully utilizes the virtual spatial positions of pixels on the video frame in the virtual environment by determining the motion vector through motion information; it verifies the motion vector through a color threshold to ensure that the pixel block indicated by the motion vector is similar to the first pixel block, simplifying the complexity of the information in the color residual information; and it constructs encoding information based on the motion vector, color residual information, and color information of the i-th video frame; compared to directly storing the color information of the (i+1)-th video frame, it reduces the space required to store video data.
[0144] Next, we will further explain the color threshold. For example, in... Figure 5 In addition to, it also includes, Figure 6 Steps 524 and 526 are shown:
[0145] Step 524: If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds the color threshold, determine the correction pixel block;
[0146] For example, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
[0147] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds the color threshold, a correction pixel block with a smaller difference from the first pixel block is determined in the i-th video frame.
[0148] Similar to the neighboring pixel block mentioned above, the corrected pixel block is located in the peripheral region of the reference pixel block, and the size of the peripheral region is larger than that of the reference pixel block; Figure 7 This illustration shows a schematic diagram of video data processing provided by an exemplary embodiment of this application. Color information of a first video frame 302 and a second video frame 304 is obtained. The first video frame 302 and the second video frame 304 are two adjacent video frames in the video data, with the first video frame 302 being the video frame displayed before the second video frame 304.
[0149] The reference vector 305 of the first pixel block 304a points to the upper left and is obtained statistically from each pixel in the first pixel block 304a (such as the average value of the motion information of each pixel). It is the representation information of the first pixel block 304a as a whole.
[0150] Based on the position of the first pixel block 304a in the second video frame 304, a reference position is obtained by translating it in the direction and distance indicated by the reference vector 305. This reference position in the first video frame 302 is the position of the reference pixel block 302b. The peripheral region 315 of the reference pixel block 302b is centered on the reference pixel block 302b, and its area is four times the area of the reference pixel block 302b. It is understandable that, in different examples, the area of the peripheral region 315 and the area of the reference pixel block 302b may have other different proportional relationships.
[0151] refer to Figure 2 Taking the first pixel block 304a, which comprises 64 pixels in eight rows and eight columns, as an example, the surrounding region 315 comprises 256 pixels in sixteen rows and sixteen columns. Within the surrounding region 315, (16-8+1)*(16-8+1) = 81 candidate pixel blocks can be sampled, each candidate pixel block having a size of 64 pixels in eight rows and eight columns. Any two candidate pixel blocks are distinct, but some overlap is possible. With the goal of minimizing the difference between the color sub-information of the candidate pixel blocks and the color sub-information of the first pixel block 304a, a color difference search is performed to obtain the corrected pixel block 316. Based on the positional difference between the first pixel block 304a in the first region of the second video frame 304 and the corrected pixel block 316 in the second region of the first video frame 302, the motion vector 304b of the first pixel block 304a is determined.
[0152] For example, there are usually overlapping pixel blocks between the corrected pixel block and the reference pixel block, but it is not impossible for the pixels in the corrected pixel block to be completely different from those in the reference pixel block. For example, the corrected pixel block is obtained by searching in the peripheral region of the reference pixel block according to the block matching algorithm; for example, the block matching algorithm can be implemented as at least one of the following: Sum of Absolute Differences (SAD) algorithm, Sum of Squared Differences (SSD) algorithm, Zero-mean Normalized Cross-Correlation (ZNCC) algorithm, and Semi-Global Block Matching (SGBM) algorithm.
[0153] Step 526: Determine the motion vector of the first pixel block based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th video frame;
[0154] For example, when performing step 524, if the difference between the reference pixel block and the first pixel block is too large, the motion vector of the first pixel block is determined based on the positional difference between the correction pixel block and the first pixel block, which is smaller in difference.
[0155] For example, steps 524 and Figure 5 Step 522c consists of two parallel steps. As described above, if the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds a color threshold, step 524 is executed; if the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed the color threshold, step 522c is executed. For example, in this embodiment, step 524 is executed after step 522b above.
[0156] In summary, the method provided in this embodiment fully utilizes the virtual spatial positions of pixels on video frames in the virtual environment by determining motion vectors through motion information; it verifies motion vectors through a color threshold, and when the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds the color threshold, it redetermines and corrects the motion vectors of the pixel blocks; it ensures that the pixel blocks indicated by the motion vectors are similar to the first pixel blocks, simplifying the complexity of information in the color residual information; and it constructs encoding information based on motion vectors, color residual information, and color information of the i-th video frame; compared to directly storing the color information of the (i+1)-th video frame, it reduces the space required to store video data.
[0157] Figure 8 This illustration shows a schematic diagram of video data processing provided by an exemplary embodiment of this application.
[0158] Obtain the color information of the current frame 602 and the historical frame 604, and determine the motion vector (MV) 606 of at least one pixel block in the current frame 602; for an example of how the motion vector 606 is determined, please refer to the above text. Figure 3 , Figure 4 The corresponding implementation examples will not be repeated here.
[0159] For example, taking the first pixel block in the current frame 602 as an example, the motion compensator (MC) 612 determines the associated pixel block corresponding to the first pixel block in the current frame 602 in the historical frame 604 based on the motion vector 606. The motion compensator 612 obtains the color difference between the first pixel block and the associated pixel block, that is, the motion compensator 612 outputs color residual information. The switch 615 is placed in the inter-frame prediction branch and connected to the motion compensator 612; the transform (T) 622 performs a transform based on the current frame 602 and the color residual information, transforming the information to the frequency domain space. The quantization (Q) 624 is used to reduce the dynamic range of the transformed information. The reorder 626 rearranges the information output by the quantizer 624. The entropy encoder 628 performs entropy encoding on the information output by the quantizer 624 to obtain encoded information 632, thereby reducing the amount of stored data. For example, the encoding methods include, but are not limited to, at least one of Shannon coding, Huffman coding, and arithmetic coding. For example, the encoded information 632 is the information obtained by processing the motion vector 606 and color residual information through the transformer, quantizer, and encoder. However, it is not excluded that in some examples, the encoded information directly stores the motion vector 606 and / or color residual information.
[0160] For example, in another example, switch 615 is placed in the intra-prediction branch and connected to intra-prediction 614; intra-prediction 614 is used to remove spatial redundancy within video frames and reduce the amount of data in video frames based on the correlation between adjacent pixels within the same video frame.
[0161] Next, we will provide further information about the exercise.
[0162] Figure 9 A flowchart illustrating a video data processing method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 Based on the illustrated embodiment, step 515 is also included:
[0163] Step 515: Perform data normalization and / or coordinate axis transformation on the motion information of at least one pixel to obtain the corrected motion information of at least one pixel;
[0164] For example, the purpose of performing data normalization on the motion information of pixels is to constrain the motion information of pixels within a data range, such as [-1, 1] or [0, 255]. When there is motion information for at least two pixels, this eliminates the differences in dimensions and scales in the motion information of each pixel, making the motion information of different pixels comparable.
[0165] For example, the methods for performing data normalization on the motion information of at least one pixel include, but are not limited to, at least one of the following: Min-Max Normalization, Z-Score Normalization, and Log Transform.
[0166] For example, the pixel motion information is referenced to a first Cartesian coordinate system, such as one provided by the rendering engine of the virtual environment; the motion vector of the pixel block is referenced to a second Cartesian coordinate system; in order to avoid the origin positions and / or coordinate axis directions of the first and second Cartesian coordinate systems being different, coordinate axis transformation is performed on the motion information of at least one pixel to avoid the motion information failing to correctly indicate the positional differences between the pixel blocks in the i-th video frame and the i+1-th video frame.
[0167] In this embodiment, the motion vector of at least one pixel block is determined based on the corrected motion information of at least one pixel.
[0168] The following section introduces the data format for motion information.
[0169] In one alternative design of this application, the motion information of the second pixel includes distance sub-information in the first direction and the second direction;
[0170] For example, the first direction and the second direction are the directions of the two perpendicular sides of the i-th video frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first and second directions. Describing the motion information of the second pixel point with distance sub-information in two mutually perpendicular directions is beneficial for determining the corresponding pixel point in a matrix of pixels in an integer manner based on the distance sub-information.
[0171] In another alternative design of this application, the motion information of the second pixel includes direction sub-information and distance sub-information;
[0172] For example, the direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position. This can intuitively describe the direction and distance of motion between two pixel frames.
[0173] In summary, the method provided in this embodiment, by performing data normalization and / or coordinate axis transformation on the motion information in preprocessing, ensures the elimination of dimensional and scale differences in the motion information of each pixel, and / or avoids the motion information failing to correctly indicate the positional differences between pixel blocks in the i-th video frame and the i+1-th video frame; the motion vector determined by the motion information fully utilizes the corresponding virtual spatial positions of pixels on the video frame in the virtual environment, and constructs encoded information based on the motion vector, color residual information, and color information of the i-th video frame; compared to directly storing the color information of the i+1-th video frame, it reduces the space required to store video data.
[0174] Figure 10 A flowchart illustrating a video data processing method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 510 can be implemented as steps 511 to 513; step 530 can be implemented as steps 532 and 534.
[0175] Step 511: Obtain the color information of the i-th video frame and the (i+1)-th video frame;
[0176] For example, the i-th video frame and the i+1-th video frame are two adjacent video frames in a video information. The color information can be directly indicated by information in the form of red-green-blue (RGB) channel, or indirectly indicated by information in the form of hue-saturation-value (HSV) channel. This application does not limit the data format of the color information.
[0177] For example, step 511 can be performed before, after, or simultaneously with any step in the first step group, which includes steps 512 and 513. This application does not limit the above execution sequence.
[0178] Step 512: Obtain control information in the virtual environment;
[0179] For example, the control information is used to control virtual objects and / or virtual cameras in a virtual environment; the control information is used to indicate at least one of the following: rotating the angle of the virtual camera, moving the position of the virtual camera, moving the position of the virtual object, rotating the viewpoint of the virtual object, controlling the virtual object to perform virtual actions, etc.
[0180] Step 513: Based on the control information and the color information of the i-th video frame, determine the motion information of at least one pixel in the (i+1)-th video frame;
[0181] For example, the control information directly or indirectly indicates how the color information of the i-th video frame is moved to obtain the (i+1)-th video frame. Taking the rotation of a virtual character's viewpoint as an example, the visual effect of rotating the virtual character's viewpoint is achieved by rotating and moving the pixels in the i-th video frame around the main virtual character in the video frame. Taking the movement of a virtual camera's position as an example, the visual effect of moving the virtual camera is achieved by translating the pixels in the i-th video frame.
[0182] For example, the control information, categorized by type, includes at least one of the following described in step 512: rotating the virtual camera angle, moving the virtual camera position, moving the virtual object position, rotating the virtual object's viewpoint, and controlling the virtual object to perform virtual actions. Different types of control information correspond to different transformation methods for the color information of the i-th video frame (such as translation, rotation, etc., as described above). Based on the mapping relationship between candidate control information and the transformation methods for the color information of the i-th video frame, the first transformation method for the color information of the i-th video frame corresponding to the control information is found; based on the scale information in the control information, the scale of the transformation performed on the color information of the i-th video frame is determined, such as the rotation angle, translation distance, and other scale information.
[0183] For example, the first transformation mode and scale information indicated by the control information are converted into the motion information of the pixel; referring to the description of step 515 above, the motion information of the pixel may include distance sub-information in the first direction and the second direction, or direction sub-information and distance sub-information.
[0184] For example, taking the second pixel in the (i+1)th video frame as an example, the motion information of the second pixel is used to indicate the positional difference between the first pixel and the second pixel; the first pixel is a pixel in the i-th video frame, and the first pixel and the second pixel correspond to the same position in the virtual environment.
[0185] It is understandable that steps 511 to 513 in this embodiment can be compared with... Figure 3 Steps 520 to 540 shown in the figure can be combined into new embodiments and implemented separately, and this application does not limit this.
[0186] Step 532: Based on the motion vector of the first pixel block, determine the color information of the associated pixel block in the color information of the i-th video frame;
[0187] For example, the first pixel block is located in a first region in the (i+1)th video frame. Based on this first region location, it is moved by a motion vector to obtain the region location (also called the second region location) of the associated pixel block in the i-th video frame. The color information of the associated pixel block is a sub-part of the color information of the i-th video frame, specifically the color information of the second region location in the i-th video frame. For example, the associated pixel block determined in the color information of the i-th video frame based on the motion vector of the first pixel block is also called the second pixel block.
[0188] For example, Figure 3 The at least one pixel block described herein includes a first pixel block, which is any pixel block that determines the motion vector.
[0189] Step 534: Determine the color residual information of the first pixel block based on the difference between the color information of the associated pixel block and the color information of the first pixel block;
[0190] For example, the color residual information of the first pixel block is used to indicate the difference in color dimension between the associated pixel block and the first pixel block, such as the difference in color values in the RGB channel; however, it is not excluded that the difference in color dimension is recorded in the HSV channel or other ways.
[0191] It is understandable that steps 532 and 534 in this embodiment can be compared with... Figure 3 Steps 510, 520 and 540 shown in the figure can be combined into new embodiments and implemented separately, and this application does not limit them.
[0192] In summary, the method provided in this embodiment fully utilizes the virtual spatial positions of pixels on a video frame in a virtual environment by determining motion vectors through motion information. Encoding information is constructed based on motion vectors, color residual information, and the color information of the i-th video frame. Compared to directly storing the color information of the (i+1)-th video frame, this reduces the space required to store video data. By using the motion vectors and color residual information of pixel blocks, the color information of the (i+1)-th video frame can be obtained based on the color information of the i-th video frame. While reducing the storage space of video data, the integrity of the color information of the (i+1)-th video frame is ensured, guaranteeing the completeness of the image of the (i+1)-th video frame.
[0193] In one application scenario of this application, a computer device is equipped with a cloud gaming application that provides video data processing capabilities. Further, the cloud gaming application sends encoded information to a terminal, which also has the cloud gaming application installed. The terminal is used to display the running screen of the cloud game. Understandably, the terminal is the user's terminal, used to send control information to the computer device for virtual characters in the cloud game, receive encoded information, and display game frames obtained after decoding the encoded information to the user. Accordingly,
[0194] • Obtain the color information of the i-th game frame and the (i+1)-th game frame, and obtain the motion information of at least one pixel in the (i+1)-th game frame;
[0195] For example, the i-th game frame and the (i+1)-th game frame are two adjacent video frames in the game video information of a cloud game.
[0196] For example, motion information is used to indicate the positional change of a pixel in the (i+1)th game frame relative to a pixel in the ith game frame.
[0197] Optionally, the motion information of the second pixel includes distance sub-information in the first and second directions;
[0198] Wherein, the first direction and the second direction are the directions of the two vertical sides of the i-th game frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first direction and the second direction.
[0199] Optionally, the motion information of the second pixel includes direction sub-information and distance sub-information;
[0200] The direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position.
[0201] • Determine the motion vector of at least one pixel block in the (i+1)th game frame based on the motion information of at least one pixel.
[0202] For example, for motion information of at least one pixel, the at least one pixel includes a second pixel; the motion information of the second pixel is used to indicate the difference between a second position of the second pixel in the (i+1)th game frame and a first position of the first pixel in the i-th game frame. For example, the motion information of the second pixel is used to indicate the positional difference between two pixels on two game frames.
[0203] For example, the i-th game frame and the (i+1)-th game frame are used to display a virtual environment, which is provided by a cloud gaming application installed on the computer device. The virtual environment provides a virtual space for virtual characters to compete in virtual battles. For example, the second pixel and the first pixel correspond to the same position in the virtual environment; the motion vector of the second pixel is used to indicate changes in the viewing screen of the virtual environment.
[0204] For example, a pixel block includes at least two pixels in the (i+1)th game frame, and the motion vector of the pixel block is obtained based on the motion information of the pixels; for example, the motion information of the pixels is used to characterize the motion vector of the pixel block.
[0205] Optionally, the motion vector of the first pixel block can be determined based on the motion information of the pixels in the first pixel set;
[0206] The first pixel set is a subset of at least one pixel, and each pixel in the first pixel set is a pixel in the first pixel block; further, any one of the average, median, or mode of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block.
[0207] Optionally, the reference vector of the first pixel block can be determined based on the motion information of the pixels in the first pixel set;
[0208] Based on the position of the first pixel block in the first region of the (i+1)th game frame and the reference vector, the color sub-information of the reference pixel block is determined in the color information of the i-th game frame.
[0209] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed the color threshold, the reference vector is determined as the motion vector of the first pixel block.
[0210] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds the color threshold, a correction pixel block is determined.
[0211] The motion vector of the first pixel block is determined based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th game frame.
[0212] Specifically, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
[0213] Furthermore, the color threshold is the minimum of at least one reference difference degree; the reference difference degree is used to indicate the difference between the color sub-information of the neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
[0214] For example, the motion vector of a pixel block can be determined based on the motion information of one or at least two pixels.
[0215] • Based on the motion vector of at least one pixel block, determine the color residual information corresponding to at least one pixel block in the color information of the i-th game frame and the (i+1)-th game frame, respectively;
[0216] For example, the motion vector of a pixel block is used to indicate the positional difference between two pixel blocks on two game frames; for example, the color residual information is the color difference between a pixel block in the (i+1)th game frame and its corresponding pixel block in the ith game frame. For example, the color residual information indicates the color differences at each pixel point within a pixel block.
[0217] Optionally, based on the motion vector of the first pixel block, the color information of the associated pixel block is determined in the color information of the i-th game frame, and at least one pixel block includes the first pixel block;
[0218] The color residual information of the first pixel block is determined based on the difference between the color information of the associated pixel block and the color information of the first pixel block.
[0219] • Construct the encoding information of the video data based on the color information of the i-th game frame, as well as the motion vector and color residual information of at least one pixel block;
[0220] For example, for the (i+1)th game frame, the motion vector of at least one pixel block indicates the difference from the i-th game frame in the pixel block dimension, which is a difference information in a large field of view; the color residual information, based on the motion vector, describes a difference information in a small field of view in the dimension of each pixel. The differences between the i-th game frame and the (i+1)-th game frame are described from different dimensions. Based on the motion vector and the color residual information, the color information of the (i+1)-th game frame can be indicated based on the color information of the i-th game frame, ensuring the integrity of the color information of the (i+1)-th game frame; the encoded information of the constructed video data does not need to directly carry the color information of the (i+1)-th game frame, reducing the volume of the encoded information of the video data and improving the efficiency of video data storage and / or transmission.
[0221] In one application scenario of this application, the computer device is equipped with a game application or a tool application with game recording functionality, and the cloud gaming application provides video data processing; further, the game application or the tool application with game recording functionality provides the function of recording game footage from past matches. It is understood that the game can be a single-player game or an online game requiring a network connection. Furthermore, the application has the function of decoding encoded information and playing game footage. Accordingly,
[0222] • Obtain the color information of the i-th game frame and the (i+1)-th game frame, and obtain the motion information of at least one pixel in the (i+1)-th game frame;
[0223] For example, the i-th game frame and the (i+1)-th game frame are two adjacent video frames in the game video information of a cloud game.
[0224] For example, motion information is used to indicate the positional change of a pixel in the (i+1)th game frame compared to the pixel in the ith game frame.
[0225] Optionally, the motion information of the second pixel includes distance sub-information in the first and second directions;
[0226] Wherein, the first direction and the second direction are the directions of the two vertical edges of the i-th game frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first direction and the second direction.
[0227] Optionally, the motion information of the second pixel includes direction sub-information and distance sub-information;
[0228] The direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position.
[0229] • Determine the motion vector of at least one pixel block in the (i+1)th game frame based on the motion information of at least one pixel.
[0230] For example, for motion information of at least one pixel, the at least one pixel includes a second pixel; the motion information of the second pixel is used to indicate the difference between a second position of the second pixel in the (i+1)th game frame and a first position of the first pixel in the i-th game frame. For example, the motion information of the second pixel is used to indicate the positional difference between two pixels on two game frames.
[0231] For example, the i-th game frame and the (i+1)-th game frame are used to display a virtual environment, which is a virtual environment provided by a cloud gaming application installed on a computer device. The virtual environment provides a virtual space for the virtual characters' virtual competition. For example, the second pixel and the first pixel correspond to the same position in the virtual environment; the motion vector of the second pixel is used to indicate changes in the viewing screen of the virtual environment.
[0232] For example, a pixel block includes at least two pixels in the (i+1)th game frame, and the motion vector of the pixel block is obtained based on the motion information of the pixels; for example, the motion information of the pixels is used to characterize the motion vector of the pixel block.
[0233] Optionally, the motion vector of the first pixel block can be determined based on the motion information of the pixels in the first pixel set;
[0234] The first pixel set is a subset of at least one pixel, and each pixel in the first pixel set is a pixel in the first pixel block; further, any one of the average, median, or mode of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block.
[0235] Optionally, the reference vector of the first pixel block can be determined based on the motion information of the pixels in the first pixel set;
[0236] Based on the first region position and reference vector of the first pixel block in the (i+1)th game frame, the color sub-information of the reference pixel block is determined in the color information of the i-th game frame.
[0237] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed the color threshold, the reference vector is determined as the motion vector of the first pixel block.
[0238] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds the color threshold, a correction pixel block is determined.
[0239] The motion vector of the first pixel block is determined based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th game frame.
[0240] Specifically, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
[0241] Furthermore, the color threshold is the minimum of at least one reference difference degree; the reference difference degree is used to indicate the difference between the color sub-information of the neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
[0242] For example, the motion vector of a pixel block can be determined based on the motion information of one or at least two pixels.
[0243] • Based on the motion vector of at least one pixel block, determine the color residual information corresponding to at least one pixel block in the color information of the i-th game frame and the (i+1)-th game frame, respectively;
[0244] For example, the motion vector of a pixel block is used to indicate the positional difference between two pixel blocks on two game frames; for example, the color residual information is the color difference between a pixel block in the (i+1)th game frame and its corresponding pixel block in the ith game frame. For example, the color residual information indicates the color differences at each pixel point within a pixel block.
[0245] Optionally, based on the motion vector of the first pixel block, the color information of the associated pixel block is determined in the color information of the i-th game frame, and at least one pixel block includes the first pixel block;
[0246] The color residual information of the first pixel block is determined based on the difference between the color information of the associated pixel block and the color information of the first pixel block.
[0247] • Construct the encoding information of the video data based on the color information of the i-th game frame, as well as the motion vector and color residual information of at least one pixel block;
[0248] For example, for the (i+1)th game frame, the motion vector of at least one pixel block indicates the difference from the i-th game frame in terms of pixel block dimension, which is a difference information in a large field of view; the color residual information, based on the motion vector, describes a difference information in a small field of view in terms of the dimension of each pixel point. The differences between the i-th game frame and the (i+1)-th game frame are described from different dimensions. Based on the motion vector and the color residual information, the color information of the (i+1)-th game frame can be indicated based on the color information of the i-th game frame, ensuring the integrity of the color information of the (i+1)-th game frame; the encoded information of the constructed video data does not need to directly carry the color information of the (i+1)-th game frame, reducing the volume of the encoded information of the video data and improving the efficiency of video data storage and / or transmission.
[0249] Those skilled in the art will understand that the above embodiments can be implemented independently, or the above embodiments can be freely combined to create new embodiments to implement the video data processing method of this application.
[0250] Figure 11 A structural block diagram of a video data processing apparatus provided in an exemplary embodiment of this application is shown. The apparatus includes:
[0251] The acquisition module 810 is used to acquire the color information of the i-th video frame and the (i+1)-th video frame, and to acquire the motion information of at least one pixel in the (i+1)-th video frame, where i is a positive integer;
[0252] The processing module 820 is configured to determine the motion vector of at least one pixel block in the (i+1)th video frame based on the motion information of the at least one pixel, wherein the at least one pixel includes a second pixel, and the motion information of the second pixel is used to indicate the difference between the second position of the second pixel in the (i+1)th video frame and the first position of the first pixel in the i-th video frame, wherein the i-th video frame and the (i+1)th video frame are used to display a virtual environment, and the pixel data of the second pixel and the pixel data of the first pixel are used to characterize the visual information of the same position in the virtual environment;
[0253] The processing module 820 is further configured to determine, based on the motion vector of the at least one pixel block, color residual information corresponding to the at least one pixel block in the color information of the i-th video frame and the (i+1)-th video frame, respectively;
[0254] The construction module 830 is used to construct the encoding information of the video data based on the color information of the i-th video frame, the motion vector of the at least one pixel block, and the color residual information.
[0255] In an optional implementation of this embodiment, the first pixel set is a subset of the at least one pixel, and each pixel in the first pixel set is a pixel in the first pixel block;
[0256] The processing module 820 is further configured to:
[0257] The motion vector of the first pixel block is determined based on the motion information of the pixels in the first pixel set.
[0258] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0259] The average value of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block;
[0260] Alternatively, the median of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block;
[0261] Alternatively, the modulo value of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block.
[0262] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0263] Based on the motion information of the pixels in the first pixel set, the reference vector of the first pixel block is determined;
[0264] Based on the first region position of the first pixel block in the (i+1)th video frame and the reference vector, the color sub-information of the reference pixel block is determined in the color information of the i-th video frame;
[0265] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed a color threshold, the reference vector is determined as the motion vector of the first pixel block.
[0266] In an optional implementation of this embodiment, the color threshold is the minimum value among at least one reference difference degree; the reference difference degree is used to indicate the difference between the color sub-information of the neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
[0267] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0268] If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds a color threshold, a correction pixel block is determined.
[0269] The motion vector of the first pixel block is determined based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th video frame.
[0270] Wherein, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
[0271] In an optional implementation of this embodiment, the motion information of the second pixel includes distance sub-information in the first direction and the second direction;
[0272] Wherein, the first direction and the second direction are the directions of the two vertical sides of the i-th video frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first direction and the second direction.
[0273] In an optional implementation of this embodiment, the motion information of the second pixel includes direction sub-information and distance sub-information;
[0274] The direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position.
[0275] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0276] Perform data normalization and / or coordinate axis transformation on the motion information of the at least one pixel to obtain the corrected motion information of the at least one pixel;
[0277] The motion vector of the at least one pixel block is determined based on the corrected motion information of the at least one pixel.
[0278] In an optional implementation of this embodiment, the acquisition module 810 is further configured to:
[0279] Acquire control information in the virtual environment, the control information being used to control virtual objects and / or virtual cameras in the virtual environment;
[0280] Based on the control information, the motion information of at least one pixel in the (i+1)th video frame is determined.
[0281] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0282] Based on the motion vector of the first pixel block, the color information of the associated pixel block is determined in the color information of the i-th video frame, wherein the at least one pixel block includes the first pixel block;
[0283] The color residual information of the first pixel block is determined based on the difference between the color information of the associated pixel block and the color information of the first pixel block.
[0284] It should be noted that the device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0285] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments of the relevant method; the technical effects achieved by each module performing its operation are the same as the technical effects in the embodiments of the relevant method, and will not be elaborated here.
[0286] This application also provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the video data processing method provided in the above method embodiments.
[0287] Alternatively, the computer device is a server. For example, Figure 12 This is a structural block diagram of a server provided in an exemplary embodiment of this application.
[0288] Typically, server 2300 includes a processor 2301 and memory 2302.
[0289] Processor 2301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2301 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 2301 may also include a main processor and a coprocessor. The main processor, also known as a central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2301 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 2301 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0290] The memory 2302 may include one or more computer-readable storage media, which may be non-transitory. The memory 2302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2302 are used to store at least one instruction, which is executed by the processor 2301 to implement the video data processing method provided in the method embodiments of this application.
[0291] In some embodiments, the server 2300 may optionally include an input interface 2303 and an output interface 2304. The processor 2301, memory 2302, and input interfaces 2303 and 2304 can be connected via a bus or signal lines. Various peripheral devices can be connected to the input interfaces 2303 and 2304 via a bus, signal lines, or a circuit board. The input interfaces 2303 and 2304 can be used to connect at least one input / output (I / O) related peripheral device to the processor 2301 and memory 2302. In some embodiments, the processor 2301, memory 2302, and input interfaces 2303 and 2304 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2301, memory 2302, and input interfaces 2303 and 2304 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.
[0292] Those skilled in the art will understand that the structure shown above does not constitute a limitation on server 2300, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.
[0293] In an exemplary embodiment, a chip is also provided, the chip including programmable logic circuitry and / or program instructions, which, when the chip is run on a computer device, are used to implement the video data processing method described above.
[0294] In an exemplary embodiment, a computer program product is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the video data processing method provided in the above-described method embodiments.
[0295] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores a computer program and a video stream. When executed by a processor, the computer program implements the aforementioned video data processing method to generate a video stream. Optionally, the computer-readable storage medium may include ROM (Read-Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0296] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0297] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0298] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for processing video data, characterized in that, The method is performed by a computer device, and the method includes: Obtain the color information of the i-th video frame and the (i+1)-th video frame, and obtain the motion information of at least one pixel in the (i+1)-th video frame. The i-th video frame and the (i+1)-th video frame are used to display the virtual environment, where i is a positive integer. Based on the motion information of the at least one pixel, the motion vector of at least one pixel block in the (i+1)th video frame is determined. The at least one pixel includes a second pixel. The motion information of the second pixel is used to indicate the difference between the second position of the second pixel in the (i+1)th video frame and the first position of the first pixel in the i-th video frame. The second pixel and the first pixel correspond to the same position in the virtual environment. Based on the motion vector of the at least one pixel block, color residual information corresponding to the at least one pixel block is determined in the color information of the i-th video frame and the (i+1)-th video frame, respectively. The encoding information of the video data is constructed based on the color information of the i-th video frame, the motion vector of the at least one pixel block, and the color residual information.
2. The method according to claim 1, characterized in that, The first set of pixels is a subset of the at least one pixel, and each pixel in the first set of pixels is a pixel in the first pixel block; determining the motion vector of at least one pixel block in the (i+1)th video frame based on the motion information of the at least one pixel includes: The motion vector of the first pixel block is determined based on the motion information of the pixels in the first pixel set.
3. The method according to claim 2, characterized in that, Determining the motion vector of the first pixel block based on the motion information of the pixels in the first pixel set includes: The average value of the motion information of at least two pixels in the first pixel set is determined as the motion vector of the first pixel block; Alternatively, the median of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block; Alternatively, the modulo value of the motion information of at least two pixels in the first pixel set can be determined as the motion vector of the first pixel block.
4. The method according to claim 2, characterized in that, Determining the motion vector of the first pixel block based on the motion information of the pixels in the first pixel set includes: Based on the motion information of the pixels in the first pixel set, the reference vector of the first pixel block is determined; Based on the first region position of the first pixel block in the (i+1)th video frame and the reference vector, the color sub-information of the reference pixel block is determined in the color information of the i-th video frame; If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block does not exceed a color threshold, the reference vector is determined as the motion vector of the first pixel block.
5. The method according to claim 4, characterized in that, The color threshold is the minimum of at least one reference difference degree; the reference difference degree is used to indicate the difference between the color sub-information of the neighboring pixel blocks around the reference pixel block and the color sub-information of the first pixel block.
6. The method according to claim 4, characterized in that, The method further includes: If the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block exceeds a color threshold, a correction pixel block is determined. The motion vector of the first pixel block is determined based on the positional difference between the first region position and the second region position of the corrected pixel block in the i-th video frame. Wherein, the difference between the color sub-information of the corrected pixel block and the color sub-information of the first pixel block is smaller than the difference between the color sub-information of the reference pixel block and the color sub-information of the first pixel block.
7. The method according to any one of claims 1 to 6, characterized in that, The motion information of the second pixel includes distance sub-information in the first and second directions; Wherein, the first direction and the second direction are the directions of the two vertical sides of the i-th video frame, and the distance sub-information is used to indicate the components between the second position and the first position in the first direction and the second direction.
8. The method according to any one of claims 1 to 6, characterized in that, The motion information of the second pixel includes direction sub-information and distance sub-information; The direction sub-information is used to indicate the direction between the second position and the first position, and the distance sub-information is used to indicate the distance between the second position and the first position.
9. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Perform data normalization and / or coordinate axis transformation on the motion information of the at least one pixel to obtain the corrected motion information of the at least one pixel; The motion vector of the at least one pixel block is determined based on the corrected motion information of the at least one pixel.
10. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining motion information of at least one pixel in the (i+1)th video frame includes: Acquire control information in the virtual environment, the control information being used to control virtual objects and / or virtual cameras in the virtual environment; Based on the control information, the motion information of at least one pixel in the (i+1)th video frame is determined.
11. The method according to any one of claims 1 to 6, characterized in that, Based on the motion vector of the at least one pixel block, the color residual information corresponding to the at least one pixel block is determined from the color information of the i-th video frame and the (i+1)-th video frame, respectively, including: Based on the motion vector of the first pixel block, the color information of the associated pixel block is determined in the color information of the i-th video frame, wherein the at least one pixel block includes the first pixel block; The color residual information of the first pixel block is determined based on the difference between the color information of the associated pixel block and the color information of the first pixel block.
12. A video data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the color information of the i-th video frame and the (i+1)-th video frame, and to acquire the motion information of at least one pixel in the (i+1)-th video frame. The i-th video frame and the (i+1)-th video frame are used to display the virtual environment, where i is a positive integer. The processing module is configured to determine the motion vector of at least one pixel block in the (i+1)th video frame based on the motion information of the at least one pixel, wherein the at least one pixel includes a second pixel, and the motion information of the second pixel is used to indicate the difference between the second position of the second pixel in the (i+1)th video frame and the first position of the first pixel in the i-th video frame, wherein the second pixel and the first pixel correspond to the same position in the virtual environment; The processing module is further configured to determine, based on the motion vector of the at least one pixel block, the color residual information corresponding to the at least one pixel block in the color information of the i-th video frame and the (i+1)-th video frame, respectively; A construction module is used to construct the encoding information of the video data based on the color information of the i-th video frame, the motion vector of the at least one pixel block, and the color residual information.
13. A computer device, characterized in that, The computer device includes: a processor and a memory, wherein the memory stores at least one program; the processor is configured to execute the at least one program in the memory to implement the video data processing method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program and a video stream, the computer program being executed by a processor to implement the video data processing method as described in any one of claims 1 to 11, to generate the video stream.
15. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the video data processing method as described in any one of claims 1 to 11.