Method, apparatus, electronic device and storage medium for video processing
By differentiating between dynamic and static areas in video processing, and utilizing virtual scene models and viewpoint parameters, the problems of network resource waste and latency in traditional video encoding methods are solved, achieving efficient video frame generation and rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VASTAI TECH (SHANGHAI) INC
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional video coding methods cannot fully utilize the static structural information in a scene, resulting in wasted network resources and bandwidth and latency bottlenecks in real-time video streaming.
By decoding the bitstream to determine mask and image information, using a virtual scene model and viewpoint parameters to distinguish between dynamic and static regions, constructing reference video data, and updating the pixel values of dynamic regions based on the mask and image information, video frames are generated.
It reduces the processing and transmission of static data, ensures the accuracy and timeliness of dynamic data, and improves the real-time generation and rendering of video frames.
Smart Images

Figure CN121585886B_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, electronic devices, and storage media for video processing. Background Technology
[0002] With the development of cloud rendering, XR (Extended Reality), and immersive multi-user interactive systems, real-time video streaming is facing bottlenecks in bandwidth and latency. Traditional video coding methods mainly rely on inter-frame prediction and transform coding, which cannot fully utilize the static structural information in the scene. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for video processing is provided. The method includes: determining mask information and image information corresponding to a video frame by decoding a bitstream received from a serving device, the mask information being determined by the serving device based on a difference between rendered video data and reference video data associated with a virtual scene model, the reference video data being determined based on viewpoint parameters of the virtual scene model; acquiring viewpoint parameters of a virtual camera associated with the virtual scene model; constructing reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model; determining a first group of pixels corresponding to dynamic regions and a second group of pixels corresponding to static regions in the reference video data based on the mask information; updating the pixel values of the first group of pixels corresponding to the dynamic regions based on the image information; and generating a video frame based on the second group of pixels and the updated first group of pixels.
[0004] In a second aspect of this disclosure, an apparatus for video processing is provided. The apparatus includes: a stream decoding module, a parameter acquisition module, a data construction module, a region determination module, a pixel update module, and a video generation module. The stream decoding module is configured to determine mask information and image information corresponding to video frames by decoding a stream received from a service device. The mask information is determined by the service device based on the difference between rendered video data and reference video data associated with a virtual scene model. The reference video data is determined based on viewpoint parameters of the virtual scene model. The parameter acquisition module is configured to acquire viewpoint parameters of a virtual camera associated with the virtual scene model. The data construction module is configured to construct reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model. The region determination module is configured to determine, based on the mask information, a first group of pixels corresponding to dynamic regions and a second group of pixels corresponding to static regions in the reference video data. The pixel update module is configured to update the pixel values of the first group of pixels corresponding to dynamic regions based on the image information. The video generation module is configured to generate video frames based on the second group of pixels and the updated first group of pixels.
[0005] In a third aspect of this disclosure, there is an electronic device. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores computer-executable instructions that can be executed by a processing unit to implement the method of the first aspect.
[0007] The solution provided in this disclosure can utilize the virtual scene model and the viewpoint parameters of the virtual camera to construct the pixels corresponding to the static area, reducing the processing and transmission of static data during video processing; then, based on the mask information and image information in the bitstream, the pixels corresponding to the dynamic area are determined and combined with the static pixels, and the data of the dynamic area is superimposed on the static area in real time to generate video frames, which can effectively ensure the accuracy and timeliness of the dynamic data, as well as the real-time generation and rendering of video frames.
[0008] It should be understood that this summary is provided to present, in a simplified form, the selection of concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0009] The above and other objects, features, and advantages of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the exemplary embodiments of this disclosure, the same or similar reference numerals denote the same or similar elements.
[0010] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown;
[0011] Figure 2 A block diagram of an example video encoder according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;
[0014] Figure 5 A flowchart illustrating an example process for processing video according to some embodiments of the present disclosure is shown;
[0015] Figure 6 A flowchart of a bitstream generation process according to some embodiments of the present disclosure is shown;
[0016] Figure 7 A schematic diagram of an example apparatus for processing video according to some embodiments of the present disclosure is shown;
[0017] Figure 8 A block diagram of an electronic device in which various embodiments of the present disclosure may be implemented is shown. Detailed Implementation
[0018] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0019] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0020] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0021] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0023] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0024] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0025] As mentioned above, with the development of cloud rendering, XR (Extended Reality), and immersive multi-user interactive systems, real-time video streaming faces bottlenecks in bandwidth and latency. Traditional video coding methods mainly rely on inter-frame prediction and transform coding, which cannot fully utilize the static structural information in the scene.
[0026] In multi-person or dynamic interactive scenarios, most pixels in a video frame typically belong to the static background, while dynamic content such as people, gestures, and props occupy only a small local area. Traditional video encoding and decoding methods process the entire video frame, repeatedly transmitting static pixels, resulting in significant network resource consumption and waste.
[0027] Embodiments of this disclosure propose a scheme for video processing. The scheme includes: decoding a bitstream received from a service device to determine mask information and image information corresponding to a video frame, wherein the mask information is determined by the service device based on the difference between rendered video data and reference video data associated with a virtual scene model, the reference video data being determined based on viewpoint parameters of the virtual scene model; obtaining viewpoint parameters of a virtual camera associated with the virtual scene model; constructing reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model; determining a first group of pixels corresponding to dynamic regions and a second group of pixels corresponding to static regions in the reference video data based on the mask information; updating the pixel values of the first group of pixels corresponding to dynamic regions based on the image information; and generating a video frame based on the second group of pixels and the updated first group of pixels.
[0028] The embodiments of this disclosure can construct reference video data corresponding to the viewpoint parameters based on the viewpoint parameters associated with the virtual scene model, decode the mask information and image information corresponding to the video frame from the bitstream, then determine the pixels in the reference video data corresponding to the dynamic area and the static area respectively based on the mask information, update the pixels corresponding to the dynamic area using the image information, and then generate a video frame by combining the updated pixels with the pixels corresponding to the static area.
[0029] In this way, the disclosed solution can use the virtual scene model and the viewpoint parameters of the virtual camera to construct the pixels corresponding to the static area, reducing the processing and transmission of static data during video processing; then, based on the mask information and image information in the bitstream, the pixels corresponding to the dynamic area are determined and combined with the static pixels. The data of the dynamic area is superimposed on the static area in real time to generate video frames, which can effectively ensure the accuracy and timeliness of the dynamic data, as well as the real-time generation and rendering of video frames.
[0030] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0031] Example environment:
[0032] Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include encoding devices (e.g., source device 110) and decoding devices (e.g., destination device 120). The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and a first I / O interface 116.
[0033] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0034] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. First I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through first I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0035] The destination device 120 may include a second I / O interface 126, a video decoder 124, and a display device 122. The second I / O interface 126 may include a receiver and / or a modem. The second I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0036] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.
[0037] Figure 2 This is an example block diagram illustrating a video encoder 114 according to some embodiments of the present disclosure. The video encoder 114 can be configured to implement any or all of the technologies disclosed herein.
[0038] exist Figure 2In the example, video encoder 114 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 114. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure. As an example, the processor can include, but is not limited to, an implementable architecture capable of running on heterogeneous platforms such as CPU / GPU / NPU / AI accelerator / hardware decoder.
[0039] In some embodiments, the video encoder 114 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, a first inverse quantization unit 210, a first inverse transform unit 211, a first reconstruction unit 212, a first buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a first motion compensation unit 205, and a first intra-frame prediction unit 206.
[0040] In other examples, the video encoder 114 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0041] Furthermore, although some components (such as motion estimation unit 204 and first motion compensation unit 205) can be integrated, for interpretive purposes, these components are... Figure 2 The examples are shown separately.
[0042] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 114 and the video decoder 124 can support various video block sizes.
[0043] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the first reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0044] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the first buffer 213 with the current video block. First motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from the first buffer 213 other than the image associated with the current video block.
[0045] Motion estimation unit 204 and first motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.
[0046] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. First motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0047] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. First motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0048] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0049] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 124 that the current video block has the same motion information as another video block.
[0050] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 124 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0051] As discussed above, the video encoder 114 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 114 include Advanced Motion Vector Prediction (AMVP) and Merge Mode Signaling.
[0052] The first intra-prediction unit 206 can perform intra-prediction on the current video block. When the first intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0053] The residual generation unit 207 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the current video block.
[0054] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.
[0055] Transform unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0056] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0057] The first inverse quantization unit 210 and the first inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The first reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the first buffer 213.
[0058] After the video block is reconstructed in the first reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0059] Entropy encoding unit 214 can receive data from other functional components of video encoder 114. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0060] Figure 3 This is a block diagram illustrating an example of a video decoder 124 according to some embodiments of the present disclosure. The video decoder 124 may be configured to perform any or all of the techniques of the present disclosure.
[0061] exist Figure 3 In the example, video decoder 124 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 124. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0062] exist Figure 3 In one example, the video decoder 124 includes an entropy decoding unit 301, a second motion compensation unit 302, a second intra-frame prediction unit 303, a second inverse quantization unit 304, a second inverse transform unit 305, a second reconstruction unit 306, and a second buffer 307. In some examples, the video decoder 124 may perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 114.
[0063] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and second motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Second motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0064] The second motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.
[0065] The second motion compensation unit 302 can use the interpolation filter used by the video encoder 114 during the encoding of the video block to calculate the interpolated values of the sub-integer pixels for the reference block. The second motion compensation unit 302 can determine the interpolation filter used by the video encoder 114 based on the received syntax information, and the second motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0066] The second motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0067] The second intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. The second dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The second inverse transform unit 305 applies an inverse transform.
[0068] The second reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the second motion compensation unit 302 or the second intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a second buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0069] As an example, such a display device can be any terminal device with display capabilities, such as AR (Augmented Reality) devices, VR (Virtual Reality) devices, MR (Mixed Reality) devices, and other XR devices.
[0070] Figure 4 This is a schematic diagram illustrating an example environment 400 according to some embodiments of the present disclosure. For example... Figure 4 As shown, example environment 400 includes server 410.
[0071] Server 410 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 410 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0072] In example environment 400, server 410 may include at least an encoding device to process and encode video frames in a video stream to generate a corresponding bitstream. As an example, server 410 may send the bitstream to a terminal device (e.g., at least one electronic device corresponding to at least one user), whereby the terminal device decodes the bitstream, generates video frames, and renders the video frames to a display device.
[0073] refer to Figure 4As shown, such a terminal device may include at least a first electronic device 421 and a second electronic device 422, and such a display device may include a first XR device 441 and a second XR device 442.
[0074] In the example environment 400, multiple users may be included, such as the first user 431 and the second user 432 illustrated in the figure. The first user 431 wears a first XR device 441, and the second user 432 wears a second XR device 442. In some implementations, the first XR device 441 can communicate with a first electronic device 421 to reconstruct a virtual scene for the first user 431 or to merge virtual content with a real scene, for example, to reconstruct and render a virtual scene 450 for the first user 431; the second XR device 442 can communicate with a second electronic device 422 to reconstruct a virtual scene for the second user 432 or to merge virtual content with a real scene, for example, to reconstruct and render a virtual scene 450 for the second user 432. In some implementations, the first user 431 can operate the first electronic device 421, and the second user 432 can operate the second electronic device 422.
[0075] As an example, server 410 can wirelessly communicate with first electronic device 421 and second electronic device 422 to transmit code streams associated with the same virtual scene 450, so that the first virtual scene constructed by first electronic device 421 for first user 431 is the same as or corresponds to the virtual scene constructed by second electronic device 422 for second user 432. Thus, this solution does not restrict multiple users to be in the same physical space.
[0076] In this disclosure, virtual scenes reconstructed based on VR technology, and scenes that integrate virtual content with real-world scenes based on AR or MR technology, are collectively referred to as virtual scene 450. The first XR device 441 and the second XR device 442 (collectively referred to as XR devices) can be head-mounted or wearable near-eye display devices, such as head-mounted displays, smart glasses, etc., supporting VR, AR, MR, and other technologies. Such XR devices may include image generation components and optical display components for reconstructing the virtual scene 450 in a monocular or binocular field of view and displaying virtual objects and / or avatars.
[0077] For example, in a virtual scene 450, a first virtual avatar 451 corresponding to a first user 431 and a second virtual avatar 452 corresponding to a second user 432 can be presented.
[0078] Virtual objects can include three-dimensional virtual objects and / or two-dimensional virtual objects. For example, a two-dimensional virtual object can include a two-dimensional window without thickness, used to display various content in the virtual scene 450, similar to an electronic screen. For example, a display block 453 in the virtual scene 450. The display block 453 can be a window used to load content such as web pages and documents, also known as a "panel". For example, a three-dimensional virtual object can include a part corresponding to a user's virtual avatar, such as a virtual hand. Furthermore, a three-dimensional virtual object can also include virtual devices presented in the virtual scene 450, such as a table or chair.
[0079] In some embodiments, the first process performed by the first electronic device 421 includes, but is not limited to: receiving a bitstream from the server 410, decoding the bitstream and generating a video frame, and then rendering the video frame to the first XR device 441; the second process performed by the second electronic device 422 includes, but is not limited to: receiving a bitstream from the server 410, decoding the bitstream and generating a video frame, and then rendering the video frame to the second XR device 442.
[0080] In this scheme, the first electronic device 421 executes the first process in the same way as the second electronic device 422 executes the second process. Below, this document uses the implementation of the first electronic device 421 executing the first process as an example to illustrate the embodiments of this disclosure.
[0081] In some embodiments, the first electronic device 421 can determine the relative position between one virtual avatar and / or object and another virtual avatar and / or object in a virtual scene 450. The first electronic device 421 can be a separate device capable of communicating with the first XR device 441 and / or other image capture devices, such as a server, computing node, etc., for image or data processing, or it can be integrated with the first XR device 441 and / or other image capture devices. In some embodiments, the first electronic device 421 can be implemented as the first XR device 441, that is, in this case, the first XR device 441 can perform all the functions of the first electronic device 421. It should be understood that the above description of the first electronic device 421 is merely exemplary and not restrictive; the first electronic device 421 can be implemented as a device of various forms, structures, or categories, and the embodiments of this disclosure are not limited in this regard.
[0082] As an example, the first electronic device 421 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the first electronic device 421 may also support any type of user-facing interface (such as "wearable" circuitry).
[0083] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes, but is not limited to: video encoding or compression, video decoding or decompression, video transcoding, and video enhancement (e.g., super-resolution processing) or post-processing, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0084] Example process:
[0085] Figure 5 A flowchart of an example process 500 for processing video according to some embodiments of the present disclosure is shown. Reference is made below. Figure 4 The example environment 400 shown is used to describe process 500. As an example, process 500 may be implemented at a first electronic device 421 and / or a second electronic device 422. The following exemplifies process 500 by taking the implementation of process 500 at the first electronic device 421 as an example.
[0086] refer to Figure 5As shown, in step 510, the first electronic device 421 determines the mask information and image information corresponding to the video frame by decoding the bitstream received from the service device. The mask information is determined by the service device based on the difference between the rendered video data and the reference video data associated with the virtual scene model. The reference video data is determined based on the viewpoint parameters of the virtual scene model.
[0087] As an example, the service device can provide Figure 1 The encoding device in the video encoding / decoding system 100 shown can also be Figure 4 Server 410 in example environment 400 shown.
[0088] The rendered video data can indicate the real-time rendered data associated with the virtual scene 450, and the virtual scene model indicates the static 3D scene model corresponding to the virtual scene 450.
[0089] When the first electronic device 421 acts as the executing entity, the reference video data corresponds to the first perspective of the first user 431 associated with the virtual scene 450.
[0090] Taking the video frame at time t as an example, the first electronic device 421 can determine the mask information and image information corresponding to the video frame by decoding the bitstream. Such mask information and image information can be correlated with the viewpoint parameters corresponding to the first user 431's first viewpoint at time t.
[0091] As an example, such masking information can indicate dynamic and / or static regions associated with the virtual scene 450. Based on the masking information, the first electronic device 421 can determine static and / or dynamic video data in the virtual scene 450 at time t. For example, such static video data can indicate background elements in the virtual scene 450, and such dynamic video data can indicate dynamic objects in the virtual scene 450 (e.g., at least parts of the first virtual avatar 451 and the second virtual avatar 452, such as hands, heads, legs, etc.).
[0092] As an example, such image information can indicate dynamic video data or image data corresponding to a dynamic region. For instance, the image information can exist in the form of a residual (e.g., the pixel difference between a dynamic region and a static region) or in the form of pixel values (e.g., the specific pixel value corresponding to a dynamic region).
[0093] In some embodiments, such mask information and image information may be information that has been compressed and encoded by a service device (e.g., server 410 or encoding device).
[0094] In step 520, the first electronic device 421 acquires the viewpoint parameters of the virtual camera associated with the virtual scene model.
[0095] In some embodiments, the first electronic device 421 may acquire a virtual scene model corresponding to the virtual scene 450 before receiving, decoding, or after receiving the bitstream, and then the first electronic device 421 may acquire the viewpoint parameters of the virtual camera associated with the virtual scene model. As an example, such a virtual camera may correspond to the first viewpoint of the first user 431.
[0096] As an example, viewpoint parameters may include, but are not limited to: the virtual camera's rotation matrix (R), translation vector (T), field of view (FOV), etc.
[0097] In some embodiments, process 500 further includes: the first electronic device 421 acquiring a virtual scene model from scene data maintained locally.
[0098] As an example, before receiving the bitstream, the first electronic device 421 obtains and loads a virtual scene model matching the virtual scene 450 from the scene data maintained locally, based on the operation instructions of the first user 431; alternatively, after receiving the bitstream, it can obtain and load the corresponding virtual scene model from the scene data maintained locally based on the information in the bitstream.
[0099] In some embodiments, the virtual scene model can be a 3D meeting model, a virtual game model, a video viewing scene model, or a virtual interactive model corresponding to a scene such as singing or dancing. As an example, such a virtual scene model can be constructed based on at least one of the following 3D scene modeling methods: NeRF (Neural Radiance Fields), 3DGS (3D Gaussian Splatting), volume rendering, point clouds, meshes, etc. Such a virtual scene model can be compatible with multi-user, multi-view virtual scenes.
[0100] In some embodiments, the virtual scene model is associated with a virtual reality conference scene, and the viewpoint parameters are determined based on the orientation information of the virtual reality device used to present the virtual reality conference scene.
[0101] As an example, such a virtual reality meeting scenario can be referenced. Figure 4 The virtual scene 450 shown can be referenced to the first XR device 441.
[0102] In some embodiments, the first electronic device 421 can also obtain the viewpoint parameters of the virtual camera by decoding the bitstream. These viewpoint parameters are associated with the first viewpoint of the virtual camera in the video frame at time t. For example, for a video frame at time t, the first electronic device 421 can simultaneously obtain the viewpoint parameters, mask information, and image information corresponding to that video frame by decoding the bitstream.
[0103] As an example, in a non-interactive or perspective-following scenario (e.g., the first user 431), the first electronic device 421 can obtain the perspective parameters of the virtual camera by decoding the bitstream.
[0104] In this way, the embodiments of this disclosure can effectively ensure the timeliness consistency between the viewpoint parameters and the mask information and image information in the bitstream, thereby ensuring the synchronization of the picture information and viewpoint in the rendered video frame, avoiding delays, and ensuring the generation and rendering quality of the video frame.
[0105] In some embodiments, if the first electronic device 421 is implemented as a first XR device 441, the first electronic device 421 can directly obtain the viewpoint parameters of the virtual camera from the detection data of the local sensor.
[0106] As an example, the first electronic device 421 can directly read the detection data of its built-in inertial measurement unit, positioning system and other sensors, and calculate the head orientation and position of the first user 431 based on the read data. Then, based on the mapping relationship between the world coordinates of the actual scene and the scene marker of the virtual scene 450, it converts such head orientation and position into viewpoint parameters associated with the virtual scene 450.
[0107] In this way, the embodiments of this disclosure can effectively ensure the synchronization between the viewpoint parameters and the user's real-time actions, and avoid action delay.
[0108] In step 530, the first electronic device 421 constructs reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model.
[0109] In some embodiments, the first electronic device 421 may use viewpoint parameters to render a virtual scene 450 corresponding to the virtual scene model, so that the rendered reference video data matches the first viewpoint corresponding to the visual parameters.
[0110] As an example, such reference video data can be static image data. For instance, such reference video data may not include any dynamic objects (such as the first virtual avatar 451 and the second virtual avatar 452) that participate in the interaction in the virtual scene 450.
[0111] In some embodiments, after loading the virtual scene model, the first electronic device 421 may further initialize its rendering pipeline so as to use the virtual scene model to perform the rendering process in real time, thereby rendering the generated video frames to the corresponding display device (e.g., the first XR device 441).
[0112] As an example, such a rendering pipeline can be an image rendering engine, which includes, but is not limited to, at least one of the following: OpenGL, Vulkan, or a game engine.
[0113] As an example, the first electronic device 421 can use the viewpoint parameters to drive the initialized rendering pipeline, and construct and render reference video data corresponding to the viewpoint parameters based on the loaded virtual scene model.
[0114] In this scheme, the construction and rendering of the reference video data can be completed locally on the first electronic device 421 (e.g., the local graphics processing unit GPU) without consuming network bandwidth, and can achieve extremely high rendering resolution and frame rate, effectively ensuring the rendering quality of the reference video data.
[0115] In step 540, the first electronic device 421 determines, based on the mask information, a first group of pixels corresponding to the dynamic region and a second group of pixels corresponding to the static region in the reference video data.
[0116] Based on the mask information decoded from the bitstream, the first electronic device 421 can divide the reference video data to determine the corresponding dynamic and static regions in the video frame, and then determine the first group of pixels corresponding to the dynamic region and the second group of pixels corresponding to the static region.
[0117] As an example, in response to the completion of the decoding of the bitstream, the first electronic device 421 may first determine whether the dynamic region indicated by the mask information is empty.
[0118] In some embodiments, in response to determining that the dynamic region indicated by the mask information is empty or the image information is empty, the first electronic device 421 may determine that no dynamic data needs to be rendered, and then the first electronic device 421 may generate a video frame based on the reference video data and render the generated video frame to the first XR device 441.
[0119] In this way, the embodiments of this disclosure can quickly generate video frames when the dynamic area indicated by the mask information is empty or the image information is empty, and only a very small amount of control information or viewpoint parameters need to be transmitted over the network, without transmitting the data corresponding to the dynamic area, thereby greatly saving bandwidth and improving video rendering efficiency.
[0120] In some embodiments, in response to determining that the dynamic region indicated by the mask information is not empty or the image information is not empty, the first electronic device 421 may determine the dynamic region and the static region in the video frame based on the mask information; and determine, from the reference video data, a first group of pixels corresponding to the dynamic region and a second group of pixels corresponding to the static region.
[0121] As an example, the dynamic region indicates the area in the video frame where dynamic content in the video is presented in the first viewpoint corresponding to the viewpoint parameter, and the first set of pixels corresponding to the dynamic region can indicate the part of the reference video data that will be replaced or modified by the dynamic image content.
[0122] As an example, a static region can indicate a region in a video frame that does not involve dynamic content, and a second set of pixels corresponding to the static region can indicate the pixels in the reference video data that can remain unchanged relative to the video frame.
[0123] In step 550, the first electronic device 421 updates the pixel values of the first group of pixels corresponding to the dynamic region based on the image information.
[0124] The first electronic device 421 uses the image information decoded from the bitstream to update the pixel value of the first group of pixels corresponding to the dynamic region, thereby obtaining the dynamic video data or image data corresponding to the dynamic region in the video frame, and rendering the dynamic video data or image data corresponding to the dynamic region into the virtual scene 450.
[0125] In some embodiments, the image information may exist in the form of a residual, for example, the image information may indicate the residual value of the first pixel in the first group of pixels. As an example, such a residual value may indicate the difference between the static pixel value corresponding to the first pixel in the reference video data and the dynamic pixel value corresponding to the first pixel in the video frame.
[0126] In this scheme, the first electronic device 421 can determine the first pixel value of the first pixel based on reference video data; and determine the pixel value of the first pixel in the video frame by accumulating the first pixel value and the residual value, and update the pixel value of the first pixel based on the pixel value of the first pixel in the video frame, thereby updating the pixel value of the first group of pixels corresponding to the dynamic area.
[0127] As an example, the first pixel value of the first pixel can indicate the background color corresponding to the virtual scene 450, and the residual value indicated by the image information can indicate the difference between the color of the dynamic object corresponding to the first pixel in the video frame and the background color.
[0128] In this way, the embodiments of this disclosure can accurately update the pixel values of the first group of pixels corresponding to the dynamic region based on the residual values corresponding to each pixel in the first group of pixels, effectively ensuring the accuracy and efficiency of the pixel update corresponding to the dynamic region; moreover, by representing image information in the form of residual values, the amount of image information data can be effectively reduced, thereby reducing the bit stream and the bandwidth occupied by image information, saving resources.
[0129] In some embodiments, image information may exist in the form of pixel values, for example, image information may indicate the second pixel value of a second pixel in a first group of pixels. As an example, such a second pixel value may be the pixel value corresponding to the second pixel in a dynamic image (e.g., the pixel value of any pixel in the first virtual avatar 451 and the second virtual avatar 452).
[0130] In this scheme, the first electronic device 421 can determine the third pixel value of the second pixel from the reference image data; and determine the pixel value of the second pixel in the video frame by fusing the second pixel value and the third pixel value.
[0131] For example, the second pixel value can indicate the color of the dynamic object corresponding to the second pixel, and the third pixel value can indicate the background color corresponding to the second pixel.
[0132] As an example, the first electronic device 421 can merge the second pixel value and the third pixel value by replacing the third pixel value with the second pixel value; it can also perform alpha blending on the second pixel value and the third pixel value to handle semi-transparency, shadows or anti-aliased edges, making the dynamic object blend more naturally with the background.
[0133] In this way, the embodiments of this disclosure can directly fuse the pixel values of the dynamic region in the video frame based on the pixel values corresponding to the dynamic object, thereby effectively ensuring the efficiency of determining the pixel values of the dynamic region, ensuring the video rendering efficiency, and ensuring the pixel quality and rendering quality corresponding to the dynamic region.
[0134] In step 560, the first electronic device 421 generates a video frame based on the second set of pixels and the updated first set of pixels.
[0135] The first electronic device 421 can use the second set of pixels corresponding to the static area and the first set of pixels corresponding to the updated dynamic area to reassemble data to generate a complete video frame. Such a video frame contains both a faithful static scene and real-time dynamic content. The first electronic device 421 can then send or render such a video frame to a display device (e.g., the first XR device 441) to present it to the user, providing the user with an immersive experience associated with the virtual scene 450.
[0136] In this way, the embodiments of this disclosure can use the virtual scene model and the viewpoint parameters of the virtual camera to construct the pixels corresponding to the static area, reducing the processing and transmission of static data during video processing; then, based on the mask information and image information in the bitstream, the pixels corresponding to the dynamic area are determined and combined with the static pixels, and the data of the dynamic area is superimposed on the static area in real time to generate video frames, which can effectively ensure the accuracy and timeliness of the dynamic data, and also ensure the real-time generation and rendering of video frames.
[0137] Figure 6 A flowchart of a stream generation process 600 according to some embodiments of the present disclosure is shown. As an example, the stream generation process 600 may indicate the encoding process of a video. The stream generation process 600 may be implemented at a serving device, such a serving device can be... Figure 1 The encoding device shown can also be Figure 4 Server 410 is shown below. See below for reference. Figure 4 The example environment 400 shown is used as an example to illustrate the bitstream generation process 600, which is implemented at server 410.
[0138] In some embodiments, before executing the bitstream generation process 600, the server 410 may first load the virtual scene model corresponding to the virtual scene 450. As an example, the virtual scene 450 can be any implementable virtual interactive scene, such as a complex game world, a virtual conference room, a virtual movie theater or dance hall, etc.
[0139] As an example, such a virtual scene model can be stored in the storage space of server 410 for the rendering engine to call at any time.
[0140] In some scenarios, server 410 can be equipped with a powerful graphics processing unit (GPU) cluster and configured with a complete graphics rendering pipeline, capable of rendering the final image containing all dynamic elements (such as the first virtual avatar 451, the second virtual avatar 452, special effects, etc.) at high frame rate and high resolution.
[0141] In some scenarios, server 410 can initialize a video encoder and a binary image encoder for compressing the mask, so as to quickly compress the extracted mask information.
[0142] As an example, the video encoder in this solution may include, but is not limited to, any of the following: H.265 / HEVC high-efficiency video encoder, AV1 (AOMedia Video 1) encoder, VVC (H.266 / Versatile Video Coding) encoder, etc.
[0143] In this scheme, server 410 can perform bitstream generation process 600 for each video frame that needs to be distributed to a client (e.g., first electronic device 421 and / or second electronic device 422).
[0144] refer to Figure 6 As shown, in step 610, server 410 determines the viewpoint parameters corresponding to the virtual scene model.
[0145] Taking a video frame at time t as an example, server 410 needs to first determine the viewpoint parameters corresponding to that video frame, such as the viewpoint parameters corresponding to the first user 431's first viewpoint. As an example, such viewpoint parameters can be adapted to the virtual camera corresponding to the first user 431's first viewpoint and associated with the virtual scene 450.
[0146] In some embodiments, the view parameters may include, but are not limited to, the rotation matrix (R), translation vector (T), and field of view (FOV) of the virtual camera.
[0147] In some embodiments, the virtual scene model corresponds to a virtual game scene provided by a cloud gaming service, and determining the viewpoint parameters corresponding to the virtual scene model includes: receiving operation instructions associated with the virtual game scene from a terminal device (e.g., a first electronic device 421 and / or a second electronic device 422) by the service device (server 410); and determining the viewpoint parameters based on the operation instructions.
[0148] As an example, server 410 can receive operation instructions associated with virtual scene 450 by first user 431 in real time via first electronic device 421. Such operation instructions may include input information received via at least one of the following: keyboard keys, mouse movement and / or clicking, gamepad joystick or button, etc.
[0149] In response to receiving an operation command, server 410 can calculate information such as the orientation and position of the virtual camera corresponding to the current video frame based on the operation command. For example, moving the mouse can cause the virtual camera's viewpoint to rotate, and pressing keys such as W / A / S / D can cause the virtual camera's position to translate. Then, server 410 can determine the virtual camera's viewpoint parameters in the current video frame based on the virtual camera's orientation and position information.
[0150] In this way, the embodiments of this disclosure can determine the viewpoint parameters associated with the virtual scene model in the current video frame based on the user's operation instructions in the virtual game scene, thereby effectively ensuring the correlation between the viewpoint parameters and the user's operation, ensuring the real-time performance and accuracy of the viewpoint parameters, and thus ensuring the accuracy and timeliness of the data in the encoded bitstream.
[0151] In some embodiments, taking a VR conference or movie-following scenario as an example, the server 410 can obtain the pose information of the speaker's (e.g., the first user 431) VR device (e.g., the first XR device 441) through the network, and determine the viewpoint parameters corresponding to the virtual scene model in the current video frame based on such pose information.
[0152] In step 620, server 410 determines the reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model.
[0153] Server 410 can construct reference video data corresponding to the acquired viewpoint parameters and the loaded virtual scene model. As an example, such reference video data can indicate static element data in the virtual scene 450, such as the shape, color, and texture of the static elements.
[0154] In some scenarios, such reference video data can indicate static scene images in virtual scene 450 that are "empty" or "time frozen".
[0155] In step 630, server 410 generates rendered video data based on the rendering pipeline.
[0156] In some embodiments, server 410 can use the rendering pipeline to render a complete video frame from the same viewpoint parameters as in step 620, and generate the rendered video data corresponding to that video frame.
[0157] As an example, such a rendering pipeline can handle all scene elements associated with the virtual scene 450, including static models, dynamic characters, user interfaces, lighting effects, visual effects (such as depth of field), etc.
[0158] In step 640, server 410 determines mask information based on the difference between rendered video data and reference video data. The mask information indicates the dynamic regions where differences exist.
[0159] Server 410 can perform pixel-by-pixel comparisons between the rendered video data and reference video data corresponding to the video frame to determine the dynamic region corresponding to the video frame, thereby generating mask information corresponding to the video frame.
[0160] In some embodiments, server 410 may convert the rendered video data and reference video data to the same color space, then calculate the pixel value difference for each pixel, and determine the mask information of the rendered video data relative to the reference video data based on such differences.
[0161] As an example, such a color space could be the YUV color space, focusing on the difference in the luminance Y component between the rendered video data and the reference video data. This difference could be the absolute value of the difference between the two sets of data.
[0162] As an example, server 410 can compare the determined pixel value difference with a preset difference threshold. If the pixel value difference of a certain pixel is greater than the difference threshold, server 410 can determine that the pixel is a pixel corresponding to a dynamic region; otherwise, it is a pixel of a static region. Thus, server 410 can accurately distinguish between static and dynamic regions in the rendered video data and generate mask information corresponding to the dynamic region.
[0163] In some scenarios, server 410 can also generate mask information maps (e.g., binary mask maps) corresponding to dynamic and static regions. For example, a pixel with a value of 1 indicates that the pixel is a dynamic region, and a pixel with a value of 0 indicates that the pixel is a static region. Then, server 410 can perform at least one of the following processing operations on the generated binary mask map: erosion, dilation, opening, closing, etc., to remove noise points, fill small holes, smooth region boundaries, etc., thereby obtaining a more accurate dynamic region contour and further ensuring the accuracy and effectiveness of the mask information.
[0164] In step 650, server 410 determines the image information corresponding to the dynamic region based on the rendered video data.
[0165] Server 410 can determine the image information corresponding to the mask information based on the rendered video data and mask information, so as to use the image information corresponding to the dynamic region.
[0166] In some embodiments, server 410 may determine a set of pixel values corresponding to the mask information from the rendered video data based on the mask information, and use this set of pixel values as the image information corresponding to the dynamic region.
[0167] In some embodiments, server 410 may determine the residual values of a first group of pixels corresponding to a dynamic region based on the difference between rendered video data and reference video data, as image information corresponding to the dynamic region.
[0168] As an example, to reduce computation, server 410 can extract the first set of pixel values corresponding to the mask information from the rendered video data and the second set of pixel values corresponding to the mask information from the reference video data based on the mask information. Then, it can calculate the difference between the first set of pixel values and the second set of pixel values one by one to obtain the residual value of the first set of pixels corresponding to the dynamic region, and use it as the image information corresponding to the dynamic region.
[0169] In this way, embodiments of the present disclosure can determine the image information corresponding to the dynamic region based on the difference between the rendered video data and the reference video data, thereby effectively ensuring the accuracy and reliability of the image information.
[0170] In step 660, server 410 generates a bitstream based on mask information and image information.
[0171] As an example, server 410 compresses the mask information or binary mask image obtained in step S640 (e.g., using a standard video encoder) and the image information obtained in step S650 (e.g., using a binary image encoder). Then, server 410 can encapsulate the compressed mask information and compressed image information into a data packet or bitstream segment corresponding to the current video frame according to a preset format. Thus, for multiple video frames, server 410 can generate corresponding bitstreams.
[0172] In some embodiments, server 410 can generate a bitstream segment corresponding to the current video frame based on the viewpoint parameters, mask information, and image information corresponding to the current video frame.
[0173] As an example, server 410 can serialize or differentially encode the viewpoint parameters corresponding to the current video frame, and then encapsulate the processed viewpoint parameters with compressed mask information and image information to generate a data packet or bitstream segment corresponding to the current video frame.
[0174] In some embodiments, if the mask information indicates that the dynamic region is empty, the bitstream may not include the mask information and image information. That is, the server 410 can generate the corresponding bitstream segment based on the viewpoint parameters or empty frame identifier corresponding to the current video frame. In this way, the embodiments of this disclosure can effectively reduce the amount of bitstream data corresponding to static video frames, thereby saving bandwidth.
[0175] This embodiment of the disclosure makes full use of the reference video data corresponding to the virtual scene model and the rendered video data corresponding to the video frame. By using differential technology, it effectively identifies the mask information and image information corresponding to the dynamic region, thereby reducing the processing and transmission of static pixels in the video, effectively reducing the data volume of the bitstream, realizing the separate transmission between static pixels and dynamic pixels, and saving network resources.
[0176] Example devices and equipment:
[0177] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 7A schematic structural block diagram of an apparatus 700 for processing video according to certain embodiments of the present disclosure is shown. As an example, apparatus 700 may be implemented as a first electronic device 421 or a second electronic device 422, or as a combined device including the first electronic device 421 and / or the second electronic device 422. The various modules / components in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0178] like Figure 7 As shown, the device 700 for processing video includes: a stream decoding module 710, a parameter acquisition module 720, a data construction module 730, a region determination module 740, a pixel update module 750, and a video generation module 760. The stream decoding module 710 is configured to determine mask information and image information corresponding to video frames by decoding the stream received from the service device. The mask information is determined by the service device based on the difference between rendered video data and reference video data associated with a virtual scene model. The reference video data is determined based on the viewpoint parameters of the virtual scene model. The parameter acquisition module 720 is... The system is configured to: acquire the viewpoint parameters of a virtual camera associated with a virtual scene model; construct reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model; determine the first group of pixels corresponding to the dynamic region and the second group of pixels corresponding to the static region in the reference video data based on mask information; update the pixel values of the first group of pixels corresponding to the dynamic region based on image information; and generate a video frame based on the second group of pixels and the updated first group of pixels.
[0179] In some embodiments, the region determination module 740 is configured to: determine a dynamic region and a static region based on the mask information in response to determining that the dynamic region indicated by the mask information is not empty or the image information is not empty; and determine a first group of pixels corresponding to the dynamic region and a second group of pixels corresponding to the static region from the reference video data.
[0180] In some embodiments, the apparatus 700 further includes a module that performs the following process: in response to determining that the dynamic region indicated by the mask information is empty or the image information is empty, generating a video frame based on reference video data.
[0181] In some embodiments, image information indicates the residual value of a first pixel in a first group of pixels, and the pixel update module 750 is configured to: determine the first pixel value of the first pixel based on reference video data; and determine the pixel value of the first pixel in a video frame by accumulating the first pixel value and the residual value.
[0182] In some embodiments, image information indicates the second pixel value of the second pixel in the first group of pixels, and the pixel update module 750 is configured to: determine the third pixel value of the second pixel based on reference video data; and determine the pixel value of the second pixel in the video frame by fusing the second pixel value and the third pixel value.
[0183] In some embodiments, the bitstream is generated by the service device based on the following process: determining viewpoint parameters corresponding to a virtual scene model; determining reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model; generating rendered video data based on the rendering pipeline; determining mask information based on the differences between the rendered video data and the reference video data, the mask information indicating dynamic regions where differences exist; determining image information corresponding to the dynamic regions based on the rendered video data; and generating the bitstream based on the mask information and the image information.
[0184] In some embodiments, determining the image information corresponding to the dynamic region based on the rendered video data includes: determining the residual values of a first group of pixels corresponding to the dynamic region based on the difference between the rendered video data and the reference video data, as the image information corresponding to the dynamic region.
[0185] In some embodiments, the virtual scene model corresponds to a virtual game scene provided by a cloud gaming service, and determining the viewpoint parameters corresponding to the virtual scene model includes: receiving operation instructions associated with the virtual game scene from the terminal device by the service device; and determining the viewpoint parameters based on the operation instructions.
[0186] In some embodiments, the apparatus 700 further includes a model acquisition module configured to acquire a virtual scene model from locally maintained scene data.
[0187] In some embodiments, the virtual scene model is associated with a virtual reality conference scene, and the viewpoint parameters are determined based on the orientation information of the virtual reality device used to present the virtual reality conference scene.
[0188] In some embodiments, the parameter acquisition module 720 is configured to acquire the viewpoint parameters of the virtual camera by decoding the bitstream.
[0189] Figure 8 A block diagram of a computing device 800 in which various embodiments of the present disclosure may be implemented is shown. The computing device 800 may be implemented as a first electronic device 421 or a second electronic device 422, or may be implemented as a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124), or may be included in a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124).
[0190] It should be understood that, Figure 8 The computing device 800 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0191] like Figure 8 As shown, the computing device 800 includes a general-purpose computing device 800. The computing device 800 may include at least one or more processors or processing units 810, memory 820, storage units 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.
[0192] In some embodiments, the computing device 800 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 800 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0193] Processing unit 810 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 800. Processing unit 810 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0194] Computing device 800 typically includes various computer storage media. Such media can be any media accessible by computing device 800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 830 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 800.
[0195] The computing device 800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 8 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0196] The communication unit 840 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 800 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0197] Input device 850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 860 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 840, computing device 800 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 800 can also communicate with one or more devices that enable a user to interact with computing device 800, or, if needed, with any device (e.g., network card, modem, etc.) that enables computing device 800 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).
[0198] In some embodiments, some or all of the components of computing device 800 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0199] In embodiments of this disclosure, computing device 800 may be used to implement video encoding / decoding. Memory 820 may include one or more video codec modules 825 having one or more program instructions. These modules are accessible and executable by processing unit 810 to perform the functions of the various embodiments described herein.
[0200] In an example embodiment of performing video encoding, input device 850 may receive video data as input 870 to be encoded. The video data may be processed, for example, by video codec module 825 to generate an encoded bitstream. The encoded bitstream may be provided as output 880 via output device 860.
[0201] In an example embodiment of performing video decoding, input device 850 may receive an encoded bitstream as input 870. The encoded bitstream may be processed, for example, by video codec module 825 to generate decoded video data. The decoded video data may be provided as output 880 via output device 860.
[0202] While this disclosure has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. Such variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, characterized in that, The method includes: By decoding the bitstream received from the service device, mask information and image information corresponding to the video frame are determined. The mask information is determined by the service device based on the difference between the rendered video data and the first reference video data associated with the virtual scene model. The first reference video data is determined by the service device based on the viewpoint parameters of the virtual scene model. The image information is determined by the service device based on the rendered video data. The image information indicates the dynamic video data or image data corresponding to the dynamic region in the video frame. Obtain the viewpoint parameters of the virtual camera associated with the virtual scene model; Based on the viewpoint parameters and the virtual scene model, construct second reference video data corresponding to the viewpoint parameters; Based on the mask information, determine the first group of pixels corresponding to the dynamic region and the second group of pixels corresponding to the static region in the second reference video data; Based on the image information, update the pixel values of the first group of pixels corresponding to the dynamic region; and The video frame is generated based on the second set of pixels and the updated first set of pixels.
2. The method according to claim 1, characterized in that, The step of determining the first group of pixels corresponding to the dynamic region and the second group of pixels corresponding to the static region in the second reference video data based on the mask information includes: In response to determining that the dynamic region indicated by the mask information is not empty or the image information is not empty, the dynamic region and the static region are determined based on the mask information; and From the second reference video data, determine the first group of pixels corresponding to the dynamic region and the second group of pixels corresponding to the static region.
3. The method according to claim 2, characterized in that, The method further includes: In response to determining that the dynamic region indicated by the mask information is empty or the image information is empty, the video frame is generated based on the second reference video data.
4. The method according to claim 1, characterized in that, The image information indicates the residual value of the first pixel in the first group of pixels, and updating the first group of pixels corresponding to the dynamic region based on the image information includes: Based on the second reference video data, determine the first pixel value of the first pixel; and The pixel value of the first pixel in the video frame is determined by summing the first pixel value and the residual value.
5. The method according to claim 1, characterized in that, The image information indicates the second pixel value of the second pixel in the first group of pixels, and updating the first group of pixels corresponding to the dynamic region based on the image information includes: Based on the second reference video data, determine the third pixel value of the second pixel; and The pixel value of the second pixel in the video frame is determined by fusing the second pixel value and the third pixel value.
6. The method according to claim 1, characterized in that, The bitstream is generated by the service device based on the following process: Determine the viewpoint parameters corresponding to the virtual scene model; Based on the viewpoint parameters and the virtual scene model, determine the first reference video data corresponding to the viewpoint parameters; The rendered video data is generated based on the rendering pipeline; Based on the difference between the rendered video data and the first reference video data, the mask information is determined, and the mask information indicates the dynamic regions where the differences exist; Based on the rendered video data, determine the image information corresponding to the dynamic region; as well as The bitstream is generated based on the mask information and the image information.
7. The method according to claim 6, characterized in that, The step of determining the image information corresponding to the dynamic region based on the rendered video data includes: Based on the difference between the rendered video data and the first reference video data, the residual value of the first group of pixels corresponding to the dynamic region is determined as the image information corresponding to the dynamic region.
8. The method according to claim 6, characterized in that, The virtual scene model corresponds to a virtual game scene provided by a cloud gaming service, and determining the viewpoint parameters corresponding to the virtual scene model includes: The service device receives operation instructions associated with the virtual game scene from the terminal device; and The viewpoint parameters are determined based on the operation instructions.
9. The method according to claim 1, characterized in that, The method further includes: The virtual scene model is obtained from the scene data maintained locally.
10. The method according to claim 9, characterized in that, The virtual scene model is associated with the virtual reality conference scene, and the viewpoint parameters are determined based on the orientation information of the virtual reality device, which is used to present the virtual reality conference scene.
11. The method according to claim 1, characterized in that, The acquisition of the viewpoint parameters of the virtual camera associated with the virtual scene includes: The viewpoint parameters of the virtual camera are obtained by decoding the bitstream.
12. An apparatus for video processing, characterized in that, The device includes: The bitstream decoding module is configured to determine mask information and image information corresponding to video frames by decoding the bitstream received from the service device. The mask information is determined by the service device based on the difference between rendered video data and first reference video data associated with a virtual scene model. The first reference video data is determined by the service device based on the viewpoint parameters of the virtual scene model. The image information is determined by the service device based on the rendered video data. The image information indicates dynamic video data or image data corresponding to dynamic regions in the video frame. The parameter acquisition module is configured to acquire the viewpoint parameters of the virtual camera associated with the virtual scene model; The data construction module is configured to construct second reference video data corresponding to the viewpoint parameters based on the viewpoint parameters and the virtual scene model; The region determination module is configured to determine, based on the mask information, a first group of pixels corresponding to the dynamic region and a second group of pixels corresponding to the static region in the second reference video data; A pixel update module is configured to update the pixel values of the first group of pixels corresponding to the dynamic region based on the image information; and The video generation module is configured to generate the video frame based on the second set of pixels and the updated first set of pixels.
13. An electronic device, characterized in that, The electronic device includes: At least one processing unit; and At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer-executable instructions thereon, characterized in that, The computer-executable instructions can be executed by a processing unit to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video generation method and device in three-dimensional scene, equipment and storage medium
CN120201258A
Method and apparatus for encoding and decoding moving image
JP2010011075A