Video processing method and device

By using inter-prediction frames as reference frames, the problem of high view angle switching delay of VR devices is solved, and more efficient view angle switching and user experience are achieved.

CN115499634BActive Publication Date: 2025-08-29HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110678265.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-18
Publication Date
2025-08-29
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

In VR devices, the problem of high viewing angle switching delay is mainly because VR devices need to download large data intra-coded frames (I frames) to display high-quality panoramic video content, resulting in a long viewing angle switching delay.

Method used

The transmitting end sends inter-prediction frames as reference frames, and the receiving end acquires the enhancement layer image based on the basic layer image, reducing the data transmission amount and improving the decoding efficiency, and reducing the viewing angle switching delay.

Benefits of technology

It reduces the delay of viewing angle switching, improves the efficiency and user experience of viewing angle switching, and reduces the need for storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115499634B_ABST
    Figure CN115499634B_ABST
Patent Text Reader

Abstract

The present application provides a video processing method and device, which relates to the field of video processing technology. The method includes: a transmitting end sends a first code stream of a source video to a receiving end, and the receiving end obtains a basic layer image of the source video based on the first code stream; the receiving end obtains perspective switching information, and obtains the content information to be displayed of the source video based on the perspective switching information, and then the receiving end decodes the content information to be displayed based on the basic layer image to obtain an enhanced layer image of the source video, and the video quality of the enhanced layer image is higher than the video quality of the basic layer image. The reference frame included in the above-mentioned content information to be displayed is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to a partial perspective area after the perspective is switched. When the perspective range is switched, the reference frame is an inter-frame prediction frame, and the data volume of the inter-frame prediction frame is less than the data volume of the I frame. Therefore, the transmission delay of the reference frame is reduced, thereby reducing the perspective switching delay of the receiving end and improving the perspective switching efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a video processing method and device. Background Art

[0002] 360° video is a video captured by multiple cameras in 360°. Also known as panoramic video, it provides users with an immersive experience from all angles. For example, in a virtual reality (VR) scenario, a VR device displays high-quality footage from the user's perspective and lower-quality footage from outside the user's perspective.

[0003] At present, during the user's perspective switching process, the VR device obtains the code stream and random access reference frame (RARF) of the high-quality panoramic video content within the switched perspective range from the server, so that the VR device decodes the code stream based on the RARF frame and displays the high-quality panoramic video content within the switched perspective range. However, the RARF frame is an intra-picture (I frame), and the I frame includes all the information of the first frame of the image that the VR device needs to display within the switched perspective range. Due to the large amount of data in the RARF frame, the VR device is exposed to the low-quality panoramic video content within the switched perspective range for too long, and the perspective switching delay is high. Therefore, how to reduce the perspective switching delay has become a problem that needs to be solved urgently. Summary of the Invention

[0004] The present application provides a video processing method and device, which solves the problem of high view switching delay.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions.

[0006] In a first aspect, the present application provides a video processing method that can be applied to a receiving end, or the method can be applied to an apparatus that can support a terminal device to implement the method, for example, the apparatus includes a chip system, and the method includes: a transmitting end sends a first code stream of a source video to a receiving end, and the receiving end obtains a base layer image of the source video based on the first code stream; the receiving end also obtains perspective switching information, and obtains content information to be displayed of the source video based on the perspective switching information, and then the receiving end decodes the content information to be displayed based on the base layer image to obtain an enhanced layer image of the source video, the video quality of the enhanced layer image being higher than the video quality of the base layer image. The above-mentioned content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to a partial perspective area after the perspective switching.

[0007] In the video processing method provided in the embodiment of the present application, when the user's viewing angle range is switched at the receiving end, the reference frame obtained from the sending end is an inter-frame prediction frame. Since the data amount of the inter-frame prediction frame is smaller than the data amount of the I frame and the independently decodable frame, the transmission delay of the reference frame is reduced, thereby reducing the viewing angle switching delay at the receiving end and improving the viewing angle switching efficiency.

[0008] In one possible example, the aforementioned video quality includes at least one or a combination of image signal-to-noise ratio, resolution, and frame rate. For example, if video quality is expressed in terms of resolution, the resolution of the base layer image is lower than that of the enhancement layer image. Therefore, the bitrate of the first bitstream corresponding to the base layer image is lower than the bitrate of the content information to be displayed corresponding to the enhancement layer image. The receiving end can first display the base layer image based on the first bitstream with the lower bitrate. This reduces transmission latency at the receiving end while ensuring that the user can view the image, thereby improving the user experience.

[0009] In one possible implementation, the aforementioned content information to be displayed also includes at least one second bitstream. The receiving end decodes the content information to be displayed based on the base layer image to obtain the enhancement layer image of the source video, including: the receiving end obtains a first image of the partial viewing area after the perspective switch based on the base layer image, wherein the timestamp of the first image matches the timestamp of the reference frame; the receiving end obtains a second image based on the first image and the reference frame, wherein the second image corresponds to the partial viewing area after the perspective switch in the enhancement layer image; and further, the receiving end obtains the enhancement layer image based on the second image and the at least one second bitstream. In this manner, since the reference frame is an inter-frame predicted frame, the data volume of the reference frame is smaller than that of an I-frame. Therefore, the transmission delay of the reference frame between the transmitting and receiving ends is reduced, and the perspective switch delay at the receiving end is reduced. In addition, during the decoding process of the enhancement layer image, the receiving end can reuse the content of the first bitstream to decode the reference frame, thereby reducing the amount of video data required to be stored by the receiving end and conserving storage resources at the receiving end.

[0010] In another possible implementation, the receiving end obtains the second image based on the first image and the reference frame, including: the receiving end encodes the first image to obtain a video frame, which is an intra-frame coded frame; and then, the receiving end decodes the reference frame based on the video frame to obtain the second image. It is worth noting that the video frame can also be other independently decodable frames. In an embodiment of the present application, the receiving end can encode the first image into a video frame, and connect the video frame with the reference frame to obtain a code stream to be decoded. The decoder in the receiving end decodes the code stream to be decoded to obtain the second image, thereby avoiding the process of the receiving end modifying the reference image in the decoder to be the first image, improving the decoding efficiency of the receiving end, and reducing the perspective switching delay of the receiving end.

[0011] In another possible implementation, the receiving end obtains the enhancement layer image based on the second image and at least one second codestream, including: the receiving end splicing the at least one second codestream to obtain an application codestream, and obtaining the enhancement layer image based on the second image and the application codestream. When multiple second codestreams are spliced ​​into one application codestream, the receiving end can use only one decoder to decode the multiple second codestreams, thereby saving processing resources required for decoding at the receiving end, reducing decoding latency at the receiving end, and improving decoding efficiency at the receiving end.

[0012] In another possible implementation, the receiving end obtains the content information to be displayed of the source video based on the perspective switching information, including: the receiving end obtains the partial perspective area after the perspective switching based on the current perspective information and the perspective switching information; the receiving end sends a first instruction including an identifier, which is used to indicate the position of the partial perspective area after the perspective switching in the basic layer image; finally, the receiving end obtains a reference frame in response to the first instruction.

[0013] In another possible implementation, the receiving end obtains a partial viewing area after the viewing angle is switched based on the current viewing angle information and the viewing angle switching information, including: the receiving end determines multiple first sub-area identifiers based on the current viewing angle information and a preset first relationship, where the first relationship is used to indicate sub-area division information and the identifier of each sub-area in the source video; the receiving end determines multiple second sub-area identifiers based on the viewing angle switching information and the first relationship; and further, the receiving end obtains at least one sub-area identifier after the viewing angle is switched based on the multiple first sub-area identifiers and the multiple second sub-area identifiers, where the sub-area indicated by the at least one sub-area identifier after the viewing angle is the partial viewing area after the viewing angle is switched. In an embodiment of the present application, the receiving end can detect a change in the user's viewing angle and obtain new viewing angle information and corresponding video sub-area information. In this way, the receiving end can obtain content information to be displayed from the sending end based on the new viewing angle information and the corresponding video sub-area information, avoiding the sending end from processing the user's viewing angle switching information and saving processing resources on the sending end.

[0014] In another possible implementation, the receiving end may also display the base layer image and / or the enhancement layer image. In the embodiments of the present application, the receiving end may display at least one of the base layer image and the enhancement layer image, allowing the user to view the source video content through the receiving end, avoiding the need to reconnect to the display device to display the above-mentioned images, reducing the view switching delay, and improving the user experience.

[0015] In another possible implementation, the receiving end obtains the base layer image of the video based on the first bitstream, including: the receiving end obtains the first bitstream from the transmitting end and decodes the first bitstream to obtain the base layer image. It is worth noting that the receiving end may obtain the first bitstream by requesting and downloading it from the transmitting end, or by actively sending it to the receiving end.

[0016] In the second aspect, the present application provides a video processing method, which can be applied to a sending end, or the method can be applied to a device that can support a terminal device to implement the method, for example, the device includes a chip system, and the method includes: the sending end sends a first code stream of the source video to the receiving end, and the first code stream is used to indicate the basic layer image of the source video; the sending end obtains perspective switching information, and sends content information to be displayed of the source video to the receiving end according to the perspective switching information, and the content information to be displayed is used to indicate the enhanced layer image corresponding to the perspective area after the perspective switching in the source video, and the video quality of the enhanced layer image is higher than the video quality of the basic layer image. The above-mentioned content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to part of the perspective area after the perspective switching. In the video processing method provided in the embodiment of the present application, when the user's viewing angle range is switched at the receiving end, the transmitting end sends the content information to be displayed of the source video to the receiving end. The reference frame included in the content information to be displayed is an inter-frame prediction frame. Since the data amount of the inter-frame prediction frame is smaller than the data amount of the I frame and the independently decodable frame, the transmission delay of the reference frame is reduced, thereby reducing the viewing angle switching delay at the receiving end and improving the viewing angle switching efficiency.

[0017] In a possible implementation, the to-be-displayed content information further includes at least one second code stream, and the at least one second code stream corresponds to other viewing angle areas in the enhancement layer image except for the partial viewing angle area after the viewing angle is switched.

[0018] In another possible implementation, before the transmitter sends the source video content information to the receiver, the method further includes: the transmitter obtaining a third bitstream representing the partial viewing area after the perspective switch based on the source video, and obtaining a reference frame based on the first bitstream and the third bitstream. For example, in a live broadcast scenario, the transmitter can generate the reference frame for the content information to be displayed from the source video based on the perspective switch information sent by the receiver, thereby enabling rapid perspective switching at the receiver and improving the user experience.

[0019] In another possible implementation, the transmitter obtains a reference frame based on the first and third bitstreams, including: the transmitter obtains a first image of the partial viewing area after the perspective switch based on the first bitstream, and the timestamp of the first image matches the timestamp of the perspective switch; then, the transmitter encodes a second image based on the first image to obtain a reference frame, where the second image is an image of the partial viewing area in the source video, or an image obtained by decoding the third bitstream, and the timestamp of the second image matches the timestamp of the first image. It is worth noting that in the embodiments of the present application, the generation of the reference frame is determined by the transmitter based on the partial viewing area after the perspective switch. In other words, the transmitter can generate the reference frame in real time based on the perspective change of the receiver, effectively reducing the data volume of the reference frame and improving the user experience.

[0020] In another possible implementation, the transmitting end encodes the second image based on the first image to obtain a reference frame, including: the transmitting end encodes the first image to obtain a video frame, and the video frame is an intra-frame coded frame; the transmitting end also encodes the second image based on the video frame to obtain a reference frame. It is worth noting that the video frame can also be an independently decodable frame. The transmitting end can encode the first image into a video frame, and encode the second image based on the video frame to obtain a reference frame. Therefore, during the decoding process, the receiving end can connect the video frame with the reference frame to obtain a connected code stream, and the decoder in the receiving end decodes the connected code stream to obtain the second image, thereby avoiding the process of the receiving end modifying the reference image in the decoder to the first image, improving the decoding efficiency of the receiving end, and reducing the perspective switching delay of the receiving end.

[0021] In another possible implementation, the sending end obtains a third code stream of the partial viewing area after the viewing angle is switched based on the source video, including: the sending end divides the source video into multiple sub-areas, and encodes the video corresponding to each sub-area in the multiple sub-areas to obtain multiple sub-area code streams; the sending end also matches the partial viewing area after the viewing angle is switched with a preset first relationship to determine multiple sub-area identifiers, where the first relationship is used to indicate the sub-area division information and the identifier of each sub-area in the source video, and the sub-area code stream corresponding to the multiple sub-area identifiers is the third code stream.

[0022] In another possible implementation, the video processing method further includes: a transmitter encoding the video corresponding to each of the multiple sub-regions in the source video to obtain multiple sub-region code streams; the transmitter also obtaining a sub-region reference image sequence within a first time range based on the first code stream, wherein the sub-region reference image sequence includes a reference image for each sub-region; and further, the transmitter encoding the image sequence to be encoded based on the sub-region reference image sequence to obtain a reference code stream. The reference code stream includes a reference frame, and the image sequence to be encoded includes multiple images to be encoded. The images to be encoded are sub-region images in the source video that match the reference images, or images obtained by decoding a code stream in the sub-region code stream that matches the reference images, and the timestamps of the images to be encoded match those of the reference images. In an embodiment of the present application, the transmitting end can generate a reference frame corresponding to each timestamp before the user's viewing angle range is switched. After the user's viewing angle at the receiving end is switched, the transmitting end can obtain the content information to be displayed in the viewing area after the viewing angle switching from the reference code stream based on the timestamp of the viewing angle switching. Since the reference frame is an inter-frame prediction frame, the data volume of the reference frame is small and the transmission delay is low. In addition, since the transmitting end does not need to generate the reference frame immediately, this reduces the acquisition time of the content information to be displayed and reduces the viewing angle switching delay.

[0023] In another possible implementation, the transmitting end encodes the image sequence to be coded based on the sub-region reference image sequence to obtain a reference code stream, including: the transmitting end interleaves and recombines the sub-region reference image sequence and the image sequence to be coded to obtain first recombined data and second recombined data, the first recombined data including the sub-region reference image and the image to be coded whose sub-region identifier meets the first condition, and the second recombined data including the sub-region reference image and the image to be coded whose sub-region identifier meets the second condition.

[0024] Furthermore, the transmitter obtains a fourth code stream based on the first reconstructed data and a fifth code stream based on the second reconstructed data, wherein the fourth code stream includes all reference frames whose sub-region identifiers meet the first condition within the first time range, and the fifth code stream includes all reference frames whose sub-region identifiers meet the second condition within the first time range. Finally, the transmitter obtains a reference code stream based on the fourth and fifth code streams. In an embodiment of the present application, the transmitter rearranges the sub-region reference image sequence and the image sequence to be encoded within the first time range in a frame-number-interleaved manner into two image sequences (first reconstructed data and second reconstructed data). In each reconstructed data, the sub-region reference image at the same time is located before the image to be encoded. After the transmitter obtains the two image sequences, it can encode the two image sequences using a low-latency encoding configuration with one key frame every two frames to obtain a reference code stream within the first time range. Because the header information of the code stream of each reference frame in the code stream obtained after the transmitter encodes the reconstructed data is consistent with the frame number of the image to be encoded, the transmitter does not need to modify the frame number of each reference frame in the reference code stream, thereby improving the encoding efficiency of the transmitter.

[0025] In another possible implementation, before the transmitter sends the first bitstream to the receiver, the video processing method further includes: the transmitter downsampling the source video to obtain a base layer image, and encoding the base layer image to obtain the first bitstream; wherein the video quality of the source video is higher than the video quality of the base layer image. Downsampling the source video before encoding by the transmitter can reduce the data volume of the first bitstream, reduce the transmission delay of the first bitstream, and improve the user experience.

[0026] In another possible implementation, the above-mentioned video quality includes at least one of image signal-to-noise ratio, resolution, and frame rate, or a combination of several of them.

[0027] In a third aspect, the present application provides a video processing device, which includes various modules for executing the video processing method in the first aspect or any possible implementation of the first aspect.

[0028] The beneficial effects can be found in the description of any aspect of the first aspect, which will not be repeated here. The video processing device has the function of implementing the behavior in the method example of any aspect of the first aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the video processing device includes: a processing module for obtaining a basic layer image of the source video based on the first bit stream; a communication module for obtaining perspective switching information, and obtaining content information to be displayed of the source video according to the perspective switching information, the content information to be displayed includes a reference frame, the reference frame is an inter-frame prediction frame obtained based on the first bit stream, and the reference frame corresponds to a partial perspective area after the perspective switching; the processing module is also used to decode the content information to be displayed based on the basic layer image to obtain an enhanced layer image of the source video, and the video quality of the enhanced layer image is higher than the video quality of the basic layer image.

[0029] In a fourth aspect, the present application provides a video processing device, which includes various modules for executing the video processing method in the second aspect or any possible implementation of the second aspect.

[0030] The beneficial effects can be found in the description of any aspect in the second aspect, which will not be repeated here. The video processing device has the function of implementing the behavior in the method instance of any aspect in the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the video processing device includes: a communication module for sending a first code stream of the source video to the receiving end, the first code stream is used to indicate the basic layer image of the source video; a processing module for obtaining perspective switching information, and sending the content information to be displayed of the source video to the receiving end according to the perspective switching information; wherein the content information to be displayed is used to indicate the enhanced layer image corresponding to the perspective area after the perspective switching in the source video, the video quality of the enhanced layer image is higher than the video quality of the basic layer image, and the content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to a part of the perspective area after the perspective switching.

[0031] In a fifth aspect, the present application provides a video processing system, which includes modules for executing the video processing method in the first aspect or any possible implementation of the first aspect, and the second aspect or any possible implementation of the second aspect.

[0032] The beneficial effects can be found in the description of either the first aspect or the second aspect, and will not be repeated here. The video processing system has the function of implementing the methods in the first aspect and the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the video processing system includes: a transmitting end and a receiving end. The transmitting end sends a first code stream of the source video to the receiving end, and the first code stream is used to indicate the basic layer image of the source video; the receiving end obtains the basic layer image of the source video based on the first code stream. The receiving end obtains the perspective switching information, and obtains the content information to be displayed of the source video from the receiving end according to the perspective switching information; the content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to the partial perspective area after the perspective switching; the receiving end decodes the content information to be displayed based on the basic layer image, and obtains the enhanced layer image of the source video, and the video quality of the enhanced layer image is higher than the video quality of the basic layer image.

[0033] In a sixth aspect, the present application provides a terminal device comprising a processor and an interface circuit, the interface circuit being used to receive signals from terminal devices other than the terminal device and transmit them to the processor, or to send signals from the processor to terminal devices other than the terminal device, the processor being used to implement the operating steps of the method of the first aspect and any possible implementation method of the first aspect through a logic circuit or executing code instructions, or being used to implement the operating steps of the method of the second aspect and any possible implementation method of the second aspect.

[0034] In a seventh aspect, the present application provides a computer-readable storage medium, which stores a computer program or instructions. When the computer program or instructions are executed, the operating steps of the method described in the above-mentioned various aspects or possible implementation methods of each aspect are implemented.

[0035] In an eighth aspect, the present application provides a computer program product, which includes instructions. When the computer program product runs on a management node or a processor, the management node or the processor executes the instructions to implement the operating steps of the method described in any one of the above aspects or any possible implementation methods of any aspect.

[0036] In a ninth aspect, the present application provides a chip comprising a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to execute the operating steps of the method described in any one of the above aspects or any possible implementation of any one of the aspects.

[0037] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A schematic diagram of a panoramic video transmission system provided in this application;

[0039] Figure 2 A schematic diagram of a video processing system provided in this application;

[0040] Figure 3 A schematic diagram of the process of a video processing method provided in this application Figure 1 ;

[0041] Figure 4 A schematic diagram of a source video provided for this application;

[0042] Figure 5 A schematic diagram of the process of a video processing method provided in this application Figure 2 ;

[0043] Figure 6 A transmission diagram of a video processing provided by this application;

[0044] Figure 7 A schematic diagram of a video encoding provided by this application;

[0045] Figure 8 A schematic diagram of a bitstream transmission for perspective switching provided by this application;

[0046] Figure 9 A schematic diagram of the structure of a video processing device provided in this application;

[0047] Figure 10 A schematic diagram of the structure of a terminal device provided in this application. DETAILED DESCRIPTION

[0048] In order to make the description of the following embodiments clear and concise, a brief introduction to the relevant technology is first given.

[0049] Figure 1 This is a schematic diagram of a panoramic video transmission system provided by this application. The panoramic video processing process includes video acquisition, video encoding, video transmission, video decoding and display processes. The panoramic video transmission system includes multiple terminal devices (such as Figure 1 The terminal devices 111 to 115 are shown and a network, wherein the network can realize the function of video transmission, and the network may include one or more network devices, which may be routers or switches, etc.

[0050] Figure 1The terminal device shown may be, but is not limited to, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The terminal device may be a mobile phone (such as Figure 1 ), tablet computers, computers with wireless transceiver functions (such as the terminal 114 shown in Figure 1 ), Virtual Reality (VR) terminal devices (such as the terminal 115 shown in FIG) Figure 1 The terminal 113 shown in the figure), augmented reality (AR) terminal equipment, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in smart cities, wireless terminals in smart homes, etc.

[0051] like Figure 1 As shown, in different processing processes of panoramic videos, the terminal devices are different. For example, in the video acquisition process, the terminal device 111 can be a camera device (such as a camera) for road monitoring, a mobile phone with video acquisition function, a tablet computer or a smart wearable device, etc. For another example, in the video encoding process, the terminal device 112 can be a server or a data center, which can include one or more physical devices with encoding functions. For another example, in the video decoding and display process, the terminal device 113 can be VR glasses, and the user can control the viewing angle range by turning; the terminal device 114 can be a mobile phone, and the user can control the viewing angle range on the mobile phone 114 through touch operation or air operation; the terminal device 115 can be a personal computer, and the user can control the viewing angle range displayed on the display screen through input devices such as a mouse or keyboard.

[0052] It is understood that the embodiments of this application use 360° video as an example to illustrate the video processing method provided by this application, and this does not limit this application. Panoramic video is a general term that can refer to 360° video or 180° video. In some possible situations, panoramic video can also refer to "large" range video that exceeds the viewing angle range of the human eye (110° to 120°), such as 270° video.

[0053] Figure 1 This is just a schematic diagram. The panoramic video transmission system can also include other devices. Figure 1 The embodiments of the present application do not limit the number and type of terminal devices included in the system.

[0054] During the panoramic video processing process, multiple terminal devices 111 capture the original panoramic video of the target scene from different angles, and the terminal device 112 encodes the original panoramic video into a base layer (BL) code stream and an enhancement layer (EL) code stream. The code stream refers to a binary file or binary data stream obtained after the original video data is encoded.

[0055] BL code stream refers to the process of encoding the entire panoramic video into a video stream with lower bit rate or lower resolution in a panoramic video transmission system with viewport priority, which serves as the basic layer of the panoramic video transmission system. The view angle can be called a viewport, which refers to the window corresponding to the direction of the user's eye line of sight when the user is watching a panoramic video. The image the user sees at any time is the panoramic image corresponding to the user's view angle. The bit rate refers to the amount of data per unit time obtained by compression during video encoding, and can also be considered as the amount of data transmitted per unit time during video transmission. Generally speaking, when the same encoder is used to encode the same video source, the image quality obtained by decoding the code stream with a high bit rate is high, and the image quality obtained by decoding the code stream with a low bit rate is low. For example Figure 1 As shown, the BL code stream is continuously transmitted by the terminal device 112 to the client (such as Figure 1 As shown in the terminal devices 113 to 115, the client displays images of any viewing angle of the panoramic video to the user based on the BL code stream, ensuring that the image is presented within the viewing angle range when the user turns around.

[0056] In a panoramic video transmission system that prioritizes viewing angles, an EL stream is a high-quality video stream corresponding to the user's viewing angle, as requested by the client. The viewing angle refers to the area visible to the user, and can be referred to as the viewport or field of view.

[0057] In the current technical solution, the access reference frames of the enhancement layer code stream appear at fixed intervals in the enhancement layer code stream, and the user's perspective switching time generally does not match the fixed interval. Therefore, after the perspective is switched, the VR device downloads the next access reference frame at the current moment in the enhancement layer code stream within the new perspective range from the server, resulting in a large perspective switching delay for the VR device. One solution is to encode a RARF frame at any time, which is used for fast access to high-quality video when the user switches perspectives. However, the RARF frame must be encoded as a high-quality I frame or an independently decodable frame, and the amount of data in the I frame or the independently decodable frame is large, which makes the transmission delay of the RARF frame high, resulting in a large perspective switching delay in the VR device.

[0058] To address the above-mentioned issues, an embodiment of the present application provides a video processing method, comprising: a transmitting end transmitting a first bitstream of a source video to a receiving end, the receiving end obtaining a base layer image of the source video based on the first bitstream; the receiving end further obtaining perspective switching information and, based on the perspective switching information, obtaining information about the content to be displayed of the source video. Furthermore, the receiving end decodes the information about the content to be displayed based on the base layer image to obtain an enhanced layer image of the source video, wherein the video quality of the enhanced layer image is higher than that of the base layer image. The aforementioned information about the content to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first bitstream. The reference frame corresponds to a portion of the viewing angle area after the perspective switching. In the video processing method provided in an embodiment of the present application, when the user switches the viewing angle range, the reference frame obtained from the transmitting end by the receiving end is an inter-frame prediction frame. Because the data volume of the inter-frame prediction frame is smaller than that of an I-frame or an independently decodable frame, the transmission delay of the reference frame is reduced, thereby reducing the perspective switching delay at the receiving end and improving the perspective switching efficiency.

[0059] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0060] Figure 2 This is a schematic diagram of a video processing system provided in the present application. The video processing system 200 includes a transmitting end 210 and a receiving end 220 . The transmitting end 210 establishes a communication connection with the receiving end 220 via a communication channel 230 .

[0061] The transmitting end 210 can realize the video encoding function, such as Figure 1 As shown, the sending end 210 may be the terminal device 112, or the sending end 210 may be a data center with video encoding capabilities, for example, the data center includes multiple servers.

[0062] The transmitting end 210 may include a data source 211 , a pre-processing module 212 , an encoder 213 and a communication interface 214 .

[0063] Data source 211 may include or may be any type of electronic device for capturing video and / or any type of source video generating device, such as a computer graphics processor for generating computer animation scenes or any type of device for acquiring and / or providing source video or computer-generated source video. Data source 211 may be any type of memory or storage for storing the aforementioned source video. The aforementioned source video may include multiple video streams captured by multiple video capture devices (e.g., cameras).

[0064] The pre-processing module 212 is used to receive the source video and pre-process the source video to obtain a panoramic video. For example, the pre-processing performed by the pre-processing module 212 may include color format conversion (e.g., from RGB to YCbCr), octree structuring, video splicing, etc.

[0065] The encoder 213 is configured to receive a panoramic video and encode the panoramic video to obtain a first code stream (base layer code stream) and a second code stream (enhancement layer code stream).

[0066] The communication interface 214 in the transmitting end 210 can be used to receive the base layer code stream and the enhancement layer code stream, and send the base layer code stream and the enhancement layer code stream (or versions of the base layer code stream and the enhancement layer code stream after performing any other processing) to another device such as the receiving end 220 or any other device through the communication channel 230 for storage, display, or direct reconstruction.

[0067] The receiving end 220 can realize the function of video decoding, such as Figure 1 As shown, the receiving end 220 can be Figure 1 Any one of the terminal devices 113 to 115 is shown.

[0068] The receiving end 220 may include a display device 221 , a post-processing module 222 , a decoder 223 , and a communication interface 224 .

[0069] The communication interface 224 in the receiving end 220 is used to receive the base layer code stream and the enhancement layer code stream (or any processed versions of the base layer code stream and the enhancement layer code stream) from the transmitting end 210 or any other transmitting end such as a storage device.

[0070] The communication interface 214 and the communication interface 224 may be configured to send or receive the base layer code stream and the enhancement layer code stream via a direct communication link between the transmitter 210 and the receiver 220, such as a direct wired or wireless connection, or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network, a public network, or any combination thereof.

[0071] The communication interface 224 corresponds to the communication interface 214, for example, and can be used to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain a base layer code stream and an enhancement layer code stream.

[0072] The communication interface 224 and the communication interface 214 can be configured as follows Figure 2 The unidirectional communication interface or the bidirectional communication interface indicated by the arrow pointing from the sending end 210 to the corresponding communication channel 230 of the receiving end 220 can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded compressed data transmission, etc.

[0073] The decoder 223 is configured to receive the base layer code stream and the enhancement layer code stream, and decode the base layer code stream and the enhancement layer code stream to obtain decoded data.

[0074] The post-processing module 222 is configured to perform post-processing on the decoded data to obtain post-processed data (e.g., an image to be displayed). The post-processing performed by the post-processing module 222 may include, for example, color format conversion (e.g., from YCbCr to RGB), octree reconstruction, video splitting and fusion, or any other processing for generating data for display on the display device 221.

[0075] Display device 221 is configured to receive the post-processed data for display to a user or viewer. Display device 221 may be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen.

[0076] As an optional implementation, the transmitting end 210 and the receiving end 220 may transmit the base layer code stream and the enhancement layer code stream via a data forwarding device, such as a router or a switch.

[0077] Below is Figure 2 Taking the transmitter 210 and the receiver 220 as an example, the video processing method provided by this application is described. Figure 3 A schematic diagram of the process of a video processing method provided in this application Figure 1 , the video processing method includes the following steps.

[0078] S310 , the sending end 210 sends a first code stream of a source video to the receiving end 220 .

[0079] The first code stream is used to indicate the base layer image of the source video. For example, the source video may be a panoramic video, and the first code stream may be the base layer code stream of the panoramic video.

[0080] Optionally, during the playback of the panoramic video, the sending end 210 may actively send the first code stream to the receiving end 220 ; the sending end 210 may also send the first code stream to the receiving end 220 after receiving a request sent by the receiving end 220 .

[0081] Before transmitting the first code stream and the content information to be displayed, the sending end 210 also encodes the source video to obtain the first code stream and the content information to be displayed. Before S310, the video processing method also includes the following steps: the sending end 210 downsamples the source video to obtain a base layer image, and encodes the base layer image to obtain the first code stream.

[0082] The video quality of the source video is higher than the video quality of the base layer image. The video quality may refer to or include at least one of a signal-to-noise ratio (SNR), a resolution, and a frame rate.

[0083] Among them, the image signal-to-noise ratio refers to the ratio of the signal mean of the image to the background standard deviation. For images, the "signal mean" here generally refers to the grayscale average of the image, and the background standard deviation can be expressed by the variance of the background signal value of the image. The variance of the background signal value refers to the noise power. For images, the larger the image signal-to-noise ratio, the better the image quality. Resolution is the number of pixels per unit area of ​​a single-frame image. The higher the resolution of the image, the better the image quality. Frame rate refers to the number of frames of the video per unit time. The larger the frame rate of the video, the better the video quality of the video content. For more information about image signal-to-noise ratio, resolution, and frame rate, please refer to the relevant description of the prior art and will not be elaborated here.

[0084] The above-mentioned "downsampling" can refer to the process of generating a thumbnail of the corresponding image based on the source video in order to make the image conform to the size of the display area. For example, for an image I with a size of M*N, it is downsampled s times, that is, a resolution image of (M / s)*(N / s) size is obtained. Of course, s should be a common divisor of M and N. If the image is in matrix form, the image in the s*s window of the original image is converted into a pixel, and the value of this pixel is the mean of all pixels in the window. The downsampling method is one or more of the following methods: nearest neighbor interpolation, bilinear interpolation, mean interpolation, median interpolation and other methods. Regarding the specific method adopted for downsampling, reference can be made to the relevant description of the prior art and will not be repeated here.

[0085] Optionally, the source video content may be in the form of a panoramic video signal or an encodable panoramic video image format. If it is a panoramic video signal, the transmitter 210 may first sample and map the panoramic video signal of the base layer image into an encodable video signal, and then perform video encoding to obtain the first bitstream. For example, the "encodable video signal" is in a latitude and longitude map format.

[0086] In the video processing method provided in the embodiment of the present application, the sending end downsamples and encodes the source video to obtain a first code stream. Since the video quality of the base layer image is lower than the video quality of the source video, the bit rate of the first code stream is lower than the bit rate of the code stream obtained by encoding the source video. The amount of data required to be transmitted per unit time by the sending end and the receiving end is reduced, thereby reducing the delay for the receiving end to obtain the base layer image.

[0087] S320: The receiving end 220 obtains a base layer image of the source video according to the first code stream.

[0088] Optionally, the source video may include multiple sub-areas of the video. For example, in the High Efficiency Video Coding (HEVC) standard, the image to be coded may be divided into multiple block-shaped coding areas. For example, each frame of the image to be coded may be divided into multiple tiles. The multiple tiles together constitute the image to be coded, and each tile may be coded independently. Figure 1 The panoramic video shown is divided into multiple sub-areas. Any two adjacent sub-areas may be independent, and any two adjacent sub-areas may also have a partially overlapping area.

[0089] For example, a source video may be processed into a spherical signal in a panoramic image format using a latitude and longitude diagram, and a two-dimensional panoramic image capable of storage and transmission is obtained by uniformly sampling and mapping the source video according to longitude and latitude. The horizontal and vertical coordinates of the panoramic image may be expressed in longitude and longitude, for example, the width direction may be expressed in longitude, spanning 360°, and the height direction may be expressed in latitude, spanning 180°.

[0090] like Figure 4 As shown, Figure 4 A schematic diagram of a source video provided for this application, wherein the source video can be divided into multiple sub-areas, and the position of each sub-area in the source video can be represented by longitude and latitude or a span range of longitude and latitude, or a combination of the two. For example, after the receiving end 220 obtains the first code stream from the sending end 210, it decodes the first code stream to obtain a base layer image, which can be a low-quality picture including all sub-areas of the source video. The "low quality" here can be understood as the video quality of the base layer image being lower than a threshold, which can be preset or adjusted according to the needs of the user. For content on video quality, please refer to the relevant description of S310 and will not be elaborated here.

[0091] S330 , the receiving end 220 obtains the view switching information, and obtains the to-be-displayed content information of the source video from the sending end 210 according to the view switching information.

[0092] The viewing angle switching information may refer to the viewing angle range of the user of the receiving end 220 being switched from the first viewing angle range to the second viewing angle range, such as Figure 4 As shown, the second viewing angle range mentioned above includes a new area that is not included in the first viewing angle range in the source video.

[0093] The first viewing angle range is used to indicate that when the user's viewing angle is in the first direction, multiple first sub-areas in the source video have an associated relationship of being displayed continuously. Figure 4 As shown, the first viewing angle range includes a plurality of first sub-areas surrounded by bold solid lines.

[0094] The second viewing angle range is used to indicate that when the user's viewing angle is in the second direction, multiple second sub-areas in the source video have an associated relationship of being displayed continuously. Figure 4 As shown, the second viewing angle range includes a plurality of second sub-areas surrounded by bold dotted lines.

[0095] in, Figure 4 The overlapping sub-region shown is the overlapping region of the first viewing angle range and the second viewing angle range, and the other sub-regions except the overlapping sub-region in the plurality of second sub-regions are the aforementioned newly added regions.

[0096] Figure 4 The explanation is given by taking the overlapping area of ​​the first viewing angle range and the second viewing angle range as an example, but in some possible examples, the first viewing angle range and the second viewing angle range may not have an overlapping area, then the area composed of all the second sub-areas in the second viewing angle range are all new areas.

[0097] In addition, in the embodiment of the present application, the sub-region may also be referred to as an image sub-region. For example, the code stream obtained by the transmitter 210 encoding the video of each sub-region may be referred to as a sub-region code stream (or sub-region stream).

[0098] In one possible scenario, the sending end 210 divides the image to be encoded with a timestamp of t included in the source video, and obtains a portion of the image to be encoded, which is called a sub-picture of the image to be encoded. The code stream obtained by encoding all images in the area corresponding to the sub-picture can be called a sub-picture code stream.

[0099] In another possible scenario, such as in the HEVC standard, each sub-image is called a tile, and the code stream obtained by encoding multiple tiles of the same area within a period of time by the transmitter 210 is a tile stream.

[0100] In various embodiments provided in this application, unless otherwise specified, sub-region code stream, sub-region stream, sub-image code stream, and tile stream are all represented by sub-region code stream.

[0101] The above-mentioned information of the content to be displayed is used to indicate the enhanced layer image of the viewing area after the viewing angle is switched, such as Figure 4 As shown, the enhanced layer image refers to the video screen displayed by multiple second sub-areas in the source video. The video quality of the enhanced layer image is higher than that of the base layer image. For details about the video quality, please refer to the relevant description of S320 above and will not be repeated here.

[0102] The content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first bitstream. The reference frame corresponds to a partial viewing area after the viewing angle is switched. For example, the reference frame can be Figure 4 The reference frame of the newly added area shown; herein, without causing misunderstanding, the reference frame may also be referred to as a RARF frame. The reference frame may be an inter-frame prediction frame obtained by the transmitter 210 based on the first bitstream, such as a forward prediction coding frame (P frame). In addition, the reference frame may refer to a reference frame image or a bitstream of a reference frame included in the content information to be displayed.

[0103] In the video processing method provided in the embodiments of the present application, RARF frames are replaced with inter-frame prediction frames (P frames) based on the first bitstream, replacing the conventional I-frames. Because the data size of P frames is smaller than that of I frames, the transmission latency of RARF frames and the view switching latency at the receiving end are reduced. Furthermore, P frames in the content information to be displayed can be decoded based on the first bitstream, allowing the receiving end to effectively reuse the first bitstream, reducing redundancy in video transmission and improving video transmission efficiency.

[0104] As an optional implementation, the receiving end 220 may obtain the content information to be displayed by sending an instruction (or request) to the sending end 210, such as Figure 5 As shown, Figure 5 A schematic diagram of the process of a video processing method provided in this application Figure 2 , Figure 5 Given Figure 3 As shown in a possible implementation manner of S330, the above S330 may include the following steps S3301 to S3303.

[0105] S3301: The receiving end 220 obtains a partial viewing area after the viewing angle is switched according to the current viewing angle information and the viewing angle switching information.

[0106] The partial viewing angle area after the viewing angle is switched may refer to a newly added area of ​​the second viewing angle range relative to the first viewing angle range. Figure 4 As shown, the newly added area refers to the other second sub-areas in the second viewing angle range except the overlapping sub-areas.

[0107] As an optional implementation method, the above-mentioned S3301 may include: the receiving end 220 matches the area information included in the first viewing angle range with the preset first relationship to determine multiple first sub-area identifiers, and the receiving end 220 matches the area information included in the second viewing angle range with the first relationship to determine multiple second sub-area identifiers. Furthermore, the receiving end 220 compares the multiple first sub-area identifiers with the multiple second sub-area identifiers to determine at least one sub-area identifier after the viewing angle is switched, and the sub-area indicated by the at least one sub-area identifier after the viewing angle is the partial viewing angle area after the viewing angle is switched. The above-mentioned first relationship is used to indicate the sub-area division information in the source video and the identifier of each sub-area, and the identifier can be used to indicate the position of the sub-area in the source video.

[0108] For example, the first relationship mentioned above may refer to the description file information of the video content of the source video, and the description file information may be a set of metadata describing the sub-region division information of the source video, and the set of metadata records the sub-region division information in the source video, and the identification of each sub-region. For example, in the case where the source video is a latitude and longitude map, the identification may be expressed in longitude and latitude, such as the longitude and latitude coordinates of a specific point in the sub-region, and the specific point may refer to the center point of the sub-region, an edge point (such as the point in the lower left corner of the sub-region or the point in the lower right corner of the sub-region), or other points (such as the point located at 2 / 3 of the symmetry axis in the sub-region), etc.

[0109] In one possible scenario, the range of a sub-region can be represented by multiple edge points of the sub-region. For example, if the sub-region is a quadrilateral, the range of the sub-region can be represented by the latitude and longitude coordinates of the four corner points of the quadrilateral. For example, the transmitting end 210 or the receiving end 220 can locate the entire sub-region by locating the positions of these four corner points.

[0110] In another possible scenario, the range of the sub-region can also be represented by a longitude and latitude span range, which serves as the sub-region identifier. For example, if the sub-region identifier is "270°-280°, 110°-115°", the longitude and latitude span range of the sub-region in the source video is "270°-280°" in longitude and "110°-115°" in latitude.

[0111] In another possible scenario, the range of the sub-region can also be represented by the longitude and latitude coordinates of the corner points of the sub-region, as well as the range of the sub-region in the longitude and latitude directions, as the identifier of the sub-region. For example, if the sub-region identifier is "(40°, 50°), 70°, 60°", then the sub-region in the source video is represented by the longitude and latitude coordinates (40°, 50°) as the marker point (such as the upper left corner of the region), and the range in the longitude direction is 70° longitude and the range in the latitude direction is 60° latitude.

[0112] In an embodiment of the present application, the receiving end 220 can detect changes in the user's perspective and obtain new perspective information and corresponding video sub-area information. Specifically, the receiving end 220 detects changes in the user's current perspective direction, obtains a new perspective direction (such as the second direction mentioned above), and based on the new perspective direction and the description file information of the source video downloaded from the sending end 210, calculates the sub-area information required to obtain the new perspective direction, and obtains the sub-area newly added to the new perspective direction compared to the old perspective. In this way, the receiving end 220 can obtain the content information to be displayed from the sending end 210 based on the new perspective information and the corresponding video sub-area information, avoiding the sending end 210 from processing the user's perspective information and saving the processing resources of the sending end 210.

[0113] In one possible scenario, the receiving end 220 saves the description file information of the source video. For example, the receiving end 220 sends a video playback request to the sending end 210. The sending end 210 sends the description file information and the basic layer code stream (first code stream) of the source video to the receiving end 220 based on the video playback request. The receiving end 220 plays the basic layer image of the source video according to the first code stream.

[0114] S3302 , the receiving end 220 sends a first instruction including an identifier to the sending end 210 .

[0115] Optionally, each sub-region in the source video has its own identifier, and the first instruction may include the identifier of the newly added region. The identifier included in the first instruction may be used to indicate the position of the partial viewing area after the perspective switch in the base layer image. For example, the first instruction may also be used to indicate to the user of the receiving end 220 that the viewing range has been switched from the first viewing range to the second viewing range.

[0116] S3303: The sending end 210 sends the content information to be displayed to the receiving end 220 based on the first instruction.

[0117] The to-be-displayed content information also includes a second bitstream corresponding to a second viewing angle range in the source video, the second bitstream being used to indicate other frames in the enhancement layer image except the reference frame. For example, the second bitstream may be an enhancement layer bitstream corresponding to the second viewing angle range.

[0118] In one possible scenario, the reference frame and the second code stream are transmitted separately. For example, the transmitter 210 may first send the reference frame included in the content information to be displayed to the receiver 220 according to the first instruction, and then send the at least one second code stream included in the content information to be displayed to the receiver 220.

[0119] In another possible scenario, the reference frame and the second code stream are transmitted together. For example, the transmitter 210 splices the reference frame and the second code stream to obtain content information to be displayed, and sends the content information to be displayed to the receiver 220.

[0120] For example, the transmitter 210 may encapsulate the reference frame and the second code stream to obtain a video data packet, and send the video data packet to the receiver 220. The protocol used by the transmitter 210 for encapsulation may be, but is not limited to, the dynamic streaming (DASH) protocol, the real-time message protocol (RTMP), and the hypertext transfer protocol (HTTP)-based live streaming (HLS) protocol. For example, in RTMP, messages with a message type identifier (identity document, ID) of 8 or 9 are used to transmit audio (identity document 8) and video (identity document 9) data, respectively.

[0121] In a possible example, the second code stream includes one or more sub-region code streams. Figure 4 As shown, the second code stream may include sub-region code streams of any two second sub-regions in the second viewing angle range.

[0122] In another possible example, the transmitter 210 may splice the at least one second code stream into an application code stream. For example, the receiver 220 splices the sub-region code streams of each second sub-region in the second viewing angle range to obtain an application code stream, and decodes the application code stream based on the second image to obtain the enhanced layer image of the partial viewing angle region after the viewing angle switch. It is worth noting that if each second code stream includes a sub-region code stream, when the transmitter 210 splices multiple second sub-region code streams, the transmitter 210 may also generate position information of each second sub-region code stream in the second viewing angle range, so that the receiver 220 can decode and display the application code stream based on this position information. In this way, when the receiver 220 splices multiple second code streams into one application code stream, the receiver 220 can use only one decoder to decode the second code stream, saving the processing resources required for decoding at the receiver 220, reducing the decoding latency at the receiver 220, and improving the decoding efficiency at the receiver.

[0123] Please continue to see Figure 3 After the receiving end 220 obtains the content information to be displayed, the video processing method provided in the embodiment of the present application further includes the following step S340.

[0124] S340: The receiving end 220 decodes the to-be-displayed content information based on the base layer image to obtain an enhancement layer image of the source video.

[0125] In an embodiment of the present application, the receiving end 220 decodes the reference frame included in the content information to be displayed based on the basic layer image, and reuses the content of the basic layer image, so that even if the reference frame is not an independently decodable frame (such as an I frame), the receiving end 220 can also decode to obtain the enhanced layer image.

[0126] Thus, in the video processing method provided in the embodiment of the present application, the reference frame (RARF frame) is replaced by the inter-frame prediction frame (P frame) based on the first code stream instead of the intra-frame coding frame (I frame) in the prior art. Since the inter-frame prediction frame contains partial information of the image and the intra-frame coding frame contains all information of the image, the data amount of the P frame is smaller than the data amount of the I frame, which reduces the transmission delay of the reference frame between the sending end and the receiving end, as well as the perspective switching delay at the receiving end.

[0127] Since the reference frame is not an independently decodable frame, the receiving end 220 decodes the content information to be displayed. The embodiment of the present application provides a possible implementation method. Please continue to refer to Figure 5 The process of decoding the display content information at the receiving end 220 may include: Figure 5 Steps S3401 to S3403 are shown.

[0128] S3401: The receiving end 220 obtains a first image of a partial viewing area after the viewing angle is switched based on the base layer image.

[0129] The timestamp of the first image matches the timestamp of the reference frame. This timestamp indicates the temporal position of the image or reference frame in the source video. The temporal domain refers to the time range in the source video, which is based on time and scaled by time. For more information about the field of view, please refer to the relevant content of the prior art and will not be elaborated here.

[0130] For example, in an embodiment of the present application, in order to accurately describe the timestamp of an image, a code stream, or a reference frame, the timestamp can be represented by a frame number. For example, if the frame rate of the source video is 30 Hz and the total duration of the source video is 10 seconds (s), then the source video has a total of 30×10=300 images, the frame number of the t-th frame of the 300 images is t, and the frame number of the code stream corresponding to the t-th frame is t, where t is a positive integer.

[0131] It is worth noting that the above examples are merely examples provided for illustrating timestamps in the embodiments of this application and should not be construed as limiting this application. In other examples, timestamps may also be represented by time stamps, such as a time stamp of "02:01" for an image in a source video, indicating that the image is the first frame of multiple frames in the second second of the source video.

[0132] S3402: The receiving end 220 obtains a second image based on the first image and the reference frame.

[0133] The second image corresponds to a partial viewing angle area of ​​the enhancement layer image after the viewing angle is switched.

[0134] like Figure 4 As shown, in one possible scenario, if the second viewing angle range and the first viewing angle range have an overlapping area, and the reference frame only includes the inter-frame prediction frame of the newly added area, then the process of the receiving end 220 decoding the reference frame based on the first image may include: the receiving end 220 decodes the reference frame based on the partial image corresponding to the newly added area in the first image to obtain the image of the newly added area, and the video quality of the image of the newly added area is higher than the video quality of the first image. In addition, the receiving end 220 also decodes the code stream of other areas (overlapping areas) in at least one second code stream to obtain the image of the overlapping area, and the timestamp of the code stream of the overlapping area is consistent with the timestamp of the reference frame. Then, the receiving end 220 splices the image of the newly added area and the image of the overlapping area to obtain the second image.

[0135] In another possible scenario, if the second viewing angle range and the first viewing angle range have an overlapping area, the reference frame includes the overlapping area and the inter-frame prediction frame of the newly added area, or the second viewing angle range and the first viewing angle range do not have an overlapping area, and the reference frame only includes the inter-frame prediction frame of the newly added area, then the process of the receiving end 220 decoding the reference frame based on the first image may include: the receiving end 220 decodes the reference frame based on the first image to obtain the second image.

[0136] S3403: The receiving end 220 obtains an enhancement layer image based on the second image and at least one second code stream.

[0137] For example, if the reference frame corresponds to time t+Δt, the receiver 220 obtains the sub-region stream starting at time t+Δt+1 for the newly added region within the second viewing angle; if the reference frame corresponds to time t, the receiver 220 obtains the sub-region stream starting at time t+1 for the newly added region within the second viewing angle. These sub-region streams are decoded with reference to the decoded image (second image) of the reference frame corresponding to the current sub-region, and the enhancement layer image obtained by decoding the sub-region streams is sent to the display device 221 of the receiver 220 for display.

[0138] In a possible implementation, if the content information to be displayed includes multiple second code streams, the receiving end 220 uses multiple decoders to decode the multiple second code streams.

[0139] In another possible implementation, if multiple second code streams included in the content information to be displayed are spliced ​​into one application code stream by the transmitting end 210 or the receiving end 220, the receiving end 220 may use only one decoder to decode the application code stream.

[0140] For example, the encoding process of the second bitstream can adopt the motion-constrained tile sets (MCTS) encoding method. The MCTS encoding method refers to the process of encoding the source video, which restricts the motion vector within the tile so that the tiles at the same position in the source video will not refer to the image pixels outside the tile area in the time domain. Therefore, each tile in the time domain can be decoded independently. For more information about the MCTS encoding method, please refer to the relevant description of the prior art and will not be repeated here. In this way, during the decoding process of the second bitstream, the receiving end 220 can splice multiple second bitstreams to obtain an application bitstream, so that the receiving end 220 can decode the application bitstream with only one decoder to obtain the enhanced layer image within the second viewing angle.

[0141] It is worth noting that the above example is illustrated by taking the receiving end 220 as an example in which the multiple sub-region code streams within the second viewing angle are spliced ​​together. In some possible scenarios, during the process of encoding the source video to obtain the second code stream, the sending end 210 may also splice the multiple sub-region code streams within the second viewing angle to obtain an application code stream, and the sending end 210 sends the application code stream to the receiving end 220. Subsequently, the receiving end 220 uses a decoder to decode the reference frame included in the content information to be displayed and the application code stream to obtain an enhanced layer image.

[0142] In this way, the embodiment of the present application uses a tile-based MCTS encoding method. After the receiving end obtains different tile code streams within the viewing angle at time t, these tile code streams can be spliced ​​into a code stream that can be decoded by a decoder by the receiving end. Therefore, the receiving end can use one decoder to decode all tile images within the viewing angle, saving the decoding system resources of the receiving end, reducing the decoding delay of the receiving end, and improving the decoding efficiency of the receiving end.

[0143] In the video processing method provided by the embodiments of the present application, during the user's viewing angle switching process, after receiving the to-be-displayed content information including the reference frame, the receiver 220 can rapidly decode the to-be-displayed content information based on the first bitstream and obtain the enhancement layer image within the second viewing angle. Specifically, first, the receiver 220 decodes the bitstream in the first bitstream that matches the timestamp of the reference frame to obtain the first image. Second, the receiver 220 uses the first image as a reference image to decode the reference frame to obtain the second image. Finally, the receiver 220 decodes at least one second bitstream based on the second image to obtain the enhancement layer image. In this manner, since the reference frame is an inter-frame predicted frame, the data volume of the reference frame is smaller than that of an I-frame. Therefore, the transmission delay of the reference frame between the transmitter 210 and the receiver 220 is reduced, and the viewing angle switching delay at the receiver 220 is reduced. Furthermore, during the decoding of the enhancement layer image, the receiver 220 can reuse the content of the first bitstream to decode the reference frame, reducing the amount of video data required to be stored by the receiver 220 and conserving storage resources at the receiver 220.

[0144] As an optional implementation, the receiving end 220 may also display at least one of the base layer image and the enhancement layer image of the source video. Source videos have various playback scenarios, such as live, which refers to "a broadcast method in which the post-production synthesis and broadcast of a radio and television program are synchronized." In some video playback scenarios, live may also refer to a broadcast method in which the interval between the post-production synthesis and broadcast of a radio and television program is less than a delay threshold (e.g., 5 seconds or 1 minute). For example, on-demand refers to "a broadcast method in which the post-production synthesis and broadcast of a radio and television program are asynchronous."

[0145] In a possible example, if the playback scene of the source video is live broadcast, the above-mentioned second code stream can be obtained from the source video according to the second viewing angle range by the sending end 210 after obtaining the first instruction; for example, the sending end 210 divides the source video into multiple sub-areas, and encodes the video corresponding to each sub-area in the multiple sub-areas to obtain multiple sub-area code streams; the sending end 210 matches the area information included in the second viewing angle range with the above-mentioned first relationship to determine multiple second sub-area identifiers, and the first relationship is used to indicate the sub-area division information and the identifier of each sub-area in the source video, and then the sending end 210 determines the sub-area code stream corresponding to each second sub-area identifier as the second code stream.

[0146] In this example, the playback scene of the source video is live broadcast, and the sending end 210 can quickly obtain the second stream based on the source video according to the second viewing angle range determined by the receiving end 220.

[0147] In another possible example, if the source video is played on demand, the second bitstream may be prepared by the transmitter 210 before playback of the source video begins. For example, the transmitter 210 encodes the video corresponding to each of the multiple sub-regions in the source video to obtain multiple sub-region bitstreams, including the second bitstream corresponding to the second viewing angle range.

[0148] In addition, the plurality of sub-region code streams may further include a third code stream, which is a sub-region code stream of a portion of the viewing angle area after the viewing angle is switched. Figure 4 As shown, the third code stream may include sub-region code streams corresponding to other sub-regions except the overlapping region in the second viewing angle range.

[0149] After the transmitter 210 obtains the third code stream, the transmitter 210 may also obtain a reference frame. In one possible example, the transmitter 210 further obtains a reference frame based on the first code stream and the third code stream. For example, the transmitter 210 decodes the first code stream to obtain a first image, and the timestamp of the first image is consistent with the timestamp of the perspective switching, such as the timestamp of the first image matches the timestamp of the user of the receiving end 220 switching from the first perspective range to the second perspective range; secondly, the transmitter 210 encodes the second image based on the first image to obtain a reference frame. The second image is an image of a partial perspective area (such as the newly added area mentioned above) after the perspective switching in the source video, or an image obtained by decoding the third code stream, and the timestamp of the second image matches the timestamp of the first image.

[0150] In the video processing method provided in the embodiment of the present application, the transmitter 210 obtains a reference image (first image) that matches the timestamp of the view switching based on the first bitstream, and encodes the image to be encoded (second image) based on the reference image to obtain a reference frame for the newly added area. Because the reference frame is an inter-frame prediction frame, the data volume of the inter-frame prediction frame is smaller than the data volume of the I-frame for the same image to be encoded. Therefore, the transmission delay of the reference frame is reduced, and the view switching delay at the receiving end is reduced.

[0151] In the above embodiment of the present application, the generation of the reference frame is determined by the transmitter based on the newly added area in the second viewing angle range, that is, the reference frame is generated in real time by the transmitter based on the viewing angle change of the receiver. However, in some possible situations, the reference frame included in the content information to be displayed may also be generated by the transmitter 210 before S310. For example, the transmitter 210 decodes the first code stream to obtain a sub-region reference image sequence within the first time range, and the sub-region reference image sequence includes a reference image of each sub-region; the transmitter 210 encodes the image sequence to be encoded based on the sub-region reference image sequence to obtain a reference code stream, and the image sequence to be encoded includes multiple images to be encoded, and the image to be encoded is a sub-region image in the source video that matches the reference image, or an image obtained by decoding a code stream in the sub-region code stream that matches the reference image, and the timestamp of the image to be encoded matches that of the reference image. The above-mentioned reference code stream includes the reference frame in the content information to be displayed.

[0152] Based on the above-mentioned transmitting end 210 and receiving end 220, the present application provides a possible implementation of a video processing method, such as Figure 6 As shown, Figure 6 This is a transmission diagram of a video processing provided in the present application. In the video processing method provided in the embodiment of the present application, the transmission of the video code stream includes the following video encoding process and video decoding process.

[0153] like Figure 6 As shown, the video encoding process includes the following steps A1 to A3.

[0154] A1, the sending end 210 downsamples the source video to obtain a base layer image, and encodes the base layer image to obtain a base layer code stream (such as the first code stream mentioned above).

[0155] A2, the sending end 210 divides the source video into sub-regions to obtain multiple source sub-region images, and then the sending end 210 encodes the multiple source sub-region images to obtain an enhancement layer code stream of each source sub-region image (such as the above-mentioned sub-region code stream).

[0156] A3: The transmitter 210 encodes the multiple images to be encoded based on the first code stream to obtain a RARF code stream (reference code stream).

[0157] The above process of obtaining the RARF code stream requires a sub-region reference image sequence and a to-be-encoded image sequence.

[0158] The sub-region reference image sequence includes multiple sub-region reference images, and the sub-region reference images are images obtained by decoding the base layer code stream (first code stream) by the transmitter 210. Figure 6As shown, for a basic layer code stream of a certain timestamp t (corresponding to the t-th frame image in the source video), the transmitter 210 decodes the basic layer code stream (code stream 1) to obtain the decoded image of the source video (such as Figure 6 The decoded image corresponds to the entire panoramic image obtained after downsampling the source video. The decoded image is upsampled to restore it to the same resolution as the t-th frame image of the source video; then, the transmitting end 210 divides the upsampled decoded image (such as Figure 6 The first image in the transmitting end 210 shown in the figure is divided into sub-regions to obtain the t-th frame reference image of each sub-region.

[0159] It is worth noting that if the transmitter 210 does not downsample the source video when generating the base layer code stream, then during the generation of the RARF code stream, the transmitter 210 does not need to upsample the decoded image obtained by decoding the base layer code stream; similarly, the same is true for the decoding process of the RARF code stream.

[0160] The image sequence to be coded includes multiple images to be coded. The images to be coded are sub-region images in the source video that match the reference image, or images obtained by decoding the sub-region code stream that matches the reference image. The timestamps of the images to be coded match those of the reference image. Figure 6 As shown, for sub-region X, the transmitter 210 decodes the enhancement layer sub-region code stream (code stream 2) of the sub-region X to obtain the t-th frame decoded image of the sub-region X (image to be encoded).

[0161] Furthermore, after the transmitting end 210 obtains the t-th frame reference image of each sub-region, the t-th frame decoded image of the sub-region X is used as a reference according to the t-th frame reference image of the sub-region X, and is encoded in the form of a P frame to obtain the t-th frame code stream, which is the t-th frame RARF code stream of the sub-region X.

[0162] In this way, the transmitter 210 performs operations consistent with the t-th frame RARF codestream of sub-region X on all sub-regions, obtaining the RARF codestream corresponding to the t-th frame (reference codestream) for all sub-regions. It is worth noting that during the RARF codestream acquisition process, if the corresponding frame number in the RARF codestream is not the t-th frame, the transmitter 210 can modify the frame number information in the codestream header to the t-th frame before saving.

[0163] In order to improve the coding efficiency of the transmitter 210, the present application also provides a possible implementation method for the above-mentioned A3, such as Figure 7 As shown, Figure 7This application provides a schematic diagram of video coding. A sub-region reference image sequence includes five sub-region reference images within a first time range, with sub-region identifiers 1 to 5. A sequence of images to be coded includes five images to be coded within the first time range, with sub-region identifiers 1 to 5. The first time range refers to a continuous period of time, such as 10 seconds or an hour. The transmitter 210 encodes the sequence of images to be coded based on the sub-region reference image sequence to obtain a reference bitstream, which may include the following steps S71 and S72.

[0164] S71 , the transmitting end 210 interleaves and recombines the sub-region reference image sequence and the to-be-encoded image sequence to obtain first recombined data and second recombined data.

[0165] The first reconstructed data includes a subregion reference image and an image to be encoded whose subregion identifiers meet a first condition, and the second reconstructed data includes a subregion reference image and an image to be encoded whose subregion identifiers meet a second condition.

[0166] like Figure 7 As shown, the first condition may refer to that the sequence number of the sub-region identifier is an odd number, and the second condition may refer to that the sequence number of the sub-region identifier is an even number.

[0167] S72 , the transmitter 210 encodes the first reassembled data and the second reassembled data to obtain a reference bit stream.

[0168] like Figure 7 As shown, the reference code stream refers to the RARF code stream, and the reference code stream includes the reference frame in the above-mentioned content information to be displayed.

[0169] S72 specifically includes: the transmitting end 210 obtains a fourth code stream based on the first reconstructed data, where the fourth code stream includes all reference frames whose sub-region identifiers within the first time range meet the first condition; the transmitting end 210 obtains a fifth code stream based on the second reconstructed data, where the fifth code stream includes all reference frames whose sub-region identifiers within the first time range meet the second condition; and further, the transmitting end 210 obtains a reference code stream based on the fourth code stream and the fifth code stream.

[0170] In an embodiment of the present application, the transmitter 210 rearranges the sub-region reference image sequence and the image sequence to be encoded within the first time range into two image sequences (first reconstructed data and second reconstructed data) in an interleaved frame number manner. In each reconstructed data, the sub-region reference image at the same time is located before the image to be encoded. After obtaining the two image sequences, the transmitter can encode the two image sequences using a low-latency encoding configuration with one key frame every two frames to obtain a reference bitstream within the first time range. Because the header information of each reference frame in the bitstream obtained after the transmitter encodes the reconstructed data is consistent with the frame number of the image to be encoded, the transmitter does not need to modify the frame number of each reference frame in the reference bitstream, thereby improving the encoding efficiency of the transmitter.

[0171] For example, a frame in a bitstream is a picture in a video. Video encoding is performed based on groups of pictures (GOPs). GOPs are not linked to each other; encoding relationships only occur within a GOP. Each GOP group begins with a keyframe, which is a complete picture. Frames other than the keyframe in a GOP are incomplete. Therefore, the pictures corresponding to these frames can be referenced based on the keyframe. For example, a keyframe can be an I-frame or a P-frame encoded in all-I-block mode.

[0172] For example, the above-mentioned “low-delay encoding configuration with one key frame every two frames” can be that the number of reference frames is 1. For example, when the transmitter 210 encodes the first reconstructed data or the second reconstructed data, the image at the K+1 frame (such as Figure 7 The square image to be encoded (included in the first reconstructed data) refers only to K frames, where K is a positive integer. Furthermore, based on the first reconstructed data, transmitter 210 obtains a third bitstream of reference frames whose sub-region identifiers are odd numbers within the first time range. Based on the second reconstructed data, transmitter 210 obtains a fourth bitstream of reference frames whose sub-region identifiers are even numbers within the first time range. Finally, transmitter 210 extracts the frames encoded by the image data of the sub-region to be encoded from the two encoded bitstreams and reconstructs them to obtain the RARF bitstream for the first time range.

[0173] Please continue to see Figure 6 , the video decoding process includes the following steps B1 to B3.

[0174] B1, the receiving end 220 obtains the base layer code stream (first code stream) from the sending end 210, and decodes the base layer code stream to obtain a base layer image.

[0175] B2: When the viewing angle of the user of the receiving end 220 switches from the first viewing angle to the second viewing angle, the content information to be displayed is obtained from the sending end 210.

[0176] The content information to be displayed includes the reference frame of the partial viewing area after the viewing angle is switched (such as the newly added area in the second viewing angle range) and the enhancement layer code stream (second code stream) of the second viewing angle range.

[0177] B3. The receiving end 220 decodes the to-be-displayed content information based on the first code stream to obtain an enhanced layer image.

[0178] First, the receiving end 220 decodes the code stream 1 in the first code stream to obtain a first image, and the timestamp of the first image is consistent with the timestamp of the reference frame.

[0179] Next, the receiving end 220 decodes the reference frame based on the first image to obtain a second image, where the second image corresponds to a partial viewing angle area after the viewing angle is switched.

[0180] Finally, the receiving end 220 decodes the enhancement layer code stream (at least one second code stream) based on the second image to obtain the enhancement layer image.

[0181] It is worth noting that before decoding the reference frame, the receiving end 220 will also determine the coding method of the reference frame with the transmitting end 210.

[0182] In the first possible example, Figure 6 As shown, if the reference frame is encoded using the first image as the encoding reference image at the transmitter 210, then at the receiver 220, the decoder searches the base layer stream (first stream) for stream 1 that matches the timestamp of the reference stream (RARF stream). The decoder decodes and upsamples stream 1 to obtain the first image. Furthermore, the decoder uses the first image as the reference image to decode the reference stream (RARF stream) to obtain the second image. This improves the coding efficiency of the reference stream by directly using the first image as the reference image.

[0183] In the second possible example, Figure 6As shown, if the reference frame is encoded by the transmitter 210 using the video frame obtained by encoding the first image as a reference, for example, the transmitter 210 encodes the t-th reference image into a video frame, which is an intra-coded frame or an independent video frame. Furthermore, the transmitter 210 uses the t-th decoded image of sub-region X as a reference based on the video frame of the t-th reference image of sub-region X to obtain the t-th RARF codestream of sub-region X. Then, the receiver 220 can use two decoders to decode the base layer codestream and the RARF codestream respectively. For example, decoder 1 decodes and upsamples codestream 1 in the base layer codestream (first codestream) that matches the timestamp of the reference codestream (RARF codestream) to obtain the first image. Furthermore, the encoder at the receiver 220 encodes the first image into a video frame, concatenates the video frame with the reference codestream (RARF codestream), and transmits the concatenated codestream to decoder 2. Finally, decoder 2 decodes the reference frame in the content information to be displayed based on the concatenated codestream to obtain the second image.

[0184] Because the decoder requires reference images to decode reference frames in the RARF codestream, in this example, the decoder concatenates the video frame and the RARF codestream to obtain a concatenated codestream and performs decoding based on this concatenated codestream. This eliminates the need for the receiver 220 to perform additional operations to update the reference image queue required for decoding in the decoder, thereby improving decoding efficiency and reducing view switching latency at the receiver. Specifically, in the second possible example described above, the receiver can encode the first image into a video frame and concatenate the video frame with the RARF codestream to obtain a concatenated codestream. The decoder then decodes this concatenated codestream to obtain a second image. Thus, in this embodiment of the present application, the decoder can use the concatenated codestream obtained from the video frame and the RARF codestream to obtain the second image, and based on the second image and at least one second codestream, obtain the enhancement layer image for the partial view area after view switching. This avoids the need for the receiver to modify the reference image in the decoder to the first image, improving decoding efficiency and reducing view switching latency at the receiver.

[0185] In order to more intuitively demonstrate the difference between this application and the prior art, the embodiment of this application provides a possible implementation method, such as Figure 8 As shown, Figure 8 A schematic diagram of a bitstream transmission for perspective switching provided by this application. When a user watches a panoramic video without moving their head, the receiving end obtains the base layer bitstream and the enhancement layer bitstream corresponding to the current perspective (first perspective range) from the sending end.

[0186] The receiving end can obtain a relatively low-quality panoramic image (basic layer image) by decoding the basic layer code stream; Figure 8The two images corresponding to media file segments t (seg) and seg t+1 in the base layer stream are shown. After decoding the enhancement layer stream, the receiver obtains high-quality sub-region images within the first viewing angle. Sub-region images within all viewing angles are displayed in the field of view.

[0187] When the receiving end detects that the user's viewing angle has changed (such as Figure 8 The perspective switching timestamp t shown), in the prior art, the receiving end can only obtain other frames corresponding to the perspective switching timestamp t in the second perspective range, and the other frames are unidirectional prediction coding frames or bidirectional prediction coding frames (bi-directional interpolated prediction frame, B frame). Among them, the unidirectional prediction coding frame can refer to a forward prediction coding frame or a backward prediction coding frame. For example, the prediction method of using K-1 frame image to predict K frame image is called forward prediction, K is a positive integer greater than or equal to 1; the prediction method of using K+1 frame image to predict K frame image is called backward prediction; the prediction method of using K-1 frame image and K+1 frame image to predict K frame image is called bidirectional prediction. For more information about forward prediction, backward prediction and bidirectional prediction, please refer to the relevant explanation of the prior art, which will not be repeated here.

[0188] However, the other frames (P frames or B frames) cannot be decoded independently, resulting in the receiver being unable to obtain high-quality panoramic video content during the view switching timestamp t~t+1, and the user being unable to watch the high-quality panoramic video content during this period, resulting in a poor user experience (Quality of Experience, QoE).

[0189] In the video processing method provided in this application, Figure 8 After the perspective switching timestamp t shown, the content information to be displayed obtained by the receiving end from the sending end includes the reference frame obtained by the sending end based on the basic layer code stream. The receiving end can decode the reference frame based on the basic layer code stream to obtain a second image of the partial perspective area after the perspective switching, and obtain the enhanced layer image based on the second image and at least one second code stream (other perspective areas in the second perspective range except the partial perspective area). During the period from the perspective switching timestamp t to t+1, other frames (P frames or B frames) can be decoded based on the enhanced layer image obtained by decoding the reference frame, thereby obtaining high-quality panoramic video content within the second perspective range for the user. Therefore, in the video processing method provided in the embodiment of the present application, when the user's perspective range switches from the first perspective range to the second perspective range, the receiving end can quickly display high-quality video content within the user's current perspective range, reducing the perspective switching delay and improving QoE.

[0190] It is understandable that in order to implement the functions in the above embodiments, the computing device and chip include hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0191] Combined with the above Figures 1 to 8 , describes in detail the video processing method provided by this application, and will be combined with Figure 9 and Figure 10 , describing the video processing apparatus and terminal device provided according to this embodiment.

[0192] Figure 9 This is a schematic diagram of the structure of a video processing device provided by the present application. The video processing device 900 includes a communication module 910 and a processing module 920. The processing module 920 includes an encoding module 921 and a decoding module 922. The video processing device 900 can implement Figure 3 、 Figure 5 、 Figure 6 and Figure 7 The functions of the sending end 210 and the receiving end 220.

[0193] When the video processing device 900 is used to implement Figure 3 In the method embodiment shown, when the transmitting end 210 performs the functions, the communication module 910 is used to implement S310 , and the processing module 920 is used to implement S330 .

[0194] When the video processing device 900 is used to implement Figure 3 In the method embodiment shown, when the receiving end 220 performs the functions, the communication module 910 is used to implement S310 and S330, and the processing module 920 is used to implement S320 and S340.

[0195] When the video processing device 900 is used to implement Figure 5 In the method embodiment shown, when the transmitting end 210 performs the functions, the communication module 910 is used to implement S310 , and the processing module 920 is used to implement S330 .

[0196] When the video processing device 900 is used to implement Figure 5 In the embodiment of the method shown, the functions of the receiving end 220 are as follows: the communication module 910 is used to implement S310 and S3301 to S3303, the processing module 920 is used to implement S320 and S3401 to S3403. Specifically, the encoding module 921 and the decoding module 922 are used to collaboratively implement S3401 to S3403.

[0197] Optionally, the video processing device 900 may further include a storage module, which may be used to store the aforementioned arbitrary code streams or reference frames. It should be understood that this embodiment merely provides an exemplary division of the structure and functional modules of the video processing device 900, and this application does not impose any limitation on the specific division.

[0198] For a more detailed description of the video processing device 900, please refer to the Figures 3 to 8 The relevant description in the method embodiment shown is directly obtained and will not be repeated here.

[0199] Figure 10 This is a schematic diagram of the structure of a terminal device provided in the present application. The terminal device 1000 includes a processor 1010 and a communication interface 1020. The processor 1010 and the communication interface 1020 are coupled to each other. It is understood that the communication interface 1020 can be a transceiver or an input / output interface. Optionally, the terminal device 1000 may also include a memory 1030 for storing instructions executed by the processor 1010, or storing input data required by the processor 1010 to execute instructions, or storing data generated after the processor 1010 executes instructions.

[0200] When the terminal device 1000 is used to implement Figure 3 、 Figure 5 、 Figure 6 or Figure 7 When performing the method shown in FIG. 1 , the processor 1010, the communication interface 1020 and the memory 1030 can also cooperate to implement the various operation steps in the video processing method executed by the transmitting end or the receiving end. The terminal device 1000 can also execute Figure 9 The functions of the video processing device 900 shown are not described in detail here.

[0201] The specific connection medium between the communication interface 1020, the processor 1010 and the memory 1030 is not limited in the embodiment of the present application. Figure 10 The communication interface 1020, the processor 1010 and the memory 1030 are connected via a bus 1040. Figure 10 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0202] The memory 1030 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video processing method provided in the embodiments of the present application. The processor 1010 executes the software programs and modules stored in the memory 1030 to perform various functional applications and video processing. The communication interface 1020 can be used to communicate signaling or data with other devices. In this application, the terminal device 1000 can have multiple communication interfaces 1020.

[0203] The memory in the embodiments of the present application can be a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium known in the art.

[0204] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), a neural processing unit (NPU) or a graphics processing unit (GPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0205] For example, a storage medium may be coupled to a processor, such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be an integral part of the processor. The processor and storage medium may be located in an ASIC. Alternatively, the ASIC may be located in a network device or a terminal device.

[0206] In some possible examples, the video processing device of the present application can be a video codec, the encoder can be in the form of a software encoder or a hardware chip and run on mobile terminal devices such as smartphones, tablets, laptops, and general-purpose computers, microcomputing devices, and servers, and the decoder can be in the form of a software decoder or a hardware chip, such as the decoder can exist in VR viewing devices such as mobile terminal devices, microcomputing devices, VR helmets, VR glasses, etc.

[0207] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).

[0208] In the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0209] The terms "first", "second" and "third" in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects rather than to limit a specific order.

[0210] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0211] "Multiple" means two or more, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, for elements (element) in the singular form "a", "an" and "the", unless the context clearly stipulates otherwise, it does not mean "one or only one", but means "one or more than one". For example, "a device" means one or more such devices. Furthermore, at least one (at least one of)..." means one or any combination of the subsequent associated objects, for example, "at least one of A, B and C" includes A, B, C, AB, AC, BC, or ABC. In the text description of this application, the character " / " generally indicates that the previous and next associated objects are in an "or" relationship; in the formula of this application, the character " / " indicates that the previous and next associated objects are in a "division" relationship.

[0212] It is understood that the various numbers used in the embodiments of this application are merely for ease of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not necessarily imply a specific order of execution; the order of execution of the processes should be determined by their functions and inherent logic.

Claims

1. A video processing method, characterized in that: Applied to a receiving end, the method includes: Obtaining a base layer image of a source video according to the first code stream; Obtaining perspective switching information, and obtaining to-be-displayed content information of the source video according to the perspective switching information, wherein the to-be-displayed content information includes a reference frame, the reference frame being an inter-frame prediction frame obtained based on the first bitstream, the reference frame corresponding to a partial perspective area after the perspective switching, and the to-be-displayed content information further including at least one second bitstream; The reference frame and the at least one second code stream are decoded based on the base layer image to obtain an enhancement layer image of the source video, where the video quality of the enhancement layer image is higher than the video quality of the base layer image.

2. The method according to claim 1, characterized in that Decoding the to-be-displayed content information based on the base layer image to obtain an enhancement layer image of the source video includes: Acquire, based on the base layer image, a first image of the partial viewing area after the viewing angle is switched, wherein a timestamp of the first image matches a timestamp of the reference frame; Obtaining a second image based on the first image and the reference frame, where the second image corresponds to a partial viewing angle area after the viewing angle is switched in the enhancement layer image; The enhancement layer image is obtained according to the second image and the at least one second code stream.

3. The method according to claim 2, characterized in that Obtaining a second image based on the first image and the reference frame, comprising: Encoding the first image to obtain a video frame, wherein the video frame is an intra-frame coded frame; The second image is obtained by decoding the reference frame based on the video frame.

4. The method according to claim 2 or 3, characterized in that Obtaining the enhancement layer image according to the second image and the at least one second code stream includes: splicing the at least one second code stream to obtain an application code stream; The enhancement layer image is obtained according to the second image and the application code stream.

5. The method according to any one of claims 1 to 3, characterized in that Acquiring the to-be-displayed content information of the source video according to the perspective switching information, including: Obtaining a partial viewing angle area after the viewing angle is switched according to the current viewing angle information and the viewing angle switching information; Sending a first instruction including an identifier, where the identifier is used to indicate a position of the partial viewing angle area after the viewing angle is switched in the base layer image; The reference frame is acquired in response to the first instruction.

6. The method according to claim 5, characterized in that The obtaining of the partial viewing area after the viewing angle switching according to the current viewing angle information and the viewing angle switching information includes: Determining a plurality of first sub-region identifiers according to the current viewing angle information and a preset first relationship, wherein the first relationship is used to indicate sub-region division information in the source video and an identifier of each sub-region; determining a plurality of second sub-region identifiers according to the perspective switching information and the first relationship; At least one sub-region identifier after the perspective switching is obtained according to the multiple first sub-region identifiers and the multiple second sub-region identifiers. The sub-region indicated by the at least one sub-region identifier after the perspective switching is the partial perspective region after the perspective switching.

7. The method according to any one of claims 1, 2, 3 or 6, characterized in that The video quality includes at least one of image signal-to-noise ratio, resolution and frame rate, or a combination of several of them.

8. A video processing method, characterized in that: Applied to a sending end, the method includes: Sending a first code stream of a source video to a receiving end, where the first code stream is used to indicate a base layer image of the source video; Obtaining perspective switching information, and sending content information to be displayed of the source video to a receiving end based on the perspective switching information, the content information to be displayed is used to indicate an enhancement layer image corresponding to a perspective area after the perspective switching in the source video, wherein the video quality of the enhancement layer image is higher than the video quality of the base layer image, the content information to be displayed includes a reference frame, which is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to a partial perspective area after the perspective switching. The content information to be displayed also includes at least one second code stream, and the at least one second code stream corresponds to other perspective areas in the enhancement layer image except the partial perspective area after the perspective switching.

9. The method according to claim 8, characterized in that Before sending the to-be-displayed content information of the source video to the receiving end, the method further includes: Obtaining a third stream of the partial viewing area after the viewing angle is switched according to the source video; The reference frame is obtained according to the first code stream and the third code stream.

10. The method according to claim 9, characterized in that Obtaining the reference frame according to the first code stream and the third code stream includes: Acquire, according to the first code stream, a first image of the partial viewing area after the viewing angle is switched, wherein a timestamp of the first image matches a timestamp of the viewing angle switching; The reference frame is obtained by encoding a second image based on the first image, where the second image is an image of the partial viewing area in the source video, or an image obtained by decoding the third code stream, and the timestamp of the second image matches the timestamp of the first image.

11. The method according to claim 10, characterized in that Encoding the second image based on the first image to obtain the reference frame includes: Encoding the first image to obtain a video frame, wherein the video frame is an intra-frame coded frame; The second image is encoded based on the video frame to obtain the reference frame.

12. The method according to any one of claims 9 to 11, characterized in that Obtaining a third stream of the partial viewing area after the viewing angle switching according to the source video includes: Dividing the source video into multiple sub-regions; Encoding the video corresponding to each of the multiple sub-regions to obtain multiple sub-region code streams; Match the partial viewing area after the viewing angle switching with a preset first relationship to determine multiple sub-area identifiers, where the first relationship is used to indicate the sub-area division information and the identifier of each sub-area in the source video, and the sub-area code streams corresponding to the multiple sub-area identifiers are the third code streams.

13. The method according to claim 8, characterized in that The method further comprises: Encoding a video corresponding to each of the multiple sub-regions in the source video to obtain multiple sub-region code streams; Obtaining a sub-region reference image sequence within a first time range according to the first code stream, wherein the sub-region reference image sequence includes a reference image of each sub-region; A reference code stream is obtained by encoding a sequence of images to be encoded based on the sub-region reference image sequence, where the reference code stream includes the reference frame, the sequence of images to be encoded includes multiple images to be encoded, and the images to be encoded are sub-region images in the source video that match the reference image, or images obtained by decoding a code stream in the sub-region code stream that matches the reference image, and the timestamps of the images to be encoded match those of the reference images.

14. The method according to claim 13, characterized in that Encoding a to-be-encoded image sequence based on the sub-region reference image sequence to obtain a reference bitstream, comprising: Interleaving and recombining the sub-region reference image sequence and the to-be-encoded image sequence to obtain first recombined data and second recombined data, wherein the first recombined data includes the sub-region reference image and the to-be-encoded image whose sub-region identifiers meet a first condition, and the second recombined data includes the sub-region reference image and the to-be-encoded image whose sub-region identifiers meet a second condition; obtaining a fourth bitstream according to the first recombined data, the fourth bitstream including all reference frames within the first time range whose sub-region identifiers meet the first condition; obtaining a fifth bitstream according to the second reorganized data, the fifth bitstream including all reference frames within the first time range whose sub-region identifiers meet the second condition; The reference bitstream is obtained according to the fourth bitstream and the fifth bitstream.

15. The method according to any one of claims 8, 9, 10, 11, 13 or 14, characterized in that The video quality includes at least one of image signal-to-noise ratio, resolution and frame rate, or a combination of several of them.

16. A video processing device, characterized in that: include: A processing module, configured to obtain a base layer image of a source video according to the first code stream; a communication module, configured to obtain perspective switching information, and obtain content information to be displayed of the source video based on the perspective switching information, wherein the content information to be displayed includes a reference frame, the reference frame being an inter-frame prediction frame obtained based on the first bitstream, the reference frame corresponding to a partial perspective area after the perspective switching, and the content information to be displayed further includes at least one second bitstream; The processing module is further configured to decode the reference frame and the at least one second code stream based on the base layer image to obtain an enhanced layer image of the source video, wherein the video quality of the enhanced layer image is higher than the video quality of the base layer image.

17. A video processing device, characterized in that: include: A communication module, configured to send a first code stream of a source video to a receiving end, wherein the first code stream is used to indicate a base layer image of the source video; a processing module, configured to obtain perspective switching information, and send content information to be displayed of the source video to a receiving end according to the perspective switching information; The content information to be displayed is used to indicate the enhanced layer image corresponding to the viewing area after the viewing angle is switched in the source video, the video quality of the enhanced layer image is higher than the video quality of the basic layer image, the content information to be displayed includes a reference frame, the reference frame is an inter-frame prediction frame obtained based on the first code stream, and the reference frame corresponds to the partial viewing area after the viewing angle is switched. The content information to be displayed also includes at least one second code stream, and the at least one second code stream corresponds to other viewing areas in the enhanced layer image except the partial viewing area after the viewing angle is switched.

18. A video processing system, characterized in that: include: sender and receiver; The transmitting end sends a first code stream of a source video to the receiving end, where the first code stream is used to indicate a base layer image of the source video; The receiving end obtains the base layer image of the source video according to the first code stream; The receiving end obtains the view switching information, and obtains the to-be-displayed content information of the source video from the receiving end according to the view switching information; The content information to be displayed includes a reference frame, the reference frame is an inter-frame prediction frame obtained based on the first bitstream, the reference frame corresponds to a partial viewing area after the viewing angle is switched, and the content information to be displayed also includes at least one second bitstream; The receiving end decodes the reference frame and the at least one second code stream based on the base layer image to obtain an enhanced layer image of the source video, where the video quality of the enhanced layer image is higher than the video quality of the base layer image.

19. A terminal device, characterized in that: The device comprises a processor and an interface circuit, wherein the interface circuit is used to receive signals from terminal devices other than the terminal device and transmit them to the processor, or to send signals from the processor to terminal devices other than the terminal device, and the processor is used to implement the method according to any one of claims 1 to 7 or any one of claims 8 to 15 through a logic circuit or executing code instructions.

20. A computer-readable storage medium, characterized in that The storage medium stores a computer program or instruction. When the computer program or instruction is executed by a terminal device, the method according to any one of claims 1 to 7 or any one of claims 8 to 15 is implemented.

Citation Information

Patent Citations

  • Method and an apparatus and a computer program for encoding media content

    CN109155861A