Method and apparatus for decoding and encoding image or video data
Patent Information
- Application Number
- PCT/EP2026/056704
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-11
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026056704_01102026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND APPARATUS FOR DECODING AND ENCODING IMAGE OR VIDEO DATA
[0002] TECHNICAL FIELD
[0003] Embodiments of the present disclosure generally relate to the field of encoding and decoding image or video data from a bitstream. In particular, some embodiments relate to methods and apparatuses for such encoding and decoding images and / or videos captured using multiple lenses.
[0004] BACKGROUND
[0005] Image and video codecs have been used for decades to compress image and video data. In such codecs, a signal is typically encoded block-wisely by predicting a block and by further coding only the difference between the original bock and its prediction. For example, in video encoding, this may involve identifying similarities between adjacent / successive frames in the video. In particular, such coding may include transformation, quantization and generating the bitstream, usually including some entropy coding. Typically, the three components of hybrid coding methods - transformation, quantization, and entropy coding - are separately optimized. Modem video compression standards like High-Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use transformed representation to code residual signal after prediction. Other video compression standards still in use include Advanced Video Coding (AVC) standard.
[0006] AVC has been defined by the H.264 video coding standard, HEVC has been defined by the H.265 video coding standard, and VVC has been defined by the H.266 video coding standard. Each of these video compression standards define a compression standard to be implemented by a suitable codec (where a codec is a software or hardware implementation that performs the compression as defined by the standard). Codecs therefore transform and compress a raw image or video data input into a bitstream, that may then be decoded by a suitable decoder.
[0007] ‘Supplemental Enhancement Information’, SEI, messages may be included in bitstreams to further enhance the data when it is being decoded by a user. SEU signalling is known from its specification in at least the AVC and HEVC standards. Specifically, SEI is metadata, that can be included into a bitstream, to convey extra information. For example, SEI can be inserted during the encoding and transmission of the relevant content. The types of metadata that SEI may contains include parameters of the camera or encoder such as time code; closed captions; lyrics (for audio).
[0008] A newer standard that includes yet more metadata types suitable for use in conjunction with the above-named video compression standards is the ‘Versatile Supplemental Enhancement Information’ . VSEI is partner specification for, and thus specifically designed for use with, the VVC encoding standard (defined in the H.266 video coding standard specification). The VSEI standard contains more nuanced and potentially richer information for devices and videos types that extend beyond the standard or traditional video coding. For example, VSEI can contain specialist information to help adapt decoded video for a certain display type. For example, VSEI introduces two SEI messages for neural-network-based post-processing. VSEI may also contain metadata for cert certain camera or lens types, to help display an image or video correctly whilst accounting for distortions that a camera or lens may have impacted.
[0009] Although current VSEI specification design supports the signaling of distortion parameters of lenses, as more modem video and camera technologies emerge, yet further capabilities for encoding and display may be necessary. New VSEI specifications should also be prepared to handle any encoding scenarios required by future video encoding standards that will be defined by the upcoming H.267 video coding standard.The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0010] Particular embodiments are outlined in the attached independent claims, with other embodiments in the dependent claims.
[0011] SUMMARY
[0012] Particular embodiments are outlined in the attached independent claims, with other embodiments in the dependent claims.
[0013] According to a first aspect, the present disclosure relates to a decoding method, implemented by a decoder, for decoding encoded image or video data, the method comprising:
[0014] receiving a bitstream including an encoded input signal defining image or video data,
[0015] wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;
[0016] parsing the bitstream, the parsing comprising:
[0017] identifying, in the bitstream, a supplementary signal comprising perspective information configured to allow the decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective;
[0018] in dependence on the perspective information, deriving at least two image or video output signals from the encoded input signal.
[0019] The foregoing and other embodiments can each optionally include one or more of the following features, alone or in combination.
[0020] In a possible implementation, the perspective information comprises a parameter indicating a number of unique lenses that contribute to the plurality of visual signals, wherein each unique lens indicated by the perspective information is associated with a unique camera perspective.
[0021] In a possible implementation, the perspective information identifies one or more lens attributes for each unique lens.
[0022] In a possible implementation, for each unique lens, the one or more lens attributes comprise attributes selected from the group of: focal parameters; focal length; radial distortion parameters; horizontal and / or vertical location in the image or video of the focal point; transversal chromatic aberration parameters; vignetting parameters.
[0023] In a possible implementation, the one or more lens attributes are encoded as an array comprising a plurality of elements, wherein each element of the plurality of elements corresponds to a unique lens.
[0024] In a possible implementation, the image or video data represents two visual signals captured by a single camera device comprising two unique lenses, wherein each visual signal of the two visual signals is captured by a unique lens on the single camera device.
[0025] In a possible implementation, the single camera device comprises a single sensor for capturing the image or video data.In a possible implementation, each lens of the two unique lenses is oriented in a same direction, and where each unique lens is configured to capture a different visual perspective.
[0026] In a possible implementation, the image or video data is configured to be output on a stereo ocular display, wherein the stereo ocular display comprises two output displays, one display per eye.
[0027] In a possible implementation, the stereo ocular display comprises a virtual reality, VR, headset.
[0028] In a possible implementation, the supplementary signal is defined according to the Versatile Supplemental Enhancement Information, VSEI, standard specification.
[0029] According to a second aspect, the present disclosure relates a device for decoding encoded image or video data, the device comprising:
[0030] a receiving unit configured to:
[0031] receive a bitstream including an encoded input signal defining image or video data,
[0032] wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;
[0033] a parsing unit configured to parse the bitstream, the parsing comprising:
[0034] parsing the bitstream, the parsing unit configured to:
[0035] identify, in the bitstream, a supplementary signal comprising perspective information configured to allow the decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective;
[0036] in dependence on the perspective information, deriving at least two image or video output signals from the encoded input signal.
[0037] In a possible implementation the device is further configured to output the two image or video output signals to two different displays.
[0038] According to a third aspect, the present disclosure relates to an encoding method, implemented by an encoder, for encoding a plurality of visual signals, the method comprising:
[0039] receiving an input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective;
[0040] receiving perspective information that identifies each unique camera perspective;
[0041] encoding the input signal into a bitstream, wherein the encoding comprises:
[0042] encoding the input signal into a bitstream;
[0043] in dependence on the perspective information, including a supplementary signal into the bitstream configured to indicate, to a decoder, i) information identifying each unique camera perspective and ii) that the bitstream is to be decoded into at least two image or video output signals.
[0044] In a possible implementation the perspective information included in the supplementary signal comprises a parameter indicating a number of unique lenses that contribute to the plurality of visual signals, wherein each unique lens indicated by the perspective information is associated with a unique camera perspective.
[0045] In a possible implementation, the perspective included in the information supplementary signal identifies one or more lens attributes for each unique lens.In a possible implementation, for each unique lens, the one or more lens attributes comprise attributes selected from the group of: focal parameters; focal length; radial distortion parameters; horizontal and / or vertical location in the image or video of the focal point; transversal chromatic aberration parameters; vignetting parameters.
[0046] In a possible implementation, the method further comprises encoding the one or more lens as an array comprising a plurality of elements, wherein each element of the plurality of elements corresponds to a unique lens.
[0047] According to a fourth aspect, the present disclosure relates to a device for encoding a plurality of visual signals, the device comprising:
[0048] a receiving unit configured to receive: i) an input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective; and ii) perspective information that identifies each unique camera perspective;
[0049] an encoding unit configured to:
[0050] encode the input signal into a bitstream;
[0051] in dependence on the perspective information, include a supplementary signal into the bitstream configured to indicate, to a decoder, i) information identifying each unique camera perspective and ii) that the bitstream is to be decoded into at least two image or video output signals.
[0052] According to a fifth aspect, the present disclosure relates to a decoding apparatus comprising processing circuitry configured to execute steps of any of the methods disclosed herein.
[0053] According to a sixth aspect, the present disclosure relates to an encoding apparatus comprising processing circuitry configured to execute steps of any of the methods disclosed herein.
[0054] According to a seventh aspect, the present disclosure relates to a decoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors, wherein the storage medium stores programming for execution by the one or more processors, wherein the programming, when executed by the one or more processors, configures the decoder to carry out any of the methods disclosed herein.
[0055] According to an eighth aspect, the present disclosure relates to a non-transitory storage medium comprising a bitstream encoded by any of the methods disclosed herein.
[0056] According to a ninth aspect, the present disclosure relates to a computer program stored on a non-transitory medium and including code instructions, which, when executed on one or more processor, causes the one or more processor to execute any of the methods disclosed herein.
[0057] According to a tenth aspect, the present disclosure relates to a system for delivering a bitstream, the system including at least one storage medium configured to store at least one bitstream generated by the encoding method described in any of the methods disclosed herein.
[0058] According to an eleventh aspect, the present disclosure relates to a system for delivering a bitstream, the system comprising:
[0059] at least one storage medium configured to store at least one bitstream generated by any of the methods disclosed herein.; and
[0060] a video streaming device configured to obtain the bitstream from one of the at least one storage medium and send the bitstream to a terminal device, wherein the video streaming device comprises a content server or a content delivery server.In a possible implementation, the system further comprises one or more processors configured to perform encryption processing on at least one bitstream to obtain at least one encrypted bitstream, the at least one storage medium configured to store the encrypted bitstream; or
[0061] the one or more processors configured to convert the bitstream in a first format into a bitstream in a second format, the at least one storage medium configured to store the bitstream in the second format.
[0062] In a possible implementation, the system further comprises:
[0063] a receiver configured to receive a first operation request; wherein the one or more processor is configured to determine a target bitstream in the at least one storage medium in response to the first operation request; and
[0064] a transmitter configured to send the target bitstream to a terminal-side apparatus.
[0065] In a possible implementation, the one or more processors is further configured to encapsulate the bitstream to obtain a transport stream in a first format, wherein the transmitter is further configured to:
[0066] send the transport stream in the first format to a terminal-side apparatus for display; or
[0067] send the transport stream in the first format to storage space for storage.
[0068] According to an twelfth aspect, the present disclosure relates to a non-transitory computer-readable medium having processorexecutable instructions stored thereon for reconstructing an image or video from a bitstream, wherein the bitstream includes an encoded input signal defining image or video data, the image or video data representing a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective;
[0069] wherein the processor-executable instructions, when executed, cause performance of the following:
[0070] identifying, in the bitstream, a supplementary signal comprising perspective information of the image or video data;
[0071] interpreting, in dependence on the supplementary signal, that the perspective information identifies each unique camera perspective;
[0072] in dependence on the identified unique camera perspectives, deriving at least two image or video output signals from the encoded input signal.
[0073] According to a thirteenth aspect, the present disclosure a bitstream, characterized by comprising:
[0074] encoded input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;
[0075] a supplementary signal comprising perspective information configured to allow a decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective
[0076] In one possible embodiment, the system further includes: one or more processor, configured to perform encryption processing on at least one bitstream to obtain at least one encrypted bitstream; the at least one storage medium, configured to store the encrypted bitstream.
[0077] In a possible implementation the system further comprises one or more processor configured to perform encryption processing on at least one bitstream to obtain at least one encrypted bitstream, the at least one storage medium configured to store the encrypted bitstream; or the one or more processor configured to convert the bitstream in a first format into a bitstream in a second format, the at least one storage medium configured to store the bitstream in the second format. The storage medium can then store a usefully encoded form of data.In a possible implementation the system further comprises: a receiver configured to receive a first operation request; wherein the one or more processor is configured to determine a target bitstream in the at least one storage medium in response to the first operation request; a transmitter configured to send the target bitstream to a terminal-side apparatus. The target bitstream can then be processed by the terminal side apparatus.
[0078] In a possible implementation the one or more processor is further configured to encapsulate the bitstream to obtain a transport stream in a first format, wherein the transmitter is further configured to: send the transport stream in the first format to a terminal-side apparatus for display; or send the transport stream in the first format to storage space for storage. In this way the transport stream can be processed as appropriate.
[0079] BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In the following embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which
[0081] Fig. 1 illustrates an example of a camera device having multiple lenses configured to capture image or video for encoding consistent with presently disclosed embodiments;
[0082] Fig. 2 shows a table lustring an example syntax for a supplementary signal for incorporation in a bitstream for indicating a plurality of image or video signals consistent with presently disclosed embodiments; Fig. 3 shows an example of a bitstream according to presently disclosed embodiments;
[0083] Fig. 4 is a block diagram showing an example of a video coding system configured to implement embodiments of the present disclosure.
[0084] Fig. 5 is a block diagram showing another example of a video coding system configured to implement embodiments of the present disclosure.
[0085] Fig. 6 is a block diagram illustrating an example of an encoding apparatus or a decoding apparatus.
[0086] Fig. 7 is a block diagram illustrating another example of an encoding apparatus or a decoding apparatus.
[0087] Fig. 8 is a block diagram illustrating another example of an encoding apparatus or a decoding apparatus.
[0088] Like reference numbers and designations in different drawings may indicate similar elements.
[0089] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0090] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0091] For instance, it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performingthe functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.
[0092] In the following, an overview of some of the used technical terms and framework within which the embodiments of the present disclosure may be employed is provided.
[0093] Versatile supplemental enhancement information, VSEI, messages maybe used for coded video bitstreams, e.g., to modify existing SEI messages, specify additional SEI messages, or specify other functionalities. VSEI are used in conjunction with the VVC video compression standard. The VSEI specification design supports the signaling of distortion parameters of lenses. The reason for this is that, in order to successfully dewarp an image or video and represent / display it in the proper way (i.e., with a perspective that represents the original perspective of the scene from which the image or video was captured), a decoder needs information about lenses, e.g., lens attributes.
[0094] However, the current VSEI standard is not configured to signal lens properties for all types of modem lenses and camera devices, nor is it configured to decode bitstream containing image or video data for modem displays, such as virtual reality displays comprising two (or more) stereo output displays (e.g., one display for each eye). Consequently, the present inventors have found a method of adapting the VSEI signalling protocol and syntax to enable modem cameras, such as those containing multiple lenses, to be signalled using VSEI consistent with the current VVC video compression protocol and / or the nextgeneration video compression as defined in the future H.267 compression specification.
[0095] Figure 1 illustrates a camera device 100 having two lenses. Such a camera device’s lens is currently not supported by the VSEI signaling syntax and protocol, because multiple lenses are not supported. The camera device 100 has two lenses, a left lens 102a and a right lens 102b, which each project an image onto one sensor (not shown). For example, each lens might be a fish eye lens in order to gain a wide perspective. The two lenses shown in the device 100 have the same orientation. However, irrespective of the attributes of the lenses and their orientation, due to the difference in position of the two lenses, the images projected two lenses will form a different perspective. Consequently, decoding the image projected onto the sensor by the two lenses 102a, 102b, will involve decoding the parallax error (i.e., the apparent difference in displacement of any object in the scene viewed by the lenses). In other words, the signal projected onto the camera sensor by the two lenses will comprise two distinct visual signals, where each visual signal has been captured from a different unique camera perspective corresponding to each lens.
[0096] Furthermore, the camera in Figure 1 is intended to record 3D perspective images or video intended for output via two separate displays, i.e., in a stereo ocular display, wherein the stereo ocular display comprises two output displays, one display per eye, such as a virtual reality, VR, headset. VSEI currently does not allow the decoding of video or image signals for output onto two separate displays.
[0097] In order for a decoder to properly decode an image or video recorded by such a multi-lens camera, according to the solution provided in the present disclosure, information about each lens is encoded in a supplementary signal forming part of the bitstream encoding of the image or video signal.
[0098] Fig. 2 shows an example syntax table 200 for a supplementary signal for incorporation in a bitstream for indicating a plurality of image or video signals consistent with presently disclosed embodiments. The supplementary signal may be according to the SEI or VSEI standard. This example contains lens optical correction messages, which provide the decoder with attributes thatenable the decoder to infer a lens distortion model, and thereby enable image correction. When parameters for multiple distortion types are available in the lens optical correction SEI message, the parameters may be expressed in a particular order, such that the decoder is able to infer which attribute is being decoded without the attribute type being explicitly signalled.
[0099] In the known VSEI syntax, lens attribute data may only be conveyed using scalar values. In order words, for each lens attribute only one value or set of values is allowed, corresponding to a single lens. This is because the known VSEI standard only allows the signalling of lens attribute data for a single lens.
[0100] Consequently, the present inventors have established that lens attribute data for multiple lenses may be signalled using a ‘for’ loop, which iterates over different lens attributes which are expressed in arrays. Specifically, for each lens attribute which is expressed, the lens attribute array comprises a plurality of elements, wherein each element of the plurality of elements corresponds to a unique lens. Thus, the ‘for loop’ iterates over each lens, and extracts the appropriate value for each lens attribute corresponding to each lens.
[0101] In Fig. 2, the “lens_number” attribute specifies the number of lenses that have projections on the decoded picture. For example, if the lens according to Fig. 1 were being encoded into a bitstream using the VSEI syntax, the “lens_number” attribute would indicate two different lenses. Then, for each attribute, the attribute indicates the appropriate attribute for the “ith” lens. The corresponding “descriptor” field identified the format and size of the attribute, e.g., in a number of bits or bytes.
[0102] For example, the “lens_attribute_l [i]” indicates the value of “attribute_l” for the “ith” lens. The lens attribute can be any suitable attribute that defines the characteristics of the lens. Merely as an example, the one or more lens attributes may be selected from the group of: focal parameters; focal length; radial distortion parameters; horizontal and / or vertical location in the image or video of the focal point; transversal chromatic aberration parameters; and vignetting parameters.
[0103] It will be understood that more than three lens attributes can be conveyed, and indeed an arbitrary number of lens attributes may be conveyed by the VSEI signalling.
[0104] When decoding the bitstream, the decoder therefore performed the loop as indicated in Fig. 2, i.e., “For ( i = 0; i <= lens_number; i++ ) {”. This involves iterating over every lens attribute for every number of lenses indicated by the “lens number” attribute. The elements in each array of each lens attributes correspond to the respective lenses. For example, when i=0, lens attributes for the first of the lenses is read from the lens attribute arrays. When i= 1 , lens attributes for the second of the lenses is read from the lens attribute arrays, and so on.
[0105] Consequently, lens correction parameters can be determined via the VSEI syntax messages in order to properly display the image or video data as recorded by the plurality of lenses. For example, for the dual lens shown in Fig. 1, the different perspectives from each of the two lenses can be properly extracted from the video or image data using the VSEI syntax shown in Fig. 2 Thus, the raw video or image data, once decoded from the bitstream, can be properly de- warped and to extracted to obtain two separate images or video data, each separate image / video corresponding to the image / video recorded via each separate lens 102a, 102b. In other examples, the camera device may contain more than two lenses, and more than two corresponding image or video outputs may be extracted.
[0106] It will also be appreciated that other flags and attributes, other than les characteristics, may be included in the SEI / VSEI message. For example, flags may be included to indicate that the current SEI message cancels the persistence of any previous lens optical correction SEI message. Other flags may also specify the persistence of the lens optical correction SEI message forthe current layer, i.e., whether the lens optical correction SEI message applies to the current decoded picture only or persists for all subsequent pictures of the current layer until certain conditions are met (e.g., the bitstream ends).
[0107] Bitstream structure
[0108] Fig. 3 shows an example structure of a bitstream 300 comprising substream indicating a start of the bitstream, the picture / video header, the entropy encoded data substream (i.e., the main ‘payload’ of the bitstream comprising the picture / video data), a portion comprising the supplemental signal of the bitstream, and a portion indicating the end of the picture / video. In embodiments, the supplemental signal is a signal according to the Versatile Supplemental Enhancement Information, VSEI, signalling specification for signalling perspective information configured to allow the decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective. Thus, the supplemental signal may contain information such as the SEI messaging described in respect of Fig. 2.
[0109] Some exemplary implementations in hardware and software
[0110] The corresponding system which may deploy the above-mentioned encoder-decoder processing chain is illustrated in Fig. 4. Fig. 4. is a schematic block diagram illustrating an example coding system, e.g. a video, image, audio, and / or other coding system (or short coding system) that may utilize techniques of this present application. Video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of video coding system 10 represent examples of devices that may be configured to perform techniques in accordance with various examples described in the present application. For example, the video coding and decoding may employ neural network such which may be distributed and which may apply the above-mentioned bitstream parsing and / or bitstream generation to convey feature maps between the distributed computation nodes (two or more).
[0111] As shown in Fig. 4., the coding system 10 comprises a source device 12 configured to provide encoded picture data 21 e.g. to a destination device 14 for decoding the encoded picture data 13.
[0112] The source device 12 comprises an encoder 20, and may additionally, i.e. optionally, comprise a picture source 16, a preprocessor (or pre-processing unit) 18, e.g. a picture pre-processor 18, and a communication interface or communication unit 22.
[0113] The picture source 16 may comprise or be any kind of picture capturing device, for example a camera for capturing a real-world picture, and / or any kind of a picture generating device, for example a computer-graphics processor for generating a computer animated picture, or any kind of other device for obtaining and / or providing a real-world picture, a computer generated picture (e.g. a screen content, a virtual reality (VR) picture) and / or any combination thereof (e.g. an augmented reality (AR) picture). The picture source may be any kind of memory or storage storing any of the aforementioned pictures.
[0114] In distinction to the pre-processor 18 and the processing performed by the pre-processing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.
[0115] Pre-processor 18 is configured to receive the (raw) picture data 17 and to perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. Pre-processing performed by the pre-processor 18 may, e.g., comprise trimming, color format conversion (e.g. from RGB to YCbCr), color correction, or de-noising. It can be understood that the pre-processing unit 18 may be optional component. It is noted that the pre-processing may also employ a neural network (such as in any of Figs. 1 to 7) which uses the presence indicator signaling.
[0116] The video encoder 20 is configured to receive the pre-processed picture data 19 and provide encoded picture data 21.Communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and to transmit the encoded picture data 21 (or any further processed version thereof) over communication channel 13 to another device, e.g. the destination device 14 or any other device, for storage or direct reconstruction.
[0117] The destination device 14 comprises a decoder 30 (e.g. a video decoder 30), and may additionally, i.e. optionally, comprise a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32) and a display device 34.
[0118] The communication interface 28 of the destination device 14 is configured receive the encoded picture data 21 (or any further processed version thereof), e.g. directly from the source device 12 or from any other source, e.g. a storage device, e.g. an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.
[0119] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g. a direct wired or wireless connection, or via any kind of network, e.g. a wired or wireless network or any combination thereof, or any kind of private and public network, or any kind of combination thereof.
[0120] The communication interface 22 may be, e.g., configured to package the encoded picture data 21 into an appropriate format, e.g. packets, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.
[0121] The communication interface 28, forming the counterpart ofthe communication interface 22, may be, e.g., configured to receive the transmitted data and process the transmission data using any kind of corresponding transmission decoding or processing and / or de-packaging to obtain the encoded picture data 21.
[0122] Both, communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces as indicated by the arrow for the communication channel 13 in Fig. M-4 pointing from the source device 12 to the destination device 14, or bi-directional communication interfaces, and may be configured, e.g. to send and receive messages, e.g. to set up a connection, to acknowledge and exchange any other information related to the communication link and / or data transmission, e.g. encoded picture data transmission. The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or a decoded picture 31.
[0123] The post-processor 32 of destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), e.g. the decoded picture 31, to obtain post-processed picture data 33, e.g. a post-processed picture 33. The post-processing performed by the post-processing unit 32 may comprise, e.g. color format conversion (e.g. from YCbCr to RGB), color correction, trimming, or re-sampling, or any other processing, e.g. for preparing the decoded picture data 31 for display, e.g. by display device 34.
[0124] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g. to a user or viewer. The display device 34 may be or comprise any kind of display for representing the reconstructed picture, e.g. an integrated or external display or monitor. The displays may, e.g. comprise liquid crystal displays (LCD), organic light emitting diodes (OLED) displays, plasma displays, projectors , micro LED displays, liquid crystal on silicon (LCoS), digital light processor (DLP) or any kind of other display.
[0125] Although Fig. 4 depicts the source device 12 and the destination device 14 as separate devices, embodiments of devices may also comprise both or both functionalities, the source device 12 or corresponding functionality and the destination device 14 orcorresponding functionality. In such embodiments the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0126] As will be apparent for the skilled person based on the description, the existence and (exact) split of functionalities of the different units or functionalities within the source device 12 and / or destination device 14 as shown in Fig. 4 may vary depending on the actual device and application.
[0127] The encoder 20 (e.g. a video encoder 20) or the decoder 30 (e.g. a video decoder 30) or both encoder 20 and decoder 30 may be implemented via processing circuitry, such as one or more microprocessors, digital signal processors (DSPs), applicationspecific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video coding dedicated or any combinations thereof. The encoder 20 may be implemented via processing circuitry 46 to embody the various modules including the neural network or its parts. The decoder 30 may be implemented via processing circuitry 46 to embody any coding system or subsystem described herein. The processing circuitry may be configured to perform the various operations as discussed later. If the techniques are implemented partially in software, a device may store instructions for the software in a suitable, non-transitory computer-readable storage medium and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either of video encoder 20 and video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, for example, as shown in Fig. 5.
[0128] Source device 12 and destination device 14 may comprise any of a wide range of devices, including any kind of handheld or stationary devices, e.g. notebook or laptop computers, mobile phones, smart phones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video gaming consoles, video streaming devices(such as content services servers or content delivery servers), broadcast receiver device, broadcast transmitter device, or the like and may use no or any kind of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0129] In some cases, video coding system 10 illustrated in Fig. 4 is merely an example and the techniques of the present application may apply to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between the encoding and decoding devices. In other examples, data is retrieved from a local memory, streamed over a network, or the like. A video encoding device may encode and store data to memory, and / or a video decoding device may retrieve and decode data from memory. In some examples, the encoding and decoding is performed by devices that do not communicate with one another, but simply encode data to memory and / or retrieve and decode data from memory.
[0130] Fig. 6 is a schematic diagram of a video coding device 600 according to an embodiment of the disclosure. The video coding device 600 is suitable for implementing the disclosed embodiments as described herein. In an embodiment, the video coding device 600 may be a decoder such as video decoder 30 of Fig. 4 or an encoder such as video encoder 20 of Fig. 4.
[0131] The video coding device 600 comprises ingress ports 610 (or input ports 610) and receiver units (Rx) 620 for receiving data; a processor, logic unit, or central processing unit (CPU) 630 to process the data; transmitter units (Tx) 640 and egress ports 650 (or output ports 650) for transmitting the data; and a memory 660 for storing the data. The video coding device 600 may also comprise optical-to-electrical (OE) components and electrical- to-optical (EO) components coupled to the ingress ports 610, the receiver units 620, the transmitter units 640, and the egress ports 650 for egress or ingress of optical or electrical signals.The processor 630 is implemented by hardware and software. The processor 630 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 630 is in communication with the ingress ports 610, receiver units 620, transmitter units 640, egress ports 650, and memory 660. The processor 630 comprises a neural network based codec 670. The neural network based codec 670 implements the disclosed embodiments described above. For instance, the neural network based codec 670 implements, processes, prepares, or provides the various coding operations. The inclusion of the neural network based codec 670 therefore provides a substantial improvement to the functionality of the video coding device 600 and effects a transformation of the video coding device 600 to a different state. Alternatively, the neural network based codec 670 is implemented as instructions stored in the memory 660 and executed by the processor 630.
[0132] The memory 660 may comprise one or more disks, tape drives, and solid-state drives and may be used as an over-flow data storage device, to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. The memory 660 may be, for example, volatile and / or non-volatile and may be a read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0133] Fig. 7 is a simplified block diagram of an apparatus that may be used as either or both of the source device 12 and the destination device 14 from Fig. 4 according to an exemplary embodiment.
[0134] A processor 702 in the apparatus 700 can be a central processing unit. Alternatively, the processor 702 can be any other type of device, or multiple devices, capable of manipulating or processing information now-existing or hereafter developed. Although the disclosed implementations can be practiced with a single processor as shown, e.g., the processor 702, advantages in speed and efficiency can be achieved using more than one processor.
[0135] A memory 704 in the apparatus 700 can be a read only memory (ROM) device or a random access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 704. The memory 704 can include code and data 706 that is accessed by the processor 702 using a bus 712. The memory 704 can further include an operating system 708 and application programs 710, the application programs 710 including at least one program that permits the processor 702 to perform the methods described here. For example, the application programs 710 can include applications 1 through N, which further include a video coding application that performs the methods described here.
[0136] The apparatus 700 can also include one or more output devices, such as a display 718. The display 718 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 718 can be coupled to the processor 702 via the bus 712.
[0137] Although depicted here as a single bus, the bus 712 of the apparatus 700 can be composed of multiple buses. Further, a secondary storage can be directly coupled to the other components of the apparatus 700 or can be accessed via a network and can comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. The apparatus 700 can thus be implemented in a wide variety of configurations.
[0138] Fig. 8 is a block diagram of a video coding system 800 according to an embodiment of the disclosure.
[0139] A platform 802 in the system 800 can be could sever or local sever. Alternatively, the platform 802 can be any other type of device, or multiple devices, capable of calculation, storing, transcoding, encryption, rendering, decoding or encoding. Although the disclosed implementations can be practiced with a single platform as shown, e.g., the platform 802, advantages in speed and efficiency can be achieved using more than one platform.A content delivery network (CDN) 804 in the system 800 can be a group of geographically distributed servers. Alternatively, the CDN 804 can be any other type of device, or multiple devices, capable of data buffering, scheduling, dissemination or speed up the delivery of web content by bringing it closer to where users are. Although the disclosed implementations can be practiced with a single CDN as shown, e.g., the CDN 804, advantages in speed and efficiency can be achieved using more than one CDN.
[0140] A terminal 806 in the apparatus 800 can be a mobile phone, computer, television, laptop, camera. Alternatively, the terminal 806 can be any other type of device, or multiple devices, capable of displaying video or image.
Claims
CLAIMS1. A decoding method, implemented by a decoder, for decoding encoded image or video data, the method comprising:receiving a bitstream including an encoded input signal defining image or video data,wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;parsing the bitstream, the parsing comprising:identifying, in the bitstream, a supplementary signal comprising perspective information configured to allow the decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective;in dependence on the perspective information, deriving at least two image or video output signals from the encoded input signal.
2. The decoding method of claim 1 , wherein the perspective information comprises a parameter indicating a number of unique lenses that contribute to the plurality of visual signals, wherein each unique lens indicated by the perspective information is associated with a unique camera perspective.
3. The decoding method of claim 2, wherein the perspective information identifies one or more lens attributes for each unique lens.
4. The decoding method of claim 3, wherein, for each unique lens, the one or more lens attributes comprise attributes selected from the group of: focal parameters; focal length; radial distortion parameters; horizontal and / or vertical location in the image or video of the focal point; transversal chromatic aberration parameters; vignetting parameters.
5. The decoding method of claim 3 or 4, wherein the one or more lens attributes are encoded as an array comprising a plurality of elements, wherein each element of the plurality of elements corresponds to a unique lens.
6. The decoding method of any preceding claim, wherein the image or video data represents two visual signals captured by a single camera device comprising two unique lenses, wherein each visual signal of the two visual signals is captured by a unique lens on the single camera device.
7. The decoding method of claim 6, wherein the single camera device comprises a single sensor for capturing the image or video data.
8. The decoding method of claim 6 or 7, wherein each lens of the two unique lenses is oriented in a same direction, and where each unique lens is configured to capture a different visual perspective.
9. The decoding method of any of claims 6 to 8, wherein the image or video data is configured to be output on a stereo ocular display, wherein the stereo ocular display comprises two output displays, one display per eye.
10. The decoding method of claim 9, wherein the stereo ocular display comprises a virtual reality, VR, headset.
11. The decoding method of any preceding claim, wherein the supplementary signal is defined according to the Versatile Supplemental Enhancement Information, VSEI, standard specification.
12. A device for decoding encoded image or video data, the device comprising:a receiving unit configured to:receive a bitstream including an encoded input signal defining image or video data,wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;a parsing unit configured to parse the bitstream, the parsing comprising:parsing the bitstream, the parsing unit configured to:identify, in the bitstream, a supplementary signal comprising perspective information configured to allow the decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective;in dependence on the perspective information, deriving at least two image or video output signals from the encoded input signal.
13. The device of claim 12, wherein the device is further configured to output the two image or video output signals to two different displays.
14. An encoding method, implemented by an encoder, for encoding a plurality of visual signals, the method comprising:receiving an input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective;receiving perspective information that identifies each unique camera perspective;encoding the input signal into a bitstream, wherein the encoding comprises:encoding the input signal into a bitstream;in dependence on the perspective information, including a supplementary signal into the bitstream configured to indicate, to a decoder, i) information identifying each unique camera perspective and ii) that the bitstream is to be decoded into at least two image or video output signals.
15. The encoding method of claim 14, wherein the perspective information included in the supplementary signal comprises a parameter indicating a number of unique lenses that contribute to the plurality of visual signals, wherein each unique lens indicated by the perspective information is associated with a unique camera perspective.
16. The encoding method of claim 15, wherein the perspective included in the information supplementary signal identifies one or more lens attributes for each unique lens.
17. The encoding method of claim 16, wherein, for each unique lens, the one or more lens attributes comprise attributes selected from the group of: focal parameters; focal length; radial distortion parameters; horizontal and / or vertical location in the image or video of the focal point; transversal chromatic aberration parameters; vignetting parameters.
18. The encoding method of claim 16 or 17, comprising encoding the one or more lens as an array comprising a plurality of elements, wherein each element of the plurality of elements corresponds to a unique lens.
19. A device for encoding a plurality of visual signals, the device comprising:a receiving unit configured to receive: i) an input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective; and ii) perspective information that identifies each unique camera perspective;an encoding unit configured to:encode the input signal into a bitstream;in dependence on the perspective information, include a supplementary signal into the bitstream configured to indicate, to a decoder, i) information identifying each unique camera perspective and ii) that the bitstream is to be decoded into at least two image or video output signals.
20. A decoding apparatus comprising processing circuitry configured to execute steps of the method according to any claims 1 to 12.
21. An encoding apparatus comprising processing circuitry configured to execute steps of the method according to any claims 14 to 18.
22. A decoder comprising one or more processors and a non- transitory computer-readable storage medium coupled to the one or more processors, wherein the storage medium stores programming for execution by the one or more processors, wherein the programming, when executed by the one or more processors, configures the decoder to carry out the method according to any one of claims 1 to 12.
23. A non-transitory storage medium comprising a bitstream encoded by the method of any of claims 14 to 18.
24. A computer program stored on a non-transitory medium and including code instructions, which, when executed on one or more processor, causes the one or more processor to execute the method according to any of claims 1 to 12, or 14 to 18.
25. A system for delivering a bitstream, the system including at least one storage medium configured to store at least one bitstream generated by the encoding method described in any of claims 14 to 18.
26. A system for delivering a bitstream, the system comprising:at least one storage medium configured to store at least one bitstream generated by the method of any one of claims 14 to 18; anda video streaming device configured to obtain the bitstream from one of the at least one storage medium and send the bitstream to a terminal device, wherein the video streaming device comprises a content server or a content delivery server.
27. The system according to claim 26, further comprising one or more processors configured to perform encryption processing on at least one bitstream to obtain at least one encrypted bitstream, the at least one storage medium configured to store the encrypted bitstream; orthe one or more processors configured to convert the bitstream in a first format into a bitstream in a second format, the at least one storage medium configured to store the bitstream in the second format.
28. The system according to claim 26 or 27, further comprising:a receiver configured to receive a first operation request; wherein the one or more processor is configured to determine a target bitstream in the at least one storage medium in response to the first operation request; anda transmitter configured to send the target bitstream to a terminal-side apparatus.
29. The system according to claim 28, wherein the one or more processors is further configured to encapsulate the bitstream to obtain a transport stream in a first format, wherein the transmitter is further configured to:send the transport stream in the first format to a terminal-side apparatus for display; orsend the transport stream in the first format to storage space for storage.1630. A non-transitory computer-readable medium having processor-executable instructions stored thereon for reconstructing an image or video from a bitstream, wherein the bitstream includes an encoded input signal defining image or video data, the image or video data representing a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective;wherein the processor-executable instructions, when executed, cause performance of the following:identifying, in the bitstream, a supplementary signal comprising perspective information of the image or video data;interpreting, in dependence on the supplementary signal, that the perspective information identifies each unique camera perspective;in dependence on the identified unique camera perspectives, deriving at least two image or video output signals from the encoded input signal.
31. A bitstream, characterized by comprising :encoded input signal defining image or video data, wherein the image or video data represents a plurality of visual signals, each visual signal of the plurality of visual signals captured from a unique camera perspective, and wherein the image or video data is configured, once decoded, to be displayed via at least two different display outputs;a supplementary signal comprising perspective information configured to allow a decoder to interpret the bitstream, wherein the perspective information identifies each unique camera perspective.