Technique for reusing a part of an encoded original video when encoding a localized video
By encoding localized video using a method that calculates predicted frames from decoded original video and encodes residual frames, the technique addresses the issue of redundant encoding in existing methods, reducing memory usage and improving cache efficiency, thus enhancing the quality of experience for end users.
Patent Information
- Application Number
- JP2024563836
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-07
- Filing Date
- 2023-05-22
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Existing video encoding techniques redundantly encode original video data in localized video chunks, leading to excessive memory usage and decreased cache efficiency in content delivery networks (CDNs), which results in increased transmission delay and reduced quality of experience (QoE) for end users.
A computer-implemented method for encoding localized video that calculates a predicted frame based on a target frame of the localized video and a reference frame of a decoded original video, generates a residual frame, performs encoding operations on the residual frame to create an encoded localization layer, and transmits this layer along with encoded original video frames for decoding.
This approach reduces the amount of redundant original video data encoded, significantly reducing memory usage and improving cache efficiency in CDNs, thereby enhancing the quality of experience for end users during video streaming.
Smart Images

Figure 2025519328000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] This application claims the benefit of U.S. Patent Application No. 17 / 834,399, filed on June 7, 2022, the content of which is incorporated herein by reference.
Technical Field
[0002] Various embodiments of the present disclosure generally relate to computer science and streaming video technology, and more particularly to a technique for reusing a part of an encoded original video (hereinafter, "encoded original video") when encoding a localized video (hereinafter, "localized video").
Background Art
[0003] In a process known as "video localization", the original video is modified to generate a more suitable localized video for the target audience. For example, by modifying the lip movements of the speakers in a movie made in English to match the French dubbed audio of the movie using lip - animation technology, a more suitable localized version of the movie for French - speaking audiences can be generated. Usually, most of the video data included in the localized video is "original" video data that is the same as the corresponding original video, and the other video data becomes "localized" video data customized for the target audience. In a general video streaming service, when streaming a video to an end - user device, the localized video and the original video are not treated differently.
[0004] In some streaming implementation modes, a video encoder is used to encode "chunks" of video to generate encoded video chunks that are smaller in size than the original video chunks. Then, these multiple encoded video chunks are stored in the origin server device and streamed on demand to the end-user device via a content delivery network (CDN). In many implementation modes, when playing a predetermined video, the end-user device runs a playback application, and from this playback application, a series of requests are sequentially sent to the edge server device included in the CDN to request multiple encoded video chunks associated with the predetermined video. The edge server device is a server device that exists closer to the end-user device than the origin server device. If the edge server device stores a copy of the encoded video chunk requested by each such request in the cache memory associated with itself, a "cache hit" occurs for the request, and the edge server device sends the requested encoded video chunk to the playback application. On the other hand, if not, a "cache miss" occurs, and the edge server device acquires the requested encoded video chunk from the origin server device and then sends this to the playback application. And as the playback application receives various encoded video chunks, the video decoder decodes these encoded video chunks to generate corresponding decoded video chunks, and the generated decoded video chunks are played back via the end-user device.
[0005] When considering the process of streaming localized videos instead of original videos, one of the problems when encoding localized video chunks using a conventional encoder is that the original video data replicated in the localized video chunks is usually encoded redundantly. More specifically, the original video data held in a certain localized video is encoded once when generating the encoded video chunks of the original video, and is encoded again when generating the encoded video chunks of the localized video. Therefore, the total size of the conventional encoded localized videos (hereinafter, "encoded localized videos") becomes comparable to the size of the corresponding encoded original videos, and thus, the amount of localized video data present in the encoded localized videos is excessively large. As a result, the amount of memory used to store the encoded video chunks of a certain localized video has become unnecessarily large. Furthermore, since the size of the cache memory used in the edge server device is usually quite small compared to the total size of all the encoded video chunks associated with the streamable video library stored in the origin server device, the ratio of cache hits to cache misses in the CDN ("cache efficiency") has often decreased excessively with respect to the amount of localized video data present in the encoded localized videos. A decrease in cache efficiency may cause an increase in transmission delay, a decrease in effective transmission speed, and a decrease in transmission reliability during video streaming. Therefore, if the cache efficiency of the CDN decreases, it may degrade the perceived quality of experience (QoE) of the end user.
[0006] In order to reduce the amount of original video data that is redundantly encoded when encoding localized video, in some implementations, a simple process is performed where localized video chunks identical to the original video chunks are not encoded. Instead of encoding such localized video chunks, metadata is used to instruct that the corresponding encoded original video chunks be reused as encoded localized video chunks. On the other hand, for localized video chunks that are not identical to the original video chunks, they are encoded, and corresponding encoded localized video chunks are generated. However, this approach also has a problem in that many of the localized video chunks often have a very slight difference from the corresponding original video chunks. In such cases, when encoding the localized video chunks, a relatively large amount of original video data will still be redundantly encoded. For example, when generating localized video chunks using lip animation technology, almost all of the localized video chunks are modified with a very slight change to the speaker's lips in at least one frame from the corresponding original video chunks. The data other than the lip modification parts included in the localized video chunks is usually the same as the data included in the original video chunks. Therefore, when encoding the localized video chunks, most of the original video data is still redundantly encoded, and thus, there is also a need to transmit or store these redundantly and redundantly. Summary of the Invention Problems to be Solved by the Invention
[0007] As is clear from the above description, in the art, a more effective technique for encoding localized video is needed. Means for Solving the Problems
[0008] In one embodiment, a computer-implemented method for encoding a localized video is defined. The method includes calculating a predicted frame based on a target frame of the localized video and at least a part of a reference frame of a decoded original video, calculating a residual frame based on the predicted frame and the target frame of the localized video, performing one or more encoding operations on the residual frame to generate a frame of an encoded localization layer, and transmitting a frame of the encoded localization layer (hereinafter, the "encoded localization layer") and at least one frame of the encoded original video to another device for decoding.
[0009] One of the technical advantages of the technology of the present disclosure over the prior art is that the technology of the present disclosure can reduce, at least, the amount of original video data that is redundantly encoded when encoding a localized video. In this sense, for each frame of the encoded localization layer, any number of parts (including the entire frame) of the corresponding frame of the decoded original video can be specified for reuse. Thereby, the amount of memory used to store the encoded localization layer chunks used to construct the encoded localized video chunks can be significantly reduced compared to the amount of memory required to store the encoded localized video chunks using prior art techniques. Another technical advantage of the technology of the present disclosure is that by storing the encoded localization layer chunks instead of the encoded localized video chunks for streaming the localized video via a CDN, it is possible to improve the cache efficiency of the CDN and thus the QoE of the end user. These technical advantages result in one or more technical improvements over prior art approaches.
Brief Description of the Drawings
[0010] The concept of the present invention has been briefly summarized above. In order to enable a more detailed understanding of each feature of the various embodiments described above, the concept of the present invention will be described more specifically with reference to the various embodiments. Some of these embodiments are also illustrated in the accompanying drawings. However, it should be noted that the accompanying drawings merely exemplify representative embodiments of the concept of the present invention, and thus should not be considered as limiting the scope of the present disclosure in any sense, and there are other embodiments having similar effects.
Figure 1
Figure 2
Figure 3
Mode for Carrying Out the Invention
[0011] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the concept of the present invention can be practiced without one or more of these specific details.
[0012] In video localization processing, the original video is modified to generate one or more localized videos that are more suitable for the target audience. For example, by using lip animation technology to modify the lip movements of the speakers in a movie produced in English to match the dubbed voices in 30 languages of that movie, localized versions of the movie in 30 countries can be generated. Usually, most of the video data included in the localized video is the original video data that is the same as the corresponding original video, and the other video data becomes localized video data customized for the target audience. In a general video streaming service, when streaming a video to an end-user device, the localized video and the original video are not treated differently.
[0013] In some streaming implementation modes, a video encoder is used to encode multiple chunks of a video to generate multiple encoded video chunks corresponding to each of multiple pre-encoded versions of a certain video. Usually, multiple pre-encoded versions of a certain video each correspond to a different combination of average bitrate and resolution, that is, a different "bitrate-resolution pair", and are associated with different average picture quality levels. By storing multiple encoded video chunks corresponding to multiple pre-encoded versions of a certain video, regardless of the throughput obtained in the network connection involved during streaming, the possibility of streaming that video to an end-user device with good picture quality and without interruption in playback is increased.
[0014] In some embodiments, these multiple encoded video chunks are stored in the origin server device and are made available for on-demand streaming to the end-user device via the CDN. In many implementations, when playing a given video, the end-user device runs a playback application, and from this playback application, a series of requests are sequentially sent to the edge server device included in the CDN, requesting the multiple encoded video chunks associated with the given video. The edge server device is a server device that is closer to the end-user device than the origin server device. If the edge server device stores a copy of the encoded video chunk requested by each such request in the cache memory associated with itself, a cache hit occurs for that request, and the edge server device sends the requested encoded video chunk to the playback application. On the other hand, if not, a cache miss occurs, and the edge server device obtains the requested encoded video chunk from the origin server device (optionally, in addition to this, storing the requested encoded video chunk in the cache memory associated with the edge server device) and sends the requested encoded video chunk to the playback application. Then, as the playback application receives various encoded video chunks, the video decoder decodes these encoded video chunks to generate corresponding decoded video chunks, and the generated decoded video chunks are played via the end-user device.
[0015] One of the problems when encoding localized video chunks using a conventional encoder is that the original video data replicated in the localized video chunks is usually redundantly encoded multiple times. More specifically, when generating multiple encoded video chunks associated with a combination of a predetermined bitrate and resolution, the encoding of the original video data held in the localized video associated with a certain original video is performed once for the original video and, in addition, as many times as the number of localized videos associated with the original video. Therefore, the size of the pre-encoded version of each conventional localized video is approximately the same as the size of the pre-encoded version of the corresponding original video, and thus, the amount of localized video data present in the pre-encoded version of the localized video is excessively large. As a result, the amount of memory used to store all the encoded video chunks of each localized video has become unnecessarily large.
[0016] For example, when encoding an original video and 30 localized videos derived from the original video with 10 different "bitrate and resolution combinations", the original video data contained in the localized videos will be redundantly encoded 300 times. And the amount of memory used to store all the encoded video chunks associated with the 30 localized videos is approximately 30 times that of the amount of memory used to store all the encoded video chunks associated with the original video.
[0017] Furthermore, since the size of the cache memory used in the edge server device is usually much smaller than the total size of all the encoded video chunks associated with the video library stored in the origin server device, the cache efficiency of the CDN often excessively decreases with respect to the amount of localized video data present in the encoded localized video. When the cache efficiency of the CDN decreases, it may lead to an increase in transmission delay during video streaming, a decrease in the actual transmission speed, and a decrease in transmission reliability, which may reduce the QoE of the end user.
[0018] To reduce the amount of original video data that is redundantly encoded when encoding localized video, in some implementation modes, a simple process is performed where the same localized video chunk as the original video chunk is not redundantly encoded. Instead of encoding such a localized video chunk, metadata is used to instruct the corresponding encoded original video chunk to be reused as an encoded localized video chunk. On the other hand, for localized video chunks that are not the same as the original video chunk, they are encoded to generate corresponding encoded localized video chunks. However, this method also has a problem that many localized video chunks often have a very slight difference from the corresponding original video chunk. In such cases, when encoding the localized video chunk, a relatively large amount of original video data will still be redundantly encoded.
[0019] For example, when generating localized video chunks corresponding to a playback time of 2 to 10 seconds using lip animation technology, almost all of the localized video chunks are modified from the corresponding original video chunks by a very small amount in at least one frame with respect to the speaker's lips. The data other than the lip modification parts included in the localized video chunks is usually the same as the data included in the original video chunks. Therefore, when encoding the localized video chunks, most of the original video data is still encoded redundantly.
[0020] However, according to the technology of the present disclosure, when encoding a localized video derived from an original video, the localized video encoding application does not generate an encoded version of the localized video, but instead reuses a part of the frames of the encoded version of the original video to generate an encoded localization layer. In some embodiments, the localized video encoding application decodes the encoded version of the original video to generate a decoded original video (hereinafter, "decoded original video"). The localized video encoding application uses the decoded original video as a baseline for generating the encoded localization layer. Then, the localized video encoding application compares each frame of the localized video with the corresponding frame of the decoded original video.
[0021] When the localization video encoding application determines that there is no recognizable difference between a predetermined frame of the localized video and the corresponding frame of the decoded original video, the localization video encoding application instructs the encoding localization layer to reuse the frame of the encoded version of the original video as the corresponding frame of the encoded version of the localized video without modification. Therefore, in this case, the encoding localization layer indirectly instructs to reuse the corresponding frame of the decoded original video as the corresponding frame of the decoded localized video (hereinafter, "decoded localized video") without modification.
[0022] On the other hand, when the localization video encoding application determines that there is a recognizable difference between a predetermined frame of the localized video and the corresponding frame of the decoded original video, the localization video encoding application divides the frame of the localized video into a plurality of localized parts (hereinafter, "localized parts"), and divides the corresponding frame of the decoded original video into a plurality of original parts. Then, for each localized part, when the localization video encoding application determines that the corresponding original part is the "part with the highest degree of match (best matched part)" with the localized part, the localization video encoding application reuses this original part at the position in the prediction frame corresponding to the position of the localized part within the frame of the localized video. The localization video encoding application subtracts the prediction frame from the frame of the localized video to generate a residual frame representing the prediction error. Also, the localization video encoding application encodes the instructions for reconstructing the prediction frame and encodes the residual frame to generate an encoded localization layer frame corresponding to the frame of the localized video. Then, the localization video encoding application adds the encoded localization layer frame to the encoding localization layer.
[0023] In some embodiments, a plurality of encoded video chunks for the original video and a plurality of chunks of the encoded localization layer for the localized video (referred to as "encoded localization layer chunks") are stored in the origin server device. In the same or other embodiments, when streaming the localized video to the end-user device, the video delivery application running on the origin server device or the edge server device CDN delivers a sequence of encoded video chunks for the original video to the playback application running on the end-user device via one stream, and delivers a sequence of encoded localization layer chunks for the localized video via another stream. Then, the playback application generates a sequence of encoded video chunks for the localized video based on the sequence of video chunks for the original video and the sequence of encoded localization layer chunks for the localized video. As the playback application generates various encoded video chunks for the localized video, the video decoder decodes these encoded video chunks to generate corresponding decoded video chunks for the localized video, and the generated decoded video chunks for the localized video are played back via the end-user device.
[0024] As one of the technical advantages of the technology of the present disclosure over the prior art, at least, it can reduce the amount of original video data that the localization video encoding application duplicates and encodes when encoding the localized video. In this sense, in each frame of the encoding localization layer, any number of portions (including the entire frame) of the corresponding frame of the decoded original video can be specified to be reused, so that the size of the encoding localization layer can be significantly reduced compared to the size of the encoded version of the localized video. Therefore, in order to stream the localized video via the CDN, by storing the encoding localization layer chunks instead of the encoded localized video chunks, the total amount of memory used to store the encoded representation of the video library on the origin server device can be significantly reduced. Further, as another technical advantage of the technology of the present disclosure over the prior art, since the total amount of memory used to store the encoded representation of the video library on the origin server device is reduced, it is also possible to improve the cache efficiency of the CDN and thus improve the QoE of the end user. Due to these technical advantages, one or more technical improvements over the prior art methods can be obtained.
[0025] Overview of the system FIG. 1 is a conceptual diagram of a system 100 configured to implement one or more aspects of various embodiments. As shown, in some embodiments, system 100 includes, but is not limited to, compute instances 110(1), …, compute instances 110(4), a display device 102, and a cloud-based video service 104. For convenience of explanation, in this specification, compute instances 110(1), …, compute instances 110(4) may be referred to individually as “compute instance 110” or collectively as “compute instances 110”. In some embodiments, system 100 can include, but is not limited to, any number of compute instances 110, any number of display devices, any number and / or type of cloud-based services, or any combination thereof. In the same or other embodiments, display device 102, cloud-based video service 104, or both may be omitted from system 100.
[0026] Any number of components of these systems 100 can be distributed across multiple geographical locations or implemented in one or more cloud computing environments (i.e., encapsulated shared resources, software, data, etc.) in any combination. In some embodiments, any number of compute instances 110 can be implemented in a cloud computing environment, as part of another distributed computing environment, or stand-alone.
[0027] As shown, in some embodiments, compute instance 110(1) includes, without limitation, processor 112(1) and memory 116(1). Also, in the same or other embodiments, compute instance 110(2) includes, without limitation, processor 112(2) and memory 116(2). For convenience of explanation, in this specification, processor 112(1) and processor 112(2) may be referred to individually as "processor 112" or collectively as "processors 112". Also, in this specification, memory 116(1) and memory 116(2) may be referred to individually as "memory 116" or collectively as "memories 116". Although not shown, compute instance 110(3) and compute instance 110(4) may also include, without limitation, any number of processors and any number of memories, respectively.
[0028] Each processor 112 can be any instruction execution system, apparatus, or device capable of executing instructions. For example, each processor 112 can include a central processing unit, a graphics processing unit, a controller, a microcontroller, a state machine, or any combination thereof. The memory 116 of each compute instance 110 stores contents such as software applications and data used by the processor 112 of that compute instance 110. In some embodiments, each compute instance 110 can include any number of processors 112 and any number of memories 116 in any combination. Specifically, any number of compute instances 110 (including 1) can provide any number of multiprocessing environments in any technically feasible way.
[0029] Each memory 116 can be one or more local or remote digital storage devices and can take the form of any readily available memory (e.g., random access memory, read-only memory, floppy disk, hard disk). In some embodiments, a storage device (not shown) can complement or replace any number of memories 116. The storage device can comprise any number and / or type of external memory accessible by any number of processors 112. For example, but not limited to, examples of the storage device can include a secure digital card, an external flash memory, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0030] In the figure, as shown in italics, in some embodiments, the compute instance 110(3) is an origin server device included in a set of one or more origin server devices (not shown). In some embodiments, the entire set of origin server devices stores any number of pre-encoded copies of each video in the video library. This video is for streaming to an end-user device, either directly or via one or more CDNs, and at least one pre-encoded copy is stored for each pre-encoded version. In the same or other embodiments, each origin server device included in the set of origin server devices stores one or more pre-encoded chunks of each video in any part of the video library. In some embodiments, each origin server device may be included in any number (including zero) of CDNs.
[0031] Each video can include any amount and / or type of video data, although not limited thereto. Examples of videos include feature films, television programs, music videos, podcasts, and the like. In some embodiments, typically, multiple pre-encoded versions of a video each correspond to a different combination of average bitrate and resolution, i.e., a different "bitrate and resolution combination", and are associated with different average picture quality levels. As used herein, a set of different "bitrate and resolution combinations" for a video is referred to as an "encode rate ladder" for that video. By making multiple pre-encoded versions of a video available, regardless of the throughput obtained in the network connection through which it is streamed, the likelihood that the video can be streamed to an end-user device with good picture quality and without interruption in playback is increased. The "low-quality" encoded version is typically streamed to an end-user device when relatively low throughput is obtained in the network connection, and the "high-quality" encoded version is typically streamed to an end-user device when relatively high throughput is obtained in the network connection. As used herein, in some embodiments, each pre-encoded version of a video may be referred to as an "encoded video".
[0032] In some embodiments, a video, although not limited thereto, includes any number of discontinuous fragmentary portions of video data. In this specification, such discontinuous fragmentary portions may be referred to as "video chunks". Also, in the same or other embodiments, an encoded video, although not limited thereto, includes any number of discontinuous fragmentary portions of encoded video data. In this specification, such discontinuous fragmentary portions may be referred to as "encoded video chunks". In some embodiments, an encoded version of a video, although not limited thereto, includes different encoded video chunks for each video chunk of the video.
[0033] In the figure, as shown in italics, in some embodiments, compute instance 110(4) is one of a plurality of CDN edge server devices within the CDN. The CDN edge server device can obtain a plurality of encoded video chunks from one or more origin server devices and receive and respond to a request for an encoded video chunk from an end-user device. In the same or other embodiments, each CDN edge server device receives and responds to a request from an end-user device that is closer to the CDN edge server device than to the origin server device. As also described in the above explanation, in some embodiments, each CDN edge server device can temporarily store copies of a limited number of encoded video chunks in a cache memory associated therewith. In some embodiments, the CDN, although not limited thereto, also includes any number of other cache memories that are distributed and provided at intermediate positions throughout the CDN.
[0034] In some embodiments, when the CDN edge server device receives a request to request an encoded video chunk from an end-user device, in response, it finds a copy of the encoded video chunk at the closest location and sends this to the end-user device. In some embodiments, if a copy of the requested encoded video chunk is stored in the cache memory associated with the CDN edge server device itself, a "cache hit" occurs for the request, and the CDN edge server device sends a copy of the requested encoded video chunk. On the other hand, in the same or other embodiments, if the requested encoded video chunk is not stored in the cache memory associated with the CDN edge server device, a "cache miss" occurs. When a cache miss occurs, the CDN server device needs to obtain the requested encoded video chunk from an intermediate cache memory or an origin server device before sending the requested encoded video chunk to the end-user device.
[0035] In the figure, as shown in italics, in some embodiments, compute instance 110(2) is an end-user device. In some embodiments, compute instance 110(4) can stream video to compute instance 110(2) via a network connection (not shown). In the same or other embodiments, compute instance 110(2) can display video on display device 102. In some embodiments, display device 102 can be any type of device configured to display any amount and / or type of video data in any technically feasible way. In the same or other embodiments, compute instance 110(2), zero or more other compute instances 110, display device 102, and zero or more other display devices are incorporated into a user device (not shown). Examples of user devices include, but are not limited to, desktop computers, laptop computers, smartphones, smart TVs, gaming consoles, tablet terminals, and the like.
[0036] In some embodiments, cloud-based video service 104 includes microservices, databases, and storage devices related to streaming video services activities and content that are not assigned to any of the origin server devices, CDNs, or end-user devices, such as microservices, databases, or storage devices. Examples of functions that can be provided by cloud-based video service 104 include, but are not limited to, login and billing, logging, video title recommendations tailored to individual users, video transcoding, server and connection health monitoring, client-specific CDN guidance, and the like. In some embodiments, cloud-based video service 104 monitors the health of compute instance 110(4), compute instance 110(2), and the network connections associated therewith.
[0037] Each compute instance 110 is configured to implement one or more software applications. For the sake of convenience of explanation, each software application is illustrated as residing in the memory 116 of one compute instance 110 and being executed on the processor 112 of the one compute instance 110. However, as would be recognized by those skilled in the art, the functions of each software application can be distributed among any number of other software applications, and these other software applications can be resident in the memories 116 of any number of compute instances 110 and executed on the processors 112 of any number of compute instances 110 in any combination. Also, the functions of any number of software applications can be integrated into one application or subsystem.
[0038] In particular, in some embodiments, compute instance 110(1) is configured to encode an original video and a localized video derived from the original video by video localization processing. In the video localization processing, one or more copies of the original video are modified to generate one or more localized videos. Each localized video is made more suitable for a different target audience. Usually, most of the video data included in the localized video is the original video data that is the same as the corresponding original video, and the other video data becomes localized video data customized for the target audience of the localized video. As used herein, the "target audience" of a given localized video includes, but is not limited to, end users who are expected to select the localized video based on any number and / or type of criteria. For example, the target audience of a localized video modified to match dubbing in a given language includes, but is not limited to, end users who are expected to select the given language from among multiple available languages. As described in detail above, in a typical video streaming service, the localized video and the original video are not treated differently when streaming the video to an end user device. This is especially true when encoding chunks of the localized video using a conventional encoder to generate encoded video chunks of the localized video.
[0039] One of the problems when encoding localized video chunks using a conventional encoder is that the original video data replicated in a chunk of a certain localized video is usually encoded once when generating the encoded video chunk of the original video, and is encoded again when generating the encoded video chunk of the localized video. Therefore, the total size of the conventionally encoded localized video becomes comparable to the size of the encoded original video corresponding thereto, and thus, it has been excessively large with respect to the amount of localized video data existing in the encoded localized video. As a result, the amount of memory used to store the encoded video chunk of a certain localized video has become unnecessarily large. Furthermore, since the size of the cache memory used in the edge server device is usually considerably smaller than the total size of all the encoded video chunks associated with the streamable video library stored in the origin server device, the cache efficiency of the CDN, and thus the QoE of the end user, has often been excessively degraded with respect to the amount of localized video data existing in the encoded localized video.
[0040] To reduce the amount of original video data that is redundantly encoded when encoding a localized video, in some implementation modes, a simple process of not encoding a localized video chunk identical to the original video chunk is performed. And instead of encoding such a localized video chunk, metadata is used to instruct to reuse the corresponding encoded original video chunk as the encoded localized video chunk. However, this method also has a problem that many localized video chunks often have a very slight difference from the corresponding original video chunk. In such a case, when encoding the localized video chunk, most of the original video data will still be redundantly encoded.
[0041] Encoding the difference between the localized video and the original video To address the above problems, in some embodiments, compute instance 110(1) includes, but is not limited to, a localized video encoding application 120. As described in more detail in the following description, in some embodiments, the localized video encoding application 120 does not encode the video data of the localized video, but rather encodes the difference between the video data of the localized video and the video data of the corresponding decoded original video to generate an encoded localization layer. Then, the generated encoded localization layer can be combined with the encoded original video and decoded to generate a decoded localized video. Also, in some embodiments, each encoded chunk included in the encoded localization layer can be combined with the corresponding encoded chunk included in the encoded original video and decoded to generate the corresponding video chunk included in the decoded localized video.
[0042] For the sake of convenience of explanation, in this specification, the video chunks included in the original video may be referred to as "original video chunks", and the video chunks included in the localized video may be referred to as "localized video chunks". Also, in this specification, the encoded video chunks included in the encoded original video may be referred to as "encoded original video chunks", the encoded video chunks included in the encoded localization layer may be referred to as "encoded localization layer chunks", and the encoded video chunks included in the encoded localized video may be referred to as "encoded localized video chunks". Further, in this specification, the video chunks included in the decoded original video may be referred to as "decoded original video chunks", and the video chunks included in the decoded localized video may be referred to as "decoded localized video chunks".
[0043] As described in more detail in the following explanation, in some embodiments, compute instance 110(2) includes, but is not limited to, playback application 170. In some embodiments, playback application 170 constructs an encoded localized video chunk based on the encoded localization layer chunk and the encoded original video chunk. Then, playback application 170 decodes the encoded localized video chunk to generate a decoded localized video chunk. In some embodiments, playback application 170 plays the corresponding localized video by sequentially displaying the decoded localized video chunks on display device 102.
[0044] As an advantage, when encoding localized video, the amount of original video data that the localized video encoding application 120 would redundantly encode can be significantly reduced compared to the amount of original video data that a conventional encoder redundantly encodes when encoding localized video. Therefore, the size of the encoding localization layer for the localized video can be made significantly smaller compared to the encoded version of the localized video. Further, when deriving a plurality of localized videos from the same single original video, the corresponding plurality of encoded localization layers can be decoded in combination with the single encoded original video. Therefore, each time a localized video is added, the amount of memory that the origin server device additionally uses to store the encoded video data thereof can be reduced from the size of the encoded version of the localized video to the size of the encoded localization layer corresponding to the localized video. As a result, the cache efficiency of the associated CDN can be improved, and thus the QoE of the end user can be improved.
[0045] As shown, in some embodiments, the localized video encoding application 120 resides in the memory 116(1) of the compute instance 110(1) and is executed on the processor 112(1) of the compute instance 110(1). Also, in the same or other embodiments, the playback application 170 resides in the memory 116(2) of the compute instance 110(2) and is executed on the processor 112(2) of the compute instance 110(2). In some other embodiments, any number of portions (including the entire function) of the functions described in the description of the localized video encoding application 120 and the playback application 170 herein can be distributed among any number of compute instances in any technically feasible manner. Also, in some embodiments, the playback application 170 can be omitted from the system 100.
[0046] For the sake of convenience of explanation, taking as an example the situation where one encoded version of the original video 122 is generated and different encoded localization layers are generated for each of the localized videos 124(1), …, localized videos 124(K), the functions of the localized video encoding application 120 of several embodiments will be described. Note that K can be any positive integer. In the same or other embodiments, the localized video encoding application 120 encodes the original video 122 and the localized videos 124(1), …, localized videos 124(K) using the same set of encoding parameter values.
[0047] However, as those skilled in the art will recognize, the technologies described in the description of the localized video encoding application 120 and the playback application 170 in this specification are exemplary and not intended to be limiting, and can be changed as long as they do not depart from the spirit and scope of the present invention broadly construed. For the functions of the localized video encoding application 120 and the playback application 170 described in this specification, many changes and variations that do not depart from the scope and spirit of the embodiments described in this specification will be apparent to those skilled in the art. For example, if J is an integer greater than 1, in some embodiments, the localized video encoding application 120 can also generate, for each of the localized videos 124(1), …, localized videos 124(K), J encoded versions of the original video 122 and J encoded localization layers based on J sets of encoding parameter values.
[0048] As shown, in some embodiments, the localize video encoding application 120 generates an encoded original video 132 and encoded localization layers 150(1),..., encoded localization layers 150(K) based on the original video 122 and the localized videos 124(1),..., localized videos 124(K). For convenience of explanation, in this specification, the localized videos 124(1),..., localized videos 124(K) may be individually referred to as "localized video 124" or collectively as "localized videos 124" and "localized videos 124(1)-124(K)". Also, in this specification, the encoded localization layers 150(1),..., encoded localization layers 150(K) may be individually referred to as "encoded localization layer 150" or collectively as "encoded localization layers 150" and "encoded localization layers 150(1)-150(K)".
[0049] As shown, in some embodiments, the localization video encoding application 120 includes, but is not limited to, a video encoder 130, an encoded original video 132, localization encoders 140(1),..., localization encoders 140(K), and an encoded localization layer 150. The localization encoders 140(1),..., localization encoders 140(K) are different instances of one localization encoder (this encoder may be referred to herein as the "localization encoder 140"). For convenience of explanation, herein, the localization encoders 140(1),..., localization encoders 140(K) may be individually referred to as the "localization encoder 140", or collectively referred to as the "localization encoders 140" and the "localization encoders 140(1) to 140(K)".
[0050] As shown, in some embodiments, video encoder 130 encodes original video 122 to produce encoded original video 132. Thus, encoded original video 132 is an encoded version of original video 122. For integers k from 1 to K, localization encoder 140(k) generates encoded localization layer 150(k) based on localized video 124(k) and encoded original video 132. As described in more detail in the following discussion of FIG. 2, in some embodiments, localization encoder 140(k) can optionally obtain and use any amount and / or type of additional data to identify localized video data. For example, in some embodiments, localized video encoding application 120 inputs original video 122, metadata generated by the localization application in generating localized video 124(k), or both, to localization encoder 140(k).
[0051] The localization encoder 140 can generate an encoded localization layer 150 by implementing any number and / or type of encoding techniques and any number and / or type of decoding techniques. In some embodiments, the localization encoder 140 generates the encoded localization layer 150 by implementing scalable encoding techniques. In scalable encoding, a modified version of the original video frame can be encoded in the form of the difference between the original video frame and the modified video frame. Examples of codecs that implement scalable encoding techniques for different encoded versions of a single video include, but are not limited to, the AV1 (AOMedia Video 1) codec, some H.245 / HEVC (High Efficiency Video Coding) codecs, and some H.264 / AVC (Advanced Video Coding) codecs. In contrast, the localization encoder 140 implements scalable encoding techniques for encoded versions of multiple videos that are different but related to each other.
[0052] For the sake of convenience of explanation, hereinafter, the functions of the localization encoder 140 in some embodiments will be described by taking the localization encoder 140(1) as an example. As will be described in more detail in the following description of FIG. 2, in some embodiments, the localization encoder 140(1) decodes the encoded original video 132 to generate a decoded original video (not shown in FIG. 1). The localization encoder 140(1) uses the frames of the decoded original video as basic data for generating the frames of the encoded localization layer 150(1) according to the frames of the localized video 124(1). In the same or other embodiments, in each frame of the encoded localization layer 150(1), although not limited thereto, encoded prediction metadata (hereinafter, "encoded prediction metadata") is specified, and optionally, in addition to this, encoded residual data (hereinafter, "encoded residual data") is specified. In some embodiments, the encoded prediction metadata instructs the video decoder to copy one or more portions of the corresponding frame of the decoded original video to the reconstructed prediction frame (optionally, in addition to this, move the one or more portions with respect to the reconstructed prediction frame). In the same or other embodiments, the encoded residual data specifies the corrections that the video decoder needs to perform on the reconstructed prediction frame to generate the frames of the decoded localized video. In some embodiments, the decoded localized video is not an exact replica of the localized video 124(1).
[0053] Although illustrations are omitted, in some embodiments, any software application can generate any portion of the encoded localized video based on the corresponding portion of the encoded original video corresponding to any portion of the encoded localized video and the corresponding portion of the encoded localization layer corresponding to any portion of the encoded localized video. For the sake of convenience in explanation, in this specification, the process of generating a chunk of the encoded version of the localized video 124(1) (the "encoded localized video chunk") will be taken as an example to explain the process of generating a part of the encoded localized video. However, the same technique can be applied to generate any portion of the encoded localized video based on the corresponding portion of the encoded original video corresponding to any portion of the encoded localized video and the corresponding portion of the encoded localization layer corresponding to any portion of the encoded localized video.
[0054] Although illustrations are omitted, in some embodiments, any software application or any hardware module such as the localization encoder 140, the localized video encoding application 120, the playback application 170, etc. can execute a "layer interleaving algorithm" to generate any portion (including the entire video) of the encoded localized video. In the same or other embodiments, the layer interleaving algorithm is not limited, but is a series of instructions that specify generating any portion of the encoded localized video based on the corresponding portion of the encoded original video corresponding to any portion of the encoded localized video, the corresponding portion of the encoded localization layer corresponding to any portion of the encoded localized video, and optionally, any amount and / or type of metadata.
[0055] In some embodiments, a software application or a hardware module executes a layer interleaving algorithm to generate an encoded localized video chunk based on an encoded original video chunk associated with an original video chunk and an encoded localization layer chunk associated with a localized video chunk. In the same or other embodiments, using the techniques of the present disclosure, based on a portion of an encoded original video corresponding to any number (including 1) of portions of an encoded localized video of any localized video, and a portion of an encoded localization layer corresponding to any number of portions of the encoded localized video, any number of portions of the encoded localized video can be generated.
[0056] In some embodiments, to generate an encoded localized video chunk, a software application or a hardware module identifies a frame of the localized video chunk that has a difference with respect to a corresponding frame of the original video chunk. The software application or the hardware module can determine whether a frame of the localized video chunk has a difference with respect to the corresponding frame of the original video chunk in any technically feasible manner. For example, in some embodiments, based on a frame of the encoded localization layer chunk corresponding to a frame of the localized video chunk, any amount and / or type of metadata, or any combination thereof, the software application or the hardware module can determine whether the frame of the localized video chunk has a difference with respect to the corresponding frame of the original video chunk.
[0057] In some embodiments, a software application or a hardware module generates an encoded base layer (hereinafter, the "encoded base layer") by marking, for each frame identified in a localized video chunk, the corresponding frame of the encoded original video chunk as a reference-only frame. Accordingly, each frame included in the encoded base layer is a selectively marked copy of a frame included in the encoded original video chunk. As used herein, a "reference-only frame" refers to a frame that can be used for reference by other frames but is neither displayed nor presented. Then, the software application or the hardware module interleaves the frames of the encoded base layer with the frames of the encoded localization layer chunk to generate an encoded localized video chunk.
[0058] As shown, in some embodiments, a localized video encoding application 120 separately transmits an encoded original video 132 and each encoded localization layer 150 to any number of origin server devices (e.g., compute instances 110(3)). In the same or other embodiments, any number of CDN edge server devices (e.g., compute instances 110(4)) can separately stream the encoded original video 132 and each encoded localization layer 150 to any number of end-user devices (e.g., compute instances 110(2)).
[0059] In some embodiments, when streaming one of the localized videos 124 to an end-user device (e.g., compute instance 110(2)), the CDN edge server device delivers two separate streams, an encoded original video stream and an encoded localization layer stream, to the end-user device. In some embodiments, the CDN edge server device delivers chunks of the encoded original video 132 via the encoded original video stream and delivers chunks of the encoded localization layer 150 corresponding to the localized video 124 via the encoded localization layer stream.
[0060] Although not shown, in some other embodiments, chunks of the encoded original video 132 and chunks of the encoded localization layer 150 corresponding to the localized video 124 can be delivered to the end-user device in one stream or two separate streams in any technically feasible manner by one or two edge server devices, one or two origin server devices, one or two other devices, or any combination thereof. In some embodiments, it is also possible to configure the encoded original video 132 to be downloaded to the end-user device at some point before the encoded localization layer 150 is streamed to the end-user device, and the techniques described in this specification can be modified accordingly.
[0061] For the sake of convenience in explanation, FIG. 1 shows an example of a high-level event executed by the playback application 170 to play the localized video 124(1). In some embodiments, the playback application 170 enables an end user of the playback application 170 to select a video to be streamed to the compute instance 110(2) by opening one or more network connections to the cloud-based video service 104. In the same or other embodiments, the playback application 170 can transmit and receive any amount and / or any type of data (including zero) to and from the cloud-based video service 104 in any technically feasible way.
[0062] In some embodiments, the end user or the playback application 170 selects the localized video 124(1) for streaming in any technically feasible way. For example, in some embodiments, the playback application 170 can select the localized video 124(1) for streaming based on any number and / or type of end-user initial settings, any number and / or type of end-user preferences, the language selected by the end user during the playback of the previously selected localized video, or any combination thereof. In some embodiments, when the end user or the playback application 170 selects the localized video 124(1) for streaming, the playback application 170 sends a manifest request specifying the localized video 124(1) to the cloud-based video service 104. In response, the cloud-based video service 104 creates a manifest file 172 based on the localized video 124(1), optionally the CDN, and optionally the compute instance 110(2). Then, the cloud-based video service 104 sends the manifest file 172 to the playback application 170.
[0063] In some embodiments, the manifest file 172 specifies, without limitation, a combination of bitrate and resolution, an average picture quality level, and location data associated with the encoded version of the localized video 124(1). In the same or other embodiments, the location data specifies, without limitation, the location of the encoded original video chunks for the original video 122 and the location of the encoded localization layer chunks for the localized video 124(1) at each of one or more CDN edge server devices proximate to the compute instance 110(2). As also described in the above description, the encoded original video chunks for the original video 122 are chunks of the encoded original video 132. Also, the encoded localization layer chunks for the localized video 124(1) are chunks of the encoded localization layer 150(1). One or more CDN edge server devices proximate to the compute instance 110(2) include, without limitation, the compute instance 110(4).
[0064] In some embodiments, the playback application 170 selects, based on the manifest file 172, a sequence of encoded original video chunks to be transmitted from the compute instance 110(4) to the compute instance 110(2) and a corresponding sequence of encoded localization layer chunks. The order of the selected sequence of encoded original video chunks corresponds to the display order of the video chunks (and frames within the video chunks) of the original video 122, and the order of the selected sequence of encoded localization layer chunks corresponds to the display order of the video chunks (and frames within the video chunks) of the localized video 124(1).
[0065] For the sake of convenience in explanation, although FIG. 1 shows an exemplary period, in some embodiments, the starting point of the period shown in FIG. 1 can be set to the time when the playback application 170 selects the position of the encoded original video chunk 186 according to the sequence of the selected encoded original video chunks and selects the position of the encoded localization layer chunk 188 according to the sequence of the selected encoded localization layer chunks. The encoded original video chunk 186 and the encoded localization layer chunk 188 are chunks corresponding to the same localized video chunk.
[0066] As shown, the playback application 170 issues an encoded video chunk request 182 and an encoded layer chunk request 184 to a video delivery application (not shown) running on the compute instance 110(4). In some embodiments, the encoded video chunk request 182 requests a plurality of bytes corresponding to the encoded original video chunk 186. Also, in the same or other embodiments, the encoded layer chunk request 184 requests a plurality of bytes corresponding to the encoded localization layer chunk 188.
[0067] In response to the encoded video chunk request 182 and the encoded layer chunk request 184, in some embodiments, the video delivery application acquires the encoded original video chunk 186 and the encoded localization layer chunk 188 and transmits them in separate streams. Also, in some other embodiments, the video delivery application provides the encoded original video chunk 186 and the encoded localization layer chunk 188 in one stream (not shown) in any technically feasible way. As the playback application 170 incrementally receives the encoded original video chunk 186 and the encoded localization layer chunk 188 little by little, it stores the encoded original video chunk 186 and the encoded localization layer chunk 188 in the playback buffer 174 incrementally little by little.
[0068] And in some embodiments, as the playback application 170 receives and / or stores in the playback buffer 174 multiple frames of the encoded original video chunk 186 and multiple frames of the encoded localization layer chunk 188, the layer interleaving algorithm described above is incrementally and gradually executed to incrementally and gradually construct the encoded localized video chunk 176. Also, in the same or other embodiments, the playback application 170 incrementally uses the multi-layer video decoder 178 to incrementally decode the encoded localized video chunk 176, thereby incrementally and gradually generating the decoded localized video chunk 190.
[0069] The multi-layer video decoder 178 can include, but is not limited to, zero or more software applications, zero or more hardware modules, or any combination thereof, which can cooperate to decode the encoded localized video chunk 176 and generate the decoded localized video chunk 190. In some embodiments, the multi-layer video decoder 178 decodes the associated multiple layers of encoded video data represented in any technically feasible manner by implementing scalable decoding techniques, any number and / or type of other decoding techniques, or both. For example, in some embodiments, the multi-layer video decoder 178 supports the Scalable Video Coding (SVC) extension of H.264 / AVC.
[0070] In some embodiments, the playback application 170 incrementally renders the decoded localized video chunk 190 and zero or one or more other decoded localized video chunks (not shown) little by little in the display order associated with the localized video 124(1) to the display device 102, thereby playing back the localized video 124(1) chunk by chunk. In the same or other embodiments, the playback application 170 can incrementally render and play back the frames of the intermediate version of the decoded localized video chunk 190 little by little before receiving all the frames of the encoded original video chunk 186 and / or all the frames of the encoded localization layer chunk 188.
[0071] Note that the technology described in this specification is exemplary and not intended to be limiting, and it should be noted that it can be changed without departing from the spirit and scope of the present invention broadly construed. It will be apparent to those skilled in the art that many changes and variations can be made to the functions of the localized video encoding application 120, the localization encoder 140, and the playback application 170 described in this specification without departing from the scope and spirit of the embodiments described herein. Similarly, many changes and variations can be made to the storage and distribution of the encoded original video, the encoded localization layer, and the encoded localized video described in this specification without departing from the scope and spirit of the embodiments described herein.
[0072] For example, in some embodiments, the original video 122 includes any amount of base video data not intended to be presented directly to the end user, although not limited thereto. Instead, the localized video 124 can be generated using the original video 122, and the encoded localization layer 150 can be generated using the decoded encoded original video 132. Note that the encoded original video 132 and the encoded localization layer 150 are not decoded and presented independently, but at least a portion of the localized encoded video (e.g., the encoded localized video chunk 176) is generated using these and decoded and presented.
[0073] In particular, in some embodiments, the original video 122 is set to be a version with the lip portions of the original language video blurred, and one of the plurality of localized videos 124 is set to be the original language video. In the same or other embodiments, each of the encoded localization layers 150 (including the encoded localization layer 150 corresponding to the original language video) performs coding to add the lip appearance data of the corresponding language to the original video 122, but does not need to perform coding to remove the lip appearance data of the original language from the original video 122. Therefore, the number of different pixels can be reduced compared to embodiments where the original video is the original language video.
[0074] In some embodiments, the original video 122 is made into a version (but not limited to) with the text bubbles depicted in the original language video blanked out, and one of the plurality of localized videos 124 is set to be the original language video. In the same or other embodiments, each of the encoded localization layers 150 (including the encoded localization layer 150 corresponding to the original language video) performs coding to add the text of the corresponding language to the text bubbles, but does not need to perform coding to remove the original language from the text bubbles of the original video 122. Therefore, compared with the embodiments where the original video is the original language video, the number of different pixels can be reduced.
[0075] It should be understood that the system 100 shown in this specification is exemplary and can be modified and changed. For example, the functions provided by the localization video encoding application 120, the localization encoder 140, the playback application 170, and the cloud-based video service 104 described in this specification can be integrated into or distributed among any number of software applications (including 1) and any number of components of the system 100. Furthermore, the connection topology between the various units shown in FIG. 1 can be changed as needed.
[0076] FIG. 2 is a diagram showing one of the plurality of localization encoders 140 shown in FIG. 1 in more detail according to various embodiments. More precisely, FIG. 2 shows the localization encoder 140(1), and in some embodiments, the localization encoder 140(1) generates the encoded localization layer 150(1) based on the localized video 124(1) and the encoded original video 132. As described in the above description of FIG. 1, the localized video 124(1) is a modified version of the original video 122. The localized video 124(1) can be derived from the original video 122 in any technically feasible way.
[0077] Although illustration is omitted, in some embodiments, the original video 122 includes a frame sequence consisting of N frames, where N can be any positive integer. For the sake of convenience of explanation, in this specification, the frames of the original video 122 may be individually referred to as "original video frames" or collectively as "original video frames". Each original video frame is an image including any amount and / or type of video data (but not limited thereto). In this specification, the video data included in the original video frame may be referred to as "original video data". For the sake of convenience of explanation, a frame index is associated with each original video frame. The frame index is a number from 1 to N assigned according to the display order which is the order in which the frames are sequentially displayed during playback. More precisely, during normal playback of the original video 122, the original video frames corresponding to the frame indices "1" to "N" are sequentially displayed.
[0078] As shown, in some embodiments, the localized video 124(1) includes, but is not limited to, a frame sequence consisting of localized video frames 224(1), …, localized video frames 224(N). Here, N is the total number of original video frames. In the same or other embodiments, the localized video frames 224(1), …, localized video frames 224(N) are respectively modified versions of the original video frames with frame indices of "1", …, the original video frames with frame indices of "N", and thus correspond to them respectively. For convenience of explanation, the localized video frames 224(1), …, localized video frames 224(N) are respectively associated with frame indices from 1 to N. In some embodiments, the number of frames of the localized video 124(1) can be configured to be different from the number of frames of the original video, and the techniques described in this specification can be changed according to this configuration.
[0079] For convenience of explanation, in this specification, the localized video frames 224(1), …, localized video frames 224(N) may be individually referred to as "localized video frame 224", or collectively referred to as "localized video frames 224" and "localized video frames 224(1) to 224(N)". In some embodiments, each localized video frame 224 is a frame of the localized video 124(1) and includes, but is not limited to, any amount of original video data that matches the original video data in the corresponding original video frame, any amount of localized video data obtained by modifying the original video data in the corresponding original video frame, or both.
[0080] The encoded original video 132 is an encoded version of the original video 122. The encoded original video 132 can be generated by any technically feasible method. As described in the above description regarding FIG. 1, in some embodiments, the video encoder 130 encodes the original video 122 to generate the encoded original video 132. As shown, in some embodiments, the encoded original video 132 includes, but is not limited to, encoded original video frames 232(1),..., encoded original video frames 232(N).
[0081] The encoded original video frames 232(1),..., encoded original video frames 232(N) are respectively encoded versions of the original video frames with frame indices "1",..., "N", and thus correspond to them respectively. For convenience of explanation, the encoded original video frames 232(1),..., encoded original video frames 232(N) are respectively associated with frame indices from 1 to N. In this specification, the encoded original video frames 232(1),..., encoded original video frames 232(N) may be individually referred to as "encoded original video frame 232", or collectively as "encoded original video frames 232" and "encoded original video frames 232(1) to 232(N)".
[0082] As shown, in some embodiments, the localization encoder 140(1) includes, but is not limited to, a video decoder 202, a decoded original video 210, a frame comparison engine 240, a modified frame index list 250, comparison metadata 252, an incremental encoding engine 260, localization frame encoders 270(1), …, localization frame encoders 270(M), and an encoded localization layer 150(1). In some embodiments, M can be any integer greater than or equal to 1 and less than or equal to N.
[0083] In some embodiments, the localization frame encoders 270(1), …, localization frame encoders 270(M) are different instances of a localization frame encoder (which may be referred to herein as the "localization frame encoder 270"). For convenience of explanation, herein, the localization frame encoders 270(1), …, localization frame encoders 270(M) may be individually referred to as "localization frame encoder 270", or collectively referred to as "localization frame encoders 270" and "localization frame encoders 270(1) - 270(M)".
[0084] As shown, video decoder 202 decodes the encoded original video 132 to generate a decoded original video 210. The video decoder 202 can be any video decoder or any part of a codec that has the function of decoding the encoded original video 132. In some embodiments, the video decoder 202 decodes single-layer encoded video data by implementing any number and / or type of decoding techniques. In the same or other embodiments, the video decoder 202 decodes multi-layer encoded video data represented in any technically feasible way by implementing any number and / or type of decoding techniques. In some embodiments, the video decoder 202 is an instance of any multi-layer video encoder such as the multi-layer video decoder 178 shown in FIG. 1. In the same or other embodiments, the video decoder 202 supports the SVC extension of H.264 / AVC.
[0085] As shown, the decoded original video 210 includes, but is not limited to, decoded original video frames 212(1), …, decoded original video frames 212(N). The decoded original video frames 212(1), …, decoded original video frames 212(N) are respectively the decoded versions of the encoded original video frames 232(1) to 232(N), and thus respectively correspond to the original video frames with frame indices "1", …, "N". For the sake of convenience of explanation, the decoded original video frames 212(1), …, decoded original video frames 212(N) are respectively associated with frame indices 1 to N. In this specification, the decoded original video frames 212(1), …, decoded original video frames 212(N) may be individually referred to as "decoded original video frame 212", or collectively referred to as "decoded original video frames 212" and "decoded original video frames 212(1) to 212(N)".
[0086] As will be recognized by those skilled in the art, since video encoders typically implement lossy compression techniques, the decoded video data may contain any amount and / or type of noise resulting from the compression. Therefore, when the decoded video data is displayed, such noise may cause visible distortions known as "encoding artifacts". Thus, in some embodiments, the decoded original video frame 212 is not an exact replica of the corresponding original video frame, and the decoded original video 210 is not an exact replica of the original video 122.
[0087] As shown, in some embodiments, the frame comparison engine 240 generates a modified frame index list 250 and comparison metadata 252 based on the localized video 124(1) and the decoded original video 210. In some embodiments, the modified frame index list 250 specifies, but is not limited to, the frame indices of a subset of the localized video frames 224 that have differences with respect to the corresponding frames of the decoded original video frames 212(1) - 212(N) among the localized video frames 224(1) - 224(N). In the same or other embodiments, the modified frame index list 250 specifies, but is not limited to, the frame indices of a subset of the localized video frames 224 that contain visible modification locations resulting from the video localization process among the localized video frames 224(1) - 224(N). For convenience of explanation, in some embodiments, the modified frame index list 250 contains, but is not limited to, M frame indices. In this specification, such frame indices are represented as idx1 - idxM in the display order. Here, 1 ≦ idx1 ≦ idxM ≦ M ≦ N.
[0088] In some embodiments, the comparison metadata 252 identifies any amount and / or type of original video data included in any portion (including the entire frame) of the localized video frames 224(1) to 224(N), although not limited thereto. In the same or other embodiments, the comparison metadata 252 identifies any number and / or type of differences between any portion (including the entire frame) of the localized video frames 224(1) to 224(N) and the corresponding portions of the decoded original video frames 212(1) to 212(N). For example, in some embodiments, although not limited thereto, the comparison metadata 252 specifies a frame index corresponding to a subset of the localized video frames 224 that do not contain localized video data, and in addition to or instead of this, for each of any number of localized video frames 224 included in a subset of the localized video frames 224 that contain localized video data, specifies the positions of one or more discontinuous localized video data fragment portions.
[0089] The frame comparison engine 240 can identify any number and / or type of differences that the localized video frames 224 have, and / or can detect localized video data in any technically feasible way. The frame comparison engine 240 can generate a modified frame index list 250, comparison metadata 252, any amount and / or type of other comparison data, or any combination thereof, in any technically feasible way, based on any number and / or type of differences associated with the localized video frames 224, the localized video data, or both.
[0090] In some embodiments, for frame indices x from "1" to "N", the frame comparison engine 240 compares the localized video frame 224(x) with the decoded original video frame 212(x) to determine whether to add the frame index x to the modified frame index list 250. Also, in the same or other embodiments, the frame comparison engine 240 calculates the pixel-by-pixel difference value between the localized video frame 224(x) and the decoded original video frame 212(x) to identify the "modification" frame associated with the localized video frame 224(x). In some embodiments, the frame comparison engine 240 can implement any minimum difference criterion (e.g., minimum pixel distance value, minimum size of the modified area, etc.) and can exclude or ignore any differences that do not meet the minimum difference criterion.
[0091] In some embodiments, the frame comparison engine 240 can distinguish differences corresponding to the localized video data from differences caused by noise or coding artifacts by performing any number and / or type of operations on any amount and / or type of data in any technically feasible way. Also, in the same or other embodiments, the frame comparison engine 240 can ignore or exclude differences caused by noise or coding artifacts. In some embodiments, the frame comparison engine 240 can generate the modified frame index list 250, the comparison metadata 252, any amount and / or type of other output data, or any combination thereof, based on the differences corresponding to the localized video data (i.e., the differences caused by the video localization process).
[0092] Although illustration is omitted, in some embodiments, the frame comparison engine 240 can also compare the original video 122 and the localized video 124(1) instead of the decoded original video 210 to identify differences corresponding to the localized video data. In the same or other embodiments, the frame comparison engine 240 can use any amount and / or type of metadata to identify differences corresponding to the localized video data. For example, in some embodiments, when the localization application modifies the original video 122 to generate the localized video 124(1), metadata specifying the modified locations of each original video frame can also be generated. And in the same or other embodiments, according to any amount and / or type of metadata, the frame comparison engine 240 excludes or ignores any differences that do not correspond to the modified locations of the original video frames.
[0093] As shown, in some embodiments, the incremental encoding engine 260 generates the encoded localization layer 150(1) based on the decoded original video 210, the localized video 124(1), the modified frame index list 250, and the comparison metadata 252. In some embodiments, the encoded localization layer 150(1) includes, but is not limited to, any amount and / or any encoded metadata (hereinafter, "encoded metadata") 256 (including zero) and the encoded localization layer frames 258(1),..., the encoded localization layer frames 258(N). Here, N is the total number of frames included in the localized video 124(1).
[0094] In some embodiments, the encoded metadata 256 can instruct to reuse any number of decoded original video frames 212 (without modification) as part of the decoded localized video (not shown). In the same or other embodiments, the encoded localization layer frames 258(1), …, the encoded localization layer frames 258(N) are respectively the encoded versions of the localized video frames 224(1) to 224(N). For convenience of explanation, the encoded localization layer frames 258(1), …, the encoded localization layer frames 258(N) are respectively associated with frame indices from 1 to N. Also, in this specification, the encoded localization layer frames 258(1), …, the encoded localization layer frames 258(N) may be individually referred to as the “encoded localization layer frame 258”, or collectively referred to as the “encoded localization layer frames 258” and the “encoded localization layer frames 258(1) to 258(N)”.
[0095] Although illustrations are omitted, in some embodiments, each of the encoded localization layer frames 258 specifies, but is not limited to, any amount (including zero) of encoded prediction metadata and any amount (including zero) of encoded residual data. In some embodiments, the encoded prediction metadata included in the encoded localization layer frame 258(x) instructs the video decoder to copy one or more portions of the decoded original video frame 212(x) to the reconstructed prediction frame (optionally, in addition to this, move the one or more portions with respect to the reconstructed prediction frame). Here, x can be any integer from 1 to N. Also, in the same or other embodiments, the encoded residual data specifies the corrections that the video decoder needs to make to the reconstructed prediction frame to generate the decoded localized video frame of frame index "x" in the decoded localized video (not shown). In some embodiments, the decoded localized video frame of frame index "x" is not an exact replica of the localized video frame 224(x).
[0096] In some other embodiments, one or more encoded localization layer frames 258 can be omitted from the encoded localization layer 150(1), and this omission serves as an instruction to reuse, without modification, the frames of the decoded original video 210 corresponding to each omitted encoded localization layer frame 258 as the corresponding decoded localized video frames included in the decoded localized video. For example, if the encoded localization layer 150(1) does not include the encoded localization layer frame 258 of frame index "x", the encoded localization layer 150(1) implicitly gives an instruction to the video decoder to set the decoded localized video frame of frame index "x" as the decoded original video frame 212(x). Note that the technology described in this specification can be modified to represent the omission of any number of encoded localization layer frames 258.
[0097] As described in detail in the above description of FIG. 1, in some embodiments, any software application or any hardware module, such as the localization encoder 140, the localized video encoding application 120, the playback application 170, etc., can execute a layer interleaving algorithm to generate any portion (including the entire video) of the encoded localized video. Also, in the same or other embodiments, using the layer interleaving algorithm, based on the corresponding portion of the encoded original video 132, the corresponding portion of the encoded localization layer 150(1), and optionally any amount and / or type of metadata, any portion of the encoded localized video can be generated. In some embodiments, any portion of the encoded localized video can be decoded to generate the corresponding portion of the decoded localized video.
[0098] In some embodiments, by combining and using any portion of the encoded localization layer 150(1) with the corresponding portion of the decoded original video 210, the corresponding portion of the decoded localized video can be generated. Also, in the same or other embodiments, by decoding any portion of the encoded localization layer 150(1) in combination with the corresponding portion of the encoded original video 132, the corresponding portion of the decoded localized video can be generated.
[0099] In some embodiments, the incremental encoding engine 260 incrementally generates the encoded localization layer 150(1) by sequentially encoding the localized video frames 224 in the "encode / decode" order. In the same or other embodiments, each of the encoded localization layer frames 258 is a P frame (predicted frame). Generally, during encoding, if a predicted frame can be "backward predicted" using any part of the decoded localized video frame or the decoded original video frame 212 that is later in the display order, the encode / decode order may be different from the display order. In the same or other embodiments, zero or more of the encoded localization layer frames 258 are I frames (intra-coded frames), zero or more of the encoded localization layer frames 258 are P frames, and zero or more of the encoded localization layer frames 258 are B frames (bidirectional predicted frames).
[0100] For the sake of convenience of explanation, in this specification, some embodiments will be taken as examples where there are no predicted frames by backward prediction and thus the encode / decode order coincides with the display order of the localized video 124(1), and the function of the incremental encoding engine 260 will be described. However, in some other embodiments, the prediction of any number and / or type of parts of each predicted frame can be performed by intra prediction, forward prediction, backward prediction, or any combination thereof, so that the encode / decode order can be different from the display order, and the techniques described in this specification can be changed according to this configuration.
[0101] In some embodiments, the incremental encoding engine 260 can initialize the encoded localization layer 150(1) in any technically feasible manner based on any amount and / or type of data. For example, in some embodiments, the incremental encoding engine 260 can specify any amount and / or type of encoding metadata 256 that identifies any number and / or type of characteristics of the localized video 124(1), any amount and / or type of encoding-related information, and the like.
[0102] In the same or other embodiments, the incremental encoding engine 260 sequentially generates the encoded localization layer frames 258(1) to 258(N) by sequentially encoding the localized video frames 224(1) to 224(N). For the sake of convenience of explanation, taking the process of encoding the localized video frame 224(x) as an example, the functions of the incremental encoding engine 260 in some embodiments will be described. Here, x can be any integer from 1 to N.
[0103] In some embodiments, if the localized video frame 224(x) is not included in the modified frame index list 250, the incremental encoding engine 260 issues an instruction via the encoded localization layer 150(1) to "reuse the decoded original video frame 212(x) as the decoded localized video frame of frame index 'x' without modification". The incremental encoding engine 260 can instruct the encoded localization layer 150(1) to reuse the decoded original video frame 212(x) without modification by making any number of modifications (including zero).
[0104] In some embodiments, the incremental encoding engine 260 generates an encoding localization layer frame 258(x) that designates the decoded original video frame 212(x) to be reused as the decoded localized video frame for frame index "x" (without modification), and adds this encoding localization layer frame 258(x) to the encoding localization layer 150(1). The incremental encoding engine 260 can specify the reuse of the decoded original video frame 212(x) via the encoding localization layer frame 258(x) in any technically feasible way.
[0105] In some embodiments, to specify the reuse of the decoded original video frame 212(x) via the encoding localization layer frame 258(x), the incremental encoding engine 260 generates one or more "skip" instructions and / or any amount and / or type of skip metadata. The skip instructions and / or skip metadata indicate that the encoded localized video frame for frame index "x" is the same as the encoded original video frame for frame index "x". Then, the incremental encoding engine 260 encodes this skip instruction and / or skip metadata to generate encoded prediction metadata. Then, the incremental encoding engine 260 generates an encoding localization layer frame 258(x) that includes (but is not limited to) the encoded prediction metadata 256.
[0106] In some embodiments, if the frame index "x" is not included in the modified frame index list 250, the incremental encoding engine 260 does not generate the encoded localization layer frame 258 of the frame index "x". Thus, in that case, the encoded localization layer 150(1) does not include the encoded localization layer frame of the frame index "x". In the same or other embodiments, implicitly by omitting the encoded localization layer frame of the frame index "x" from the encoded localization layer 150(1), it is instructed to reuse the decoded original video frame 212(x) as the decoded localized video frame of the frame index "x" without modification.
[0107] On the other hand, as shown, in some embodiments, if the frame index "x" is included in the modified frame index list 250, the incremental encoding engine 260 causes an instance of the localization frame encoder 270 to perform one or more encoding operations on the localized video frame 224(x) to generate the encoded localization layer frame 258(x). And in the same or other embodiments, an instance of the incremental encoding engine 260 or the localization frame encoder 270 adds the encoded localization layer frame 258(x) to the encoded localization layer 150(1).
[0108] As described in the above explanation, in some embodiments, the modified frame index list 250 includes, but is not limited to, M frame indexes. In this specification, such frame indexes are represented as idx1 to idxM in the display order. Here, 1 ≦ idx1 ≦ idxM ≦ M ≦ N. In the same or other embodiments, the incremental encoding engine 260 configures the localization frame encoders 270(1) to 270(M) to encode the localized video frames 224(idx1) to 224(idxM) and generate the encoded localization layer frames 258(idx1) to 258(idxM), respectively. Note that in some other embodiments, the number of instances of the localization frame encoder 270 can be changed, and the technology described in this specification can be changed according to the changed configuration. For the sake of convenience of explanation, the function of the localization frame encoder 270 will be illustrated and described in detail by taking the localization frame encoder 270(1) as an example.
[0109] As shown in the illustration, in some embodiments, the localization frame encoder 270(1) generates the encoded localization layer frame 258(idx1) and the decoded localized video frame 292(idx1) based on the localized video frame 224(idx1), the buffer of the decoded frames (hereinafter, "decoded frame buffer") 262(idx1), and the comparison metadata 252. The localization frame encoder 270(1) can generate the encoded localization layer frame 258(idx1) by implementing any number and / or type of encoding ("coding") techniques. Also, in some embodiments, the localization frame encoder 270(1) can generate the decoded localized video frame 292(idx1) by implementing any number and / or type of decoding techniques. In some other embodiments, the localization frame encoder 270(1) does not generate the decoded localized video frame.
[0110] Although illustration is omitted, the localization frame encoder 270(1) can receive or identify any number and / or type of encoding parameter values, and optionally, in addition to this, any number and / or type of decoding parameter values, in any technically feasible way. Returning to FIG. 1, in some embodiments, the localized video encoding application 120 generates the encoded original video 132 and the encoded localization layers 150(1) to 150(K) based on a set of encoding parameter values. In the same or other embodiments, the localized video encoding application 120 generates the decoded original video 210 and the decoded localized video frame 292(idx1) based on a set of decoding parameter values.
[0111] In some embodiments, the decoded frame buffer 262(idx1) designates, but is not limited to, the decoded original video frame 212(idx1), zero or one or more other decoded original video frames 212, and zero or one or more decoded localized video frames. In this specification, the (one or more) decoded frames designated by the decoded frame buffer 262(idx1) (hereinafter, "decoded frames") may be individually referred to as "reference frame" or collectively as "reference frames".
[0112] In some embodiments, the localization frame encoder 270(1) can identify and effectively utilize zero or one or more spatial redundancies within the localized video frame 224(idx1), zero or one or more spatial redundancies between the localized video frame 224(idx1) and the decoded original video frame 212(idx1), and zero or one or more temporal redundancies between the localized video frame 224(idx1) and each of zero or one or more other reference frames. As used herein, "exploit" redundancy means reducing the number of bits allocated to represent the redundancy in the encoded localized video frame of frame index "idx1".
[0113] For the sake of convenience of explanation, in this specification, the function of the localization frame encoder 270(1) will be described by taking as an example the decoded frame buffer 262(idx1) that includes (but is not limited to) the decoded frame buffer 262(idx1) and at most one other reference frame. More precisely, when idx1 is "1", that is, when the localized video frame 224(idx1) is the first localized video frame 224 in the display order, the decoded frame buffer 262(idx1) designates, although not limited to, the decoded original video frame 212(idx1). On the other hand, when idx1 is not "1", the decoded frame buffer 262(idx1) designates, although not limited to, the decoded original video frame 212(idx1) and the decoded localized video frame 292(idx1 - 1: the frame index obtained by subtracting 1 from idx1), which is the decoded localized video frame immediately preceding the localized video frame 224(idx1) in the display order.
[0114] As shown, in some embodiments, the localization frame encoder 270(1) includes, but is not limited to, a prediction engine 280, prediction metadata 282, a residual frame 284, quantized coefficients 286, an entropy encoding engine 290, an encoded localization layer frame 258(idx1), and a partial decoder 278. In some embodiments, the localization frame encoder 270(1) divides each of the localized video frame 224(idx1), the decoded original video frame 212(idx1), and any other reference frame into non-overlapping portions, such as any number and / or type of non-overlapping processing units, in any technically feasible way. In the same or other embodiments, the localization frame encoder 270(1) specifies the size and / or type of the portions, such as processing units, according to the video compression format and / or image format.
[0115] For example, in some embodiments, the localization frame encoder 270(1) divides each of the localized video frame 224(idx1), the decoded original video frame 212(idx1), and any other reference frame into processing units known as macroblocks. In the same or other embodiments, each macroblock can include, but is not limited to, one or more sample blocks for each of any number and / or type of color components. For example, in some embodiments, each macroblock includes, but is not limited to, four 8×8 sample blocks of the luma component and different 8×8 sample blocks for each of the two chroma components.
[0116] In some embodiments, the prediction engine 280 independently processes any number of non-overlapping portions of the localized video frame 224(idx1) to identify corresponding portions of a predicted frame (not shown) and corresponding portions of the prediction metadata 282 at the same location. In this way, in some embodiments, the prediction engine 280 identifies non-overlapping portions of the predicted frame and its prediction metadata 282. As used herein, the "co-located" portions of a frame refer to portions at the same location within the frame. In some embodiments, a portion of the prediction metadata 282 corresponding to a portion of the localized video frame 224(idx1) specifies, without limitation, one or more instructions for reconstructing the co-located portion of the predicted frame. In the same or other embodiments, the one or more instructions for reconstructing a portion of the predicted frame specify, without limitation, one or more predictors for that portion of the predicted frame. Each of these predictors specifies a different portion for each predictor for the localized video frame 224, the decoded original video frame 212(idx1), or any other reference frame.
[0117] For the sake of convenience of explanation, in the following, the function of the prediction engine 280 in some embodiments will be described in detail by taking as an example the case where the non-overlapping portions of the frame are non-overlapping macroblocks included in the frame. Accordingly, in some embodiments, the prediction engine 280 independently processes any number of non-overlapping macroblocks included in the localized video frame 224(idx1) to identify corresponding macroblocks at the same location in a predicted frame (not shown) and corresponding portions of the prediction metadata 282. In this way, in some embodiments, the prediction engine 280 identifies non-overlapping macroblocks included in the predicted frame and its prediction metadata 282.
[0118] In some embodiments, a portion of the prediction metadata 282 corresponding to a macroblock included in the localized video frame 224(idx1) specifies, but is not limited to, one or more instructions for reconstructing the co-located macroblock of the prediction frame. As used herein, a "co-located" macroblock refers to a macroblock having the same position and size within a frame in different frames. For the sake of convenience of description, in this specification, a macroblock included in the prediction frame may be referred to as a "prediction macroblock". In some embodiments, the one or more instructions for reconstructing the prediction macroblock specify, but are not limited to, one or more predictors for the prediction macroblock. Each of these predictors specifies a different macroblock for each predictor with respect to the localized video frame 224, the decoded original video frame 212(idx1), or any other reference frame.
[0119] In some embodiments, the one or more instructions specifying the predictors of the prediction macroblock include, but are not limited to, a reference frame identifier and a motion vector. In the same or other embodiments, the reference frame identifier and the motion vector identify, as predictors used to construct the prediction macroblock in any technically feasible manner, a macroblock within the localized video frame 224, the decoded original video frame 212(idx1), or any other reference frame. In some embodiments, the reference frame identifier specifies, in any technically feasible manner, one of the localized video frame 224, the decoded original video frame 212(idx1), or any other reference frame. Also, in the same or other embodiments, the motion vector specifies the distance and direction from the prediction macroblock included in the prediction frame to the macroblock within the frame identified by the reference frame identifier.
[0120] For the sake of convenience in explanation, from the perspective of processing one macroblock included in the localized video frame 224(idx1), the function of the prediction engine 280 in some embodiments will be described. In this specification, the localized video frame 224 being processed by the prediction engine 280 may be referred to as the "target frame", and the macroblock of the target frame may be referred to as "a part of the target frame" and "target macroblock".
[0121] In some embodiments, to process the target macroblock, the prediction engine 280 determines, based on the comparison metadata 252, whether the target macroblock contains visible localized video data. In some other embodiments, the prediction engine 280 can also identify or estimate whether the target macroblock contains visible localized video data and / or localized video data in any other technically feasible way. For example, in some embodiments, the prediction engine 280 compares the target macroblock with the co-located macroblock in the decoded original video frame 212(idx1) to estimate whether the target macroblock contains visible localized video data.
[0122] In some embodiments, if the prediction engine 280 determines that the target macroblock does not contain visible localized video data, the prediction engine 280 sets the predicted macroblock in the prediction frame to be the co-located macroblock in the decoded original video frame 212(idx1). Also, in the same or other embodiments, the prediction engine 280 adds to the prediction metadata 282 one or more instructions that direct the reuse of the co-located macroblock in the decoded original video frame 212(idx1) as the predicted macroblock in the prediction frame. For example, in some embodiments, the one or more instructions direct the decoder to set the predicted macroblock to be the co-located macroblock in the decoded original video frame 212(idx1).
[0123] In some embodiments, the one or more instructions that direct the decoder to set the predicted macroblock to be the co-located macroblock in the decoded original video frame 212(idx1) are not limited, but are those that specify a reference frame identifier corresponding to the decoded original video frame 212(idx1) and a motion vector with a distance of zero. In the same or other embodiments, the motion vector with a distance of zero indicates that the predictor of the predicted macroblock in the prediction frame is the co-located macroblock in the frame corresponding to the reference frame identifier.
[0124] On the one hand, when the target macroblock contains localizable video data that is visible, the prediction engine 280 selects zero or one or more of the macroblocks included in the localizable video frame 224(idx1), the decoded original video frame 212(idx1), and any number of other reference frames as candidate macroblocks. The prediction engine 280 can select candidate macroblocks in any technically feasible way. For example, in some embodiments, the prediction engine 280 can select candidate macroblocks by implementing any number and / or type of search algorithms. Next, the prediction engine 280 evaluates the selected candidate macroblocks to identify the macroblock with the highest degree of match (best match) with the target macroblock (hereinafter, the "best match macroblock"). The prediction engine 280 can evaluate candidate macroblocks in any technically feasible way.
[0125] In some embodiments, for each candidate macroblock, the prediction engine 280 calculates the mean squared error, mean absolute difference, peak signal-to-noise ratio, or any combination thereof between the pixel values of the target macroblock and the pixel values of the corresponding candidate prediction. For example, if the candidate macroblock is a macroblock included in the decoded original video frame 212(idx1), the corresponding candidate prediction is the candidate macroblock shifted to be in the same position as the target macroblock. Then, when evaluating the candidate macroblock, the prediction engine 280 sets the candidate macroblock with the least error, or the highest similarity, or the lowest dissimilarity compared to the target macroblock as the best match macroblock.
[0126] In some embodiments, the prediction engine 280 sets the predicted macroblock within the prediction frame as the best match macroblock. Also, in the same or other embodiments, the prediction engine 280 adds to the prediction metadata 282 one or more instructions that direct the decoder to set the predicted macroblock as the best match macroblock. In some embodiments, the instructions that direct the decoder to set the predicted macroblock as the best match macroblock are not limited, but specify a reference frame identifier corresponding to the frame containing the best match macroblock and a motion vector from the predicted macroblock to the best match macroblock. More specifically, in some embodiments, the motion vector specifies the distance and direction from the position of the predicted macroblock within the prediction frame to the position of the best match macroblock within the frame corresponding to the reference frame identifier. In some embodiments, when the best match macroblock is in the same position as the target macroblock, the prediction engine 280 sets the distance of the motion vector to zero.
[0127] In some embodiments, the video decoder identifies a reference frame based on the reference frame identifier and a portion of the prediction metadata 282 that specifies a motion vector for identifying a predictor of a predicted macroblock included in the prediction frame. Next, the video decoder maps the position of the predicted macroblock within the prediction frame to the best match macroblock within the reference frame according to the motion vector. In some embodiments, the decoder implements a virtually reconstructed prediction frame and sets the virtual predicted macroblocks included in the virtually reconstructed prediction frame as the best match macroblocks. In the same or other embodiments, the decoder copies the best match macroblocks to the predicted macroblocks within the prediction frame reconstructed from the reference frame.
[0128] In some embodiments, upon generating the last predicted macroblock in the prediction frame, the prediction engine 280 generates a residual frame 284 by subtracting the prediction frame from the target frame. Thus, in some embodiments, the residual frame 284 represents the prediction error between the prediction frame and the target frame in the spatial domain. Next, the localization frame encoder 270(1) generates an encoded localization layer frame 258(idx1). The encoded localization layer frame 258(idx1) encodes, but is not limited to, instructions for reconstructing the prediction frame and zero or one or more modifications to the reconstructed prediction frame to eliminate or reduce the prediction error of the prediction frame reconstructed by this instruction. The localization frame encoder 270(1) can generate the encoded localization layer frame 258(1) in any technically feasible manner.
[0129] In some embodiments, the localization frame encoder 270(1) converts the residual frame 284 from the spatial domain to the frequency domain by performing any number and / or type of conversion operations on the residual frame 284. For example, in some embodiments, the localization frame encoder 270(1) divides the residual frame 284 into any number of non-overlapping blocks in any technically feasible way. In the same or other embodiments, the localization frame encoder 270(1) converts the pixel values within each block of the residual frame 284 from the spatial domain to the frequency domain to generate conversion coefficients (not shown). In the same or other embodiments, the localization frame encoder 270(1) can perform any type of two-dimensional linear transformation on the pixel values within each block of the residual frame 284 to generate conversion coefficients. For example, in some embodiments, the localization frame encoder 270(1) uses, but is not limited to, a two-dimensional discrete cosine transform (DCT), discrete Fourier transform, discrete sine transform, or discrete Haar transform to convert the pixel values to conversion coefficients.
[0130] In some embodiments, the localization frame encoder 270(1) generates the quantized coefficients 286 by performing any number and / or type of quantization operations on the conversion coefficients. The localization frame encoder 270(1) can quantize the conversion coefficients in any technically feasible way by performing any number and / or any type of quantization operations on the conversion coefficients. And in some embodiments, the localization frame encoder 270(1) generates the encoded residual data (not shown) by executing an instance of the entropy encoding engine 290 on the quantized coefficients 286. Also, in the same or other embodiments, the localization frame encoder 270(1) generates the encoded prediction metadata (not shown) by executing an instance of the entropy encoding engine 290 on the prediction metadata 282.
[0131] In some embodiments, the entropy encoding engine 290 can perform any number and / or type of reversible compression operations on any amount and / or type of data or symbols, and in addition to / or instead of, by applying any number and / or type of reversible compression techniques, encoded data (i.e., code) of a smaller size can be generated for the data or symbols. For example, in some embodiments, the entropy encoding engine 290 uses run-length encoding techniques to replace continuously appearing symbols with one such symbol and its repetition count. In the same or other embodiments, the entropy encoding engine 290 uses variable-length encoding techniques to assign short codes to symbols with a high occurrence probability and long codes to symbols with a low occurrence probability.
[0132] More specifically, in some embodiments, the entropy encoding engine 290 performs any number and / or type of run-length encoding operations, any number and / or type of variable-length encoding operations, any number and / or type of other reversible compression operations, or any combination thereof on the quantized coefficients 286 to generate encoded residual data of a smaller size for the quantized coefficients 286. Also, in the same or other embodiments, the entropy encoding engine 290 performs any number and / or type of run-length encoding operations, any number and / or type of variable-length encoding operations, any number and / or type of other reversible compression operations, or any combination thereof on the prediction metadata 282 to generate encoded prediction metadata (not shown) of a smaller size for the prediction metadata 282. As shown, in some embodiments, the localization frame encoder 270(1) generates an encoded localization layer frame 258(idx1) that includes, but is not limited to, the encoded residual data and the encoded prediction metadata.
[0133] In some embodiments, the localization frame encoder 270(1) generates an encoded localization layer frame 258(idx1) according to one or more video standards, encoding standards, and / or compression standards. For example, in some embodiments, the encoded localization layer frame 258(idx1) is a P-frame. When the localization frame encoder 270(1) generates the encoded localization layer frame 258(idx1), the localization frame encoder 270(1) or the incremental encoding engine 260 adds the encoded localization layer frame 258(idx1) to the encoded localization layer 150(1).
[0134] In some embodiments, the localization frame encoder 270(1) uses a partial decoder 278 to generate a decoded localized video frame 292(idx1) that matches the encoded data included in the encoded localization layer frame 258(idx1) in any technically feasible way. In the same or other embodiments, the partial decoder 278 generates the decoded localized video frame 292(idx1) based on the quantized coefficients 286 and the predicted frame by implementing any number and / or type of decoding techniques.
[0135] More specifically, in some embodiments, the partial decoder 278 inverse quantizes the quantized coefficients 286 in any technically feasible manner to generate decoded transform coefficients (not shown). For example, in some embodiments, the partial decoder 278 can generate decoded transform coefficients by performing any number and / or type of inverse quantization operations on the quantized coefficients 286. Then, the partial decoder 278 performs an inverse transform (e.g., inverse 2D DCT) on each block of the decoded transform coefficients to generate a decoded residual frame (not shown). In some embodiments, the decoded residual frame is not an exact replica of the residual frame 284. Next, the partial decoder 278 generates a decoded localized video frame 292(idx1) by adding the decoded residual frame to the predicted frame. Although not shown, in some embodiments, the incremental encoding engine 260 includes the decoded localized video frame 292(idx1) in the decoded frame buffer 262 associated with the frame index (idx1 + 1) obtained by adding 1 to idx1.
[0136] In some embodiments, the functions of other instances of the localization frame encoder 270 are also similar to the functions described above. Specifically, as shown, the localization frame encoder 270(M) generates an encoded localization layer frame 258(idxM) and optionally a decoded localized video frame 292(idxM) based on the localized video frame 224(idxM), the decoded frame buffer 262(idxM), and the comparison metadata 252. Also, in the same or other embodiments, the decoded frame buffer 262(idxM) specifies, but is not limited to, the decoded original video frame 212(idxM) and the decoded localized video frame 292(idxM - 1: the frame index obtained by subtracting 1 from idxM). As will be appreciated by those skilled in the art, the techniques described herein can be modified to reflect any number and / or type of reference frames.
[0137] Although illustration is omitted, in some embodiments, the localization encoder 140(1) can generate any amount and / or type of "localization" metadata related to streaming the localized video 124(1), constructing the corresponding encoded localized video, or both. For example, in some embodiments, the localization metadata can include any amount and / or type of data that can be used by the cloud-based video service 104 or any other software application to generate a manifest file. In the same or other embodiments, the localization metadata can include any amount and / or type of data that can be used in combination with a chunk of the encoded original video 132 and the corresponding chunk of the encoded localization layer 150(1) by the playback application 170, or any other software application, or any hardware module executing a layer interleaving algorithm to generate chunks of the encoded version of the corresponding localized video 124(1). Also, in some embodiments, the localization encoder 140(1), the localized video encoding application 120, or both can transmit any portion of the localization metadata to any number of software applications, hardware modules, or any combination thereof.
[0138] Figure 3 is a flowchart showing method steps for encoding a localized video according to various embodiments. Note that the description of the method steps is made with reference to the systems shown in FIGS. 1 and 2, but those skilled in the art will understand that the scope of the various embodiments includes any system configured to perform the method steps in any order.
[0139] As shown, method 300 begins at step 302 where localization encoder 140 obtains the encoded original video 132 associated with the original video 122 and the localized video 124. Then, at step 304, the localization encoder 140 decodes the encoded original video 132 to generate the decoded original video 210. Next, at step 306, the localization encoder 140 evaluates the localized video 124 together with one or more of the decoded original video 210, the original video 122, and the metadata to identify the visible localized video data included in the frames of the localized video 124. And at step 308, the localization encoder 140 initializes the encoded localization layer 150 and selects the first frame of the localized video 124.
[0140] Next, at step 310, the localization encoder 140 determines whether the selected frame contains perceivable localized video data. If at step 310 the localization encoder 140 determines that the selected frame does not contain visible localized video data, method 300 proceeds to step 312. And at step 312, the localization encoder 140 instructs the encoded localization layer 150 to reuse the corresponding frame of the decoded original video 210 as is (without modification). Next, method 300 proceeds directly to step 318.
[0141] On the one hand, in step 310, if the localization encoder 140 determines that the selected frame contains visible localized video data, method 300 proceeds directly to step 314. Then, in step 314, the localization frame encoder 270 generates an encoded localization layer frame 258 corresponding to the selected frame by encoding the selected frame while using zero or one or more decoded original video frames 212, zero or one or more decoded localized video frames, or both as reference frames. Next, in step 316, the localization encoder 140 adds the encoded localization layer frame 258 corresponding to the selected frame to the encoded localization layer 150.
[0142] Then, in step 318, the localization encoder 140 determines whether the selected frame is the last frame of the localized video 124. If, in step 318, the localization encoder 140 determines that the selected frame is not the last frame of the localized video 124, method 300 proceeds to step 320. Then, in step 320, the localization encoder 140 selects the next frame of the localized video 124, and method 300 returns to step 310, where the localization encoder 140 determines whether the selected frame contains visible localized video data.
[0143] On the other hand, if, in step 318, the localization encoder 140 determines that the selected frame is the last frame of the localized video 124, method 300 proceeds directly to step 322. Then, in step 322, the localized video encoding application 120 transmits the encoded original video 132 and the encoded localization layer 150 to the one or more server devices to stream the localized video 124 to the end-user device via the one or more server devices. Thereafter, method 300 ends.
[0144] That is, by using the technology of the present disclosure, it is possible to reduce the amount of original video data that is double-encoded when encoding localized video. In some embodiments, the localized video encoding application encodes the original video to generate an encoded original video. Then, the localized video encoding application executes a localization encoder on the localized video and the encoded original video. The localization encoder decodes the encoded original video to generate a decoded original video. The localization encoder compares the frames of the localized video with the frames of the decoded original video to generate a list of frames of the localized video that contain visible localized video data (and thus look different from the corresponding portions of the decoded original video when displayed). In some embodiments, the localization encoder also generates comparison metadata indicating zero or one or more portions in each frame of the localized video that contain these perceptible localized video data.
[0145] The localization encoder initializes an encoded localization layer and then encodes each frame of the localized video in display order. The localization encoder performs the following processing for each frame of the localized video. If the frame of the localized video does not contain visible localized video data, the localization encoder instructs the encoded localization layer to reuse the corresponding frame of the decoded original video as part of the decoded localized video without modification.
[0146] On the one hand, when the frame of the localized video contains visible localized video data, the localization encoder adds the corresponding frame of the decoded original video to the decoded frame buffer of a plurality of decoded frames that can be reused during encoding. Next, the localization encoder encodes the frame of the localized video based on the decoded frame buffer to generate an encoded localization layer frame. In some embodiments, the encoded localization layer frame includes, but is not limited to, encoded prediction metadata and, optionally in addition to this, encoded residual data. In some embodiments, the encoded prediction metadata encodes instructions for constructing a reconstructed prediction frame based on, but not limited to, one or more portions of the corresponding frame of the decoded original video, one or more portions of one or more other decoded frames specified in the decoded frame buffer, or both. Note that the reconstructed prediction frame can include, but is not limited to, copies of any number of portions of the corresponding decoded original video frame, and optionally, for each copy of these portions of the corresponding decoded original video frame, the position within the reconstructed prediction frame can be moved from the position in the decoded original video frame. In some embodiments, the encoded residual data specifies a correction for reducing the residual error performed on the reconstructed prediction frame for generating the frame of the decoded localized video.
[0147] In some embodiments, the encoded original video and the encoded localization layer are each stored independently and delivered to the end-user device via separate streams. The playback application running on the end-user device constructs each encoded localized video chunk based on the encoded original video chunk corresponding to each encoded localized video chunk and the corresponding encoded localization layer chunk. More specifically, the playback application identifies, based on the encoded localized video chunk, the frames of the localized video chunk that have differences with respect to the corresponding frames of the original video chunk. For each frame identified in the localized video chunk, the playback application generates an encoded base layer chunk by marking the frame of the corresponding encoded original video chunk as a reference-only frame. Then, the playback application interleaves the frames of the encoded base layer chunk with the frames of the encoded localization layer chunk to generate the encoded localized video chunk.
[0148] One of the technical advantages of the technology of the present disclosure over the prior art is that the technology of the present disclosure can at least reduce the amount of original video data that is multiplexed and encoded when encoding localized video. In this sense, in each frame of the encoded localization layer, any number of parts (including the entire frame) of the corresponding reconstructed original video frame can be specified for reuse. Thereby, the amount of memory used to store the encoded localization layer chunks used to construct the encoded localized video chunks can be significantly reduced compared to the amount of memory required to store the encoded localized video chunks using prior art techniques. Another technical advantage of the technology of the present disclosure is that by storing the encoded localization layer chunks instead of the encoded localized video chunks for streaming localized video via a CDN, it is possible to improve the cache efficiency of the CDN and thus improve the QoE of the end user. These technical advantages result in one or more technical improvements over prior art approaches.
[0149] 1. In some embodiments, a computer-implemented method for encoding localized video includes calculating a predicted frame based on a target frame of the localized video and at least a portion of a reference frame of a decoded original video, calculating a residual frame based on the predicted frame and the target frame of the localized video, performing one or more encoding operations on the residual frame to generate a frame of an encoded localization layer, and transmitting the frame of the encoded localization layer and at least one frame of an encoded original video to another device for decoding.
[0150] 2. The step of calculating the predicted frame includes determining that the similarity between the part of the reference frame of the decoded original video and the part of the target frame of the localized video is higher than the similarity between the part of the frame of the decoded localized video and the part of the target frame of the localized video, and generating prediction metadata indicating that the part of the reference frame of the decoded original video is a predictor of the part of the predicted frame corresponding to the part of the target frame. The computer-implemented method according to claim 1.
[0151] 3. The computer-implemented method according to claim 1 or 2 further includes the step of performing one or more decoding operations on the encoded residual data included in the frame of the encoded localization layer to generate a frame of the decoded localized video.
[0152] 4. The step of performing the one or more encoding operations on the residual frame includes performing at least one of a transform operation, a quantization operation, and an invertible compression operation on the residual frame to generate encoded residual data. The computer-implemented method according to any one of claims 1 to 3.
[0153] 5. The frame of the encoded localization layer includes at least one of the encoded residual data associated with the residual frame and the encoded prediction metadata associated with the predicted frame. The computer-implemented method according to any one of claims 1 to 4.
[0154] 6. The computer-implemented method according to any one of claims 1 to 5 further includes the step of decoding the at least one frame of the encoded original video to generate the reference frame of the decoded original video.
[0155] 7. The method implemented on a computer according to any one of claims 1 to 6, wherein the step of calculating the residual frame includes subtracting the predicted frame from the target frame of the localized video.
[0156] 8. The method implemented on a computer according to any one of claims 1 to 7, further comprising: performing one or more comparison operations between the frames of the localized video and the frames of the decoded original video to determine that the frames of the localized video do not contain visible correction locations due to video localization processing; and instructing the encoded localization layer to reuse the frames of the decoded original video as part of the decoded localized video.
[0157] 9. The method implemented on a computer according to any one of claims 1 to 8, further comprising: generating an encoded base layer by marking the at least one frame of the encoded original video as a reference-only frame; and interleaving a plurality of frames included in the encoded base layer with a plurality of frames included in the encoded localization layer to generate at least one chunk of the encoded localized video.
[0158] 10. The method implemented on a computer according to any one of claims 1 to 9, wherein the frames of the encoded localization layer include P frames (predicted frames) or B frames (bidirectional predicted frames).
[0159] 11. In some embodiments, one or more non-transitory computer-readable media include instructions that, when executed by one or more processors, cause the one or more processors to encode localized video. The instructions cause the one or more processors to calculate a predicted frame based on a target frame of the localized video and at least a portion of a reference frame of the decoded original video, calculate a residual frame based on the predicted frame and the target frame of the localized video, perform one or more encoding operations on the residual frame to generate a frame of an encoded localization layer, and transmit the frame of the encoded localization layer and at least one frame of the encoded original video to another device for decoding.
[0160] 12. The one or more non-transitory computer-readable media of claim 11, wherein the step of calculating the predicted frame includes determining that a similarity between the portion of the reference frame of the decoded original video and the portion of the target frame of the localized video is higher than a similarity between a second portion of the reference frame of the decoded original video and the portion of the target frame of the localized video.
[0161] 13. The one or more non-transitory computer-readable media of claim 11 or 12, further including generating encoded prediction metadata included in the frame of the encoded localization layer by encoding a motion vector from a portion of the predicted frame to the portion of the reference frame of the decoded original video.
[0162] 14. The one or more non-transitory computer-readable media of any one of claims 11 to 13, wherein the step of performing the one or more encoding operations on the residual frame includes performing at least one of a transform operation, a quantization operation, and a reversible compression operation on the residual frame to generate encoded residual data.
[0163] 15. The frame of the encoded localization layer includes at least one of the encoded residual data associated with the residual frame and the encoded prediction metadata associated with the prediction frame, and the one or more non-transitory computer-readable media according to any one of claims 11 to 14.
[0164] 16. The one or more non-transitory computer-readable media according to any one of claims 11 to 15, wherein the instructions further cause the one or more processors to execute a step of decoding the encoded original video to generate the decoded original video.
[0165] 17. The one or more non-transitory computer-readable media according to any one of claims 11 to 16, wherein the step of calculating the residual frame includes a step of subtracting the prediction frame from the target frame of the localized video.
[0166] 18. The frame of the encoded localization layer is then decoded in combination with the at least one frame of the encoded original video when generating a first chunk of the decoded localized video, and the one or more non-transitory computer-readable media according to any one of claims 11 to 17.
[0167] 19. The instructions cause the one or more processors to generate an encoded base layer by marking the at least one frame of the encoded original video as a reference-only frame, and to interleave a plurality of frames included in the encoded base layer with a plurality of frames included in the encoded localization layer to generate at least one chunk of the encoded localized video, and the one or more non-transitory computer-readable media according to any one of claims 11 to 18.
[0168] 20. In some embodiments, the system comprises one or more memories storing instructions and one or more processors coupled to the one or more memories. By executing the instructions, the one or more processors perform the steps of calculating a predicted frame based on a target frame of a localized video and at least a portion of a reference frame of a decoded original video, calculating a residual frame based on the predicted frame and the target frame of the localized video, performing one or more encoding operations on the residual frame to generate a frame of an encoded localization layer, and transmitting the frame of the encoded localization layer and at least one frame of an encoded original video to another device for decoding.
[0169] Any component recited in any claim of the claims and / or any combination of elements described herein, in any form of combination whatsoever, all such combinations are intended to be included within the scope of the present invention and protection.
[0170] Although various embodiments have been described, these are presented for illustrative purposes only and are not intended to cover or limit all embodiments of the present disclosure. Many changes and modifications will be apparent to those skilled in the art that do not depart from the scope and spirit of the embodiments described herein.
[0171] Aspects of the present embodiment can be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure can take the form of an embodiment that is entirely hardware only, an embodiment that is entirely software only (including firmware, resident software, microcode, etc.), or an embodiment that combines aspects of software and aspects of hardware, and all of these are collectively referred to herein as "modules" or "systems" in some cases. Further, aspects of the present disclosure can also take the form of a computer program product and can be embodied as computer-readable program code on one or more computer-readable media.
[0172] Also, one or more computer-readable media can be used in any combination. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination thereof. More specific examples of the computer-readable storage medium (but not an exhaustive list) include, without limitation, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, the computer-readable storage medium can be any tangible medium that can store (memorize) a program for use by or in connection with a system, apparatus, or device that executes instructions.
[0173] Above, aspects of the present disclosure have been described with reference to flowchart diagrams and / or block diagrams showing methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart diagrams and / or block diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, can be implemented by computer program instructions. And by providing these computer program instructions to a processor of a programmable data processing apparatus such as a general-purpose computer or a dedicated computer, a machine can be manufactured. The instructions, when executed by a processor of a programmable data processing apparatus such as a computer, implement the functions / operations specified in one or more blocks of the flowchart and / or block diagram. Such processors can include, but are not limited to, general-purpose processors, dedicated processors, application-specific processors, or field programmable gate arrays.
[0174] The flowcharts and block diagrams shown in the drawings can represent the architectures, functionality, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure, which can be embodiments. In this sense, each block in the flowchart or block diagram can be considered to represent a module, segment, or part of code that includes one or more executable instructions for implementing the (one or more) logical functions specified therein. Also, note that in some alternative embodiments, the functions described in the blocks can be executed in an order different from the order described in the drawings. For example, two blocks shown as consecutive blocks can actually be executed substantially simultaneously, or alternatively, depending on the relevant functions, these blocks can also be executed in the reverse order. Also, note that each block of the block diagram and / or flowchart diagram, and combinations of blocks of the block diagram and / or flowchart diagram, can be implemented by a dedicated hardware-based system that executes the functions or operations specified therein, or alternatively, can also be implemented by a combination of dedicated hardware and computer instructions.
[0175] Note that the above description is directed to embodiments of the present disclosure, but it is also possible to devise other embodiments and further embodiments of the present disclosure without departing from the basic scope of the present disclosure, and such scope is defined by the following claims.
Description of Reference Numerals
[0176] 100 System 102 Display Device 104 Cloud-Based Video Service 110 Compute Instance 112 Processor 116 Memory 120 Localized Video Encoding Application 122 Original Video 124 Localized Video 130 Video Encoder 132 Encoded Original Video 140 Localization Encoder 150 Encoded Localization Layer 170 Playback Application 172 Manifest File 174 Playback Buffer 176 Encoded Localized Video Chunk 178 Multilayer Video Decoder 182 Encoded Video Chunk Request 184 Encoded Layer Chunk Request 186 Encoded Original Video Chunk 188 Encoded Localization Layer Chunk 190 Decoded Localized Video Chunk 202 Video Decoder 210 Decoded Original Video 212 Decoded Original Video Frame 224 Localized Video Frame 232 Encoded Original Video Frame 240 Frame Comparison Engine 250 Modified Frame Index List 252 Comparison Metadata 256 Encoded Metadata, Encoded Prediction Metadata 258 Encoded Localization Layer Frame 260 Incremental Encoding Engine 262 Decoded Frame Buffer 270 Localization Frame Encoder 278 Partial Decoder 280 Prediction Engine 282 Prediction Metadata 284 Residual Frame 286 Quantized Coefficients 290 Entropy Encoding Engine 292 Decoded Localized Video Frame
Claims
1. A computer-implemented method for encoding a localized video, comprising: calculating a predicted frame based on a target frame of the localized video and at least a part of a reference frame of a decoded original video; calculating a residual frame based on the predicted frame and the target frame of the localized video; performing one or more encoding operations on the residual frame to generate a frame of an encoded localization layer; transmitting the frame of the encoded localization layer and at least one frame of an encoded original video to another device for decoding; A method comprising the above steps.
2. The step of calculating the predicted frame includes: determining that a similarity between the part of the reference frame of the decoded original video and a part of the target frame of the localized video is higher than a similarity between a part of a frame of a decoded localized video and the part of the target frame of the localized video; generating prediction metadata indicating that the part of the reference frame of the decoded original video is a predictor of a part of the predicted frame corresponding to the part of the target frame; The computer-implemented method according to claim 1, comprising the above steps.
3. The computer-implemented method according to claim 1, further comprising performing one or more decoding operations on the encoded residual data included in the frame of the encoded localization layer to generate a frame of a decoded localized video.
4. The step of performing the one or more encoding operations on the residual frame includes performing at least one of a transform operation, a quantization operation, and an invertible compression operation on the residual frame to generate encoded residual data. The computer-implemented method according to claim 1, comprising the above steps.
5. The computer-implemented method according to claim 1, wherein the frame of the encoded localization layer includes at least one of encoded residual data associated with the residual frame and encoded prediction metadata associated with the predicted frame.
6. The computer-implemented method according to claim 1, further comprising the step of decoding the at least one frame of the encoded original video to generate the reference frame of the decoded original video.
7. The computer-implemented method according to claim 1, wherein the step of calculating the residual frame includes the step of subtracting the predicted frame from the target frame of the localized video.
8. Performing one or more comparison operations between the frame of the localized video and the frame of the decoded original video to determine that the frame of the localized video does not contain visible correction locations due to video localization processing; Instructing the encoded localization layer to reuse the frame of the decoded original video as part of the decoded localized video; The computer-implemented method according to claim 1, further comprising.
9. Generating an encoded base layer by marking the at least one frame of the encoded original video as a reference-only frame; Interleaving a plurality of frames included in the encoded base layer with a plurality of frames included in the encoded localization layer to generate at least one chunk of the encoded localized video; The computer-implemented method according to claim 1, further comprising.
10. The computer-implemented method according to claim 1, wherein the frame of the encoded localization layer includes a P frame (predicted frame) or a B frame (bidirectional predicted frame).
11. One or more non-transitory computer-readable media including instructions that, when executed by one or more processors, cause the one or more processors to Calculate a predicted frame based on a target frame of the localized video and at least a portion of a reference frame of the decoded original video; Calculate a residual frame based on the predicted frame and the target frame of the localized video; Performing one or more encoding operations on the residual frame to generate a frame of an encoded localization layer; Transmitting the frame of the encoded localization layer and at least one frame of the encoded original video to another device for decoding; One or more non-transitory computer-readable media for causing the above to be executed.
12. The step of calculating the prediction frame includes determining that the similarity between the part of the reference frame of the decoded original video and the part of the target frame of the localized video is higher than the similarity between the second part of the reference frame of the decoded original video and the part of the target frame of the localized video. The one or more non-transitory computer-readable media according to claim 11.
13. The one or more non-transitory computer-readable media according to claim 11, further comprising generating encoded prediction metadata included in the frame of the encoded localization layer by encoding a motion vector from a part of the prediction frame to the part of the reference frame of the decoded original video.
14. The step of performing the one or more encoding operations on the residual frame includes performing at least one of a conversion operation, a quantization operation, and a reversible compression operation on the residual frame to generate encoded residual data. The one or more non-transitory computer-readable media according to claim 11.
15. The frame of the encoded localization layer includes at least one of the encoded residual data associated with the residual frame and the encoded prediction metadata associated with the prediction frame. The one or more non-transitory computer-readable media according to claim 11.
16. The one or more non-transitory computer-readable media according to claim 11, wherein the instructions further cause the one or more processors to execute the step of decoding the encoded original video to generate the decoded original video.
17. The step of calculating the residual frame includes subtracting the prediction frame from the target frame of the localized video. The one or more non-transitory computer-readable media according to claim 11.
18. The one or more non-transitory computer-readable media of claim 11, wherein the frame of the encoded localization layer is then decoded in combination with the at least one frame of the encoded original video when generating a first chunk of the decoded localized video.
19. The instructions cause the one or more processors to generate an encoded base layer by marking the at least one frame of the encoded original video as a reference-only frame; generate at least one chunk of the encoded localized video by interleaving a plurality of frames included in the encoded base layer with a plurality of frames included in the encoded localization layer; The one or more non-transitory computer-readable media of claim 11, further causing the above to be executed.
20. One or more memories storing instructions; One or more processors coupled to the one or more memories; A system comprising: By executing the instructions, the one or more processors calculate a predicted frame based on a target frame of the localized video and at least a part of a reference frame of the decoded original video; calculate a residual frame based on the predicted frame and the target frame of the localized video; perform one or more encoding operations on the residual frame to generate a frame of the encoded localization layer; send the frame of the encoded localization layer and at least one frame of the encoded original video to another device for decoding; A system that performs the above.
Citation Information
Patent Citations
Video compression apparatus, video reproduction apparatus and video distribution system
JP2016092837A
Transmission device and reception device
JP2022053534A
Video compression apparatus, video playback apparatus and video delivery system
US20160127728A1
Hybrid broadcast and communication system, data generation device, and receiver
WO2013021643A1
Delivery device and receiving device
WO2023106259A1