Adaptive image encoding and decoding for resolution limited systems

By segmenting high-resolution images into multiple regions and generating metadata, the challenge of encoding and decoding high-resolution images on devices with limited computing resources is solved, achieving an efficient encoding and decoding process, reducing resource requirements and artifacts.

CN121814959APending Publication Date: 2026-04-07NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

On devices with limited computing resources, high-resolution image encoding and decoding operations become infeasible due to a lack of contiguous memory and processing resources, especially on edge devices or systems lacking large amounts of video memory.

Method used

High-resolution images or video frames are segmented into multiple regions, metadata about the relative position and size of each region is generated, and each region is encoded using a suitable encoding algorithm. After generating a media packet, it is stored or transmitted. During decoding, the high-resolution image is reconstructed using the metadata, and filtering is applied to reduce artifacts.

Benefits of technology

It enables efficient and adaptive encoding and decoding of high-resolution images on systems with limited computing resources, reducing the demand for storage and processing resources, improving encoding efficiency and reducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814959A_ABST
    Figure CN121814959A_ABST
Patent Text Reader

Abstract

The invention relates to adaptive image encoding and decoding for resolution limited systems. In various examples, a system and method are disclosed that pertains to encoding and decoding high resolution images on a system that supports limited resolution. The system may identify an image to be encoded using an image encoding process. The system may extract a plurality of regions from the image and generate a plurality of encoded regions by encoding each region using an image encoding process. The system may generate a media package using a plurality of encoded regions and metadata corresponding to the plurality of regions.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] High resolution image encoding involves compressing large amounts of pixel data into a smaller file size for ease of transmission and storage. This compression process uses algorithms to reduce redundancy in the image data, thereby minimizing the overall file size without significantly impacting visual quality. However, encoding ultra-high resolution images can quickly exhaust the processing power and memory resources of a computing device. SUMMARY

[0002] Embodiments of the present disclosure provide techniques for adaptively encoding high resolution media (e.g., images / videos) on platforms characterized by limited computing resources. Certain media systems (e.g., systems used for industrial inspection or medical imaging) can generate media with resolutions up to several million pixels. While it is feasible to encode and decode these high resolution images for processing on high performance distributed computing systems equipped with specialized hardware, it proves impractical when using regular computing devices with limited resources.

[0003] Conventional high resolution media encoding and decoding operations involve processing the entire high resolution image at once. However, conventional computing systems lack the continuous available memory and sufficient processing resources to meet the storage and processing demands of the entire high resolution image. This processing challenge is particularly evident on edge devices or computing systems that lack large amounts of video memory for image data manipulation. These inherent limitations make it infeasible to encode or decode high resolution media on conventional computing systems.

[0004] The techniques described herein can be used to efficiently and adaptively encode and decode high resolution media on computing systems with limited computing power. To enable efficient encoding of media data, an input high resolution image or frame from a high resolution video can be partitioned into a set of regions. Each region can contain a region of interest (ROI), a tile, or any other portion of the image / video intended for encoding. Metadata can be generated that indicates the relative position and size of each region, thereby enabling the mapping of data for each region back to its original position and dimensions within the high resolution image / video frame. Once the high resolution media is divided into multiple regions, each region can be encoded using a suitable media encoding algorithm. Subsequently, the encoded regions are combined with the generated metadata to form a media package (e.g., an image / video file) that can then be stored or transmitted for further processing.

[0005] At least one aspect relates to one or more processors. The one or more processors can include one or more circuits. The one or more circuits can extract a plurality of regions from an image to be encoded using an image encoding process. The one or more circuits can generate a plurality of encoded regions by encoding each of the plurality of regions using the image encoding process. The one or more circuits can generate a media package using the plurality of encoded regions and metadata corresponding to the plurality of regions.

[0006] In some embodiments, the one or more circuits can generate the metadata to include / indicate / track / record respective locations of each of the plurality of regions in the image. In some embodiments, the one or more circuits can generate the media package by stitching each of the plurality of encoded regions. In some embodiments, the one or more circuits can generate a header for the media package to include the metadata corresponding to the plurality of regions. In some embodiments, the metadata is provided as exchangeable image file format (EXIF) data in the header of the media package.

[0007] In some embodiments, the media package includes one of a joint photographic experts group (JPEG) file, a portable network graphics (PNG) file, a tagged image file format (TIFF) file, or a WEBP file. In some embodiments, each of the plurality of regions has a different size. In some embodiments, the one or more circuits can encode a first region of the plurality of regions using a first set of encoding parameters. In some embodiments, the one or more circuits can encode a second region of the plurality of regions using a second set of encoding parameters.

[0008] At least one aspect relates to a system. The system can include one or more processors. The system can extract at least a plurality of encoded regions from a media package. The system can generate a plurality of regions of an image by decoding the plurality of encoded regions. The system can generate the image using at least the plurality of regions.

[0009] In some embodiments, the system can apply a filter to the image to eliminate encoding artifacts that can appear on boundaries of ROIs, tiles, etc. In some embodiments, the system can parse a header of the media package to identify metadata corresponding to the plurality of encoded regions. In some embodiments, the system can determine one or more offsets of the plurality of encoded regions in the media package based at least on the metadata corresponding to the plurality of encoded regions. In some embodiments, the media package includes an image file in which each of the plurality of encoded regions is stitched and stored as image data in the image file.

[0010] In some embodiments, the system can decode a first encoded region of the plurality of encoded regions using a first set of decoding parameters. In some embodiments, the system can decode a second encoded region of the plurality of encoded regions using a second set of decoding parameters. In some embodiments, one or more of the first set of decoding parameters and the second set of decoding parameters are stored in the metadata.

[0011] At least one aspect is directed to a method. The method can include obtaining, using one or more processors, a plurality of regions from an image to be encoded. The method can include applying, using the one or more processors, a plurality of levels of compression to the plurality of regions to generate a plurality of encoded regions. The method can include stitching, using the one or more processors, the plurality of encoded regions to generate a media package.

[0012] In some embodiments, the method can include generating, using the one or more processors, metadata corresponding to the plurality of regions, the metadata including a respective location of each region of the plurality of regions within the image. In some embodiments, the media package is generated using at least the metadata corresponding to the plurality of regions.

[0013] The processors, systems, and / or methods described herein can be implemented by or included in at least one of a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twin operations, a system for performing optical transport simulations, a system for performing collaborative content creation of 3D assets, a system for performing deep learning operations, a system for performing generative AI operations using large language models (LLMs), a system for performing generative AI operations using small language models (SLMs), a system for performing one or more conversational AI operations, a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content, a system implemented using edge devices, a system implemented using robots, a system for performing conversational AI operations, a system for generating synthetic data, a system containing one or more virtual machines (VMs), a system implemented at least partially in a data center, or a system implemented at least partially using cloud computing resources. BRIEF DESCRIPTION OF DRAWINGS

[0014] The present systems and methods for adaptive image encoding and decoding for resolution constrained systems are described in detail below with reference to the attached drawing figures, wherein:

[0015] Figure 1 is a block diagram of an example system for encoding and decoding high resolution images in accordance with some embodiments of the present disclosure.

[0016] Figure 2An example dataflow diagram showing high resolution image encoding processing is shown in accordance with some embodiments of the present disclosure;

[0017] Figure 3 An example diagram showing decoding processing of a high resolution image encoded according to the techniques described herein is shown in accordance with some embodiments of the present disclosure;

[0018] Figure 4 A flow diagram of an example method for encoding and decoding high resolution images on a system that supports limited resolution is shown in accordance with some embodiments of the present disclosure;

[0019] Figure 5 A block diagram of an example content streaming system suitable for implementing some embodiments of the present disclosure is shown;

[0020] Figure 6 A block diagram of an example computing device suitable for implementing some embodiments of the present disclosure is shown; and

[0021] Figure 7 A block diagram of an example data center suitable for implementing some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The present disclosure relates to systems and methods for adaptively encoding high resolution media (e.g., images / video) on platforms with limited computing resources. Certain media systems, such as industrial inspection systems or medical imaging systems, produce media with resolutions up to several million pixels. While such images can be encoded and decoded on high performance distributed computing systems with specialized hardware, it is impractical to perform such operations on regular computing devices with limited resources.

[0023] Conventional high resolution media encoding and decoding operations involve processing the entire high resolution image at once. However, conventional computing systems lack the continuous available memory and processing resources to store and process the entire high resolution image. Such processing is particularly challenging on edge devices or computing systems that do not include large amounts of video memory for processing image data. These limitations make it impossible to encode or decode high resolution media on conventional computing systems.

[0024] The systems and methods described herein provide techniques for efficiently and adaptively encoding and decoding high resolution media on computing systems with limited computing resources. To efficiently encode media data, an input high resolution image or high resolution video frame can be partitioned into a set of regions. Each region can include a region of interest (ROI), a tile, or other portion of the image / video to be encoded. Metadata can be generated that indicates the relative position and size of the regions so that the data for each region can be later / subsequently mapped back to its original position and size in the high resolution image / video frame using the metadata.

[0025] The regions extracted from the high resolution media can each have the same size, or can have different sizes. In one example, a high resolution image can be subdivided into four tile regions of equal size. In another example, the size of each region can be determined based on the content depicted by the pixels of the high resolution image / video frame. Once the high resolution media is subdivided into regions, each region can be encoded using a suitable media encoding algorithm (e.g., JPEG, PNG, or JPEG-2000, etc.). The encoded regions can then be combined with the generated metadata into a media package (e.g., an image / video file) that can be stored or transmitted for further processing.

[0026] To decode the encoded media package, the metadata can be extracted from the package and used to enumerate the number, size, and binary offset of each region in the original media data. Using a suitable decoding process, the image data of each region can be individually decoded and then combined to reconstruct the original high resolution image. The metadata extracted from the media package can indicate the mapping of the size and location of each region to the high resolution image. In some embodiments, a filtering process can be performed to reduce potential artifacts between each region (particularly, occurring on tile boundaries) in the reconstructed high resolution image.

[0027] Referring to Figure 1 , Figure 1 is an example computing environment in accordance with some embodiments of the present disclosure, including a system 100 for implementing high resolution image encoding and decoding on systems that support limited resolution. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be wholly omitted or consolidated. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combination and location. The various functions described herein as being performed by the entities can be stored as instructions in memory and executed by a processor. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0028] The system 100 is shown to include an encoder system 102 that can receive a high resolution image 106 and generate media packages 114, and a decoder system 104 that can receive the media packages 114 and generate a decoded high resolution image 130. The encoder system 102 can implement various techniques described herein to encode the high resolution image 106 by extracting one or more regions 110 from the high resolution image 106 and separately encoding each region 110. The encoder system 102 can receive the high resolution image 106 from, for example, one or more computer networks. In some embodiments, the encoder system 102 can access the high resolution image 106 from a data store or storage system. The storage system can be an external server, a distributed storage / computing environment (e.g., a cloud storage system), or any other type of storage device or system in communication with the encoder system 102 and / or the decoder system 104. In some embodiments, the storage system can constitute or be located internal to the encoder system 102 and / or the decoder system 104. In such embodiments, the encoder system 102 can access the high resolution image 106 from internal memory.

[0029] In accordance with the techniques described herein, the encoder system 102 can access, retrieve, or otherwise receive the high resolution image 106 in response to receiving a request to encode the high resolution image 106. The request can be provided by a device external to and in communication with the encoder system 102 (e.g., a client device in communication over a network). In some embodiments, the request can be provided in response to input at the encoder system 102, such as provided by an operator of the encoder system 102. The request can specify the high resolution image 106 to be processed or a location from which the encoder system 102 is to retrieve the high resolution image 106.

[0030] The high-resolution image 106 can include any type of digital image information that includes a large number of pixels. Examples of the high-resolution image 106 can include images with a number of pixels greater than a threshold, including but not limited to 50 megapixels, 64 megapixels, 100 million pixels, 150 million pixels, or 200 million pixels, etc. The number of pixels (e.g., resolution) of the high-resolution image 106 can vary depending on the source of the high-resolution image. For example, an industrial inspection system can generate images with a resolution exceeding hundreds of millions of pixels, enabling effective capture of minute details of manufactured parts or machine surfaces. The high-resolution image 106 can be an image generated by one or more medical imaging systems, such as systems for high-resolution computed tomography (CT) or magnetic resonance imaging (MRI), which can have similarly high resolutions to accurately render internal anatomical structures. The high-resolution image 106 can be an image generated by other systems, such as surveillance cameras for city infrastructure or building security, which capture detailed visual data for monitoring and security purposes. These systems employ advanced imaging technology to produce high-quality images and videos for real-time viewing and post-event analysis. Furthermore, high-resolution imaging can be used in various fields, including satellite imagery for geographic mapping or environmental monitoring.

[0031] In some embodiments, the high-resolution image 106 can represent a single frame extracted from a high-resolution video sequence. Such high-resolution videos can be provided by surveillance systems or entertainment / communication systems, which can transmit or otherwise process / access high-resolution video data. In such embodiments, the encoder system 102 can process the high-resolution image 106 data of each frame individually according to the techniques described herein. The high-resolution image 106 can be stored, provided, or otherwise accessed in any number of image formats, including but not limited to raw image data (e.g., bitmap (BMP) images), portable network graphics (PNG) images, tagged image file format (TIFF) images, or other lossless image formats.

[0032] To encode the high-resolution image 106, the encoder system 102 can execute a region extractor 108 to extract (e.g., determine, obtain, isolate, retrieve, identify, etc.) one or more regions 110 of the high-resolution image 106 for individual encoding by the encoder 112. In one example, the region extractor 108 can extract the regions 110 from the high-resolution image 106, for example, by subdividing the high-resolution image 106 into a plurality of tiles. In one example, the tiles can be equal in size and can form a grid structure across the high-resolution image 106. The number and size of the regions 110 can be determined as a configurable parameter of the encoding process, which can be specified in a request to encode the high-resolution image 106 and / or can be specified in configuration settings of the encoder system 102.

[0033] In some embodiments, the region extractor 108 can extract regions 110 that are not all of the same size. For example, the region extractor 108 can extract regions as non-uniform partitions of the high-resolution image 106 that can be identified based on configuration settings of the encoder system 102, based on a request to encode the high-resolution image 106, or based on content of the high-resolution image 106 or other factors (e.g., pre-defined templates, external input parameters, or real-time analysis of the image). These additional factors can include contextual information, user preferences, or specific algorithmic criteria, including how to segment regions for extraction. In some embodiments, the size and shape of each region 110 can be dynamically determined according to the content of the high-resolution image 106. The regions 110 extracted by the region extractor 108 can contain contiguous areas of pixels. In some embodiments, the regions 110 can be square or rectangular regions of the high-resolution image 106. The regions 110 can be extracted such that each pixel is not contained in more than one region 110, and such that all regions 110 can be combined to reconstitute the high-resolution image 106.

[0034] In some embodiments, the region extractor 108 can implement one or more image processing techniques, including but not limited to image classification machine learning models, object / feature detection models, or image segmentation models, to dynamically determine the size and shape of different regions 110 in the high-resolution image 106. For example, an object detection model can be used to identify objects or features of interest in the high-resolution image 106. The region extractor 108 can then extract regions 110 that contain these detected objects or features, thereby minimizing the splitting of any single object or feature across multiple regions 110. In some embodiments, the high-resolution image 106 can include data (e.g., segmentations, bounding boxes, classifications, etc.) indicating objects / features of interest present in the high-resolution image 106.

[0035] In some embodiments, the region extractor 108 can receive dimensions / sizes of regions 110 to extract from the high-resolution image 106 from one or more external computing systems. For example, the external computing systems can be systems that perform downstream processing using the encoded high-resolution image. Depending on the performance / quality of the downstream processing or specific needs of the application, the external computing systems can provide feedback indicating the size / dimensions of different regions 110, or update to configuration settings to increase, decrease, or otherwise modify the number or size of regions 110 to extract from the high-resolution image 106. In examples where the high-resolution image 106 is a frame of a video stream, processing of previous frames (e.g., decoding and use in downstream processing tasks) can provide feedback for subsequent frames of the video stream to be encoded by the encoder system 102.

[0036] Once extracted, the encoder system 102 can execute the encoder 112 to encode the individual regions 110 into a set of encoded regions 118. The encoder 112 can be executed to encode each region 110 sequentially, or in parallel in certain embodiments. To encode each region 110, the encoder 112 can execute any suitable image encoding algorithm. Non-limiting examples of such encoding algorithms include, but are not limited to, Joint Photographic Experts Group (JPEG) encoding, JPEG-2000 encoding, PNG encoding, Graphics Interchange Format (GIF) encoding, or WEBP encoding, among others. The regions 110 can be encoded into a compressed format suitable for storage or transmission in the media packets 114 using any suitable lossy or lossless encoding algorithm.

[0037] In some embodiments, the encoder 112 can encode the regions 110 according to one or more quality parameters. These quality parameters can include a compression level or quality level applied to each region 110 during encoding. Using different quality or compression levels can optimize storage and transmission efficiency while maintaining visual fidelity in regions 110 that present the most relevant visual information. In one example, regions 110 containing high frequency details (e.g., edges, textures, or sharp transitions) can be encoded using a higher quality setting (e.g., a lower quantization parameter (QP) in JPEG encoding), reducing compression artifacts and better preserving fine details.

[0038] In another example, regions 110 presenting smoother content (e.g., uniform backgrounds or regions with gradual color changes) can be encoded using a lower quality setting (e.g., a higher QP). This approach allows for greater compression rates without significantly impacting the perceived visual quality of these regions. In some embodiments, the encoder 112 can determine the quality parameters for one or more regions 110 based on a predefined rule set, based on content-based analysis (e.g., edge detection, texture classification), or specified preferences (e.g., in a request to process the high resolution image 106, in configuration settings of the encoder system 102, etc.). Determining encoding quality on a per-region 110 basis enables the encoder to minimize the size of the encoded high resolution image (e.g., the media packets 114) while maintaining visual fidelity upon decoding.

[0039] Once the regions 110 are encoded, the encoder 112 can generate metadata 116 associated with each encoded region 118. The metadata 116 can contain the following information: the relative position of the corresponding region 110 within the original high-resolution image 106 (e.g., row and column coordinates), the pixel size of the region 110, a unique identifier assigned to each region 110, and other metadata. The metadata 116 can also specify other relevant attributes, such as the encoding algorithm used to generate a particular encoded region 118, the encoding / compression parameters employed, or any other information related to the encoding process.

[0040] The generated metadata 116 is combined with the encoded regions 118 to form the media package 114. In one example, the encoder system 102 can generate the media package 114 as an image file by stitching the individual encoded regions 118 as part of the image file image data and including the metadata 116 corresponding to the encoded regions 118 in one or more headers of the image file. In the example using JPEG encoding, the media package 114 can be generated as a JPEG file, with the metadata 116 stored as part of the Exchangeable Image File Format (EXIF) header data embedded in the JPEG file. Similar approaches can be used for different formats of the media package 114. In some embodiments, the encoder system 102 can further compress the media package 114, for example, to reduce its size for transmission via one or more networks. Non-limiting examples of compression include ZIP compression, GZIP compression, bzip2 compression, or LZMA compression, among others.

[0041] The media package 114 can contain an identifier of the high-resolution image 106 from which the media package 114 was generated. In some embodiments, the media package 114 can be stored and / or provided to other downstream processing systems or processes. In the present example, the media package 114 is shown as being provided to a decoder system 104. In some embodiments, the media package 114 can be stored in one or more repositories, from which it can be subsequently retrieved for processing by the decoder system 104.

[0042] The decoder system 104 can be any type of computing system for decoding the media package 114 to generate a decoded high-resolution image 130. To generate the decoded high-resolution image 130, the decoder system 104 can execute a media parser 120. The media parser 120 can include hardware, software, or a combination of hardware and software. The media parser 120 can access the media package 114 generated from the high-resolution image 106 and parse the metadata 116 contained therein. Parsing the metadata 116 can include decompressing the media package 114, identifying header information stored in the media package 114, and parsing the header information to extract the metadata 116.

[0043] As described herein, the metadata 116 can include the relative location, size, identifier, and / or other attributes of the encoded regions 118. Using the metadata 116, the media parser 120 can parse the image data of the media package 114 to extract each encoded region 118 stored therein. For example, in some embodiments, the media parser 120 can use the location data, size data, or dimension data to identify an offset in the media package 114 where each encoded region 118 is stored. This offset information can then be used to extract the binary data representing each encoded region 118. In some embodiments, the media parser 120 can parse all of the encoded regions 118 in the media package 114 before decoding all of the encoded regions 118 in the media package 114. In some embodiments, the media parser 120 can sequentially provide one or more encoded regions 118 to the decoder 122 as input to be decoded sequentially before parsing the next encoded region 118 in the media package 114.

[0044] The encoded regions 118 parsed by the media parser 120 can be provided as input to the decoder 122 for decoding. The decoder 122 can include hardware, software, or a combination of hardware and software. The decoder 122 can decode each encoded region 118 by reversing the operations applied to generate the encoded region 118. In some embodiments, the decoder 122 can access the metadata 116 of the media package 114 to identify the encoding parameters used to encode the corresponding encoded region 118. The encoded regions 118 can be decoded using any suitable decoding process to reconstruct the pixel data of the encoded region 118 to produce an output decoded pixel data region. The decoder 122 can repeat this process for each encoded region 118 to generate a set of decoded regions that can be used to generate the decoded high-resolution image 130.

[0045] To generate the decoded high-resolution image 130, the media generator 124 can access the metadata 116 of the media package 114 to identify or otherwise determine the location of each decoded region in the original high-resolution image 106 used to generate the media package 114. For example, the metadata 116 can specify the row and column coordinates, size, or other identifier of each encoded region 118 relative to its location in the original high-resolution image 106. In some embodiments, the identifier of each encoded region 118 can encode its relative location in the original high-resolution image 106. Once the location of each region is determined, the media generator 124 can use the decoded regions to generate the decoded high-resolution image 130.

[0046] To this end, the media generator 124 can assemble the decoded regions into the correct locations in the decoded high-resolution image 130, which can be stored in one or more memory regions in the decoder system 104. The media generator 124 can map each decoded region to a corresponding location in the original high-resolution image 106 based on the location / size information derived from the metadata 116. The decoded high-resolution image 130 can include pixel data for all decoded regions mapped to their respective locations such that the decoded high-resolution image 130 is similar to the high-resolution image 106.

[0047] In some embodiments, generating the decoded high-resolution image 130 can result in one or more artifacts at the boundaries between decoded regions used to generate the decoded high-resolution image 130. These artifacts can manifest as visible discontinuities, color bands, or blurring along the edges of different encoded region combinations. To address any artifacts, the media generator 124 can apply one or more filtering operations, such as a tiling filter or a deblocking filter, in portions of the decoded high-resolution image corresponding to the edges of the decoded regions. Applying filtering operations can smooth the transitions between adjacent decoded regions and can minimize visual discontinuities.

[0048] The tiling filter applied by the media generator 124 can include a smoothing or blending technique applied to pixel values near the boundaries of two or more decoded regions, which can reduce visible block edges. The deblocking filter can be used to address various types of artifacts that can arise when generating the decoded high-resolution image 130, including quantization errors introduced during encoding. In some embodiments, the media generator 124 can perform an edge detection algorithm to identify sharp transitions between decoded regions and then apply a local filtering operation to refine the detected boundaries containing discontinuities.

[0049] The decoded high-resolution image 130 generated by the decoder system 104 can be stored in the decoder system’s memory or provided to one or more computing systems or processes for further processing. For example, the decoder system 104 can be implemented as part of a computing system that processes image data using one or more machine learning techniques or other image processing techniques. In such embodiments, the decoded high-resolution image 130 can be provided as input to one or more machine learning models or image processing algorithms. In some embodiments, the high-resolution image can be rendered or displayed through one or more display devices.

[0050] Referring to Figure 2 In conjunction with Figure 1Within the context of the components described herein, an example data flow diagram 200 is shown according to some embodiments of this disclosure, illustrating encoding processing for a high-resolution image. As shown, this encoding processing can be used to encode a high-resolution image 202 (e.g., high-resolution image 106). Region extraction processing 204 can be applied to the high-resolution image to extract one or more pixel data regions (e.g., region 110), which can be individually encoded using image encoding operation 206.

[0051] As shown in the figure, image encoding operations can be used to encode each region individually. Although shown in the figure as parallel execution, it should be understood that in some implementations, at least some encoding operations can be performed sequentially (e.g., processing one or more regions at a time). Encoding regions may include performing... Figure 1 The encoder 112 operates to generate an encoded region (e.g., encoded region 118) corresponding to the region generated using the region extraction process 204. The image combining and header generation process 208 can then be used to generate a stitched encoded image 210 (e.g., media package 114).

[0052] Image combining and header generation process 208 can generate metadata and store it as header information in the stitched encoded image 210. As described herein, the metadata can indicate the size, location, dimensions, and / or identifier of each encoded region generated from the high-resolution image 202. Image combining and header generation process 208 can generate the stitched encoded image by stitching each encoded region into a single file or data structure. In some embodiments, image combining and header generation process 208 can compress the stitched encoded image using a suitable encoding algorithm.

[0053] Reference Figure 3 In combination Figure 1 and Figure 2 Within the context of the components described herein, example figure 300 is shown according to some embodiments of the present disclosure, illustrating decoding processing of a high-resolution image encoded according to the techniques described herein. Decoding processing can be used to decode a stitched encoded image 302 (e.g., a stitched encoded image 210, media package 114, etc.). To decode the stitched encoded image 302, header parsing and tile extraction processing 304 can be performed, which can use metadata contained in the stitched encoded image 302 to identify each encoded tile (e.g., encoded region 118) within the stitched encoded image 302. As described herein, metadata can specify or provide information that can be used to derive the location (e.g., binary position, binary offset, etc.) of each encoded tile within the stitched encoded image 302.

[0054] As shown, the encoded tiles / regions extracted from the stitched encoded image 302 can be decoded using corresponding image decoding operations 306. Although shown as being performed in parallel, it should be understood that, in some embodiments, at least some of the decoding operations can be performed sequentially (e.g., processing one or more regions at a time). Decoding the regions can include performing inverse operations of the encoder 112 operations to generate corresponding decoded regions, which can be used to reassemble the high resolution image 310. Figure 1

[0055] The decoded regions are provided as input to an image tiling and artifact filtering process 308, which is used to assemble each of the decoded regions / tiles into the high resolution image 310 according to the metadata. In some embodiments, filtering can be performed to reduce artifacts that can be generated as a result of separately encoding each region to generate the stitched encoded image 302. The high resolution image 310, once generated, can be provided to one or more computing systems for further processing.

[0056] Figure 4 A flowchart of a method 400 for implementing high resolution image encoding and decoding on systems that support limited resolutions in accordance with certain embodiments of the present disclosure. The various operations of the method 400 can be implemented by the same or different devices or entities at different points in time. For example, one or more first devices can implement operations related to encoding a high resolution image to generate a media package, and one or more second devices can implement operations related to decoding the media package to reconstruct the high resolution image.

[0057] Each block of the method 400 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For instance, various functions can be carried out by a processor executing instructions stored in memory. The method 400 can also be embodied as computer-usable instructions stored on computer storage media. The method 400 can be provided by a standalone application, a service or hosted service (standalone or in combination with other hosted services), a plug-in to another product, or some other form of computational implementation. Further, the methods described herein can be implemented in Figure 1 The method 400 is described by way of example with respect to the systems in Figure 2 However, the method 400 can additionally or alternatively be performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0058] ​The method 400 includes, at block B402, identifying an image (e.g., the high resolution image 106) to be encoded using the image encoding process. Identifying the image can include retrieving the image from one or more data stores, receiving the image from an external computing system (e.g., in a request to encode the image), or receiving the image by another application or process (e.g., through inter-process communication). In some embodiments, the resolution of the image can be determined based on header information or content data of the image. To further illustrate this example, if the image resolution exceeds a threshold (e.g., exceeds the processing capabilities of the computing system to encode the image), the method can continue to perform step B404. In some embodiments, the method can continue to perform step B404 regardless of whether the image resolution exceeds a threshold.

[0059] The method 400 includes, at step B404, extracting a plurality of regions (e.g., the regions 110) from the image. The regions can be extracted to subdivide the image into a plurality of tiles that can be individually encoded according to the techniques described herein. Extracting the regions can include performing any of the operations described in relation to the region extractor 108 of Figure 1 In some embodiments, the size of each region can be the same. In some embodiments, the number of regions can depend on the resolution of the image. In some embodiments, the size of the regions can be different. The parameters of the regions (e.g., size, location, number, etc.) can be provided as a configuration setting of the encoding process. The parameters can be stored in a configuration file, provided in a request to encode the image, or provided through operator input. Once the parameters of the regions are determined, the regions can be extracted by extracting the pixel data of each region from the image.

[0060] The method 400 includes, at block B406, applying a plurality of levels of compression to generate a plurality of encoded regions (e.g., the encoded regions 118). The levels of compression can be applied by encoding each of the plurality of regions using the image encoding process. The regions of pixel data extracted in step B406 can be encoded using any suitable encoding process, including but not limited to JPEG encoding, JPEG-2000 encoding, PNG encoding, TIFF encoding, GIF encoding, or WEBP encoding, etc. In some embodiments, the encoding parameters of each region can be determined according to the image content. Regions with a higher level of detail (e.g., more number of images, objects / features of interest, etc.) can be encoded using higher quality parameters than regions with a relatively lower level of detail (e.g., smoother background, fewer edges / transitions between colors or features, etc.). In some embodiments, one or more regions can be encoded in parallel. In some embodiments, one or more regions can be encoded sequentially. Encoding the regions can include performing any of the operations of the encoder 112 of Figure 1 ​

[0061] Method 400 at block B408 includes: concatenating multiple encoded regions to generate a media packet. The generated media packet can contain metadata (e.g., metadata 116) corresponding to the multiple regions. The media packet can be generated by concatenating each encoded region and storing the concatenated regions as part of the image data in the media packet. In one example, the media packet can be an image file (e.g., a JPEG file, a JPEG-2000 file, a TIFF file, a PNG file, a WEBP file, etc.). The media packet can contain metadata indicating the corresponding location of each region extracted from the image and other parameters of each region (e.g., size, dimensions, identifier, etc.). The metadata can be stored at least partially in the header of the media packet (e.g., EXIF ​​data of the JPEG file, other header fields, etc.). In some embodiments, the metadata can specify encoding parameters used to encode each region. Once generated, the media packet can be provided or stored for further processing by a computational system (e.g., decoder system 104, etc.) that decodes the media packet.

[0062] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models (e.g., one or more large language models (LLMs) and / or one or more small language models (SLMs)), systems for performing one or more conversational AI operations, systems for presenting at least one of virtual reality content, augmented reality content, or mixed reality content, systems for performing optical transmission simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other system types.

[0063] Example content streaming system

[0064] See now Figure 5 , Figure 5 This is an example system diagram of a content streaming system 500 according to some embodiments of the present disclosure. Figure 5 Includes one or more application servers 502 (which may include...) Figure 7 Example computing device 700 (similar components, features and / or functions), one or more client devices 504 (which may include similar components, features and / or functions to the example computing device 700), and one or more client devices 504.Figure 7 similar components, features, and / or functions of the example computing device 700 described herein) and one or more networks 506 (which can be similar to one or more networks described herein). In some embodiments of the present disclosure, the system 500 can be implemented to generate audio-driven facial animations with varying identities and speaking styles, including techniques to train / update various machine learning models described herein. The application session can correspond to a game streaming application (e.g., NVIDIA GeForce NOW), a remote desktop application, a simulation application (e.g., autonomous or semi-autonomous vehicle simulation), a computer-aided design (CAD) application, a virtual reality (VR) and / or augmented reality (AR) streaming application, a deep learning application, and / or other application types. For example, the system 500 can be implemented to receive input indicating one or more features of an output generated using a neural network model, provide the input to the model to cause the model to generate the output, and use the output for various operations including display or simulation operations.

[0065] In the system 500, for an application session, the one or more client devices 504 can receive input data in response to input to the one or more input devices 526, send the input data to the one or more application servers 502, receive encoded display data from the one or more application servers 502, and display the display data on the display 524. Thus, computationally more intensive computations and processing are offloaded to the one or more application servers 502 (e.g., rendering of graphical output for the application session, particularly ray or path tracing, performed by one or more GPUs of the one or more application servers 502). In other words, the application session is streamed from the one or more application servers 502 to the one or more client devices 504, thereby reducing the requirements of the one or more client devices 504 for graphics processing and rendering.

[0066] For example, with respect to instantiation of an application session, the client device 504 can display a frame of the application session on the display 524 based at least on receiving display data from the one or more application servers 502. The client device 504 can receive input to one of the one or more input devices and, in response, generate input data. The client device 504 can send the input data to the one or more application servers 502 via the communication interface 520 and over the network 506 (e.g., the Internet), and the one or more application servers 502 can receive the input data via the communication interface 518. The one or more CPUs 510 can receive the input data, process the input data, and send data to the one or more GPUs 510 that causes the one or more GPUs 510 to generate a rendering of the application session. For example, the input data can represent movement of a user’s character in a game session of a game application, firing a weapon, reloading, passing a ball, turning a vehicle, etc. The rendering component 512 can render the application session (e.g., representing the results of the input data), and the rendering capture component 514 can capture the rendering of the application session as display data (e.g., as image data that captures rendered frames of the application session). The rendering of the application session can include lighting and / or shadow effects computed using ray or path tracing of the one or more application servers 502 one or more parallel processing units such as GPUs, which can further employ the use of one or more specialized hardware accelerators or processing cores to perform ray or path tracing techniques. In some embodiments, one or more virtual machines (VMs) — e.g., including one or more virtual components such as vGPUs, vCPUs, etc. — can be used by the one or more application servers 502 to support the application session. The encoder 516 can then encode the display data to generate encoded display data, and the encoded display data can be sent to the client device 504 via the communication interface 518 over the network 506. The client device 504 can receive the encoded display data via the communication interface 520, and the decoder 522 can decode the encoded display data to generate the display data. The client device 504 can then display the display data via the display 524.

[0067] Example computing device

[0068] Figure 6is a block diagram of an example computing device 600 suitable for implementing some embodiments of the present disclosure. The computing device 600 can include an interconnection system 602 coupling the following components: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., one or more displays), and one or more logic units 620. In at least one embodiment, one or more computing devices 600 can include one or more virtual machines (VMs), and / or any component thereof can include a virtual component (e.g., a virtual hardware component). For a non-limiting example, one or more of GPUs 608 can include one or more vGPUs, one or more of CPUs 606 can include one or more vCPUs, and / or one or more of logic units 620 can include one or more virtual logic units. As such, one or more computing devices 600 can include discrete components (e.g., a full GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof.

[0069] Although Figure 6 various blocks of are shown as being connected to a line with the interconnection system 602, this is not intended to be limiting, and is for clarity only. For example, in some embodiments, a presentation component 618 such as a display device can be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, a CPU 606 and / or GPU 608 can include memory (e.g., in addition to the memory of GPU 608, the memory 604 can represent a storage device). In other words, Figure 6 computing devices of are merely illustrative. There is no distinction, in terms of Figure 6 scope, between "workstation" "server" "laptop" "desktop" "tablet" "client device" "mobile device" "hand-held device" "game console" "electronic control unit (ECU)" "virtual reality system" and other device or system types as the scope of

[0070] The interconnection system 602 can represent one or more links or buses linking the various components, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 602 can be arranged in various topologies including, but not limited to, bus, star, ring, mesh, tree, or hybrid topologies. The interconnection system 602 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. For example, the CPU 606 can be directly connected to the memory 604. Further, the CPU 606 can be directly connected to the GPU 608. Where there are direct connections or point-to-point connections between components, the interconnection system 602 can include a PCIe link to perform the connection. In these examples, the PCI bus need not be included in the computing device 600.

[0071] The memory 604 can include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 600. Computer-readable media can include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media can comprise computer storage media and communication media.

[0072] Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, the memory 604 can store computer readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 600. Computer storage media, as used herein, does not include signals per se.

[0073] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of the any of the above should also be included within the scope of computer readable media.

[0074] The CPU(s) 606 can be configured to execute computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. The CPU(s) 606 can each include one or more cores capable of handling multiple software threads concurrently (e.g., 1, 2, 4, 8, 28, 72, etc.). The CPU(s) 606 can include any type of processors, and can include different types of processors depending on the type of computing device 600 being implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 600, the processor can be an Advanced RISC Machines (ARM) processor implemented using a Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 600 can include the CPU(s) 606 in addition to one or more microprocessors or supplemental co-processors such as math co-processors.

[0075] In addition or alternative to CPU 606, one or more GPUs 608 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. One or more of GPUs 608 can be integrated GPUs (e.g., with one or more of CPUs 606) and / or one or more of GPUs 608 can be discrete GPUs. In embodiments, one or more of GPUs 608 can be a co-processor of one or more of CPUs 606. GPUs 608 can be used by computing device 600 to render graphics (e.g., 3D graphics) or to perform general purpose computing. For example, GPUs 608 can be used for general purpose computing on GPUs (GPGPU). GPUs 608 can include hundreds or thousands of cores capable of handling hundreds or thousands of software threads concurrently. GPUs 608 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from CPUs 606 received via a host interface). GPUs 608 can include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory can be included as part of memory 604. GPUs 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using an NVSwitch). When combined together, each of GPUs 608 can generate pixel data or GPGPU data for different portions of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0076] In addition to or in place of CPU(s) 606 and / or GPU(s) 608, one or more logic units 620 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. In embodiments, CPU(s) 606, GPU(s) 608, and / or logic unit(s) 620 can perform any combination of methods, processes, and / or portions thereof, discretely or jointly. One or more of logic units 620 can be part of and / or integrated within one or more of CPU(s) 606 and / or GPU(s) 608, and / or one or more of logic units 620 can be discrete components or otherwise external to CPU(s) 606 and / or GPU(s) 608. In embodiments, one or more of logic units 620 can be a co-processor of one or more of CPU(s) 606 and / or one or more of GPU(s) 608.

[0077] Examples of logic units 620 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel visual core (PVC), a visual processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multi-processor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating-point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) element, and the like.

[0078] The communication interface 610 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 610 can include components and functionality to enable communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 620 and / or the communication interface 610 can include one or more data processing units (DPUs) to transfer data received over a network and / or over the interconnect system 602 directly to the one or more GPUs 608 (e.g., memory of the one or more GPUs 608). In some embodiments, multiple computing devices 600 or components thereof (which can be similar or different from one another in various aspects) can be communicatively coupled to send and receive data to perform various operations described herein, for example, to facilitate a reduction in latency.

[0079] The I / O ports 612 can enable the computing device 600 to be logically coupled to other devices including I / O components 614, presentation components 618, and / or other components, some of which can be built-in (e.g., integrated) to the computing device 600. Illustrative I / O components 614 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 614 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some instances, inputs can be transmitted to an appropriate network element for further processing, such as modification and registration of images. A NUI can implement any combination of speech recognition, gesture recognition, facial recognition, biometric recognition, posture recognition, gesture recognition within, and on, a touchscreen, air gestures, head and eye tracking, and touch recognition associated with a display of the computing device 600, as described in more detail below. The computing device 600 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. Additionally, the computing device 600 can include an accelerometer or a gyroscope (e.g., as part of an inertial measurement unit (IMU)) to detect motion. In some examples, the computing device 600 can use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.

[0080] Power supply 616 can include a hard-wired power supply, a battery power supply, or a combination thereof. Power supply 616 can provide power to computing device 600 to enable operation of components of computing device 600.

[0081] One or more presentation components 618 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Presentation components 618 can receive data from other components (e.g., GPU 608, CPU 606, DPU, etc.) and output the data (e.g., as images, video, sound, etc.).

[0082] Example data center

[0083] Figure 7 An example data center 700 that can be used in at least one embodiment of the present disclosure, such as to implement system 100, is shown, e.g., in conjunction with Figure 2 The operations described can be performed in one or more examples of data center 700. Data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.

[0084] As Figure 7 shown, data center infrastructure layer 710 can include resource orchestrator 712, grouped computing resources 714, and node computing resources (“node C.R.s”) 716(1)-716(N), where “N” represents any integer, positive integer. In at least one embodiment, node C.R.s 716(1)-716(N) can include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and / or cooling modules, etc. In some embodiments, one or more of node C.R.s 716(1)-716(N) can correspond to a server having one or more of the above-described computing resources. Further, in some embodiments, node C.R.s 716(1)-716(N) can include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of node C.R.s 716(1)-716(N) can correspond to a virtual machine (VM).

[0085] In at least one embodiment, the grouped computing resources 714 may include separate groups of nodes CR716 housed within one or more racks (not shown) or within a plurality of racks in data centers (also not shown) located in different geographical locations. The separate groups of nodes CR716 within the grouped computing resources 714 may include the group's computing resources, network resources, memory resources, or storage resources, which may be configured or allocated to support one or more workloads. In at least one embodiment, a plurality of nodes CR716, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0086] Resource coordinator 712 may be configured or otherwise control one or more nodes CR716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may include a Software Design Infrastructure (“SDI”) management entity for data center 700. Resource coordinator 712 may include hardware, software, or some combination thereof.

[0087] In at least one embodiment, such as Figure 7 As shown, framework layer 720 may include a job scheduler 728, a configuration manager 734, a resource manager 736, and / or a distributed file system 738. Framework layer 720 may include a framework for software 732 supporting software layer 730 and / or one or more applications 742 of application layer 740. Software 732 or application 742 may respectively include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 720 may be, but is not limited to, a type of free and open-source software web application framework that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"), such as Apache Spark. TM(hereinafter “Spark”). In at least one embodiment, job scheduler 728 can include a Spark driver to facilitate scheduling of workloads supported by layers of data center 700. Configuration manager 734 can be capable of configuring different layers, such as software layer 730 and framework layer 720 including Spark and distributed file system 738 for supporting large scale data processing. Resource manager 736 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting distributed file system 738 and job scheduler 728. In at least one embodiment, clustered or grouped computing resources can include grouped computing resources 714 at data center infrastructure layer 710. Resource manager 736 can coordinate with resource orchestrator 712 to manage these mapped or allocated computing resources.

[0088] In at least one embodiment, software 732 included in software layer 730 can include software used by at least portions of node C.R.s 716(1)-716(N), grouped computing resources 714, and / or distributed file system 738 of framework layer 720. One or more types of software can include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.

[0089] In at least one embodiment, applications 742 included in application layer 740 can include one or more types of application programs used by at least portions of node C.R.s 716(1)-716(N), grouped computing resources 714, and / or distributed file system 738 of framework layer 720. One or more types of application programs can include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications (including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0090] In at least one embodiment, any of configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modifying actions based at least on any number and type of data acquired in any technically feasible manner. Self-modifying actions can free data center 700’s data center operator from making possibly poor configuration decisions and can avoid underutilization and / or poor performance of portions of data center.

[0091] According to one or more embodiments described herein, data center 700 may include tools, services, software, or other resources for updating / training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, one or more machine learning models may be updated / trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 700. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 700 by using weight parameters calculated through one or more training techniques (such as, but not limited to, those described herein).

[0092] In at least one embodiment, the data center 700 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured to allow users to update / train or perform information inference services, such as image recognition, speech recognition, or other artificial intelligence services.

[0093] Example network environment

[0094] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network attached storage (NAS), other back-end devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 6 The implementation is carried out on one or more instances of computing device 600, for example, each device may include similar components, features and / or functions of computing device 600. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 700, the examples of which in this document are relative to... Figure 7 To describe in more detail.

[0095] Components of a network environment can communicate with each other via one or more networks, which may be wired, wireless, or both. A network can include multiple networks or a network of networks. For example, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0096] A compatible network environment can include one or more peer-to-peer network environments - in which case no servers can be included in the network environment - and one or more client-server network environments - in which case one or more servers can be included in the network environment. In a peer-to-peer network environment, functionality described herein with respect to one or more servers can be implemented on any number of client devices.

[0097] In at least one embodiment, the network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports a software layer and / or one or more applications of an application layer. The software or applications can include web-based service software or applications, respectively. In an embodiment, one or more of the client devices can use the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, without limitation, a type of free and open-source software web application framework such as Apache® that can use the distributed file system for large-scale data processing (e.g., “big data”).

[0098] The cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functionality described herein (or one or more portions thereof). Any of these different functionalities can be distributed across multiple locations from central or core servers (e.g., one or more data centers that can be distributed across states, regions, countries, globally, etc.). The core servers can designate at least a portion of the functionality to edge servers if the connection to the user (e.g., client device) is relatively close to the edge servers. The cloud-based network environment can be private (e.g., limited to a single organization), can be public (e.g., available to many organizations), and / or combinations thereof (e.g., a hybrid cloud environment).

[0099] One or more client devices can include those described herein with respect to Figure 6At least some of the components, features and functionalities of the depicted one or more example computing devices 600. By way of example, and not limitation, a client device can be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a watercraft, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these depicted devices, or any other suitable device.

[0100] The present disclosure can be described in the general context of computer code or machine-useable instructions, including computer-executable instructions such as program modules, being executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules including routines, programs, objects, components, data structures, and the like, refer to code that performs particular tasks or implements particular abstract data types. The present disclosure can be practiced in a variety of system configurations, including hand-held devices, consumer electronics, general- purpose computers, more specialty computing devices, and the like. The present disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network.

[0101] As used herein, recitations of “and / or” with respect to two or more elements shall be interpreted to mean any element alone, or in combination with any of the other elements. For example, “element A, element B, and / or element C” can include just element A, just element B, just element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0102] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have contemplated that the claimed subject matter can also be embodied in other ways, including different steps or combinations of steps in a different order, to include in the claims the different steps or combinations of steps disclosed in this document, without departing from the spirit and essential characteristics of the subject matter disclosed herein. Additionally, it is contemplated that individual steps or combinations of steps of the claimed subject matter can be practiced by one or more entities and / or over one or more locations. Moreover, it is contemplated that individual steps or combinations of steps of the claimed subject matter can be practiced by one or more entities and / or over one or more locations.

Claims

1. One or more processors, said one or more processors comprising: One or more circuits are used for: Image encoding processing is used to extract multiple regions from the image to be encoded; Multiple encoded regions are generated by encoding each of the multiple regions using the image encoding process. as well as The media package is generated using the plurality of encoded regions and the metadata corresponding to the plurality of regions.

2. The processors according to claim 1, wherein, The one or more circuits are used for: The metadata is generated to include the corresponding location of each of the plurality of regions within the image.

3. The processors according to claim 1, wherein, The one or more circuits are used for: The media package is generated by concatenating each of the plurality of encoded regions.

4. The processors according to claim 1, wherein, The one or more circuits are used for: Generate a header for the media package to include the metadata corresponding to the plurality of regions.

5. The processors according to claim 4, wherein, The metadata is provided as EXIF ​​data in the interchangeable image file format in the header of the media package.

6. The processors according to claim 1, wherein, The media package includes one of the following: a Joint Image Experts Group JPEG file, a Portable Web Graphics PNG file, a Tagged Image File Format TIFF file, or a WEBP file.

7. The processors according to claim 1, wherein, Each of the multiple regions has a different size.

8. The processors according to claim 1, wherein, The one or more circuits are used for: The first region among the plurality of regions is encoded using the first set of encoding parameters; as well as The second region among the plurality of regions is encoded using the second set of encoding parameters.

9. The processors according to claim 1, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system used to perform deep learning operations; Systems implemented using edge devices; Systems implemented using robots; Systems used to perform conversational AI operations; A system for performing generative AI operations using large language model LLM; A system for performing generative AI operations using a small language model (SLM); A system for performing one or more conversational AI operations; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

10. A system comprising: One or more processors are used for: Extract at least several encoded regions from the media package; Multiple regions of the image are generated by decoding the multiple encoded regions; as well as The image is generated using at least the plurality of regions.

11. The system according to claim 10, wherein, The one or more processors are used for: A filter is applied to the image to remove coded artifacts.

12. The system according to claim 10, wherein, The one or more processors are used for: The header of the media packet is parsed to identify metadata corresponding to the plurality of encoded regions.

13. The system according to claim 10, wherein, The one or more processors are used for: Based on the metadata corresponding to the plurality of encoded regions, one or more offsets of the plurality of encoded regions in the media package are determined.

14. The system according to claim 10, wherein, The media package includes an image file, wherein each of the plurality of encoded regions is stitched together and stored as image data in the image file.

15. The system according to claim 10, wherein, The one or more processors are used for: Decode the first encoded region among the plurality of encoded regions using the first set of decoding parameters; and The second set of decoding parameters is used to decode the second encoded region among the plurality of encoded regions.

16. The system according to claim 15, wherein, One or more of the first set of decoding parameters and the second set of decoding parameters are stored in the metadata.

17. The system according to claim 10, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system used to perform deep learning operations; Systems implemented using edge devices; Systems implemented using robots; Systems used to perform conversational AI operations; A system for performing generative AI operations using large language model LLM; A system for performing generative AI operations using a small language model (SLM); A system for performing one or more conversational AI operations; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

18. A method, the method comprising: Use one or more processors to extract multiple regions from the image to be encoded; Multiple levels of compression are applied to the multiple regions using one or more processors to generate multiple encoded regions; as well as The plurality of encoded regions are spliced ​​together using one or more processors to generate a media package.

19. The method of claim 18, further comprising: Metadata corresponding to the plurality of regions is generated using one or more processors, the metadata including the corresponding location of each of the plurality of regions within the image.

20. The method according to claim 19, wherein, The media package is generated using at least the metadata corresponding to the plurality of regions.