Prediction using compression network

By processing input values ​​and conditional inputs through the encoder and decoder parts of the compression network, encoding data is generated and prediction values ​​are generated, which solves the problem of resource waste in portable computing devices and achieves more efficient data processing and transmission.

CN120752914APending Publication Date: 2025-10-03QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480016954.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-14
Filing Date
2024-03-04
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies have the problem of wasting resources (such as memory and bandwidth) when processing large amounts of data, especially in portable computing devices, especially when generating and transmitting image frames.

Method used

A compression network is used for data compression and prediction. The input values ​​and conditional inputs are processed through the encoder and decoder parts to generate encoded data and predicted values. The information at the decoder is used to estimate and reduce transmission and storage resources.

Benefits of technology

It effectively reduces the use of storage and transmission resources and improves data processing efficiency, especially in image frame reconstruction, future image frame prediction, classification and related data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752914A_ABST
    Figure CN120752914A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to obtain encoded data associated with one or more motion values. The one or more processors are further configured to obtain a conditional input of the compression network, where the conditional input is based on the one or more first predicted motion values. The one or more processors are further configured to process the encoded data and conditional inputs using a compression network to generate one or more second predicted motion values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to commonly owned U.S. non-provisional patent application serial number 18 / 183,867, filed on March 14, 2023, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to using compressed networks to perform prediction. Background Art

[0004] Advances in technology have led to smaller and more powerful computing devices. For example, there are currently a variety of portable personal computing devices, including wireless phones such as mobile and smart phones, tablet computers and laptop computers, which are small, lightweight and easy to carry by users. These devices can communicate voice and data packets over wireless networks. In addition, many such devices incorporate additional functions such as digital still cameras, digital video cameras, digital recorders and audio file players. In addition, such devices can process executable instructions, including software applications that can be used to access the Internet, such as web browser applications. Therefore, these devices can include significant computing power.

[0005] Such computing devices often include functionality to process large amounts of data. Compressing data before storage or transmission can save resources such as memory and bandwidth. For example, a computing device can generate an encoded version of an image frame that uses fewer bits than the original image frame. Techniques to reduce the size of compressed data can further conserve resources. Summary of the Invention

[0006] According to an embodiment of the present disclosure, a device includes one or more processors configured to obtain encoded data associated with one or more motion values. The one or more processors are further configured to obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The one or more processors are further configured to process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0007] According to another embodiment of the present disclosure, a method includes obtaining, at a device, encoded data associated with one or more motion values. The method also includes obtaining, at the device, a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The method also includes processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0008] According to another embodiment of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain encoded data associated with one or more motion values. The instructions, when executed by the one or more processors, further cause the one or more processors to obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The instructions, when executed by the one or more processors, further cause the one or more processors to process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0009] According to another embodiment of the present disclosure, an apparatus includes means for obtaining encoded data associated with one or more motion values. The apparatus also includes means for obtaining a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The apparatus also includes means for processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0010] According to another embodiment of the present disclosure, a device includes one or more processors configured to obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The one or more processors are further configured to process the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0011] According to another embodiment of the present disclosure, a method includes obtaining, at a device, a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The method also includes processing, using the compression network, the conditional input and the one or more motion values ​​to generate encoded data associated with the one or more motion values.

[0012] According to another embodiment of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values. The instructions, when executed by the one or more processors, further cause the one or more processors to process the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0013] According to another embodiment of the present disclosure, an apparatus includes means for obtaining a conditional input for a compression network, wherein the conditional input is based on one or more first predicted motion values. The apparatus also includes means for processing the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0014] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections: Brief Description of the Figures, Detailed Description, and Claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram of certain illustrative aspects of a compression network operable to perform prediction according to some examples of the present disclosure.

[0016] Figure 2 is a diagram of schematic aspects of a system operable to perform prediction using a compression network according to some examples of the present disclosure.

[0017] Figure 3 is a diagram of schematic aspects of a system operable to perform prediction using a compression network according to some examples of the present disclosure.

[0018] Figure 4 is a diagram of schematic aspects of an example of an encoder portion of a compression network operable to perform prediction according to some examples of the present disclosure.

[0019] Figure 5 is a diagram of schematic aspects of an example of a decoder portion of a compression network operable to perform prediction according to some examples of the present disclosure.

[0020] Figure 6 is a diagram of schematic aspects of an example of an encoder portion of a compression network operable to perform prediction corresponding to reconstruction according to some examples of the present disclosure.

[0021] Figure 7 is a diagram of schematic aspects of an example of a decoder portion of a compression network operable to perform prediction corresponding to reconstruction according to some examples of the present disclosure.

[0022] Figure 8 is a diagram of schematic aspects of an example of an encoder portion of a compression network operable to perform prediction corresponding to reconstruction based on auxiliary prediction, according to some examples of the present disclosure.

[0023] Figure 9 is a diagram of schematic aspects of an example of a decoder portion of a compression network operable to perform prediction corresponding to reconstruction based on an auxiliary prediction, according to some examples of the present disclosure.

[0024] Figure 10 is a diagram of schematic aspects of an example of an encoder portion of a compression network operable to perform prediction corresponding to low-latency reconstruction according to some examples of the present disclosure.

[0025] Figure 11is a diagram of schematic aspects of an example of a decoder portion of a compression network operable to perform prediction corresponding to low-latency reconstruction according to some examples of the present disclosure.

[0026] Figure 12 is a diagram of schematic aspects of an example of an encoder portion of a compression network operable to perform predictions corresponding to future predictions according to some examples of the present disclosure.

[0027] Figure 13 is a diagram of schematic aspects of an example of a decoder portion of a compression network operable to perform predictions corresponding to future predictions according to some examples of the present disclosure.

[0028] Figure 14 is a diagram of schematic aspects of an example of a motion estimation encoder portion of a compression network operable to perform prediction corresponding to image reconstruction according to some examples of the present disclosure.

[0029] Figure 15 is a diagram of schematic aspects of an example of a motion estimation decoder portion of a compression network operable to perform prediction corresponding to image reconstruction according to some examples of the present disclosure.

[0030] Figure 16 is a diagram of schematic aspects of an example of an image estimator of a compression network operable to perform prediction corresponding to image reconstruction according to some examples of the present disclosure.

[0031] Figure 17 is a diagram of schematic aspects of an example of an image encoder portion of a compression network operable to perform prediction corresponding to image reconstruction according to some examples of the present disclosure.

[0032] Figure 18 is a diagram of schematic aspects of an example of an image decoder portion of a compression network operable to perform prediction corresponding to image reconstruction according to some examples of the present disclosure.

[0033] Figure 19 is a diagram of schematic aspects of an example of a motion estimation encoder portion of a compression network operable to perform prediction corresponding to low-latency image reconstruction according to some examples of the present disclosure.

[0034] Figure 20 is a diagram of schematic aspects of an example of a motion estimation decoder portion of a compression network operable to perform prediction corresponding to low-latency image reconstruction according to some examples of the present disclosure.

[0035] Figure 21is a diagram of schematic aspects of an example of an image estimator of a compression network operable to perform prediction corresponding to low-latency image reconstruction according to some examples of the present disclosure.

[0036] Figure 22 is a diagram of schematic aspects of an example of an image encoder portion of a compression network operable to perform prediction corresponding to low-latency image reconstruction according to some examples of the present disclosure.

[0037] Figure 23 is a diagram of schematic aspects of an example of an image decoder portion of a compression network operable to perform prediction corresponding to low-latency image reconstruction according to some examples of the present disclosure.

[0038] Figure 24 An example of an integrated circuit operable to perform prediction using a compression network according to some examples of the present disclosure is shown.

[0039] Figure 25 is a diagram of a mobile device operable to perform prediction using a compressed network according to some examples of the present disclosure.

[0040] Figure 26 is a diagram of a headset operable to perform prediction using a compression network according to some examples of the present disclosure.

[0041] Figure 27 is a diagram of a wearable electronic device operable to perform prediction using a compression network according to some examples of the present disclosure.

[0042] Figure 28 is a diagram of a voice-controlled speaker system operable to perform prediction using a compression network according to some examples of the present disclosure.

[0043] Figure 29 is a diagram of a camera operable to perform prediction using a compression network according to some examples of the present disclosure.

[0044] Figure 30 is a diagram of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) operable to perform prediction using a compression network according to some examples of the present disclosure.

[0045] Figure 31 is a diagram of a first example of a vehicle operable to perform prediction using a compression network according to some examples of the present disclosure.

[0046] Figure 32 is a diagram of a second example of a vehicle operable to perform prediction using a compression network according to some examples of the present disclosure.

[0047] Figure 33 Some examples of the use of Figure 1 FIGURE 1 illustrates a specific embodiment of a method for generating encoded data using a compression network.

[0048] Figure 34 Some examples of use according to the present disclosure Figure 1 A diagram of a specific embodiment of a method for generating predicted data from encoded data using a compression network.

[0049] Figure 35 is a block diagram of a particular illustrative example of a device operable to perform prediction using a compression network according to some examples of the present disclosure. DETAILED DESCRIPTION

[0050] Computing devices often include functionality for processing large amounts of data. Compressing data before storage or transmission can save resources such as memory and bandwidth. For example, a computing device can generate an encoded version of an image frame that uses fewer bits than the original image frame. Techniques for reducing the size of compressed data can further conserve resources. The compressed data can be processed to generate predicted data. For example, the predicted data can correspond to a reconstructed version of an image frame, a predicted future image frame in a sequence of images including the image frame, a classification of the image frame, other types of data associated with the image frame, or a combination thereof.

[0051] Disclosed are systems and methods for performing prediction using a compression network. For example, the compression network includes an encoder portion and a decoder portion. The encoder portion is configured to process input values ​​from a sequence of input values ​​and encoder conditional inputs to generate encoded data. The encoded data corresponds to a compressed version of the input values ​​(having fewer bits than the input values). The decoder portion is configured to process the encoded data and decoder conditional inputs to generate a predicted value associated with the input values.

[0052] The encoder conditional input is an estimate of the decoder conditional input. For example, the decoder conditional input is based on a previously predicted value generated at the decoder portion, a local decoder portion at the encoder portion is used to generate an estimate of the previously predicted value, and the encoder conditional input is based on the estimate of the previously predicted value. Generating encoded data based on an estimate of information available at the decoder portion (e.g., the previously predicted value) can reduce the size of the information (e.g., encoded data) that must be provided to the decoder portion to generate the predicted value.

[0053] The following describes certain aspects of the present disclosure with reference to the accompanying drawings. In the description, common features are represented by common reference numerals. As used herein, various terms are used only for the purpose of describing specific embodiments and are not intended to limit the embodiments. For example, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In addition, some features described herein are singular in some embodiments and plural in other embodiments. For illustration, Figure 2 Describes a process comprising one or more processors ( Figure 2 290) indicates that in some embodiments, the device 202 includes a single processor 290, while in other embodiments, the device 202 includes multiple processors 290. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular unless aspects related to multiple features are described.

[0054] In some figures, multiple instances of a particular type of feature are used. Although the features are physically and / or logically distinct, each feature uses the same reference numeral, and different instances are distinguished by adding a letter to the reference numeral. When features are referred to herein as a group or type (e.g., when no reference is made to a specific one of the features), the reference numeral is used without a distinguishing letter. However, when a specific feature of multiple features of the same type is referred to herein, the reference numeral is used with a distinguishing letter. For example, when reference is made to Figure 1 , multiple input values ​​are shown and associated with reference numerals 105A, 105B, and 105C. When referring to a specific one of these input values ​​(such as input value 105A), the distinguishing letter "A" is used. However, when referring to any one of these input values ​​or referring to these input values ​​as a group, the reference numeral 105 is used without a distinguishing letter.

[0055] As used herein, the terms "comprise," "comprises," and "comprising" may be used interchangeably with "include," "includes," or "including." Additionally, the term "wherein" may be used interchangeably with "wherein." As used herein, "exemplary" indicates an example, implementation, and / or aspect and should not be construed as limiting or indicating a preference or preferred implementation. As used herein, ordinal terms (e.g., "first," "second," "third," etc.) used to modify an element (such as a structure, component, operation, etc.) do not, by themselves, indicate any priority or order of the element relative to other elements, but rather merely distinguish the element from other elements with the same name (but with respect to the use of the ordinal term). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to a plurality (e.g., two or more) of a particular element.

[0056] As used herein, "coupling" may include "communicative coupling," "electrical coupling," or "physical coupling," and may also (or alternatively) include any combination thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. As illustrative, non-limiting examples, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronic devices, one or more connectors, or inductive coupling. In some embodiments, two devices (or components) that are communicatively coupled (such as electrically communicating) may directly or indirectly send and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "direct coupling" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.

[0057] In this disclosure, terms such as "determine," "calculate," "estimate," "shift," "adjust," etc. may be used to describe how one or more operations are performed. It should be noted that these terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generate," "calculate," "estimate," "use," "select," "access," and "determine" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal) or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.

[0058] refer to Figure 1 , certain illustrative aspects of a compression network configured to perform prediction are disclosed and generally designated 140 . Compression network 140 includes an encoder portion 160 configured to be coupled to a decoder portion 180 .

[0059] In some aspects, the encoder portion 160 is included in a first device that is different from a second device that includes the decoder portion 180, as described with reference to FIG. Figure 2 For illustration, in these aspects, the compression network 140 can be used to reduce resource usage (e.g., memory, bandwidth, transmission time, etc.) associated with the transmission of data from the first device to the second device. In other aspects, the encoder portion 160 and the decoder portion 180 are included in a single device, as described in reference to FIG. Figure 3 To illustrate, in these aspects, compression network 140 can be used to reduce resource usage (eg, memory) associated with storing data for access by a device.

[0060] The encoder portion 160 is configured to process one or more input values ​​105 to generate one or more sets of encoded data 165. In an example, the one or more input values ​​105 include input value 105A, input value 105B, and input value 105C, one or more additional input values, or a combination thereof. The encoder portion 160 is configured to process input value 105A to generate encoded data 165A, process input value 105B to generate encoded data 165B, process input value 105C to generate encoded data 165C, and so on.

[0061] The encoder portion 160 includes a conditional input generator 162 coupled to an encoder 166 via a feature generator 164. In some examples, the encoder 166 is configured to process some of the one or more input values ​​105 (e.g., key values) independently from other input values ​​in the input values ​​105 to generate corresponding encoded data. In an example, the input value 105A corresponds to a key image frame (e.g., an intra frame or an I-frame), and the encoder 166 processes the input value 105A independently from the other input values ​​in the one or more input values ​​105 to generate the encoded data 165A. In some examples, the encoder 166 is configured to process some of the one or more input values ​​105 based on at least one other input value in the one or more input values ​​105 (e.g., a non-key value) to generate the corresponding encoded data. In the example, input value 105B corresponds to a non-key picture frame, such as a predicted frame (P frame) or a bidirectional frame (B frame), and encoder 166 processes input value 105B based on encoded data 165A (corresponding to input value 105A) to generate encoded data 165B. For illustration, conditional input generator 162 is configured to process encoded data 165A to generate conditional input 167B, feature generator 164 is configured to process conditional input 167B to generate feature data 163B, and encoder 166 is configured to process input value 105B and feature data 163B to generate encoded data 165B.

[0062] Decoder portion 180 is configured to process one or more sets of encoded data 165 to generate one or more predicted values ​​195. In an example, decoder portion 180 is configured to process encoded data 165A to generate predicted value 195A, process encoded data 165B to generate predicted value 195B, process encoded data 165C to generate predicted value 195C, and so on.

[0063] Decoder portion 180 includes a conditional input generator 182 coupled to a decoder 186 via a feature generator 184. In some examples, decoder 186 is configured to process some sets of encoded data 165 independently from other sets of encoded data 165 to generate corresponding predicted values. In an example, decoder 186 processes encoded data 165A independently from other sets of encoded data 165 to generate predicted values ​​195A. In some examples, decoder 186 is configured to process some sets of encoded data 165 based on at least one other set of encoded data 165 to generate corresponding predicted values. For example, decoder 186 processes encoded data 165B based on predicted values ​​195A (corresponding to encoded data 165A) to generate predicted values ​​195B. To illustrate, conditional input generator 182 is configured to process predicted value 195A to generate conditional input 187B, feature generator 184 is configured to process conditional input 187B to generate feature data 183B, and decoder 186 is configured to process encoded data 165B and feature data 183B to generate predicted value 195B.

[0064] Conditional input 167B at encoder portion 160 corresponds to an estimate of conditional input 187B available at decoder portion 180. Technical advantages of generating encoded data 165B based on an estimate of information available at decoder portion 180 (e.g., conditional input 187B) may include reducing the size of the information (e.g., encoded data 165B) that must be provided to decoder portion 180 to generate predicted value 195B.

[0065] In some embodiments, the compression network components (e.g., encoder portion 160, decoder portion 180, or both) of compression network 140 correspond to or are included in one of various types of devices. In the illustrative example, the compression network components are integrated into a head mounted device (such as a head mounted device). Figure 26 In other examples, the compression network component is integrated into a mobile phone or tablet computer device (such as the reference Figure 25 As described), wearable electronic devices (such as reference Figure 27 As described), voice control speaker system (as referenced Figure 28 As described), camera equipment (as referenced Figure 29 ) or virtual reality, mixed reality, and augmented reality headsets (as referenced Figure 30 In another illustrative example, the compression network component is integrated into a carrier (such as the reference Figure 31 and Figure 32 further described).

[0066] During operation, the encoder portion 160 obtains one or more input values ​​105. In some embodiments, the one or more input values ​​105 are based on the output of one or more sensors, such as reference signals. Figure 2 As further described. As illustrative, non-limiting examples, the one or more sensors may include an image sensor, an inertial measurement unit (IMU), a motion sensor, a temperature sensor, other types of sensors, or a combination thereof. In some embodiments, the encoder portion 160 obtains one or more input values ​​105 from a storage device, a processor, other devices, or a combination thereof.

[0067] In some embodiments, one or more input values ​​105 indicate motion, such as reference Figure 14 Further described. For example, input value 105A includes a first image frame captured by a camera of a vehicle, and input value 105B includes a second image frame captured by the camera. The difference between the first image frame and the second image frame can indicate the speed and direction of movement of the vehicle. One or more input values ​​105 including image frames indicating motion are provided as an illustrative example. In other examples, one or more input values ​​105 can include other types of data indicating motion. One or more input values ​​105 indicating motion are provided as an illustrative example. In other examples, one or more input values ​​105 can correspond to data indicating other types of information.

[0068] Encoder portion 160 generates a set of encoded data 165 corresponding to one or more input values ​​105. For example, encoder 166 processes input value 105A to generate encoded data 165A. In some embodiments, encoder 166 processes (e.g., encodes) input value 105A independently from other input values ​​in one or more input values ​​105 to generate encoded data 165A based on a determination that input value 105A satisfies an independent encoding criterion. In examples, encoder 166 determines that the independent encoding criterion is satisfied based on determining that input value 105A is an initial value for one or more input values ​​105, input value 105A corresponds to a key value (e.g., an I-frame), at least a threshold number of input values ​​have been encoded since the most recently independently encoded input value, a difference between input value 105A and a previous input value of one or more input values ​​105 is greater than a difference threshold (e.g., due to a scene change in an image frame), or a combination thereof.

[0069] As another example, encoder 166 processes input value 105B to generate encoded data 165B. In some embodiments, encoder 166 processes input value 105B based at least in part on encoded data 165A of input value 105A in response to determining that input value 105B fails to meet independent encoding criteria. In some aspects, encoder 166 selects encoded data 165A for processing input value 105B based on determining that input value 105A corresponds to a value (e.g., a key value) that was most recently independently encoded to generate encoded data 165A among one or more input values ​​105. In some aspects, encoder 166 selects encoded data 165A for processing input value 105B based on determining that input value 105B corresponds to the next value after input value 105A among one or more input values ​​105. In these aspects, input value 105A can correspond to a key value that was independently encoded to generate encoded data 165A or to a non-key value that was encoded based on at least one other input value among one or more input values ​​105.

[0070] In response to determining that input value 105B fails to meet the independent encoding criteria, encoder portion 160 provides encoded data 165A to conditional input generator 162 to obtain conditional input 167B for compression network 140. Conditional input generator 162 includes a local decoder portion 168, an estimator 170, or both. In some aspects, local decoder portion 168 is configured to perform operations similar to decoder portion 180. For example, local decoder portion 168 performs one or more operations described herein with reference to decoder portion 180 to process encoded data 165A to generate predicted value 169A. Predicted value 169A corresponds to an estimate of predicted value 195A that can be generated at decoder portion 180 by processing encoded data 165A.

[0071] Optionally, in some embodiments, estimator 170 processes predicted value 169A to generate estimated value 171 B. In some aspects, estimated value 171B corresponds to an estimate of input value 105B that can be generated at decoder portion 180 based on predicted value 195A. In some aspects, the closer estimated value 171B is to input value 105B, the less information must be provided to decoder portion 180 as encoded data 165B.

[0072] Conditional input 167B is based on predicted value 169A, estimated value 171B, or both. In some embodiments, conditional input 167B can be based on one or more additional predicted values, one or more additional estimated values, or a combination thereof associated with one or more other input values ​​in one or more input values ​​105.

[0073] In response to determining that the input value 105B fails to meet the independent encoding criteria, the encoder portion 160 processes the input value 105B and the conditional input 167B using the compression network 140 to generate encoded data 165B associated with the input value 105B. For example, the feature generator 164 processes the conditional input 167B to generate feature data 163B, as shown in FIG. Figure 4 As further described. In some embodiments, the feature data 163B corresponds to multi-scale feature data having different resolutions (e.g., different spatial resolutions). In some embodiments, the feature data 163B includes multi-scale wavelet transform data. The encoder 166 processes the input value 105B and the feature data 163B to generate encoded data 165B, as shown in FIG. Figure 4 Further described.

[0074] The decoder portion 180 obtains one or more of a set of encoded data 165 associated with the one or more input values ​​105. In some embodiments, the encoder portion 160 at the first device provides the encoded data 165 to the decoder portion 180 at the second device via a bitstream, as described with reference to FIG. Figure 2 In other embodiments, the encoder portion 160 stores the encoded data 165 at a storage device, and the decoder portion 180 retrieves the encoded data 165 from the storage device, as described in reference to Figure 3 Further described.

[0075] Decoder portion 180 generates one or more predicted values ​​195 corresponding to a set of encoded data 165. For example, decoder 186 processes encoded data 165A to generate predicted values ​​195A. In some embodiments, decoder 186 processes (e.g., decodes) encoded data 165A independently from other predicted values ​​in one or more predicted values ​​195 to generate predicted values ​​195A based on a determination that encoded data 165A meets independent decoding criteria. In examples, decoder 186 determines that the independent encoding criteria are met based on a determination that metadata associated with encoded data 165 indicates that encoded data 165 corresponds to a key value (e.g., an I-frame), that encoded data 165 is to be independently decoded, or both.

[0076] As another example, decoder 186 processes coded data 165B to generate predicted value 195B. In some embodiments, decoder 186 processes coded data 165B based at least in part on predicted value 195A corresponding to coded data 165A associated with input value 105A in response to determining that coded data 165B fails to meet independent decoding criteria. In some aspects, decoder 186 selects predicted value 195A for processing coded data 165B based on determining that predicted value 195A corresponds to a value (e.g., a key value) that was most recently independently decoded. In some aspects, decoder 186 selects predicted value 195A for processing coded data 165B based on determining that predicted value 195B corresponds to the next value after predicted value 195A in one or more predicted values ​​195. In these aspects, predicted value 195A can correspond to an independently decoded key value or to a non-critical value decoded based on at least one other predicted value in one or more predicted values ​​195.

[0077] In response to determining that the encoded data 165B fails to meet the independent decoding criteria, the decoder portion 180 provides the predicted value 195A to the conditional input generator 182 to obtain the conditional input 187B for the compression network 140. In some embodiments, the conditional input generator 182 outputs the predicted value 195A as the conditional input 187B. Optionally, in some embodiments, the conditional input generator 182 includes an estimator 188 that processes the predicted value 195A to generate an estimated value 193B. In some aspects, the estimated value 193B corresponds to an estimate of the input value 105B based on the predicted value 195A. In some aspects, the closer the estimated value 193B is to the input value 105B, the less information that must be processed by the decoder portion 180 as the encoded data 165B to generate the predicted value 195B.

[0078] Conditional input 187B is based on predicted value 195A, estimated value 193B, or both. In some implementations, conditional input 187B can be based on at least one additional value of one or more predicted values ​​195, one or more additional estimated values, or a combination thereof.

[0079] In response to determining that the encoded data 165B fails to meet the independent decoding criteria, the decoder portion 180 processes the encoded data 165B and the conditional input 187B using the compression network 140 to generate a predicted value 195B associated with the input value 105B. For example, the feature generator 184 processes the conditional input 187B to generate feature data 183B, as shown in FIG. Figure 5As further described. In some embodiments, the feature data 183B corresponds to multi-scale feature data having different resolutions (e.g., different spatial resolutions). In some embodiments, the feature data 183B includes multi-scale wavelet transform data. The decoder 186 processes the encoded data 165B and the feature data 183B to generate a predicted value 195B, as shown in FIG. Figure 5 Further described.

[0080] In some examples, predicted value 195B corresponds to a reconstructed version of input value 105B. In some examples, predicted value 195B corresponds to a predicted future value, such as a prediction of a future value of one or more input values ​​105. For illustration, a future value may be an input value 105C that is not yet available at encoder portion 160 when encoded data 165B is generated. In some examples, predicted value 195B corresponds to a classification associated with input value 105B. For illustration, predicted value 195B may indicate whether input value 105B corresponds to an alert condition.

[0081] In some examples, predicted value 195B corresponds to a detection result associated with input value 105B. For illustration, predicted value 195B indicates whether a face was detected in the image frame associated with input value 105B. In some examples, predicted value 195B corresponds to a collision avoidance output. For example, input value 105B indicates a first position of a vehicle relative to a second position of an object. In some aspects, predicted value 195B indicates a first predicted future position of the vehicle relative to a second predicted future position of the object. In some aspects, predicted value 195B indicates whether the first predicted future position of the vehicle is within a collision threshold of the second future position of the object. In some examples, one or more predicted values ​​195 are processed by one or more downstream applications, and compression network 140 is trained (e.g., configured) based on a performance metric associated with the downstream application. For example, compression network 140 is trained to reduce a loss metric associated with the downstream application.

[0082] Technical advantages of using compression network 140 to generate one or more predicted values ​​195 may include reduced resource usage. For example, generating encoded data 165B based on conditional input 167B as an estimate of information available at decoder portion 180 (e.g., conditional input 187B) reduces the information that must be provided to decoder portion 180 as encoded data 165B to maintain the accuracy of predicted values ​​195B.

[0083] refer to Figure 2, schematic aspects of a system operable to perform prediction using a compression network are shown and generally designated 200. System 200 includes a device 202 configured to be coupled to or include one or more sensors 240. Device 202 is configured to communicate with device 260. For example, device 202 is configured to couple to device 260 via a network (e.g., a wired network, a wireless network, or both).

[0084] The encoder portion 160 is included in one or more processors 290 of the device 202, and the decoder portion 180 is included in one or more processors 292 of the device 260. The one or more processors 290 are coupled to the one or more sensors 240 and the modem 270. The one or more processors 292 are coupled to the modem 280. Optionally, in some embodiments, the one or more processors 292 include one or more applications 262. Optionally, in some embodiments, the device 260 is configured to be coupled to a device 264.

[0085] During operation, the encoder portion 160 receives sensor data 226 from one or more sensors 240 as one or more input values ​​105. As illustrative, non-limiting examples, the one or more sensors 240 may include an image sensor, an IMU, a motion sensor, an accelerometer, a speedometer, a gyroscope, a radar, a temperature sensor, a microphone, other types of sensors, or combinations thereof. The encoder portion 160 processes the one or more input values ​​105 (e.g., sensor data 226) to generate a set of encoded data 165, as shown in FIG. Figure 1 The encoder portion 160 provides the set of encoded data 165 to the modem 270 for transmission to the device 260 as a bitstream 235. The modem 280 receives the bitstream 235 and provides the set of encoded data 165 to the decoder portion 180.

[0086] The decoder portion 180 processes the set of encoded data 165 to generate one or more predicted values ​​195, as shown in FIG. Figure 1 The one or more processors 292 generate an output 295 based on the one or more predicted values ​​195. Optionally, in some embodiments, the decoder portion 180 provides the one or more predicted values ​​195 to one or more applications 262, and the one or more applications 262 generate an output 263 based on the one or more predicted values ​​195. For example, the one or more predicted values ​​195 may correspond to an image frame, and the application 262 processes the predicted values ​​195B corresponding to the image frame to generate an output 263 indicating a classification associated with the image frame. The output 295 is based on the output 263, the one or more predicted values ​​195, or a combination thereof.

[0087] In some embodiments, the encoder portion 160, the decoder portion 180, or both are trained (e.g., configured) based on performance metrics associated with one or more applications 262. For example, a network trainer trains the encoder portion 160, the decoder portion 180, or both (e.g., configures the network weights and biases of the encoder portion 160, the decoder portion 180, or both) based on a loss metric associated with a classification output generated by the application 262. In some embodiments, the one or more processors 292 provide an output 295 to a device 264. The device 264 may include a display device, a network device, a storage device, a user device, or a combination thereof.

[0088] In some aspects, output 295 initiates one or more actions at device 264. In the illustrative example, device 264 and device 202 are the same device (e.g., a vehicle), and the one or more predicted values ​​195 correspond to collision avoidance outputs. Device 260 transmits output 295 to initiate one or more collision avoidance actions (e.g., braking) at device 202 in response to determining that predicted value 195B indicates that a first predicted future position of device 260 is expected to be within a threshold distance of a second predicted future position of an object. In some embodiments, device 260 is the same as or included in device 202 (e.g., a vehicle). In other embodiments, device 260 is external to device 202 and generates output 295 based on one or more predicted values ​​195, based on one or more additional predicted values ​​(e.g., associated with an object, another vehicle, or both), or a combination thereof.

[0089] Technical advantages of using encoder portion 160 to generate encoded data 165 based on an estimate of information available at decoder portion 180 may include reduced resource usage (eg, memory, bandwidth, and transmission time) associated with sending bitstream 235 to device 260 .

[0090] refer to Figure 3 , schematic aspects of a system operable to perform prediction using a compression network are shown and generally designated 300. System 300 includes a device 302 configured to be coupled to or include one or more sensors 240. Device 302 is configured to be coupled to or include a storage device 392.

[0091] The encoder portion 160 and the decoder portion 180 are included in one or more processors 390 of the device 302. The encoder portion 160 is configured to store one or more of the sets of encoded data 165 in the storage device 392. The decoder portion 180 is configured to obtain one or more of the sets of encoded data 165 from the storage device 392. Storing one or more of the sets of encoded data 165 in the storage device 392 may use less memory than storing corresponding values ​​of the sensor data 226 in the storage device 392.

[0092] The decoder portion 180 processes the set of encoded data 165 to generate one or more predicted values ​​195, as shown in FIG. Figure 1 The one or more processors 390 generate an output 295 based on the one or more predicted values ​​195. Optionally, in some embodiments, the decoder portion 180 provides the one or more predicted values ​​195 to one or more applications 262, and the one or more applications 262 generate an output 263 based on the one or more predicted values ​​195. The output 295 is based on the output 263, the one or more predicted values ​​195, or a combination thereof. In some embodiments, the one or more processors 390 provide the output 295 to the device 264.

[0093] In the illustrative example, device 264 and device 302 are the same device (e.g., a vehicle), and one or more predicted values ​​195 correspond to collision avoidance outputs. Device 302 generates output 295 to initiate one or more collision avoidance maneuvers (e.g., braking) at device 302 in response to determining that predicted value 195B indicates that a first predicted future position of device 302 is expected to be within a threshold distance of a second predicted future position of the object.

[0094] refer to Figure 4 , shows an example 400 of the encoder portion 160 of the compression network 140 operable to perform prediction. For example, the encoder portion 160 is configured to generate encoded data 165B, which can be processed at the decoder portion 180 of the compression network 140 to generate predicted values ​​195B, as shown in FIG. Figure 5 Further described.

[0095] The compression network 140 comprises a neural network having multiple layers. For example, the feature generator 164 comprises one or more feature layers 404 coupled to one or more encoder layers 402 of the encoder 166. The one or more feature layers 404 include feature layer 404A, feature layer 404B, feature layer 404C, one or more additional feature layers including feature layer 404N, or a combination thereof. The one or more encoder layers 402 include encoder layer 402A, encoder layer 402B, encoder layer 402C, one or more additional encoder layers including encoder layer 402N, or a combination thereof. The output of each feature layer 404 is coupled to the input of the corresponding encoder layer 402. For example, the output of feature layer 404A is coupled to the input of encoder layer 402A, the output of feature layer 404B is coupled to the input of encoder layer 402B, and so on.

[0096] The output of each previous feature layer 404 is coupled to the input of a subsequent feature layer 404. For example, the output of feature layer 404A is coupled to the input of feature layer 404B, the output of feature layer 404B is coupled to the input of feature layer 404C, and so on. The output of each previous encoder layer 402 is coupled to the input of a subsequent encoder layer 402. For example, the output of encoder layer 402A is coupled to the input of encoder layer 402B, the output of encoder layer 402B is coupled to the input of encoder layer 402C, and so on.

[0097] In some embodiments, encoder 166 is configured to encode multiple orders of resolution of input values ​​105 to generate encoded data 165. In the illustrative example, encoder 166 corresponds to a video encoder configured to encode multiple orders of spatial resolution of input values ​​105 (e.g., image frames). One or more feature layers 404 and one or more encoder layers 402 correspond to network layers associated with multiple resolutions. For example, feature layer 404A and encoder layer 402A are associated with a first resolution, feature layer 404B and encoder layer 402B are associated with a second resolution, feature layer 404C and encoder layer 402C are associated with a third resolution, feature layer 404N and encoder layer 402N are associated with an Nth resolution, or a combination thereof.

[0098] During operation, the encoder portion 160 provides the conditional input 167B to the feature generator 164 and the input value 105B ( ) is provided to the encoder 166, as referenced Figure 1Feature layer 404A processes conditional input 167B to generate feature data 163BA associated with the first resolution and provides feature data 163BA to encoder layer 402A. Encoder layer 402A processes input value 105B and feature data 163BA to generate output (associated with the first resolution) that is provided to encoder layer 402B.

[0099] Each subsequent feature layer 404 processes the output of the previous feature layer 404 to generate feature data 163 and provides the feature data 163 to the corresponding encoder layer 402. For example, feature layer 404B provides feature data 163BB associated with the second resolution to encoder layer 402B, feature layer 404C provides feature data 163BC associated with the third resolution to encoder layer 402C, feature layer 404N provides feature data 163BN associated with the Nth resolution to encoder layer 402N, etc. The feature data 163B (e.g., feature data 163BA, feature data 163BB, feature data 163BC, feature data 163BN, or a combination thereof) thus corresponds to multi-scale feature data having different resolutions.

[0100] Each subsequent encoder layer 402 processes the output of the previous encoder layer 402 and the feature data 163 from the corresponding feature layer to generate an output. For example, encoder layer 402B processes the output of encoder layer 402A and feature data 163BB to generate an output (associated with the second resolution) that is provided to encoder layer 402C, and so on. The output of encoder layer 402N corresponds to encoded data 165B.

[0101] Optionally, in some embodiments, feature generator 164 may use various techniques to generate feature data 163B. For example, feature generator 164 may perform a wavelet transform to generate feature data 163B corresponding to multi-scale wavelet transform data. To illustrate, feature generator 164 may perform a wavelet transform based on conditional input 167B to generate feature data 163BA associated with a first resolution (e.g., first wavelet transform data). Feature generator 164 may perform a wavelet transform based on feature data 163BA, conditional input 167B, or both to generate feature data 163BB associated with a second resolution (e.g., second wavelet transform data). Similarly, feature generator 164 may generate feature data 163BC associated with a third resolution (e.g., third wavelet transform data), and so on.

[0102] In certain aspects, the conditional input 167B may be associated with Figure 5The feature data 163B generated by the decoder portion 180 may correspond to an estimate of the conditional input 187B generated at the decoder portion 180. Technical advantages of using an estimate of the information available at the decoder portion 180 (e.g., the conditional input 167B) to generate the feature data 163B used to generate the encoded data 165B may include a reduced size of the encoded data 165B to maintain the accuracy of the predicted value 195B generated at the decoder portion 180.

[0103] refer to Figure 5 , shows an example 500 of a decoder portion 180 of a compression network 140 operable to perform prediction. For example, the decoder portion 180 is configured to process (as shown in FIG. Figure 4 The data 165B is encoded at the encoder portion 160 to generate the same value as the input value 105B ( ) associated with the predicted value 195B ( ).

[0104] The compression network 140 comprises a neural network having multiple layers. For example, the feature generator 184 comprises one or more feature layers 504 coupled to one or more decoder layers 508 of the decoder 186. The one or more feature layers 504 include feature layer 504A, feature layer 504B, feature layer 504C, one or more additional feature layers including feature layer 504N, or a combination thereof. The one or more decoder layers 508 include decoder layer 508A, decoder layer 508B, decoder layer 508C, one or more additional decoder layers including decoder layer 508N, or a combination thereof. The output of each feature layer 504 is coupled to the input of the corresponding decoder layer 508. For example, the output of feature layer 504A is coupled to the input of decoder layer 508A, the output of feature layer 504B is coupled to the input of decoder layer 508B, and so on.

[0105] The output of each previous feature layer 504 is coupled to the input of the subsequent feature layer 504. For example, the output of feature layer 504A is coupled to the input of feature layer 504B, the output of feature layer 504B is coupled to the input of feature layer 504C, and so on. The output of each previous decoder layer 508 is coupled to the input of the subsequent decoder layer 508. The one or more decoder layers 508 are ordered according to reference numerals ending in higher letters to lower letters. For example, decoder layer 508B follows decoder layer 508C, and decoder layer 508A follows decoder layer 508B. The output of decoder layer 508B is coupled to the input of decoder layer 508A, the output of decoder layer 508C is coupled to the input of decoder layer 508B, and so on. The output of the last decoder layer 508 of decoder 186 (e.g., decoder layer 508A) corresponds to the predicted value 195 associated with the input value 105.

[0106] In some embodiments, decoder 186 is configured to decode multiple levels of resolution of encoded data 165B. In the illustrative example, decoder 186 corresponds to a video decoder configured to decode multiple levels of spatial resolution of encoded data 165B associated with input values ​​105 (e.g., image frames). One or more feature layers 504 and one or more decoder layers 508 correspond to network layers associated with multiple resolutions. For example, feature layer 504A and decoder layer 508A are associated with a first resolution, feature layer 504B and decoder layer 508B are associated with a second resolution, feature layer 504C and decoder layer 508C are associated with a third resolution, feature layer 504N and decoder layer 508N are associated with an Nth resolution, or a combination thereof.

[0107] During operation, the decoder portion 180 provides the conditional input 187B to the feature generator 184 and the encoded data 165B to the decoder 186, as shown in FIG. Figure 1 As described above, feature layer 504A processes conditional input 187B to generate feature data 183BA associated with a first resolution and provides feature data 183BA to decoder layer 508A. Each subsequent feature layer 504 processes the output of the previous feature layer 504 to generate feature data 183 and provides feature data 183 to the corresponding decoder layer 508. For example, feature layer 504B provides feature data 183BB associated with the second resolution to decoder layer 508B, feature layer 504C provides feature data 183BC associated with the third resolution to decoder layer 508C, feature layer 504N provides feature data 183BN associated with the Nth resolution to decoder layer 508N, and so on. Feature data 183B (e.g., feature data 183BA, feature data 183BB, feature data 183BC, feature data 183BN, or a combination thereof) thus corresponds to multi-scale feature data having different resolutions.

[0108] Decoder layer 508N processes the encoded data 165B and feature data 183BN to generate an output (e.g., associated with the Nth resolution) that is provided to a subsequent decoder layer 508. Each subsequent decoder layer 508 processes the output of the previous decoder layer 508 and feature data 183 from the corresponding feature layer to generate an output. For example, decoder layer 508B processes the output of decoder layer 508C and feature data 183BB to generate an output (e.g., associated with the second resolution). Decoder layer 508A processes the output of decoder layer 508B and feature data 183BA to generate an output corresponding to the input value 105B ( ) associated with the predicted value 195B ( ) output.

[0109] Optionally, in some embodiments, the feature generator 184 may use Figure 4 Feature data 183B may be generated using various techniques similar to those performed by feature generator 164. For example, feature generator 184 may perform a wavelet transform to generate feature data 183B corresponding to multi-scale wavelet transformed data. To illustrate, feature generator 184 may perform a wavelet transform based on conditional input 187B to generate feature data 183BA associated with a first resolution (e.g., first wavelet transformed data). Feature generator 184 may perform a wavelet transform based on feature data 183BA, conditional input 187B, or both to generate feature data 183BB associated with a second resolution (e.g., second wavelet transformed data). Similarly, feature generator 184 may generate feature data 183BC associated with a third resolution (e.g., third wavelet transformed data), and so on.

[0110] Optionally, in some embodiments, decoder portion 180 may include a feature generator 582 configured to process decoder information 587 available at decoder portion 180 to generate feature data 583, which is used by decoder 186 to generate predicted values ​​195B. By way of example, feature generator 582 includes one or more feature layers 506 coupled to one or more decoder layers 508. One or more feature layers 506 include feature layer 506A, feature layer 506B, feature layer 506C, one or more additional feature layers including feature layer 506N, or a combination thereof. The output of each feature layer 506 is coupled to the input of a corresponding decoder layer 508. For example, the output of feature layer 506A is coupled to the input of decoder layer 508A, the output of feature layer 506B is coupled to the input of decoder layer 508B, and so on.

[0111] The output of each previous feature layer 506 is coupled to the input of a subsequent feature layer 506. For example, the output of feature layer 506A is coupled to the input of feature layer 506B, the output of feature layer 506B is coupled to the input of feature layer 506C, and so on. In particular aspects, one or more feature layers 506 correspond to network layers associated with multiple resolutions. For example, feature layer 506A is associated with a first resolution, feature layer 506B is associated with a second resolution, feature layer 506C is associated with a third resolution, feature layer 506N is associated with an Nth resolution, or a combination thereof.

[0112] In addition to providing conditional input 187B to feature generator 184 and encoded data 165B to decoder 186 (e.g., while providing conditional input 187B to feature generator 184 and encoded data 165B to decoder 186), decoder portion 180 also provides decoder information 587 to feature generator 582. Feature layer 506A processes decoder information 587 to generate feature data 583A associated with the first resolution and provides feature data 583A to decoder layer 508A. Each subsequent feature layer 506 processes the output of the previous feature layer 506 to generate feature data 583 and provides feature data 583 to the corresponding decoder layer 508. For example, feature layer 506B processes the output of feature layer 506A to generate feature data 583B and provides feature data 583B associated with the second resolution to decoder layer 508B. Similarly, feature layer 506C provides feature data 583C associated with the third resolution to decoder layer 508C, feature layer 506N provides feature data 583N associated with the Nth resolution to decoder layer 508N, etc. Feature data 583 (e.g., feature data 583A, feature data 583B, feature data 583C, feature data 583N, or a combination thereof) thus corresponds to multi-scale feature data having different resolutions.

[0113] Decoder layer 508N processes encoded data 165B, feature data 183BN, and feature data 583N to generate an output (e.g., associated with the Nth resolution) that is provided to a subsequent decoder layer 508. Each subsequent decoder layer 508 processes the output of the previous decoder layer 508, feature data 183 of the corresponding feature layer from feature generator 184, and feature data 583 of the corresponding feature layer from feature generator 582 to generate an output. For example, decoder layer 508B processes the output of decoder layer 508C, feature data 183BB, and feature data 583B to generate an output (e.g., associated with the second resolution). Decoder layer 508A processes the output of decoder layer 508B, feature data 183BA, and feature data 583A to generate a feature value corresponding to the input value 105B ( ) associated with the predicted value 195B ( ) output.

[0114] Alternatively, in some embodiments, feature generator 582 may use various techniques (e.g., similar to those performed by feature generator 184) to generate feature data 583. For example, feature generator 582 may perform a wavelet transform to generate feature data 583 corresponding to multi-scale wavelet transformed data.

[0115] In some embodiments, the predicted value 195B ( ) corresponds to an input value of 105B ( ) of the reconstructed version, as referenced Figures 6 to 11 For example, if the input value is 105B ( ) may correspond to an image unit, and the predicted value 195B ( ) may correspond to a reconstructed version of the image unit, as in reference Figures 14 to 23 As further described. In some embodiments, the predicted value 195B ( ) corresponds to a predicted future value of one or more input values ​​105, as referenced Figures 12 to 13 As further described. In other embodiments, the predicted value 195B ( ) may correspond to a predicted future value associated with the input value 105B. In some aspects, the predicted value 195B also includes a prediction of a future value of the one or more input values ​​105. In other aspects, the predicted value 195B does not include a prediction of a future value of the one or more input values ​​105. In an example, the compression network 140 tracks an object associated with one or more motion values ​​(e.g., one or more input values ​​105) across one or more frames of pixels. To illustrate, the input value 105B may include tracking data (e.g., associated with an image frame), and the predicted value 195B ( ) can include a collision avoidance output indicating whether the future position of the vehicle relative to the object (e.g., the future position of the object) is predicted to be less than a threshold. The collision avoidance output (e.g., collision or no collision) is a future prediction associated with the tracking data. In some embodiments, the collision avoidance output also indicates predicted future tracking data, such as the predicted future position of the vehicle relative to the predicted future position of the object.

[0116] refer to Figure 6 , shows an example 600 of the encoder portion 160 of the compression network 140 operable to perform prediction corresponding to the reconstruction. For example, the encoder portion 160 is configured to generate encoded data 165B, which can be processed at the decoder portion 180 of the compression network 140 to generate an image corresponding to the input value 105B ( ) corresponds to the predicted value 195B ( ), as referenced Figure 7 Further described.

[0117] Conditional input generator 162 processes encoded data 165A using local decoder portion 168 to generate predicted values ​​169A, as shown in FIG. Figure 1 In this particular example, the predicted value 169A ( ) and the input value 105A ( ) corresponds to the estimate of the reconstructed version.

[0118] In some implementations, conditional input generator 162 generates one or more additional predicted values. The one or more additional predicted values ​​can be associated with at least one input value before input value 105B in one or more input values ​​105, at least one input value after input value 105B in one or more input values ​​105, or both.

[0119] In an example, the one or more additional predicted values ​​may include a predicted value 169C that is associated with an input value 105C that follows an input value 105B in the one or more input values ​​105 and that may be encoded as encoded data 165C independently of (e.g., before) the encoded input value 105B. In the illustrative example, input value 105C corresponds to a key value (e.g., an I-frame) that may be encoded independently of other key values ​​in one or more input values ​​105. Conditional input generator 162 uses local decoder portion 168 to process encoded data 165C to generate predicted value 169C, process one or more additional sets of encoded data to generate one or more additional predicted values, or a combination thereof.

[0120] Optionally, in some embodiments, conditional input generator 162 uses estimator 170 to generate input value 105B based on one or more predicted values ​​(e.g., predicted value 169A, predicted value 169C, one or more additional predicted values, or a combination thereof). ) of the estimated value 171B ( For illustration, the estimated value 171B ( ) and the input value 105B that may be generated at the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values ​​and independently of the encoded data 165B ( ) corresponds to the estimate of .

[0121] In certain embodiments, estimator 170 uses one or more estimation techniques to determine estimated value 171B. For example, predicted value 169A corresponds to input value 105A (e.g., a previous image frame), predicted value 169C corresponds to input value 105C (e.g., a subsequent image frame), and estimated value 171B corresponds to an estimated input value between predicted value 169A and predicted value 169C. In some embodiments, estimator 170 includes a neural network configured to process one or more predicted values ​​to generate estimated value 171B.

[0122] Conditional input 167B includes estimated value 171B ( ), one or more predicted values ​​(such as predicted value 169A ( ), predicted value 169C ( ), one or more additional predicted values, or a combination thereof. The feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes the input value 105B based on the feature data 163B to generate encoded data 165B, as shown in FIG. Figure 4 described.

[0123] refer to Figure 7 , shows an example 700 of a decoder portion 180 of a compression network 140 operable to perform prediction corresponding to a reconstruction. For example, the decoder portion 180 is configured to process (as referenced) Figure 6 The data 165B is encoded at the encoder portion 160 to generate the same value as the input value 105B ( ) corresponds to the predicted value 195B ( ).

[0124] Conditional input generator 182 obtains predicted value 195A, as shown in FIG. Figure 1 For example, decoder portion 180 processes encoded data 165A to generate predicted value 195A. In a specific example, predicted value 195A ( ) corresponds to an input value of 105A ( ) of the rebuilt version.

[0125] In some embodiments, conditional input generator 182 obtains one or more additional predicted values. The one or more additional predicted values ​​can be associated with at least one input value before input value 105B in one or more input values ​​105, at least one input value after input value 105B in one or more input values ​​105, or both.

[0126] In an example, the one or more additional predicted values ​​may include a predicted value 195C associated with an input value 105C that follows the input value 105B in the one or more input values ​​105 ( ), and corresponding encoded data 165C can be decoded independently of (e.g., before) decoding the encoded data 165B associated with the input value 105B to generate the predicted value 195C. In the illustrative example, the input value 105C corresponds to a key value (e.g., an I-frame) and is associated with the encoded data 165C, which can be decoded independently of the rest of the set of encoded data 165.

[0127] Optionally, in some embodiments, conditional input generator 182 uses estimator 188 to generate input value 105B based on one or more predicted values ​​(e.g., predicted value 195A, predicted value 195C, one or more additional predicted values, or a combination thereof). ) of the estimated value 193B ( For illustration, the estimated value 193B ( ) and the input value 105B that may be generated at the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values ​​and independently of the encoded data 165B ( ) corresponds to the estimate of .

[0128] In certain embodiments, estimator 188 uses one or more estimation techniques (similar to the estimation techniques performed by estimator 170, see Figure 6 ) to determine estimated value 193B. For example, predicted value 195A corresponds to input value 105A (e.g., a previous image frame), predicted value 195C corresponds to input value 105C (e.g., a subsequent image frame), and estimated value 193B corresponds to an estimated input value between predicted value 195A and predicted value 195C. In some embodiments, estimator 188 includes a neural network configured to process one or more predicted values ​​to generate estimated value 193B.

[0129] Conditional input 187B includes estimated value 193B ( ), one or more predicted values ​​(such as predicted value 195A ( ), predicted value 195C ( ), one or more additional predicted values, or a combination thereof. Feature generator 184 generates feature data 183B based on conditional input 187B, and decoder 186 processes encoded data 165B based on feature data 183B to generate predicted values ​​195B ( ), as referenced Figure 5 The predicted value is 195B ( ) corresponds to an input value of 105B ( ) of the rebuilt version.

[0130] Technical advantages of using an estimate (e.g., conditional input 167B) of information available at decoder portion 180 (e.g., conditional input 187B) to generate encoded data 165B may include reducing the number of bytes that must be encoded as encoded data 165B to generate input value 105B ( ) of the rebuilt version ( ) amount of information.

[0131] refer to Figure 8 , shows an example 800 of an encoder portion 160 of a compression network 140 operable to perform prediction corresponding to a reconstruction based on an auxiliary prediction. For example, the encoder portion 160 is configured to generate encoded data 165B that can be processed at a decoder portion 180 of the compression network 140 based on the auxiliary prediction data to generate a value corresponding to the input value 105B ( ) corresponds to the predicted value 195B ( ), as referenced Figure 9 Further described.

[0132] It will be appreciated that the use of auxiliary prediction data to generate the same ) corresponds to the predicted value 195B ( ) is used as an illustrative example. In other examples, the compression network 140 can use the auxiliary prediction data to generate various other types of predicted values ​​195, such as predicted future values, collision avoidance outputs, classification outputs, detection outputs, etc.

[0133] The encoder portion 160 includes or is coupled to one or more auxiliary prediction layers 806. In certain aspects, the encoder portion 160 is configured to provide domain-specific data 805 to the one or more auxiliary prediction layers 806, while providing input values ​​105B to the encoder 166 and conditional input 167B to the feature generator 164. The one or more auxiliary prediction layers 806 process the domain-specific data 805 to generate auxiliary prediction data 807.

[0134] Feature generator 164 processes conditional input 167B and auxiliary prediction data 807 to generate feature data 163B. For example, feature layer 404A processes auxiliary prediction data 807 and conditional input 167B to generate feature data 163BA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163BB, and so on.

[0135] In some embodiments, auxiliary prediction data 807 corresponds to an estimate of auxiliary prediction data available at decoder portion 180 that may be used to assist in generating predicted value 195B corresponding to input value 105B.

[0136] refer to Figure 9 , shows an example of a decoder portion 180 of a compression network 140 operable to perform prediction corresponding to reconstruction based on the auxiliary prediction. For example, the decoder portion 180 is configured to process the auxiliary prediction data (e.g., Figure 8The data 165B is encoded at the encoder portion 160 to generate the same value as the input value 105B ( ) corresponds to the predicted value 195B ( ).

[0137] The decoder portion 180 includes or is coupled to one or more auxiliary prediction layers 906. In a particular aspect, the decoder portion 180 is configured to provide domain-specific data 905 to the one or more auxiliary prediction layers 906 while providing the encoded data 165B to the decoder 186 and providing the conditional input 187B to the feature generator 184. The one or more auxiliary prediction layers 906 process the domain-specific data 905 to generate auxiliary prediction data 907.

[0138] Feature generator 184 processes conditional input 187B and auxiliary prediction data 907 to generate feature data 183B. For example, feature layer 504A processes auxiliary prediction data 907 and conditional input 187B to generate feature data 183BA. Feature layer 504B processes the output of feature layer 506A to generate feature data 183BB, and so on.

[0139] In some embodiments, the auxiliary prediction data 907 is the same as the auxiliary prediction data 807. For example, the encoder portion 160 and the decoder portion 180 may access the same auxiliary prediction data. For illustration, the encoder portion 160 and the decoder portion 180 may be included in the same device (e.g., a reference Figure 3 ), obtain the auxiliary prediction data from the same source, or both. In some embodiments, the auxiliary prediction data 907 is a predicted (eg, reconstructed) version, a decoded version, or both of the auxiliary prediction data 807.

[0140] In certain aspects, auxiliary prediction data 907 can be used to assist in generating predicted values ​​195B corresponding to input values ​​105B. In a reconstruction example, auxiliary prediction data 907 can indicate the locations of facial features of a person, and predicted values ​​195B correspond to a reconstructed version of input values ​​105B (e.g., an image frame) representing the person's face. Accessing the locations of facial features can improve the accuracy of the reconstruction. In a collision avoidance example, auxiliary prediction data 907 can indicate a predicted future path of a first vehicle, and predicted values ​​195B can correspond to a collision avoidance output indicating whether a future predicted position of a second vehicle is expected to be within a threshold distance of the future predicted position of an object given the predicted future path of the first vehicle.

[0141] refer to Figure 10, shows an example 1000 of an encoder portion 160 of a compression network operable to perform prediction corresponding to low-latency reconstruction. For example, the encoder portion 160 is configured to generate encoded data 165B corresponding to an input value 105B independently of subsequent values ​​of the one or more input values ​​105 (e.g., before obtaining the subsequent values ​​of the one or more input values ​​105), and the encoded data 165B can be processed at the decoder portion 180 of the compression network 140 independently of the encoded data associated with the subsequent values ​​of the one or more input values ​​105, as described with reference to Figure 11 Further described.

[0142] Figure 1 The conditional input generator 162 generates a conditional input 167B corresponding to the input value 105B independently of subsequent input values ​​of the one or more input values ​​105 (including the input value 105C) (e.g., before obtaining the subsequent input values ​​of the one or more input values ​​105 (including the input value 105C)). For example, the conditional input generator 162 uses the estimator 170 to generate an estimated value 171B based on the predicted value 169A, one or more additional predicted values ​​corresponding to one or more input values ​​preceding the input value 105B in the one or more input values ​​105, or a combination thereof. The conditional input 167B includes the estimated value 171B, the predicted value 169A, the one or more additional predicted values, or a combination thereof. The feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes the input value 105B based on the feature data 163B to generate the encoded data 165B, as shown in FIG. Figure 4 As described above, the encoder portion 160 can generate the encoded data 165B independently of obtaining the input value 105C (e.g., before obtaining the input value 105C). Technical advantages of generating the encoded data 165B independently of the input value 105C can include reduced latency associated with generating the encoded data 165B without having to wait for access to the input value 105C.

[0143] refer to Figure 11 , shows an example 1100 of a decoder portion 180 of a compression network operable to perform prediction corresponding to low-latency reconstruction. For example, the decoder portion 180 is configured to generate a predicted value 195B corresponding to an input value 105B independently of a set of encoded data corresponding to subsequent values ​​of the one or more input values ​​105 (e.g., before obtaining the set of encoded data corresponding to the subsequent values ​​of the one or more input values ​​105).

[0144] Conditional input generator 182 generates conditional input 187B corresponding to input value 105B independently of (e.g., before obtaining) the encoded data corresponding to subsequent input values ​​of one or more input values ​​105 (including encoded data 165C). Conditional input generator 182 obtains predicted value 195A, as shown in FIG. Figure 1 For example, decoder portion 180 processes encoded data 165A to generate predicted value 195A. Conditional input generator 182 uses estimator 188 to generate estimated value 193B based on predicted value 195A, one or more additional predicted values ​​corresponding to one or more input values ​​preceding input value 105B in one or more input values ​​105, or a combination thereof. Conditional input 187B includes estimated value 193B, predicted value 195A, one or more additional predicted values, or a combination thereof.

[0145] The feature generator 184 generates feature data 183B based on the conditional input 187B, and the decoder 186 processes the encoded data 165B based on the feature data 183B, as shown in FIG. Figure 5 As described above, decoder portion 180 can generate predicted value 195B independently of obtaining encoded data 165C (e.g., before obtaining encoded data 165C). Technical advantages of generating predicted value 195B independently of encoded data 165C can include reduced latency associated with generating predicted value 195B without having to wait for access to encoded data 165C.

[0146] refer to Figure 12 , shows an example 1200 of an encoder portion 160 of a compression network 140 operable to perform predictions corresponding to future predictions. For example, the encoder portion 160 is configured to generate a value corresponding to the input value 105B ( ), which may be processed at a decoder portion 180 of the compression network 140 to generate a predicted value 195C corresponding to a predicted future value ( In certain aspects, the input value is 105B ( ) corresponds to the motion value, and the predicted value 195C ( ) corresponds to a predicted future motion value (e.g., a predicted future motion vector).

[0147] Conditional input generator 162 uses local decoder portion 168 to generate predicted value 169B ( For example, the local decoder portion 168 processes the input value 105A ( ) to generate the encoded data 165A associated with the input value 105B ( ) corresponds to the predicted future value of the predicted value 169B. In the specific example, the predicted value 169B ( ) corresponds to an estimate of a predicted future value that can be generated at the decoder portion 180 based on the encoded data 165A.

[0148] In some implementations, conditional input generator 162 generates one or more additional predicted values. The one or more additional predicted values ​​can be associated with at least one input value before input value 105B in one or more input values ​​105, at least one input value after input value 105B in one or more input values ​​105, or both.

[0149] Optionally, in some embodiments, conditional input generator 162 uses estimator 170 to generate input value 105C based on one or more predicted values ​​(eg, predicted value 169B). ) of the estimated value of the predicted future value 171C ( For illustration, the estimated value 171C ( ) corresponds to an input value 105C that may be generated at the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values ​​and independently of the encoded data 165B ( ) is an estimate of the predicted future value of .

[0150] In particular embodiments, estimator 170 uses one or more estimation techniques to determine estimated value 171C. For example, predicted value 169B corresponds to an estimate of a predicted future value of input value 105B, one or more additional predicted values ​​correspond to estimates of predicted future values ​​of other input values ​​of one or more input values ​​105, and estimated value 171C corresponds to an estimated input value subsequent to predicted value 169B. In some embodiments, estimator 170 includes a neural network configured to process the one or more predicted values ​​to generate estimated value 171C.

[0151] Conditional input 167B includes estimated value 171C ( ), one or more predicted values ​​(such as predicted value 169B ( ), one or more additional predicted values, or a combination thereof. The feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes the input value 105B based on the feature data 163B to generate encoded data 165B, as shown in FIG. Figure 4 described.

[0152] refer to Figure 13, shows an example 1300 of a decoder portion 180 of a compression network 140 operable to perform predictions corresponding to future predictions. For example, the decoder portion 180 is configured to process the input value 105B ( ) (as referenced Figure 12 , generated at the encoder portion 160 of the compression network 140) to generate the encoded data 165B to generate the predicted value 195C ( ).

[0153] Conditional input generator 182 obtains predicted value 195B ( For example, decoder portion 180 processes encoded data 165A to generate predicted values ​​195B ( In this particular example, the predicted value 195B corresponds to the input value 105B ( ) is the predicted future value of .

[0154] In some embodiments, conditional input generator 182 obtains one or more additional predicted values. The one or more additional predicted values ​​can be associated with at least one input value before input value 105B in one or more input values ​​105, at least one input value after input value 105B in one or more input values ​​105, or both.

[0155] Optionally, in some embodiments, the conditional input generator 182 uses the estimator 188 to generate the input value 105C based on one or more predicted values ​​(eg, the predicted value 195B). ) of the estimated value of the predicted future value 193C ( For illustration, the estimated value of 193C ( ) and an input value 105C that may be generated at the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values ​​and independently of the encoded data 165B ( ) corresponds to an estimate of the predicted future value of .

[0156] In certain embodiments, estimator 188 uses one or more estimation techniques (similar to the estimation techniques performed at estimator 170, as described in reference to FIG. Figure 12 ) to determine estimated value 193C. For example, predicted value 195B corresponds to an estimate of a predicted future value of input value 105B, one or more additional predicted values ​​correspond to estimates of predicted future values ​​of other input values ​​in one or more input values ​​105, and estimated value 193C corresponds to an estimated input value subsequent to predicted value 195B. In some embodiments, estimator 188 includes a neural network configured to process the one or more predicted values ​​to generate estimated value 193C.

[0157] Conditional input 187B includes estimated value 193C ( ), one or more predicted values ​​(such as predicted value 195B ( ), one or more additional predicted values, or a combination thereof. Feature generator 184 generates feature data 183B based on conditional input 187B, and decoder 186 processes encoded data 165B based on feature data 183B to generate predicted values ​​195C ( ), as referenced Figure 5 described.

[0158] Technical advantages of using compression network 140 to generate predicted future values ​​can include initiating actions based on the predicted future values. For example, based on determining that predicted value 195C corresponds to an alert condition (e.g., a collision), the action can include a preventive action, such as initiating braking. It should be understood that predicted value 195C corresponding to a predicted future value of one of one or more input values ​​105 is provided as an illustrative example. In other examples, predicted value 195C can correspond to a predicted future value (e.g., collision or no collision) associated with input value 105B (e.g., an image frame).

[0159] Figures 14-18 An example of a compression network 140 configured to perform prediction corresponding to image reconstruction is described. The compression network 140 includes a first compression network 140 associated with motion estimation and a second compression network 140 associated with image reconstruction based on motion estimation.

[0160] Figure 14 An example of an encoder portion 160 comprising a first compression network 140 associated with motion estimation. Figure 15 An example of the decoder portion 180 of the first compression network 140 associated with motion estimation is included. Figure 17 An example of the encoder portion 160 of the second compression network 140 associated with motion estimation based image reconstruction is included. Figure 18 An example of the decoder portion 180 of the second compression network 140 associated with motion estimation based image reconstruction is included. Figure 16 Examples of image estimation that may be performed at the encoder portion 160 and the decoder portion 180 of the second compression network 140 in association with motion estimation based image reconstruction are included.

[0161] In some examples, encoder portion 160 of compression network 140 includes first and second compression network 140 encoder portion 160. Similarly, decoder portion 180 of compression network 140 includes first and second compression network 140 decoder portion 180, 180.

[0162] refer to Figure 14 , shows an example 1400 of an encoder portion 160 of a compression network 140 operable to perform prediction corresponding to image reconstruction. For example, the encoder portion 160 is included in a first compression network 140 associated with motion estimation.

[0163] The encoder portion 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The one or more image units 1407 include image unit 1407A, image unit 1407B, image unit 1407C, one or more additional image units, or a combination thereof. As illustrative, non-limiting examples, the image unit 1407 can correspond to a coding unit, a block of pixels, a frame of pixels, an image frame, or a combination thereof.

[0164] The encoder portion 160 is configured to determine one or more motion values ​​1405 associated with one or more image units 1407. For example, the encoder portion 160 is responsive to determining that the image unit 1407B ( ) is to be encoded and determines one or more motion values ​​1405 based on a comparison of image unit 1407B with one or more other image units in one or more image units 1407. For illustration, encoder portion 160 determines motion value 1405A ( ), the motion value 1405B is determined based on the comparison of the image unit 1407C with the image unit 1407B ( ), determine one or more additional motion values, or a combination thereof.

[0165] In certain aspects, the one or more motion values ​​1405 represent motion vectors associated with one or more image units 1407. For example, motion value 1405A represents motion vectors associated with image unit 1407A and image unit 1407B. In certain aspects, the one or more motion values ​​1405 indicate one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, or an angular position associated with one or more image units 1407. For example, motion value 1405A indicates one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position associated with image unit 1407A and image unit 1407B. In certain aspects, the compression network 140 is configured to track an object associated with the one or more motion values ​​1405 across the one or more image units. In certain aspects, the one or more motion values ​​1405 are based on sensor data 226 received from one or more sensors 240, as described with reference to FIG. Figure 2 For example, the motion value 1405A ( ) is based on first sensor data 226 associated with image unit 1407A (eg, captured simultaneously with image unit 1407A) and second sensor data 226 associated with image unit 1407B (eg, captured simultaneously with image unit 1407B).

[0166] Conditional input generator 162 of encoder portion 160 generates conditional input 167B. For example, conditional input generator 162 processes the encoded data associated with image unit 1407A using local decoder portion 168 to generate a first predicted image unit, and processes the encoded data associated with image unit 1407C using local decoder portion 168 to generate a second predicted image unit. In certain aspects, the first predicted image unit ( ) and can be Figure 18 The decoder portion 180 generates the image unit 1407A ( ) corresponds to the predicted estimate of . In certain aspects, the second predicted image unit ( ) and can be Figure 18 The decoder portion 180 generates the image unit 1407C ( ) corresponds to the predicted estimate of ). The conditional input generator 162 is based on the first predicted image unit ( ) and the second predicted image unit ( ) to determine the estimated motion value 1469A ( ) and the estimated motion value 1469B ( ). Conditional input 167B includes estimated motion value 1469A ( ), estimated motion value 1469B ( ) or both.

[0167] Feature generator 164 processes conditional input 167B to generate feature data 163B, as shown in FIG. Figure 4 The encoder 166 processes the motion value 1405A ( )、Movement value 1405B( )、Image unit 1407B( ) and feature data 163B to generate encoded data 165B associated with image unit 1407B. For example, encoder layer 402A processes motion value 1405A ( )、Movement value 1405B( )、Image unit 1407B( ) and feature data 163BA to generate an output. Encoder layer 402B processes the output of encoder layer 402A and feature data 163BB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165B associated with motion estimation (e.g., encoded data 1465B). One or more motion values ​​1405, one or more weights, or a combination thereof may be used to generate an estimated image unit ( ), as referenced Figure 16 Further described.

[0168] refer to Figure 15 , shows an example 1500 of a decoder portion 180 of a compression network 140 operable to perform prediction corresponding to image reconstruction. For example, the decoder portion 180 is included in a first compression network associated with motion estimation.

[0169] The decoder portion 180 is configured to process the set of encoded data 165 to generate one or more predicted motion values ​​1595, one or more weights, or a combination thereof. The one or more predicted motion values ​​1595, one or more weights, or a combination thereof may be used to generate an estimated image unit ( ), as referenced Figure 16 Further described.

[0170] The conditional input generator 182 of the decoder part 180 generates the conditional input 187B. For example, the conditional input generator 182 obtains the conditional input 187B obtained by the decoder part 180 of the second compression network 140 through processing with the image unit 1407A ( ) associated with the coded data to generate the first predicted image unit ( ), as referenced Figure 18As further described. In some aspects, the conditional input generator 182 is obtained by the decoder portion 180 of the second compression network 140 by processing with the image unit 1407C ( ) associated with the coded data to generate the second predicted image unit ( ), as referenced Figure 18 As further described. In certain aspects, image unit 1407A ( )、Image unit 1407C( ) or both and can be decoded independently with image unit 1407B ( ) is associated with the encoded data and is decoded to correspond to the key unit (e.g., I frame).

[0171] The conditional input generator 182 is based on the first predicted image unit ( ) and the second predicted image unit ( ) determines the estimated motion value 1593A ( ) and the estimated motion value 1593B ( ). Conditional input 187B includes estimated motion value 1593A ( ), estimated motion value 1593B ( ) or both.

[0172] Feature generator 184 processes conditional input 187B to generate feature data 183B. For example, feature layer 504A processes estimated motion values ​​1593A ( ), estimated motion value 1593B ( ) or both to generate feature data 183BA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183BB, and so on. Decoder 186 processes encoded data 165B (e.g., encoded data 1465B) and feature data 183B, as shown in FIG. Figure 5 to generate predicted motion values ​​1595A ( ), predicted motion value 1595B ( )、weight 1565( ) or a combination thereof.

[0173] In a specific example, the decoder portion 180 uses the image unit associated with the first prediction ( ) and the second predicted image unit ( ) corresponds to the estimated motion value 1593A ( ) and the estimated motion value 1593B ( ) as a conditional input 187B, and the generated image unit 1407B can be used to generate ) associated with the predicted image unit ( ) of the predicted motion value 1595A ( ), predicted motion value 1595B ( ), and weight 1565 ( ), as referenced Figure 16-18 Further described.

[0174] refer to Figure 16 , shows an example of an image estimator 1600 of a compression network 140 operable to perform prediction corresponding to image reconstruction. For example, the image estimator 1600 is included in the estimator 170 of the encoder portion 160, the estimator 188 of the decoder portion 180, or both, of the second compression network 140 associated with image reconstruction.

[0175] The image estimator 1600 is based on the predicted motion value 1695A ( ) performs prediction of the image unit 1607A ( ) is warped 1602A to generate an estimated image unit 1609A ( ). The image estimator 1600 is based on the predicted motion value 1695B ( ) performs prediction of the image unit 1607C ( ) to generate an estimated image unit 1609B ( ). The image estimator 1600 performs the estimated image unit 1609A ( ) and estimated image unit 1609B ( ) combination 1604 to generate an estimated image unit 1671B ( For example, the estimated image unit 1671B ( ) and the estimated image unit 1609A ( ) and estimated image unit 1609B ( ). To illustrate, image estimator 1600 applies a first weight (eg, weight 1665) to estimated image element 1609A ( ) to generate a first weighted image unit, applying a second weight (eg, 1-weight 1665) to the estimated image unit 1609B ( ) to generate a second weighted image unit, and combining the first weighted image unit and the second image unit to generate an estimated image unit 1671B ( ).

[0176] An illustrative example of image estimator 1600 generating estimated image unit 1671B based on two predicted image units 1607 is provided. In other examples, image estimator 1600 may generate estimated image unit 1671B based on more than two predicted image units 1607. For example, image estimator 1600 may warp one or more additional predicted image units 1607 to generate one or more additional estimated image units, and combine the estimated image units based on various weights to generate estimated image unit 1671B.

[0177] In certain aspects, the encoder portion 160 of the second compression network 140 uses the image estimator 1600 to generate a first estimated image unit ( ), as referenced Figure 17 In certain aspects, decoder portion 180 of second compression network 140 uses image estimator 1600 to generate a second estimated image unit ( ), as referenced Figure 18 Further described.

[0178] refer to Figure 17 , shows an example 1700 of an encoder portion 160 of a compression network 140 operable to perform prediction corresponding to image reconstruction. For example, the encoder portion 160 corresponds to an image encoder portion of the second compression network 140 associated with motion estimation based image reconstruction.

[0179] The encoder portion 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The conditional input generator 162 uses the local decoder portion 168 to generate a predicted image unit 1769A ( For example, the local decoder portion 168 processes the image unit 1407A ( ) associated with the encoded data to generate the predicted image unit 1769A ( In this specific example, the predicted image unit 1769A ( ) corresponds to the Figure 18 Similarly, the conditional input generator 162 uses the local decoder portion 168 to generate an estimate of the predicted image unit associated with the image unit 1407A. For example, the local decoder portion 168 processes the image unit 1407C ( ) associated with the coded data to generate the predicted image unit 1769C ( ).

[0180] Conditional input generator 162 uses estimator 170 (eg, image estimator 1600) to generate a conditional input based on the Figure 14 The encoder portion 160 performs motion estimation to generate motion values ​​1405A ( ) and motion value 1405B ( ) generates estimated image unit 1771B ( For example, the predicted image unit 1769A ( ) corresponds to the predicted image unit 1607A ( ), predicted image unit 1769C ( ) corresponds to the predicted image unit 1607C ( ), movement value 1405A ( ) corresponds to a predicted motion value of 1695A ( ), and the movement value 1405B ( ) corresponds to a predicted motion value of 1695B ( ). The conditional input generator 162 also provides the weight as weight 1665. In certain aspects, the weight is based on configuration settings, default data, user input, or a combination thereof. Estimated image unit 1671B ( ) corresponds to the estimated image unit 1771B ( ). Conditional input 167B includes predicted image unit 1769A ( ), estimated image unit 1771B ( ) and predicted image unit 1769C ( ).

[0181] Feature generator 164 processes conditional input 167B to generate feature data 163B, as shown in FIG. Figure 4 For example, feature layer 404A processes the predicted image unit 1769A ( ), estimated image unit 1771B ( ), and predicted image unit 1769C ( ), to generate feature data 163BA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163BB, and so on. Encoder 166 processes image unit 1407B ( ) and feature data 163B to generate encoded data 165B, as shown in reference Figure 4 For example, the encoder layer 402A processes the image unit 1407B ( ) and feature data 163BA to generate an output. Encoder layer 402B processes the output of encoder layer 402A and feature data 163BB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165B associated with image reconstruction (e.g., encoded data 1765B).

[0182] refer to Figure 18 , shows an example 1800 of a decoder portion 180 of a compression network 140 operable to perform prediction corresponding to image reconstruction. For example, the decoder portion 180 corresponds to an image decoder portion of the second compression network 140 associated with motion estimation based image reconstruction.

[0183] The decoder portion 180 is configured to process a set of encoded data 165 to generate one or more predicted image units 1895. The conditional input generator 182 obtains the predicted image unit 1895A ( For example, the decoder portion 180 processes the image unit 1407A ( ) associated with the encoded data to generate the predicted image unit 1895A ( ). Similarly, the conditional input generator 182 obtains the predicted image unit 1895C ( For example, the decoder portion 180 processes the image unit 1407C ( ) associated with the encoded data to generate the predicted image unit 1895C ( ).

[0184] The conditional input generator 182 uses an estimator 188 (e.g., image estimator 1600) to generate the conditional input based on the Figure 15 The decoder portion 180 performs motion estimation to generate predicted motion values ​​1595A ( ) and the predicted motion value 1595B ( ) generates estimated image unit 1893B ( For example, the predicted image unit 1895A ( ) corresponds to the predicted image unit 1607A ( ), predicted image unit 1895C ( ) corresponds to the predicted image unit 1607C ( ), predicted motion value 1595A ( ) corresponds to a predicted motion value of 1695A ( ), predicted motion value 1595B ( ) corresponds to a predicted motion value of 1695B ( ),and Figure 15 The weight is 1565 ( ) corresponds to weight 1665. Estimated image unit 1671B ( ) corresponds to the estimated image unit 1893B ( ). Conditional input 187B includes predicted image unit 1895A ( ), estimated image unit 1893B ( ) and predicted image unit 1895C ( ).

[0185] Feature generator 184 processes conditional input 187B to generate feature data 183B, as shown in FIG. Figure 5 For example, feature layer 504A processes the predicted image unit 1895A ( ), estimated image unit 1893B ( ), predicted image unit 1895C ( ), to generate feature data 183BA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183BB, and so on. Decoder 186 processes encoded data 165B (e.g., encoded data 1765B) and feature data 183B, as shown in FIG. Figure 5 As described, to generate the predicted image unit 1895B ( Technical advantages of using compression network 140 to process image unit 1407B to generate encoded data 1465B and encoded data 1765B and processing encoded data 1465B and encoded data 1765B to generate predicted image unit 1895B may include a reduced size of encoded data 1465B, encoded data 1765B, or both.

[0186] Figures 19-23 An example of a compression network 140 configured to perform prediction corresponding to low-latency image reconstruction is described. The compression network 140 includes a first compression network 140 associated with low-latency motion estimation and a second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.

[0187] Figure 19 An example of the encoder portion 160 of the first compression network 140 is included in association with low-latency motion estimation. Figure 20 An example of the decoder portion 180 of the first compression network 140 associated with low-latency motion estimation is included. Figure 22 An example of an encoder portion 160 of a second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation is included. Figure 23 An example of the decoder portion 180 of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation is included. Figure 21Examples of image estimation that may be performed at the encoder portion 160 and the decoder portion 180 of the second compression network 140 are included in association with low-latency motion estimation based low-latency image reconstruction.

[0188] In some examples, encoder portion 160 of compression network 140 includes first and second compression network 140 encoder portion 160. Similarly, decoder portion 180 of compression network 140 includes first and second compression network 140 decoder portion 180, 180.

[0189] refer to Figure 19 , Figure 19 is a diagram of schematic aspects of an example of an encoder portion 160 of a compression network 140 operable to perform prediction corresponding to low-latency image reconstruction. For example, the encoder portion 160 is included in a first compression network 140 associated with low-latency motion estimation.

[0190] The encoder portion 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The encoder portion 160 is configured to determine one or more motion values ​​1905 associated with the one or more image units 1407. For example, the encoder portion 160 may be configured to process one or more image units 1407 to generate a set of encoded data 165. ) and determines one or more motion values ​​1905 based on a comparison of picture unit 1407D with one or more other picture units in one or more picture units 1407. For illustration, the encoder portion 160 determines one or more motion values ​​1905 based on a comparison of picture unit 1407D with one or more other picture units in one or more picture units 1407 in response to determining that picture unit 1407D is to be encoded. ) and image unit 1407D ( ) is compared to determine the motion value 1905A ( ), one or more additional motion values ​​are determined based on a comparison of image unit 1407D with one or more image units preceding image unit 1407D in one or more image units 1407, or a combination thereof.

[0191] In certain aspects, the one or more motion values ​​1905 represent motion vectors associated with one or more image units 1407. For example, motion value 1905A ( ) represents a motion vector associated with image unit 1407C and image unit 1407D. In certain aspects, one or more motion values ​​1905 indicate one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position associated with one or more image units 1407. For example, motion value 1905A indicates one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position associated with image unit 1407C and image unit 1407D. In certain aspects, one or more motion values ​​1905 are based on sensor data 226 received from one or more sensors 240, as described with reference to FIG. Figure 2 For example, the motion value 1905A ( ) is based on first sensor data 226 associated with image unit 1407C (eg, captured simultaneously with image unit 1407C) and second sensor data 226 associated with image unit 1407D (eg, captured simultaneously with image unit 1407D).

[0192] Conditional input generator 162 of encoder portion 160 generates conditional input 167D. For example, conditional input generator 162 processes the encoded data associated with image unit 1407A using local decoder portion 168 to generate a first predicted image unit ( ), using local decoder portion 168 to process the encoded data associated with image unit 1407B to generate a second predicted image unit ( ), and processing the encoded data associated with image unit 1407C using local decoder portion 168 to generate a third predicted image unit ( In certain aspects, the first predicted image unit ( ) and can be Figure 23 The decoder portion 180 generates the image unit 1407A ( ) corresponds to the predicted estimate of the second predicted image unit ( ) and the third predicted image unit ( ) to determine the estimated motion value 1969A ( ). The conditional input generator 162 is based on the first predicted image unit ( ) and the second predicted image unit ( ) to determine the estimated motion value 1969B ( ). Conditional input 167D includes estimated motion value 1969A ( ), estimated motion value 1969B ( ) or both.

[0193] Feature generator 164 processes conditional input 167D to generate feature data 163D, as shown in FIG. Figure 4 For example, feature layer 404A processes estimated motion values ​​1969A and estimated motion values ​​1969B to generate feature data 163DA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163DB, and so on. Feature data 163D includes feature data 163DA, feature data 163DB, feature data 163DC, one or more additional sets of feature data including feature data 163DN, or a combination thereof. Encoder 166 processes motion value 1905A ( )、Image unit 1407D( ) and feature data 163D to generate encoded data 165D associated with image unit 1407D. For example, encoder layer 402A processes motion value 1905A ( )、Image unit 1407D( ) and feature data 163DA to generate an output. Encoder layer 402B processes the output of encoder layer 402A and feature data 163DB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165D associated with motion estimation (e.g., encoded data 1965D). One or more motion values ​​1905, one or more weights, or a combination thereof may be used to generate an estimated image unit ( ), as referenced Figure 21 Further described.

[0194] Thus, the encoder portion 160 can generate the encoded data 165D independently of obtaining any image unit after the image unit 1407D in the one or more image units 1407 (e.g., before obtaining any image unit after the image unit 1407D in the one or more image units 1407). Technical advantages of generating the encoded data 165D independently of subsequent image units can include reduced latency associated with generating the encoded data 165D without having to wait for access to subsequent image units.

[0195] refer to Figure 20 , shows an example 2000 of a decoder portion 180 of a compression network 140 operable to perform prediction corresponding to low-latency image reconstruction. For example, the decoder portion 180 is included in a first compression network associated with low-latency motion estimation.

[0196] The decoder portion 180 is configured to process the set of encoded data 165 to generate one or more predicted motion values ​​2095, one or more weights, or a combination thereof. The one or more predicted motion values ​​2095, one or more weights, or a combination thereof may be used to generate an estimated image unit ( ), as referenced Figure 21 Further described.

[0197] The conditional input generator 182 of the decoder part 180 generates the conditional input 187D. For example, the conditional input generator 182 obtains the conditional input 187D obtained by the decoder part 180 of the second compression network 140 through processing with the image unit 1407A ( ) associated with the coded data to generate the first predicted image unit ( ), as referenced Figure 23 Similarly, in some aspects, the conditional input generator 182 obtains the image unit 1407B ( ) and image unit 1407C ( ) associated with the second predicted image unit ( ) and the third predicted image unit ( ), as referenced Figure 23 As further described. In certain aspects, image unit 1407A ( )、Image unit 1407B( )、Image unit 1407C( ) or a combination thereof in the image unit being encoded 1407D ( )Before.

[0198] The conditional input generator 182 is based on the second predicted image unit ( ) and the third predicted image unit ( ) to determine the estimated motion value 2093A ( ). Similarly, the conditional input generator 182 is based on the first predicted image unit ( ) and the second predicted image unit ( ) to determine the estimated motion value 2093B ( ). Conditional input 187D includes estimated motion value 2093A ( ), estimated motion value 2093B ( ) or both.

[0199] Feature generator 184 processes conditional input 187D to generate feature data 183D, as shown in FIG. Figure 5 For example, feature layer 504A processes estimated motion value 2093A ( ), estimated motion value 2093B ( ) or both to generate feature data 183DA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183DB, and so on. Feature data 183D includes feature data 183DA, feature data 183DB, feature data 183DC, one or more additional sets of feature data including feature data 183DN, or a combination thereof. Decoder 186 processes encoded data 165D (e.g., encoded data 1965D) and feature data 183D, as described with reference to FIG. Figure 5 to generate the predicted motion value 2095A ( )、weight 2065( ) or both. For example, decoder layer 508N processes the encoded data 165D and feature data 183DN to generate an output. Decoder layer 508B processes the output of decoder layer 508C and feature data 183DB to generate an output, and so on. The output of decoder layer 508A corresponds to the predicted motion value 2095A ( associated with image unit 1407D )、weight 2065( ) or both.

[0200] In a specific example, the decoder portion 180 uses the estimated motion value 2093A ( ) and the estimated motion value 2093B ( ) as a conditional input 187D, and generates a predicted motion value 2095A ( ) and weight 2065 ( ), which can be used to generate an image with unit 1407D ( ) associated with the predicted image unit ( ), as referenced Figure 21-23 As further described. The predicted motion value 2095A is generated independently of the coded data associated with any picture unit following the picture unit 1407D ( ) technical advantages may include generating predicted motion values ​​2095A ( ) associated with reduced latency.

[0201] refer to Figure 21 , shows an example of an image estimator 2100 of a compression network 140 operable to perform prediction corresponding to low-latency image reconstruction. For example, the image estimator 2100 is included in the estimator 170 of the encoder portion 160, the estimator 188 of the decoder portion 180, or both, of the second compression network 140 associated with low-latency image reconstruction.

[0202] The image estimator 2100 is based on the predicted motion value 2195A ( ) performs prediction of the image unit 2107C ( ) of the warp 2102 to generate an estimated image unit 2109A ( ). The image estimator 2100 uses the predicted image unit 2107C ( ) as the estimated image unit 2109B ( ). The image estimator 2100 performs the estimated image unit 2109A ( ) and estimated image unit 2109B ( ) combination 2104 to generate an estimated image unit 2171D ( For example, the estimated image unit 2171D ( ) and the estimated image unit 2109A ( ) and estimated image unit 2109B ( ). To illustrate, the image estimator 2100 applies a first weight (eg, weight 2165) to the estimated image element 2109A ( ) to generate a first weighted image unit, applying a second weight (eg, 1-weight 2165) to the estimated image unit 2109B ( ) to generate a second weighted image unit, and combining the first weighted image unit and the second image unit to generate an estimated image unit 2171D ( ).

[0203] As an illustrative example, the image estimator 2100 is provided to generate an estimated image unit 2171D based on a single predicted image unit 2107. In other examples, the image estimator 2100 may generate an estimated image unit 2171D based on multiple predicted image units 2107. For example, the image estimator 2100 may warp one or more additional predicted image units 2107 to generate one or more additional estimated image units, and combine the estimated image units based on various weights to generate the estimated image unit 2171D.

[0204] In certain aspects, the encoder portion 160 of the second compression network 140 uses the image estimator 2100 to generate a first estimated image unit ( ), as referenced Figure 22 As further described. In certain aspects, the decoder portion 180 of the second compression network 140 uses the image estimator 2100 to generate a second estimated image unit ( ), as referenced Figure 23 Further described.

[0205] refer to Figure 22 , an example 2200 of an encoder portion 160 of a compression network 140 operable to perform prediction corresponding to low-latency image reconstruction is shown. For example, the encoder portion 160 corresponds to an image encoder portion of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.

[0206] The encoder portion 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The conditional input generator 162 uses the local decoder portion 168 to generate a predicted image unit 2269C ( For example, the local decoder portion 168 processes the image unit 1407C ( ) associated with the encoded data to generate the predicted image unit 2269C ( In this specific example, the predicted image unit 2269C ( ) corresponds to the Figure 23 The decoder portion 180 generates an estimate of the predicted picture unit associated with the picture unit 1407C.

[0207] The conditional input generator 162 uses the estimator 170 (e.g., the image estimator 2100) to generate the conditional input based on the Figure 19 The motion estimation performed at the encoder portion 160 generates motion values ​​1905A ( ) generates estimated image unit 2271D ( For example, the predicted image unit 2269C ( ) corresponds to the predicted image unit 2107C ( ) and movement value 1905A ( ) corresponds to a predicted motion value of 2195A ( ). The conditional input generator 162 also provides the weight as weight 2165. In certain aspects, the weight is based on configuration settings, default data, user input, or a combination thereof. Estimated image unit 2171D ( ) corresponds to the estimated image unit 2271D ( ). The conditional input 167D associated with image unit 1407D includes the predicted image unit 2269C ( ) and estimated image unit 2271D ( ).

[0208] Feature generator 164 processes conditional input 167D to generate feature data 163D, as shown in FIG. Figure 4For example, feature layer 404A processes conditional input 167D to generate feature data 163DA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163DB, and so on. Feature data 163D includes feature data 163DA, feature data 163DB, feature data 163DC, one or more additional sets of feature data including feature data 163DN, or a combination thereof. Encoder 166 processes image unit 1407D ( ) to generate encoded data 165D. For example, encoder layer 402A processes image unit 1407D ( ) and feature data 163DA to generate an output. Encoder layer 402B processes the output of encoder layer 402A and feature data 163DB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165D associated with low-latency image reconstruction (e.g., encoded data 2265D).

[0209] Thus, the encoder portion 160 can generate the encoded data 165D independently of obtaining any image unit after the image unit 1407D in the one or more image units 1407 (e.g., before obtaining any image unit after the image unit 1407D in the one or more image units 1407). Technical advantages of generating the encoded data 165D independently of subsequent image units can include reduced latency associated with generating the encoded data 165D without having to wait for access to subsequent image units.

[0210] refer to Figure 23 , shows an example 2300 of a decoder portion 180 of a predictive compression network 140 operable to perform corresponding low-latency image reconstruction. For example, the decoder portion 180 corresponds to an image decoder portion of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.

[0211] The decoder portion 180 is configured to process a set of encoded data 165 to generate one or more predicted image units 2395. The conditional input generator 182 obtains the predicted image unit 2395C ( For example, the decoder portion 180 processes the image unit 1407C ( ) associated with the encoded data to generate the predicted image unit 2395C ( ).

[0212] The conditional input generator 182 uses an estimator 188 (e.g., image estimator 2100) to generate the conditional input based on the Figure 20 The decoder portion 180 performs motion estimation to generate predicted motion values ​​2095A ( ) generates an estimated image unit 2393D ( For example, the predicted image unit 2395C ( ) corresponds to the predicted image unit 2107C ( ), predicted motion value 2095A ( ) corresponds to a predicted motion value of 2195A ( ),as well as Figure 20 The weight is 2065 ( ) corresponds to weight 2165. Estimated image unit 2171D ( ) corresponds to the estimated image unit 2393D ( ). The conditional input 187D associated with image unit 1407D includes the predicted image unit 2395C ( ) and the estimated image unit 2393D ( ).

[0213] Feature generator 184 processes conditional input 187D to generate feature data 183D, as shown in FIG. Figure 5 For example, feature layer 504A processes the predicted image unit 2395C ( ) and the estimated image unit 2393D ( ) to generate feature data 183DA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183DB, and so on. Feature data 183D includes feature data 183DA, feature data 183DB, feature data 183DC, one or more additional sets of feature data including feature data 183DN, or a combination thereof. Decoder 186 processes encoded data 165B (e.g., encoded data 2265D) and feature data 183D, as described with reference to FIG. Figure 5 The unit 2395D ( Technical advantages of using compression network 140 may include a reduced size of encoded data 1965D, encoded data 2265D, or both. The predicted image unit 2395D is generated independently of the encoded data of any subsequent image units of one or more image units 1407 ( ) can include generating a predicted image unit 2395D ( ) associated with reduced latency.

[0214] Figure 24 An embodiment 2400 of an integrated circuit 2402 including one or more processors 2490 is depicted. The one or more processors 2490 include a compression network component 2460, such as Figure 1160, the decoder portion 180, or both of the compression network 140. Integrated circuit 2402 also includes a signal input 2404, such as one or more bus interfaces, to enable input data 2428 to be received for processing. Integrated circuit 2402 also includes a signal output 2406, such as a bus interface, to enable output data 2450 to be transmitted. For example, in embodiments where compression network component 2460 includes encoder portion 160, input data 2428 may include input value 105, and output data 2450 may include a set of encoded data 165. In embodiments where compression network component 2460 includes decoder portion 180, input data 2428 may include a set of encoded data 165, and output data 2450 may include predicted value 195. In embodiments where compression network component 2460 includes both encoder portion 160 and decoder portion 180, input data 2428 may include input value 105, and output data 2450 may include predicted value 195. Integrated circuit 2402 enables implementation of at least a portion of compression network 140 as a component of various systems such as Figure 25 The depicted mobile phone or tablet computer, Figure 26 The head-mounted equipment depicted, Figure 27 The wearable electronic devices depicted, Figure 28 The voice-controlled speaker system depicted, Figure 29 The camera depicted, Figure 30 Virtual reality, mixed reality, or augmented reality headsets depicted, or Figure 31 or Figure 32 components in the vehicle depicted).

[0215] As an illustrative, non-limiting example, Figure 25An embodiment 2500 is depicted in which a compression network component 2460 is implemented in a mobile device 2502, such as a phone or tablet computer. Mobile device 2502 includes one or more microphones 2510, one or more speakers 2520, and a display screen 2504. Furthermore, compression network component 2460 is integrated into mobile device 2502 and is illustrated using dashed lines to indicate internal components that are generally not visible to the user of mobile device 2502. In a specific example, compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, compression network component 2460 operates to encode video data captured by a camera of mobile device 2502 for transmission to other devices, decode video data received from other devices, or a combination thereof. In another illustrative, non-limiting example, compression network component 2460 operates to encode data for storage at mobile device 2502 and decode data when retrieved from a storage device.

[0216] Figure 26 An embodiment 2600 is depicted in which a compression network component 2460 is implemented in a head-mounted equipment device 2602. The head-mounted equipment device 2602 includes a microphone 2610 and a camera 2620. In a specific example, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, the compression network component 2460 operates to encode data (such as audio data captured by the microphone 2610, video data captured by the camera 2620, and / or head tracking data (e.g., sensor data from an IMU of the head-mounted equipment device 2602)) for transmission to other devices, decode data (such as audio data received from other devices for playback on headphone speakers of the head-mounted equipment device 2602), or a combination thereof. In another illustrative, non-limiting example, the compression network component 2460 operates to encode data for storage at the head-mounted equipment device 2602 and decode data when retrieved from a storage device.

[0217] Figure 27An embodiment 2700 is depicted in which a compression network component 2460 is implemented in a wearable electronic device 2702 (illustrated as a "smartwatch"). Compression network component 2460, microphone 2710, speaker 2720, and display 2704 are integrated into wearable electronic device 2702. In certain examples, compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, compression network component 2460 operates to encode data (such as audio data captured by microphone 2710) for transmission to another device; alternatively, compression network component 2460 operates to decode data, such as audio data received from another device for playback at speaker 2720 or video data for playback at display 2704, or both. In another illustrative, non-limiting example, compression network component 2460 operates to encode data for storage at wearable electronic device 2702 and decode the data when retrieved from the storage device.

[0218] In a specific example, the wearable electronic device 2702 includes a haptic device that provides a tactile notification (e.g., a vibration) in response to detecting activity associated with the operation of the compression network component 2460. For example, the tactile notification can cause a user viewing the wearable electronic device 2702 to see a displayed notification that data (e.g., audio or video data) has been received from a remote device and is available for playback at the wearable electronic device 2702. Thus, the wearable electronic device 2702 can alert a user with a hearing impairment or a user wearing head-mounted equipment to such a notification.

[0219] Figure 282800 is an embodiment in which a compression network component 2460 is implemented in a wireless speaker and voice-activated device 2802. The wireless speaker and voice-activated device 2802 may have a wireless network connection and be configured to perform auxiliary operations. One or more processors 2890, including the compression network component 2460, a microphone 2810, a camera 2820, a speaker 2804, or a combination thereof, are included in the wireless speaker and voice-activated device 2802. In certain examples, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In illustrative, non-limiting examples, the compression network component 2460 operates to encode video data captured by the camera 2820 for transmission to other devices, decode audio data for playback at the speaker 2804, decode video data received from other devices for playback on a display screen (not shown) of the wireless speaker and voice-activated device 2802, or a combination thereof. In another illustrative, non-limiting example, the compression network component 2460 operates to encode data for storage at the wireless speaker and voice-activated device 2802 and decode the data when retrieved from the storage device.

[0220] During operation, the wireless speaker and voice-activated device 2802 can perform auxiliary operations, such as those performed by a voice-activated system (e.g., an integrated auxiliary application), in response to receiving a verbal command from a user via the microphone 2810. The auxiliary operations can include adjusting the temperature, playing music, turning on lights, etc. For example, the auxiliary operations can include initiating the transmission of data to a remote device, receiving data from a remote device, and / or storing / retrieving data from the local memory of the wireless speaker and voice-activated device 2802, each of which is performed more efficiently due to the operation of the compressed network component 2460.

[0221] Figure 29An embodiment 2900 is depicted in which a compression network component 2460 is implemented in a portable electronic device corresponding to a camera device 2902. The compression network component 2460 and a microphone 2910 are integrated into the camera device 2902. In a specific example, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, the compression network component 2460 operates to encode data (such as image data, video data, audio data, or a combination thereof) captured at the camera device 2902 for transmission to another device. In another illustrative, non-limiting example, the compression network component 2460 operates to encode data (such as video or image data captured at the camera device 2902) for storage in a memory of the camera device 2902 and also to decode the data when retrieved from the memory.

[0222] Figure 30 An embodiment 3000 is depicted in which a compression network component 2460 is implemented in a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality head-mounted device 3002. The compression network component 2460, a microphone 3010, and a camera 3020 are integrated into the head-mounted device 3002. A visual interface device, such as a display screen, is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while wearing the head-mounted device 3002. In a specific example, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In this illustrative, non-limiting example, the compression network component 2460 operates to decode data received at the head-mounted device 3002, such as video data, audio data, other data associated with providing a virtual, augmented, or mixed reality experience for the user of the head-mounted device 3002, or a combination thereof. In an illustrative, non-limiting example, the compression network component 2460 operates to encode data captured at the head-mounted equipment 3002, such as head tracking data from an IMU of the head-mounted equipment 3002, for transmission to a remote device (e.g., a virtual reality (VR) session server). In another illustrative, non-limiting example, the compression network component 2460 operates to encode data, such as video, audio, or image data captured at or received by the head-mounted equipment 3002, for storage in a memory of the head-mounted equipment 3002, and also operates to decode the data when retrieved from the memory.

[0223] Figure 31An embodiment 3100 is depicted in which the compression network component 2460 corresponds to or is integrated within a vehicle 3102, which is shown as a manned or unmanned aerial vehicle (e.g., a package delivery drone). The compression network component 2460, microphone 3110, speaker 3120, camera 3104, or a combination thereof are integrated into the vehicle 3102. User voice activity detection, such as for delivery instructions from an authorized user of the vehicle 3102, can be performed based on audio signals received from the microphone 3110 of the vehicle 3102.

[0224] In a specific example, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, the compression network component 2460 operates to encode video data captured by the camera 3104 of the vehicle 3102 for transmission to other devices, decode video data received from other devices, or a combination thereof. In another illustrative, non-limiting example, the compression network component 2460 operates to encode data for storage at the vehicle 3102 and decode the data when retrieved from the storage device.

[0225] Figure 32 Another embodiment 3200 is depicted in which the compression network component 2460 corresponds to or is integrated within a vehicle 3202, which is shown as an automobile. The vehicle 3202 includes one or more processors 290, one or more processors 292, one or more processors 390, or a combination thereof, including the compression network component 2460. The vehicle 3202 also includes a microphone 3210, a camera 3204, or both.

[0226] User voice activity detection can be performed based on audio signals received from microphone 3210 of vehicle 3202. In some embodiments, user voice activity detection can be performed based on audio signals received from an internal microphone (e.g., microphone 3210), such as for voice commands from authorized passengers. For example, user voice activity detection can be used to detect voice commands from the operator of vehicle 3202 (e.g., from a parent to set the volume to 5 or set the destination of the self-driving vehicle) and ignore the voice of another passenger (e.g., from a child to set the volume to 10 or from other passengers discussing other locations). In some embodiments, user voice activity detection can be performed based on audio signals received from an external microphone (e.g., microphone 3210), such as for authorized users of the vehicle. In particular embodiments, in response to receiving a verbal command recognized as user speech, the voice activation system initiates one or more operations of the vehicle 3202 based on one or more detected keywords (e.g., "unlock," "start engine," "play music," "show weather forecast," or other voice commands), such as by providing feedback or information via the display 3220 or one or more speakers (e.g., speaker 3230).

[0227] In a specific example, the compression network component 2460 operates to improve decoding efficiency to reduce the amount of resources used for transmission or storage of encoded data. In an illustrative, non-limiting example, the compression network component 2460 operates to encode video data captured by the camera 3204 of the vehicle 3202 for transmission to other devices, decode video data received from other devices, or a combination thereof. In another illustrative, non-limiting example, the compression network component 2460 operates to encode data for storage at the vehicle 3202 and decode the data when retrieved from the storage device.

[0228] The camera 3204 can capture one or more image frames while the vehicle 3202 is operating. The compression network component 2460 can process one or more motion values ​​(e.g., one or more input values ​​105) associated with the one or more image frames to generate a set of encoded data 165. The one or more motion values ​​can include one or more image frames, one or more speed measurements, one or more acceleration measurements, other sensor data, a navigation route of the vehicle 3202, or a combination thereof. The vehicle 3202 can store the encoded data 165 at the vehicle 3202, transmit the set of encoded data 165 to another device (e.g., a server), or both.

[0229] In some embodiments, the compression network component 2460 generates one or more predicted values ​​195 based on the set of encoded data 165. In a collision avoidance example, the compression network component 2460 can track an object in one or more image frames. The compression network component 2460 generates one or more predicted values ​​195 corresponding to predicted future motion values. For example, the one or more predicted values ​​195 indicate whether the predicted future position of the vehicle 3202 is within a distance threshold of the predicted future position of the object. In some embodiments, the compression network component 2460 initiates one or more collision avoidance actions, such as generating an alert, initiating braking, activating an alarm, sending an alert to an emergency vehicle, or a combination thereof, in response to determining that the predicted future position of the vehicle 3202 is within a distance threshold of the predicted future position of the object.

[0230] refer to Figure 33 , illustrates a specific implementation of a method 3300 for generating encoded data using a compression network. In certain aspects, one or more operations of the method 3300 are performed by at least one of the following: Figure 1 Conditional input generator 162, feature generator 164, encoder 166, local decoder part 168, estimator 170, encoder part 160, compression network 140, Figure 2 One or more processors 290, device 202, system 200, Figure 3 One or more processors 390, device 302, system 300, Figure 24 compression network component 2460, or a combination thereof.

[0231] Method 3300 includes obtaining a conditional input to a compression network at 3302, wherein the conditional input is based on one or more first predicted motion values. For example, Figure 1 The conditional input generator 162 obtains the conditional input 167B of the compression network 140, as shown in FIG. Figure 1 The conditional input 167B is based on the predicted value 169A (eg, the predicted motion value).

[0232] Method 3300 also includes processing the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values ​​at 3304. For example, encoder portion 160 processes conditional input 167B using feature generator 164 of compression network 140 to generate feature data 163B, and processes feature data 163B and input value 105B using encoder 166 of compression network 140 to generate encoded data 165B associated with input value 105B, as described with reference to FIG. Figure 1 described.

[0233] Thus, method 3300 enables reducing the amount of information to be provided as encoded data 165A to decoder portion 180 by generating encoded data 165A based on an estimate (e.g., conditional input 167B) of information that can be generated at decoder portion 180 (e.g., conditional input 187B).

[0234] Figure 33 The method 3300 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a digital signal processor (DSP), a controller, other hardware devices, firmware devices, or any combination thereof. As an example, Figure 33 The method 3300 may be performed by a processor executing instructions, such as reference Figure 35 described.

[0235] refer to Figure 34 , illustrates a specific implementation of a method 3400 for generating predicted data from encoded data using a compression network. In certain aspects, one or more operations of the method 3400 are performed by at least one of the following: Figure 1 Conditional input generator 182, feature generator 184, decoder 186, estimator 188, decoder part 180, compression network 140, Figure 2 One or more processors 292, device 260, system 200, Figure 3 One or more processors 390, device 302, system 300, Figure 24 compression network component 2460, or a combination thereof.

[0236] Method 3400 includes obtaining encoded data associated with one or more motion values ​​at 3402. For example, decoder portion 180 obtains encoded data 165B associated with input value 105B from encoder portion 160, as shown in FIG. Figure 1 described.

[0237] The method 3400 also includes obtaining a conditional input to the compression network at 3404, wherein the conditional input is based on the one or more first predicted motion values. For example, the conditional input generator 182 obtains the conditional input 187B of the compression network 140, as shown in FIG. Figure 1 The conditional input 187B is based on the predicted value 195A.

[0238] Method 3400 also includes processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values ​​at 3406. For example, decoder portion 180 processes conditional input 187B using feature generator 184 of compression network 140 to generate feature data 183B, and processes feature data 183B and encoded data 165B using decoder 186 of compression network 140 to generate predicted values ​​195B, as described with reference to FIG. Figure 1 described.

[0239] Thus, method 3400 enables reducing the amount of information to be obtained by decoder portion 180 as encoded data 165A by generating predicted value 195B based on information (e.g., conditional input 167B) that can be estimated at encoder portion 160 (e.g., conditional input 187B).

[0240] Figure 34 The method 3400 may be implemented by an FPGA device, an ASIC, a processing unit (such as a CPU), a DSP, a controller, other hardware devices, a firmware device, or any combination thereof. As an example, Figure 34 The method 3400 may be performed by a processor executing instructions, such as reference Figure 35 described.

[0241] refer to Figure 35 , a block diagram of a particular exemplary embodiment of a device is depicted and generally designated 3500. In various embodiments, the device 3500 may have Figure 35 In an exemplary embodiment, the device 3500 may be configured with more or fewer components than those shown. Figure 2 Device 202, device 260 and Figure 3 In an exemplary embodiment, the device 3500 may execute the reference Figure 1-Figure 34 Describes one or more operations.

[0242] In certain embodiments, device 3500 includes a processor 3506 (e.g., a CPU). Device 3500 may include one or more additional processors 3510 (e.g., one or more DSPs). In certain aspects, Figure 2 One or more processors 290, one or more processors 292, Figure 3The one or more processors 390, or a combination thereof, may correspond to the processor 3506, the processor 3510, or a combination thereof. The processor 3510 may include a speech and music coder-decoder (CODEC) 3508, which includes a speech coder ("vocoder") encoder 3536, a vocoder decoder 3538, or both. The processor 3510 may include the compression network component 2460, one or more applications 262, or a combination thereof.

[0243] The device 3500 may include a memory 3586 and a CODEC 3534. The memory 3586 may include instructions 3556 that may be executed by one or more additional processors 3510 (or processor 3506) to implement the functionality described with reference to the compression network component 2460. The device 3500 may include a modem 3570 coupled to an antenna 3552 via a transceiver 3550. In certain aspects, the modem 3570 corresponds to Figure 2 modem 270, modem 280, or both.

[0244] Device 3500 may include a display 3528 coupled to a display controller 3526. One or more speakers 3520, one or more microphones 3524, or a combination thereof may be coupled to a CODEC 3534. CODEC 3534 may include a digital-to-analog converter (DAC) 3502, an analog-to-digital converter (ADC) 3504, or both. In particular embodiments, CODEC 3534 may receive an analog signal from one or more microphones 3524, convert the analog signal to a digital signal using ADC 3504, and provide the digital signal to a voice and music decoder 3508. Voice and music decoder 3508 may process the digital signal, and the digital signal may be further processed by compression network component 2460, one or more applications 262, or a combination thereof. In particular embodiments, voice and music decoder 3508 may provide the digital signal to CODEC 3534. The CODEC 3534 may convert the digital signal into an analog signal using the digital-to-analog converter 3502 and may provide the analog signal to one or more speakers 3520 .

[0245] In particular embodiments, device 3500 may be included in a system-in-package or system-on-chip device 3522. In particular embodiments, memory 3586, processor 3506, processor 3510, display controller 3526, CODEC 3534, and modem 3570 are included in the system-in-package or system-on-chip device 3522. In particular embodiments, input device 3530, one or more sensors 240, and power supply 3544 are coupled to the system-in-package or system-on-chip device 3522. Additionally, in particular embodiments, as shown in FIG. Figure 35 As shown, the display 3528, input device 3530, one or more speakers 3520, one or more microphones 3524, one or more sensors 240, antenna 3552, and power supply 3544 are external to the system-in-package or system-on-chip device 3522. In particular embodiments, each of the display 3528, input device 3530, one or more speakers 3520, one or more microphones 3524, one or more sensors 240, antenna 3552, and power supply 3544 may be coupled to a component of the system-in-package or system-on-chip device 3522, such as an interface or controller.

[0246] Device 3500 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop, a computer, a tablet computer, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a VR device, an extended reality (XR) device, a base station, a mobile device, or any combination thereof.

[0247] In conjunction with the described embodiments, the apparatus comprises means for obtaining encoded data associated with one or more motion values. For example, the means for obtaining the encoded data may correspond to Figure 1 The decoder part 180, the encoder part 160, the compression network 140, Figure 2 modem 280, one or more processors 292, device 260, system 200, Figure 3The storage device 392, the device 302, the system 300, the processor 3506, the processor 3510, the antenna 3552, the transceiver 3550, the modem 3570, one or more other circuits or components configured to obtain encoded data associated with one or more motion values, or any combination thereof.

[0248] The apparatus further comprises means for obtaining a conditional input to the compression network, wherein the conditional input is based on one or more first predicted motion values. For example, the means for obtaining the conditional input may correspond to Figure 1 Estimator 188, conditional input generator 182, decoder part 180, compression network 140, Figure 2 one or more processors 292, device 260, system 200, Figure 3 Device 302, system 300, Figure 16 Image estimator 1600, Figure 21 The image estimator 2100, the processor 3506, the processor 3510, one or more other circuits or components configured to obtain the conditional input, or any combination thereof.

[0249] The apparatus further includes means for processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values. For example, the means for processing the encoded data and the conditional input may correspond to Figure 1 The feature generator 184, the decoder 186, the decoder part 180, the compression network 140, Figure 2 one or more processors 292, device 260, system 200, Figure 3 The device 302, the system 300, the processor 3506, the processor 3510, one or more other circuits or components configured to process the encoded data and the conditional input, or any combination thereof.

[0250] Also in conjunction with the embodiment described above, the apparatus includes means for obtaining a conditional input to the compression network, wherein the conditional input is based on one or more first predicted motion values. For example, the means for obtaining the conditional input may correspond to Figure 1 The local decoder part 168, the estimator 170, the conditional input generator 162, the encoder part 160, the compression network 140, Figure 2 One or more processors 290, device 202, system 200, Figure 3 Device 302, system 300, Figure 16 Image estimator 1600, Figure 21 The image estimator 2100, the processor 3506, the processor 3510, one or more other circuits or components configured to obtain the conditional input, or any combination thereof.

[0251] The apparatus further includes means for processing the conditioned input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values. For example, the means for processing the conditioned input and the one or more motion values ​​may correspond to Figure 1 The feature generator 164, the encoder 166, the encoder part 160, the compression network 140, Figure 2 One or more processors 290, device 202, system 200, Figure 3 The device 302, the system 300, the processor 3506, the processor 3510, one or more other circuits or components configured to process the conditional input and the one or more motion values, or any combination thereof.

[0252] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 3586) includes instructions (e.g., instructions 3556) that, when executed by one or more processors (e.g., one or more processors 3510 or processor 3506), cause the one or more processors to obtain encoded data (e.g., encoded data 165B) associated with one or more motion values ​​(e.g., input values ​​105B). The instructions, when executed by the one or more processors, further cause the one or more processors to obtain a conditional input (e.g., conditional input 187B) to a compression network (e.g., compression network 140), where the conditional input is based on one or more first predicted motion values ​​(e.g., predicted values ​​195A). The instructions, when executed by the one or more processors, further cause the one or more processors to process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values ​​(e.g., predicted values ​​195B).

[0253] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 3586) includes instructions (e.g., instructions 3556) that, when executed by one or more processors (e.g., one or more processors 3510 or processor 3506), cause the one or more processors to obtain a conditional input (e.g., conditional input 167B) for a compression network (e.g., compression network 140), where the conditional input is based on one or more first predicted motion values ​​(e.g., predicted values ​​169A). The instructions, when executed by the one or more processors, further cause the one or more processors to process the conditional input and one or more motion values ​​(e.g., input values ​​105B) using the compression network to generate encoded data associated with the one or more motion values ​​(e.g., encoded data 165B).

[0254] Certain aspects of the disclosure are described below in a collection of interrelated examples:

[0255] According to Example 1, a device includes: one or more processors configured to obtain encoded data associated with one or more motion values; obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0256] Example 2 includes the device of Example 1, wherein the one or more motion values ​​are based on outputs of one or more sensors.

[0257] Example 3 includes the device of Example 1 or Example 2, wherein the one or more sensors include an inertial measurement unit (IMU).

[0258] Example 4 includes the apparatus of any of Examples 1 to 3, wherein the one or more motion values ​​represent one or more motion vectors associated with one or more image units.

[0259] Example 5 includes the apparatus of Example 4, wherein an image unit of the one or more image units comprises a decoding unit.

[0260] Example 6 includes the apparatus of Example 4 or Example 5, wherein an image unit of the one or more image units comprises a block of pixels.

[0261] Example 7 includes the apparatus of any of Examples 4 to 6, wherein an image unit of the one or more image units comprises a frame of pixels.

[0262] Example 8 includes the apparatus of any of Examples 1 to 7, wherein the one or more second predicted motion values ​​represent future motion vectors.

[0263] Example 9 includes the apparatus of any of Examples 1 to 8, wherein the one or more second predicted motion values ​​correspond to a reconstructed version of the one or more motion values.

[0264] Example 10 includes the device of any of Examples 1 to 9, wherein the one or more processors are integrated into at least one of a head-mounted device, a mobile communication device, an extended reality (XR) device, or a vehicle.

[0265] Example 11 includes the apparatus of any of Examples 1 to 10, wherein the one or more motion values ​​indicate one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position.

[0266] Example 12 includes the apparatus of any of Examples 1 to 11, wherein the compression network comprises a neural network having a plurality of layers.

[0267] Example 13 includes the apparatus of any of Examples 1 to 12, wherein the compression network includes a video decoder, and wherein the video decoder has multiple decoder layers configured to decode multiple levels of resolution of the encoded data associated with the one or more motion values.

[0268] Example 14 includes the device of any of Examples 1 to 13, wherein the one or more processors are configured to track an object associated with the one or more motion values ​​across one or more frames of pixels.

[0269] Example 15 includes the apparatus of Example 14, wherein the one or more second predicted motion values ​​represent a collision avoidance output associated with the vehicle.

[0270] Example 16 includes the apparatus of Example 15, wherein the collision avoidance output indicates a predicted future position of the vehicle relative to the object.

[0271] Example 17 includes the apparatus of example 15 or example 16, wherein the collision avoidance output indicates a predicted future position of the vehicle and a predicted future position of the object.

[0272] Example 18 includes the apparatus of any of Examples 1 to 17, further comprising a modem configured to receive a bitstream from the encoder device, wherein the bitstream includes encoded data.

[0273] Example 19 includes the apparatus of any of Examples 1 to 18, wherein the one or more processors are configured to process the encoded data and the feature data to generate one or more second predicted motion values, and wherein the feature data is based on a conditional input.

[0274] Example 20 includes the device of Example 19, further comprising a modem configured to receive a bitstream from the encoder device, wherein the bitstream includes the feature data and the encoded data.

[0275] Example 21 includes the apparatus of Example 19 or Example 20, wherein the one or more processors are configured to process the conditional input using a compression network to generate the feature data.

[0276] Example 22 includes the apparatus of any of Examples 19 to 21, wherein the feature data corresponds to multi-scale feature data having different spatial resolutions.

[0277] Example 23 includes the apparatus of any of Examples 19 to 22, wherein the feature data comprises multi-scale wavelet transform data.

[0278] Example 24 includes the apparatus of any of Examples 1 to 23, wherein the encoded data includes data associated with an image unit (e.g., ) associated motion estimation coded data, and wherein the one or more processors are configured to: based on a reconstructed previous image unit (e.g., a ) and subsequent image units reconstructed (e.g., c ) to determine the estimated motion value (e.g., 、 ); processing the estimated motion values ​​to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation coded data and the motion estimation feature data using the compression network to generate one or more second predicted motion values, wherein the one or more second predicted motion values ​​include one or more reconstructed motion values ​​corresponding to the one or more motion values ​​(e.g., 、 ).

[0279] Example 25 includes the apparatus of Example 24, wherein the encoded data comprises reconstruction encoded data, and wherein the one or more processors are configured to: 、 ) generates estimated image units (e.g., ); processing the estimated image unit to generate reconstruction feature data corresponding to the image unit; and processing the reconstructed encoded data and the reconstruction feature data using the compression network to generate a reconstructed image unit (e.g., b ).

[0280] Example 26 includes the apparatus of Example 25, wherein the one or more first predicted motion values ​​include first reconstructed motion values ​​(e.g., ) and the second reconstructed motion value (e.g., ), and wherein the one or more processors are configured to: based on the first reconstructed motion value (eg, ) is applied to the previous image unit of the reconstruction (e.g., a ) to determine a first estimated version of the image unit; based on the second reconstructed motion value (eg, ) is applied to the subsequent image units of the reconstruction (e.g., c) to determine a second estimated version of the image unit; and generating an estimated image unit based on a combination of the first estimated version of the image unit and the second estimated version of the image unit (eg, ).

[0281] Example 27 includes the apparatus of any of Examples 1 to 23, wherein the encoded data includes data associated with an image unit (e.g., ) associated motion estimation coded data, and wherein the one or more processors are configured to: based on a first reconstructed previous image unit (e.g., c ) and the second reconstructed previous image unit (e.g., b ) to determine a first estimated motion value (eg, ); processing at least the first estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation coded data and the motion estimation feature data using the compression network to generate one or more second predicted motion values, wherein the one or more second predicted motion values ​​include one or more reconstructed motion values ​​corresponding to the one or more motion values ​​(e.g., ).

[0282] Example 28 includes the apparatus of Example 27, wherein the encoded data comprises reconstructed encoded data, and wherein the one or more processors are configured to: ) generates estimated image units (e.g., ); processing the estimated image unit to generate reconstructed feature data corresponding to the image unit; and processing the reconstructed encoded data and the reconstructed feature data using the compression network to generate a reconstructed image unit (e.g., d ).

[0283] Example 29 includes the apparatus of Example 28, wherein the one or more processors are configured to generate a second reconstructed image based on a previous image unit (e.g., b ) and the third reconstructed previous image unit (e.g., a ) to determine a second estimated motion value (eg, ), wherein the motion estimation feature data corresponding to the image unit is further based on processing the second estimated motion value.

[0284] Example 30 includes the apparatus of Example 28 or Example 29, wherein the one or more first predicted motion values ​​include first reconstructed motion values ​​(e.g., ), and wherein the one or more processors are configured to generate a first reconstructed motion value (eg, ) is applied to the previous image unit of the first reconstruction (e.g., c ) to determine the estimated image unit (e.g., ), where the reconstructed feature data is based on the estimated image units (e.g., ).

[0285] According to Example 31, a method includes: obtaining encoded data associated with one or more motion values ​​at a device; obtaining a conditional input to a compression network at the device, wherein the conditional input is based on one or more first predicted motion values; and processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0286] Example 32 includes the method of Example 31, wherein the one or more motion values ​​are based on outputs of one or more sensors.

[0287] Example 33 includes the method of Example 31 or Example 32, wherein the one or more sensors include an inertial measurement unit (IMU).

[0288] Example 34 includes the method of any one of Examples 31 to 33, wherein the one or more motion values ​​represent one or more motion vectors associated with one or more image units.

[0289] Example 35 includes the method of Example 34, wherein an image unit of the one or more image units comprises a decoding unit.

[0290] Example 36 includes the method of Example 34 or Example 35, wherein an image unit of the one or more image units comprises a block of pixels.

[0291] Example 37 includes the method of any of Examples 34 to 36, wherein an image unit in the one or more image units comprises a frame of pixels.

[0292] Example 38 includes the method of any one of Examples 31 to 37, wherein the one or more second predicted motion values ​​represent future motion vectors.

[0293] Example 39 includes the method of any one of Examples 31 to 38, wherein the one or more second predicted motion values ​​correspond to a reconstructed version of the one or more motion values.

[0294] Example 40 includes the method of any one of Examples 31 to 39, wherein the device comprises at least one of a head-mounted device, a mobile communication device, an extended reality (XR) device, or a vehicle.

[0295] Example 41 includes the method of any of Examples 31 to 40, wherein the one or more motion values ​​indicate one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position.

[0296] Example 42 includes the method of any of Examples 31 to 41, wherein the compression network comprises a neural network having a plurality of layers.

[0297] Example 43 includes the method of any of Examples 31 to 42, wherein the compression network includes a video decoder, and wherein the video decoder has multiple decoder layers configured to decode multiple levels of resolution of the encoded data associated with the one or more motion values.

[0298] Example 44 includes the method of any one of Examples 31 to 43, further comprising tracking an object associated with the one or more motion values ​​across one or more frames of pixels.

[0299] Example 45 includes the method of Example 44, wherein the one or more second predicted motion values ​​represent a collision avoidance output associated with the vehicle.

[0300] Example 46 includes the method of Example 45, wherein the collision avoidance output indicates a predicted future position of the vehicle relative to the object.

[0301] Example 47 includes the method of Example 45 or Example 46, wherein the collision avoidance output indicates a predicted future position of the vehicle and a predicted future position of the object.

[0302] Example 48 includes the method of any one of Examples 31 to 47, further comprising receiving a bitstream from an encoder device via a modem, wherein the bitstream includes encoded data.

[0303] Example 49 includes the method of any one of Examples 31 to 48, further comprising processing the encoded data and the feature data to generate one or more second predicted motion values, wherein the feature data is based on the conditional input.

[0304] Example 50 includes the method of Example 49, further comprising receiving a bitstream from an encoder device via a modem, wherein the bitstream includes the feature data and the encoded data.

[0305] Example 51 includes the method of Example 49 or Example 50, further comprising processing the conditional input using a compression network to generate feature data.

[0306] Example 52 includes the method of any one of Examples 31 to 51, wherein processing the encoded data and the conditional input comprises: processing the conditional input using a compression network to generate feature data; and processing the encoded data and the feature data to generate the one or more second predicted motion values.

[0307] Example 53 includes the method of any one of Examples 49 to 52, wherein the feature data corresponds to multi-scale feature data having different spatial resolutions.

[0308] Example 54 includes the method of any one of Examples 49 to 53, wherein the feature data includes multi-scale wavelet transform data.

[0309] Example 55 includes the method of any one of Examples 31 to 54, further comprising: based on the reconstructed previous image unit (e.g., a ) and subsequent image units reconstructed (e.g., c ) to determine the estimated motion value (e.g., ), wherein the encoded data includes image units (e.g., ) associated with the image unit; processing the estimated motion values ​​to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation coded data and the motion estimation feature data using the compression network to generate one or more second predicted motion values, wherein the one or more second predicted motion values ​​include one or more reconstructed motion values ​​corresponding to the one or more motion values ​​(e.g., ).

[0310] Example 56 includes the method of Example 55, further comprising: reconstructing the image based on the one or more reconstructed motion values ​​(e.g., ) generates estimated image units (e.g., ); processing the estimated image unit to generate reconstructed feature data corresponding to the image unit; and processing the reconstructed encoded data and the reconstructed feature data using the compression network to generate a reconstructed image unit (e.g., b ), wherein the encoded data includes reconstructed encoded data.

[0311] Example 57 includes the method of Example 56, further comprising: based on reconstructing the first reconstructed motion value (e.g., ) is applied to the previous image unit of the reconstruction (e.g., a ) to determine a first estimated version of the image unit, wherein the one or more first predicted motion values ​​include a first reconstructed motion value (eg, ); Based on the second reconstructed motion value (eg, ) is applied to the subsequent image units of the reconstruction (e.g., c ) to determine a second estimated version of the image unit, wherein the one or more first predicted motion values ​​include a second reconstructed motion value (eg, ); and generating an estimated image unit based on a combination of the first estimated version of the image unit and the second estimated version of the image unit (eg, ).

[0312] Example 58 includes the method of any of Examples 31 to 54, further comprising: reconstructing a previous image unit (e.g., c ) and the second reconstructed previous image unit (e.g., b ) to determine the first estimated motion value ( ), wherein the encoded data includes image units (e.g., ) associated with the image unit; processing at least the first estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation coded data and the motion estimation feature data using the compression network to generate one or more second predicted motion values, wherein the one or more second predicted motion values ​​include one or more reconstructed motion values ​​corresponding to the one or more motion values ​​(e.g., ).

[0313] Example 59 includes the method of Example 58, further comprising: reconstructing the image based on the one or more reconstructed motion values ​​(e.g., ) generates estimated image units (e.g., ); processing the estimated image unit to generate reconstructed feature data corresponding to the image unit; and processing the reconstructed encoded data and the reconstructed feature data using the compression network to generate a reconstructed image unit (e.g., ), wherein the encoded data includes reconstructed encoded data.

[0314] Example 60 includes the method of Example 59, further comprising: reconstructing a previous image unit (e.g., b ) and the third reconstructed previous image unit (e.g., a ) is compared to determine the second estimated motion value ( ), wherein the motion estimation feature data corresponding to the image unit is further based on processing the second estimated motion value.

[0315] Example 61 includes the method of Example 59 or Example 60, further comprising: ) is applied to the previous image unit of the first reconstruction (e.g., c ) to determine the estimated image unit (e.g., ), wherein the one or more first predicted motion values ​​include first reconstructed motion values ​​(eg, ), and wherein the reconstructed feature data is based on the estimated image units (e.g., ).

[0316] According to Example 62, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any one of Examples 31 to 61.

[0317] According to Example 63, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any one of Examples 31 to 61.

[0318] According to Example 64, an apparatus includes components for implementing the method of any one of Examples 31 to 61.

[0319] According to Example 65, a device includes: one or more processors configured to obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and processing the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0320] Example 66 includes the device of Example 65, wherein the one or more motion values ​​are based on outputs of one or more sensors.

[0321] Example 67 includes the device of Example 66, wherein the one or more sensors include an inertial measurement unit (IMU).

[0322] Example 68 includes the apparatus of any of Examples 65 to 67, wherein the one or more motion values ​​represent one or more motion vectors associated with one or more image units.

[0323] Example 69 includes the apparatus of Example 68, wherein an image unit of the one or more image units comprises a decoding unit.

[0324] Example 70 includes the apparatus of Example 68 or Example 69, wherein an image unit of the one or more image units comprises a block of pixels.

[0325] Example 71 includes the apparatus of any of Examples 68 to 70, wherein an image unit in the one or more image units comprises a frame of pixels.

[0326] Example 72 includes the apparatus of any of Examples 65 to 71, further comprising a modem configured to send a bit stream to the decoder device, wherein the bit stream includes the encoded data.

[0327] Example 73 includes the device of any of Examples 65 to 72, wherein the one or more processors are integrated into at least one of a head-mounted device, a mobile communication device, an extended reality (XR) device, and a vehicle.

[0328] Example 74 includes the apparatus of any of Examples 65 to 73, wherein the one or more motion values ​​indicate one or more of a linear velocity, a linear acceleration, a linear position, an angular velocity, an angular acceleration, and an angular position.

[0329] Example 75 includes the apparatus of any of Examples 65 to 74, wherein the compression network comprises a neural network having a plurality of layers.

[0330] Example 76 includes the apparatus of any of Examples 65 to 75, wherein the compression network includes a video encoder, and wherein the video encoder has multiple encoder layers configured to encode higher-order resolutions of the video data associated with the one or more motion values.

[0331] Example 77 includes the apparatus of any of Examples 65 to 76, further comprising a modem configured to send a bit stream to the decoder device, wherein the bit stream includes the encoded data.

[0332] Example 78 includes the apparatus of any of Examples 65 to 77, wherein the one or more processors are configured to: process the conditional input using a compression network to generate feature data; and process the one or more motion values ​​and the feature data using the compression network to generate the encoded data.

[0333] Example 79 includes the device of Example 78, further comprising a modem configured to send a bit stream to the decoder device, wherein the bit stream includes the characteristic data and the encoded data.

[0334] Example 80 includes the apparatus of Example 78 or Example 79, wherein the feature data comprises multi-scale feature data having different spatial resolutions.

[0335] Example 81 includes the apparatus of any of Examples 78 to 80, wherein the feature data comprises multi-scale wavelet transform data.

[0336] Example 82 includes the apparatus of any of Examples 65 to 81, wherein the one or more processors are configured to: b ) and the previous image unit (e.g., a ) to determine a first motion value of one or more motion values ​​(eg, ); based on the image unit and subsequent image units (for example, c ) to determine a second motion value (eg, ); based on the reconstruction of the previous image unit (e.g., a ) and subsequent image units reconstructed (e.g., c ) to determine the estimated motion value (e.g., 、 ); processing the estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation feature data, the first motion value, the second motion value, and the image unit using a compression network to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes the motion estimation encoded data.

[0337] Example 83 includes the apparatus of Example 82, wherein the one or more first predicted motion values ​​include first reconstructed motion values ​​(e.g., ) and the second reconstructed motion value (e.g., ), and wherein the one or more processors are configured to: based on the first reconstructed motion value (eg, ) is applied to the previous image unit of the reconstruction (e.g., a ) to determine a first estimated version of the image unit; based on the second reconstructed motion value (eg, ) is applied to the subsequent image units of the reconstruction (e.g., c ) to determine a second estimated version of the image unit; generating an estimated image unit based on a combination of the first estimated version of the image unit and the second estimated version of the image unit (eg, ); process the estimated image units (e.g., ) to generate reconstructed feature data; and processing the image unit and the reconstructed feature data using a compression network to generate reconstructed encoded data associated with the image unit, wherein the encoded data includes the reconstructed encoded data.

[0338] Example 84 includes the apparatus of any of Examples 65 to 81, wherein the one or more processors are configured to: d ) and the previous image unit (e.g., c ) to determine a motion value of one or more motion values ​​(e.g., ); Based on the previous image unit of the first reconstruction (eg, c ) with the previous element of the second reconstruction (e.g., b ) to determine a first estimated motion value (eg, ); processing at least the first estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation feature data, the motion value, and the image unit using a compression network to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes motion estimation encoded data.

[0339] Example 85 includes the apparatus of Example 84, wherein the one or more processors are configured to generate a second reconstruction of the previous element (e.g., b ) with the previous elements of the third reconstruction (e.g., a ) to determine a second estimated motion value (eg, ), wherein the motion estimation feature data corresponding to the image unit is further based on processing the second estimated motion value.

[0340] Example 86 includes the apparatus of Example 84 or Example 85, wherein the one or more first predicted motion values ​​include first reconstructed motion values ​​(e.g., ), and wherein the one or more processors are configured to: based on the first reconstructed motion value (eg, ) is applied to the previous image unit of the first reconstruction (e.g., c ) to determine the estimated image unit (e.g., ); process the estimated image units (e.g., ) to generate reconstructed feature data; and processing the image unit and the reconstructed feature data using a compression network to generate reconstructed encoded data associated with the image unit, wherein the encoded data includes the reconstructed encoded data.

[0341] According to Example 87, a method includes: obtaining a conditional input to a compression network at a device, wherein the conditional input is based on one or more first predicted motion values; and processing the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0342] Example 88 includes the method of Example 87, further comprising processing the conditional input using a compression network at the device to generate feature data; and processing the one or more motion values ​​and the feature data using the compression network to generate the encoded data.

[0343] Example 89 includes the method of Example 87 or Example 88, wherein the feature data includes multi-scale feature data having different spatial resolutions.

[0344] Example 90 includes the method of any of Examples 87 to 89, wherein the one or more motion values ​​are based on outputs of one or more sensors.

[0345] Example 91 includes the method of Example 90, wherein the one or more sensors include an inertial measurement unit (IMU).

[0346] Example 92 includes the method of any one of Examples 87 to 91, wherein the one or more motion values ​​represent one or more motion vectors associated with one or more image units.

[0347] Example 93 includes the method of Example 92, wherein an image unit of the one or more image units comprises a decoding unit.

[0348] Example 94 includes the method of Example 92 or Example 93, wherein an image unit in the one or more image units comprises a block of pixels.

[0349] Example 95 includes the method of any of Examples 92 to 94, wherein an image unit in the one or more image units comprises a frame of pixels.

[0350] Example 96 includes the method of any one of Examples 87 to 95, further comprising sending a bitstream to a decoder device via a modem, wherein the bitstream includes the encoded data.

[0351] Example 97 includes the method of any one of Examples 87 to 96, wherein the device comprises at least one of a head-mounted device, a mobile communication device, an extended reality (XR) device, and a vehicle.

[0352] Example 98 includes the method of any of Examples 87 to 97, wherein the one or more motion values ​​indicate one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, and angular position.

[0353] Example 99 includes the method of any of Examples 87 to 98, wherein the compression network comprises a neural network having a plurality of layers.

[0354] Example 100 includes the method of any of Examples 97 to 99, wherein the compression network includes a video encoder, and wherein the video encoder has multiple encoder layers configured to encode higher-order resolutions of the video data associated with the one or more motion values.

[0355] Example 101 includes the method of any one of Examples 87 to 100, further comprising sending a bitstream to a decoder device via a modem, wherein the bitstream includes the encoded data.

[0356] Example 102 includes the method of any one of Examples 87 to 101, further comprising: processing the conditional input using a compression network to generate feature data; and processing the one or more motion values ​​and the feature data using the compression network to generate encoded data.

[0357] Example 103 includes the method of Example 102, further comprising sending a bitstream to a decoder device via a modem, wherein the bitstream includes the feature data and the encoded data.

[0358] Example 104 includes the method of Example 102 or Example 103, wherein the feature data includes multi-scale feature data having different spatial resolutions.

[0359] Example 105 includes the method of any one of Examples 102 to 104, wherein the feature data includes multi-scale wavelet transform data.

[0360] Example 106 includes the method of any of Examples 87 to 105, comprising: based on an image unit (e.g., b ) and the previous image unit (e.g., a ) to determine a first motion value of one or more motion values ​​(eg, ); based on the image unit and subsequent image units (for example, c ) to determine a second motion value of the one or more motion values ​​(eg, ); based on the reconstruction of the previous image unit (e.g., a ) and subsequent image units reconstructed (e.g., c ) to determine the estimated motion value (e.g., 、 ); processing the estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation feature data, the first motion value, the second motion value, and the image unit using a compression network to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes the motion estimation encoded data.

[0361] Example 107 includes the method of Example 106, further comprising: based on reconstructing the first reconstructed motion value (e.g., ) is applied to the previous image unit of the reconstruction (e.g., a ) to determine a first estimated version of the image unit, wherein the one or more first predicted motion values ​​include a first reconstructed motion value (eg, ); Based on the second reconstructed motion value (eg, ) is applied to the subsequent image units of the reconstruction (e.g., c ) to determine a second estimated version of the image unit, wherein the one or more first predicted motion values ​​include a second reconstructed motion value (eg, ); generating an estimated image unit based on a combination of the first estimated version of the image unit and the second estimated version of the image unit (eg, ); process the estimated image units (e.g., ) to generate reconstructed feature data; and processing the image unit and the reconstructed feature data using a compression network to generate reconstructed encoded data associated with the image unit, wherein the encoded data includes the reconstructed encoded data.

[0362] Example 108 includes the method of any of Examples 87 to 105, further comprising: d ) and the previous image unit (e.g., c ) to determine a motion value of the one or more motion values ​​(eg, ); Based on the previous image unit of the first reconstruction (eg, c ) and the previous elements of the second reconstruction (e.g., b ) to determine a first estimated motion value (eg, ); processing at least the first estimated motion value to generate motion estimation feature data corresponding to the image unit; and processing the motion estimation feature data, the motion value, and the image unit using a compression network to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes motion estimation encoded data.

[0363] Example 109 includes the method of Example 108, further comprising: b ) with the previous elements of the third reconstruction (e.g., a ) to determine a second estimated motion value (eg, ), wherein the motion estimation feature data corresponding to the image unit is further based on processing the second estimated motion value.

[0364] Example 110 includes the method of Example 108 or Example 109, further comprising: based on reconstructing the first reconstructed motion value (e.g., ) is applied to the previous image unit of the first reconstruction (e.g., c ) to determine the estimated image unit (e.g., ), wherein the one or more first predicted motion values ​​include a first reconstructed motion value (eg, ); process the estimated image units (e.g., ) to generate reconstructed feature data; and processing the image unit and the reconstructed feature data using a compression network to generate reconstructed encoded data associated with the image unit, wherein the encoded data includes the reconstructed encoded data.

[0365] According to Example 111, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any one of Examples 87 to 110.

[0366] According to Example 112, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any one of Examples 87 to 110.

[0367] According to Example 113, an apparatus includes components for implementing the method of any of Examples 87 to 110.

[0368] According to example 114, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to: obtain encoded data associated with one or more motion values; obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0369] According to Example 115, an apparatus includes: components for obtaining encoded data associated with one or more motion values; components for obtaining a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and components for processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.

[0370] According to example 116, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to: obtain a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and process the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0371] According to Example 117, an apparatus includes: means for obtaining a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and means for processing the conditional input and the one or more motion values ​​using the compression network to generate encoded data associated with the one or more motion values.

[0372] Those skilled in the art will further appreciate that the various schematic logic blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of the two. Various schematic components, blocks, configurations, modules, circuits, and steps have been generally described above with respect to their functions. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functions in different ways for each specific application, and such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure.

[0373] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disk read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. In an alternative, the processor and storage medium may reside as discrete components in the computing device or user terminal.

[0374] The previous description of the disclosed aspects is provided to enable those skilled in the art to implement or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be consistent with the broadest possible scope consistent with the principles and novel features defined by the appended claims.

Claims

1. A device comprising: One or more processors configured to: obtaining encoded data associated with one or more motion values; obtaining a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; as well as The encoded data and the conditional input are processed using the compression network to generate one or more second predicted motion values.

2. The device according to claim 1, wherein The one or more motion values ​​are based on output of one or more sensors.

3. The device according to claim 2, wherein The one or more sensors include an inertial measurement unit (IMU).

4. The apparatus according to claim 1, wherein The one or more motion values ​​represent one or more motion vectors associated with one or more image units.

5. The device according to claim 4, wherein An image unit of the one or more image units comprises a decoding unit.

6. The device according to claim 4, wherein An image unit of the one or more image units comprises a block of pixels.

7. The apparatus according to claim 4, wherein An image unit of the one or more image units comprises a frame of pixels.

8. The apparatus according to claim 1, wherein The one or more second predicted motion values ​​represent future motion vectors.

9. The apparatus according to claim 1, wherein The one or more second predicted motion values ​​correspond to reconstructed versions of the one or more motion values.

10. The apparatus according to claim 1, wherein The one or more processors are integrated into at least one of a head-mounted device, a mobile communication device, an extended reality (XR) device, or a vehicle.

11. The apparatus according to claim 1, wherein The one or more motion values ​​indicate one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.

12. The apparatus according to claim 1, wherein The compression network includes a neural network having a plurality of layers.

13. The apparatus according to claim 1, wherein The compression network includes a video decoder, and wherein the video decoder has a plurality of decoder layers configured to decode multiple resolutions of encoded data associated with the one or more motion values.

14. The apparatus according to claim 1, wherein The one or more processors are configured to track an object associated with the one or more motion values ​​across one or more frames of pixels.

15. The apparatus according to claim 14, wherein The one or more second predicted motion values ​​represent collision avoidance outputs associated with the vehicle.

16. The apparatus according to claim 15, wherein The collision avoidance output indicates a predicted future position of the vehicle relative to the object.

17. The apparatus according to claim 15, wherein The collision avoidance output indicates a predicted future position of the vehicle and a predicted future position of the object.

18. The apparatus of claim 1, further comprising a modem configured to receive the bitstream from the encoder device, wherein The bitstream includes the encoded data.

19. A method comprising: obtaining, at a device, encoded data associated with one or more motion values; obtaining, at the device, a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; as well as The encoded data and the conditional input are processed using the compression network to generate one or more second predicted motion values.

20. The method according to claim 19, wherein Processing the encoded data and the conditional input includes: Processing the conditional input using the compression network to generate feature data; and The encoded data and the feature data are processed to generate the one or more second predicted motion values.

21. The method according to claim 20, wherein The feature data corresponds to multi-scale feature data with different spatial resolutions.

22. The method according to claim 20, wherein The feature data includes multi-scale wavelet transform data.

23. A device comprising: One or more processors configured to: obtaining a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and The conditional input and one or more motion values ​​are processed using the compression network to generate encoded data associated with the one or more motion values.

24. The apparatus according to claim 23, wherein The one or more motion values ​​are based on output of one or more sensors.

25. The apparatus of claim 24, wherein: The one or more sensors include an inertial measurement unit (IMU).

26. The apparatus of claim 23, wherein: The one or more motion values ​​represent one or more motion vectors associated with one or more image units.

27. The device of claim 23, further comprising a modem configured to send a bit stream to the decoder device, wherein The bitstream includes the encoded data.

28. A method comprising: obtaining, at a device, a conditional input to a compression network, wherein the conditional input is based on one or more first predicted motion values; and The conditional input and one or more motion values ​​are processed using the compression network to generate encoded data associated with the one or more motion values.

29. The method according to claim 28, further comprising: processing the conditional input using the compression network at the device to generate feature data; as well as The one or more motion values ​​and the feature data are processed using the compression network to generate the encoded data.

30. The method according to claim 29, wherein The feature data includes multi-scale feature data with different spatial resolutions.