Prediction using a compressed network
The compression network efficiently processes and transmits data by generating encoded and predicted values, addressing resource constraints in computing devices through reduced data transmission and storage requirements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2026-03-25
AI Technical Summary
Computing devices face challenges in efficiently processing and transmitting large amounts of data, leading to resource constraints such as memory and bandwidth usage.
A compression network is utilized to process input values, generating encoded data and predicted values, where the encoder portion processes input values to generate encoded data, and the decoder portion processes this encoded data along with an estimate of decoder conditional input to generate predicted values, reducing the amount of information needed for accurate prediction.
This approach reduces resource usage by minimizing the amount of data required for accurate prediction, thereby optimizing memory and bandwidth consumption.
Smart Images

Figure 2026509831000001_ABST
Abstract
Description
[Technical Field]
[0001] (Cross-reference of related applications)
[0001] This application claims the benefit of priority from U.S. Nonprovisional Patent Application No. 18 / 183,867, filed on 14 March 2023 and owned by the same applicant, the entirety of which is expressly incorporated herein by reference.
[0002]
[0002] This disclosure generally relates to performing predictions using a compression network. [Background technology]
[0003]
[0003] Technological advancements have led to smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless phones such as mobile phones and smartphones, tablets and laptop computers, which are small and lightweight and easily carried by users. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functions such as digital still cameras, digital video cameras, digital recorders, and audio file players. Also, such devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Thus, these devices can contain considerable computing power.
[0004]
[0004] Such computing devices often incorporate functions for processing large amounts of data. By compressing data before storing or transmitting it, resources such as memory and bandwidth can be saved. For example, a computing device can generate an encoded version of an image frame that uses fewer bits than the original image frame. Techniques to reduce the size of compressed data can save resources even further. [Overview of the Initiative]
[0005]
[0005] According to one implementation of the present disclosure, the device includes one or more processors configured to acquire encoded data associated with one or more motion values. The one or more processors are also configured to acquire a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. The one or more processors are further configured to use the compression network to process the encoded data and the conditional input to generate one or more second predicted motion values.
[0006]
[0006] According to another implementation of the present disclosure, the method includes, in a device, obtaining encoded data associated with one or more motion values. The method also includes, in a device, obtaining a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. The method also includes, using the compression network, processing the encoded data and the conditional input to generate one or more second predicted motion values.
[0007]
[0007] According to another implementation of the present disclosure, the non-temporary computer-readable medium includes an instruction, which, when executed by one or more processors, causes one or more processors to acquire encoded data associated with one or more motion values. The instruction, when executed by one or more processors, also causes one or more processors to acquire a conditional input to a compressed network, which is based on one or more first predicted motion values. The instruction further, when executed by one or more processors, causes one or more processors to use the compressed network to process the encoded data and the conditional input to generate one or more second predicted motion values.
[0008]
[0008] According to another implementation of the present disclosure, the apparatus includes means for acquiring encoded data associated with one or more motion values. The apparatus also includes means for acquiring a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. The apparatus further includes means for using the compression network to process the encoded data and the conditional input to generate one or more second predicted motion values.
[0009]
[0009] According to another implementation of the present disclosure, the device includes one or more processors configured to acquire a conditional input to a compressed network, the conditional input being based on one or more first predicted motion values. The one or more processors are further configured to use the compressed network to process the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0010]
[0010] According to another implementation of the present disclosure, the method comprises, in a device, obtaining a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. The method also comprises, using the compression network, processing the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0011]
[0011] According to another implementation of the present disclosure, the non-temporary computer-readable medium includes instructions, which, when executed by one or more processors, cause one or more processors to acquire a conditional input to a compressed network, the conditional input being based on one or more first predicted motion values. The instructions, when executed by one or more processors, also cause one or more processors to use the compressed network to process the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0012]
[0012] According to another implementation of the present disclosure, the device includes means for acquiring a conditional input to a compressed network, the conditional input being based on one or more first predicted motion values. The device also includes means for using the compressed network to process the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0013]
[0013] After reviewing the following sections, namely the brief description of the drawings, the modes for carrying out the invention, and the claims, other aspects, advantages, and features of the present disclosure will become apparent. [Brief explanation of the drawing]
[0014] [Figure 1]
[0014] This is a block diagram of a particular exemplary embodiment of a compression network capable of performing prediction, according to some embodiments of the present disclosure. [Figure 2]
[0015] This figure illustrates an exemplary embodiment of a system capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 3]
[0016] This figure illustrates an exemplary embodiment of a system capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 4]
[0017] This figure illustrates an exemplary embodiment of one embodiment of an encoder portion of a compression network that is operable to perform predictions, according to some embodiments of the present disclosure. [Figure 5]
[0018] This figure illustrates an exemplary embodiment of one embodiment of a decoder portion of a compression network that is operable to perform predictions, according to some embodiments of the present disclosure. [Figure 6]
[0019] This figure illustrates an exemplary embodiment of one embodiment of an encoder portion of a compression network that is operable to perform predictions corresponding to reconstruction, according to some embodiments of the present disclosure. [Figure 7]
[0020] This figure illustrates an exemplary embodiment of one embodiment of a decoder portion of a compression network that is operable to perform predictions corresponding to reconstruction, according to some embodiments of the present disclosure. [Figure 8]
[0021] This figure illustrates an exemplary embodiment of one embodiment of an encoder portion of a compression network that is operable to perform predictions corresponding to reconstruction based on auxiliary predictions, according to some embodiments of the present disclosure. [Figure 9]
[0022] This figure illustrates an exemplary embodiment of one embodiment of a decoder portion of a compression network that is operable to perform predictions corresponding to reconstruction based on auxiliary predictions, according to some embodiments of the present disclosure. [Figure 10]
[0023] This figure illustrates an exemplary embodiment of one embodiment of the encoder portion of a compression network that is operable to perform predictions corresponding to low latency reconstruction, according to some embodiments of the present disclosure. [Figure 11]
[0024] This figure illustrates an exemplary embodiment of one embodiment of the decoder portion of a compression network that is operable to perform predictions corresponding to low latency reconstruction, according to some embodiments of the present disclosure. [Figure 12]
[0025] This figure illustrates an exemplary embodiment of one embodiment of an encoder portion of a compression network that is operable to perform predictions corresponding to future predictions, according to some embodiments of the present disclosure. [Figure 13]
[0026] This figure illustrates an exemplary embodiment of one embodiment of a decoder portion of a compression network that is operable to perform predictions corresponding to future predictions, according to some embodiments of the present disclosure. [Figure 14]
[0027] This figure illustrates an exemplary embodiment of one embodiment of a motion estimation encoder portion of a compressed network that can operate to perform predictions corresponding to image reconstruction, according to some embodiments of the present disclosure. [Figure 15]
[0028] This figure shows an exemplary embodiment of one embodiment of a motion estimation decoder portion of a compressed network that can operate to perform predictions corresponding to image reconstruction, according to some embodiments of the present disclosure. [Figure 16]
[0029] This figure illustrates an exemplary embodiment of one example of a compressed network image estimator capable of performing predictions corresponding to image reconstruction, according to some embodiments of the present disclosure. [Figure 17]
[0030] This figure illustrates an exemplary embodiment of an image encoder portion of a compression network that can operate to perform predictions corresponding to image reconstruction, according to some embodiments of the present disclosure. [Figure 18]
[0031] This figure illustrates an exemplary embodiment of an image decoder portion of a compression network that can operate to perform predictions corresponding to image reconstruction, according to some embodiments of the present disclosure. [Figure 19]
[0032] This figure illustrates an exemplary embodiment of one embodiment of a motion estimation encoder portion of a compressed network capable of performing predictions corresponding to low-latency image reconstruction, according to some embodiments of the present disclosure. [Figure 20]
[0033] This figure shows an exemplary embodiment of one embodiment of a motion estimation decoder portion of a compressed network capable of performing predictions corresponding to low-latency image reconstruction, according to some embodiments of the present disclosure. [Figure 21]
[0034] This figure illustrates an exemplary embodiment of one example of a compressed network image estimator capable of performing predictions corresponding to low-latency image reconstruction, according to some embodiments of the present disclosure. [Figure 22]
[0035] This figure illustrates an exemplary embodiment of one embodiment of the image encoder portion of a compression network capable of performing predictions corresponding to low-latency image reconstruction, according to some embodiments of the present disclosure. [Figure 23]
[0036] This figure illustrates an exemplary embodiment of one embodiment of the image decoder portion of a compression network capable of performing predictions corresponding to low-latency image reconstruction, according to some embodiments of the present disclosure. [Figure 24]
[0037] Figure 24 shows one embodiment of an integrated circuit that can operate to perform predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 25]
[0038] Figure 25 shows a mobile device capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 26]
[0039] Figure 26 shows a headset that can operate to perform predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 27]
[0040] Figure 27 shows a wearable electronic device, according to some embodiments of the present disclosure, that is capable of performing predictions using a compressed network. [Figure 28]
[0041] Figure 28 shows a voice-controlled speaker system, according to some embodiments of the present disclosure, that can operate to perform predictions using a compressed network. [Figure 29]
[0042] Figure 29 shows a camera capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 30]
[0043] Figure 30 shows a headset, such as a virtual reality headset, a mixed reality headset, or an augmented reality headset, that is capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 31]
[0044] Figure 31 is a diagram of a first embodiment of a vehicle capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 32]
[0045] This is a diagram of a second embodiment of a vehicle capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Figure 33]
[0046] Figure 33 illustrates a specific implementation of a method for generating encoded data using the compression network of Figure 1, according to some embodiments of the present disclosure. [Figure 34]
[0047] Figure 34 illustrates specific implementations of a method for generating predictive data from encoded data using the compression network of Figure 1, according to some embodiments of the present disclosure. [Figure 35]
[0048] This is a block diagram of a specific exemplary embodiment of a device capable of performing predictions using a compressed network, according to some embodiments of the present disclosure. [Modes for carrying out the invention]
[0015]
[0049] Computing devices often incorporate features for processing large amounts of data. Compressing data before storage or transmission can save resources such as memory and bandwidth. For example, a computing device can generate an encoded version of an image frame that uses fewer bits than the original image frame. Techniques to reduce the size of compressed data can further save resources. Compressed data can be processed to generate predictive data. For example, predictive data could represent a reconfigured version of an image frame, predicted future image frames in a sequence of images containing the image frame, the classification of the image frame, other types of data associated with the image frame, or a combination of these.
[0016]
[0050] A system and method for performing predictions using a compressed network are disclosed. For example, the compressed network includes an encoder portion and a decoder portion. The encoder portion is configured to process the input values and encoder conditional input of a sequence of input values to generate encoded data. The encoded data corresponds to a compressed version (having fewer bits) of the input values. The decoder portion is configured to process the encoded data and decoder conditional input to generate predicted values associated with the input values.
[0017]
[0051] The encoder conditional input is an estimate of the decoder conditional input. For example, the decoder conditional input is based on the previous predicted value generated in the decoder section, the local decoder section of the encoder section is used to generate an estimate of the previous predicted value, and the encoder conditional input is based on the estimate of the previous predicted value. By generating encoded data based on an estimate of the information available in the decoder section (e.g., the previous predicted value), the size of the information (e.g., encoded data) that must be provided to the decoder section to generate the predicted value can be reduced.
[0018]
[0052] Specific aspects of this disclosure are described below with reference to the drawings. In this description, common features are indicated by common reference numerals. Where used herein, various terms are used solely for the purpose of describing specific implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural form unless the context otherwise indicates. Furthermore, some features described herein are singular in some implementations and plural in others. For example, Figure 2 shows a device 202 containing one or more processors ("Processors (singular or plural)" 290 in Figure 2), which indicates that in some implementations, device 202 contains a single processor 290, and in other implementations, device 202 contains multiple processors 290. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular form unless an aspect relating to multiple features is described.
[0019]
[0053] In some drawings, multiple instances of a particular type of feature are used. These features are physically and / or logically distinct, but the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When a feature is referred herein as a group or type (for example, when no particular feature is referred), the reference number is used without a distinguishing letter. However, when a particular feature of one of several features of the same type is referred herein, the reference number is used with a distinguishing letter. For example, referring to Figure 1, several input values are shown and associated with reference numbers 105A, 105B, and 105C. When referring to a particular one of these input values, such as input value 105A, the distinguishing letter "A" is used. However, when referring to any one of these input values or these input values as a group, the reference number 105 is used without a distinguishing letter.
[0020]
[0054] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” indicates an example, implementation, and / or aspect, and should not be construed as limiting, or indicating a preferred or suitable implementation. As used herein, order terms used to modify elements such as structure, components, and behavior (e.g., “first,” “second,” “third,” etc.) do not themselves indicate priority or order of an element relative to another element, but merely distinguish an element from another element having the same name (apart from the use of order terms). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple of a particular element (e.g., two or more).
[0021]
[0055] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and may also include any combination thereof. Two devices (or components) may be directly or indirectly coupled via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or a combination thereof) (e.g., they may be communicatively coupled, electrically coupled, or physically coupled). Two electrically coupled devices (or components) may be contained in the same device or in different devices, and may be connected via electronic components, one or more connectors, or inductive coupling, as exemplary and non-limiting examples. Two communicatively coupled devices (or components), such as those communicating telecommunicates, may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled without any intermediary components (e.g., communicatively coupled, electrically coupled, or physically coupled).
[0022]
[0056] In this disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” and “adjusting” may be used to describe how one or more actions are performed. Such terms should not be construed as restrictive, and it should be noted that other techniques may be used to perform similar actions. Additionally, as used herein, “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or it may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device.
[0023]
[0057] Referring to Figure 1, a particular exemplary embodiment of a compression network configured to perform predictions is disclosed and is designated 140 overall. The compression network 140 includes an encoder portion 160 configured to be coupled to a decoder portion 180.
[0024]
[0058] In some embodiments, the encoder portion 160 is contained in a first device distinct from a second device containing the decoder portion 180, as further illustrated with reference to Figure 2. For example, in these embodiments, the compression network 140 can be used to reduce resource usage (e.g., memory, bandwidth, transmission time, etc.) associated with transmitting data from the first device to the second device. In other embodiments, the encoder portion 160 and the decoder portion 180 are contained in a single device, as further illustrated with reference to Figure 3. For example, in these embodiments, the compression network 140 can be used to reduce resource usage (e.g., memory) associated with storing data for access by the device.
[0025]
[0059] The encoder section 160 is configured to process one or more input values 105 to generate one or more sets of encoded data 165. In one embodiment, one or more input values 105 include input value 105A, input value 105B, and input value 105C, one or more additional input values, or a combination thereof. The encoder section 160 is configured to process input value 105A to generate encoded data 165A, process input value 105B to generate encoded data 165B, process input value 105C to generate encoded data 165C, and so on.
[0026]
[0060] The encoder portion 160 includes a conditional input generator 162 coupled to the encoder 166 via a feature generator 164. In some embodiments, the encoder 166 is configured to process some of one or more input values 105 (e.g., key values) to generate corresponding encoded data, independently of the other input values 105. In one embodiment, input value 105A corresponds to a key image frame (e.g., an intra-frame or I-frame), and the encoder 166 processes input value 105A to generate encoded data 165A, independently of the other input values 105. In some embodiments, the encoder 166 is configured to process some of one or more input values 105 based on at least one other input value (e.g., a non-key value) to generate corresponding encoded data. In one embodiment, the input value 105B corresponds to a non-key image frame such as a prediction frame (P frame) or a bidirectional frame (B frame), and the encoder 166 processes the input value 105B based on the encoded data 165A (corresponding to the input value 105A) to generate encoded data 165B. For example, the conditional input generator 162 is configured to process the encoded data 165A to generate a conditional input 167B, the feature generator 164 is configured to process the conditional input 167B to generate feature data 163B, and the encoder 166 is configured to process the input value 105B and the feature data 163B to generate encoded data 165B.
[0027]
[0061] The decoder unit 180 is configured to process one or more sets of encoded data 165 to generate one or more predicted values 195. In one embodiment, the decoder unit 180 is configured to process encoded data 165A to generate predicted value 195A, process encoded data 165B to generate predicted value 195B, process encoded data 165C to generate predicted value 195C, and so on.
[0028]
[0062] The decoder section 180 includes a conditional input generator 182 coupled to the decoder 186 via a feature generator 184. In some embodiments, the decoder 186 is configured to process some of the sets of encoded data 165 to generate corresponding predicted values, independently of other sets of encoded data 165. In one embodiment, the decoder 186 processes encoded data 165A to generate predicted value 195A, independently of other sets of encoded data 165. In some embodiments, the decoder 186 is configured to process some of the sets of encoded data 165 based on at least one other set of encoded data 165 to generate corresponding predicted values. For example, the decoder 186 processes encoded data 165B based on predicted value 195A (corresponding to encoded data 165A) to generate predicted value 195B. For example, the conditional input generator 182 is configured to process the predicted value 195A to generate the conditional input 187B, the feature generator 184 is configured to process the conditional input 187B to generate the feature data 183B, and the decoder 186 is configured to process the encoded data 165B and the feature data 183B to generate the predicted value 195B.
[0029]
[0063] The conditional input 167B in the encoder section 160 corresponds to an estimate of the conditional input 187B available in the decoder section 180. A technical advantage of generating encoded data 165B based on an estimate of the information available in the decoder section 180 (e.g., conditional input 187B) is that it can reduce the size of the information (e.g., encoded data 165B) that must be provided to the decoder section 180 to generate the predicted value 195B.
[0030]
[0064] In some implementations, the compression network components of the compression network 140 (e.g., encoder portion 160, decoder portion 180, or both) correspond to or are included in one of various types of devices. In an exemplary embodiment, the compression network components are incorporated into a headset device, as further described with reference to Figure 26. In other embodiments, the compression network components are incorporated into at least one of the following: a mobile phone or tablet computer device, as described with reference to Figure 25; a wearable electronic device, as described with reference to Figure 27; a voice-controlled speaker system, as described with reference to Figure 28; a camera device, as described with reference to Figure 29; or a virtual reality, mixed reality, or augmented reality headset, as described with reference to Figure 30. In yet another exemplary embodiment, the compression network components are incorporated into a vehicle, as further described with reference to Figures 31 and 32.
[0031]
[0065] During operation, the encoder portion 160 acquires one or more input values 105. In some implementations, one or more input values 105 are based on the output of one or more sensors, as further described with reference to Figure 2. In exemplary, non-limiting embodiments, one or more sensors may include an image sensor, an inertial measurement unit (IMU), a motion sensor, a temperature sensor, another type of sensor, or a combination thereof. In some implementations, the encoder portion 160 acquires one or more input values 105 from a memory device, a processor, another device, or a combination thereof.
[0032]
[0066] In some implementations, one or more input values 105 indicate motion, as further illustrated with reference to Figure 14. For example, input value 105A includes a first image frame captured by the vehicle's camera, and input value 105B includes a second image frame captured by the camera. The difference between the first and second image frames may indicate the speed and direction of the vehicle's movement. The one or more input values 105 containing motion-indicating image frames are provided as exemplary examples, and in other embodiments, one or more input values 105 may include other types of data indicating motion.
[0033]
[0067] The encoder portion 160 generates a set of encoded data 165 corresponding to one or more input values 105. For example, encoder 166 processes input value 105A to generate encoded data 165A. In some implementations, encoder 166 processes (e.g., encodes) input value 105A to generate encoded data 165A independently of other input values among the one or more input values 105, based on the determination that input value 105A satisfies the independent encoding criterion. In one embodiment, encoder 166 determines that the independent encoding criterion is satisfied based on the determination that input value 105A is the initial value of one or more input values 105, input value 105A corresponds to a key value (e.g., an I-frame), at least a threshold count of input values has been encoded since the most recently independently encoded input value, the difference between input value 105A and the previous input value among the one or more input values 105 is greater than a difference threshold (e.g., due to a scene change in an image frame), or a combination thereof.
[0034]
[0068] In another embodiment, encoder 166 processes the input value 105B to generate encoded data 165B. In some implementations, encoder 166 processes the input value 105B based at least in part on the encoded data 165A of input value 105A, depending on whether it has determined that the input value 105B does not satisfy an independent encoding criterion. In some embodiments, encoder 166 selects the encoded data 165A for processing the input value 105B based on whether it has determined that the input value 105A corresponds to one of the most recently and independently encoded input values 105 (e.g., a key value). In some embodiments, encoder 166 selects the encoded data 165A for processing the input value 105B based on whether it has determined that the input value 105B corresponds to the next value of one of the subsequent input values 105 after input value 105A. In these embodiments, the input value 105A may correspond to a key value that is encoded independently to generate encoded data 165A, or a non-key value that is encoded based on at least one other input value among one or more input values 105.
[0035]
[0069] The encoder section 160, upon determining that the input value 105B does not satisfy the independent coding criterion, provides the coded data 165A to the conditional input generator 162 to obtain the conditional input 167B of the compression network 140. The conditional input generator 162 includes a local decoder section 168, an estimator 170, or both. In some embodiments, the local decoder section 168 is configured to perform operations similar to those of the decoder section 180. For example, the local decoder section 168 processes the coded data 165A to generate a predicted value 169A by performing one or more operations described herein with reference to the decoder section 180. The predicted value 169A corresponds to an estimate of the predicted value 195A that can be generated in the decoder section 180 by processing the coded data 165A.
[0036]
[0070] In some implementations, the estimator 170 processes the predicted value 169A to generate an estimated value 171B. In some embodiments, the estimated value 171B corresponds to an estimate of the input value 105B that can be generated in the decoder section 180 based on the predicted value 195A. In some embodiments, the closer the estimated value 171B is to the input value 105B, the less information needs to be provided to the decoder section 180 as encoded data 165B.
[0037]
[0071] The conditional input 167B is based on the predicted value 169A, the estimated value 171B, or both. In some implementations, the conditional input 167B may be based on one or more additional predicted values, one or more additional estimated values, or a combination thereof, associated with one or more of the other input values among the one or more input values 105.
[0038]
[0072] The encoder section 160, upon determining that the input value 105B does not satisfy the independent coding criterion, uses the compression network 140 to process the input value 105B and the conditional input 167B to generate coded data 165B associated with the input value 105B. For example, the feature generator 164 processes the conditional input 167B to generate feature data 163B, as further explained with reference to Figure 4. In some implementations, the feature data 163B corresponds to multiscale feature data with different resolutions (e.g., different spatial resolutions). In some implementations, the feature data 163B includes multiscale wavelet transform data. The encoder 166 processes the input value 105B and the feature data 163B to generate coded data 165B, as further explained with reference to Figure 4.
[0039]
[0073] The decoder unit 180 obtains one or more sets of encoded data 165 associated with one or more input values 105. In some implementations, as further described with reference to Figure 2, the encoder unit 160 in the first device provides the encoded data 165 to the decoder unit 180 in the second device via a bitstream. In other implementations, as further described with reference to Figure 3, the encoder unit 160 stores the encoded data 165 in a storage device, and the decoder unit 180 retrieves the encoded data 165 from the storage device.
[0040]
[0074] The decoder section 180 generates one or more predicted values 195 corresponding to a set of encoded data 165. For example, decoder 186 processes encoded data 165A to generate predicted value 195A. In some implementations, decoder 186 processes (e.g., decodes) encoded data 165A based on the determination that it satisfies the independent decoding criterion to generate predicted value 195A independently of other predicted values among the one or more predicted values 195. In one embodiment, decoder 186 determines that the independent coding criterion is satisfied based on the determination that the metadata associated with the encoded data 165 indicates that the encoded data 165 corresponds to a key value (e.g., an I-frame), that the encoded data 165 is decoded independently, or both.
[0041]
[0075] In another embodiment, the decoder 186 processes the encoded data 165B to generate a predicted value 195B. In some implementations, the decoder 186 processes the encoded data 165B based at least in part on a predicted value 195A that corresponds to the encoded data 165A associated with the input value 105A, depending on whether the encoded data 165B does not satisfy the independent decoding criterion. In some embodiments, the decoder 186 selects a predicted value 195A for processing the encoded data 165B based on whether the predicted value 195A corresponds to the most recently decoded independently value (e.g., key value). In some embodiments, the decoder 186 selects a predicted value 195A for processing the encoded data 165B based on whether the predicted value 195B corresponds to the next value among one or more predicted values 195 that follow the predicted value 195A. In these embodiments, the predicted value 195A may correspond to a key value that is decoded independently, or a non-key value that is decoded based on at least one other predicted value among one or more predicted values 195.
[0042]
[0076] The decoder unit 180, upon determining that the encoded data 165B does not satisfy the independent decoding criterion, provides the predicted value 195A to the conditional input generator 182 to obtain the conditional input 187B of the compression network 140. In some implementations, the conditional input generator 182 outputs the predicted value 195A as the conditional input 187B. Optionally, in some implementations, the conditional input generator 182 includes an estimator 188 that processes the predicted value 195A to produce an estimated value 193B. In some embodiments, the estimated value 193B corresponds to an estimate of the input value 105B based on the predicted value 195A. In some embodiments, the closer the estimated value 193B is to the input value 105B, the less information needs to be processed by the decoder unit 180 as encoded data 165B to generate the predicted value 195B.
[0043]
[0077] The conditional input 187B is based on the predicted value 195A, the estimated value 193B, or both. In some implementations, the conditional input 187B may be based on one or more predicted values 195, one or more additional estimated values, or at least one of a combination thereof.
[0044]
[0078] The decoder section 180, upon determining that the encoded data 165B does not satisfy the independent decoding criterion, uses the compression network 140 to process the encoded data 165B and the conditional input 187B to generate a predicted value 195B associated with the input value 105B. For example, the feature generator 184 processes the conditional input 187B to generate feature data 183B, as further explained with reference to Figure 5. In some implementations, the feature data 183B corresponds to multiscale feature data with different resolutions (e.g., different spatial resolutions). In some implementations, the feature data 183B includes multiscale wavelet transform data. The decoder 186 processes the encoded data 165B and the feature data 183B to generate a predicted value 195B, as further explained with reference to Figure 5.
[0045]
[0079] In some embodiments, the predicted value 195B corresponds to a reconfigured version of the input value 105B. In some embodiments, the predicted value 195B corresponds to a predicted future value, such as a prediction of future values for one or more input values 105. For example, the future value could be an input value 105C that is not yet available in the encoder portion 160 at the time of generating the encoded data 165B. In some embodiments, the predicted value 195B corresponds to a classification associated with the input value 105B. For example, the predicted value 195B may indicate whether the input value 105B corresponds to an alert condition.
[0046]
[0080] In some embodiments, the predicted value 195B corresponds to a detection result associated with the input value 105B. For example, the predicted value 195B indicates whether a face was detected in the image frame associated with the input value 105B. In some embodiments, the predicted value 195B corresponds to a collision avoidance output. For example, the input value 105B indicates a first position of a vehicle relative to a second position of an object. In some embodiments, the predicted value 195B indicates a first predicted future position of a vehicle relative to a second predicted future position of an object. In some embodiments, the predicted value 195B indicates whether the first predicted future position of the vehicle is within a collision threshold for the second future position of the object. In some embodiments, one or more predicted values 195B are processed by one or more downstream applications, and the compression network 140 is trained (e.g., configured) based on performance metrics associated with the downstream applications. For example, the compression network 140 is trained to reduce loss metrics associated with the downstream applications.
[0047]
[0081] The technical advantages of using the compressed network 140 to generate one or more predicted values 195 include reduced resource usage. For example, by generating encoded data 165B based on the conditional input 167B as an estimate of the information available in the decoder section 180 (e.g., the conditional input 187B), the amount of information that must be provided to the decoder section 180 as encoded data 165B in order to maintain the accuracy of the predicted value 195B is reduced.
[0048]
[0082] Referring to Figure 2, an exemplary embodiment of a system capable of performing predictions using a compressed network is shown, designated collectively as 200. System 200 includes a device 202 that is coupled to, or configured to include, one or more sensors 240. Device 202 is configured to communicate with device 260. For example, device 202 is configured to be coupled to device 260 via a network (e.g., a wired network, a wireless network, or both).
[0049]
[0083] The encoder portion 160 is included in one or more processors 290 of device 202, and the decoder portion 180 is included in one or more processors 292 of device 260. One or more processors 290 are coupled to one or more sensors 240 and modem 270. One or more processors 292 are coupled to modem 280. Optionally, in some implementations, one or more processors 292 include one or more applications 262. Optionally, in some implementations, device 260 is configured to be coupled to device 264.
[0050]
[0084] During operation, the encoder unit 160 receives sensor data 226 from one or more sensors 240 as one or more input values 105. In an exemplary, non-limiting embodiment, one or more sensors 240 may include an image sensor, IMU, motion sensor, accelerometer, speedometer, gyroscope, radar, temperature sensor, microphone, other type of sensor, or a combination thereof. The encoder unit 160 processes one or more input values 105 (e.g., sensor data 226) to generate a set of encoded data 165, as described with reference to Figure 1. The encoder unit 160 provides the set of encoded data 165 to the modem 270 for transmission as a bitstream 235 to the device 260. The modem 280 receives the bitstream 235 and provides the set of encoded data 165 to the decoder unit 180.
[0051]
[0085] The decoder section 180 processes a set of encoded data 165 to generate one or more predicted values 195, as illustrated with reference to Figure 1. One or more processors 292 generate an output 295 based on one or more predicted values 195. Optionally, in some implementations, the decoder section 180 provides one or more predicted values 195 to one or more applications 262, and one or more applications 262 generate an output 263 based on one or more predicted values 195. For example, one or more predicted values 195 may correspond to image frames, and application 262 processes the predicted values 195B corresponding to the image frames to generate an output 263 indicating the classification associated with the image frames. The output 295 is based on output 263, one or more predicted values 195, or a combination thereof.
[0052]
[0086] In some implementations, the encoder portion 160, the decoder portion 180, or both are trained (e.g., configured) based on performance metrics associated with one or more applications 262. For example, a network trainer trains the encoder portion 160, the decoder portion 180, or both (e.g., configures its network weights and biases) based on loss metrics associated with the classification output generated by the application 262. In some implementations, one or more processors 292 provide outputs 295 to a device 264. The device 264 may include a display device, a network device, a storage device, a user device, or a combination thereof.
[0053]
[0087] In some embodiments, output 295 initiates one or more actions in device 264. In an exemplary embodiment, device 264 and device 202 are the same device (e.g., a vehicle), and one or more predicted values 195 correspond to collision avoidance outputs. Device 260 transmits output 295 to initiate one or more collision avoidance actions (e.g., braking) in device 202, in response to determining that predicted value 195B indicates that a first predicted future position of device 260 is expected to be within a threshold distance of a second predicted future position of an object. In some implementations, device 260 is the same as or included in device 202 (e.g., a vehicle). In other implementations, device 260 is external to device 202 and generates output 295 based on one or more predicted values 195, based on one or more additional predicted values (e.g., associated with an object, another vehicle, or both), or a combination thereof.
[0054]
[0088] A technical advantage of using the encoder section 160 to generate encoded data 165 based on an estimate of the information available in the decoder section 180 is that it can reduce the resource usage (e.g., memory, bandwidth, and transmission time) associated with transmitting the bitstream 235 to the device 260.
[0055]
[0089] Referring to Figure 3, an exemplary embodiment of a system capable of performing predictions using a compressed network is shown, collectively designated as 300. System 300 includes a device 302 that is coupled to or configured to include one or more sensors 240. Device 302 is coupled to or configured to include a storage device 392.
[0056]
[0090] The encoder portion 160 and the decoder portion 180 are included in one or more processors 390 of device 302. The encoder portion 160 is configured to store one or more sets of encoded data 165 in the storage device 392. The decoder portion 180 is configured to retrieve one or more sets of encoded data 165 from the storage device 392. Storing one or more sets of encoded data 165 in the storage device 392 uses less memory than storing the corresponding values of sensor data 226 in the storage device 392.
[0057]
[0091] The decoder section 180 processes a set of encoded data 165 to generate one or more predicted values 195, as described with reference to Figure 1. One or more processors 390 generate an output 295 based on one or more predicted values 195. Optionally, in some implementations, the decoder section 180 provides one or more predicted values 195 to one or more applications 262, and one or more applications 262 generate an output 263 based on one or more predicted values 195. The output 295 is based on output 263, one or more predicted values 195, or a combination thereof. In some implementations, one or more processors 390 provide the output 295 to a device 264.
[0058]
[0092] In an exemplary embodiment, devices 264 and 302 are the same device (e.g., a vehicle), and one or more predicted values 195 correspond to collision avoidance outputs. Device 302 generates an output 295 to initiate one or more collision avoidance actions (e.g., braking) in device 302 when it determines that predicted value 195B indicates that a first predicted future position of device 302 is expected to be within a threshold distance of a second predicted future position of an object.
[0059]
[0093] Referring to Figure 4, an embodiment 400 of the encoder portion 160 of a compression network 140 that can operate to perform predictions is shown. For example, the encoder portion 160 is configured to generate encoded data 165B that can be processed by the decoder portion 180 of the compression network 140 to generate a predicted value 195B, as will be further explained with reference to Figure 5.
[0060]
[0094] The compression network 140 includes a neural network having multiple layers. For example, the feature generator 164 includes one or more feature layers 404 coupled to one or more encoder layers 402 of the encoder 166. The one or more feature layers 404 include one or more additional feature layers, including feature layer 404A, feature layer 404B, feature layer 404C, feature layer 404N, or a combination thereof. The one or more encoder layers 402 include one or more additional encoder layers, including encoder layer 402A, encoder layer 402B, encoder layer 402C, encoder layer 402N, or a combination thereof. The output of each feature layer 404 is coupled to the input of the corresponding encoder layer 402. For example, the output of feature layer 404A is coupled to the input of encoder layer 402A, the output of feature layer 404B is coupled to the input of encoder layer 402B, and so on.
[0061]
[0095] The output of each preceding feature layer 404 is coupled to the input of the subsequent feature layer 404. For example, the output of feature layer 404A is coupled to the input of feature layer 404B, the output of feature layer 404B is coupled to the input of feature layer 404C, and so on. The output of each preceding encoder layer 402 is coupled to the input of the subsequent encoder layer 402. For example, the output of encoder layer 402A is coupled to the input of encoder layer 402B, the output of encoder layer 402B is coupled to the input of encoder layer 402C, and so on.
[0062]
[0096] In some implementations, the encoder 166 is configured to encode multiple resolutions of an input value 105 to generate encoded data 165. In an exemplary embodiment, the encoder 166 corresponds to a video encoder configured to encode multiple spatial resolutions of an input value 105 (e.g., an image frame). One or more feature layers 404 and one or more encoder layers 402 correspond to network layers associated with multiple resolutions. For example, feature layer 404A and encoder layer 402A are associated with a first resolution, feature layer 404B and encoder layer 402B are associated with a second resolution, feature layer 404C and encoder layer 402C are associated with a third resolution, feature layer 404N and encoder layer 402N are associated with an Nth resolution, or a combination thereof.
[0063]
[0097] During operation, the encoder section 160 provides the conditional input 167B to the feature generator 164, as described with reference to Figure 1, and input value 105B(x b The ) is provided to the encoder 166. The feature layer 404A processes the conditional input 167B to generate feature data 163BA associated with the first resolution and provides the feature data 163BA to the encoder layer 402A. The encoder layer 402A processes the input value 105B and the feature data 163BA to generate an output (associated with the first resolution) that is provided to the encoder layer 402B.
[0064]
[0098] Each subsequent feature layer 404 processes the output of the previous feature layer 404 to generate feature data 163 and provides the feature data 163 to the corresponding encoder layer 402. For example, feature layer 404B provides feature data 163BB associated with a second resolution to encoder layer 402B, feature layer 404C provides feature data 163BC associated with a third resolution to encoder layer 402C, feature layer 404N provides feature data 163BN associated with an Nth resolution to encoder layer 402N, and so on. Thus, feature data 163B (e.g., feature data 163BA, feature data 163BB, feature data 163BC, feature data 163BN, or a combination thereof) corresponds to multiscale feature data with different resolutions.
[0065]
[0099] Each subsequent encoder layer 402 processes the output of the previous encoder layer 402 and the feature data 163 from the corresponding feature layer to generate an output. For example, encoder layer 402B processes the output of encoder layer 402A and the feature data 163BB to generate an output (associated with a second resolution) provided to encoder layer 402C, and so on. The output of encoder layer 402N corresponds to encoded data 165B.
[0066]
[0100] Optionally, in some implementations, the feature generator 164 can generate feature data 163B using various techniques. For example, the feature generator 164 can perform a wavelet transform to generate feature data 163B corresponding to multiscale wavelet transform data. For example, the feature generator 164 can perform a wavelet transform based on a conditional input 167B to generate feature data 163BA (e.g., first wavelet transform data) associated with a first resolution. The feature generator 164 can perform a wavelet transform based on feature data 163BA, conditional input 167B, or both to generate feature data 163BB (e.g., second wavelet transform data) associated with a second resolution. Similarly, the feature generator 164 can generate feature data 163BC (e.g., third wavelet transform data) associated with a third resolution, and so on.
[0067]
[0101] In certain embodiments, the conditional input 167B corresponds to an estimate of the conditional input 187B that can be generated in the decoder section 180 of Figure 5. A technical advantage of using an estimate of the information available in the decoder section 180 (e.g., conditional input 167B) to generate the feature data 163B used to generate the encoded data 165B is that it can reduce the size of the encoded data 165B in order to maintain the accuracy of the predicted value 195B generated in the decoder section 180.
[0068]
[0102] Referring to Figure 5, an embodiment 500 of the decoder portion 180 of the compression network 140, which is capable of performing predictions, is shown. For example, the decoder portion 180 processes the encoded data 165B (generated by the encoder portion 160 as described with reference to Figure 4) to input value 105B(x b It is configured to generate the predicted value 195B(γ) associated with ).
[0069]
[0103] The compressed network 140 includes a neural network having multiple layers. For example, the feature generator 184 includes one or more feature layers 504 coupled to one or more decoder layers 508 of the decoder 186. The one or more feature layers 504 include one or more additional feature layers, including feature layer 504A, feature layer 504B, feature layer 504C, feature layer 504N, or a combination thereof. The one or more decoder layers 508 include one or more additional decoder layers, including decoder layer 508A, decoder layer 508B, decoder layer 508C, decoder layer 508N, or a combination thereof. The output of each feature layer 504 is coupled to the input of the corresponding decoder layer 508. For example, the output of feature layer 504A is coupled to the input of decoder layer 508A, the output of feature layer 504B is coupled to the input of decoder layer 508B, and so on.
[0070]
[0104] The output of each preceding feature layer 504 is coupled to the input of the following feature layer 504. For example, the output of feature layer 504A is coupled to the input of feature layer 504B, the output of feature layer 504B is coupled to the input of feature layer 504C, and so on. The output of each preceding decoder layer 508 is coupled to the input of the following decoder layer 508. One or more decoder layers 508 are ordered from higher-alphabetical reference numbers to lower-alphabetical ones. For example, decoder layer 508B follows decoder layer 508C, and decoder layer 508A follows decoder layer 508B. The output of decoder layer 508B is coupled to the input of decoder layer 508A, the output of decoder layer 508C is coupled to the input of decoder layer 508B, and so on. The output of the last decoder layer 508 of decoder 186 (e.g., decoder layer 508A) corresponds to the predicted value 195 associated with the input value 105.
[0071]
[0105] In some implementations, the decoder 186 is configured to decode multiple resolutions of encoded data 165B. In an exemplary embodiment, the decoder 186 corresponds to a video decoder configured to decode multiple spatial resolutions of encoded data 165B associated with an input value 105 (e.g., an image frame). One or more feature layers 504 and one or more decoder layers 508 correspond to network layers associated with multiple resolutions. For example, feature layer 504A and decoder layer 508A are associated with a first resolution, feature layer 504B and decoder layer 508B are associated with a second resolution, feature layer 504C and decoder layer 508C are associated with a third resolution, feature layer 504N and decoder layer 508N are associated with an Nth resolution, or a combination thereof.
[0072]
[0106] During operation, the decoder section 180 provides the conditional input 187B to the feature generator 184 and the encoded data 165B to the decoder 186, as described with reference to Figure 1. The feature layer 504A processes the conditional input 187B to generate feature data 183BA associated with the first resolution and provides the feature data 183BA to the decoder layer 508A. Each subsequent feature layer 504 processes the output of the previous feature layer 504 to generate feature data 183 and provides the feature data 183 to the corresponding decoder layer 508. For example, feature layer 504B provides feature data 183BB associated with the second resolution to the decoder layer 508B, feature layer 504C provides feature data 183BC associated with the third resolution to the decoder layer 508C, feature layer 504N provides feature data 183BN associated with the nth resolution to the decoder layer 508N, and so on. Therefore, feature data 183B (e.g., feature data 183BA, feature data 183BB, feature data 183BC, feature data 183BN, or a combination thereof) corresponds to multiscale feature data with different resolutions.
[0073]
[0107] Decoder layer 508N processes encoded data 165B and feature data 183BN to produce an output (e.g., associated with the Nth resolution) that is provided to subsequent decoder layers 508. Each subsequent decoder layer 508 processes the output of the previous decoder layer 508 and feature data 183 from the corresponding feature layer to produce an output. For example, decoder layer 508B processes the output of decoder layer 508C and feature data 183BB to produce an output (e.g., associated with the second resolution). Decoder layer 508A processes the output of decoder layer 508B and feature data 183BA to produce an input value 105B(x b This generates an output corresponding to the predicted value 195B(γ) associated with ).
[0074]
[0108] Optionally, in some implementations, the feature generator 184 can generate feature data 183B using a variety of techniques similar to those performed by the feature generator 164 in Figure 4. For example, the feature generator 184 can perform a wavelet transform to generate feature data 183B corresponding to multiscale wavelet transform data. For example, the feature generator 184 can perform a wavelet transform based on a conditional input 187B to generate feature data 183BA (e.g., first wavelet transform data) associated with a first resolution. The feature generator 184 can perform a wavelet transform based on feature data 183BA, conditional input 187B, or both to generate feature data 183BB (e.g., second wavelet transform data) associated with a second resolution. Similarly, the feature generator 184 can generate feature data 183BC (e.g., third wavelet transform data) associated with a third resolution, and so on.
[0075]
[0109] Optionally, in some implementations, the decoder section 180 may include a feature generator 582 configured to process decoder information 587 available in the decoder section 180 to generate feature data 583 used by the decoder 186 to generate a predicted value 195B. In one embodiment, the feature generator 582 includes one or more feature layers 506 coupled to one or more decoder layers 508. The one or more feature layers 506 include one or more additional feature layers, including feature layer 506A, feature layer 506B, feature layer 506C, feature layer 506N, or a combination thereof. The output of each feature layer 506 is coupled to the input of the corresponding decoder layer 508. For example, the output of feature layer 506A is coupled to the input of decoder layer 508A, the output of feature layer 506B is coupled to the input of decoder layer 508B, and so on.
[0076]
[0110] The output of each preceding feature layer 506 is coupled to the input of a subsequent feature layer 506. For example, the output of feature layer 506A is coupled to the input of feature layer 506B, the output of feature layer 506B is coupled to the input of feature layer 506C, and so on. In certain embodiments, one or more feature layers 506 correspond to network layers associated with multiple resolutions. For example, feature layer 506A is associated with a first resolution, feature layer 506B is associated with a second resolution, feature layer 506C is associated with a third resolution, feature layer 506N is associated with an Nth resolution, or a combination thereof.
[0077]
[0111] The decoder section 180 provides the conditional input 187B to the feature generator 184 and the encoded data 165B to the decoder 186, and (for example, simultaneously) provides the decoder information 587 to the feature generator 582. The feature layer 506A processes the decoder information 587 to generate feature data 583A associated with the first resolution and provides the feature data 583A to the decoder layer 508A. Each subsequent feature layer 506 processes the output of the previous feature layer 506 to generate feature data 583 and provides the feature data 583 to the corresponding decoder layer 508. For example, feature layer 506B processes the output of feature layer 506A to generate feature data 583B and provides the feature data 583B associated with the second resolution to the decoder layer 508B. Similarly, feature layer 506C provides feature data 583C associated with a third resolution to decoder layer 508C, feature layer 506N provides feature data 583N associated with an Nth resolution to decoder layer 508N, and so on. Thus, feature data 583 (e.g., feature data 583A, feature data 583B, feature data 583C, feature data 583N, or a combination thereof) corresponds to multiscale feature data with different resolutions.
[0078]
[0112] Decoder layer 508N processes encoded data 165B, feature data 183BN, and feature data 583N to produce an output (e.g., associated with the Nth resolution) that is provided to subsequent decoder layers 508. Each subsequent decoder layer 508 processes the output of the previous decoder layer 508, feature data 183 from the corresponding feature layer of feature generator 184, and feature data 583 from the corresponding feature layer of feature generator 582 to produce an output. For example, decoder layer 508B processes the output of decoder layer 508C, feature data 183BB, and feature data 583B to produce an output (e.g., associated with the second resolution). Decoder layer 508A processes the output of decoder layer 508B, feature data 183BA, and feature data 583A to produce an input value 105B(x b This generates an output corresponding to the predicted value 195B(γ) associated with ).
[0079]
[0113] Optionally, in some implementations, the feature generator 582 can generate feature data 583 using various techniques (similar to the techniques performed by, for example, the feature generator 184). For example, the feature generator 582 can perform a wavelet transform to generate feature data 583 corresponding to multiscale wavelet transform data.
[0080]
[0114] In some implementations, the predicted value 195B(γ) is obtained from the input value 105B(x), as further explained with reference to Figures 6 to 11. b This corresponds to a reconfigured version of ). For example, input value 105B(x b ) can correspond to an image unit, and the predicted value 195B(γ) can correspond to a reconfigured version of the image unit, as further illustrated with reference to Figures 14 to 23. In some implementations, the predicted value 195B(γ) corresponds to the predicted future values of one or more input values 105, as further illustrated with reference to Figures 12 to 13. In other implementations, the predicted value 195B(γ) can correspond to the predicted future values associated with input value 105B. In some embodiments, the predicted value 195B also includes predictions of future values for one or more input values 105. In other embodiments, the predicted value 195B does not include predictions of future values for one or more input values 105. In one embodiment, the compression network 140 tracks objects associated with one or more motion values (e.g., one or more input values 105) across one or more frames of pixels. For example, the input value 105B may include tracking data (e.g., associated with an image frame), and the prediction value 195B(γ) may include a collision avoidance output indicating whether the vehicle's future position relative to an object (e.g., the object's future position) is predicted to be below a threshold. The collision avoidance output (e.g., collision or no collision) is a future prediction associated with the tracking data. In some implementations, the collision avoidance output also indicates predicted future tracking data, such as the vehicle's predicted future position relative to the object's predicted future position.
[0081]
[0115] Referring to Figure 6, an embodiment 600 of the encoder portion 160 of the compression network 140, which is capable of performing predictions corresponding to reconstruction, is shown. For example, the encoder portion 160 takes an input value 105B(x) as further described with reference to Figure 7. b Predicted value 195B corresponding to the reconfigured version of )
[0082]
number
[0083] The system is configured to generate encoded data 165B that can be processed in the decoder section 180 of the compression network 140 in order to generate the above.
[0084]
[0116] The conditional input generator 162 uses the local decoder section 168 to process the encoded data 165A to generate a predicted value 169A, as described with reference to Figure 1. In a particular embodiment, the predicted value 169A
[0085]
number
[0086] The input value 105A(x) that can be generated in the decoder section 180 is... a This corresponds to the estimate of the reconfigured version of ).
[0087]
[0117] In some implementations, the conditional input generator 162 generates one or more additional predicted values. These one or more additional predicted values may be associated with at least one input value prior to input value 105B among the one or more input values 105, at least one input value after input value 105B among the one or more input values 105, or both.
[0088]
[0118] In one embodiment, one or more additional predicted values are after input value 105B among one or more input values 105 and can be encoded as encoded data 165C independently of (e.g., before) encoding input value 105B, and are associated with a predicted value 169C associated with input value 105C
[0089]
Number
[0090] can include. In an exemplary embodiment, input value 105C corresponds to a key value (e.g., an I-frame) that can be encoded independently of other input values among one or more input values 105. The conditional input generator 162 uses the local decoder portion 168 to process the encoded data 165C to generate the predicted value 169C, process one or more additional sets of encoded data to generate one or more additional predicted values, or perform a combination thereof.
[0091]
[0119] Optionally, in some implementations, the conditional input generator 162 uses the estimator 170 to generate an estimated value 171B of the input value 105B (x b ) based on one or more predicted values (e.g., predicted value 169A, predicted value 169C, one or more additional predicted values, or a combination thereof).
[0092]
Number
[0093] To illustrate, the estimated value 171B
[0094]
Number
[0095] can be generated by the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values and independently of the encoded data 165B, for the input value 105B (x bThis corresponds to the estimated value of ).
[0096]
[0120] In certain implementations, the estimator 170 uses one or more estimation techniques to determine the estimate 171B. For example, the predicted value 169A corresponds to the input value 105A (e.g., the previous image frame), the predicted value 169C corresponds to the input value 105C (e.g., the subsequent image frame), and the estimate 171B corresponds to the estimated input value between the predicted values 169A and 169C. In some implementations, the estimator 170 includes a neural network configured to process one or more predicted values to generate the estimate 171B.
[0097]
[0121] The conditional input 167B gives the estimated value 171B.
[0098]
number
[0099] Predicted value: 169A
[0100]
number
[0101] Predicted value: 169C
[0102]
number
[0103] This includes one or more predicted values, such as one or more additional predicted values, or a combination thereof. As illustrated with reference to Figure 4, the feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes the input value 105B based on the feature data 163B to generate encoded data 165B.
[0104]
[0122] Referring to Figure 7, an embodiment 700 of the decoder portion 180 of the compression network 140, which is capable of performing predictions corresponding to reconstruction, is shown. For example, the decoder portion 180 processes the encoded data 165B (generated by the encoder portion 160 as described with reference to Figure 6) to input value 105B(x b Predicted value 195B corresponding to the reconfigured version of )
[0105]
number
[0106] It is configured to generate.
[0107]
[0123] The conditional input generator 182 obtains the predicted value 195A, as described with reference to Figure 1. For example, the decoder section 180 processes the encoded data 165A to generate the predicted value 195A. In a particular embodiment, the predicted value 195A
[0108]
number
[0109] The input value is 105A(x a This corresponds to a reconfigured version of ).
[0110]
[0124] In some implementations, the conditional input generator 182 obtains one or more additional predicted values. These one or more additional predicted values may be associated with at least one input value prior to input value 105B among the one or more input values 105, at least one input value after input value 105B among the one or more input values 105, or both.
[0111]
[0125] In one embodiment, one or more additional predicted values are associated with the predicted value 195C which follows the input value 105B among the one or more input values 105.
[0112]
number
[0113] It may include the corresponding coded data 165C, and the predicted value 195C can be generated independently of (e.g., before) decoding the coded data 165B associated with the input value 105B. In an exemplary embodiment, the input value 105C is associated with coded data 165C that corresponds to a key value (e.g., an I-frame) and can be decoded independently of the others in the set of coded data 165.
[0114]
[0126] In some implementations, the conditional input generator 182 uses the estimator 188 to calculate the input value 105B(x) based on one or more predicted values (e.g., predicted value 195A, predicted value 195C, one or more additional predicted values, or a combination thereof). b ) Estimated value 193B
[0115]
number
[0116] This generates an estimated value of 193B.
[0117]
number
[0118] Based on a set of coded data corresponding to one or more predicted values, and independent of the coded data 165B, the decoder unit 180 can generate an input value 105B(x b This corresponds to the estimated value of ).
[0119]
[0127] In certain implementations, the estimator 188 determines the estimate 193B using one or more estimation techniques (similar to the estimation techniques performed by the estimator 170 as described with reference to Figure 6). For example, the predicted value 195A corresponds to the input value 105A (e.g., the previous image frame), the predicted value 195C corresponds to the input value 105C (e.g., the subsequent image frame), and the estimate 193B corresponds to the estimated input value between the predicted values 195A and 195C. In some implementations, the estimator 188 includes a neural network configured to process one or more predicted values to generate the estimate 193B.
[0120]
[0128] The conditional input 187B gives the estimated value 193B.
[0121]
number
[0122] Predicted value: 195A
[0123]
number
[0124] Predicted value: 195C
[0125]
number
[0126] This includes one or more predicted values, such as one or more additional predicted values, or a combination thereof. As illustrated with reference to Figure 5, the feature generator 184 generates feature data 183B based on the conditional input 187B, and the decoder 186 processes the encoded data 165B based on the feature data 183B to obtain the predicted value 195B
[0127]
number
[0128] Generates a predicted value of 195B.
[0129]
number
[0130] The input value is 105B(x b This corresponds to a reconfigured version of ).
[0131]
[0129] The technical advantage of using an estimate of the available information (e.g., conditional input 167B) in the decoder section 180 to generate the encoded data 165B is that the input value 105B(x b ) a reconfigured version
[0132]
number
[0133] This may include reducing the amount of information that must be encoded as encoded data 165B in order to generate the data.
[0134]
[0130] Referring to Figure 8, an embodiment 800 of the encoder portion 160 of the compression network 140 that is operable to perform predictions corresponding to reconstruction based on auxiliary predictions is shown. For example, the encoder portion 160 takes an input value 105B(x b Predicted value 195B corresponding to the reconfigured version of )
[0135]
number
[0136] To generate the encoded data 165B that can be processed in the decoder portion 180 of the compression network 140 based on auxiliary prediction data, it is configured to generate encoded data 165B that can be processed in the decoder portion 180 of the compression network 140.
[0137]
[0131] Input value 105B(xb Predicted value 195B corresponding to the reconfigured version of )
[0138]
number
[0139] It should be understood that the compressed network 140, which uses auxiliary prediction data to generate the predicted values 195, is provided as an exemplary embodiment. In other embodiments, the compressed network 140 may use auxiliary prediction data to generate various other types of predicted values 195, such as predicted future values, collision avoidance outputs, classification outputs, and detection outputs.
[0140]
[0132] The encoder section 160 includes or is coupled to one or more auxiliary prediction layers 806. In a particular embodiment, the encoder section 160 is configured to provide an input value 105B to the encoder 166, a conditional input 167B to the feature generator 164, and at the same time provide domain-specific data 805 to one or more auxiliary prediction layers 806. One or more auxiliary prediction layers 806 process the domain-specific data 805 to generate auxiliary prediction data 807.
[0141]
[0133] The feature generator 164 processes the conditional input 167B and the auxiliary prediction data 807 to generate feature data 163B. For example, the feature layer 404A processes the auxiliary prediction data 807 and the conditional input 167B to generate feature data 163BA. The feature layer 404B processes the output of the feature layer 404A to generate feature data 163BB, and so on.
[0142]
[0134] In some implementations, the auxiliary prediction data 807 corresponds to an estimated value of the auxiliary prediction data available in the decoder section 180, which can be used to help generate a predicted value 195B corresponding to the input value 105B.
[0143]
[0135] Referring to Figure 9, an embodiment of the decoder portion 180 of the compression network 140 is shown, which is capable of performing predictions corresponding to reconstruction based on auxiliary predictions. For example, the decoder portion 180 processes encoded data 165B (generated by the encoder portion 160 as described with reference to Figure 8) based on the auxiliary prediction data to input value 105B(x b Predicted value 195B corresponding to the reconfigured version of )
[0144]
number
[0145] It is configured to generate.
[0146]
[0136] The decoder section 180 includes or is coupled to one or more auxiliary prediction layers 906. In a particular embodiment, the decoder section 180 is configured to provide coded data 165B to the decoder 186 and conditional input 187B to the feature generator 184, while simultaneously providing domain-specific data 905 to one or more auxiliary prediction layers 906. One or more auxiliary prediction layers 906 process the domain-specific data 905 to generate auxiliary prediction data 907.
[0147]
[0137] The feature generator 184 processes the conditional input 187B and the auxiliary prediction data 907 to generate feature data 183B. For example, the feature layer 504A processes the auxiliary prediction data 907 and the conditional input 187B to generate feature data 183BA. The feature layer 504B processes the output of the feature layer 506A to generate feature data 183BB, and so on.
[0148]
[0138] In some implementations, the auxiliary prediction data 907 is the same as the auxiliary prediction data 807. For example, the encoder part 160 and the decoder part 180 have access to the same auxiliary prediction data. For example, the encoder part 160 and the decoder part 180 can be contained in the same device, obtain the auxiliary prediction data from the same source, or both, as illustrated with reference to Figure 3. In some implementations, the auxiliary prediction data 907 is the predicted (e.g., reconstructed) version, the decoded version, or both of the auxiliary prediction data 807.
[0149]
[0139] In certain embodiments, auxiliary prediction data 907 can be used to help generate a prediction value 195B corresponding to an input value 105B. In a reconstruction embodiment, the auxiliary prediction data 907 may indicate the location of a person's facial features, and the prediction value 195B may correspond to a reconstructed version of the input value 105B (e.g., an image frame) representing a person's face. By accessing the location of facial features, the accuracy of the reconstruction can be improved. In a collision avoidance embodiment, the auxiliary prediction data 907 may indicate the predicted future path of a first vehicle, and the prediction value 195B may correspond to a collision avoidance output indicating whether, taking into account the predicted future path of the first vehicle, the predicted future position of a second vehicle is expected to be within a threshold distance of the object's predicted future position.
[0150]
[0140] Referring to Figure 10, an embodiment 1000 of an encoder portion 160 of a compression network that can operate to perform predictions corresponding to low latency reconstruction is shown. For example, the encoder portion 160 is configured to generate encoded data 165B corresponding to an input value 105B independently of the subsequent values of one or more input values 105 (e.g., before acquiring them), and the encoded data 165B can be processed in the decoder portion 180 of the compression network 140 independently of the encoded data associated with the subsequent values of one or more input values 105, as further described with reference to Figure 11.
[0151]
[0141] The conditional input generator 162 in Figure 1 generates a conditional input 167B corresponding to input value 105B, independently of any subsequent input value containing input value 105C (for example, before it is acquired) from among one or more input values 105. For example, the conditional input generator 162 uses the estimator 170 to generate an estimate 171B based on an estimate 169A, one or more additional estimates corresponding to one or more input values preceding input value 105B from among one or more input values 105, or a combination thereof. The conditional input 167B includes the estimate 171B, the estimate 169A, one or more additional estimates, or a combination thereof. As illustrated with reference to Figure 4, the feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes input value 105B based on the feature data 163B to generate encoded data 165B. Therefore, the encoder section 160 can generate encoded data 165B independently of (for example, before) obtaining the input value 105C. A technical advantage of generating encoded data 165B independently of the input value 105C is that it reduces the latency associated with generating encoded data 165B without having to wait for access to the input value 105C.
[0152]
[0142] Referring to Figure 11, an embodiment 1100 of a decoder portion 180 of a compressed network capable of performing predictions corresponding to low-latency reconstruction is shown. For example, the decoder portion 180 is configured to generate a predicted value 195B corresponding to an input value 105B independently of (e.g., before acquiring) a set of encoded data corresponding to the subsequent values of one or more input values 105.
[0153]
[0143] The conditional input generator 182 generates a conditional input 187B corresponding to input value 105B independently of (for example, before obtaining) the encoded data (including encoded data 165C) corresponding to the subsequent input values of one or more input values 105. The conditional input generator 182 obtains a predicted value 195A, as described with reference to Figure 1. For example, the decoder section 180 processes the encoded data 165A to generate a predicted value 195A. The conditional input generator 182 uses the estimator 188 to generate an estimated value 193B based on the predicted value 195A, one or more additional predicted values corresponding to one or more input values preceding input value 105B among the one or more input values 105, or a combination thereof. The conditional input 187B includes the estimated value 193B, the predicted value 195A, one or more additional predicted values, or a combination thereof.
[0154]
[0144] As explained with reference to Figure 5, the feature generator 184 generates feature data 183B based on the conditional input 187B, and the decoder 186 processes the encoded data 165B based on the feature data 183B. Thus, the decoder part 180 can generate the predicted value 195B independently of (for example, before) obtaining the encoded data 165C. A technical advantage of generating the predicted value 195B independently of the encoded data 165C is that it reduces the latency associated with generating the predicted value 195B without having to wait for access to the encoded data 165C.
[0155]
[0145] Referring to Figure 12, an embodiment 1200 of the encoder portion 160 of the compression network 140 is shown, which is capable of performing predictions corresponding to future predictions. For example, the encoder portion 160 performs prediction values 195C corresponding to predicted future values.
[0156]
number
[0157] The input value 105B(x) can be processed in the decoder section 180 of the compression network 140 to generate the input value 105B(x b It is configured to generate encoded data 165B corresponding to the input value 105B(x b ) corresponds to the movement value, and the predicted value is 195C
[0158]
number
[0159] This corresponds to the predicted future movement value (e.g., the predicted future movement vector).
[0160]
[0146] The conditional input generator 162 uses the local decoder section 168 to predict the value 169B
[0161]
number
[0162] It generates the following. For example, the local decoder part 168 uses the input value 105A(x a The encoded data 165A associated with the input value 105B(x b ) generates a predicted value 169B corresponding to the predicted future value. In a particular embodiment, the predicted value 169B
[0163]
number
[0164] This corresponds to an estimated predicted future value that can be generated in the decoder section 180 based on the encoded data 165A.
[0165]
[0147] In some implementations, the conditional input generator 162 generates one or more additional predicted values. The one or more additional predicted values may be associated with at least one input value prior to input value 105B among the one or more input values 105, at least one input value after input value 105B among the one or more input values 105, or both.
[0166]
[0148] Optionally, in some implementations, the conditional input generator 162 uses the estimator 170 to calculate the input value 105C(x) based on one or more predicted values (e.g., predicted value 169B). c Estimated future value of ) 171C
[0167]
number
[0168] This generates an estimated value of 171C.
[0169]
number
[0170] Based on a set of coded data corresponding to one or more predicted values, and independently of coded data 165B, the decoder unit 180 can generate the input value 105C(x c This corresponds to an estimate of the predicted future value of ).
[0171]
[0149] In certain implementations, the estimator 170 uses one or more estimation techniques to determine the estimate 171C. For example, the predicted value 169B corresponds to an estimate of the predicted future value of the input value 105B, one or more additional predicted values correspond to estimates of the predicted future values of the other input values among the one or more input values 105, and the estimate 171C corresponds to the estimated input value following the predicted value 169B. In some implementations, the estimator 170 includes a neural network configured to process one or more predicted values to generate the estimate 171C.
[0172]
[0150] The conditional input 167B is the estimated value 171C
[0173]
number
[0174] Predicted value: 169B
[0175]
number
[0176] This includes one or more predicted values, such as one or more additional predicted values, or a combination thereof. As illustrated with reference to Figure 4, the feature generator 164 generates feature data 163B based on the conditional input 167B, and the encoder 166 processes the input value 105B based on the feature data 163B to generate encoded data 165B.
[0177]
[0151] Referring to Figure 13, an embodiment 1300 of the decoder portion 180 of the compression network 140 is shown, which is capable of performing predictions corresponding to future predictions. For example, the decoder portion 180 takes an input value 105B(x b The encoded data 165B (generated in the encoder portion 160 of the compression network 140 as described with reference to Figure 12) corresponding to the predicted future value 195C is processed to obtain the predicted value 195C that corresponds to the predicted future value.
[0178]
number
[0179] It is configured to generate.
[0180]
[0152] The conditional input generator 182 predicts value 195B
[0181]
number
[0182] For example, the decoder part 180 processes the encoded data 165A to obtain the predicted value 195B.
[0183]
number
[0184] This generates the following. In a particular embodiment, the predicted value 195B is the input value 105B(x b This corresponds to the predicted future value of ).
[0185]
[0153] In some implementations, the conditional input generator 182 obtains one or more additional predicted values. The one or more additional predicted values may be associated with at least one input value prior to input value 105B among the one or more input values 105, at least one input value after input value 105B among the one or more input values 105, or both.
[0186]
[0154] Optionally, in some implementations, the conditional input generator 182 uses the estimator 188 to calculate the input value 105C(x) based on one or more predicted values (e.g., predicted value 195B). c Estimated future value of ) 193C
[0187]
number
[0188] is generated. By way of example, the estimated value 193C
[0189]
Number
[0190] is generated in the decoder portion 180 based on a set of encoded data corresponding to one or more predicted values and independently of the encoded data 165B, and corresponds to an estimated value of a predicted future value of the input value 105C(x c ).
[0191]
[0155] In certain implementations, the estimator 188 uses one or more estimation techniques (similar to the estimation techniques executed in the estimator 170 as described with reference to FIG. 12) to determine the estimated value 193C. For example, the predicted value 195B corresponds to an estimated value of a predicted future value of the input value 105B, one or more additional predicted values correspond to estimated values of predicted future values of other input values among the one or more input values 105, and the estimated value 193C corresponds to an estimated input value following the predicted value 195B. In some implementations, the estimator 188 includes a neural network configured to process one or more predicted values to generate the estimated value 193C.
[0192]
[0156] The conditional input 187B is the estimated value 193C
[0193]
Number
[0194] the predicted value 195B
[0195]
Number
[0196] This includes one or more predicted values, such as one or more additional predicted values, or a combination thereof. As illustrated with reference to Figure 5, the feature generator 184 generates feature data 183B based on the conditional input 187B, and the decoder 186 processes the encoded data 165B based on the feature data 183B to produce the predicted value 195C
[0197]
number
[0198] Generates.
[0199]
[0157] A technical advantage of using the compressed network 140 to generate predicted future values is that an action may be initiated based on the predicted future values. For example, the action may include a preventative action such as initiating braking based on the determination that the predicted value 195C corresponds to an alert condition (e.g., collision). It should be understood that the predicted value 195C corresponding to one of the predicted future values of one or more input values 105 is provided as an exemplary embodiment. In other embodiments, the predicted value 195C may correspond to a predicted future value (e.g., collision or no collision) associated with an input value 105B (e.g., image frame).
[0200]
[0158] Figures 14 to 18 illustrate an embodiment of a compression network 140 configured to perform predictions corresponding to image reconstruction. The compression network 140 includes a first compression network 140 associated with motion estimation and a second compression network 140 associated with image reconstruction based on motion estimation.
[0201]
[0159] Figure 14 includes an embodiment of the encoder portion 160 of the first compression network 140 associated with motion estimation. Figure 15 includes an embodiment of the decoder portion 180 of the first compression network 140 associated with motion estimation. Figure 17 includes an embodiment of the encoder portion 160 of the second compression network 140 associated with image reconstruction based on motion estimation. Figure 18 includes an embodiment of the decoder portion 180 of the second compression network 140 associated with image reconstruction based on motion estimation. Figure 16 includes an embodiment of image estimation that can be performed in the encoder portion 160 and decoder portion 180 of the second compression network 140 associated with image reconstruction based on motion estimation.
[0202]
[0160] In some embodiments, the encoder portion 160 of the compression network 140 includes the encoder portion 160 of the first compression network 140 and the encoder portion 160 of the second compression network 140. Similarly, the decoder portion 180 of the compression network 140 includes the decoder portion 180 of the first compression network 140 and the decoder portion 180 of the second compression network 140.
[0203]
[0161] Referring to Figure 14, an embodiment 1400 of an encoder portion 160 of a compression network 140 that is operable to perform predictions corresponding to image reconstruction is shown. For example, the encoder portion 160 is included in a first compression network 140 associated with motion estimation.
[0204]
[0162] The encoder section 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. One or more image units 1407 include image unit 1407A, image unit 1407B, image unit 1407C, one or more additional image units, or a combination thereof. In an exemplary, non-limiting embodiment, an image unit 1407 may correspond to a coding unit, a block of pixels, a frame of pixels, an image frame, or a combination thereof.
[0205]
[0163] The encoder section 160 is configured to determine one or more motion values 1405 associated with one or more image units 1407. For example, the encoder section 160 determines the image unit 1407B(x b Depending on whether it is determined that ) should be encoded, the encoder part 160 determines one or more motion values 1405 based on a comparison between image unit 1407B and one or more other image units from the one or more image units 1407. For example, depending on whether it is determined that image unit 1407B should be encoded, the encoder part 160 determines one or more motion values 1405A (m a→b ) is determined, and based on the comparison between image unit 1407C and image unit 1407B, the motion value 1405B(m c→b ) is determined, one or more additional motion values are determined, or a combination thereof is performed.
[0206]
[0164] In certain embodiments, one or more motion values 1405 represent motion vectors associated with one or more image units 1407. For example, motion value 1405A represents motion vectors associated with image units 1407A and 1407B. In certain embodiments, one or more motion values 1405 indicate one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position associated with one or more image units 1407. For example, motion value 1405A indicates one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position associated with image units 1407A and 1407B. In certain embodiments, the compression network 140 is configured to track objects associated with one or more motion values 1405 across one or more image units. In certain embodiments, one or more motion values 1405 are based on sensor data 226 received from one or more sensors 240, as described with reference to Figure 2. For example, motion value 1405A(m a→bis based on the first sensor data 226 associated with (e.g., captured simultaneously with) the image unit 1407A and the second sensor data 226 associated with (e.g., captured simultaneously with) the image unit 1407B.
[0207]
[0165] The conditional input generator 162 of the encoder portion 160 generates a conditional input 167B. For example, the conditional input generator 162 uses the local decoder portion 168 to process the encoded data associated with the image unit 1407A to generate a first predicted image unit, and uses the local decoder portion 168 to process the encoded data associated with the image unit 1407C to generate a second predicted image unit. In a particular aspect, the first predicted image unit
[0208]
Number
[0209] corresponds to an estimated value of the prediction of the image unit 1407A(x a ) that can be generated in the decoder portion 180 of FIG. 18. In a particular aspect, the second predicted image unit
[0210]
Number
[0211] corresponds to an estimated value of the prediction of the image unit 1407C(x c ) that can be generated in the decoder portion 180 of FIG. 18. The conditional input generator 162 uses the first predicted image unit
[0212]
Number
[0213] and the second predicted image unit
[0214]
number
[0215] Based on this, the estimated movement value is 1469A
[0216]
number
[0217] and estimated movement value 1469B
[0218]
number
[0219] Determine each of these. The conditional input 167B is the estimated motion value 1469A
[0220]
number
[0221] Estimated movement value: 1469B
[0222]
number
[0223] Or both.
[0224]
[0166] The feature generator 164 processes the conditional input 167B to generate feature data 163B, as described with reference to Figure 4. The encoder 166 generates motion values 1405A(m a→b ), motion value 1405B(m c→b ), Image Unit 1407B(x b ), and feature data 163B are processed to generate encoded data 165B associated with image unit 1407B. For example, the encoder layer 402A processes motion value 1405A(m a→b ), motion value 1405B(mc→b ), Image Unit 1407B(x b The encoder layer 402B processes the output of encoder layer 402A and feature data 163BB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165B (e.g., encoded data 1465B) associated with motion estimation. As further explained with reference to Figure 16, one or more motion values 1405, one or more weights, or a combination thereof are used to estimate the image unit.
[0225]
number
[0226] It can generate [this].
[0227]
[0167] Referring to Figure 15, an embodiment 1500 of a decoder portion 180 of a compression network 140 that is operable to perform predictions corresponding to image reconstruction is shown. For example, the decoder portion 180 is included in a first compression network associated with motion estimation.
[0228]
[0168] The decoder section 180 is configured to process a set of encoded data 165 to generate one or more predicted motion values 1595, one or more weights, or a combination thereof. As will be further explained with reference to Figure 16, one or more predicted motion values 1595, one or more weights, or a combination thereof are used to estimate the image unit
[0229]
number
[0230] It can generate [this].
[0231]
[0169] The conditional input generator 182 of the decoder section 180 generates a conditional input 187B. For example, the conditional input generator 182 generates an image unit 1407A(x a The first predictive image unit generated by the decoder portion 180 of the second compression network 140 is obtained by processing the encoded data associated with the second predictive image unit.
[0232]
number
[0233] To obtain. In some embodiments, the conditional input generator 182, as further described with reference to Figure 18, the image unit 1407C(x c By processing the encoded data associated with the second prediction image unit generated by the decoder portion 180 of the second compression network 140
[0234]
number
[0235] Obtain the image unit 1407A(x a ), Image Unit 1407C (x c ), or both) are image unit 1407B(x b This corresponds to a key unit (e.g., an I-frame) that can be decoded independently of decoding the encoded data associated with it.
[0236]
[0170] The conditional input generator 182 is the first predictive image unit
[0237]
number
[0238] and the second predictive image unit
[0239]
number
[0240] Based on this, the estimated movement value is 1593A
[0241]
number
[0242] and estimated movement value 1593B
[0243]
number
[0244] Each of these is determined. The conditional input 187B is the estimated motion value 1593A
[0245]
number
[0246] Estimated movement value: 1593B
[0247]
number
[0248] Or both.
[0249]
[0171] The feature generator 184 processes the conditional input 187B to generate feature data 183B. For example, the feature layer 504A generates the estimated motion value 1593A.
[0250]
number
[0251] Estimated movement value: 1593B
[0252]
number
[0253] Alternatively, both are processed to generate feature data 183BA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183BB, and so on. Decoder 186 processes the encoded data 165B (e.g., encoded data 1465B) and feature data 183B to generate predicted motion value 1595A, as explained with reference to Figure 5.
[0254]
number
[0255] Predicted movement value: 1595B
[0256]
number
[0257] Generate weights 1565 (∝), or combinations thereof.
[0258]
[0172] In a particular embodiment, the decoder portion 180 is a first predictive image unit, as will be further described with reference to Figures 16 to 18.
[0259]
number
[0260] and the second predictive image unit
[0261]
number
[0262] Estimated motion value 1593A corresponding to the motion value (e.g., motion vector) between [the specified point] and [the specified point].
[0263]
number
[0264] and estimated movement value 1593B
[0265]
number
[0266] Use as conditional input 187B, and image unit 1407B(x b ) Predictive image unit associated with
[0267]
number
[0268] Predicted motion value 1595A can be used to generate this.
[0269]
number
[0270] Predicted movement value: 1595B
[0271]
number
[0272] And generate weights 1565 (∝).
[0273]
[0173] Referring to Figure 16, an embodiment of an image estimator 1600 of a compression network 140 that can operate to perform predictions corresponding to image reconstruction is shown. For example, the image estimator 1600 is included in the estimator 170 of the encoder portion 160, the estimator 188 of the decoder portion 180, or both, of a second compression network 140 associated with image reconstruction.
[0274]
[0174] The image estimator 1600 predicts motion value 1695A
[0275]
number
[0276] Based on this, predictive image unit 1607A
[0277]
number
[0278] Perform warp 1602A and estimate image unit 1609A
[0279]
number
[0280] The image estimator 1600 generates the predicted motion value 1695B.
[0281]
number
[0282] Based on this, predictive image unit 1607C
[0283]
number
[0284] Perform warp 1602B and estimate image unit 1609B
[0285]
number
[0286] The image estimator 1600 generates the estimated image unit 1609A.
[0287]
number
[0288] Estimated image unit 1609B
[0289]
number
[0290] Perform synthesis 1604 with the estimated image unit 1671B
[0291]
number
[0292] This generates the estimated image unit 1671B.
[0293]
number
[0294] This is the estimated image unit 1609A
[0295]
number
[0296] Estimated image unit 1609B
[0297]
number
[0298] This corresponds to weighted synthesis with respect to the first weight (e.g., weight 1665) of the estimated image unit 1609A.
[0299]
number
[0300] Applying this to generate a first weighted image unit, the second weight (e.g., 1 - weight 1665) is used to estimate image unit 1609B
[0301]
number
[0302] Applying this to generate a second weighted image unit, the first weighted image unit and the second image unit are combined to obtain the estimated image unit 1671B.
[0303]
number
[0304] Generates.
[0305]
[0175] An image estimator 1600 that generates an estimated image unit 1671B based on two predicted image units 1607 is provided as an exemplary embodiment. In other embodiments, the image estimator 1600 can generate an estimated image unit 1671B based on three or more predicted image units 1607. For example, the image estimator 1600 can warp one or more additional predicted image units 1607 to generate one or more additional estimated image units and synthesize the estimated image units based on various weights to generate the estimated image unit 1671B.
[0306]
[0176] In a particular embodiment, the encoder portion 160 of the second compression network 140 uses an image estimator 1600 to determine the first estimated image unit corresponding to the image unit 1407B, as will be further explained with reference to Figure 17.
[0307]
number
[0308] This generates the second estimated image unit corresponding to the image unit 1407B, using the image estimator 1600, as further described with reference to Figure 18.
[0309]
number
[0310] Generates.
[0311]
[0177] Referring to Figure 17, an embodiment 1700 of an encoder portion 160 of a compression network 140 that is operable to perform predictions corresponding to image reconstruction is shown. For example, the encoder portion 160 corresponds to an image encoder portion of a second compression network 140 associated with image reconstruction based on motion estimation.
[0312]
[0178] The encoder section 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The conditional input generator 162 uses the local decoder section 168 to predict the image unit 1769A
[0313]
number
[0314] It generates the image unit 1407A(x a The encoded data associated with the predictive image unit 1769A is processed.
[0315]
number
[0316] This generates the predictive image unit 1769A in a particular embodiment.
[0317]
number
[0318] This corresponds to the estimated value of the predicted image unit associated with the image unit 1407A that can be generated in the decoder section 180 of Figure 18. Similarly, the conditional input generator 162 uses the local decoder section 168 to predict the image unit 1769C
[0319]
number
[0320] It generates the image unit 1407C(x c The encoded data associated with the predictive image unit 1769C is processed to make the image unit 1769C
[0321]
number
[0322] Generates.
[0323]
[0179] The conditional input generator 162 uses the estimator 170 (e.g., the image estimator 1600) to generate motion values 1405A (m) generated by motion estimation performed by the encoder section 160 in Figure 14. a→b ) and motion value 1405B(m c→b Based on this, estimated image unit 1771B
[0324]
number
[0325] This generates the following: For example, predictive image unit 1769A
[0326]
number
[0327] This is the predictive image unit 1607A
[0328]
number
[0329] It corresponds to the predictive image unit 1769C
[0330]
number
[0331] This is the predictive image unit 1607C
[0332]
number
[0333] Corresponding to a motion value of 1405A(m a→b ) is the predicted movement value 1695A
[0334]
number
[0335] Corresponding to a motion value of 1405B(m c→b ) is predicted to be 1695B
[0336]
number
[0337] Corresponds to the following. The conditional input generator 162 also provides weights as weights 1665. In certain embodiments, the weights are based on configuration settings, default data, user input, or a combination thereof. Estimation image unit 1671B
[0338]
number
[0339] This is the estimated image unit 1771B
[0340]
number
[0341] This corresponds to the predictive image unit 1769A.
[0342]
number
[0343] Estimated image unit 1771B
[0344]
number
[0345] and predictive image unit 1769C
[0346]
number
[0347] Includes.
[0348]
[0180] The feature generator 164 processes the conditional input 167B to generate feature data 163B, as further explained with reference to Figure 4. For example, the feature layer 404A is the predictive image unit 1769A
[0349]
number
[0350] Estimated image unit 1771B
[0351]
number
[0352] and predictive image unit 1769C
[0353]
number
[0354] The output of feature layer 404A is processed to generate feature data 163BA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163BB, and so on. Encoder 166 processes the image unit 1407B(x b ) and feature data 163B are processed to generate encoded data 165B. For example, the encoder layer 402A processes the image unit 1407B (x b The encoder layer 402B processes the output of encoder layer 402A and feature data 163BB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165B (e.g., encoded data 1765B) associated with image reconstruction.
[0355]
[0181] Referring to Figure 18, an embodiment 1800 of a decoder portion 180 of a compression network 140 that is operable to perform predictions corresponding to image reconstruction is shown. For example, the decoder portion 180 corresponds to an image decoder portion of a second compression network 140 associated with image reconstruction based on motion estimation.
[0356]
[0182] The decoder section 180 is configured to process a set of encoded data 165 to generate one or more predictive image units 1895. The conditional input generator 182 is configured to process the predictive image unit 1895A
[0357]
number
[0358] To obtain: For example, the decoder part 180 is image unit 1407A(x a The encoded data associated with the predictive image unit 1895A is processed to make the prediction image unit 1895A
[0359]
number
[0360] Similarly, the conditional input generator 182 generates the predictive image unit 1895C.
[0361]
number
[0362] To obtain: For example, the decoder part 180 is image unit 1407C(x c The encoded data associated with the predictive image unit 1895C is processed to make the image unit 1895C
[0363]
number
[0364] Generates.
[0365]
[0183] The conditional input generator 182 uses the estimator 188 (e.g., the image estimator 1600) to generate the predicted motion value 1595A generated by the motion estimation performed by the decoder section 180 in Figure 15.
[0366]
number
[0367] and predicted movement value 1595B
[0368]
number
[0369] Based on this, estimated image unit 1893B
[0370]
number
[0371] It generates the following: For example, predictive image unit 1895A
[0372]
number
[0373] This is the predictive image unit 1607A
[0374]
number
[0375] It corresponds to the predictive image unit 1895C
[0376]
number
[0377] This is the predictive image unit 1607C
[0378]
number
[0379] Corresponding to the predicted movement value of 1595A
[0380]
number
[0381] The predicted movement value is 1695A.
[0382]
number
[0383] Corresponding to a predicted movement value of 1595B
[0384]
number
[0385] The predicted movement value is 1695B.
[0386]
number
[0387] Corresponding to this, weight 1565(∝) in Figure 15 corresponds to weight 1665. Estimated image unit 1671B
[0388]
number
[0389] This is the estimated image unit 1893B
[0390]
number
[0391] This corresponds to the predictive image unit 1895A.
[0392]
number
[0393] Estimated image unit 1893B
[0394]
number
[0395] and predictive image unit 1895C
[0396]
number
[0397] Includes.
[0398]
[0184] The feature generator 184 processes the conditional input 187B to generate feature data 183B, as further explained with reference to Figure 5. For example, the feature layer 504A is the predictive image unit 1895A
[0399]
number
[0400] Estimated image unit 1893B
[0401]
number
[0402] Predictive image unit 1895C
[0403]
number
[0404] The feature layer 504B processes the output of feature layer 504A to generate feature data 183BB, and so on. The decoder 186 processes the encoded data 165B (e.g., encoded data 1765B) and feature data 183B, as explained with reference to Figure 5, to produce the predictive image unit 1895B.
[0405]
number
[0406] The technical advantages of using the compression network 140 to process the image unit 1407B to generate encoded data 1465B and encoded data 1765B, and processing the encoded data 1465B and encoded data 1765B to generate the predictive image unit 1895B, include reducing the size of encoded data 1465B, encoded data 1765B, or both.
[0407]
[0185] Figures 19 to 23 illustrate an embodiment of a compression network 140 configured to perform predictions corresponding to low-latency image reconstruction. The compression network 140 includes a first compression network 140 associated with low-latency motion estimation and a second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.
[0408]
[0186] Figure 19 includes an embodiment of the encoder portion 160 of the first compression network 140 associated with low-latency motion estimation. Figure 20 includes an embodiment of the decoder portion 180 of the first compression network 140 associated with low-latency motion estimation. Figure 22 includes an embodiment of the encoder portion 160 of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation. Figure 23 includes an embodiment of the decoder portion 180 of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation. Figure 21 includes an embodiment of image estimation that can be performed in the encoder portion 160 and decoder portion 180 of the second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.
[0409]
[0187] In some embodiments, the encoder portion 160 of the compression network 140 includes the encoder portion 160 of the first compression network 140 and the encoder portion 160 of the second compression network 140. Similarly, the decoder portion 180 of the compression network 140 includes the decoder portion 180 of the first compression network 140 and the decoder portion 180 of the second compression network 140.
[0410]
[0188] Referring to Figure 19, this is a diagram illustrating an exemplary embodiment of an encoder portion 160 of a compression network 140 that is operable to perform predictions corresponding to low-latency image reconstruction. For example, the encoder portion 160 is included in a first compression network 140 associated with low-latency motion estimation.
[0411]
[0189] The encoder section 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The encoder section 160 is configured to determine one or more motion values 1905 associated with one or more image units 1407. For example, the encoder section 160 determines one or more image units 1407D(x d In response to determining that ) should be encoded, the encoder part 160 determines one or more motion values 1905 based on a comparison between image unit 1407D and one or more other image units from one or more image units 1407. For example, in response to determining that image unit 1407D should be encoded, the encoder part 160 determines that image unit 1407C(x c ) and image unit 1407D(x d Based on a comparison with ), the motion value is 1905A(m c→d The system determines the following and, based on a comparison between image unit 1407D and one or more image units that precede image unit 1407D among the one or more image units 1407, determines one or more additional motion values, or a combination thereof.
[0412]
[0190] In certain embodiments, one or more motion values 1905 represent motion vectors associated with one or more image units 1407. For example, motion value 1905A(m c→d ) represents the motion vector associated with image unit 1407C and image unit 1407D. In certain embodiments, one or more motion values 1905 represent one or more of the linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position associated with one or more image units 1407. For example, motion value 1905A represents one or more of the linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position associated with image unit 1407C and image unit 1407D. In certain embodiments, one or more motion values 1905 are based on sensor data 226 received from one or more sensors 240, as described with reference to Figure 2. For example, motion value 1905A(m c→d This is based on first sensor data 226 associated with image unit 1407C (for example, captured simultaneously with it) and second sensor data 226 associated with image unit 1407D (for example, captured simultaneously with it).
[0413]
[0191] The conditional input generator 162 of the encoder section 160 generates a conditional input 167D. For example, the conditional input generator 162 uses the local decoder section 168 to process the encoded data associated with the image unit 1407A to generate a first predicted image unit
[0414]
number
[0415] The local decoder section 168 is used to process the encoded data associated with the image unit 1407B and generate a second predicted image unit.
[0416]
number
[0417] It generates and uses the local decoder section 168 to process the encoded data associated with the image unit 1407C to produce a third predicted image unit.
[0418]
number
[0419] Generates a first predictive image unit. In a particular embodiment, the first predictive image unit
[0420]
number
[0421] This is an image unit 1407A(x) that can be generated in the decoder section 180 of Figure 23. a The conditional input generator 162 corresponds to the estimated value of the prediction of the second prediction image unit.
[0422]
number
[0423] and the third predictive image unit
[0424]
number
[0425] Based on this, the estimated movement value is 1969A
[0426]
number
[0427] The conditional input generator 162 determines the first predictive image unit.
[0428]
number
[0429] and the second predictive image unit
[0430]
number
[0431] Based on this, the estimated movement value is 1969B.
[0432]
number
[0433] Determine the following. The conditional input 167D is the estimated motion value 1969A.
[0434]
number
[0435] Estimated movement value: 1969B
[0436]
number
[0437] Or both.
[0438]
[0192] The feature generator 164 processes the conditional input 167D to generate feature data 163D, as described with reference to Figure 4. For example, feature layer 404A processes the estimated motion values 1969A and 1969B to generate feature data 163DA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163DB, and so on. Feature data 163D includes one or more additional sets of feature data, including feature data 163DA, feature data 163DB, feature data 163DC, and feature data 163DN, or a combination thereof. Encoder 166 processes the motion value 1905A (m c→d ), Image Unit 1407D (x d ), and feature data 163D are processed to generate encoded data 165D associated with image unit 1407D. For example, the encoder layer 402A processes motion values 1905A (m c→d ), Image Unit 1407D (x d The encoder layer 402B processes the output of encoder layer 402A and feature data 163DB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165D (e.g., encoded data 1965D) associated with motion estimation. As further explained with reference to Figure 21, one or more motion values 1905, one or more weights, or a combination thereof are used to estimate the image unit.
[0439]
number
[0440] It can generate [this].
[0441]
[0193] Thus, the encoder portion 160 can generate encoded data 165D independently of (for example, before) acquiring any image unit following image unit 1407D among the one or more image units 1407. A technical advantage of generating encoded data 165D independently of subsequent image units is that it reduces the latency associated with generating encoded data 165D without having to wait for access to subsequent image units.
[0442]
[0194] Referring to Figure 20, an embodiment 2000 of a decoder portion 180 of a compression network 140 capable of performing predictions corresponding to low-latency image reconstruction is shown. For example, the decoder portion 180 is included in a first compression network associated with low-latency motion estimation.
[0443]
[0195] The decoder section 180 is configured to process a set of encoded data 165 to generate one or more predicted motion values 2095, one or more weights, or a combination thereof. As further explained with reference to Figure 21, one or more predicted motion values 2095, one or more weights, or a combination thereof are used to estimate the image unit
[0444]
number
[0445] It can generate [this].
[0446]
[0196] The conditional input generator 182 of the decoder section 180 generates a conditional input 187D. For example, the conditional input generator 182 generates an image unit 1407A(x a The first predictive image unit generated by the decoder portion 180 of the second compression network 140 is obtained by processing the encoded data associated with the second predictive image unit.
[0447]
number
[0448] Similarly, in some embodiments, the conditional input generator 182 obtains the image unit 1407B(x b ) and image unit 1407C(x c ) A second predictive image unit associated with each
[0449]
number
[0450] and the third predictive image unit
[0451]
number
[0452] Obtain the image unit 1407A(x a ), Image Unit 1407B(x b ), Image Unit 1407C (x c ), or a combination thereof, is an encoded image unit 1407D(x d It is before ).
[0453]
[0197] The conditional input generator 182 is the second predictive image unit
[0454]
number
[0455] and the third predictive image unit
[0456]
number
[0457] Based on the comparison, the estimated movement value is 2093A.
[0458]
number
[0459] The conditional input generator 182 determines the first predictive image unit.
[0460]
number
[0461] and the second predictive image unit
[0462]
number
[0463] Based on the comparison, the estimated movement value is 2093B.
[0464]
number
[0465] Determine the following. The conditional input 187D is the estimated motion value 2093A.
[0466]
number
[0467] Estimated movement value: 2093B
[0468]
number
[0469] Or both.
[0470]
[0198] The feature generator 184 processes the conditional input 187D to generate feature data 183D, as explained with reference to Figure 5. For example, the feature layer 504A generates the estimated motion value 2093A
[0471]
number
[0472] Estimated movement value: 2093B
[0473]
number
[0474] Alternatively, both are processed to generate feature data 183DA. Feature layer 504B processes the output of feature layer 504A to generate feature data 183DB, and so on. Feature data 183D includes one or more additional sets of feature data, including feature data 183DA, feature data 183DB, feature data 183DC, and feature data 183DN, or a combination thereof. Decoder 186 processes encoded data 165D (e.g., encoded data 1965D) and feature data 183D to obtain the predicted motion value 2095A, as described with reference to Figure 5.
[0475]
number
[0476] It generates weights 2065 (∝), or both. For example, decoder layer 508N processes encoded data 165D and feature data 183DN to produce an output. Decoder layer 508B processes the output of decoder layer 508C and feature data 183DB to produce an output, and so on. The output of decoder layer 508A is the predicted motion value 2095A associated with image unit 1407D.
[0477]
number
[0478] This corresponds to a weight of 2065 (∝), or both.
[0479]
[0199] In a particular embodiment, the decoder section 180, as further described with reference to Figures 21-23, estimates motion values 2093A corresponding to motion values (e.g., motion vectors) between the predicted image units in front of the image unit 1407D.
[0480]
number
[0481] and estimated movement value 2093B
[0482]
number
[0483] Use as conditional input 187D, and image unit 1407D(x d ) Predictive image unit associated with
[0484]
number
[0485] Predicted motion value 2095A can be used to generate this.
[0486]
number
[0487] And generate weight 2065(∝). Predict motion value 2095A regardless of the encoded data associated with any image unit after image unit 1407D.
[0488]
number
[0489] The technical advantage of generating this is the predicted movement value 2095A
[0490]
number
[0491] This could include reducing the latency associated with generating it.
[0492]
[0200] Referring to Figure 21, an embodiment of an image estimator 2100 of a compression network 140 that can operate to perform predictions corresponding to low-latency image reconstruction is shown. For example, the image estimator 2100 is included in the estimator 170 of the encoder portion 160, the estimator 188 of the decoder portion 180, or both, of a second compression network 140 associated with low-latency image reconstruction.
[0493]
[0201] The image estimator 2100 predicts motion value 2195A
[0494]
number
[0495] Based on this, predictive image unit 2107C
[0496]
number
[0497] Perform warp 2102 and estimate image unit 2109A
[0498]
number
[0499] The image estimator 2100 generates the predicted image unit 2107C.
[0500]
number
[0501] Estimated image unit 2109B
[0502]
number
[0503] It is used as follows: The image estimator 2100 is the estimated image unit 2109A
[0504]
number
[0505] Estimated image unit 2109B
[0506]
number
[0507] Perform synthesis 2104 with the estimated image unit 2171D
[0508]
number
[0509] It generates the estimated image unit 2171D.
[0510]
number
[0511] This is the estimated image unit 2109A
[0512]
number
[0513] Estimated image unit 2109B
[0514]
number
[0515] This corresponds to weighted synthesis with respect to the first weight (e.g., weight 2165) of the estimated image unit 2109A.
[0516]
number
[0517] Applying this to generate a first weighted image unit, the second weight (e.g., 1 - weight 2165) is used to estimate image unit 2109B
[0518]
number
[0519] Applying this to generate a second weighted image unit, the first weighted image unit and the second image unit are combined to obtain the estimated image unit 2171D
[0520]
number
[0521] Generates.
[0522]
[0202] An image estimator 2100 that generates an estimated image unit 2171D based on a single predicted image unit 2107 is provided as an exemplary embodiment. In other embodiments, the image estimator 2100 can generate an estimated image unit 2171D based on multiple predicted image units 2107. For example, the image estimator 2100 can warp one or more additional predicted image units 2107 to generate one or more additional estimated image units, and then synthesize the estimated image units based on various weights to generate the estimated image unit 2171D.
[0523]
[0203] In a particular embodiment, the encoder portion 160 of the second compression network 140 uses the image estimator 2100 to determine the first estimated image unit corresponding to the image unit 1407D, as will be further explained with reference to Figure 22.
[0524]
number
[0525] This generates the second estimated image unit corresponding to the image unit 1407D, using the image estimator 2100, as further described with reference to Figure 23.
[0526]
number
[0527] Generates.
[0528]
[0204] Referring to Figure 22, an embodiment 2200 of an encoder portion 160 of a compression network 140 that can operate to perform predictions corresponding to low-latency image reconstruction is shown. For example, the encoder portion 160 corresponds to an image encoder portion of a second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.
[0529]
[0205] The encoder section 160 is configured to process one or more image units 1407 to generate a set of encoded data 165. The conditional input generator 162 uses the local decoder section 168 to predict the image unit 2269C
[0530]
number
[0531] It generates the image unit 1407C(x c The encoded data associated with the predictive image unit 2269C is processed.
[0532]
number
[0533] This generates the predictive image unit 2269C in a specific embodiment.
[0534]
number
[0535] This corresponds to the estimated value of the predicted image unit associated with the image unit 1407C that can be generated in the decoder section 180 of Figure 23.
[0536]
[0206] The conditional input generator 162 uses the estimator 170 (e.g., the image estimator 2100) to generate motion values 1905A(m) generated by motion estimation performed in the encoder section 160 of Figure 19. c→d Based on this, estimated image unit 2271D
[0537]
number
[0538] It generates the following: For example, the predictive image unit 2269C
[0539]
number
[0540] This is the predictive image unit 2107C
[0541]
number
[0542] Corresponding to a motion value of 1905A(m c→d ) is the predicted movement value 2195A
[0543]
number
[0544] Corresponds to the following. The conditional input generator 162 also provides weights as weights 2165. In certain embodiments, the weights are based on configuration settings, default data, user input, or a combination thereof. Estimation image unit 2171D
[0545]
number
[0546] This is the estimated image unit 2271D
[0547]
number
[0548] This corresponds to the conditional input 167D associated with image unit 1407D, which corresponds to the predictive image unit 2269C.
[0549]
number
[0550] and estimated image unit 2271D
[0551]
number
[0552] Includes.
[0553]
[0207] The feature generator 164 processes the conditional input 167D to generate feature data 163D, as described with reference to Figure 4. For example, feature layer 404A processes the conditional input 167D to generate feature data 163DA. Feature layer 404B processes the output of feature layer 404A to generate feature data 163DB, and so on. Feature data 163D includes one or more additional sets of feature data, including feature data 163DA, feature data 163DB, feature data 163DC, and feature data 163DN, or a combination thereof. The encoder 166 processes the image unit 1407D(x d ) processes to generate encoded data 165D. For example, the encoder layer 402A processes the image unit 1407D(x d The encoder layer 402B processes the output of encoder layer 402A and feature data 163DB to generate an output, and so on. The output of encoder layer 402N corresponds to encoded data 165D (e.g., encoded data 2265D) associated with low-latency image reconstruction.
[0554]
[0208] Thus, the encoder portion 160 can generate encoded data 165D independently of (for example, before) acquiring any image unit following image unit 1407D among the one or more image units 1407. A technical advantage of generating encoded data 165D independently of subsequent image units is that it reduces the latency associated with generating encoded data 165D without having to wait for access to subsequent image units.
[0555]
[0209] Referring to Figure 23, an embodiment 2300 of a decoder portion 180 of a compression network 140 that can operate to perform predictions corresponding to low-latency image reconstruction is shown. For example, the decoder portion 180 corresponds to an image decoder portion of a second compression network 140 associated with low-latency image reconstruction based on low-latency motion estimation.
[0556]
[0210] The decoder section 180 is configured to process a set of encoded data 165 to generate one or more predictive image units 2395. The conditional input generator 182 is configured to generate predictive image units 2395C
[0557]
number
[0558] To obtain: For example, the decoder part 180 is image unit 1407C(x c The encoded data associated with the predictive image unit 2395C is processed.
[0559]
number
[0560] Generates.
[0561]
[0211] The conditional input generator 182 uses the estimator 188 (e.g., the image estimator 2100) to generate the predicted motion value 2095A generated by the motion estimation performed in the decoder section 180 of Figure 20.
[0562]
number
[0563] Based on this, estimated image unit 2393D
[0564]
number
[0565] It generates the following: For example, the predictive image unit 2395C
[0566]
number
[0567] This is the predictive image unit 2107C
[0568]
number
[0569] Corresponding to the predicted movement value of 2095A
[0570]
number
[0571] The predicted movement value is 2195A.
[0572]
number
[0573] Corresponding to this, weight 2065(∝) in Figure 20 corresponds to weight 2165. Estimated image unit 2171D
[0574]
number
[0575] This is the estimated image unit 2393D
[0576]
number
[0577] This corresponds to the conditional input 187D associated with image unit 1407D, which corresponds to the predictive image unit 2395C.
[0578]
number
[0579] and estimated image unit 2393D
[0580]
number
[0581] Includes.
[0582]
[0212] The feature generator 184 processes the conditional input 187D to generate feature data 183D, as described with reference to Figure 5. For example, the feature layer 504A is the predictive image unit 2395C
[0583]
number
[0584] and estimated image unit 2393D
[0585]
number
[0586] The feature layer 504B processes the output of feature layer 504A to generate feature data 183DB, and so on. Feature data 183D includes one or more additional sets of feature data, including feature data 183DA, feature data 183DB, feature data 183DC, and feature data 183DN, or a combination thereof. Decoder 186 processes the encoded data 165B (e.g., encoded data 2265D) and feature data 183D as described with reference to Figure 5 to generate the predictive image unit 2395D
[0587]
number
[0588] The technical advantages of using the compression network 140 include reducing the size of the encoded data 1965D, the encoded data 2265D, or both. The predicted image unit 2395D is generated independently of the encoded data of any subsequent image unit among one or more image units 1407.
[0589]
number
[0590] The technical advantage of generating this is the predictive image unit 2395D
[0591]
number
[0592] This could include reducing the latency associated with generating it.
[0593]
[0213] Figure 24 shows an implementation configuration 2400 of the integrated circuit 2402, which includes one or more processors 2490. One or more processors 2490 include compression network components 2460, such as an encoder portion 160, a decoder portion 180, or both of the compression network 140 in Figure 1. The integrated circuit 2402 also includes signal inputs 2404, such as one or more bus interfaces, which enable input data 2428 to be received for processing. The integrated circuit 2402 also includes signal outputs 2406, such as bus interfaces, which enable transmission of output data 2450. For example, in an implementation configuration in which the compression network component 2460 includes an encoder portion 160, the input data 2428 may include input values (one or more) 105, and the output data 2450 may include a set of encoded data 165. In an implementation where the compressed network component 2460 includes a decoder section 180, the input data 2428 may include a set of encoded data 165, and the output data 2450 may include a predicted value(single or multiple) 195. In an implementation where the compressed network component 2460 includes an encoder section 160 and a decoder section 180, the input data 2428 may include an input value(single or multiple) 105, and the output data 2450 may include a predicted value(single or multiple) 195. The integrated circuit 2402 enables implementations of at least a portion of the compressed network 140 as a component in various systems such as a mobile phone or tablet shown in Figure 25, a headset shown in Figure 26, a wearable electronic device shown in Figure 27, a voice-controlled speaker system shown in Figure 28, a camera shown in Figure 29, a virtual reality headset, mixed reality headset, or augmented reality headset shown in Figure 30, or a vehicle shown in Figure 31 or Figure 32.
[0594]
[0214] Figure 25 shows an implementation form 2500 in which the compression network component 2460 is implemented in a mobile device 2502 such as a telephone or tablet, as an exemplary, non-limiting embodiment. The mobile device 2502 includes one or more microphones 2510, one or more speakers 2520, and a display screen 2504. In addition, the compression network component 2460 is shown using dotted lines to indicate internal components that are built into the mobile device 2502 and are generally not visible to the user of the mobile device 2502. In certain embodiments, the compression network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode video data captured by the camera of the mobile device 2502 for transmission to another device, decode video data received from another device, or a combination thereof. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in the mobile device 2502 and to decode the data when retrieved from the storage device.
[0595]
[0215] Figure 26 shows an implementation configuration 2600 in which the compression network component 2460 is implemented in a headset device 2602. The headset device 2602 includes a microphone 2610 and a camera 2620. In certain embodiments, the compression network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data such as audio data captured by the microphone 2610, video data captured by the camera 2620, and / or head tracking data (e.g., sensor data from the IMU of the headset device 2602) for transmission to another device, and to decode data such as audio data received from another device for playback in the earphone speaker of the headset device 2602, or a combination thereof. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in the headset device 2602 and to decode the data when retrieved from the storage device.
[0596]
[0216] Figure 27 shows an implementation form 2700 in which the compressed network component 2460 is implemented in a wearable electronic device 2702, referred to as a “smartwatch”. The compressed network component 2460, microphone 2710, speaker 2720, and display screen 2704 are incorporated into the wearable electronic device 2702. In certain embodiments, the compressed network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compressed network component 2460 operates to encode data such as audio data captured by the microphone 2710 for transmission to another device. Alternatively, the compressed network component 2460 operates to decode data such as audio data received from another device for playback on the speaker 2720, or video data for playback on the display screen 2704, or both. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in the wearable electronic device 2702 and to decode the data when retrieved from the storage device.
[0597]
[0217] In certain embodiments, the wearable electronic device 2702 includes a haptic device that provides haptic notifications (e.g., vibrates) in response to the detection of activity associated with the operation of the compression network component 2460. For example, the haptic notification may cause the user to look at the wearable electronic device 2702 to see a displayed notification that data (e.g., audio or video data) has been received from a remote device and is available for playback on the wearable electronic device 2702. Thus, the wearable electronic device 2702 can alert a user who is hearing impaired or who is wearing a headset to such a notification.
[0598]
[0218] Figure 28 shows an implementation configuration 2800 in which the compressed network component 2460 is implemented in a wireless speaker and voice activation device 2802. The wireless speaker and voice activation device 2802 can have wireless network connectivity and is configured to perform assistant operations. One or more processors 2890, a microphone 2810, a camera 2820, a speaker 2804, or a combination thereof, including the compressed network component 2460, are included in the wireless speaker and voice activation device 2802. In certain embodiments, the compressed network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode video data captured by camera 2820 for transmission to another device, decode audio data for playback at speaker 2804, decode video data received from another device for playback on the display screen (not shown) of wireless speaker and audio activation device 2802, or a combination thereof. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in wireless speaker and audio activation device 2802 and decode the data upon retrieval from storage.
[0599]
[0219] During operation, in response to receiving a verbal command from a user via the microphone 2810, the wireless speaker and voice activation device 2802 can perform an assistant operation through, for example, the execution of a voice activation system (such as an integrated assistant application). The assistant operation can include adjusting the temperature, playing music, turning on the lighting, etc. For example, the assistant operation can include starting the transmission of data to a remote device, receiving data from a remote device, and / or storing / retrieving data from the local memory of the wireless speaker and voice activation device 2802, and each of these can be performed more efficiently by the operation of the compression network component 2460.
[0600]
[0220] FIG. 29 shows an implementation form 2900 in which the compression network component 2460 is implemented in a portable electronic device corresponding to the camera device 2902. The compression network component 2460 and the microphone 2910 are incorporated in the camera device 2902. In a specific embodiment, the compression network component 2460 operates to improve the coding efficiency and reduce the amount of resources used for the transmission or storage of the encoded data. In an exemplary non-limiting embodiment, the compression network component 2460 operates to encode data captured in the camera device 2902, such as image data, video data, audio data, or a combination thereof, for transmission to another device. In another exemplary non-limiting embodiment, the compression network component 2460 operates to encode data such as video or image data captured in the camera device 2902 for storage in the memory of the camera device 2902, and also to decode the data when retrieved from the memory.
[0601]
[0221] Figure 30 shows an implementation form 3000 in which the compression network component 2460 is implemented in a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 3002. The compression network component 2460, microphone 3010, and camera 3020 are incorporated into the headset 3002. A visual interface device, such as a display screen, is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the headset 3002 is being worn. In certain embodiments, the compression network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In exemplary, non-limiting embodiments, the compression network component 2460 operates to decode data received in the headset 3002, such as video data, audio data, other data associated with providing a virtual, augmented, or mixed reality experience to the user of the headset 3002, or a combination thereof. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data captured in the headset 3002, such as head tracking data from the headset 3002's IMU, for transmission to a remote device (e.g., a virtual reality (VR) session server). In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data such as video, audio, or image data captured in or received by the headset 3002 for storage in the headset 3002's memory, and also operates to decode the data upon retrieval from memory.
[0602]
[0222] Figure 31 shows an implementation configuration 3100 in which the compressed network component 2460 corresponds to or is incorporated within the vehicle 3102, which is shown as a manned or unmanned aerial device (e.g., a package delivery drone). The compressed network component 2460, microphone 3110, speaker 3120, camera 3104, or a combination thereof, is incorporated into the vehicle 3102. User voice activity detection can be performed based on audio signals received from the microphone 3110 of the vehicle 3102, such as delivery orders from an authorized user of the vehicle 3102.
[0603]
[0223] In certain embodiments, the compression network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode video data captured by the camera 3104 of the vehicle 3102 for transmission to another device, decode video data received from another device, or a combination thereof. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in the vehicle 3102 and decode the data when retrieved from the storage device.
[0604]
[0224] Figure 32 shows another implementation form 3200 in which the compressed network component 2460 corresponds to or is incorporated within a vehicle 3202, which is shown as a car. The vehicle 3202 includes one or more processors 290, one or more processors 292, one or more processors 390, or a combination thereof, including the compressed network component 2460. The vehicle 3202 also includes a microphone 3210, a camera 3204, or both.
[0605]
[0225] User voice activity detection can be performed based on audio signals received from the microphone 3210 of the vehicle 3202. In some implementations, user voice activity detection can be performed based on audio signals received from an internal microphone (e.g., microphone 3210), such as for voice commands from an authorized occupant. For example, user voice activity detection can be used to detect voice commands from the operator of the vehicle 3202 (e.g., from a parent to set the volume to 5 or to set the destination of the autonomous vehicle) and ignore voices from other occupants (e.g., from a child to set the volume to 10 or from another occupant discussing a different location). In some implementations, user voice activity detection can be performed based on audio signals received from an external microphone (e.g., microphone 3210), such as from an authorized user of the vehicle. In certain implementations, upon receiving a verbal command identified as a user utterance, the voice activation system initiates one or more actions of the vehicle 3202 based on one or more detected keywords (e.g., "unlock," "start engine," "play music," "show weather forecast," or another voice command), such as by providing feedback or information via the display 3220 or one or more speakers (e.g., speaker 3230).
[0606]
[0226] In certain embodiments, the compression network component 2460 operates to improve coding efficiency and reduce the amount of resources used to transmit or store encoded data. In an exemplary, non-limiting embodiment, the compression network component 2460 operates to encode video data captured by the camera 3204 of the vehicle 3202 for transmission to another device, decode video data received from another device, or a combination thereof. In another exemplary, non-limiting embodiment, the compression network component 2460 operates to encode data for storage in the vehicle 3202 and decode the data when retrieved from the storage device.
[0607]
[0227] The camera 3204 can capture one or more image frames while the vehicle 3202 is in motion. The compressed network component 2460 can process one or more motion values (e.g., one or more input values 105) associated with one or more image frames to generate a set of encoded data 165. One or more motion values may include one or more image frames, one or more velocity measurements, one or more acceleration measurements, other sensor data, the navigation route of the vehicle 3202, or a combination thereof. The vehicle 3202 can store the encoded data 165 in the vehicle 3202, send the set of encoded data 165 to another device (e.g., a server), or both.
[0608]
[0228] In some implementations, the compressed network component 2460 generates one or more predicted values 195 based on a set of encoded data 165. In a collision avoidance embodiment, the compressed network component 2460 can track an object in one or more image frames. The compressed network component 2460 generates one or more predicted values 195 corresponding to predicted future motion values. For example, one or more predicted values 195 indicate whether the predicted future position of vehicle 3202 is within a distance threshold of the predicted future position of an object. In some implementations, upon determining that the predicted future position of vehicle 3202 is within a distance threshold of the predicted future position of an object, the compressed network component 2460 initiates one or more collision avoidance actions, such as generating an alert, initiating braking, activating an alarm, sending an alert to an emergency vehicle, or a combination thereof.
[0609]
[0229] Referring to Figure 33, a specific implementation of method 3300 for generating encoded data using a compression network is shown. In a particular embodiment, one or more operations of method 3300 are performed by at least one of the following: conditional input generator 162, feature generator 164, encoder 166, local decoder portion 168, estimator 170, encoder portion 160, compression network 140, one or more processors 290, device 202, system 200 in Figure 2, one or more processors 390, device 302, system 300 in Figure 3, compression network component 2460 in Figure 24, or a combination thereof.
[0610]
[0230] Method 3300 includes, in 3302, obtaining a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. For example, the conditional input generator 162 in Figure 1 obtains a conditional input 167B to the compression network 140, as described with reference to Figure 1. The conditional input 167B is based on a predicted value 169A (e.g., a predicted motion value).
[0611]
[0231] Method 3300 also includes, in 3304, using a compression network to process a conditional input and one or more motion values to generate encoded data associated with one or more motion values. For example, the encoder section 160 uses the feature generator 164 of the compression network 140 to process the conditional input 167B to generate feature data 163B, and uses the encoder 166 of the compression network 140 to process the feature data 163B and the input value 105B to generate encoded data 165B associated with the input value 105B.
[0612]
[0232] Therefore, the method 3300 makes it possible to reduce the amount of information that should be provided to the decoder section 180 as encoded data 165A by generating encoded data 165A based on an estimate (e.g., conditional input 167B) of the information that can be generated in the decoder section 180 (e.g., conditional input 187B).
[0613]
[0233] The method 3300 in Figure 33 can be performed by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. In one embodiment, the method 3300 in Figure 33 can be performed by a processor that executes instructions as described with reference to Figure 35.
[0614]
[0234] Referring to Figure 34, a specific implementation of Method 3400 for generating predictive data from encoded data using a compression network is shown. In a particular embodiment, one or more operations of Method 3400 are performed by at least one of the following: the conditional input generator 182, feature generator 184, decoder 186, estimator 188, decoder section 180, compression network 140, one or more processors 292, device 260, system 200 in Figure 2, one or more processors 390, device 302, system 300 in Figure 3, compression network component 2460 in Figure 24, or a combination thereof.
[0615]
[0235] Method 3400 includes obtaining encoded data associated with one or more motion values in 3402. For example, the decoder section 180 obtains encoded data 165B associated with the input value 105B from the encoder section 160, as described with reference to Figure 1.
[0616]
[0236] Method 3400 also includes, in 3404, obtaining a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. For example, the conditional input generator 182 obtains a conditional input 187B to the compression network 140, as described with reference to Figure 1. The conditional input 187B is based on a predicted value 195A.
[0617]
[0237] Method 3400 further includes, in 3406, using a compression network to process the encoded data and conditional input to generate one or more second predicted motion values. For example, the decoder section 180 uses the feature generator 184 of the compression network 140 to process the conditional input 187B to generate feature data 183B, and uses the decoder 186 of the compression network 140 to process the feature data 183B and the encoded data 165B to generate predicted value 195B.
[0618]
[0238] Therefore, the method 3400 makes it possible to reduce the amount of information that should be acquired by the decoder part 180 as encoded data 165A by generating predicted values 195B based on information that can be estimated in the encoder part 160 (e.g., conditional input 187B) (e.g., conditional input 167B).
[0619]
[0239] Method 3400 in Figure 34 can be performed by an FPGA device, an ASIC, a CPU, a DSP or other processing unit, a controller, another hardware device, a firmware device, or any combination thereof. In one embodiment, Method 3400 in Figure 34 can be performed by a processor that executes instructions as described with reference to Figure 35.
[0620]
[0240] Referring to Figure 35, a block diagram of a specific exemplary implementation of the device is shown, which is designated as 3500 overall. In various implementations, device 3500 may have more or fewer components than those shown in Figure 35. In the exemplary implementation, device 3500 may correspond to one or more of device 202, device 260 in Figure 2, or device 302 in Figure 3. In the exemplary implementation, device 3500 may perform one or more operations described with reference to Figures 1 to 34.
[0621]
[0241] In certain implementations, device 3500 includes a processor 3506 (e.g., a CPU). Device 3500 may include one or more additional processors 3510 (e.g., one or more DSPs). In certain embodiments, one or more processors 290, one or more processors 292 in Figure 2, one or more processors 390 in Figure 3, or a combination thereof, correspond to processor 3506, processor 3510, or a combination thereof. Processor 3510 may include a speech and music coder decoder (codec) 3508, which may include a speech coder ("vocoder") encoder 3536, a vocoder decoder 3538, or both. Processor 3510 may include a compression network component 2460, one or more applications 262, or a combination thereof.
[0622]
[0242] Device 3500 may include memory 3586 and codec 3534. Memory 3586 may include instructions 3556, which are executable by one or more additional processors 3510 (or processor 3506) to perform the functions described with reference to the Compression Network Component 2460. Device 3500 may include modem 3570 coupled to antenna 3552 via transceiver 3550. In certain embodiments, modem 3570 corresponds to modem 270, modem 280, or both in Figure 2.
[0623]
[0243] Device 3500 may include a display 3528 coupled to a display controller 3526. One or more speakers 3520, one or more microphones 3524, or a combination thereof may be coupled to a codec 3534. Codec 3534 may include a digital-to-analog converter (DAC) 3502, an analog-to-digital converter (ADC) 3504, or both. In a particular implementation, codec 3534 may receive analog signals from one or more microphones 3524, convert the analog signals to digital signals using the analog-to-digital converter 3504, and provide the digital signals to a speech and music codec 3508. The speech and music codec 3508 may process the digital signals, which may be further processed by a compression network component 2460, one or more applications 262, or a combination thereof. In certain implementations, the speech and music codec 3508 may provide a digital signal to the codec 3534. The codec 3534 may use the digital-to-analog converter 3502 to convert the digital signal to an analog signal, and provide the analog signal to one or more speakers 3520.
[0624]
[0244] In certain implementations, device 3500 may be included in a system-in-package or system-on-chip device 3522. In certain implementations, memory 3586, processor 3506, processor 3510, display controller 3526, codec 3534, and modem 3570 are included in a system-in-package or system-on-chip device 3522. In certain implementations, input device 3530, one or more sensors 240, and power supply 3544 are coupled to the system-in-package or system-on-chip device 3522. Furthermore, in certain implementations, as shown in Figure 35, display 3528, input device 3530, one or more speakers 3520, one or more microphones 3524, one or more sensors 240, antenna 3552, and power supply 3544 are outside the system-in-package or system-on-chip device 3522. In certain implementations, each of the display 3528, input device 3530, one or more speakers 3520, one or more microphones 3524, one or more sensors 240, antenna 3552, and power supply 3544 may be coupled to components of a system-in-package or system-on-chip device 3522, such as an interface or controller.
[0625]
[0245] Device 3500 may include smart speakers, speaker covers, mobile communication devices, smartphones, cellular phones, laptop computers, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headsets, augmented reality headsets, mixed reality headsets, virtual reality headsets, aviation vehicles, home automation systems, voice activation devices, wireless speakers and voice activation devices, portable electronic devices, cars, computing devices, communication devices, Internet of Things (IoT) devices, VR devices, extended reality (XR) devices, base stations, mobile devices, or any combination thereof.
[0626]
[0246] In relation to the described implementation, the device includes means for acquiring encoded data associated with one or more motion values. For example, the means for acquiring encoded data may correspond to the decoder section 180, encoder section 160, compression network 140 in Figure 1, modem 280 in Figure 2, one or more processors 292, device 260, system 200, storage device 392, device 302, system 300, processor 3506, processor 3510, antenna 3552, transceiver 3550, modem 3570 in Figure 3, one or more other circuits or components configured to acquire encoded data associated with one or more motion values, or any combination thereof.
[0627]
[0247] The apparatus also includes means for acquiring a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. For example, the means for acquiring a conditional input may correspond to the estimator 188, conditional input generator 182, decoder section 180, compression network 140 in Figure 1, one or more processors 292, device 260, system 200 in Figure 2, device 302, system 300 in Figure 3, image estimator 1600 in Figure 16, image estimator 2100 in Figure 21, processor 3506, processor 3510, one or more other circuits or components configured to acquire a conditional input, or any combination thereof.
[0628]
[0248] The apparatus further includes means for processing encoded data and conditional inputs using a compression network to generate one or more second predicted motion values. For example, the means for processing encoded data and conditional inputs may correspond to the feature generator 184, decoder 186, decoder section 180, compression network 140, one or more processors 292, device 260, system 200 in Figure 2, device 302, system 300, processor 3506, processor 3510 in Figure 3, one or more other circuits or components configured to process encoded data and conditional inputs, or any combination thereof.
[0629]
[0249] In addition, in relation to the implementation described, the device includes means for acquiring a conditional input to a compression network, the conditional input being based on one or more first predicted motion values. For example, the means for acquiring a conditional input may correspond to the local decoder section 168, estimator 170, conditional input generator 162, encoder section 160, compression network 140 in Figure 1, one or more processors 290, device 202, system 200 in Figure 2, device 302, system 300 in Figure 3, image estimator 1600 in Figure 16, image estimator 2100 in Figure 21, processor 3506, processor 3510, one or more other circuits or components configured to acquire a conditional input, or any combination thereof.
[0630]
[0250] The apparatus also includes means for processing a conditional input and one or more motion values using a compression network to generate encoded data associated with one or more motion values. For example, the means for processing a conditional input and one or more motion values may correspond to the feature generator 164, encoder 166, encoder section 160, compression network 140, one or more processors 290, device 202, system 200 in Figure 2, device 302, system 300, processor 3506, processor 3510 in Figure 3, one or more other circuits or components configured to process a conditional input and one or more motion values, or any combination thereof.
[0631]
[0251] In some implementations, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 3586) includes an instruction (e.g., instruction 3556) that, when executed by one or more processors (e.g., one or more processors 3510 or processor 3506), causes one or more processors to obtain encoded data (e.g., encoded data 165B) associated with one or more motion values (e.g., input value 105B). The instruction also, when executed by one or more processors, causes one or more processors to obtain a conditional input (e.g., conditional input 187B) of a compression network (e.g., compression network 140), the conditional input being based on one or more first predicted motion values (e.g., predicted value 195A). The instruction, when executed by one or more processors, causes one or more processors to use a compression network to process the encoded data and conditional input to generate one or more second predicted motion values (e.g., predicted value 195B).
[0632]
[0252] In some implementations, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 3586) includes an instruction (e.g., instruction 3556) that, when executed by one or more processors (e.g., one or more processors 3510 or processor 3506), causes one or more processors to acquire a conditional input (e.g., conditional input 167B) of a compressed network (e.g., compressed network 140), the conditional input being based on one or more first predicted motion values (e.g., predicted value 169A). The instruction also, when executed by one or more processors, causes one or more processors to use the compressed network to process the conditional input and one or more motion values (e.g., input value 105B) to generate encoded data (e.g., encoded data 165B) associated with one or more motion values.
[0633]
[0253] Specific aspects of the present disclosure are described below in a set of interrelated embodiments.
[0634]
[0254] According to Embodiment 1, the device includes one or more processors configured to acquire encoded data associated with one or more motion values, acquire a conditional input to a compression network which is based on one or more first predicted motion values, and process the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.
[0635]
[0255] Example 2 includes the device described in Example 1, wherein one or more motion values are based on the output of one or more sensors.
[0636]
[0256] Example 3 includes the device described in Example 1 or Example 2, wherein one or more sensors include an inertial measuring unit (IMU).
[0637]
[0257] Example 4 includes the device described in any one of Examples 1 to 3, wherein one or more motion values represent one or more motion vectors associated with one or more image units.
[0638]
[0258] Example 5 includes the device described in Example 4, wherein one of the one or more image units includes a coding unit.
[0639]
[0259] Example 6 includes the device described in Example 4 or Example 5, wherein one of the one or more image units includes a block of pixels.
[0640]
[0260] Example 7 includes the device described in any one of Examples 4 to 6, wherein one of the one or more image units includes a frame of pixels.
[0641]
[0261] Example 8 includes the device described in any one of Examples 1 to 7, wherein one or more second predicted motion values represent future motion vectors.
[0642]
[0262] Example 9 includes the device described in any one of Examples 1 to 8, wherein one or more second predicted motion values correspond to reconstructed versions of one or more motion values.
[0643]
[0263] Example 10 includes the device described in any one of Examples 1 to 9, wherein one or more processors are incorporated into at least one of a headset, a mobile communication device, an extended reality (XR) device, or a vehicle.
[0644]
[0264] Example 11 includes the device described in any one of Examples 1 to 10, wherein one or more motion values represent one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.
[0645]
[0265] Example 12 includes the device described in any one of Examples 1 to 11, and the compression network includes a neural network having multiple layers.
[0646]
[0266] Example 13 includes the device described in any one of Examples 1 to 12, wherein the compression network includes a video decoder, and the video decoder has multiple decoder layers configured to decode multiple resolutions of encoded data associated with one or more motion values.
[0647]
[0267] Example 14 includes the device described in any one of Examples 1 to 13, wherein one or more processors are configured to track an object associated with one or more motion values across one or more frames of pixels.
[0648]
[0268] Example 15 includes the device described in Example 14, wherein one or more second predictive motion values represent collision avoidance outputs associated with the vehicle.
[0649]
[0269] Example 16 includes the device described in Example 15, wherein the collision avoidance output indicates the predicted future position of the vehicle relative to an object.
[0650]
[0270] Example 17 includes the device described in Example 15 or Example 16, wherein the collision avoidance output indicates the predicted future position of the vehicle and the predicted future position of the object.
[0651]
[0271] Example 18 includes the device described in any one of Examples 1 to 17, and further includes a modem configured to receive a bitstream from an encoder device, the bitstream including encoded data.
[0652]
[0272] Example 19 includes the device described in any one of Examples 1 to 18, wherein one or more processors are configured to process encoded data and feature data to generate one or more second predicted motion values, the feature data being based on a conditional input.
[0653]
[0273] Example 20 includes the device described in Example 19, and further includes a modem configured to receive a bitstream from the encoder device, the bitstream including feature data and encoded data.
[0654]
[0274] Example 21 includes the device described in Example 19 or Example 20, wherein one or more processors are configured to process conditional inputs using a compression network to generate feature data.
[0655]
[0275] Example 22 includes the device described in any one of Examples 19 to 21, and the feature data corresponds to multiscale feature data having different spatial resolutions.
[0656]
[0276] Example 23 includes the device described in any one of Examples 19 to 22, and the feature data includes multiscale wavelet transform data.
[0657]
[0277] Example 24 includes the device described in any one of Examples 1 to 23, wherein the encoded data is an image unit (e.g., x b ) including motion estimation encoded data associated with it, one or more processors reconstruct the previous image unit (e.g.,
[0658]
number
[0659] and the reconstructed subsequent image unit (for example,
[0660]
number
[0661] Based on the comparison with, the estimated movement value (for example,
[0662]
number
[0663] The system is configured to determine, process estimated motion values to generate motion estimation feature data corresponding to the image unit, and use a compression network to process the motion estimation encoded data and motion estimation feature data to generate one or more second predicted motion values, where one or more second predicted motion values are one or more reconstructed motion values corresponding to one or more motion values (for example,
[0664]
number
[0665] Includes.
[0666]
[0278] Example 25 includes the device described in Example 24, wherein the encoded data includes re-encoded encoded data, and one or more processors include one or more re-encoded motion values (e.g.,
[0667]
number
[0668] Based on this, estimated image units (for example,
[0669]
number
[0670] Generates an image, processes the estimated image unit to generate reconstructed feature data corresponding to the image unit, and uses a compression network to process the reconstructed encoded data and reconstructed feature data to reconstruct the image unit (for example,
[0671]
number
[0672] It is configured to generate.
[0673]
[0279] Example 26 includes the device described in Example 25, wherein one or more first predicted motion values are first reconfigured motion values (e.g.,
[0674]
number
[0675] and a second reconfigured motion value (for example,
[0676]
number
[0677] Including, one or more processors, a first reconfigured motion value (e.g.,
[0678]
number
[0679] The previous image unit that was reconfigured (for example,
[0680]
number
[0681] Based on the application to the first estimated version of the image unit, a second reconstructed motion value (e.g.,
[0682]
number
[0683] The subsequent image unit that was reconstructed (for example,
[0684]
number
[0685] Based on the application to the second estimated version of the image unit, the estimated image unit (for example,
[0686]
number
[0687] It is configured to generate.
[0688]
[0280] Example 27 includes the device described in any one of Examples 1 to 23, wherein the encoded data is an image unit (e.g., x d Including motion estimation encoded data associated with ), one or more processors process the first reconfigured previous image unit (e.g.,
[0689]
number
[0690] and the second reconfigured previous image unit (for example,
[0691]
number
[0692] Based on the comparison with the first estimated movement value (for example,
[0693]
number
[0694] The system is configured to determine, process at least a first estimated motion value to generate motion estimation feature data corresponding to the image unit, and use a compression network to process the motion estimation coded data and motion estimation feature data to generate one or more second predicted motion values, where one or more second predicted motion values are one or more reconstructed motion values corresponding to one or more motion values (e.g.,
[0695]
number
[0696] Includes.
[0697]
[0281] Example 28 includes the device described in Example 27, wherein the encoded data includes reconstructed encoded data, and one or more processors process one or more reconstructed motion values (e.g.,
[0698]
number
[0699] Based on this, estimated image units (for example,
[0700]
number
[0701] Generates an image, processes the estimated image unit to generate reconstructed feature data corresponding to the image unit, and uses a compression network to process the reconstructed encoded data and reconstructed feature data to reconstruct the image unit (for example,
[0702]
number
[0703] It is configured to generate.
[0704]
[0282] Example 29 includes the device described in Example 28, wherein one or more processors are second reconfigured previous image units (e.g.,
[0705]
number
[0706] and the third reconfigured previous image unit (for example,
[0707]
number
[0708] Based on the comparison with the second estimated movement value (for example,
[0709]
number
[0710] It is configured to determine that the motion estimation feature data corresponding to the image unit is further based on processing a second estimated motion value.
[0711]
[0283] Example 30 includes the device described in Example 28 or Example 29, wherein one or more first predicted motion values are first reconfigured motion values (e.g.,
[0712]
number
[0713] Including, one or more processors, a first reconfigured motion value (e.g.,
[0714]
number
[0715] to the first reconfigured previous image unit (for example,
[0716]
number
[0717] Based on its application, estimated image units (e.g.,
[0718]
number
[0719] It is configured to determine the reconstructed feature data, and the estimated image units (for example,
[0720]
number
[0721] Based on.
[0722]
[0284] According to Example 31, the method includes: obtaining encoded data associated with one or more motion values in a device; obtaining a conditional input to a compression network in the device, which is based on one or more first predicted motion values; and processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.
[0723]
[0285] Example 32 includes the method described in Example 31, wherein one or more motion values are based on the output of one or more sensors.
[0724]
[0286] Example 33 comprises the method described in Example 31 or Example 32, wherein one or more sensors include an inertial measuring unit (IMU).
[0725]
[0287] Example 34 includes the method described in any one of Examples 31 to 33, wherein one or more motion values represent one or more motion vectors associated with one or more image units.
[0726]
[0288] Example 35 includes the method described in Example 34, wherein one of the one or more image units includes a coding unit.
[0727]
[0289] Example 36 comprises the method described in Example 34 or Example 35, wherein one of the one or more image units includes a block of pixels.
[0728]
[0290] Example 37 comprises the method described in any one of Examples 34 to 36, wherein one of the one or more image units includes a frame of pixels.
[0729]
[0291] Example 38 comprises the method described in any one of Examples 31 to 37, wherein one or more second predicted motion values represent future motion vectors.
[0730]
[0292] Example 39 comprises the method described in any one of Examples 31 to 38, wherein one or more second predicted motion values correspond to reconstructed versions of one or more motion values.
[0731]
[0293] Example 40 comprises the method described in any one of Examples 31 to 39, wherein the device comprises at least one of a headset, a mobile communication device, an extended reality (XR) device, or a vehicle.
[0732]
[0294] Example 41 includes the method described in any one of Examples 31 to 40, wherein one or more motion values represent one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.
[0733]
[0295] Example 42 includes the method described in any one of Examples 31 to 41, wherein the compression network includes a neural network having multiple layers.
[0734]
[0296] Example 43 comprises the method described in any one of Examples 31 to 42, wherein the compression network comprises a video decoder, the video decoder having multiple decoder layers configured to decode multiple resolutions of encoded data associated with one or more motion values.
[0735]
[0297] Example 44 comprises the method described in any one of Examples 31 to 43, and further comprises tracking an object associated with one or more motion values across one or more frames of pixels.
[0736]
[0298] Example 45 includes the method described in Example 44, wherein one or more second predicted motion values represent collision avoidance outputs associated with the vehicle.
[0737]
[0299] Example 46 includes the method described in Example 45, wherein the collision avoidance output indicates the predicted future position of the vehicle relative to an object.
[0738]
[0300] Example 47 includes the method described in Example 45 or Example 46, wherein the collision avoidance output indicates the predicted future position of the vehicle and the predicted future position of the object.
[0739]
[0301] Example 48 comprises the method described in any one of Examples 31 to 47, further comprising receiving a bitstream from an encoder device via a modem, the bitstream comprising encoded data.
[0740]
[0302] Example 49 comprises the method described in any one of Examples 31 to 48, further comprising processing encoded data and feature data to generate one or more second predicted motion values, wherein the feature data is based on a conditional input.
[0741]
[0303] Example 50 includes the method of embodiment 49, further including receiving a bitstream from an encoder device via a modem, the bitstream including feature data and encoded data.
[0742]
[0304] Example 51 includes the method described in Example 49 or Example 50, and further includes processing a conditional input using a compression network to generate feature data.
[0743]
[0305] Example 52 includes the method described in any one of Examples 31 to 51, wherein processing the encoded data and conditional input includes processing the conditional input using a compression network to generate feature data, and processing the encoded data and feature data to generate one or more second predicted motion values.
[0744]
[0306] Example 53 includes the method described in any one of Examples 49 to 52, wherein the feature data corresponds to multiscale feature data having different spatial resolutions.
[0745]
[0307] Example 54 includes the method described in any one of Examples 49 to 53, wherein the feature data includes multiscale wavelet transform data.
[0746]
[0308] Example 55 includes the method described in any one of Examples 31 to 54, wherein the reconstructed previous image unit (for example,
[0747]
number
[0748] and the reconstructed subsequent image unit (for example,
[0749]
number
[0750] Based on the comparison with, the estimated movement value (for example,
[0751]
number
[0752] The determination is made that the encoded data is an image unit (for example, x bThe process includes determining, including motion estimation coded data associated with, processing the estimated motion values to generate motion estimation feature data corresponding to the image unit, and using a compression network to process the motion estimation coded data and motion estimation feature data to generate one or more second predicted motion values, where one or more second predicted motion values are one or more reconstructed motion values corresponding to one or more motion values (e.g.,
[0753]
number
[0754] Includes.
[0755]
[0309] Example 56 includes the method described in Example 55, and comprises one or more reconstructed motion values (e.g.,
[0756]
number
[0757] Based on this, estimated image units (for example,
[0758]
number
[0759] The process involves generating an image unit, processing the estimated image unit to generate reconstructed feature data corresponding to the image unit, and using a compression network to process the reconstructed encoded data and the reconstructed feature data to obtain a reconstructed image unit (for example,
[0760]
number
[0761] This further includes generating the encoded data, and the encoded data includes the reconstructed encoded data.
[0762]
[0310] Example 57 includes the method described in Example 56, wherein the first reconfigured motion value (for example,
[0763]
number
[0764] The previous image unit that was reconfigured (for example,
[0765]
number
[0766] Determining a first estimated version of an image unit based on application to the first estimated motion values, wherein one or more first predicted motion values are first reconstructed motion values (e.g.,
[0767]
number
[0768] This includes making a determination and a second reconstructed motion value (for example,
[0769]
number
[0770] The subsequent image unit that was reconstructed (for example,
[0771]
number
[0772] Determining a second estimated version of an image unit based on its application to the second reconstructed motion value (e.g.,
[0773]
number
[0774] This includes determining the estimated image unit (for example,
[0775]
number
[0776] This includes generating and further includes.
[0777]
[0311] Example 58 includes the method described in any one of Examples 31 to 54, wherein the first reconfigured previous image unit (for example,
[0778]
number
[0779] and the second reconfigured previous image unit (for example,
[0780]
number
[0781] Based on the comparison with the first estimated movement value (for example,
[0782]
number
[0783] The determination is made that the encoded data is an image unit (for example, x dThe process further includes determining, including motion estimation coded data associated with, at least a first estimated motion value to generate motion estimation feature data corresponding to the image unit, and processing the motion estimation coded data and motion estimation feature data using a compression network to generate one or more second predicted motion values, wherein one or more second predicted motion values are one or more reconstructed motion values corresponding to one or more motion values (e.g.,
[0784]
number
[0785] Includes.
[0786]
[0312] Example 59 includes the method described in Example 58, and comprises one or more reconstructed motion values (e.g.,
[0787]
number
[0788] Based on this, estimated image units (for example,
[0789]
number
[0790] The process involves generating an image unit, processing the estimated image unit to generate reconstructed feature data corresponding to the image unit, and using a compression network to process the reconstructed encoded data and the reconstructed feature data to obtain a reconstructed image unit (for example,
[0791]
number
[0792] This further includes generating the encoded data, and the encoded data includes the reconstructed encoded data.
[0793]
[0313] Example 60 includes the method described in Example 59, wherein a second reconfigured previous image unit (for example,
[0794]
number
[0795] and the third reconfigured previous image unit (for example,
[0796]
number
[0797] Based on the comparison with the second estimated movement value (for example,
[0798]
number
[0799] This further includes determining that the motion estimation feature data corresponding to the image unit is further based on processing a second estimated motion value.
[0800]
[0314] Example 61 includes the method described in Example 59 or Example 60, and a first reconfigured motion value (for example,
[0801]
number
[0802] to the first reconfigured previous image unit (for example,
[0803]
number
[0804] Based on its application, estimated image units (e.g.,
[0805]
number
[0806] The process further includes determining that one or more first predicted motion values are first reconstructed motion values (e.g.,
[0807]
number
[0808] The reconstructed feature data includes estimated image units (e.g.,
[0809]
number
[0810] Based on.
[0811]
[0315] According to Example 62, the device includes a memory configured to store instructions and a processor configured to execute instructions in order to perform the method described in any one of Examples 31 to 61.
[0812]
[0316] According to Example 63, when executed by a processor, the non-temporary computer-readable medium stores instructions that cause the processor to perform the method described in any one of Examples 31 to 61.
[0813]
[0317] According to Example 64, the apparatus includes means for carrying out the method described in any one of Examples 31 to 61.
[0814]
[0318] According to Example 65, the device includes one or more processors configured to take a conditional input, which is a conditional input to a compression network, based on one or more first predicted motion values, and to process the conditional input and one or more motion values using the compression network to generate encoded data associated with one or more motion values.
[0815]
[0319] Example 66 includes the device described in Example 65, wherein one or more motion values are based on the output of one or more sensors.
[0816]
[0320] Example 67 includes the device described in Example 66, wherein one or more sensors include an inertial measuring unit (IMU).
[0817]
[0321] Example 68 includes the device described in any one of Examples 65 to 67, wherein one or more motion values represent one or more motion vectors associated with one or more image units.
[0818]
[0322] Example 69 includes the device described in Example 68, wherein one of the one or more image units includes a coding unit.
[0819]
[0323] Example 70 includes the device described in Example 68 or Example 69, wherein one of the one or more image units includes a block of pixels.
[0820]
[0324] Example 71 includes the device described in any one of Examples 68 to 70, wherein one of the one or more image units includes a frame of pixels.
[0821]
[0325] Example 72 includes the device described in any one of Examples 65 to 71, and further includes a modem configured to transmit a bitstream to a decoder device, the bitstream including encoded data.
[0822]
[0326] Example 73 includes the device described in any one of Examples 65 to 72, wherein one or more processors are incorporated into at least one of a headset, a mobile communication device, an extended reality (XR) device, or a vehicle.
[0823]
[0327] Example 74 includes the device described in any one of Examples 65 to 73, wherein one or more motion values represent one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.
[0824]
[0328] Example 75 includes the device described in any one of Examples 65 to 74, and the compression network includes a neural network having multiple layers.
[0825]
[0329] Example 76 includes the device described in any one of Examples 65 to 75, wherein the compression network includes a video encoder, and the video encoder has multiple encoder layers configured to encode video data at a higher resolution associated with one or more motion values.
[0826]
[0330] Example 77 includes the device described in any one of Examples 65 to 76, and further includes a modem configured to transmit a bitstream to a decoder device, the bitstream including encoded data.
[0827]
[0331] Example 78 includes the device described in any one of Examples 65 to 77, wherein one or more processors are configured to process a conditional input using a compression network to generate feature data, and to process one or more motion values and feature data using a compression network to generate encoded data.
[0828]
[0332] Example 79 includes the device described in Example 78 and further includes a modem configured to transmit a bitstream to a decoder device, the bitstream including feature data and encoded data.
[0829]
[0333] Example 80 includes the device described in Example 78 or Example 79, and the feature data includes multi-scale feature data having different spatial resolutions.
[0830]
[0334] Example 81 includes the device described in any one of Examples 78 to 80, and the feature data includes multi-scale wavelet transform data.
[0831]
[0335] Example 82 includes the device described in any one of Examples 65 to 81, and one or more processors determine a first motion value (e.g., m b ) among one or more motion values based on a comparison between an image unit (e.g., x a ) and a previous image unit (e.g., x a→b ), determine a second motion value (e.g., m c ) among one or more motion values based on a comparison between the image unit and a subsequent image unit (e.g., x c→b ), and based on a comparison between the reconstructed previous image unit (e.g.,
[0832]
Number
[0833] and the reconstructed subsequent image unit (e.g.,
[0834]
Number
[0835] determine an estimated motion value (e.g.,
[0836]
Number
[0837] The system is configured to determine the motion, process the estimated motion values to generate motion estimation feature data corresponding to the image unit, and use a compression network to process the motion estimation feature data, the first motion value, the second motion value, and the image unit to generate motion estimation encoded data associated with the image unit, the encoded data including the motion estimation encoded data.
[0838]
[0336] Example 83 includes the device described in Example 82, wherein one or more first predicted motion values are first reconfigured motion values (e.g.,
[0839]
number
[0840] and a second reconfigured motion value (for example,
[0841]
number
[0842] Including, one or more processors, a first reconfigured motion value (e.g.,
[0843]
number
[0844] The previous image unit that was reconfigured (for example,
[0845]
number
[0846] Based on the application to the first estimated version of the image unit, a second reconstructed motion value (e.g.,
[0847]
number
[0848] The subsequent image unit that was reconstructed (for example,
[0849]
number
[0850] Based on the application to the second estimated version of the image unit, the estimated image unit (for example,
[0851]
number
[0852] Generates an estimated image unit (for example,
[0853]
number
[0854] The system is configured to process the image units and the reconstructed feature data to generate reconstructed encoded data associated with the image units, and to use a compression network to process the image units and the reconstructed feature data to generate reconstructed encoded data, which includes the reconstructed encoded data.
[0855]
[0337] Example 84 includes the device described in any one of Examples 65 to 81, wherein one or more processors are image units (e.g., x d ) and the previous image unit (for example, x c Based on a comparison with ), one of the one or more motion values (e.g., m c→d) determines the first reconfigured previous image unit (for example,
[0856]
number
[0857] and the second reconfigured previous element (for example,
[0858]
number
[0859] Based on the comparison with the first estimated movement value (for example,
[0860]
number
[0861] The system is configured to determine, process at least a first estimated motion value to generate motion estimation feature data corresponding to the image unit, and use a compression network to process the motion estimation feature data, motion value, and image unit to generate motion estimation encoded data associated with the image unit, the encoded data including motion estimation encoded data.
[0862]
[0338] Example 85 includes the device described in Example 84, wherein one or more processors are second reconfigured previous elements (e.g.,
[0863]
number
[0864] and the third reconfigured previous element (for example,
[0865]
number
[0866] Based on the comparison with the second estimated movement value (for example,
[0867]
number
[0868] It is configured to determine that the motion estimation feature data corresponding to the image unit is further based on processing a second estimated motion value.
[0869]
[0339] Example 86 includes the device described in Example 84 or Example 85, wherein one or more first predicted motion values are first reconfigured motion values (e.g.,
[0870]
number
[0871] Including, one or more processors, a first reconfigured motion value (e.g.,
[0872]
number
[0873] to the first reconfigured previous image unit (for example,
[0874]
number
[0875] Based on its application, estimated image units (e.g.,
[0876]
number
[0877] Determine and estimate image units (for example,
[0878]
number
[0879] The system is configured to process the image units and the reconstructed feature data to generate reconstructed encoded data associated with the image units, and to use a compression network to process the image units and the reconstructed feature data to generate reconstructed encoded data, which includes the reconstructed encoded data.
[0880]
[0340] According to Example 87, the method includes, in a device, obtaining a conditional input which is a conditional input to a compression network, based on one or more first predicted motion values, and using the compression network to process the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0881]
[0341] Example 88 includes the method described in Example 87, and further includes, in a device, processing a conditional input using a compression network to generate feature data, and processing one or more motion values and feature data using a compression network to generate encoded data.
[0882]
[0342] Example 89 includes the method described in Example 87 or Example 88, wherein the feature data includes multiscale feature data having different spatial resolutions.
[0883]
[0343] Example 90 comprises the method described in any one of Examples 87 to 89, wherein one or more motion values are based on the output of one or more sensors.
[0884]
[0344] Example 91 comprises the method described in Example 90, wherein one or more sensors include an inertial measuring unit (IMU).
[0885]
[0345] Example 92 includes the method described in any one of Examples 87 to 91, wherein one or more motion values represent one or more motion vectors associated with one or more image units.
[0886]
[0346] Example 93 includes the method described in Example 92, wherein one of the one or more image units includes a coding unit.
[0887]
[0347] Example 94 includes the method described in Example 92 or Example 93, wherein one of the one or more image units includes a block of pixels.
[0888]
[0348] Example 95 comprises the method described in any one of Examples 92 to 94, wherein one of the one or more image units includes a frame of pixels.
[0889]
[0349] Example 96 comprises the method described in any one of Examples 87 to 95, and further comprises transmitting a bitstream to a decoder device via a modem, wherein the bitstream comprises encoded data.
[0890]
[0350] Example 97 comprises the method described in any one of Examples 87 to 96, wherein the device comprises at least one of a headset, a mobile communication device, an extended reality (XR) device, or a vehicle.
[0891]
[0351] Example 98 includes the method described in any one of Examples 87 to 97, wherein one or more motion values represent one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.
[0892]
[0352] Example 99 includes the method described in any one of Examples 87 to 98, wherein the compression network includes a neural network having multiple layers.
[0893]
[0353] Example 100 includes the method according to any one of Examples 97 to 99, the compression network includes a video encoder, and the video encoder has a plurality of encoder layers configured to encode a higher resolution of video data associated with one or more motion values.
[0894]
[0354] Example 101 includes the method according to any one of Examples 87 to 100, and further includes transmitting a bitstream to a decoder device via a modem, the bitstream including encoded data.
[0895]
[0355] Example 102 includes the method according to any one of Examples 87 to 101, and further includes using a compression network to process conditional inputs to generate feature data, and using the compression network to process one or more motion values and the feature data to generate encoded data.
[0896]
[0356] Example 103 includes the method according to Example 102, and further includes transmitting a bitstream to a decoder device via a modem, the bitstream including feature data and encoded data.
[0897]
[0357] Example 104 includes the method according to Example 102 or Example 103, and the feature data includes multi-scale feature data having different spatial resolutions. [[ID=!17]]
[0898]
[0358] Example 105 includes the method according to any one of Examples 102 to 104, and the feature data includes multi-scale wavelet transform data.
[0899]
[0359] Example 106 includes the method according to any one of Examples 87 to 105, and based on the comparison between an image unit (e.g., x b ) and the previous image unit (e.g., x a ), the first motion value (e.g., m a→b) and to determine the image unit and the subsequent image unit (for example, x c Based on a comparison with ), the second motion value among one or more motion values (for example, m c→b ) and to determine the reconstructed previous image unit (for example,
[0900]
number
[0901] and the reconstructed subsequent image unit (for example,
[0902]
number
[0903] Based on the comparison with, the estimated movement value (for example,
[0904]
number
[0905] The process includes determining, processing estimated motion values to generate motion estimation feature data corresponding to the image unit, and using a compressed network to process the motion estimation feature data, first motion values, second motion values, and image unit to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes motion estimation encoded data.
[0906]
[0360] Example 107 includes the method described in Example 106, wherein the first reconfigured motion value (for example,
[0907]
number
[0908] The previous image unit that was reconfigured (for example,
[0909]
number
[0910] Determining a first estimated version of an image unit based on application to the first estimated motion values, wherein one or more first predicted motion values are first reconstructed motion values (e.g.,
[0911]
number
[0912] This includes making a determination and a second reconstructed motion value (for example,
[0913]
number
[0914] The subsequent image unit that was reconstructed (for example,
[0915]
number
[0916] Determining a second estimated version of an image unit based on its application to the second reconstructed motion value (e.g.,
[0917]
number
[0918] This includes determining the estimated image unit (for example,
[0919]
number
[0920] To generate and estimate image units (for example,
[0921]
number
[0922] The process further includes processing to generate reconstructed feature data, and using a compression network to process the image units and the reconstructed feature data to generate reconstructed coded data associated with the image units, wherein the coded data includes the reconstructed coded data.
[0923]
[0361] Example 108 includes the method described in any one of Examples 87 to 105, and includes an image unit (e.g., x d ) and the previous image unit (for example, x c Based on a comparison with ), one of the one or more motion values (e.g., m c→d ) and the determination of the first reconfigured previous image unit (for example,
[0924]
number
[0925] and the second reconfigured previous element (for example,
[0926]
number
[0927] Based on the comparison with the first estimated movement value (for example,
[0928]
number
[0929] The method further includes determining, processing at least a first estimated motion value to generate motion estimation feature data corresponding to an image unit, and using a compressed network to process the motion estimation feature data, motion value, and image unit to generate motion estimation encoded data associated with the image unit, wherein the encoded data includes motion estimation encoded data.
[0930]
[0362] Example 109 includes the method described in Example 108, wherein a second reconfigured previous element (for example,
[0931]
number
[0932] and the third reconfigured previous element (for example,
[0933]
number
[0934] Based on the comparison with the second estimated movement value (for example,
[0935]
number
[0936] This further includes determining that the motion estimation feature data corresponding to the image unit is further based on processing a second estimated motion value.
[0937]
[0363] Example 110 includes the method described in Example 108 or Example 109, and a first reconfigured motion value (for example,
[0938]
number
[0939] to the first reconfigured previous image unit (for example,
[0940]
number
[0941] Based on its application, estimated image units (e.g.,
[0942]
number
[0943] The determination is made that one or more first predicted motion values are first reconstructed motion values (for example,
[0944]
number
[0945] This includes making a determination and estimating an image unit (for example,
[0946]
number
[0947] The process further includes processing to generate reconstructed feature data, and using a compression network to process the image units and the reconstructed feature data to generate reconstructed coded data associated with the image units, wherein the coded data includes the reconstructed coded data.
[0948]
[0364] According to Example 111, the device includes a memory configured to store instructions and a processor configured to execute instructions in order to perform the method described in any one of Examples 87 to 110.
[0949]
[0365] According to Example 112, a non-temporary computer-readable medium, when executed by a processor, stores instructions causing the processor to perform the method described in any one of Examples 87 to 110.
[0950]
[0366] According to Example 113, the apparatus includes means for carrying out the method described in any one of Examples 87 to 110.
[0951]
[0367] According to Example 114, the non-temporary computer-readable medium includes instructions, which, when executed by one or more processors, cause one or more processors to acquire encoded data associated with one or more motion values, to acquire a conditional input to a compression network, which is a conditional input based on one or more first predicted motion values, and to use the compression network to process the encoded data and the conditional input to generate one or more second predicted motion values.
[0952]
[0368] According to Example 115, the apparatus includes means for acquiring encoded data associated with one or more motion values; means for acquiring a conditional input to a compression network, which is based on one or more first predicted motion values; and means for processing the encoded data and the conditional input using the compression network to generate one or more second predicted motion values.
[0953]
[0369] According to Example 116, the non-temporary computer-readable medium, when executed by one or more processors, includes instructions that cause one or more processors to acquire a conditional input to a compression network, which is based on one or more first predicted motion values, and to use the compression network to process the conditional input and one or more motion values to generate encoded data associated with one or more motion values.
[0954]
[0370] According to Example 117, the apparatus includes means for acquiring a conditional input to a compressed network, which is based on one or more first predicted motion values, and means for processing the conditional input and one or more motion values using the compressed network to generate encoded data associated with one or more motion values.
[0955]
[0371] Those skilled in the art will further understand that various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described herein with respect to the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been described above in terms of their functions. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art will understand that the described functions can be implemented in various ways for each specific application, and the determination of such implementations should not be construed as causing a departure from the scope of this disclosure.
[0956]
[0372] Steps of methods or algorithms described in relation to the implementations disclosed herein may be embodied directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compact disc read-only memory (CD-ROM), or any other form of non-temporary storage medium known in the art. The exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and storage medium may reside within an application-specific integrated circuit (ASIC). The ASIC may reside within a computing device or user terminal. Alternatively, the processor and storage medium may reside as separate components within the computing device or user terminal.
[0957]
[0373] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to create or use the disclosed embodiments. Various modifications of these embodiments will be readily apparent to a person skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of this disclosure. Accordingly, this disclosure is not intended to be limited to the embodiments shown herein, but should be given the broadest possible scope to coincide with the principles and novel features defined by the following claims.
Claims
1. One or more processors, Obtain the encoded data associated with one or more motion values, A conditional input to a compressed network, which obtains a conditional input based on one or more first predicted motion values, The compressed network is used to process the encoded data and the conditional input in order to generate one or more second predicted motion values. One or more processors configured in such a way, A device equipped with the following features.
2. The device according to claim 1, wherein the one or more motion values are based on the output of one or more sensors.
3. The device according to claim 2, wherein the one or more sensors include an inertial measuring unit (IMU).
4. The device according to claim 1, wherein the one or more motion values represent one or more motion vectors associated with one or more image units.
5. The device according to claim 4, wherein one of the one or more image units includes a coding unit.
6. The device according to claim 4, wherein one of the one or more image units includes a block of pixels.
7. The device according to claim 4, wherein one of the one or more image units includes a frame of pixels.
8. The device according to claim 1, wherein the one or more second predicted motion values represent future motion vectors.
9. The device according to claim 1, wherein the one or more second predicted motion values correspond to reconstructed versions of the one or more motion values.
10. The device according to claim 1, wherein the one or more processors are incorporated into at least one of a headset, a mobile communication device, an extended reality (XR) device, or a vehicle.
11. The device according to claim 1, wherein the one or more motion values represent one or more of linear velocity, linear acceleration, linear position, angular velocity, angular acceleration, or angular position.
12. The device according to claim 1, wherein the compression network includes a neural network having multiple layers.
13. The device according to claim 1, wherein the compression network includes a video decoder, and the video decoder has a plurality of decoder layers configured to decode a plurality of resolutions of the encoded data associated with the one or more motion values.
14. The device according to claim 1, wherein the one or more processors are configured to track an object associated with the one or more motion values across one or more frames of pixels.
15. The device according to claim 14, wherein the one or more second predicted motion values represent collision avoidance outputs associated with a vehicle.
16. The device according to claim 15, wherein the collision avoidance output indicates the predicted future position of the vehicle relative to the object.
17. The device according to claim 15, wherein the collision avoidance output indicates the predicted future position of the vehicle and the predicted future position of the object.
18. The device according to claim 1, further comprising a modem configured to receive a bitstream from an encoder device, wherein the bitstream includes the encoded data.
19. In a device, the process involves obtaining encoded data associated with one or more motion values, The device acquires a conditional input to a compression network, which is based on one or more first predicted motion values. To generate one or more second predicted motion values, the compressed network is used to process the encoded data and the conditional input. Methods that include...
20. Processing the encoded data and the conditional input is Using the aforementioned compression network, the conditional input is processed to generate feature data. The encoded data and the feature data are processed to generate one or more second predicted motion values, including, The method according to claim 19.
21. The method according to claim 20, wherein the feature data corresponds to multiscale feature data having different spatial resolutions.
22. The method according to claim 20, wherein the feature data includes multiscale wavelet transform data.
23. One or more processors, A conditional input to a compressed network, which obtains a conditional input based on one or more first predicted motion values, To generate encoded data associated with the one or more motion values, the compression network is used to process the conditional input and the one or more motion values. One or more processors configured in such a way, A device equipped with the following features.
24. The device according to claim 23, wherein the one or more motion values are based on the output of one or more sensors.
25. The device according to claim 24, wherein the one or more sensors include an inertial measuring unit (IMU).
26. The device according to claim 23, wherein the one or more motion values represent one or more motion vectors associated with one or more image units.
27. The device according to claim 23, further comprising a modem configured to transmit a bitstream to a decoder device, wherein the bitstream includes the encoded data.
28. In the device, a conditional input to a compressed network is obtained, which is a conditional input based on one or more first predicted motion values. To generate encoded data associated with the one or more motion values, the compression network is used to process the conditional input and the one or more motion values, Methods that include...
29. In the aforementioned device, the compression network is used to process the conditional input and generate feature data. Using the compression network, the one or more motion values and the feature data are processed to generate the encoded data. The method according to claim 28, further comprising:
30. The method according to claim 29, wherein the feature data includes multiscale feature data having different spatial resolutions.