Feature propagation based on stream using range expansion warping
Patent Information
- Application Number
- CN202580016265.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-01-29
- Publication Date
- 2026-09-25
Smart Images

Figure CN122826591A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to image processing. For example, aspects of this disclosure relate to systems and techniques for feature propagation using flow-based distortion. Background Technology
[0002] Many devices and systems allow a scene to be captured by generating images (or frames) and / or video data (including multiple frames). For example, a camera or a device that includes a camera can capture a sequence of frames of a scene (e.g., video of the scene). In some cases, the frame sequence can be processed to perform one or more functions, can be output for display, can be output for processing and / or consumption by other devices, and for other purposes.
[0003] Artificial neural networks attempt to replicate, using computer technology, the logical reasoning performed by the biological neural networks that make up the animal brain. Deep neural networks (such as convolutional neural networks) are widely used in many applications, such as object detection, object classification, object tracking, and big data analysis. For example, a convolutional neural network can extract high-level features (such as facial shape) from an input image and use these features to output, for example, the probability that the input image contains a specific object. Summary of the Invention
[0004] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Thus, the sole purpose of this summary is to present, in a simplified form, certain concepts relating to one or more aspects involving the mechanisms disclosed herein, prior to the detailed description presented below.
[0005] Systems, methods, apparatuses, and computer-readable media for performing image processing based on flow warping are disclosed. For example, the flow warping described herein can be used to perform feature propagation (e.g., for two-dimensional (2D) and / or three-dimensional (3D) feature maps) and / or other field propagation. In some aspects, these systems and techniques can be used with a variety of image processing techniques and / or various field propagation techniques. For example, the systems and techniques described herein can be used to perform Virtual Range Extended Warping (VREW) to accelerate feature propagation and / or field propagation processes. In some examples, Virtual Range Extended Warping can be used to implement one or more machine learning models, such as denoising diffusion models, diffusion generation models, probabilistic latent models, Markov variational models, etc. In some examples, Virtual Range Extended Warping can be used to perform image processing tasks such as diffusion-based dense prediction and / or propagation-based dense prediction (e.g., motion and / or optical flow estimation, depth estimation, multi-view stereo correspondence and / or disparity estimation, image and / or video inpainting, image reconstruction and / or completion, image editing, neighborhood statistical fields, etc.). In some respects, these systems and techniques can be used to implement one or more diffusion-based processes in neural processors (e.g., NPUs) and / or various other hardware accelerators to reduce processing latency and / or power consumption associated with diffusion inference.
[0006] According to at least one exemplary example, a method for image processing based on feature propagation and / or flow warping is provided. The method includes: obtaining flow information corresponding to multiple flow vectors between a first feature map and a second feature map in a plurality of feature maps; performing a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; obtaining feature information of the corresponding target block from the second feature map stored in one or more memories; and obtaining corresponding feature information of a plurality of adjacent blocks included within a virtual extension range surrounding the corresponding target block in the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0007] In another example, an apparatus is provided. The apparatus includes: at least one memory; and at least one processor coupled to the at least one memory and configured to: obtain flow information corresponding to a plurality of flow vectors between a first feature map and a second feature map in a plurality of feature maps; perform a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; obtain feature information of the corresponding target block from the second feature map stored in one or more memories; and obtain corresponding feature information of a plurality of adjacent blocks included in a virtual extension range surrounding the corresponding target block within the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0008] In another example, a non-transitory computer-readable medium is provided comprising instructions that, when executed by at least one processor, cause the at least one processor to: obtain flow information corresponding to a plurality of flow vectors between a first feature map and a second feature map in a plurality of feature maps; perform a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; obtain feature information of the corresponding target block from the second feature map stored in one or more memories; and obtain, from the second feature map stored in the one or more memories, corresponding feature information of a plurality of adjacent blocks included in a virtual extension range surrounding the corresponding target block within the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0009] In another example, an apparatus is provided. The apparatus includes: components for obtaining flow information corresponding to multiple flow vectors between a first feature map and a second feature map in a plurality of feature maps; components for performing a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; components for obtaining feature information of the corresponding target block from the second feature map stored in one or more memories; and components for obtaining corresponding feature information of a plurality of adjacent blocks included in a virtual extension range surrounding the corresponding target block within the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0010] The aspects generally include, as described substantially with reference to the accompanying drawings and description and illustrated as shown in the drawings and description, methods, apparatus, systems, computer program products, non-transitory computer-readable media, user equipment, user gear, wireless communication equipment, and / or processing systems.
[0011] Some aspects include a device having a processor configured to perform one or more operations of any of the methods outlined above. Further aspects include a processing device for use in the device, configured using processor-executable instructions to perform operations of any of the methods outlined above. Further aspects include a non-transitory processor-readable storage medium storing processor-executable instructions thereon configured to cause the device's processor to perform operations of any of the methods outlined above. Further aspects include a device having components for performing functions of any of the methods outlined above.
[0012] The features and technical advantages of the examples according to this disclosure have been summarized quite extensively above in order to better understand the detailed description below. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily utilized as the basis for modifying or designing other structures for achieving the same purpose of this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and operation) and their associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each figure in the drawings is provided for illustrative and descriptive purposes and not as a definition of limitation of the claims. The foregoing, as well as other features and aspects, will become more apparent upon reference to the following specification, claims, and appended drawings.
[0013] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim. Attached Figure Description
[0014] The accompanying drawings are provided to aid in describing various aspects of this disclosure, and are provided for illustrative purposes only and not for limiting the scope of the aspects. For a more detailed understanding of the foregoing features of this disclosure, a more specific description of the invention, briefly summarized above, can be obtained by referring to the aspects, some of which are illustrated in the drawings. However, it should be noted that the drawings illustrate only certain typical aspects of this disclosure and are therefore not to be considered as limiting its scope, as other equally valid aspects are permissible in this description. The same reference numerals in different drawings may identify the same or similar elements.
[0015] Figure 1 This is a block diagram illustrating an example architecture of an image processing system based on some examples; Figure 2 This is a block diagram illustrating an example specific implementation of a system based on some examples, which may include a central processing unit (CPU) configured to perform one or more of the functions described herein. Figure 3 It is an illustration of a first set of images representing a forward diffusion process (e.g., which is fixed) of a diffusion model according to some examples, and a second set of images representing a backward diffusion process (e.g., which is learned) of a diffusion model. Figure 4 This is a diagram illustrating, based on some examples, the use of a diffusion model to distribute diffused data from initial data to noise in the forward diffusion direction; Figure 5 This is a diagram illustrating the U-Net architecture for a diffusion model based on some examples; Figure 6 This is a diagram illustrating an example of randomized nearest neighbor feature propagation based on some examples; Figure 7 This is an illustration of an example of feature map propagation based on determining a stream query for a stack of multiple offset feature map memory objects (e.g., multiple stacked copies of offset feature maps). Figure 8 This is a diagram illustrating an example of feature map propagation based on a stream query determined by a stack of multiple offset streams applied to a single copy of the feature map. Figure 9 This is an example of a feature map propagation based on reading neighborhood blocks of features expanded around a single copy of the feature map, according to some examples of using virtual range expansion distortion. Figure 10 This is a diagram illustrating an example of a virtual range expansion distortion based on some examples, where memory and query complexity do not change with the neighborhood size used for propagating the range; Figure 11 This is a flowchart illustrating an example of a feature propagation process based on some examples; and Figure 12 This is a block diagram illustrating examples of computing systems that can be adopted by the disclosed systems and technologies, based on some examples. Detailed Implementation
[0016] Certain aspects of this disclosure are provided below for illustrative purposes. Alternative aspects may be devised without departing from the scope of this disclosure. Additionally, well-known elements of this disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of this disclosure. Some aspects described herein can be applied independently, and some of them can be combined, as will be apparent to those skilled in the art. In the following description, specific details are set forth for illustrative purposes to provide a thorough understanding of various aspects of this application. However, it will be apparent that various aspects can be practiced without these specific details. The figures and descriptions are not intended to be limiting.
[0017] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the scope of this application as set forth in the appended claims.
[0018] As noted above, machine learning systems (e.g., deep neural network systems or models) can be used to perform a variety of tasks, such as, but not limited to, detection and / or recognition (e.g., scene or object detection and / or recognition, face detection and / or recognition, etc.), depth estimation, pose estimation, image reconstruction, classification, 3D modeling, dense regression tasks, data compression and / or decompression, and image processing, etc. Furthermore, machine learning models can be general-purpose and can achieve high-quality results across a wide range of tasks.
[0019] In some cases, diffusion can be used to implement and execute various machine learning models and / or machine learning tasks. For example, a machine learning diffusion model is a type of generative model that can be used to model the data generation process based on transforming a simple initial distribution (e.g., such as Gaussian noise) into a more complex data distribution with a desired form and / or content. A diffusion model can be trained to generate new data similar to data seen during the training of the diffusion model. For example, a diffusion model trained on face images can generate new, realistic face images not seen during training. Diffusion models can be trained and / or implemented based on forward and backward diffusion processes. For example, in forward diffusion, noise is added to the input (e.g., the training data input). In backward diffusion, the diffusion model learns to reverse the forward diffusion process, thereby effectively denoising the data to recover the original distribution and / or creating new samples.
[0020] In image processing and / or image generation examples, a diffusion model can be configured to add Gaussian noise to an image at several time steps in the forward diffusion process, thereby gradually transforming the input image into pure noise. In the backward diffusion process, the diffusion model learns, based on a neural network or other machine learning architecture, to reconstruct the original input image from the noise. This neural network or other machine learning architecture is trained to predict the noise and subtract the predicted noise in a backward, stepwise manner. Image processing diffusion models can be used to generate high-quality images with detail, at least in part, based on the stability of the image processing diffusion model in the backward diffusion direction. The stability of the diffusion model in the backward direction corresponds to each step in the backward diffusion process where the image is further refined to enhance fidelity and coherence. Diffusion models are a class of generative machine learning models and can be implemented using flexible architectures adaptable to various data types and / or tasks. For example, diffusion models can be used for a variety of tasks, such as image generation, natural language processing (e.g., text generation, etc.), audio processing (e.g., generating high-fidelity sound, performing speech synthesis, etc.), etc.
[0021] However, diffusion models implement the backdiffusion process as an iterative process, and the inferences performed by diffusion models can be associated with relatively high or long latency. For example, a large number of iterations performed by a diffusion model during the backdiffusion inference process can correspond to increased latency. The relatively high latency of diffusion models can limit their use for real-time processing tasks, as many diffusion models may require hundreds or thousands of iterations to achieve high-quality results. The relatively high computational complexity associated with performing inference using diffusion models can further limit their use in on-device and / or mobile processing implementations.
[0022] Systems and technologies are needed to accelerate diffusion model inference. Systems and technologies are also needed to implement diffusion models and / or perform diffusion model inference using computationally limited computing devices such as smartphones, XR / AR devices, etc.
[0023] In some cases, stream-based propagation techniques and / or stream-based propagation machine learning models can be used as alternatives to machine learning diffusion models. As noted above, diffusion models can operate by progressively adding noise to an image until it is transformed into a random noise distribution, and then learning to reverse the noise addition process, allowing the diffusion model to effectively capture complex data distributions using a series of denoising steps. Stream-based propagation models can be used to perform the same and / or similar tasks as diffusion models, where, compared to diffusion models, stream-based propagation models can be associated with shorter inference times (e.g., lower latency) and lower output quality.
[0024] Flow-based propagation models can be configured to learn the distribution of data using a series of reversible transformations. For example, flow-based propagation can be performed to transform data between states while preserving at least a portion of the underlying distribution or structure across states. In an image processing example, flow warping can be based on learning and / or determining a flow field using a flow-based model that warps one image (e.g., a first state) to align with another image (e.g., a second state). Flow-based propagation and flow warping can be based on a series of smooth, reversible transformations that preserve the structure and properties of the original data. The reversible transformations associated with flow-based propagation can correspond to the determined flow field, which may also be referred to as a flow graph and / or a reversible mapping. For example, an optical flow graph is an example of a flow field that can be used to warp one image (e.g., the first frame in an image or video frame sequence) into another image (e.g., the second frame in an image or video frame sequence), where optical flow is represented as a velocity vector field indicating the displacement of a point from one frame to another.
[0025] Given the use of fragmented memory read / write (R / W) accesses that are difficult to parallelize, existing techniques for stream-based propagation and stream twisting can be computationally complex (e.g., computationally expensive) to implement and / or execute. There is a need for systems and techniques that can be used to perform stream-based propagation and stream twisting more efficiently. There is also a need for systems and techniques that can be used to reduce the number and / or frequency of memory R / W accesses associated with performing stream-based propagation and stream twisting.
[0026] In some contexts, the term "field propagation" can be used to refer to the process of evolving or transmitting a field (e.g., a scalar field, vector field, probability field, etc.) in spatial and / or temporal dimensions. Both diffusion models and flow-based models can be considered classes or types of field propagation models. For example, a diffusion model can perform field propagation where the field being propagated is a probability distribution or noise distribution in an image or data space (e.g., in the forward diffusion direction, the image data field propagates toward the noise field; in the backward diffusion direction, the noise field propagates back to the structured data field).
[0027] Flow-based models and flow warping techniques can be associated with the propagation of spatial information. For example, the propagated field can be associated with the spatial configuration or location of each pixel in an image (e.g., a flow field can be generated as a vector field indicating how each pixel in an image should move to align with a second image). The flow field can propagate the spatial information of an image to another configuration, thereby transforming the original field of pixel location into a new or different field of pixel location (e.g., propagation of spatial location and structure within an image).
[0028] This document describes systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively, “Systems and Techniques”) that can be used to perform accelerated field propagation with reduced inference latency and reduced memory usage. For example, reduced memory usage may correspond to a reduction in the total memory space used for field propagation and / or flow-based propagation, and / or may correspond to a reduction in the total number of memory queries associated with field propagation and / or flow-based propagation. In some aspects, the systems and techniques described herein for accelerated field propagation and / or flow-based propagation can be used to implement one or more machine learning diffusion models and / or one or more diffusion-based processing operations. In some cases, these systems and techniques can be used to implement various flow-based propagation and / or flow warping processing operations, including feature map warping through various machine learning models and / or architectures.
[0029] In some examples, these systems and techniques can be used to implement field propagation and / or feature map warping for one or more machine learning models, such as denoising diffusion models, diffusion-generating models, probabilistic latent models, Markov variational models, etc. In some examples, these systems and techniques can be used to implement field propagation and / or feature map warping associated with performing image processing tasks such as diffusion-based dense prediction and / or propagation-based dense prediction (e.g., motion and / or optical flow estimation, depth estimation, multi-view stereo correspondence and / or disparity estimation, image and / or video inpainting, image reconstruction and / or completion, image editing, neighborhood statistical fields, etc.). In some aspects, these systems and techniques can be used to implement field propagation and / or feature map warping in neural processors (e.g., NPUs) and / or various other hardware accelerators to reduce processing latency and / or power consumption associated with field propagation-based inference (e.g., diffusion model inference, flow-based propagation, or flow warping-based inference, etc.).
[0030] In an exemplary example, these systems and techniques can be used to implement virtual range expansion warps using stream-based propagation. The number of memory space usages and / or memory queries can be reduced by performing virtual range expansion warps without using multiple copies of the offset feature map and without using multiple copies of the offset stream. For example, in an exemplary example, using only one copy or instance of the target feature map (e.g., the original) and using one stream query for each block of the source feature map, these systems and techniques can perform field propagation and virtual range expansion warps between a first feature map (e.g., the source) and a second feature map (e.g., the target). A block may correspond to a single pixel or pixel location within a feature map, and / or may correspond to multiple pixels or pixel locations.
[0031] Each flow query can be based on the distortion between a specific block of the source feature map (e.g., a source block) and a specific block of the target feature map (e.g., a target block) and / or on the propagation of the flow. The target block can be configured as a relative point (e.g., a center point) of the neighborhood of blocks within the target feature map. For example, the neighborhood of a block can include a configurable number of blocks, each of which is adjacent to a relative point of the target block (e.g., the center point of the target block) and / or adjacent to one or more blocks also included in the neighborhood.
[0032] In some aspects, each streaming query (e.g., from a source feature map to a target feature map) can be configured to read a relative point (e.g., a center point) corresponding to the target chunk and to read chunks of features extending around the target within a virtual radius for relevance purposes. In an exemplary example, the same (e.g., a single) streaming query and warp operation between the source and target feature maps can be used to read out the target chunk (e.g., a neighborhood center point) and chunks of features extending around the target (e.g., configured neighborhood chunks). In some aspects, streaming queries can be configured to read out the target chunk and chunks of features extending around the target simultaneously (e.g., in parallel). In some examples, multiple streaming queries can be executed in parallel between the source and target feature maps, where each of the multiple streaming queries corresponds to a corresponding source chunk, a corresponding target chunk, and a corresponding chunk of features extending around the corresponding target chunk.
[0033] Further aspects of the system and technology will be described with reference to the accompanying drawings.
[0034] Figure 1 This is a block diagram illustrating an example architecture of an image processing system 100 according to various aspects of this disclosure. The image processing system 100 includes various components for capturing and processing images, such as an image of scene 106. The image processing system 100 can capture image frames (e.g., still images or video frames). In some cases, a lens 108 and an image sensor 118 (which may include an analog-to-digital converter (ADC)) may be associated with an optical axis. In one exemplary example, both the photosensitive area of the image sensor 118 (e.g., a photodiode) and the lens 108 may be centered on the optical axis.
[0035] In some examples, the lens 108 of the image processing system 100 faces the scene 106 and receives light from the scene 106. The lens 108 bends the incident light from the scene toward the image sensor 118. The light received by the lens 108 then passes through the aperture of the image processing system 100. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 110. In other cases, the aperture may have a fixed size.
[0036] One or more control mechanisms 110 may control exposure, focus, and / or zoom based on information from image sensor 118 and / or image processor 124. In some cases, one or more control mechanisms 110 may include multiple mechanisms and components. For example, control mechanism 110 may include one or more exposure control mechanisms 112, one or more focus control mechanisms 114, and / or one or more zoom control mechanisms 116. One or more control mechanisms 110 may also include, in addition to Figure 1 Additional control mechanisms beyond those illustrated herein. For example, in some cases, one or more control mechanisms 110 may include controls for controlling analog gain, flash, HDR, depth of field, and / or other image capture characteristics.
[0037] The focus control mechanism 114 of the control mechanism 110 can obtain focus settings. In some examples, the focus control mechanism 114 stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 114 can adjust the positioning of the lens 108 relative to the positioning of the image sensor 118. For example, based on the focus settings, the focus control mechanism 114 can adjust the focus by moving the lens 108 closer to or further away from the image sensor 118 via an actuated motor or servo system (or other lens mechanism). In some cases, additional lenses may be included in the image processing system 100. For example, the image processing system 100 may include one or more microlenses on each photodiode of the image sensor 118. These microlenses can each bend light received from the lens 108 toward the corresponding photodiode before the light reaches the photodiode.
[0038] In some examples, focus settings may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. Focus settings may be determined using control mechanism 110, image sensor 118, and / or image processor 124. Focus settings may be referred to as image capture settings and / or image processing settings. In some cases, lens 108 may be fixed relative to the image sensor and focus control mechanism 114.
[0039] Exposure control mechanism 112 of control mechanism 110 can obtain exposure settings. In some cases, exposure control mechanism 112 stores exposure settings in a memory register. Based on this exposure setting, exposure control mechanism 112 can control the aperture size (e.g., aperture size or aperture value), the duration of aperture opening (e.g., exposure time or shutter speed), the duration of light collection by the sensor (e.g., exposure time or electronic shutter speed), the sensitivity of image sensor 118 (e.g., ISO speed or film speed), the analog gain applied by image sensor 118, or any combination thereof. Exposure settings may be referred to as image capture settings and / or image processing settings.
[0040] The zoom control mechanism 116 of the control mechanism 110 can obtain zoom settings. In some examples, the zoom control mechanism 116 stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 116 can control the focal length of an assembly (lens assembly) of lens elements including lens 108 and one or more additional lenses. For example, the zoom control mechanism 116 can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 108) that first receives light from scene 106, where the light then passes through a focusing zoom system between the focusing lens (e.g., lens 108) and image sensor 118 before reaching image sensor 118. In some cases, a focusing zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference between them), with a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, zoom control mechanism 116 moves one or more lenses in the focusing zoom system, such as a negative lens and one or both positive lenses. In some cases, zoom control mechanism 116 can control zoom by capturing images from an image sensor (e.g., including image sensor 118) among a plurality of image sensors at a zoom corresponding to a zoom setting. For example, image processing system 100 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, zoom control mechanism 116 may capture images from the corresponding sensor based on the selected zoom setting.
[0041] Image sensor 118 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 118. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters, and thus light matching the color of the filter covering the photodiode can be measured. Various color filter arrays can be used, such as, for example, and not limited to, Bayer color filter arrays, four-color filter arrays (QCFA), and / or any other color filter array.
[0042] In some cases, image sensor 118 may optionally or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, UV cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). Image sensor 118 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms in control mechanism 110 may alternatively or additionally be included in image sensor 118. Image sensor 118 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0043] Image processor 124 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 128), one or more host processors (including host processor 126), and / or related to Figure 12The computing device architecture 1200 may include one or more processors of any other type discussed. The host processor 126 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 124 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes the host processor 126 and the ISP 128. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 130), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™ This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 130 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer port or interface), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 126 may use the I2C port to communicate with image sensor 118, and ISP 128 may use the MIPI port to communicate with image sensor 118.
[0044] Image processor 124 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 124 can store image frames and / or processed images in random access memory (RAM) 120, read-only memory (ROM) 122, cache, memory unit, another storage device, or some combination thereof.
[0045] Various input / output (I / O) devices 132 may be connected to the image processor 124. I / O devices 132 may include displays, keyboards, keypads, touchscreens, touchpads, touch-sensitive surfaces, printers, any other output devices, any other input devices, or any combination thereof. In some cases, text may be entered into the image processing device 104 via the physical keyboard or keypad of the I / O device 132, or via a virtual keyboard or keypad on the touchscreen of the I / O device 132. I / O devices 132 may include one or more ports, jacks, or other connectors that enable wired connections between the image processing system 100 and one or more peripheral devices, through which the image processing system 100 may receive data from and / or send data to one or more peripheral devices. I / O devices 132 may include one or more wireless transceivers that enable wireless connections between the image processing system 100 and one or more peripheral devices, through which the image processing system 100 may receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 132 discussed earlier, and they can be considered I / O devices 132 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.
[0046] In some cases, the image processing system 100 may be a single device. In other cases, the image processing system 100 may be two or more separate devices, including an image capture device 102 (e.g., a camera) and an image processing device 104 (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 102 and the image processing device 104 may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 102 and the image processing device 104 may be disconnected from each other.
[0047] like Figure 1 As shown, the vertical dashed line will Figure 1The image processing system 100 is divided into two parts, namely image capture device 102 and image processing device 104. Image capture device 102 includes a lens 108, a control mechanism 110, and an image sensor 118. Image processing device 104 includes an image processor 124 (including an ISP 128 and a host processor 126), RAM 120, ROM 122, and I / O devices 132. In some cases, certain components illustrated in image capture device 102 (e.g., such as ISP 128 and / or host processor 126) may be included in image capture device 102. In some examples, image processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof.
[0048] Image processing system 100 may be part of or implemented by a single computing device or multiple computing devices. In some examples, image processing system 100 may be part of electronic devices (or multiple electronic devices), such as camera systems (e.g., digital cameras, IP cameras, video cameras, security cameras, etc.), telephone systems (e.g., smartphones, cellular phones, conferencing systems, etc.), laptops or notebook computers, tablet computers, set-top boxes, smart TVs, display devices, game consoles, XR devices (e.g., HMDs, smart glasses, etc.), IoT (Internet of Things) devices, smart wearable devices, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices.
[0049] Although the image processing system 100 is shown as including certain components, those skilled in the art will understand that the image processing system 100 may include more than [other components]. Figure 1 The components shown are further components. Components of the image processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image processing system 100 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image processing system 100.
[0050] In some examples, Figure 12The computing device architecture 1200 shown and further described below may include an image processing system 100, an image capture device 102, an image processing device 104, or a combination thereof.
[0051] As noted above, various aspects of this disclosure may utilize machine learning models or systems.
[0052] Figure 2 An example implementation of system 200 is illustrated, which may include a central processing unit (CPU 202) (which may be a multi-core CPU) configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., a weighted neural network), task information, and other information may be stored in a memory block associated with the neural processing unit (NPU 208), a memory block associated with the CPU 202, a memory block associated with the graphics processing unit (GPU 204), a memory block associated with the digital signal processor (DSP 206), memory 216, and / or may be distributed across multiple blocks. Instructions executed at the CPU 202 may be loaded from the program memory associated with the CPU 202 or may be loaded from memory 216.
[0053] System 200 may also include additional processing blocks tailored to specific functions, such as GPU 204, DSP 206, connectivity engine 218 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and multimedia processor 212 capable of, for example, detecting and recognizing gestures. In one implementation, the NPU is implemented in CPU 202, DSP 206, and / or GPU 204. System 200 may also include one or more sensor processors 214, one or more image signal processors (ISP 210), and / or navigation engine 220, which may include a global positioning system. In some examples, sensor processor 214 may be associated with or connected to one or more sensors for providing sensor input to sensor processor 214. For example, the one or more sensors and sensor processor 214 may be provided in the same computing device, coupled to the same computing device, or otherwise associated with the same computing device.
[0054] System 200 may be implemented as a system-on-a-chip (SoC). System 200 may be based on an Advanced Reduced Instruction Set Computer (RISC) machine (ARM) instruction set. System 200 and / or its components may be configured to perform machine learning techniques according to various aspects of this disclosure discussed herein. For example, system 200 and / or its components may be configured to implement machine learning models (e.g., quantized trained machine learning models) as described herein and / or according to various aspects of this disclosure.
[0055] Machine learning (ML) can be considered a subset of artificial intelligence (AI). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks through pattern-dependent inference without explicit instructions. An example of an ML system is a neural network (also known as an artificial neural network), which can include groups of interconnected artificial neurons (e.g., neuron models). Neural networks can be used in a variety of applications and / or devices, such as image and / or video decoding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and more.
[0056] Individual nodes in a neural network mimic biological neurons by taking input data and performing simple operations on that data. The results of these simple operations on the input data are selectively passed to other neurons. Weights are associated with each vector and node in the network, and these values constrain how the input data relates to the output data. For example, the input data of each node can be multiplied by its corresponding weight value, and the products can be summed. The sum of the products can be adjusted with optional biases, and activation functions can be applied to the results to produce the node's output signal or "output activation" (sometimes called a feature map or activation map). The weights can initially be determined by an iterative stream of training data through the network (e.g., weights are established during training phases where the network learns how to identify a particular category based on the characteristics of its typical input data).
[0057] There are different types of neural networks, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Generative Adversarial Networks (GANs), Multilayer Perceptron (MLP) neural networks, Transformer Neural Networks, and Diffusion-based Neural Networks. For example, a Convolutional Neural Network (CNN) is a feedforward artificial neural network. A CNN may comprise a collection of artificial neurons, each possessing a receptive field (e.g., a localized region of the input space) and collectively tiling the input space. RNNs work on the principle of storing the layer's output and feeding that output back to the input to help predict the layer's outcome. A GAN is a generative neural network that learns patterns in the input data so that the neural network model can generate new synthetic outputs, which may reasonably come from the original dataset. A GAN may comprise two neural networks operating together: a generative neural network that generates the synthetic output and a discriminative neural network that evaluates the authenticity of the output. In an MLP neural network, data is fed into the input layer, and one or more hidden layers provide an abstraction level to the data. The output layer can then be predicted based on this abstract data.
[0058] Deep learning (DL) is an example of machine learning techniques and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. Using multiple layers in a deep neural network allows for the progressive extraction of higher-level features from a given raw data input. For example, the output of the first layer of artificial neurons becomes the input of the second layer, the output of the second layer becomes the input of the third layer, and so on. The layers located between the input and output of the entire deep neural network are often called hidden layers. Hidden layers learn (e.g., are trained) by transforming intermediate inputs from previous layers into slightly more abstract and complex representations that can be provided to subsequent layers until the final or desired representation is obtained as the final output of the deep neural network.
[0059] As noted above, neural networks are examples of machine learning systems and can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes in the input layer, processed by hidden nodes in one or more hidden layers, and output is produced by output nodes in the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which can include artificial neurons (or nodes). Feature maps can include filters, kernels, etc. Nodes can include one or more weights used to indicate the importance of nodes in one or more layers. In some cases, deep learning networks may have a series of many hidden layers, where earlier layers are used to determine simple and low-level properties of the input, and later layers build a hierarchy of more complex and abstract properties.
[0060] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases. Deep learning architectures perform particularly well when applied to problems with natural hierarchical structures. For example, the classification of motorized vehicles can benefit from first learning to recognize features such as wheels, windshields, and others. These features can then be combined in different ways at higher layers to identify cars, trucks, and airplanes.
[0061] Neural networks can be designed to have multiple connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the recognition of higher-level concepts can aid in discerning specific lower-level features of the input.
[0062] As previously noted, machine learning diffusion models can be implemented and / or configured as a class of generative machine learning models that can be used to model the data generation process based on transforming a simple initial distribution (e.g., Gaussian noise) into a more complex data distribution with a desired form and / or content. Figure 3 This is a diagram illustrating examples of forward and reverse diffusion processes that can be implemented using machine learning diffusion models.
[0063] For example, Figure 3 Two sets 300 of images are provided, which illustrate the forward diffusion process (e.g., which is fixed) and the backward diffusion process (e.g., which is learned) of the diffusion model. The backward diffusion process can also be referred to as the "backward denoising process". As previously noted, the backward diffusion process can be a generative process and can be used to perform inference using a trained machine learning diffusion model.
[0064] like Figure 3As shown in the forward diffusion process, noise 304 is gradually added to the first set 302 of the image at different time steps in a total of T time steps (e.g., forming a Markov chain), thereby generating a series of noise samples X1 to X... T .
[0065] From a training perspective, the diffusion model acquires an image and slowly adds noise to it to destroy the information within the image. In some respects, the noise (304) is Gaussian noise. Each time step can be compared with... Figure 3 Each consecutive image in the first set 302 of images shown corresponds to a specific image. Figure 3 The initial image X0 is an image of a vase with flowers. Noise 304 is added to each image (with noise samples X1 to X). T Correspondingly, this causes a gradual diffusion of pixels in each image until the final image (with sample X) is reached. T Correspondingly, until the noise distribution is basically matched. For example, by adding noise, as the time step increases, each data sample X1 to X... T Gradually losing its distinguishable features, it eventually leads to the final sample X T Equivalent to the target noise distribution, such as a unit variance, zero Gaussian distribution. .
[0066] The second set of images, 306, illustrates the reverse diffusion process, where X T This is the starting point for noisy images (e.g., images with Gaussian noise). The diffusion model can be trained to reverse the diffusion process (e.g., by training model p). θ- (x t-1 | x t This generates new data. In some respects, the diffusion model can be trained by finding the inverse Markov transformation that maximizes the likelihood of the training data. By traversing backward along the time-step chain, the diffusion model can generate new data. For example, as... Figure 3 As shown, the reverse diffusion process continues to generate X0 as an image of a vase with flowers. In other cases, the input and output data may vary based on the task for which the diffusion model was trained.
[0067] As noted above, the diffusion model is trained to denoise or restore the original image X0 in a progressive process, as shown in the second set of images 306. In some respects, the neural network of the diffusion model can be trained to denoise or restore the original image X0 in a progressive process. t-1 Restore X in the case t As shown in the following example equation:
[0068] The diffusion nucleus can be defined as follows: definition
[0069] Sampling can be defined as follows: .
[0070] In some cases, Value scheduling (also known as noise scheduling) is designed to make and .
[0071] The diffusion model runs iteratively to progressively generate the input image X0. In one example, the model may have twenty steps. However, in other examples, the number of steps can vary.
[0072] Figure 4 This is an illustration of how a diffusion model can be used to distribute diffusion data from initial data to noise in the forward diffusion direction, based on some aspects. Note that the initial data q(X0) is detailed in the initial stage of the diffusion process. An illustrative example of the data q(X0) is... Figure 3 The image shown is the initial image of the vase with flowers. As the diffusion model iterates and sampling noise is iteratively added to the data from t=0 to t=T, as... Figure 4 As shown, the data becomes noisier and may eventually result in pure noise (e.g., in q(X)). T ) place). Figure 4 The example illustrates the progress of the data and how the data spreads along with noise during the forward diffusion process.
[0073] In some respects, the distribution of diffuse data (e.g., such as...) Figure 4 (As shown) can be as follows: .
[0074] In the above equation, Indicates the distribution of diffusion data. Indicates the joint distribution. This represents the distribution of the input data, and It is a diffusion kernel. In this respect, the model can be improved by first sampling... And then sample To sample (This can be called ancestor sampling). The diffusion kernel takes input and returns a vector or other data structure as output.
[0075] The following is an overview of the training and sampling algorithms for the diffusion model. The training algorithm may include the following steps: repeat Perform gradient descent steps on the following expression Until convergence The sampling algorithm may include the following steps: for Finish return
[0076] Figure 5 This is an illustration of an example U-Net machine learning architecture 500 that can be used to implement a diffusion model, based on some examples. (For example, a vase with flowers) An initial image 502 is provided to the U-Net architecture 500, which includes a series of residual network (ResNet) blocks and self-attention layers to represent the network. (x t The U-Net architecture 500 also includes a fully connected layer 510. In some cases, the time representation 512 can be a sinusoidal localization embedding or a random Fourier feature. Noise output 508 from the forward diffusion process is also shown.
[0077] U-Net architecture 500 includes, for example Figure 5 The contraction path 504 and expansion path 506 are shown, giving it a U-shaped architecture. The contraction path 504 can be a convolutional network comprising repeated convolutional layers (which apply convolutional operations), each followed by a rectified linear unit (ReLU) and max-pooling operation. While an image (e.g., image 502) is being processed during the contraction path 504, the spatial information of image 502 is reduced as features are generated. The expansion path 506 combines features and spatial information through a series of up-convolutions and concatenation with high-resolution features from the contraction path 504. Some layers can be self-attention layers, which explicitly model complete contextual information by leveraging global interactions between semantic features at the encoder ends.
[0078] As previously noted, this paper describes systems and techniques that can be used to perform accelerated field propagation with reduced inference latency and reduced memory usage. For example, reduced memory usage may correspond to a reduction in the total memory space used for field propagation and / or flow-based propagation, and / or may correspond to a reduction in the total number of memory queries associated with field propagation and / or flow-based propagation. In some aspects, the systems and techniques described herein for accelerated field propagation and / or flow-based propagation can be used to implement one or more machine learning diffusion models and / or one or more diffusion-based processing operations. In some cases, these systems and techniques can be used to implement various flow-based propagation and / or flow warping processing operations, including feature map warping through various machine learning models and / or architectures.
[0079] In some cases, field propagation and / or flow-based propagation associated with various image processing techniques and / or operations can be performed based on nearest-neighbor matching between image tiles. An image tile may comprise a single pixel or a group of multiple pixels (e.g., neighboring pixels within a tile region, such as a square tile of j×j pixels, etc.). In some examples, flow-based propagation can be performed based on determining approximate nearest-neighbor matching between image tiles.
[0080] For example, randomized nearest neighbor feature propagation can be used to determine approximate nearest neighbor matches. In randomized nearest neighbor feature propagation, random sampling is used to find or determine one or more good block matches. Based on the natural coherence in the underlying image, such matches can be quickly propagated to the surrounding region to determine approximate nearest neighbor matches for the entire input image or graph, which includes multiple blocks.
[0081] Figure 6 This is a diagram illustrating an example of randomized nearest neighbor feature propagation 600 based on some examples. Randomized nearest neighbor feature propagation 600 can be implemented using iterative steps of (a) random initialization, (b) propagation, and (c) search, followed by subsequent iterations of the repeated process to iteratively propagate the best block match of the currently identified block to one or more neighboring blocks.
[0082] For example, block matching can be performed to determine approximate nearest neighbor matches for multiple blocks between the source frame 610 and the target frame 630. The source frame 610 and the target frame 630 can be, for example, image frames, feature maps, etc., associated with a sequence of image or video frames. The source frame 610 may include multiple overlapping blocks 612-1, 614-1, 616-1, etc.
[0083] Random initialization can be performed to generate randomly initialized nearest neighbor fields (NNFs) between blocks of the source frame 610 and blocks of the target frame 630. For example, a randomly initialized NNF may include a randomly initialized stream between source frame block 612-1 and target frame block 612-2, a randomly initialized stream between source frame block 614-1 and target frame block 614-2, a randomly initialized stream between source frame block 616-1 and target frame block 616-2, etc. A randomly initialized NNF may include a randomly initialized stream between each corresponding block in a plurality of blocks in the source frame 610 and a random block in a plurality of blocks in the target frame 630. After random initialization, the NNF maps each block in the source frame 610 to a random corresponding block in the target frame 630.
[0084] Subsequently, randomized nearest neighbor feature propagation 600 can be performed to iteratively improve the randomly initialized NNF (e.g., from...). Figure 6 The initialization step (a)). Each iteration may include a propagation step (e.g., Figure 6 The propagation step (b) and the random search step (e.g., Figure 6 Search step (c)).
[0085] For example, after random initialization of NNF (e.g., the stream between each source frame 610 block and the corresponding random target frame 630 block), each iteration can traverse all blocks in the source frame 610 and perform propagation and random search at each source frame 610 block.
[0086] For example, the propagation steps can be implemented based on coherence in natural images and the assumption that if a block has a good match, its neighbors are likely to have similar matches. During propagation, the current match of the corresponding block in the source frame 610 (e.g., in the target frame 630) is checked against the matches of the corresponding block's neighbors. For example, neighboring blocks of source frame block 612-1 may include block 614-1 (e.g., a left offset or perturbation relative to block 612-1) and block 616-1 (e.g., an up offset or perturbation relative to block 612-1), etc. Neighboring blocks of source frame block 612-1 may additionally include right and down perturbations relative to block 612-1, but in Figure 6 The example is not shown.
[0087] For example, during the propagation step, the left and upper neighboring blocks of block 612-1 can be checked to determine whether the current match of one of the adjacent blocks (e.g., within the source frame 610) is better than the current match 619-2 of block 612-1.
[0088] For example, during the propagation step for source block 612-1, source block 612-1 can be associated with the current match given by the flow to target block 612-2. Matches between adjacent source blocks 614-1 and 616-1 (e.g., the left and upper neighbor blocks of block 612-1) can be tested, and this match is compared with the current match 612-2 of block 612-1. In some aspects, the propagation and testing of adjacent block matches can be based on a 2D perturbation that shifts the current match slightly. To execute, use p = {all combinations of directional perturbations} = {(-p, -p), (-p, p), (p, -p), (p, p)}.
[0089] For example, the current match 614-2 of neighboring block 614-1 can be perturbed and tested against the current match 612-2 of block 612-1, and the current match 616-2 of neighboring block 616-1 can be perturbed and tested against the current match 612-2 of block 612-1. In some cases, the current matches 614-2 and 616-2 of neighboring blocks 614-1 and 616-2 (respectively) can each be perturbed by a small 2D perturbation. p is perturbed, and then warped before determining the similarity assessment (e.g., correlation) between the original features and the shifted features: The source feature map is represented as 610. The target feature map is represented as 630, and... Indicates based on the disturbance The corresponding perturbations in the set of p are applied to the target feature map. The set of offset feature maps generated. express Source features and The flow between target features (e.g., spatial parallax). This indicates that it can be used to apply perturbations. The spatial offset function of p, and This indicates that it can be used for configuration-based streaming or spatial parallax. A distortion function used to distort spatial offset feature maps.
[0090] If a better match is found, the current match of the source frame block is updated to the better match of the neighboring block. For example, during the propagation step for source frame block 612-1, a similarity assessment (e.g., correlation) given in the above equation can be determined between the original features of block 612-1 and a set of candidate matches, which includes one or more (or all) of the following: current match 612-2, current neighboring match 614-2, and the current neighboring match 614-2. Perturbation, current adjacent match 616-2 and / or current adjacent match 616-2 Disturbance.
[0091] For example, a similarity assessment (e.g., correlation) determined during the propagation step for source block 612-1 can determine that a rightward perturbation of adjacent match 614-2 is a better match than the current match 612-2 in terms of the features of source block 612-1. Based on identifying the better similarity between this match and source block 612-1, the current match of source block 612-1 can be updated to target block 612-3 (e.g., a rightward perturbation of adjacent match 614-2).
[0092] Following the propagation steps described above, each iteration may perform a random search around the current match of each source frame 610 block. For example, after updating the current match of source block 612-1 to target block 612-3, a random search may be performed around the current match 612-3. The random search can be used to escape local minima and further improve block matching (e.g., increase the similarity and / or relevance of matching blocks of source block 612-1).
[0093] For example, a random search can be performed based on sampling additional blocks within the target frame 630 and using an exponentially decreasing window around the current best match 612-3 identified during the propagation step for the source block 612-1. Each sample can be drawn from a uniform distribution in a square neighborhood centered on the current best match 612-3, where each sample is taken from a smaller window until the sampling window becomes smaller than the configured size.
[0094] For example, a first random sample 642-1 can be obtained from a first window 640-1 located within the target frame 630 and centered on the current best match 612-3. A second random sample 642-2 can be obtained from a smaller second window 640-2 located within the target frame 630 (and within the first window 640-1). A third random sample 642-3 can be obtained from an even smaller third window 640-3 located within the target frame 630 (and within the second window 640-2), and so on.
[0095] Each random sample drawn from the exponentially decreasing windows 640-1, 640-2, 640-3, ..., 642-1, 642-2, 642-3, ... is examined to determine if any random sample provides a better match to the source block 612-1 than the current best match 612-3. If a random sample provides a better match, the current best match of the source block 612-1 is updated to the better match from the random samples 642-1, 642-2, 642-3, ..., 642-1. This random search step can be used to prevent the randomized nearest neighbor feature propagation 600 from getting trapped in local optima.
[0096] In each iteration, a propagation step (b) and a random search step (c) may be performed on each of the multiple blocks included in the source frame 610. The propagation step (b) and the random search step (c) may be repeated multiple times to improve the quality of the nearest neighbor field (NNF) corresponding to the flow between each corresponding block of the source frame 610 and the corresponding best-matching block within the target frame 630.
[0097] Figure 7 This is an illustration of an example of feature map propagation 700 based on randomized nearest neighbor. For example, feature map propagation 700 can be compared with... Figure 6 The randomized nearest neighbor feature propagation corresponds to 600. In an exemplary example, Figure 7 The first feature map F1 710 can be the source feature map, and can be compared with... Figure 6 The source frame 610 is the same as or similar to it.
[0098] Multiple stacked feature maps 730M may include a target feature map F2 (e.g., target feature map 730) and an offset feature map F'. 2,1 、F' 2,2 、F' 2,3 、F' 2,4 、F' 2,5 、F' 2,6 、F' 2,7 and F' 2,8 The set. The target feature map F2 730 can be a non-offset target feature map and can be combined with... Figure 6 The target frame 630 is the same as or similar. Offset feature map (e.g., F' 2,1 、F' 2,2 、F' 2,3 、F' 2,4 、F' 2,5 、F' 2,6 、F' 2,7 and F' 2,8The set of offset feature maps can be offset copies of the target feature map F2 730, wherein each corresponding offset feature map in the set of offset feature maps is perturbed from the target feature map F2 730 by offset or shift direction. The set corresponds to a directional perturbation.
[0099] For example, multiple stacked feature maps 730M may include a total of M=9 feature maps and can be perturbed in eight different directions. The set corresponds to this. In some respects, the value of M can represent the number of neighboring blocks used in flow-based propagation. The M=9 feature maps included in the multiple stacked feature maps 730M may include the non-offset original target feature map 730 and may include eight copies of the target feature map 730 (e.g., duplicates or additional memory object instances, etc.), wherein each corresponding copy is... One of the eight different directional perturbations in the data is offset by a different directional perturbation (e.g., the perturbation).
[0100] For each of the multiple blocks in the source feature map F1 710 (e.g., block p1 712, block p2 714, ..., etc.), corresponding stream queries f_p1, f_p2, ..., etc., are performed (respectively) for the twist operation of the randomized nearest neighbor technique. The number of stream queries may be equal to the number of blocks in the source feature map F1 710, where each stream query corresponds to multiple memory R / W operations across a stack of M=9 offset feature maps 730M.
[0101] As used herein, the terms “flow query” and “query” are used interchangeably to refer to the process of using a flow field to determine the corresponding location of pixels, features, or blocks between a first frame and a second frame. For example, a flow query may utilize the NNF or other flow field between a source feature map 710 and a stacked offset target feature map 730M to determine the corresponding location of blocks of the source feature map 710 within the stacked offset target feature map 730M.
[0102] For example, flow query f_p1 can use NNF or other flow fields to determine the corresponding location of block 712 within the stack of the offset target feature map 730M. Flow query f_p2 can use the same NNF or other flow fields to determine the corresponding location of block 714 within the stack of the offset target feature map 730M, and so on. Figure 7 At each iteration of the feature map propagation 700, for each block of the source feature map 710, each of the M replicas included in the stack of the offset target feature map 730M is queried once.
[0103] In the example, Figure 7 The feature map propagation 700 corresponds to 2D propagation, and its feature map size is... (For example, the number of pixels / blocks in the height dimension, the number of pixels / blocks in the width dimension, and the number of channels in the channel dimension), the number of neighboring blocks in the propagation is M, and downsampling is performed by a factor of 4 in both height and width. The feature map memory requirement for warping is... .
[0104] The requirements for the flow graph memory used for twisting are: , where factor 2 represents the horizontal and vertical displacement of each block during the twist. The feature size (e.g., the channel size read) of each stream query in the twist is equal to And it represents the depth of the queried target feature map volume (e.g., stacked offset feature maps 730M) for each source frame 710 block and each corresponding stream query. For example, the feature map memory requirement for warped features is... As indicated above, and in The "depth" of the queried target feature map volume of 730M at each pixel / block location within the map is equal to Therefore, it represents the feature size of each stream query in the distortion.
[0105] Based on the features included in the source feature map 710 Each block location in the twist performs a single stream query, and the number of memory queries in the twist is equal to .
[0106] use Figure 7 The total feature memory required for a 700° feature map propagation warp is equal to the feature map memory requirement for the warp (e.g., ) + Flow graph memory requirements for warping (e.g., For example, using Figure 7 The total feature memory required for feature map propagation with a 700-degree twist is equal to .
[0107] for r =2 (for example, M =25) neighborhood radius size (e.g., propagation range), and for Using INT8 to represent, Figure 7 Example feature map propagation 700 (e.g., it uses M A stack of offset feature maps, wherein a single stream query is performed on each block of the source feature map and across... M The stacking of offset feature maps (multiple memory R / W) utilizes a total feature memory of 122.9 MB for twisting and requires 19,200 stream queries.
[0108] Figure 8This is an illustration of an example of an enhanced feature map propagation technique 800, which can be based on the relationship between the source feature map F1810 and the target feature map F2830. M This is achieved by stacking offset feature streams (e.g., streaming queries). In an illustrative example, M The stacking of offset feature streams can represent multiple operations performed on a single target feature map F2 830 for each of the multiple blocks included in the source feature map F1 810 (e.g., M (Number) stream queries.
[0109] In some respects, Figure 8 The source feature map F1 810 can be compared with Figure 7 The source feature map F1 710 and / or Figure 6 The source frame 610 is the same as or similar to it. Figure 8 The target feature map F2 830 can be compared with Figure 7 Non-offset target feature map F2730 and / or Figure 7 The target frame 730 is the same as or similar to it. Figure 8 The source feature map F1 810 may include and Figure 7 The first block p1 712 is the same as or similar to the first block p1 812, and... Figure 8 The second block p2 is the same as or similar to the second block p2 814, etc.
[0110] exist Figure 7 Feature map propagation technology 700 can be used for per-source block pair M In the case of performing feature map warping by querying the target feature map memory copy once, Figure 8 Feature map propagation technology 800 can be used to query a target feature map memory copy based on per-source block. M Next, feature map warping will be performed.
[0111] In an exemplary example, Figure 8 The enhanced feature map propagation technology 800 can be compared with Figure 7 Feature map propagation technology 700 has small associated memory space requirements. M It requires twice the memory space to perform feature map warping.
[0112] For example, regarding Figure 7 The offset feature map of 730M M The stacked, distorted feature map associated with multiple copies (e.g., memory replicas) requires memory equal to... Targeting and Figure 8 Single target feature map 830 M The memory requirements of the stacked, twisted feature maps associated with the offset feature streams are reduced.M times, and equal to .
[0113] For example, Figure 8 The feature map propagation technique 800 can reduce memory costs or memory requirements for twisting when implementing stream-based propagation, where multiple copies of the offset feature map 730M are stacked instead of being stacked. M (For example, like in) Figure 7 As in the feature map propagation technique 700, a single copy of the original feature map 830 is stored in memory, and the feature map propagation technique 800 performs... M Queries for multiples of offset streams:
[0114] For example, the first feature block p1 812 can be compared with the total from the source feature map F1 810 to the single target feature map F2 830. M =9 stacked and offset stream queries. Executed in chunks for each source feature. M The set of stacked and offset streaming queries corresponds to the streaming query from the source feature chunk to each corresponding target feature chunk included in a virtual extended neighborhood of the relative point (e.g., the center point) around the target feature map F2 830.
[0115] For example, in the first neighborhood 832 (e.g., its stacking and offset flow query f_p1+ performed with respect to the first source feature block p1 812), Correspondingly, the center point in the target feature map F2 830 is represented as block "0". Block "0" in the second neighborhood 834 represents the stacked and offset flow query f_p2+ performed with respect to the second source feature block p1 814. The corresponding center point. For M =9, the first neighborhood 832 and the second neighborhood 834 each include the total within the target feature map F2 830. M =9 feature blocks. As noted above, feature block "0" in each neighborhood 832, 834 (respectively) corresponds to a non-offset stream query for the first source feature block p1 812 and the second source feature block p2 814. Feature blocks "1", "2", ..., "8" in each neighborhood 832, 834 (respectively) correspond to eight different offset stream queries for the first source feature block p1 812 and the second source feature block p2 814, where the offset stream query is based on... The set of directional perturbations in the middle is offset or perturbed.
[0116] As previously noted, Figure 8 The enhanced feature map propagation technology 800 can be compared with Figure 7 The feature map propagation technique 700 has a memory space requirement that is M times smaller than the memory space requirement to perform feature map warping.
[0117] For example, Figure 8 The stacked stream query feature map propagation technology 800 can be used with feature map memory requirements for twisted features. Related (e.g., relative to) Figure 7 Stacked target feature map technology reduces M times).
[0118] Targeting and Figure 8 The stacked flow query associated with the twisted flow graph memory requirement can be equal to (For example, relative to) Figure 7 Stacked target feature map technology increases M (times), executed in blocks based on each source feature. M A distortion, rather than like Figure 7 In this approach, each source feature is divided into blocks and a distortion is performed. (Similar to...) Figure 8 The number of memory queries in the stacked stream query associated with the twist can be equal to That is, relative to Figure 7 Stacked target feature map technology increases M Multipliers (e.g., also performed based on block partitioning of each source feature) M (Instead of executing a stream query per source feature block).
[0119] The feature size of each query in the distortion can be equal to C That is, relative to Figure 7 Stacked target feature map technology reduces M (e.g., based on using a single target feature map F2 830 instead of stacking target feature maps) M (One offset copy).
[0120] For use Figure 8 The total characteristic memory requirement of the stacked stream query distortion can be equal to For r =2 (for example, M =25) neighborhood radius size (e.g., propagation range), and for Using INT8 to represent, Figure 8 Example feature map propagation 800 (e.g., it uses a single target feature map F2 830) M The stack of offset stream queries can be twisted using 5.9MB of total feature memory and requires 480,000 stream queries.
[0121] In an exemplary example, these systems and techniques can be used to perform flow-based propagation using virtual range extension distortions, without using offsets and / or multiple copies of the target feature map (e.g., memory replicas), and without using multiple stacked and / or offset flow queries per feature block of the source feature map. For example, Figure 9 This is an example of a feature map propagation 900 based on reading neighborhood blocks of features extended by a streaming query applied to a single copy of the feature map, using virtual range extension distortion based on some examples.
[0122] In some respects, Figure 9 Feature map propagation (e.g., Virtual Range Extended Twist (VREW)) 900 can be used to perform flow-based propagation (e.g., as in memory replicas with multiple stacks and / or offsets of the target feature map) without using memory replicas of multiple stacks and / or offsets of the target feature map. Figure 7 As in the example, it utilizes M A stacked and offset target feature map memory replica (730M). For example... Figure 9 The virtual range extension distortion 900 can be performed based on and using a single target feature map F2930, which in some respects can be combined with... Figure 8 Target feature map F2 830 Figure 7 Non-offset target feature map F2 730 and / or Figure 6 The target frame 630 is the same as or similar to it.
[0123] In some respects, Figure 9 The virtual range extension twist 900 can be used to perform flow-based propagation without using multiple stacked and / or offset flow queries per feature block of the source feature map (e.g., as in...). Figure 8 As in the example, it utilizes each feature block of the source feature map F1 810. M (Single stack and offset stream query). For example... Figure 9 The virtual range extension distortion 900 can be performed based on and using a single stream query of each feature block of the source feature map F1 910, which in some examples can be compared with... Figure 8 The source feature map F1 810 Figure 7 The source feature map F1 710 and / or Figure 6 The source frame 610 is the same as or similar to the source frame. In some cases, the source feature map F1 910 may include multiple feature blocks p1 912, p2 914, ..., etc., and these multiple feature blocks may be the same as or similar to the source frame 610. Figure 8 The (corresponding) multiple feature blocks p1 812, p2 814, ... etc. and / or Figure 7 The (corresponding) multiple feature blocks p1 712, p2 714, ... are the same or similar.
[0124] In an exemplary example, Figure 9 The virtual range extension distortion 900 can be implemented using a single (e.g., original) copy or memory replica corresponding to the target feature map F2 930. For each source feature block 912, 914, ... within the source feature map F1 910, a corresponding stream query to the target feature map F2 930 can be executed. For example, the first feature block p1 912 can be associated with a corresponding first stream query f_p1 to the target feature map F2 930, the second feature block p2 914 can be associated with a corresponding second stream query f_p2 to the target feature map F2 930, and so on.
[0125] In some respects, when a corresponding streaming query is executed for a specific feature block of the source feature map F1 910, these systems and techniques are configured to read blocks of features that extend around the target within a configured virtual radius (e.g., feature blocks of the target feature map F2 930) for relevance purposes:
[0126] In an exemplary example, the virtual radius of the configuration used to perform the virtual range expansion distortion is... r It can indicate the radius of adjacent blocks within the target feature map F2 930, which are read during the streaming query and read from the center point (e.g., the center block or other relative point) corresponding to the source feature block associated with the streaming query.
[0127] For example, each flow query may be based on the distortion between specific blocks of the source feature map 910 (e.g., source feature block 912, source feature block 914, ...) and specific blocks of the target feature map 930 and / or based on flow propagation.
[0128] For each stream query, the corresponding target block "0" can be configured as a relative point (e.g., a center point) of the neighborhood of a block within the target feature map 930. For example, the neighborhood of a target feature block may include a configurable number of target feature blocks, each of which is adjacent to the target block center point "0" and / or adjacent to one or more blocks also included in the neighborhood.
[0129] In an exemplary example, a first source feature block p1 912 may be associated with a first stream query f_p1 and a corresponding first neighborhood 932 of a target feature block centered at the target block center point "0" of the first stream query f_p1. A second source feature block p2 914 may be associated with a second stream query f_p2 and a corresponding second neighborhood 934 of a target feature block centered at the target block center point "0" of the second stream query f_p2.
[0130] Each neighborhood 932, 934 may include M Each target feature is divided into blocks, where the neighborhood size is... M Configuration-based virtual radius r For example, for r The virtual radius is configured as =1, and each neighborhood 932, 934 has a size. M =9 (For example, a target block center point "0" + a radius from the target block center point "0") r =Eight neighboring blocks "1", "2", ..., "8" within a single block.
[0131] for r The virtual radius is configured as =2, and each neighborhood 932, 934 has a size. M =25 (For example, center block "0" + distance from center point) r =1 is divided into 8 blocks "1" to "8" + distance from the center point r =1 in 16 blocks “9” to “24” (not shown). For r The virtual radius is configured as =3, and each neighborhood 932, 934 can have a size of... M =49, ... etc.
[0132] In some aspects, each stream query f_p1, f_p2, etc., can be configured to read the corresponding target block center point "0" in the corresponding neighborhood 932, 934, etc., and further read within the configured virtual radius. r Expanding around the target block "0" M-1 Each neighborhood feature is used for correlation. In an exemplary example, the same (e.g., a single) flow query f_p1, f_p2, ... and the corresponding twist operation between the source feature map 910 and the target feature map 930 can be used to read out the target block "0" (e.g., the center point of neighborhoods 932, 934) and the blocks of features extending around the target (e.g., configured neighborhood blocks "1" to "8").
[0133] In some respects, the corresponding stream queries f_p1, f_p2... can be configured to read the target block "0" and simultaneously (e.g., in parallel) read blocks of neighborhood features extending around the target (e.g., a virtual radius configured from the center block "0").r (Neighborhood blocks “1” to “8”). In some examples, multiple stream queries f_p1, f_p2... can be executed in parallel between the source feature map 910 and the target feature map 930, wherein each of the multiple stream queries corresponds to a corresponding source block, a corresponding target block, and a corresponding neighborhood block of features extending around the corresponding target block.
[0134] In an exemplary example, Figure 9 The virtual range extension distortion 900 is configured to perform a single streaming query on each feature block of the source feature map F1 910, wherein the single streaming query reads the target feature block (e.g., the "0" feature block) within the target feature map F2 930 corresponding to the queried source feature block, and additionally reads the adjacent radii located in the target feature block. r within M -1 adjacent feature blocks (e.g., feature blocks "1" to "8" in neighborhood 932 / 934 centered on the "0" feature block).
[0135] As previously noted, a flow query can refer to the process of using a flow field to perform a lookup or determination of the corresponding location of a pixel or feature block within a first feature map (e.g., within a second feature map). For example, performing a flow query f_p1 on a feature block p1 912 of a source feature map F1 910 may include determining the corresponding location of feature block p1 912 within a target feature map F2 930 (e.g., where the corresponding location of flow query f_p1 / feature block p1 912 is the center point of neighborhood 932 or the target feature block "0").
[0136] In some respects, a flow query can be performed to query the flow field to determine where a specific point (e.g., a pixel, block, etc.) in the source feature map F1 has moved to in the target feature map F2. In an exemplary example, the flow query can be configured with arguments indicating source points within the source feature map F1. For example, flow query f_p1 can be configured with source feature block p1, flow query f_p2 can be configured with source feature block p2, and so on.
[0137] The corresponding flow vector of the source point can be determined by searching in the flow field. The flow vector indicates the direction and magnitude of the movement of the source point between its position in the source feature map F1 and its position in the target feature map F2.
[0138] A warping (e.g., flow warping) can be performed to warp or transform blocks of the source feature map F1 based on the corresponding flow field and flow vector information for each block. For example, a warping operation can be performed for each pixel or block of the source feature map F1, where the source feature block is warped or transformed from a first location (e.g., within the source feature map F1) to a second location within the target feature map F2. The second location of the source feature block within the target feature map F2 can be referred to as the warped location or position of the source feature block.
[0139] Distortion can be a computationally expensive operation. In some cases, the corresponding flow vectors of points (e.g., pixels, blocks, etc.) within the source feature map F1 can point to non-integer locations in the target feature map F2, where these non-integer locations are not aligned with the pixel or block grid dimensions of the target feature map F2, but rather lie at locations between two or more pixels / blocks. Various interpolation methods (e.g., bilinear or bicubic interpolation, etc.) can be used to estimate the pixel values of the distortion at these non-integer locations.
[0140] In an exemplary example, the center point or “0” feature block in the target feature map F2 is a twist position corresponding to a specific flow query and source feature block. For example, the “0” center feature block of neighborhood 932 is a twist position determined based on the source feature block p1 and the corresponding flow vector determined for the source feature block p1 (e.g., included or represented within the flow query f_p1).
[0141] As described above, the stream query f_p1 can be used to obtain or perform readouts of features included in the "0"-centered feature block of neighborhood 932. To obtain the readouts of features from "1" to "8" in neighborhood 932, Figure 8 The method performs a separate flow query for each corresponding feature block location within a neighborhood of 932 (e.g., perturbing or shifting the flow query used to obtain the center block "0" one block to the left to query block "8", perturbing or shifting one block to the right to query block "4", perturbing or shifting one block down and one block to the right to query block "5", ... etc.). Figure 8 In this approach, executing each stream query requires performing one or more twist operations, and these twist operations can be computationally expensive, as noted above.
[0142] In an illustrative example, the systems and techniques described herein can be configured to implement virtual range-expanded warps by modifying the readouts associated with each stream query and / or warp operation. For example, a warp engine can be used to implement warp operations and stream queries between a first frame (e.g., source feature map F1) and a second frame (e.g., target feature map F2). The warp engine can be configured to implement an expanded readout operation for each warp, wherein the warp engine reads the features / values corresponding to the warp center point “0” block, and, as part of the same readout and the same stream query warp operation, is additionally configured to read the corresponding features / values (e.g., within the neighborhood size) corresponding to each adjacent block “1” through “8”. M =9 and / or virtual radius r (In the example where =1).
[0143] In some aspects, virtual range expansion can be implemented by configuring the warp engine to receive additional arguments associated with one or more stream queries and / or warp operations. For example, the additional arguments passed to the warp engine could indicate information and / or characteristics associated with neighborhoods such as 932, 934, etc., which would be read in a stream query operation targeting the corresponding neighborhood center block "0". In some examples, the additional arguments passed to the warp engine could indicate the neighborhood size. M And / or the virtual radius used for virtual extension ranges (e.g., neighborhood 932, 934, ...) of adjacent target feature blocks r One or more of these adjacent target feature blocks will be read in a stream query and twist operation targeting the combination of the central block "0".
[0144] For example, Figure 10 This is an example based on some examples regarding the radius. r Different values (e.g., with neighborhood size) M A diagram illustrating an example of virtual range extension distortion of 1000 (corresponding to different values). In an illustrative example, this can be applied to the virtual radius. r and neighborhood size M Different values of 1000 perform virtual range expansion distortion, where the memory and query complexity associated with feature propagation and / or flow distortion do not change with the neighborhood size used for propagation range. M And change.
[0145] In the first range extension configuration 1010-1, the virtual radius r =1 and neighborhood size M =9 corresponds to. In the second range extension configuration 1010-2, the virtual radius... r =2 and neighborhood size M =25 corresponds to this. In the third range extension configuration 1010-3, the virtual radius... r=3 and neighborhood size M =49 corresponds to.
[0146] In each of the three range expansion configurations 1010-1, 1010-2, and 1010-3, the source feature map and source feature blocks are shown on the left, and the corresponding target feature maps 1030-1, 1030-2, and 1030-3 are shown on the right (respectively). Figure 10 The source feature map can be compared with Figures 6 to 9 One or more of the source feature maps F1 are the same or similar. Figure 10 Target feature maps 1030-1, 1030-2, and 1030-3 can be compared with Figures 6 to 9 One or more of the target feature maps F2 are the same or similar.
[0147] For each range expansion configuration example 1010-1, 1010-2, 1010-3, the corresponding search range 1025 (respectively) represents the shaded blocks within the target feature maps 1030-1, 1030-2, 1030-3. For example, for one of them... r =1 ( M =9) The first configuration 1010-1 has a search range 1025 that is a neighborhood 1035-1, which includes the center point feature block "0" and the radius at the center. r Eight adjacent feature blocks within =1.
[0148] In response to r =2( M The search range of the second configuration 1010-2 (=25) is a neighborhood 1035-2, which includes the center point feature block "0" and the radius at the center. r =24 adjacent feature blocks within 2.
[0149] In response to r =3( M The search range of the third configuration 1010-3 (=49) is a neighborhood 1035-3, which includes the center point feature block "0" and the radius at the center. r =3 is a set of 48 adjacent feature blocks.
[0150] Neighborhood size M Based on the configured radius r ,based on To determine.
[0151] In one exemplary example, the systems and techniques described herein can be used to implement virtual range extension distortions (e.g., including...). Figure 9 and / or Figure 10The virtual range extension distortion), where the required memory size and the complexity of the number of stream queries do not change with the neighborhood size used for the propagation range. M However, variations exist. For example, three different virtual range extension twist configurations 1010-1, 1010-2, and 1010-3 can each be executed with the same memory usage and the same number of queries.
[0152] Table 1 presented below describes the relationship with Figures 6 to 10 The memory usage and streaming query usage of various techniques are illustrated with corresponding example complexity information. Table 1 corresponds to examples of 2D propagation, where the source and target feature maps have the original feature map size. The number of neighboring blocks during propagation is M ; and at a height H and width W Perform a 4x downsampling on both:
[0153] Table 1: Example complexity of memory usage and stream querying for different stream propagation techniques. This example corresponds to 2D propagation, with a feature map size of [size missing]. ,against r= The neighboring blocks in the propagation of 2 are (For example, M =25), with 4x downsampling in both height and width.
[0154] In some respects, the first column (e.g., "per block") M An offset feature map and a single stream query can be compared with Figures 6 to 7 This corresponds to the flow propagation technology. The second column (e.g., "Individual feature maps of each block and...") M An offset stream query can be used with... Figure 8 This corresponds to the streaming propagation techniques. The third column (e.g., Virtual Range Extended Twist (VREW): a single feature map and a single streaming query per block) may be related to aspects of this disclosure and / or Figure 9 and / or Figure 10 This corresponds to the flow propagation technology.
[0155] As noted above, in an exemplary example, the systems and techniques described herein can perform stream propagation and / or virtual range expansion distortions, where the complexity of memory size and the number of stream queries is independent of the size of the neighboring blocks used for the propagation range. M And change. For example, with Figure 9 and / or Figure 10 The total feature memory associated with the virtual range extension distortion example can be obtained by Given, it only depends on the feature map dimension. And it changes and is independent of the size of the neighboring blocks used for propagation.M .
[0156] In an illustrative example, the systems and techniques described herein can perform stream propagation and / or virtual range expansion distortion to allocate a larger radius size virtually (e.g., a larger radius for neighborhoods 1035-1, 1035-2, 1035-3, etc., queried by a single stream query). r and a larger number of blocks M This keeps both the feature map memory size and the number of queries unaffected (e.g., constant).
[0157] For example, where the virtual radius r =1 and neighborhood size M =9's first configuration 1010-1, where the virtual radius r =2 and neighborhood size M =25, the second configuration 1010-2, and its virtual radius. r =3 and neighborhood size M =49's third configuration 1010-3 can each utilize the resources provided by... The same total feature memory is given for twisting, and each can be executed in the twist by... The same number of memory queries are given.
[0158] In one illustrative example, the systems and techniques described herein can be used to perform accelerated field and / or flow propagation based on and / or using virtual range extension distortion. For example, virtual range extension (e.g., increasing the virtual radius) r This increases the neighborhood size for each propagation iteration. M This can be achieved by using a larger radius to accelerate field or flow propagation, resulting in faster propagation in terms of the number of iterations required to reach convergence (or other threshold levels of performance, accuracy, etc.). In some aspects, a set of coordinates at each pixel / block of the target feature map F2... M The output tensor of the virtual range expansion distortion of each neighboring block can be sequentially expanded in dimension 1 to form a shape. .
[0159] The systems and techniques described herein can be used to perform field and / or flow propagation with significantly smaller memory requirements for feature queries targeting warps. For example, in some cases, virtual range-expanded warps can be associated with memory requirements an order of magnitude smaller than existing techniques and models for feature queries used in warps. In some examples, field and / or flow propagation for warp operations can be implemented on resource-constrained devices such as smartphones, AR headsets, mobile computing devices, etc., which can utilize NPUs or NSPs with relatively small vector tightly coupled memory (VTCM). These systems and techniques can be used to implement field and / or flow propagation using virtual range-expanded warps on resource-constrained devices, including NPUs or NSPs with relatively small VTCM sizes.
[0160] In some examples, the systems and techniques described herein can be used to perform field and / or flow propagation with fewer flow queries required in warping operations (e.g., mesh sampling). In some cases, the total number of flow queries and warping operations used for mesh sampling and / or flow propagation may be the largest latency bottleneck in the NPU or NSP used to implement field or flow propagation. These systems and techniques can be used to provide accelerated field or flow propagation by reducing the number of flow queries and mitigating the corresponding latency bottleneck associated with a larger number of flow queries.
[0161] In some respects, these systems and techniques can be used to perform search ranges in target feature maps. r The virtual expansion of the search range reduces the number of propagation or diffusion iterations required to achieve or reach convergence. Reducing the number of propagation or diffusion iterations can accelerate field propagation operations (e.g., flow propagation, diffusion, etc.).
[0162] The systems and techniques described in this paper for accelerating field propagation using virtual range extension warps can be applied to flow propagation, feature map warping and / or diffusion, as well as various other techniques and models. In some cases, virtual range extension warps can be used to accelerate optical flow and / or various other iterative propagation computer vision (CV) tasks, as well as general diffusion tasks.
[0163] The systems and techniques described herein can be used to perform accelerated field propagation using virtual range-extended distortion for 2D diffusion and / or propagation tasks, as well as 3D diffusion and / or propagation tasks (e.g., spatial and temporal propagation on video, on 3D geometry such as scene flow and 3DR, etc.). In some aspects, in addition to utilizing and supporting feature queries for square and cubic 2D / 3D kernel shapes, these systems and techniques can also utilize and support corresponding 2D / 3D feature queries for arbitrary 2D and / or 3D kernel shapes.
[0164] Figure 11This is a flowchart illustrating an example of a process 1100 for image processing. For example, process 1100 could be a process for performing image processing using feature propagation. Process 1100 may be executed by a computing device or apparatus or components or systems of a computing device or apparatus (e.g., one or more chipsets, one or more processors (such as one or more CPUs, DSPs, NPUs, NSPs), microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gate or transistor logic components, discrete hardware components, etc., any combination thereof and / or other components or systems). The operation of process 1100 may be implemented on one or more processors (e.g., Figure 12 Software components that execute and run on a processor (1210 or other processor).
[0165] At box 1102, the computing device (or a component thereof) may obtain flow information corresponding to multiple flow vectors between a first feature map and a second feature map among multiple feature maps. For example, the flow information may include a dense map of coordinate disparity information indicating the correspondence between the first and second feature maps. In some cases, the flow information may be associated with the first feature map... Figure 6 The source frame 610 and the second feature map Figure 6 The target frame 630 has the same or similar streaming information associated with it and the streaming information between them.
[0166] In some cases, the first feature map can be compared with... Figure 7 Feature map F1710 Figure 8 810, Figure 9 910 Figure 10 The 1010-1, 1010-2, 1010-3, etc., are the same or similar. In some cases, the second feature map can be the same as... Figure 7 F2 and / or F2' feature maps 730M Figure 8 F2 feature map 830 Figure 9 930, Figure 10 The same as or similar to 1030-1, 1030-2, 1030-3, etc. Flow information may include one or more flow vectors, such as those related to... Figures 7 to 10 The associated flow vectors f_p1 and f_p2.
[0167] At box 1104, the computing device (or a component thereof) may perform a flow query for each corresponding block among a plurality of blocks within a first feature map to determine a corresponding target block within a second feature map, wherein the flow query is based on flow information. For example, the flow query may be based on Figure 6 Stream initialization, stream propagation, and / or search can be based on Figures 7 to 10 Stream queries, etc.
[0168] In some cases, the corresponding target block within the second feature map is configured as the current best match for feature propagation, corresponding to the corresponding block within the first feature map. For example, the corresponding target block could be... Figure 8 , Figure 9 , Figure 10 Block "0" of feature map F2. In some examples, the computing device (or a component thereof) is configured to perform feature propagation based on a similarity assessment between feature information of the corresponding block within the first feature map and corresponding feature information obtained for multiple adjacent blocks within the second feature map. For example, multiple adjacent blocks could be in Figure 8 The F2 feature map 830 contains blocks "1" to "8" in the neighborhood 832 surrounding block "0" and / or in Figure 8 Blocks “1” to “8” in the neighborhood 834 surrounding block “0” in the F2 feature map 830.
[0169] In some cases, to perform feature propagation, a computing device (or a component thereof) may determine that a particular adjacent block among a plurality of adjacent blocks has a greater similarity to a corresponding block within a first feature map than the current best match, and may configure the particular adjacent block as the current best match for feature propagation corresponding to the corresponding block within the first feature map. In some examples, to perform one iteration of feature propagation, the computing device (or a component thereof) may be configured to perform a single streaming query for each corresponding block among a plurality of blocks within the first feature map, wherein the number of streaming queries performed does not vary with the virtual expansion range and does not vary with the number of blocks included in the plurality of adjacent blocks.
[0170] In some examples, a computing device (or a component thereof) may be configured to determine optical flow information corresponding to a first feature map and a second feature map, wherein the optical flow information is based on feature propagation.
[0171] At box 1106, the computing device (or a component thereof) can obtain feature information of the corresponding target block from a second feature map stored in one or more memories.
[0172] At box 1108, a computing device (or a component thereof) may obtain from a second feature map stored in one or more memories the corresponding feature information of a plurality of adjacent blocks within a virtual extended range surrounding a corresponding target block, which is included in the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on a stream query.
[0173] In some cases, the memory requirements for the twist associated with streaming queries do not change with the virtual extension range and do not change with the number of blocks included in multiple adjacent blocks.
[0174] In some examples, multiple adjacent blocks can be included in Figure 9 The virtual extension range 932 surrounding the first target block "0" in the second feature map F2 930, and / or may be included in Figure 9 Within the virtual extension range 934 surrounding the second target block "0" in the second feature map F2. In some examples, the virtual extension range can be... Figure 10 The virtual extension ranges 1035-1, 1035-2 and / or 1035-3 are the same or similar.
[0175] In some cases, the corresponding feature information of multiple adjacent blocks is obtained from a single memory copy of a second feature map stored in one or more memories. For example, Figure 9 The second feature map 930 may include and / or be associated with a single memory copy stored in memory. In some examples, the corresponding feature information of multiple adjacent blocks is obtained without using an offset memory copy corresponding to the second feature map. For example, the corresponding feature information of multiple adjacent blocks may be obtained without using an offset memory copy corresponding to the second feature map. Figure 7 The second feature map F2 Figure 7 This was obtained using the same or similar offset memory copy, such as the 730M offset memory copy.
[0176] In some cases, the computing device (or its components) may be configured to determine the virtual extension range based on a configured radius value and the location of the corresponding target block within a second feature map. In other cases, the computing device (or its components) may be configured to determine the virtual extension range based on a configured neighborhood size that indicates the number of blocks included in a plurality of adjacent blocks.
[0177] In some examples, each of the multiple flow vectors indicates a displacement from a source block within a first feature map to a target block within a second feature map. In some cases, the computing device (or a component thereof) may be configured to perform diffusion based on iterative propagation between the first and second feature maps, using corresponding feature information of multiple adjacent blocks included within a virtual extension range.
[0178] In some examples, as previously noted, the processes described herein (e.g., Figure 11 Process 1100 and / or other processes described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more processes described herein may be performed by... Figure 1 The image processing system 100, image processing device 104, and / or image capture device 102 perform the process. In another example, one or more processes described herein may be performed by... Figure 2 The system 200 is running.
[0179] In another example, these processes (e.g., Figure 11 One or more of the processes in process 1100 and / or other processes described herein may be performed by Figure 12 The computing device architecture 1200 shown is implemented wholly or partially. For example, it has Figure 12 The computing device of the computing device architecture 1200 shown may include Figure 1 Image processing system 100, image processing device 104 and / or image capture device 102, Figure 2 The system 200 and other components, or components included in these components, and capable of implementing Figure 11 The operation of process 1100 and / or other processes described herein. In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.
[0180] Configured to execute Figure 11 The components of the device in process 1100 may be implemented in a circuit. For example, the components may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0181] Process 1100 is illustrated as a logic flowchart, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.
[0182] Additionally, process 1100 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or implemented in a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0183] In some respects, the machine learning system or neural network described in this paper (e.g., implementation) Figure 4 The neural network shown in the diffusion process Figure 5 The U-Net machine learning architecture 500, and about Figures 6 to 10 Training of one or more of the various other machine learning networks described (such as those described) can be performed using online training, offline training, and / or various combinations of online and offline training. In some cases, online may refer to processing the input data during its processing (e.g., regarding...). Figures 6 to 10The term "offline" refers to one or more feature maps in the described feature maps, for example, a time period used to perform feature propagation using flow-based distortions implemented by the systems and techniques described herein. In some examples, "offline" may refer to a period of idle time or a period during which no input data is processed. Additionally, "offline" may be based on one or more time conditions (e.g., after a certain amount of time has elapsed, such as a day, a week, a month, etc.) and / or on various other conditions, such as network and / or server availability, and various other conditions. In some aspects, offline training of a machine learning model (e.g., a neural network model) may be performed by a first device (e.g., a server device) to generate a pre-trained model, and a second device may receive the trained model from the second device. In some cases, a second device (e.g., a mobile device, XR device, vehicle or system / component of a vehicle, or other device) may perform online (or on-device) training of the pre-trained model to further adapt or tune the parameters of the model.
[0184] Figure 12 Example computing device architecture 1200 illustrates example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1200 may include, implement Figure 1 Image processing system 100, image processing device 104 and / or image capture device 102, Figure 2 The system may include or be comprised of any or all of the systems 200, etc. Additionally or alternatively, the computing device architecture 1200 may be configured to perform... Figure 11 Process 1100 and / or other processes described herein.
[0185] The components of computing device architecture 1200 are shown to communicate electrically with each other using a connection 1212, such as a bus. Example computing device architecture 1200 includes a processing unit (CPU or processor) 1202 and a computing device connection 1212 that couples various computing device components, including computing device memories 1210 (such as read-only memory (ROM) 1208 and random access memory (RAM) 1206), to the processor 1202.
[0186] The computing device architecture 1200 may include a cache of high-speed memory that is directly connected to, very close to, or integrated into the processor 1202. The computing device architecture 1200 may copy data from memory 1210 and / or storage device 1214 to cache 1204 for fast access by the processor 1202. In this way, the cache can provide performance improvements by avoiding latency for the processor 1202 while waiting for data. These and other modules may control or be configured to control the processor 1202 to perform various actions. Other computing device memory 1210 may also be used. Memory 1210 may include various different types of memory with different performance characteristics. The processor 1202 may include any general-purpose processor and hardware or software services configured to control the processor 1202 (such as services 11216, 1218, and 31220 stored in storage device 1214), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1202 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0187] To enable user interaction with the computing device architecture 1200, input device 1222 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1224 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some instances, multi-mode computing devices allow users to provide multiple types of input to communicate with computing device architecture 1200. Communication interface 1226 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0188] Storage device 1214 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 1206, read-only memory (ROM) 1208, and hybrid forms thereof. Storage device 1214 may include services 1216, 1218, and 1220 for controlling processor 1202. Other hardware or software modules are envisioned. Storage device 1214 may be connected to computing device connection 1212. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 1202, connection 1212, output device 1224, etc., to perform that function.
[0189] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.
[0190] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.
[0191] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.
[0192] Specific details have been provided in the foregoing description to offer a thorough understanding of the aspects and examples presented herein, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that various inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. Various features and aspects of the applications described above may be used individually or in combination. Furthermore, without departing from the broader scope of this specification, aspects may be used in any number of environments and applications beyond those described herein. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.
[0193] For clarity, in some instances, this technology may be presented as comprising various functional blocks, which include devices, device components, steps, or routines embodied in a method, either in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the aspects.
[0194] Furthermore, those skilled in the art will understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure.
[0195] Individual aspects may be described above as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many operations within an operation may be executed in parallel or concurrently. Furthermore, the order of operations may be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process may correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, the termination of that process may correspond to the function returning to the calling function or the main function.
[0196] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that can be used to store the instructions, the information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0197] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0198] Those skilled in the art will understand that information and signals can be represented using any of a variety of different techniques and arts. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may, in some cases, be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or light particles, or any combination thereof, depending in part on the specific application, in part on the desired design, in part on the corresponding technology, etc.
[0199] The various exemplary logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein can be implemented or executed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any form factor of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks can be stored in a computer-readable or machine-readable medium. A processor can perform the necessary tasks. Examples of form factors include: laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, self-contained devices, etc. The functionality described herein can also be embodied in peripheral devices or interlocking cards. By additional examples, such functionality can also be implemented on circuit boards of different chips or different processes executed on a single device.
[0200] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0201] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0202] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0203] Those skilled in the art will understand that, without departing from the scope of this description, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“>”) respectively. ") and greater than or equal to (" The symbol ) is used instead.
[0204] When a component is described as being “configured” to perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0205] The phrase “coupled to” or “communicatively coupled to” means that any component is physically connected directly or indirectly to another component, and / or that any component is in communication with another component directly or indirectly (e.g., connected to that other component via a wired or wireless connection and / or other suitable communication interface).
[0206] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0207] Claims using phrases such as "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," or "one or more processors being configured to," or other languages, indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks involving operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" could mean that any single processor can perform only a subset of operations X, Y, and Z.
[0208] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.
[0209] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).
[0210] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.
[0211] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0212] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0213] The exemplary aspects of this disclosure include: Aspect 1. An apparatus for performing feature propagation, the apparatus comprising: one or more memories configured to store a plurality of feature maps; and one or more processors coupled to the one or more memories, the one or more processors configured to: obtain flow information corresponding to a plurality of flow vectors between a first feature map and a second feature map in the plurality of feature maps; perform a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; obtain feature information of the corresponding target block from the second feature map stored in the one or more memories; and obtain corresponding feature information of a plurality of adjacent blocks included in a virtual extension range surrounding the corresponding target block within the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0214] Aspect 2. The apparatus according to aspect 1, wherein: the corresponding target block in the second feature map is configured as the current best match for feature propagation corresponding to the corresponding block in the first feature map; and the one or more processors are configured to perform feature propagation based on a similarity assessment between feature information of the corresponding block in the first feature map and corresponding feature information obtained for the plurality of adjacent blocks in the second feature map.
[0215] Aspect 3. The apparatus according to Aspect 2, wherein, in order to perform feature propagation, the one or more processors are configured to: determine that a particular adjacent block among the plurality of adjacent blocks has a greater similarity to the corresponding block in the first feature map than the current best match; and configure the particular adjacent block as the current best match for feature propagation corresponding to the corresponding block in the first feature map.
[0216] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein the one or more processors are configured to determine the virtual extension range based on a configured radius value and the position of the corresponding target block within the second feature map.
[0217] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the one or more processors are configured to determine the virtual extension range based on a neighborhood size indicating the number of blocks included in the plurality of adjacent blocks.
[0218] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the memory requirements for the distortion associated with the stream query do not change with the virtual extension range and do not change with the number of blocks included in the plurality of adjacent blocks.
[0219] Aspect 7. The apparatus according to any one of Aspects 1 to 6, wherein, in order to perform one iteration of feature propagation, the one or more processors are configured to: perform a single stream query for each corresponding block of the plurality of blocks within the first feature map, wherein the number of stream queries performed does not vary with the virtual extension range and does not vary with the number of blocks included in the plurality of adjacent blocks.
[0220] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the corresponding feature information of the plurality of adjacent blocks is obtained from a single memory copy of the second feature map stored by the one or more memories.
[0221] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the corresponding feature information of the plurality of adjacent blocks is obtained without using an offset memory copy corresponding to the second feature map.
[0222] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the one or more processors are configured to determine optical flow information corresponding to the first feature map and the second feature map, wherein the optical flow information is based on the feature propagation.
[0223] Aspect 11. The apparatus according to any one of aspects 1 to 10, wherein each of the plurality of flow vectors indicates a displacement from a source block within the first feature map to a target block within the second feature map.
[0224] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the one or more processors are configured to perform diffusion based on iterative propagation between the first feature map and the second feature map using the corresponding feature information included in the plurality of adjacent blocks within the virtual extended range.
[0225] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the flow information includes a dense map of coordinate parallax information indicating the correspondence between the first feature map and the second feature map.
[0226] Aspect 14. The apparatus according to any one of aspects 1 to 13, the apparatus further comprising one or more cameras configured to capture corresponding images corresponding to the first feature map and the second feature map.
[0227] Aspect 15. The apparatus according to aspect 14, wherein the one or more processors are configured to: generate one or more output images corresponding to the corresponding image, wherein the one or more output images are generated based on iterative feature propagation using the virtual extended range.
[0228] Aspect 16. The apparatus according to aspect 15, the apparatus further comprising one or more displays configured to display the one or more output images.
[0229] Aspect 17. The apparatus according to any one of Aspects 14 to 16, wherein the one or more processors are configured to perform one or more of optical flow estimation, depth estimation or motion estimation between the first feature map and the second feature map based on iterative feature propagation using the virtual extended range.
[0230] Aspect 18. A method for feature propagation, the method comprising: obtaining flow information corresponding to a plurality of flow vectors between a first feature map and a second feature map in a plurality of feature maps; performing a flow query for each corresponding block in a plurality of blocks within the first feature map to determine a corresponding target block in the second feature map, wherein the flow query is based on the flow information; obtaining feature information of the corresponding target block from the second feature map stored in one or more memories; and obtaining corresponding feature information of a plurality of adjacent blocks included in a virtual extension range surrounding the corresponding target block within the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the flow query.
[0231] Aspect 19. The method according to aspect 18, wherein: the corresponding target block in the second feature map is configured as the current best match for feature propagation corresponding to the corresponding block in the first feature map; and the method further comprises: performing feature propagation based on a similarity assessment between feature information of the corresponding block in the first feature map and corresponding feature information obtained for the plurality of adjacent blocks in the second feature map.
[0232] Aspect 20. The method according to aspect 19, wherein performing feature propagation includes: determining that a particular adjacent block among the plurality of adjacent blocks has a greater similarity to the corresponding block in the first feature map than the current best match; and configuring the particular adjacent block as the current best match for feature propagation corresponding to the corresponding block in the first feature map.
[0233] Aspect 21. The method according to any one of Aspects 18 to 20, the method further comprising: determining the virtual extension range based on a configured radius value and the position of the corresponding target block within the second feature map.
[0234] Aspect 22. The method according to any one of aspects 18 to 21, the method further comprising: determining the virtual extension range based on a neighborhood size indicating the number of blocks included in the plurality of adjacent blocks.
[0235] Aspect 23. The method according to any one of Aspects 18 to 22, wherein the memory requirements for the distortion associated with the stream query do not change with the virtual extension range and do not change with the number of blocks included in the plurality of adjacent blocks.
[0236] Aspect 24. The method according to any one of Aspects 18 to 23, wherein performing one iteration of feature propagation comprises: performing a single stream query for each corresponding block of the plurality of blocks within the first feature map, wherein the number of stream queries performed does not vary with the virtual extension range and does not vary with the number of blocks included in the plurality of adjacent blocks.
[0237] Aspect 25. The method according to any one of Aspects 18 to 24, wherein the corresponding feature information of the plurality of adjacent blocks is obtained from a single memory copy of the second feature map stored by the one or more memories.
[0238] Aspect 26. The method according to any one of Aspects 18 to 25, wherein the corresponding feature information of the plurality of adjacent blocks is obtained without using an offset memory copy corresponding to the second feature map.
[0239] Aspect 27. The method according to any one of Aspects 18 to 26, the method further comprising: determining optical flow information corresponding to the first feature map and the second feature map, wherein the optical flow information is propagated based on the features.
[0240] Aspect 28. The method according to any one of aspects 18 to 27, wherein each of the plurality of flow vectors indicates a displacement from a source block within the first feature map to a target block within the second feature map.
[0241] Aspect 29. The method according to any one of Aspects 18 to 28, the method further comprising: performing diffusion based on iterative propagation between the first feature map and the second feature map using the corresponding feature information included in the plurality of adjacent blocks within the virtual extension range.
[0242] Aspect 30. The method according to any one of Aspects 18 to 29, wherein the flow information includes a dense map of coordinate disparity information indicating the correspondence between the first feature map and the second feature map.
[0243] Aspect 31. The method according to any one of aspects 18 to 30, the method further comprising: generating one or more output images corresponding to respective images captured by one or more cameras and corresponding to the first feature map and the second feature map, wherein the one or more output images are generated based on iterative feature propagation using the virtual extended range.
[0244] Aspect 32. The method according to any one of Aspects 18 to 31, the method further comprising: performing one or more of optical flow estimation, depth estimation or motion estimation between the first feature map and the second feature map based on iterative feature propagation using the virtual extended range.
[0245] Aspect 33. A non-transitory computer-readable storage medium comprising instructions stored thereon, the instructions causing the at least one processor, when executed by at least one processor, to perform an operation according to any one of Aspects 1 to 17.
[0246] Aspect 34. A non-transitory computer-readable storage medium comprising instructions stored thereon, the instructions causing the at least one processor, when executed by at least one processor, to perform any one of aspects 18 to 32.
[0247] Aspect 35. An apparatus comprising one or more components for performing operations according to any one of aspects 1 to 17.
[0248] Aspect 36. An apparatus comprising one or more components for performing operations according to any one of aspects 18 to 32.
Claims
1. An apparatus for performing feature propagation, the apparatus comprising: One or more memories, the one or more memories being configured to store a plurality of feature maps; and One or more processors, said one or more processors coupled to said one or more memories, said one or more processors being configured to: Obtain flow information corresponding to multiple flow vectors between the first and second feature maps in the plurality of feature maps; For each corresponding block among multiple blocks in the first feature map, a flow query is performed to determine the corresponding target block in the second feature map, wherein the flow query is based on the flow information; The feature information of the corresponding target block is obtained from the second feature map stored in the one or more memories; as well as The corresponding feature information of a plurality of adjacent blocks, which are included in a virtual extension range around the corresponding target block within the second feature map stored in the one or more memories, is obtained from the second feature map, wherein the corresponding feature information of the plurality of adjacent blocks and the feature information of the corresponding target block are obtained based on the stream query.
2. The apparatus according to claim 1, wherein: The corresponding target block within the second feature map is configured as the current best match for feature propagation, corresponding to the corresponding block within the first feature map; and The one or more processors are configured to perform feature propagation based on a similarity assessment between feature information of the corresponding blocks in the first feature map and corresponding feature information obtained for the plurality of adjacent blocks in the second feature map.
3. The apparatus of claim 2, wherein, in order to perform feature propagation, the one or more processors are configured to: Determine that a specific adjacent block among the plurality of adjacent blocks has a similarity greater than that of the corresponding block in the first feature map than the current best match; and The specific adjacent block is configured as the current best match for feature propagation, corresponding to the corresponding block within the first feature map.
4. The apparatus of claim 1, wherein the one or more processors are configured to determine the virtual extension range based on a configured radius value and the position of the corresponding target block within the second feature map.
5. The apparatus of claim 1, wherein the one or more processors are configured to determine the virtual extension range based on a neighborhood size indicating the number of blocks included in the plurality of adjacent blocks.
6. The apparatus of claim 1, wherein the memory requirements for the distortion associated with the stream query do not change with the virtual extension range and do not change with the number of blocks included in the plurality of adjacent blocks.
7. The apparatus of claim 1, wherein, in order to perform one iteration of feature propagation, the one or more processors are configured to: Perform a single streaming query for each of the multiple blocks within the first feature map. The number of stream queries executed does not change with the virtual extended range, nor with the number of blocks included in the plurality of adjacent blocks.
8. The apparatus of claim 1, wherein the corresponding feature information of the plurality of adjacent blocks is obtained from a single memory copy of the second feature map stored in the one or more memories.
9. The apparatus of claim 1, wherein the corresponding feature information of the plurality of adjacent blocks is obtained without using an offset memory copy corresponding to the second feature map.
10. The apparatus of claim 1, wherein the one or more processors are configured to determine optical flow information corresponding to the first feature map and the second feature map, wherein the optical flow information is propagated based on the features.
11. The apparatus of claim 1, wherein each of the plurality of flow vectors indicates a displacement from a source block within the first feature map to a target block within the second feature map.
12. The apparatus of claim 1, wherein the one or more processors are configured to perform diffusion based on iterative propagation between the first feature map and the second feature map, using the corresponding feature information included in the plurality of adjacent blocks within the virtual extension range.
13. The apparatus of claim 1, wherein the flow information includes a dense map of coordinate parallax information indicating the correspondence between the first feature map and the second feature map.
14. The apparatus of claim 1, further comprising one or more cameras configured to capture corresponding images corresponding to the first feature map and the second feature map.
15. The apparatus of claim 14, wherein the one or more processors are configured to: Generate one or more output images corresponding to the corresponding image, wherein the one or more output images are generated based on iterative feature propagation using the virtual extended range.
16. The apparatus of claim 15, further comprising one or more displays configured to display the one or more output images.
17. The apparatus of claim 14, wherein the one or more processors are configured to perform one or more of optical flow estimation, depth estimation, or motion estimation between the first feature map and the second feature map based on iterative feature propagation using the virtual extended range.
18. A method for feature propagation, the method comprising: Obtain flow information corresponding to multiple flow vectors between the first and second feature maps in multiple feature maps; For each corresponding block among multiple blocks in the first feature map, a flow query is performed to determine the corresponding target block in the second feature map, wherein the flow query is based on the flow information; The feature information of the corresponding target block is obtained from the second feature map; as well as Obtain corresponding feature information of multiple adjacent blocks within a virtual extended range surrounding the corresponding target block from the second feature map, wherein the corresponding feature information of the multiple adjacent blocks and the feature information of the corresponding target block are obtained based on the stream query.
19. The method of claim 18, wherein: The corresponding target block within the second feature map is configured as the current best match for feature propagation, corresponding to the corresponding block within the first feature map; and The method further includes performing feature propagation based on a similarity assessment between the feature information of the corresponding block in the first feature map and the corresponding feature information obtained for the plurality of adjacent blocks in the second feature map.
20. The method of claim 19, wherein performing feature propagation comprises: The similarity between a specific adjacent block among the plurality of adjacent blocks and the corresponding block in the first feature map is determined to be greater than the current best match; as well as The specific adjacent block is configured as the current best match for feature propagation, corresponding to the corresponding block within the first feature map.