Reverse parallax error correction

By generating a confidence map using reverse optical flow error correction technology, the problem of error accumulation in optical flow estimation is solved, thereby improving the accuracy of motion estimation and the quality of optical flow information.

CN122497979APending Publication Date: 2026-07-31QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-12-20
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing optical flow estimation methods are prone to introducing errors when dealing with locally inconsistent environments, leading to error accumulation and superposition, which affects the accuracy of motion estimation.

Method used

By using reverse optical flow error correction technology, a confidence map is generated to identify and clean up defective optical flow information. Dynamic thresholding and machine learning techniques are used to adjust the validity threshold to prevent errors from being superimposed in subsequent images.

Benefits of technology

It effectively reduces the accumulation of errors in optical flow estimation and improves the accuracy of motion estimation and the quality of optical flow information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497979A_ABST
    Figure CN122497979A_ABST
Patent Text Reader

Abstract

Systems and techniques for capturing images and performing inverse optical flow error correction (e.g., using image capture) are disclosed. According to some aspects, a computing system or device can obtain first disparity information associated with a current image. The first disparity information estimates a first movement of a first feature to a first destination location in the current image. The computing system or device can remap the current image based on the first disparity information to obtain an estimated previous image; determine a confidence map associated with a confidence level of the first disparity information based on the difference associated with the estimated previous image; and apply the confidence map to the first disparity information to generate updated first disparity information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image processing. For example, according to some aspects, systems and techniques for correcting errors in parallax information (e.g., optical flow information, depth information, etc.) using reverse optical flow error correction are described. Background Technology

[0002] Multimedia systems are widely deployed to provide various types of multimedia communication content, such as voice, video, packet data, messaging, and broadcasting. These multimedia systems are capable of processing, storing, generating, manipulating, and reproducing multimedia information. Examples of multimedia systems include mobile devices, gaming devices, entertainment systems, information systems, virtual reality systems, models, and simulation systems. These systems can employ a combination of hardware and software technologies to support the processing, storage, generation, manipulation, and reproduction of multimedia information, such as client devices, capture devices, storage devices, communication networks, computer systems, and display devices. Summary of the Invention

[0003] Various systems and techniques can be used to correct errors in disparity information using reverse optical flow error correction. According to at least one example, a method includes: obtaining first disparity information associated with a current image, the first disparity information estimating a first movement of a first feature to a first destination location in the current image; remapping the current image based on the first disparity information to obtain an estimated previous image; determining a confidence map associated with a confidence level of the first disparity information based on differences associated with the estimated previous image; and applying the confidence map to the first disparity information to generate updated first disparity information.

[0004] In another example, an apparatus for processing one or more images is provided, the apparatus comprising: one or more memories configured to store the one or more images; and one or more processors (e.g., implemented in a circuit) coupled to the one or more memories and configured to: obtain first disparity information associated with a current image, the first disparity information estimating a first movement of a first feature to a first destination location in the current image; remap the current image based on the first disparity information to obtain an estimated previous image; determine a confidence map associated with a confidence level of the first disparity information based on a difference associated with the estimated previous image; and apply the confidence map to the first disparity information to generate updated first disparity information.

[0005] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: obtain first disparity information associated with a current image, the first disparity information estimating a first movement of a first feature to a first destination location in the current image; remap the current image based on the first disparity information to obtain an estimated previous image; determine a confidence map associated with a confidence level of the first disparity information based on the difference associated with the estimated previous image; and apply the confidence map to the first disparity information to generate updated first disparity information.

[0006] In another example, an apparatus for processing one or more images is provided. The apparatus includes: components for: obtaining first disparity information associated with a current image, the first disparity information estimating a first movement of a first feature to a first destination location in the current image; components for: remapping the current image based on the first disparity information to obtain an estimated previous image; components for: determining a confidence map associated with a confidence level of the first disparity information based on differences associated with the estimated previous image; and components for: applying the confidence map to the first disparity information to generate updated first disparity information.

[0007] In some aspects, one or more of the devices described herein are, are part of, and / or include the following devices: wireless communication devices, mobile devices (e.g., mobile phones and / or mobile cell phones and / or so-called "smartphones" or other mobile devices), extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), such as head-mounted display (HMD) devices, vehicles or computing devices or components of vehicles, wearable devices, cameras, personal computers, laptop computers, server computers, other devices, or combinations thereof. In some aspects, the one or more devices include one or more cameras for capturing one or more images. In some aspects, the one or more devices include a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the one or more devices may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).

[0008] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0009] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0010] The exemplary aspects of this application are described in detail below with reference to the following figures: Figure 1A , Figure 1B and Figure 1C This is an illustration of an example configuration of an image sensor for an image capture device according to various aspects of this disclosure.

[0011] Figure 2 This is a block diagram illustrating the architecture of an image capture and processing apparatus according to various aspects of this disclosure.

[0012] Figure 3 This is a block diagram illustrating an example of an image capture system according to various aspects of this disclosure.

[0013] Figure 4 This is a conceptual diagram of an optical flow correction system that provides reverse optical flow error correction based on some aspects of this disclosure.

[0014] Figure 5 This is a block diagram of an optical flow correction system configured to perform reverse optical flow error correction according to some aspects of this disclosure.

[0015] Figure 6A This is a conceptual illustration of optical flow and possible errors between two images based on some aspects of this disclosure.

[0016] Figure 6B These are previous images estimated based on some aspects of this disclosure and conceptual examples of the identification of errors.

[0017] Figure 6C This is a conceptual example of a confidence graph that a cleaning engine can use to remove errors introduced into the optical flow, based on some aspects of this disclosure.

[0018] Figure 6D This is a conceptual example of optical flow after some aspects of this disclosure have been cleaned up.

[0019] Figures 7A to 7C Examples of optical flow without a reference truth value, optical flow without reverse optical flow error correction, and optical flow with reverse optical flow error correction are provided according to some aspects of this disclosure.

[0020] Figure 8 This is a flowchart illustrating an example method for capturing an image under low-light conditions according to various aspects of this disclosure.

[0021] Figure 9These are exemplary examples of deep learning neural networks that can be used to implement alignment prediction based on machine learning, according to various aspects of this disclosure.

[0022] Figure 10 These are exemplary examples of convolutional neural networks (CNNs) according to various aspects of this disclosure.

[0023] Figure 11 This is an illustration of an example system that uses optical flow to encode video frames.

[0024] Figure 12 This is a diagram illustrating an example of a system used to implement some of the aspects described in this article. Detailed Implementation

[0025] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0026] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0027] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0028] A camera is a device that uses an image sensor to receive light and capture images, such as still images or video frames. The terms "image," "image frame," and "frame" are used interchangeably herein. A camera can be configured using various image capture and image processing settings. Different settings produce images with different appearances. Camera settings, such as ISO, exposure time, aperture size, aperture value, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to the image sensor used to capture one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changes to contrast, brightness, saturation, sharpness, level, curves, or color. For example, settings or parameters can be applied to a processor (e.g., an image signal processor (ISP)) used to process one or more image frames captured by the image sensor.

[0029] Images can be used in a variety of disparity estimation applications to determine disparity information, such as optical flow estimation for determining disparity information and stereo depth estimation for determining depth information. For example, one use of images from a camera is to detect motion within an image or to use optical flow to detect motion of a device including the camera. Optical flow is a representation of motion patterns between images in a sequence of images (e.g., 310 consecutive images). For example, optical flow allows algorithms to track pixel movement by comparing pixel intensities between images. Optical flow can be useful for understanding how objects or points in an image move over time. For example, in the context of two sequential images, optical flow can estimate the displacement of pixels between the images, providing valuable information about apparent motion in the scene. For example, a camera used to capture images may be fixed in place on a device (e.g., a mobile device or vehicle), and the optical flow between images can provide information related to movement of the device or movement within the environment (such as the device's movement within the environment).

[0030] In some cases, based on the optical flow determination between the first and second images, an optical flow vector can be determined for one or more pixels in the second image. For example, a dense optical flow map may include the optical flow vector for each pixel in the image, where all pixels in the image are represented by optical flow vectors. The optical flow vector of a pixel may indicate the direction and magnitude of pixel movement from a first time when the first image is captured to a second time when the second image is captured.

[0031] Images can also be used to determine depth information. For example, a first camera can be used to capture a first image of a scene, and a second camera can be used to capture a second image of the scene. Stereo depth estimation can be performed to determine depth information between the first and second images. Depth information can represent the distance between the first and / or second camera and one or more objects depicted in the image. Depth is inversely proportional to parallax, and parallax is proportional to the baseline distance between the first and second cameras. The greater the parallax, the closer the object is to the camera's baseline. The smaller the parallax, the farther the object is from the baseline.

[0032] Parallax estimation applications (e.g., optical flow, stereo depth estimation, etc.) can be useful for many tasks, such as surveillance systems, robotics, autonomous or semi-autonomous vehicles (e.g., for autonomous or semi-autonomous navigation), video analytics, etc. For example, optical flow can be used to track trajectories and / or detect the motion of objects, identify events in the environment (e.g., to understand and respond to events in the environment), and so on. In another example, optical flow can also be used to improve video stabilization, image interpolation, image correction, extended reality (XR) applications, and so on.

[0033] Optical flow faces numerous challenges in environments with local inconsistencies, such as those caused by varying lighting, occlusion, and complex motion patterns. Local inconsistencies arise from local variations in pixel intensity, such as those caused by texture, shadows, or occlusion. For example, an object occluded in one image but visible in a subsequent image can introduce errors into the optical flow between images. In some cases, optical flow algorithms may assume that motion is consistent between consecutive images, and any abrupt changes or discontinuities in motion (e.g., interruptions in temporal coherence) can introduce errors propagating in future images, thereby introducing errors into the optical flow.

[0034] Such challenges can make it difficult to accurately estimate pixel motion between sequential images. Errors introduced into the optical flow between images can spread and accumulate when determining optical flow in subsequent images, leading to progressively larger inaccuracies in motion estimation. This accumulation of errors can be referred to as error accumulation or error propagation. Error accumulation / propagation can result in poor quality of optical flow performance. For example, if an optical flow algorithm makes an incorrect optical flow estimate at a point in the first image, this incorrect estimate can be superimposed on the previously estimated flow in subsequent images.

[0035] Errors introduced into an image can accumulate over time, and because optical flow algorithms use estimated motion vectors from previous images to initialize or guide estimations in the current image, any inaccuracies in the initial image will affect subsequent images. The cumulative effect can lead to significant deviations from the true motion trajectory. Optical flow algorithms may encounter challenges in handling occlusion or anomalies. When occlusion is not properly accounted for, the algorithm may assign incorrect velocities to pixels, leading to error aggregation in subsequent images. Furthermore, optical flow algorithms must also address dynamic changes within the environment, such as sudden object movement, scene transitions, etc. Regardless of the type of error, any errors in the optical flow from one image can propagate into subsequent images and exacerbate inaccuracies in motion estimation.

[0036] In some aspects, systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for improving disparity information for disparity estimation applications are described, including optical flow information for optical flow estimation and depth information for stereo depth estimation. For example, the systems and techniques can reverse-map a current image to generate an estimated previous image, and can compare the estimated previous image with the previous image to identify defective optical flow in the optical flow information. In some aspects, the systems and techniques generate a confidence map identifying defective optical flow based on the comparison. This confidence map can be applied to the optical flow information to remove (or "clean") the defective optical flow, thereby preventing errors from being introduced into the optical flow information.

[0037] In some aspects, different systems and techniques can use different techniques, such as dynamic thresholds (referred to herein as "validity thresholds"), to generate confidence maps. Validity thresholds can be used to identify whether optical flow is considered valid or invalid within the optical flow information and can be dynamic based on different criteria. In an illustrative example, the progress of optical flow can increase the validity threshold based on the number of iterations of the optical flow (e.g., relative to the time when the optical flow was initiated). The sparsity of features in the captured image can also be used to increase or decrease the validity threshold. In some aspects, the magnitude of the flow can also adjust the validity threshold, and the attention associated with the features can also adjust the validity threshold. In some aspects, machine learning (ML) techniques, such as using one or more neural networks, can be used to implement different types of feature detection techniques.

[0038] Various systems and techniques can apply confidence maps to optical flow information and clean up the optical flow to prevent the introduction of errors that are superimposed in subsequent optical flow determination or prediction. In some cases, systems and techniques can use optical flow information to perform timely forward remapping of images and identify objects occluded between two images.

[0039] While the examples described herein are presented in conjunction with optical flow estimation, the systems and techniques can be applied to any disparity estimation application, including stereo depth estimation. For example, the techniques used for optical flow and stereo depth are similar, except that optical flow is used for two-dimensional (2D) disparity and (corrected) stereo depth is used for one-dimensional (1D) disparity; therefore, inversion and self-cleaning operate for optical flow in 2D (e.g., in the horizontal (x) and vertical (y) directions) and for stereo depth in 1D (x direction only).

[0040] Various aspects and examples of the system and technology will be described below with respect to the accompanying figures.

[0041] Generally, an image sensor comprises one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by the image sensor. In some cases, different photodiodes can be covered by different color filters in a color filter array, and thus the light that matches the color of the color filter covering the photodiode can be measured.

[0042] Various color filter arrays can be used, including Bayer color filter arrays, four-color color filter arrays (also known as four-color Bayer filters or QCFAs) and / or other color filter arrays. Figure 1A An example of a Bayer color filter array 100 is shown. As illustrated, the Bayer color filter array 100 includes a repeating pattern of red, blue, and green color filters. Figure 1B As shown, the QCFA 110 includes a 2×2 (or "four-color") patterned color filter, comprising a 2×2 patterned rI(R) color filter, a pair of 2×2 patterned green (G) color filters, and a 2×2 patterned blue (B) color filter. For the entire array of photodiodes of a given image sensor, repeat... Figure 1B The diagram shows the pattern of a QCFA 110. Regardless of whether a QCFA 110 or a Bayer color filter array 100 is used, each pixel of the image is generated based on the following data: red light data from at least one photodiode covered by the red filter in the color filter array, blue light data from at least one photodiode covered by the blue filter in the color filter array, and green light data from at least one photodiode covered by the green filter in the color filter array. Other types of color filter arrays may use yellow, magenta, and / or cyan (also known as "emerald green") color filters as alternatives to or complements to red, blue, and / or green color filters. Different photodiodes throughout the pixel array can have different spectral sensitivity profiles, thus responding to light of different wavelengths. Monochrome image sensors may also lack color filters and therefore lack color depth.

[0043] In some cases, subgroups of multiple adjacent photodiodes (e.g., when using...) Figure 1B As shown in the QCFA 110, a 2×2 patch of photodiodes can measure light of the same color in approximately the same area of ​​a scene. For example, when the photodiodes included in each subgroup of photodiodes are physically close together, the light incident on each photodiode in the subgroup can originate from approximately the same location in the scene (e.g., a part of a leaf on a tree, a small part of the sky, etc.).

[0044] In some examples, the brightness range of light from a scene can significantly exceed the brightness levels that an image sensor can capture. For instance, a digital single-lens reflex (DSLR) camera might be able to capture light from a scene at a ratio of 1:30,000, while the brightness levels of an HDR scene can exceed a ratio of 1:1,000,000.

[0045] In some cases, HDR sensors can be used to enhance the contrast ratio of images captured by image capture devices. In some examples, HDR sensors can be used to obtain multiple exposures within an image, where such multiple exposures can include short exposure times (e.g., 5 ms) and long exposure times (e.g., 15 ms or longer). As used herein, long exposure time generally refers to any exposure time longer than short exposure time.

[0046] In some specific implementations, the HDR sensor may be able to configure individual photodiodes within a subgroup of photodiodes (e.g., from...). Figure 1B The QCFA 110 shown contains four individual R photodiodes, four individual B photodiodes, and four individual G photodiodes in each of the two 2×2 G patches with different exposure settings. A collection of photodiodes with matched exposure settings is also referred to herein as a photodiode exposure group. Figure 1C An example is shown of a portion of an image sensor array with a QCFA filter configured with four different photodiode exposure groups 1 to 4. For example... Figure 1C The example photodiode exposure group array 120 shown in the image may include photodiodes from each of the different photodiode exposure groups from a specific image sensor. Although in Figure 1C Four groups are shown in the specific groupings, but those skilled in the art will recognize that different numbers of photodiode exposure groups, different arrangements of photodiode exposure groups within subgroups, and any combination thereof may be used without departing from the scope of this disclosure.

[0047] As relative to Figure 1CAs noted, in some HDR image sensor implementations, the exposure settings corresponding to different photodiode exposure groups may include different exposure times (also known as exposure durations), such as short exposure, medium exposure, and long exposure. In some cases, the light captured by the photodiodes of each photodiode exposure group can form different images of the scene associated with different exposure settings. For example, the light captured by the photodiodes of photodiode exposure group 1 can form a first image, the light captured by the photodiodes of photodiode exposure group 2 can form a second image, the light captured by the photodiodes of photodiode exposure group 3 can form a third image, and the light captured by the photodiodes of photodiode exposure group 4 can form a fourth image. Based on the differences in the exposure settings corresponding to each group, the brightness of objects in the scene captured by the image sensor may differ in each image. For example, a well-lit object captured by a photodiode with a long exposure setting may appear saturated (e.g., completely white). In some cases, the image processor may select between pixels of images corresponding to different exposure settings to form a combined image.

[0048] In an illustrative example, the first image corresponds to a short exposure time (also known as a short exposure image), the second image corresponds to a medium exposure time (also known as a medium exposure image), and the third and fourth images correspond to long exposure times (also known as long exposure images). In such examples, pixels corresponding to the low-light portions of a scene (e.g., portions of a scene in shadow) of the combined image can be selected from the long exposure image (e.g., the third or fourth image). Similarly, pixels corresponding to the high-light portions of a scene (e.g., portions of a scene in direct sunlight) of the combined image can be selected from the short exposure image (e.g., the first image).

[0049] In some cases, image sensors can also utilize photodiode exposure arrays to capture moving objects without blurring. The length of the exposure time for the photodiode array can correspond to the distance an object in the scene moves during the exposure time. If light from a moving object is captured by photodiodes corresponding to multiple image pixels during the exposure time, the moving object may appear blurred across multiple image pixels (also known as motion blur). In some implementations, motion blur can be reduced by configuring one or more photodiode arrays with short exposure times. In some implementations, an image capture device (e.g., a camera) can determine the amount of local motion (e.g., motion gradient) within a scene by comparing the position of an object between two consecutively captured images. For example, motion can be detected in a preview image captured by the image capture device to provide a preview function to the user on a display. In some cases, machine learning models can be trained to detect local motion between consecutive images.

[0050] The various aspects of the technology described in this article will be discussed below with respect to the accompanying figures. Figure 2 This is a block diagram illustrating the architecture of an image capture and processing system 200. The image capture and processing system 200 includes various components for capturing and processing images of a scene (e.g., an image of scene 210). The image capture and processing system 200 can capture individual images (or photographs) and / or capture video comprising multiple images (or video frames) in a specific sequence. In some cases, a lens 215 and an image sensor 230 may be associated with an optical axis. In one exemplary example, both the photosensitive area of ​​the image sensor 230 (e.g., a photodiode) and the lens 215 may be centered on the optical axis. The lens 215 of the image capture and processing system 200 faces scene 210 and receives light from scene 210. The lens 215 bends the incident light from the scene toward the image sensor 230. The light received by the lens 215 passes through an aperture. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 220 and received by the image sensor 230. In some cases, the aperture may have a fixed size.

[0051] One or more control mechanisms 220 may control exposure, focus, and / or zoom based on information from image sensor 230 and / or information from image processor 250. One or more control mechanisms 220 may include multiple mechanisms and components; for example, control mechanism 220 may include one or more exposure control mechanisms 225A, one or more focus control mechanisms 225B, and / or one or more zoom control mechanisms 225C. One or more control mechanisms 220 may also include additional control mechanisms besides those illustrated, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture properties.

[0052] The focus control mechanism 225B of the control mechanism 220 can obtain the focus setting. In some examples, the focus control mechanism 225B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 225B can adjust the positioning of the lens 215 relative to the positioning of the image sensor 230. For example, based on the focus setting, the focus control mechanism 225B can move the lens 215 closer to or further away from the image sensor 230 by actuating a motor or servo system (or other lens mechanism), thereby adjusting the focus. In some cases, additional lenses (such as one or more microlenses on each photodiode of the image sensor 230) may be included in the image capture and processing system 200, each of which bends light received from the lens 215 toward the corresponding photodiode before it reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting may be determined using the control mechanism 220, the image sensor 230, and / or the image processor 250. The focus settings may be referred to as image capture settings and / or image processing settings. In some cases, lens 215 may be fixed relative to the image sensor, and the focus control mechanism 225B may be omitted without departing from the scope of this disclosure.

[0053] The exposure control mechanism 225A of the control mechanism 220 can obtain the exposure settings. In some cases, the exposure control mechanism 225A stores the exposure settings in a memory register. Based on the exposure settings, the exposure control mechanism 225A can control the aperture size (e.g., aperture size or aperture value), the duration of aperture opening (e.g., exposure time or shutter speed), the duration of light collection by the sensor (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 230 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 230, or any combination thereof. The exposure settings may be referred to as image capture settings and / or image processing settings.

[0054] The zoom control mechanism 225C of the control mechanism 220 can obtain zoom settings. In some examples, the zoom control mechanism 225C stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 225C can control the focal length of an assembly (lens assembly) of lens elements including lens 215 and one or more additional lenses. For example, the zoom control mechanism 225C can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses in the lens relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 215) that first receives light from scene 210, where the light then passes through a focusless zoom system between the focusing lens (e.g., lens 215) and image sensor 230 before reaching image sensor 230. In some cases, a focusless zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference between them), with a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, zoom control mechanism 225C moves one or more lenses in the focusless zoom system, such as a negative lens and one or both positive lenses. In some cases, zoom control mechanism 225C can control zoom by capturing images from an image sensor (e.g., including image sensor 230) among a plurality of image sensors at a zoom setting corresponding to the zoom setting. For example, image capture and processing system 200 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a higher zoom. In some cases, zoom control mechanism 225C may capture images from a corresponding sensor based on the selected zoom setting.

[0055] Image sensor 230 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 230. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters, and thus light matching the color of the filter covering the photodiode can be measured. Various color filter arrays can be used, including Bayer color filter arrays (such as...). Figure 1A As shown), QCFA (see...) Figure 1B ) and / or any other color filter array.

[0056] return Figure 1A and Figure 1BOther types of color filters can use yellow, magenta, and / or cyan (also known as "emerald green") color filters as alternatives to or supplements to red, blue, and / or green color filters. In some cases, some photodiodes can be configured to measure infrared (IR) light. In some specific implementations, the photodiode measuring IR light may not be covered by any filter, thus allowing the IR photodiode to measure both visible light (e.g., color) and IR light. In some examples, the IR photodiode may be covered by an IR filter, thus allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., image sensor 230) may lack filters entirely (e.g., color, IR, or any other part of the spectrum) and alternatively use different photodiodes (in some cases, vertically stacked) throughout the pixel array. Different photodiodes throughout the pixel array can have different spectral sensitivity profiles, thereby responding to light of different wavelengths. Monochrome image sensors may also lack filters and therefore lack color depth.

[0057] In some cases, image sensor 230 may optionally or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, opaque and / or reflective masks may be used in PDAF. In some cases, opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, ultraviolet (UV) cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). Image sensor 230 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms in control mechanism 220 may alternatively or additionally be included in image sensor 230. Image sensor 230 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0058] Image processor 250 may include one or more processors, such as one or more ISPs (e.g., ISP 254), one or more host processors (e.g., host processor 252), and / or related processors. Figure 12The computing system 1200 may include one or more processors of any other type of processor 1110 discussed herein. The host processor 252 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 250 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes the host processor 252 and the ISP 254. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 256), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™ This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 256 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer port or interface), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 252 may use the I2C port to communicate with image sensor 230, and ISP 254 may use the MIPI port to communicate with image sensor 230.

[0059] Image processor 250 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 250 can store image frames and / or processed images in random access memory (RAM) 240, read-only memory (ROM) 245, cache, memory unit, another storage device, or some combination thereof.

[0060] Various I / O devices 260 can be connected to the image processor 250. I / O devices 260 may include a display screen, keyboard, keypad, touchscreen, touchpad, touch-sensitive surface, printer, any other output device 1135, any other input device 1145, or some combination thereof. In some cases, text can be entered into the image processing device 205B via the physical keyboard or keypad of the I / O device 260, or via the virtual keyboard or keypad of the touchscreen of the I / O device 260. I / O devices 260 may include one or more ports, jacks, or other connectors that enable a wired connection between the image capture and processing system 200 and one or more peripheral devices, through which the image capture and processing system 200 can receive data from and / or send data to one or more peripheral devices. I / O devices 260 may include one or more wireless transceivers that enable a wireless connection between the image capture and processing system 200 and one or more peripheral devices, through which the image capture and processing system 200 can receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 260 discussed earlier, and they can be considered I / O devices 260 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.

[0061] In some cases, the image capture and processing system 200 may be a single device. In other cases, the image capture and processing system 200 may be two or more separate devices, including an image capture device 205A (e.g., a camera) and an image processing device 205B (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 205A and the image processing device 205B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 205A and the image processing device 205B may be disconnected from each other.

[0062] like Figure 2 As shown, the vertical dashed line will Figure 2 The image capture and processing system 200 is divided into two parts, namely image capture device 205A and image processing device 205B. Image capture device 205A includes a lens 215, a control mechanism 220, and an image sensor 230. Image processing device 205B includes an image processor 250 (including an ISP 254 and a host processor 252), RAM 240, ROM 245, and I / O devices 260. In some cases, certain components illustrated in image capture device 205A (such as ISP 254 and / or host processor 252) may be included in image capture device 205A.

[0063] Image capture and processing system 200 may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, image capture and processing system 200 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some specific implementations, image capture device 205A and image processing device 205B may be different devices. For example, image capture device 205A may include a camera device, and image processing device 205B may include a computing device, such as a mobile phone, desktop computer, or other computing device.

[0064] Although the image capture and processing system 200 is shown to include certain components, those skilled in the art will understand that the image capture and processing system 200 may include more than [other components]. Figure 2 The components shown herein are additional components. Components of the image capture and processing system 200 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image capture and processing system 200 may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits); and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image capture and processing system 200.

[0065] Figure 3This is a block diagram illustrating an example of an image capture system 300. The image capture system 300 includes various components for processing an input image or frame to generate disparity information, such as optical flow information, stereo depth (or disparity) information, or other types of disparity information. As shown, the components of the image capture system 300 include one or more image capture devices 302, a disparity engine 310, and a disparity information consumer 312. The image capture device 302 can generate an image of a scene, and the disparity information consumer 312 can analyze that image and previous images to generate disparity information, such as optical flow, depth information, and / or other types of disparity information, as described in more detail herein. In some aspects, the disparity information consumer 312 can use inverse remapping to construct a confidence map and identify true flow and defective flow. In some aspects, the disparity engine 310 can also provide depth information. For example, the disparity engine 310 can be configured to perform stereo depth estimation between a pair of images to determine disparity, and depth information can be determined based on the disparity.

[0066] Image capture system 300 may include or be part of an electronic device or system. For example, image capture system 300 may include or be part of an electronic device or system, such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles or vehicle computing devices / systems, unmanned aerial vehicle systems, autonomous robots, server computers (e.g., communicating with another device or system such as mobile devices, XR systems / devices, vehicle computing systems / devices, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, camera devices, display devices, digital media players, video streaming devices, or any other suitable electronic device. In some examples, image capture system 300 may include one or more wireless transceivers (or separate wireless receivers and transmitters) for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, WLAN communication, Bluetooth or other short-range communication, any combination thereof, and / or other communication). In some embodiments, components of the image capture system 300 may be part of the same computing device. In some embodiments, components of the image capture system 300 may be part of two or more separate computing devices.

[0067] Although the image capture system 300 is shown as including certain components, those skilled in the art will understand that the image capture system 300 may include more than [other components]. Figure 3The components shown may include more or fewer components. In some cases, additional components of the image capture system 300 may include software, hardware, or one or more combinations of software and hardware. For example, in some cases, the image capture system 300 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or... Figure 3 One or more other software and / or hardware components are not shown. In some embodiments, additional components of the image capture system 300 may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., DSP, microprocessor, microcontroller, GPU, CPU, any combination thereof and / or other suitable electronic circuitry), and / or may include computer software, firmware or any combination thereof, and / or may be implemented using computer software, firmware or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture system 300.

[0068] Image capture device 302 can capture image data and generate images (or frames) based on the image data, and / or provide the image data to disparity engine 310 to generate disparity information. For example, disparity engine 310 can determine optical flow information based on optical flow motion estimation techniques for various purposes. Optical flow motion estimation can be performed on a pixel-by-pixel basis. For example, for an image y Motion estimation for each pixel in the image. f Defines the corresponding pixel in the image x Position within the pixel. Motion estimation for each pixel. f It may include vectors (e.g., motion vectors) indicating the movement of pixels between images. In some cases, an optical flow map (e.g., also referred to as a motion vector map) may be generated based on the calculation of optical flow vectors between images. An optical flow map may include an optical flow vector for each pixel in an image, where each vector indicates the movement of pixels between images. In an exemplary example, the optical flow vector for a pixel may be a displacement vector indicating the movement of the pixel from a first image to a second image (e.g., indicating horizontal and vertical displacements, such as x and y displacements). In some aspects, optical flow estimation techniques may also include depth estimation techniques (e.g., stereo depth estimation).

[0069] In some cases, optical flow maps may include vectors for fewer than all pixels in an image. For example, dense optical flow between images may be computed to generate an optical flow vector for each pixel in at least one of the images, which may be included in the dense optical flow map. In some examples, each optical flow map may include a 2D vector field, where each vector is a displacement vector indicating the movement of a point from a first image to a second image.

[0070] As noted above, optical flow vectors or optical flow maps can be calculated between images in an image sequence. Two images can include two directly adjacent images captured consecutively, or two images separated by a specific distance in the image sequence (e.g., within two images of each other, within three images of each other, or any other suitable distance). In an illustrative example, the images... Pixels Available in image Move a certain distance or displacement .

[0071] For various purposes, parallax information (e.g., optical flow information, depth information, etc.) is provided to parallax information consumer 312. For example, parallax information consumer 312 may be a control system of an autonomous navigation system used to navigate an autonomous vehicle within the physical world. In an example of an autonomous navigation system, optical flow information can be used to identify objects in the environment, such as a person crossing a road, or to identify the movement of stationary objects to detect displacement information related to how the autonomous vehicle is moving. Parallax information consumer 312 may also be a safety system and is configured to identify human movement to trigger recording or to draw the attention of supervisors or other monitoring and control systems.

[0072] In some aspects, the parallax engine 310 may include inverse error correction (e.g., inverse optical flow error correction, inverse depth error correction, etc.) to correct errors detected between two sequential images. In one aspect, the parallax engine 310 is configured to generate optical flow information between a previous image and a current image. The parallax engine 310 performs inverse remapping of the current image based on the optical flow information to generate an estimated previous image. The parallax engine 310 compares the estimated previous image with the previous image to identify defective optical flow and may clean up the optical flow information based on the defective optical flow. In some aspects, cleaning up defective optical flow prevents the introduction of defects that could propagate into subsequent optical flow and produce errors that spread over time. For example, cleaning up defective optical flow can remove various defects that might be introduced between subsequent images (e.g., temporal coherence interruptions, local inconsistencies, etc., as described above).

[0073] As described herein, in some aspects, the systems and techniques can also be used for disparity estimation applications beyond optical flow estimation and stereo depth estimation. Similar techniques can be used to correct stereo depth information based on the one-dimensional disparity between two cameras used to capture two images (referred to as the current frame pair). The systems and techniques can generate depth information from the current frame pair, perform inverse remapping of the current frame pair based on the depth information, and identify defective depth information. In some aspects, except that optical flow is for two-dimensional disparity and stereo depth is for one-dimensional disparity, the techniques for correcting depth information are similar to those described herein for correcting optical flow information. In such aspects, the inversion and self-cleaning operations described herein operate on optical flow in 2D (e.g., in the horizontal (x) and vertical (y) directions) and on stereo depth in 1D (only in the x direction).

[0074] One or more image capture devices 302 may also provide image data to an output device for output (e.g., on a display). In some cases, the output device may also include a storage device. An image or frame may include an array of pixels representing a scene. For example, an image may be: a red-green-blue (RGB) image with red, green, and blue color components per pixel; a lightness, redness, and blueness (YCbCr) image with a lightness component and two chromaticity (redness and blueness) components per pixel; or any other suitable type of color or monochrome image. In addition to image data, the image capture device may also generate supplementary information, such as the amount of time between consecutively captured images, timestamps of image captures, etc.

[0075] Figure 4 This is a conceptual diagram of an optical flow correction system 400 that provides reverse optical flow error correction according to some aspects of this disclosure. Regarding... Figure 4 The described techniques can also be used in other disparity estimation techniques, such as for stereo depth estimation. The optical flow correction system 400 includes an optical flow engine 402 (e.g., included in the disparity engine 310), an inverse remapping engine 404, and a correction engine 406.

[0076] At time t0, optical flow engine 402 is configured to receive first image 410 (or frame). Optical flow engine 402 is configured to extract multiple features from first image 410 to generate feature map F. 1,0 The reverse remapping engine 404 is not configured to receive any information because image 410 is the first image associated with the new optical flow. For example, the optical flow correction system 400 may detect environmental changes that cause the optical flow engine 402 to discard previous data. Non-limiting examples of scene changes may include changes in the average brightness that affect the previous optical flow, such as an autonomous vehicle entering a tunnel.

[0077] At time t1, the optical flow engine 402 receives the second image 420 and the feature map F from the first image 410. 1,0 The optical flow engine 402 extracts features from the second image 420 to generate a feature map F. 2,0 Then generate a feature map F 2,0 and feature map F 2,0 Optical flow information f corresponding to the motion between them 12,i In one aspect, the optical flow engine 402 is configured to use an iterative process to estimate the dense flow field in iteration i to obtain a feature map F. 1,i and F 2,i The pixel-by-pixel displacement of features between them.

[0078] In one respect, the reverse remapping engine 404 is configured to be based on optical flow information f 12,i Remapping feature map F 2,0 For example, the reverse remapping engine 404 is configured to remap feature map F 2,0 Perform reverse remapping to generate the estimated first feature map F'. 2,i The estimated first feature map F' 2,i Based on the remapping function W(F) 2,i , f 12,i ), where W corresponds to

[0079] The reverse remapping engine 404 generates the remapped optical flow information F'. 2,1 = W(F 2,i , f 12,i ), where W corresponds to the reverse remapping operation performed at a specific time, which is the feature map F. 2,i and the corresponding optical flow information f 12,i Generate feature map output F' 2,i The time spent. For example, through F 2,i The query coordinates are correlated and interpolated according to the flow field f. 12,i For F 1,i Each output pixel p x,y Restore F 2,i Pixel-level displacement.

[0080] In some respects, the remapped optical flow information F' 2,i and feature map F 1,0 It is provided to the correction engine 406 for reverse optical flow error correction. The correction engine 406 is configured to determine the feature map F. 1,0 With the remapped optical flow information F' 2,1The difference between them. In one respect, the correction engine 406 can apply the negative of the above difference to the Gaussian kernel function (as illustrated in Equation 1 below), or apply any other appropriately defined function by design.

[0081] (Equation 1)

[0082] In equation 1, C This corresponds to the set of elements in the channel dimension of the corresponding coordinates (x, y) of feature maps F1 and F'2. The Gaussian kernel includes properties whose value range is defined as in Equation 2.

[0083] (Equation 2)

[0084] The maximum value in the range of Gaussian kernel values ​​and Correspondingly, and the minimum value is Correspondingly, the correction engine 406 is also configured to construct a confidence map based on the Gaussian kernel to determine validity (e.g., whether the optical flow is defective). In one aspect, the validity threshold (e.g., denoted as scaler T) and the feature map FF(F1, F'2) from the Gaussian kernel function are applied to the function in Equation 3 to generate the confidence map.

[0085] (Equation 3)

[0086] The confidence plot V (or validity plot) is used to estimate the flow f of iteration i. 12,i Adjustments are made, where f 12,i = V f 12,i ,and" The operator represents Hadamard multiplication. In some aspects, the confidence map V corresponds to the confidence of optical flow between various features in the first image 410 and the second image 420 (and in some cases, between the first image 410 and the third image 430 and / or between the second image 420 and the third image 430). For example, if there is no optical flow error, features from the estimated first image should strongly correspond to the first image. In some aspects, different threshold functions can be configured to improve the detection of defective optical flow. For example, the validity threshold can be based on iterations associated with optical flow (e.g., later optical flow with a higher threshold), the density of features close to the feature (e.g., sparsity), the magnitude of optical flow (e.g., flow guidance), semantic thresholds (e.g., using an attention module), and occlusion detection.

[0087] In some aspects, the optical flow correction system 400 is configured to provide reverse optical flow error correction. For example, the optical flow correction system 400 identifies defective flows based on a reverse remapping of the current image relative to time and prevents the introduction of defective flows to reduce errors. The optical flow correction system 400 compares the reverse-remapped image with a ground truth reference to determine the confidence level of the optical flow between the previous image and the current image, and this excludes the introduction of defective flows. Although Figure 4 An example is shown of the reverse remapping engine 404 receiving feature information, but the reverse remapping engine 404 can also use bitmap images to construct a confidence map and then clean up the optical flow information.

[0088] Figure 5 This is a block diagram of an optical flow correction system 500 configured to perform reverse optical flow error correction according to some aspects of this disclosure. In some aspects, the optical flow correction system 500 includes an encoder 510, an optical flow engine 520, a remapping engine 530, a confidence engine 540, and a correction engine 550. An image 502 is received by the encoder 510, which is configured to identify features in each image. These features are typically referred to as vectors or terms and refer to a transformed representation of the input data generated by the encoding process. The encoder 510 extracts relevant features from the input data (e.g., image 502) and transforms these features into a concise and meaningful format. These features encapsulate the fundamental information, patterns, and characteristics inherent in the input data (e.g., image 502) and are a refined representation more suitable for analysis by downstream components of the optical flow correction system 500.

[0089] In one aspect, encoder 510 transforms and processes information, particularly in the context of artificial intelligence and machine learning. Encoder 510 is configured to transform raw input data (e.g., image 502) into a structured format that facilitates analysis and pattern recognition. For example, encoder 510 may include, or be part of, a machine learning system (e.g., a deep neural network (DNN), a convolutional neural network (DNN), a transformer neural network, a diffuse neural network, any combination thereof, and / or other types of machine learning systems). In such an example, encoder 510 may process image 502 to generate features representing image 502. In some cases, these features are multidimensional vectors or terms. In some aspects, the features of image 502 may be used by various other engines for various purposes. For example, these features may be used by optical flow engine 520 to identify motion of features between images, or by attention engine (not shown) configured to identify the importance of features compared to other features within an image.

[0090] Optical flow engine 520 is configured to generate optical flow information based on detected motion between two images. Remapping engine 530 is configured to remap image 502 based on the optical flow information from optical flow engine 520. For example, remapping engine 530 is configured to perform a backward remapping of image 502 in time and generate an estimated previous image. In some cases, optical flow engine 520 may perform different remapping operations. For example, optical flow engine 520 may forward remap a previous image to an estimated current image and compare the estimated current image with the current image. In some aspects, combining the forward-remapped image and the backward-remapped image can be used to detect occlusion of objects.

[0091] The confidence engine 540 is configured to receive images (e.g., previous images, current images, estimated previous images, and estimated current images) and optical flow information from the optical flow engine 520 and the remapping engine 530. In some aspects, the confidence engine 540 is configured to generate confidence information based on the optical flow information and various images. In some aspects, the confidence engine 540 may use one or more engines to dynamically construct the confidence information. For example, the confidence engine 540 may include a process engine 541, a sparsity engine 542, a stream-guided engine 543, a semantic engine 544, and an occlusion detection engine 545.

[0092] Each of the process engine 541, sparsity engine 542, flow-guided engine 543, and semantic engine 544 is configured to dynamically adjust a validity threshold associated with the confidence of optical flow relative to detected features. The validity threshold identifies whether the optical flow is considered valid or invalid. For example, the difference between an estimated image (e.g., an estimated previous image) and the estimated image is the confidence (e.g., probability) that a feature in the estimated previous image is the same feature in the previous image. The confidence engine 540 uses the validity threshold to generate a mask identifying valid and invalid optical flow and provides this mask to the correction engine 550. The mask can be a two-dimensional bitmap, a multi-dimensional vector, or a matrix that can be applied to optical flow information.

[0093] In one aspect, process engine 541 dynamically adjusts the validity threshold of confidence engine 540 based on iterations (or processes) of the confidence engine relative to the initial state of the current environment. The initial state corresponds to the time when the control state is reset based on feedback. For example, Figure 4Time t0 exemplifies the initial state caused by a change within the optical flow correction system 500 that prevents optical information from being fed back into time t0. Non-limiting examples of events causing feedback include sudden changes in brightness (e.g., entering a tunnel while driving), sudden changes in background content (e.g., causing the autonomous vehicle to turn), etc. The process engine 541 is configured to increase the validity threshold based on the number of iterations to prevent the introduction of errors over longer iterations. For example, when the autonomous vehicle is traveling on a long road segment, the environment does not change significantly, and the validity threshold used to introduce new features and errors into the optical flow is increased. However, when the autonomous vehicle turns or experiences a sudden change, the optical flow correction system 500 decreases the validity threshold to ensure that useful new features and environmental characteristics are introduced into the optical flow. In some aspects, the process engine 541 may increase the validity threshold to disable self-cleaning of the optical flow correction system 500.

[0094] The sparsity engine 542 dynamically adjusts the validity threshold of the confidence engine 540 based on the sparsity of features within the current environment and / or optical flow. For example, the sparsity engine 542 may adjust the validity threshold based on the sparsity of neighboring features (e.g., neighboring pixels). For instance, in regions where validity is sparse (e.g., for most pixels in that region), V... (x,y) =0), the sparsity engine 542 can apply a more conservative validity threshold in this region.

[0095] The stream-guided engine 543 dynamically adjusts the validity threshold of the confidence engine 540 based on optical flow. The stream-guided engine 543 can be based on the flow at a pixel (e.g., f...). 12,i The validity threshold of a pixel is dynamically adjusted based on the magnitude of its optical flow. For example, when the magnitude of the pixel's flow is small, the flow-guided engine 543 can increase the validity threshold because it is relatively safer to incorrectly detect smaller movements, and it is simpler to correct optical flow in later flows than for larger movements.

[0096] The semantic engine 544 dynamically adjusts the validity threshold of the confidence engine 540 based on features, scene, or certain semantic attributes of the region or pixel. The semantic engine 544 may include an attention module configured to determine in-scene attention between two images. For example, the semantic engine 544 may use the attention module to distinguish between background and foreground content to identify features to be tracked. Based on the attention associated with a feature, the semantic engine 544 can dynamically adjust the validity threshold. For example, features with less attention (e.g., considered less important) may have a higher validity threshold, and features with more attention may have a lower validity threshold.

[0097] The confidence engine 540 may also include an occlusion detection engine 545 for detecting objects entering or leaving an occluded state. In one aspect, the occlusion detection engine 545 may estimate two optical flows (e.g., forward and backward flows) in two directions, and then perform remapping in the forward and backward directions. For example, using two encoded feature maps F1 and F2 from an input image pair, the occlusion detection engine 545 estimates the two flows f 12 and f 21 Furthermore, cyclic remapping is applied in both directions within the image pair. Cyclic remapping generates two estimated optical flows F1' = W( F 1, f 21 ) and F1''=W(F1, f 12 These estimated optical flow can be used to identify whether an object is occluded.

[0098] In some aspects, the confidence engine 540 may also include one or more ML models (not shown) configured to perform a combination of detection engines. For example, one or more ML models may be trained based on a combination of techniques to vary the validity threshold of the confidence engine 540 based on process (e.g., process engine 541), sparsity (e.g., sparsity engine 542), the magnitude of the stream (e.g., stream-guided engine 543), and attention (e.g., semantic engine 544).

[0099] The correction engine 550 receives the estimated optical flow from the confidence engine 540 and cleans and outputs it based on the optical flow information 560 from the optical flow correction system 500. By cleaning up invalid flows from the optical flow information, the correction engine 550 prevents the introduction of errors based on the analysis performed in the confidence engine 540 and improves the accuracy and quality of the optical flow information of the optical flow correction system 500.

[0100] Figure 6A This is a conceptual illustration of the optical flow and possible errors that may be introduced between two images, based on some aspects of this disclosure. In some aspects, Figure 6A A first feature 602 and a second feature 604, present at time t0 and moving across the image at time t1, are illustrated. As the illustration indicates, the fill pattern corresponds to the positioning of the first feature 602 and the second feature 604. The first feature 602 moves based on a first optical flow 612. However, the second optical flow 614 associated with the second feature 604 is incorrect and identifies a third object 606. The third object 606 may exist at time t0 and is omitted for simplicity. The error illustrated by the second optical flow 614 will accumulate in subsequent images and degrade the quality of object detection and tracking.

[0101] Figure 6BThis is a conceptual illustration of a prior image estimated according to some aspects of this disclosure and an identification of errors. In some aspects, as described above, the imaging system can perform reverse optical flow error correction by remapping the image at time t1 based on a first optical flow 612 and a second optical flow 614 to generate an estimated image corresponding to time t0. For example, for each flow information of the image at time t1, the movement of each corresponding feature is reverse-mapped. In this case, the second feature 604 can be removed from the estimated image because at time t1, the flow may not be associated with the second feature 604. The flow may be associated with the second feature 604 at time t1, and this would be incorrect and is omitted for simplicity of explanation.

[0102] Figure 6C This is a conceptual example of a confidence map that a cleaning engine can use to remove errors introduced into the optical flow, based on some aspects of this disclosure. The image capture system will (e.g., Figure 6A The previous image at time t0 and the estimated optical flow (e.g., as shown in the image) are compared with the estimated optical flow. Figure 6B (as shown) are compared to generate Figure 6C The confidence map is shown in the image. In this example, the third object 606 in the image at time t1 does not correspond to the second feature 604 of the object in the image at time t0, and the confidence map includes regions 620 indicating defective flow in the initial optical information. The density of the fill pattern in the confidence map indicates the likelihood of introducing errors into the original optical flow information before correction. As described above, the confidence map can be further modified based on various engines and ML models to improve detection based on additional criteria such as the process of process engine 541, the sparsity of sparsity engine 542, the magnitude of flow of flow of flow-guided engine 543, etc.

[0103] Figure 6D This is a conceptual illustration of cleaned optical flow based on some aspects of this disclosure. In some aspects, a confidence map is provided to a correction engine (e.g., Figure 5 Correction engine 550, Figure 4 The correction engine 406 cleans up the optical flow information of defective streams. In this case, the first optical flow 612 of the first feature 602 is retained, and the optical flow associated with the second feature 604 and the third object 606 (e.g., the second optical flow 614) is removed.

[0104] Figures 7A to 7C Examples of optical flow without a reference truth value, optical flow without reverse optical flow error correction, and optical flow with reverse optical flow error correction are provided according to some aspects of this disclosure.

[0105] In particular, Figure 7A The reference ground truth of optical flow between two consecutive images is illustrated. Figure 7BExamples of using an optical flow engine (e.g., Figure 5 The optical flow engine 520 performs optical flow correction on two consecutive images without any reverse optical flow error correction. For example... Figure 7B As shown, region 710 is Figure 7A The missing flow information exists in the optical flow, and corresponds to at least one region containing optical flow information of defective optical flow. Figure 7C An example is shown using the same optical flow engine but with additional reverse optical flow error correction (e.g., using...). Figure 5 The optical flow correction system (500) corrects the same optical flow in two consecutive images. Figure 7C In the middle, area 710 is included Figure 7B Missing flow information in the optical information. In this case, reverse optical flow error correction prevents the introduction of errors that may diffuse downstream and be superimposed in the optical information.

[0106] Figure 8 This is a flowchart illustrating an example process 800 of capturing an image and estimating optical flow information using an image capture system according to various aspects of this disclosure. Process 800 may be performed by a computing device (e.g., including an image sensor), or components or systems of the computing device (e.g., chipsets, one or more processors, one or more machine learning models (such as one or more neural networks), any combination thereof, and / or other components or systems). In some examples, the computing device may include a mobile wireless communication device, a vehicle (e.g., an autonomous or semi-autonomous vehicle, a wireless-enabled vehicle, and / or other types of vehicle), or a computing device or system of a vehicle, a robotic device or system (e.g., manufacturing), a camera, an XR device, or another computing device. In one illustrative example, a computing system (e.g., computing system 1100) may be configured to perform all or part of process 800. In one illustrative example, optical flow correction system 400 and / or optical flow correction system 500 may be configured to perform all or part of process 800. For example, computing system 1100 may include components of system 400 and / or 500 and may be configured to perform all or part of process 800. In another exemplary example, the ISP (such as ISP 254) may be configured to perform all or part of process 800.

[0107] Although example process 800 depicts a specific sequence of operations, this sequence may be changed without departing from the scope of this disclosure. For example, some of the operations depicted may be performed in parallel or in a different order that does not substantially affect the functionality of process 800. In other examples, different components of the example device or system implementing process 800 may perform their functions substantially simultaneously or in a specific order.

[0108] At box 802, the computing system (e.g., computing system 1100) may obtain first disparity information associated with the current image. In one aspect, the first disparity information may include first optical flow information estimating a first movement of a first feature to a first destination location in the current image. In another aspect, the first disparity information may include depth information representing the depth of the first feature. For example, in such an aspect, the computing system may include two or more image capture devices for stereo depth estimation. For example, the computing devices may use two image sensors to determine the distances (e.g., stereo depth information) from the two image sensors to the first feature. In some aspects, the computing system includes one or more image capture devices.

[0109] At box 804, the computing system can remap the current image based on the first disparity information to obtain an estimated previous image.

[0110] At box 806, the computing system may determine a confidence map associated with the confidence of the first disparity information based on the difference associated with the estimated previous image (e.g., the difference between the previous image and the estimated previous image). In one aspect, the computing system may determine the difference between the previous image and the estimated previous image. The confidence map includes a first region corresponding to a first feature that is valid in the first disparity information. The confidence map includes different regions corresponding to different features that are false alarms (e.g., defective stream information) in the first disparity information.

[0111] In some cases, a confidence map is determined based on a first threshold at a first time point, and then based on a second threshold at a second time point after the first time point. The second threshold includes a higher confidence level than the first threshold. For example, the second threshold increases with iterations of optical flow.

[0112] In some aspects, the computational system can determine the confidence map based on different criteria. Confidence can be based on the sparsity of features near the first feature, the magnitude of flow, and attention, etc. In one aspect, to determine the confidence map at box 806, the computational system can determine the sparsity of the region associated with the first feature in the current image or a previous image, and determine a threshold corresponding to the confidence of the first feature in the first disparity information based on the sparsity.

[0113] In one aspect, to determine the confidence map at box 806, the computing system may determine a first motion amount value associated with a first feature in the current image. Based on the first motion amount value, the computing system may determine a first threshold corresponding to the confidence of the first feature within the first disparity information. For smaller motions, the threshold may be higher. The computing system may determine, based on the first threshold, whether the region in the confidence map associated with the first feature corresponds to the true disparity information.

[0114] In one aspect, to determine the confidence map at box 806, the computational system may determine an attention associated with a first feature. This attention corresponds to the importance of the first feature associated with at least one other feature in the first disparity information. The computational system may determine a first threshold corresponding to authentication of the first disparity information of the first feature based on the attention. In some aspects, the attention includes information identifying the importance of the first feature in the previous and current images compared to other features in the previous and current images.

[0115] In one aspect, to determine the confidence map at box 806, the computational system can use a combination of the described techniques. In other aspects, the ML model can be configured to determine the threshold based on a combination of iteration, flow magnitude, and attention, etc.

[0116] At box 808, the computing system may apply a confidence map to the first disparity information to generate updated first disparity information. The computing system may also apply a confidence map to the first disparity information to remove first disparity information (e.g., defective disparity information) and generate updated first disparity information.

[0117] In some aspects, the computing system can use first disparity information to identify occluded objects. The computing system can obtain second disparity information associated with a previous image. In some aspects, the second disparity information estimates a second movement of a first feature within the current image or a previous image, or estimates the depth of the first feature. The computing system can remap the previous image based on the second disparity information to obtain an estimated current image; and generate a second confidence map associated with the second disparity information based on the differences associated with the estimated current image. The computing system can apply the second confidence map to the second disparity information to generate updated second disparity information. The computing system can determine whether the first feature is occluded in the current image or a previous image based on the updated first and second disparity information.

[0118] In some examples, the processes described herein (e.g., process 800 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 800 may be performed by a device having Figure 12 The computing architecture of the computing system 1200 shown is a computing device (e.g., Figure 2 The image capture and processing system 200 in the system is used to perform this operation.

[0119] Process 800 is illustrated as a logic flowchart, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the method.

[0120] Process 800 and / or other methods or processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or implemented in a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0121] As noted above, various aspects of this disclosure may utilize machine learning models or systems. Figure 9 This is an exemplary example of a deep learning neural network 900 used to implement the machine learning-based alignment prediction described above. Input layer 920 includes input data. In one exemplary example, input layer 920 may include data representing pixels of an input video frame. Neural network 900 includes multiple hidden layers 922a, 922b through 922n. Hidden layers 922a, 922b through 922n include “n” hidden layers, where “n” is an integer greater than or equal to one. The multiple hidden layers can include as many layers as needed for a given application. Neural network 900 also includes an output layer 924 that provides the output produced by the processing performed by hidden layers 922a, 922b through 922n. In one exemplary example, output layer 924 may provide a classification of objects in the input video frame. The classification may include categories identifying activity types (e.g., looking up, looking down, eyes closed, yawning, etc.).

[0122] Neural network 900 is a multi-layered neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, neural network 900 may include a feedforward network, in which case there is no feedback connection where the network's output is fed back into itself. In some cases, neural network 900 may include a recurrent neural network, which may have loops that allow information to be carried across nodes when reading input.

[0123] Information can be exchanged between nodes via node-to-node interconnects between layers. Nodes in input layer 920 can activate the node set in the first hidden layer 922a. For example, as shown, each input node in input layer 920 is connected to each node in the first hidden layer 922a. Nodes in the first hidden layer 922a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 922b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 922b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 922n can activate one or more nodes in output layer 924, at which the output is provided. In some cases, although a node in neural network 900 (e.g., node 926) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.

[0124] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 900. Once the neural network 900 is trained, it can be called a trained neural network, which can be used to classify one or more activities. For example, the interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 900 to adapt to the input and learn as more and more data is processed.

[0125] The neural network 900 is pre-trained to process features from the data in the input layer 920 using different hidden layers 922a, 922b to 922n, in order to provide an output through the output layer 924. In an example where the neural network 900 is used to identify features and / or objects in an image, the neural network 900 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, with each training frame having a label indicating a feature in the image (for a feature extraction machine learning system) or a label indicating the category of activity in each frame. In an example where object classification is used for illustrative purposes, a training frame may include an image of the number 2, in which case the image label could be [00 1 0 0 0 0 0 0 0].

[0126] In some cases, the neural network 900 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process includes forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. The process can be repeated a certain number of iterations for each training image set until the neural network 900 is trained well enough that the weights of each layer are accurately tuned.

[0127] For an example of identifying features and / or objects in an image, the forward pass may include passing a training image through a neural network 900. The weights are initially randomized before training the neural network 900. As an illustrative example, a frame may include a numerical array representing pixels of an image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0128] As noted above, for the first training iteration of the neural network 900, the output will likely include values ​​due to the weights being randomly chosen during initialization, without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 900 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function includes mean squared error (MSE), defined as... The loss can be set to equal to The value of .

[0129] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. The Neural Network 900 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as...) can be calculated. dL / dW ,in W These are the weights at a specific layer, used to determine the weights that contribute the most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... ,in w Indicates weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0130] Neural Network 900 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 900 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), and other examples.

[0131] Figure 10 This is an exemplary example of a CNN 1000. The input layer 1020 of the CNN 1000 includes data representing an image or frame. For example, the data may include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image may be passed through a convolutional hidden layer 1022a, an optional non-linear activation layer, a pooling hidden layer 1022b, and a fully connected hidden layer 1022c to obtain an output at the output layer 1024. Although... Figure 10Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in a CNN 1000. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0132] The first layer of CNN 1000 is a convolutional hidden layer 1022a. Convolutional hidden layer 1022a analyzes the image data input to layer 1020. Each node in convolutional hidden layer 1022a is connected to a region of the input image called a receptive field (pixel). Convolutional hidden layer 1022a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in convolutional hidden layer 1022a. For example, the region of the input image covered by the filter at each convolutional iteration will be the filter's receptive field. In an exemplary example, if the input image consists of a 28×28 array, and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in convolutional hidden layer 1022a. Each connection between a node and its receptive field learns weights, and in some cases, learns an overall bias, allowing each node to learn to analyze its specific local receptive field in the input image. Each node in hidden layer 1022a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the video frame example, the filter will have a depth of 3 (based on the three color components of the input image). An exemplary example size of the filter array is 5×5×3, corresponding to the size of the receptive field of the node.

[0133] The convolutional property of the convolutional hidden layer 1022a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 1022a may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 1022a. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 1022a. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or other suitable amounts. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus producing a sum value determined for each node of the convolutional hidden layer 1022a.

[0134] The mapping from the input layer to the convolutional hidden layer 1022a is called an activation map (or feature map). An activation map includes values ​​for each node representing the filter results at each location of the input volume. Activation maps may include arrays containing various sums of values ​​produced by the filter for each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 1022a may include several activation maps to identify multiple features in the image. Figure 10 The example shown includes three activation maps. Using the three activation maps, the convolutional hidden layer 1022a can detect three different types of features, each of which is detectable across the entire image.

[0135] In some examples, nonlinear hidden layers can be applied after the convolutional hidden layer 1022a. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can increase the nonlinearity of the CNN 1000 without affecting the receptive field of the convolutional hidden layer 1022a.

[0136] A pooling hidden layer 1022b can be applied after the convolutional hidden layer 1022a (and, in use, after the non-linear hidden layer). The pooling hidden layer 1022b is used to simplify the information in the output of the convolutional hidden layer 1022a. For example, the pooling hidden layer 1022b can take each activation map output from the convolutional hidden layer 1022a and use a pooling function to generate a compressed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 1022a uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to the activation maps included in the convolutional hidden layer 1022a. Figure 10 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 1022a.

[0137] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 1022a. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes from the previous layer (where each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 1022a, the output from pooling hidden layer 1022b will be an array of 18×12 nodes.

[0138] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0139] Intuitively, pooling functions (e.g., max pooling, L2-norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of an image, discarding the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 1000.

[0140] The final connection in the network is a fully connected layer that connects each node from the pooling hidden layer 1022b to each output node in the output layer 1024. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 1022a comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling layer 1022b comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region across each of the three feature maps. Extending this example, the output layer 1024 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 1022b is connected to each node of the output layer 1024.

[0141] The fully connected layer 1022c takes the output of the previous pooling hidden layer 1022b (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 1022c can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 1022c and the pooling hidden layer 1022b can be computed to obtain the probabilities for different classes. For example, if CNN 1000 is being used to predict whether an object in a video frame is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the upper left and upper right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0142] In some examples, the output from output layer 1024 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes the CNN 1000 must choose from when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector represents objects of ten different classes as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0143] Figure 11This is an illustration of an example of a system 1100 (e.g., a neural P-frame decoding system) that encodes video frames using optical flow. As illustrated, example system 1100 includes a motion prediction system 1102, a remapping engine 1104, and a residual prediction system 1106. Motion prediction system 1102 and residual prediction system 1106 may include any type of machine learning system (e.g., using one or more neural networks and / or other machine learning models, architectures, networks, etc.).

[0144] In some aspects, motion prediction system 1102 may include one or more machine learning systems, which in some cases may include neural networks (e.g., one or more autoencoders, deep neural networks (DNNs), convolutional neural networks (CNNs), transformer neural networks, diffuse neural networks, any combination thereof, and / or other types of neural networks). In an exemplary example, the encoder network 1105 and decoder network 1103 of motion prediction system 1102 may be implemented as an autoencoder (e.g., also referred to as a "motion autoencoder" or "motion AE"). In some cases, motion prediction system 1102 may be used to implement optical flow correction techniques. For example, encoder network 1105 may include optical flow correction system 1112, which can perform optical flow motion estimation techniques. For example, optical flow correction system 1112 can perform the above-mentioned... Figure 4 Optical flow correction system 400 and / or Figure 5 The illustrated optical flow correction system 500 describes the techniques. In some aspects, the residual prediction system 1106 may include one or more machine learning systems, which in some cases may include neural networks (e.g., one or more autoencoders, deep neural networks (DNNs), convolutional neural networks (CNNs), transformer neural networks, diffuse neural networks, any combination thereof, and / or other types of neural networks). In an illustrative example, the encoder network 1107 and decoder network 1109 of the residual prediction system 1106 may be implemented as autoencoders (e.g., also referred to as "residual autoencoders" or "residual AEs"). Although Figure 11 The example P-frame decoding system 1100 is shown to include certain components, but those skilled in the art will understand that the example P-frame decoding system 1100 may include more than Figure 11 The components shown are fewer or more components.

[0145] In an exemplary example, for a given time t System 1100 can receive input frames and reference frame In some respects, the reference frame It can be in time t Previously (e.g., at time) Previously reconstructed frames generated at (e.g., as by the hat operator ") (Instructions). Input frame and reference frame It can be associated with or otherwise obtained from the same video data sequence (e.g., as consecutive frames, etc.). For example, input frames. It can be in time t The current frame at that location, and the reference frame. It can be immediately following the input frame in time or sequence. Previous frames. In some cases, reference frames can be received from the Decoded Picture Buffer (DPB) of Example System 1100. In some cases, the input frame It can be a P-frame and a reference frame. It can be an I-frame, P-frame, or B-frame. For example, a reference frame. It can be reconstructed or generated previously by an I-frame decoding system (e.g., which may be part of a device that includes a P-frame decoding system 1100 or a device different from the device that includes a P-frame decoding system 1100), by a P-frame decoding system 1100 (or a P-frame decoding system of a device different from the device that includes a P-frame decoding system 1100), or by a B-frame decoding system (e.g., which may be part of a device that includes a P-frame decoding system 1100 or a device different from the device that includes a P-frame decoding system 1100).

[0146] like Figure 11 As depicted, the motion prediction system 1102 receives a reference frame. and the current (e.g., input) frame As input, the motion prediction system 1102 can determine the reference frame. pixels and input frames The motion between pixels (e.g., represented by vectors such as optical flow motion vectors). As described above (e.g., regarding...). Figure 4 and / or Figure 5 The optical flow correction system 1112 can use the techniques described herein to generate corrected optical flow (e.g., optical flow 560). The motion prediction system 1102 can then encode the determined motion and, in some cases, decode it for the input frame. Predicted movement .

[0147] For example, the encoder network 1105 of the motion prediction system 1102 can be used to determine the current frame. and reference frame The motion between them (e.g., motion information). The optical flow correction system 1112 can use the techniques described herein (e.g., as per the description of motion). Figure 4 and / or Figure 5 The discussed methods are used to generate a corrected version of the motion information (e.g., optical flow 560). In some aspects, the encoder network 1105 can encode the determined motion information into a latent representation (e.g., represented as latent data). For example, in some cases, the encoder network 1105 can map the determined motion information to a latent code, which can be used as latent data. The encoder network 1105 can additionally or alternatively transmit data via latent data. The associated latent code performs entropy decoding to... Converted into a bitstream. In some examples, the encoder network 1105 can quantize the latent data. (For example, before performing entropy decoding on the latent code). Quantizing latent data may include latent data. Quantitative representation of latent data. In some cases, latent data... It may include neural network data (e.g., activation maps or feature maps of neural network nodes) representing one or more quantization codes.

[0148] In some respects, the encoder network 1105 can store latent data. , will latent data The data is transmitted to the decoder network 1103 included in the motion prediction system 1102, and / or the latent data can be transmitted to the decoder network 1103. Transmit to latent data Another device or system that performs decoding. Upon receiving the latent data... At that time, decoder network 1103 can process latent data. Decoding (e.g., inverse entropy decoding, dequantization, and / or reconstruction) is performed to generate a reference frame. pixels and input frames Predicted motion between pixels For example, decoder network 1103 can process latent data. Decode to generate optical flow map The optical flow map includes a reference frame. Some (or all) of the pixels included in the input frame are mapped to the input frame. One or more motion vectors of pixels. The encoder network 1105 and decoder network 1103 can be trained and optimized using training data (e.g., training images or frames) and one or more loss functions, as will be described in more detail below.

[0149] In one exemplary example, encoder network 1105 and decoder network 1103 may be included in one or more machine learning systems, which in some cases may include neural networks (e.g., one or more autoencoders, deep neural networks (DNNs), convolutional neural networks (CNNs), transformer neural networks, diffusing neural networks, any combination thereof, and / or other types of neural networks). Encoder network 1105 may include features for quantizing latent data. (For example, generated as the output of encoder network 1105 of motion prediction system 1102) and one or more components that convert quantized latent data into a bit stream. The bit stream generated from the quantized latent data can be provided as input to decoder 1103 of motion prediction system 1102.

[0150] In some examples, predicting motion This can include optical flow information or data (e.g., an optical flow graph including one or more motion vectors), dynamically convolutional data (e.g., a matrix or kernel for data convolution), or block-based motion data (e.g., motion vectors for each block). In an illustrative example, motion is predicted. This may include optical flow maps. In some cases, as previously described, optical flow maps... May include input frames The motion vector for each pixel (e.g., the first motion vector for the first pixel, the second motion vector for the second pixel, etc.). The motion vector can represent the motion vector for the current frame. The pixels in the reference frame Motion information determined by the corresponding pixel in the image (e.g., motion information determined by encoder network 1105).

[0151] The remapping engine 1104 of system 1100 can obtain an optical flow map generated as the output of motion prediction system 1102 (e.g., generated as the output of decoder network 1103). For example, the remapping engine 1104 can retrieve optical flow maps from a storage device. Alternatively, optical flow maps can be received directly from the motion prediction system 1102. The remapping engine 1104 can use optical flow graphs. To remap (e.g., by performing motion compensation) the reference frame The pixels, thus resulting in the generation of remapped frames. In some respects, the remapped frames Also known as motion-compensated frames (For example, by using the optical flow diagram) The corresponding motion vectors in the reference frame are used to remap the reference frame. (Generated based on pixels). For example, remapping engine 1104 can be generated based on pixels included in the optical flow map. The motion vectors (and / or other motion information) in the reference frame will be used as reference frames. The pixels are moved to new positions to generate motion-compensated frames. .

[0152] As mentioned above, in order to generate the remapped frame System 1100 can perform motion compensation by predicting the input frame. and reference frame Optical flow between And subsequently, by using optical flow maps Remapping reference frame To generate motion-compensated frames However, in some cases, based on optical flow maps... Generated frame predictions (e.g., motion-compensated frames) It may not be accurate enough to capture the input frame. Represented as a reconstructed frame For example, there may exist a combination of input frames. One or more occluded areas in the depicted scene, over-illumination, under-illumination, and / or causing motion-compensated frames Not accurate enough to be used as a reconstructed input frame Other effects.

[0153] The residual prediction system 1106 can be used to correct or otherwise improve motion-compensated frames. Associated predictions. For example, residual prediction system 1106 can generate one or more residuals, which system 1100 can then correlate with motion-compensated frames. Combined to generate a more accurate reconstructed input frame (For example, to more accurately represent the underlying input frame) Reconstructed input frame In an exemplary example, such as Figure 11 As described, system 1100 can transmit data from input frames. Subtract the predicted (e.g., motion-compensated) frames (For example, using subtraction 1108 to determine) to determine the residual. For example, in motion-compensated prediction frames. After being determined by the remapping engine 1104, the P-frame decoding system 1100 can determine the motion-compensated predicted frame. and input frame The difference between them (e.g., using subtraction 1108) is used to determine the residual. .

[0154] In some respects, the encoder network 1107 of the residual prediction system 1106 can predict the residuals. Encoding into latent data Among them, latent data Residual For example, encoder network 1107 can convert residuals Mapping to data that can be used as latent data The latent code. In some cases, the encoder network 1107 can extract the latent data by performing entropy decoding on the latent code. The data is converted into a bitstream. In some examples, the encoder network 1107 may additionally or alternatively quantize the latent data. (For example, before performing entropy decoding). Quantizing latent data. May include residuals Quantitative representation of latent data. In some cases, latent data... It may include neural network data (e.g., activation maps or feature maps of neural network nodes) representing one or more quantization codes. In some aspects, the encoder network 1107 may store latent data. The latent data is sent to or otherwise provided to the decoder network 1109 of the residual prediction system 1106. And / or can latent data Transmit to latent data Another device or system that performs decoding. Upon receiving the latent data... At that time, decoder network 1109 can process latent data. Decoding (e.g., inverse entropy decoding, dequantization, and / or reconstruction) is performed to generate the prediction (e.g., post-decoding) residual. In some examples, training data (e.g., training images or frames) and one or more loss functions can be used to train and optimize the encoder network 1107 and the decoder network 1109, as described below.

[0155] In an exemplary example, encoder network 1107 and decoder network 1109 may be included in a residual prediction autoencoder. The residual autoencoder may include features for quantizing latent data. (For example, the latent data) One or more components that generate (as output of the encoder network 1107 of the residual autoencoder) and convert quantized latent data into a bit stream. The generated bitstream can be used as input to the decoder 1109 of the residual autoencoder.

[0156] Predicted residuals (For example, generated by the decoder network 1109 and / or the residual autoencoder used to implement the residual prediction system 1106) can be used with motion-compensated prediction frames. (For example, the optical flow map generated by the decoder network 1103 is used by the remapping engine 1104) (To generate) used together to generate a representation in time t Input frame at Reconstructed input frame .

[0157] For example, system 1100 can predict residuals and motion compensation prediction frames Add them (e.g., using the addition operation 1110) or otherwise combine them to generate a reconstructed input frame. In some cases, the decoder network 1109 of the residual prediction system 1106 can predict the residuals. Add to motion-compensated frame prediction In some examples, the input frame has been reconstructed. This can also be referred to as a decoded frame and / or a reconstructed current frame. Reconstructed current frame It can be output for storage (e.g., in a decoded picture buffer (DPB) or other storage device), transmission, display, further processing (e.g., as a reference frame in further inter-frame prediction, for post-processing, etc.) and / or for any other purpose.

[0158] In one exemplary example, the P-frame decoding system 1100 may transmit, in one or more bit streams, representations of optical flow maps or other motion information (e.g., latent data) to another device for decoding. () latent data and representations of residual information (e.g., latent data) Latent data. In some cases, other devices may include those configured to access latent data. and A video decoder for decoding. In an exemplary example, another device may include a video decoder that implements one or more portions of the P-frame decoding system 1100, the motion prediction system 1102, and / or the residual prediction system 1106 (e.g., a residual autoencoder), as described above.

[0159] Other devices or video decoders can use the latent data generated as the output of the decoder network 1103 included in the motion prediction system 1102. Optical flow diagram Decoding and / or other predicted motion information. Other devices or video decoders may additionally use the latent data generated as the output of the decoder network 1109 included in the residual prediction system 1106 (e.g., as the output of a residual autoencoder including the decoder network 1109). To the residual Decoding is then performed. Other devices or video decoders can subsequently use the optical flow graph. and residual To generate decoded (e.g., reconstructed) input frames .

[0160] For example, when the video decoder implements an architecture that is the same as or similar to that of the P-frame decoding system 1100 described above, the video decoder may include a remapping engine (e.g., the same as or similar to remapping engine 1104) that receives the decoded optical flow map. and reference frame As input, the video decoder remapping engine can be based on the reference frame. Motion vectors and / or other motion information determined by some (or all) pixels, based on the decoded optical flow map. To remap reference frames The video decoder remapping engine can output motion-compensated frame predictions. (For example, as described above relative to the output of remapping engine 1104). The video decoder can then decode the residual. and motion-compensated frame prediction They can be added or otherwise combined to generate a decoded (e.g., reconstructed) input frame. (For example, as described above with respect to the output of addition operation 1110).

[0161] In some examples, training data and one or more loss functions can be used to train and / or optimize the motion prediction system 1102 and / or the residual prediction system 1106. In some cases, the motion prediction system 1102 and / or the residual prediction system 1106 can be trained end-to-end (e.g., where all neural network components are trained during the same training process). In some aspects, the training data may include multiple training images and / or training frames. In some cases, loss functions (e.g., ...) can be used. Loss Training is performed using a motion prediction system 1102 and / or a residual prediction system 1106 that processes training images or frames.

[0162] In one example, the loss function (e.g., Loss ) can be given as Loss = D +β R ,in D It is a given frame (e.g., such as an input frame) ) and its corresponding reconstructed frame (e.g., Distortion between ) . For example, distortion D It can be determined as D ( , β is a hyperparameter that can be used to control the bit rate (e.g., bits per pixel), and R It is used to convert residuals (e.g., residuals) Converting data into a compressed bitstream (e.g., latent data) The number of bits. In some examples, distortion. D It can be calculated based on one or more of the following: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Multi-Scale SSIM (MS-SSIM), etc. In some aspects, the parameters (e.g., weights, biases, etc.) of the motion prediction system 1102 and / or the residual prediction system 1106 can be tuned using one or more training datasets and one or more loss functions until the example system 1100 achieves the desired video decoding result.

[0163] In some respects, the machine learning systems or neural networks described in this paper (e.g., such as...) Figure 4 System 400 Figure 5 System 500 Figure 9 Deep learning network 900, Figure 10 CNN 1000 Figure 11 Training of one or more of the systems (such as System 1100, etc.) can be performed using online training, offline training, and / or various combinations of online and offline training. In some cases, online may refer to processing input data during its operation (e.g., such as...). Figure 5 Image 502, etc., can be used as a time period for performing parallax correction processing (e.g., optical flow correction, etc.) implemented by the systems and techniques described herein. In some examples, offline may refer to an idle time period or a time period during which no input data is processed. Additionally, offline may be based on one or more time conditions (e.g., after a certain amount of time has elapsed, such as a day, a week, a month, etc.) and / or may be based on various other conditions, such as network and / or server availability, and various other conditions.

[0164] Figure 12 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. Specifically, Figure 12 An example of computing system 1200 is illustrated. This computing system can be any computing device, such as constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1205. Connection 1205 can be a physical connection using a bus, or a direct connection to processor 1210, such as in a chipset architecture. Connection 1205 can also be a virtual connection, a networking connection, or a logical connection.

[0165] In some aspects, computing system 1200 is a distributed system in which the functions described herein can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a plurality of such components, each of which performs some or all of the functions described for that component. In some aspects, the components can be physical or virtual devices.

[0166] Example computing system 1200 includes at least one processing unit (CPU or processor) 1210 and a connection 1205 that couples various system components, including system memory 1215 (such as ROM 1220 and RAM 1225), to processor 1210. Computing system 1200 may include a cache 1212 of high-speed memory that is directly connected to, closely proximate to, or integrated into processor 1210.

[0167] Processor 1210 may include any general-purpose processor and hardware or software services (such as services 1232, 1234, and 1236 stored in storage device 1230 and configured to control processor 1210), as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1210 may be a substantially completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0168] To enable user interaction, the computing system 1200 includes an input device 1245 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 1200 may also include an output device 1235, which can be one or more of a plurality of output mechanisms. In some instances, a multi-mode system allows the user to provide multiple types of input / output to communicate with the computing system 1200. The computing system 1200 may include a communication interface 1240, which typically governs and manages user input and system output. The communication interface can perform or facilitate the receipt and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple... ® Lightning ® Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Bluetooth ® Wireless signal transmission, BLE wireless signal transmission, IBEACON ®Wireless signal transmission, RFID wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short Range Communication (DSRC) wireless signal transmission, 802.11 WiFi wireless signal transmission, WLAN signal transmission, Visible Light Communication (VLC), Microwave Access Global Interoperability (WiMAX), IR communication wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or combinations thereof. The communication interface 1240 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1200 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US GPS, Russia's GLONASS, China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no limitations on operation on any particular hardware configuration, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware configurations as they are developed.

[0169] Storage device 1230 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory card, solid-state storage device, digital versatile optical disc, cartridge, floppy disk, hard disk, magnetic tape, magnetic stripe / magnetic stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state storage, CD-ROM, rewritable CD, DVD, Blu-ray Disc, holographic disc, another optical medium, secure digital (SD) card, microSD card, Memory Stick. ®Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, RAM, static RAM (SRAM), dynamic RAM (DRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette and / or combinations thereof.

[0170] Storage device 1230 may include software services, servers, services, etc., which enable the system to perform functions when the code defining such software is executed by processor 1210. In some aspects, hardware services performing a particular function may include software components for performing that function stored in a computer-readable medium connected to necessary hardware components such as processor 1210, connection 1205, output device 1235, etc. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which does not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as CDs or DVDs), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0171] In some examples, the processes described herein (e.g., process 800 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 800 may be performed by a device having Figure 12 The computing architecture of the computing system 1200 shown is a computing device (e.g., Figure 2 The image capture and processing system 200 in the system is used to perform this operation.

[0172] In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, one or more network interfaces configured to transmit and / or receive data, any combination thereof, and / or other components. One or more network interfaces may be configured to transmit and / or receive wired and / or wireless data, including data according to 3G, 4G, 5G, and / or other cellular standards, data according to the Wi-Fi (802.11x) standard, and data according to Bluetooth. ™ Standard data, data according to IP standards, and / or other types of data.

[0173] Components that enable the implementation of a computing device in a circuit. For example, each component may include and / or may be implemented using electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits)), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0174] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0175] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0176] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0177] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0178] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0179] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0180] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that various inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including such variations unless limited by prior art. The various features and aspects of the applications described above may be used individually or in combination. Furthermore, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0181] Those skilled in the art will understand that, without departing from the scope of this description, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“>”) respectively. ") and greater than or equal to (" The symbol ) is used instead.

[0182] When a component is described as being “configured” to perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0183] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0184] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, the claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language that states "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the claim language that states "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0185] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.

[0186] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium can include memory or data storage media, such as RAM (e.g., Synchronous Dynamic Random Access Memory (SDRAM)), ROM, non-volatile random access memory (NVRAM), EEPROM, flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0187] The program code can be executed by a processor, which may include one or more processors, such as one or more DSPs, general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0188] The exemplary aspects of this disclosure include: Aspect 1. An apparatus for processing one or more images, the apparatus comprising: one or more memories configured to store the one or more images; and one or more processors coupled to the one or more memories and configured to: obtain first disparity information associated with a current image in the one or more images; remap the current image based on the first disparity information to obtain an estimated previous image; determine a confidence map associated with a confidence level of the first disparity information based on a difference associated with the estimated previous image; and apply the confidence map to the first disparity information to generate updated first disparity information.

[0189] Aspect 2. The apparatus according to aspect 1, wherein the first disparity information includes at least one of: first disparity information that estimates a first movement of a first feature to a first destination location in the current image; or depth information that represents the depth of the first feature.

[0190] Aspect 3. The apparatus according to any one of Aspects 1 to 2, wherein the one or more processors are configured to: determine the difference between a previous image and the estimated previous image.

[0191] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein the confidence map includes a first region corresponding to the first feature that is valid in the first disparity information.

[0192] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the confidence map includes a first region corresponding to the first feature that is a false alarm in the first disparity information.

[0193] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the one or more processors are configured to: remove the first disparity information to generate the updated first disparity information.

[0194] Aspect 7. The apparatus according to any one of Aspects 1 to 6, wherein the confidence map is determined based on a first threshold at a first time, and wherein the confidence map is determined based on a second threshold at a second time after the first time.

[0195] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the second threshold includes a higher confidence level than the first threshold.

[0196] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the one or more processors are configured to: determine the sparsity of a region associated with the first feature in the current image or a previous image; and determine, based on the sparsity, a threshold corresponding to the confidence level of the first feature in the first disparity information.

[0197] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the one or more processors are configured to: determine a first motion amount value associated with the first feature in the current image; determine a first threshold corresponding to the confidence level of the first feature within the first disparity information based on the first motion amount value; and determine, based on the first threshold, whether a region associated with the first feature in the confidence map corresponds to the true disparity information.

[0198] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the one or more processors are configured to: determine an attention associated with the first feature, wherein the attention corresponds to the importance of the first feature associated with at least one other feature in the first disparity information; and determine, based on the attention, a first threshold corresponding to authentication of the first disparity information of the first feature.

[0199] Aspect 12. The apparatus according to any one of aspects 1 to 11, wherein the attention includes information identifying the importance of the first feature in the previous image and the current image compared to other features in the previous image and the current image.

[0200] Aspect 13. The apparatus according to any one of aspects 1 to 12, wherein the one or more processors are configured to: obtain second disparity information associated with a previous image, the second disparity information estimating a second movement of the first feature within the current image or the previous image.

[0201] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein the one or more processors are configured to: determine, based on the first parallax information and the second parallax information, that the first feature is occluded in the current image or the previous image.

[0202] Aspect 15. The apparatus according to any one of Aspects 1 to 14, wherein the one or more processors are configured to: remap the previous image based on the second disparity information to obtain an estimated current image; generate a second confidence map associated with the second disparity information based on the difference associated with the estimated current image; and apply the second confidence map to the second disparity information to generate updated second disparity information.

[0203] Aspect 16. The apparatus according to any one of aspects 1 to 15, the apparatus further comprising one or more cameras configured to capture the one or more images.

[0204] Aspect 17. The apparatus according to any one of Aspects 1 to 16, wherein, in order to obtain the first disparity information associated with the current image, the one or more processors are configured to: use one or more machine learning systems to generate features representing the current image; and generate the first disparity information based on the features representing the current image.

[0205] Aspect 18. The apparatus according to aspect 17, wherein the one or more machine learning systems include at least one of a deep neural network (DNN) or a convolutional neural network (CNN).

[0206] Aspect 19. A method of processing one or more images by an image capture device, the method comprising: obtaining first disparity information associated with a current image, the first disparity information estimating a first movement of a first feature to a first destination location in the current image; remapping the current image based on the first optical flow information to obtain an estimated previous image; determining a confidence map associated with a confidence level of the first disparity information based on a difference associated with the estimated previous image; and applying the confidence map to the first disparity information to generate updated first disparity information.

[0207] Aspect 20. The method according to aspect 19, the method further comprising: determining the difference between the previous image and the estimated previous image.

[0208] Aspect 21. The method according to any one of Aspects 19 to 20, wherein the confidence map includes a first region corresponding to the first feature valid in the first disparity information.

[0209] Aspect 22. The method according to any one of aspects 19 to 21, wherein the confidence map includes a first region corresponding to the first feature that is a false alarm in the first disparity information.

[0210] Aspect 23. The method according to any one of Aspects 19 to 22, wherein applying the confidence map to the first disparity information comprises: removing the first disparity information to generate the updated first disparity information.

[0211] Aspect 24. The method according to any one of Aspects 19 to 23, wherein the confidence map is determined based on a first threshold at a first time, and wherein the confidence map is determined based on a second threshold at a second time after the first time.

[0212] Aspect 25. The method according to any one of Aspects 19 to 24, wherein the second threshold includes a higher confidence level than the first threshold.

[0213] Aspect 26. The method according to any one of Aspects 19 to 25, wherein generating the confidence map comprises: determining the sparsity of a region associated with the first feature in the current image or a previous image; and determining a threshold corresponding to the confidence of the first feature in the first disparity information based on the sparsity.

[0214] Aspect 27. The method according to any one of Aspects 19 to 26, wherein generating the confidence map comprises: determining a first motion value associated with the first feature in the current image; determining a first threshold corresponding to the confidence of the first feature within the first disparity information based on the first motion value; and determining, based on the first threshold, whether a region associated with the first feature in the confidence map corresponds to the true disparity information.

[0215] Aspect 28. The method according to any one of Aspects 19 to 27, wherein generating the confidence map comprises: determining an attention associated with the first feature, wherein the attention corresponds to the importance of the first feature associated with at least one other feature in the first disparity information; and determining a first threshold based on the attention corresponding to authentication of the first disparity information of the first feature.

[0216] Aspect 29. The method according to any one of Aspects 19 to 28, wherein the attention includes information identifying the importance of the first feature in the previous image and the current image compared to other features in the previous image and the current image.

[0217] Aspect 30. The method according to any one of aspects 19 to 29, the method further comprising: obtaining second disparity information associated with a previous image, the second disparity information estimating a second movement of the first feature within the current image or the previous image.

[0218] Aspect 31. The method according to any one of Aspects 19 to 30, the method further comprising: determining, based on the updated first disparity information and the second disparity information, that the first feature is occluded in the current image or the previous image.

[0219] Aspect 32. The method according to any one of Aspects 19 to 31, the method further comprising: remapping the previous image based on the second disparity information to obtain an estimated current image; generating a second confidence map associated with the second disparity information based on differences associated with the estimated current image; and applying the second confidence map to the second disparity information to generate updated second disparity information.

[0220] Aspect 33. The method according to any one of Aspects 19 to 32, wherein the first disparity information comprises at least one of: first optical flow information, the first optical flow information estimating a first movement of a first feature to a first destination location in the current image; or depth information representing the depth of the first feature.

[0221] Aspect 34. The apparatus according to any one of aspects 19 to 33, wherein, in order to obtain the first disparity information associated with the current image, the one or more processors are configured to: use one or more machine learning systems to generate features representing the current image; and generate the first disparity information based on the features representing the current image.

[0222] Aspect 35. The apparatus according to aspect 34, wherein the one or more machine learning systems include at least one of a deep neural network (DNN) or a convolutional neural network (CNN).

[0223] Aspect 36. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one of aspects 19 to 45.

[0224] Aspect 37. An apparatus for processing one or more images, the apparatus comprising one or more components for performing operations according to any one of aspects 19 to 45.

Claims

1. An apparatus for processing one or more images, the apparatus comprising: One or more memories, the one or more memories being configured to store the one or more images; and One or more processors, said one or more processors being coupled to said one or more memories and configured to: Obtain first disparity information associated with the current image in the one or more images; The current image is remapped based on the first disparity information to obtain an estimated previous image; A confidence map associated with the confidence of the first disparity information is determined based on the difference associated with the previous image estimated thereto; as well as The confidence map is applied to the first disparity information to generate updated first disparity information.

2. The apparatus of claim 1, wherein the first parallax information comprises at least one of: first optical flow information, the first optical flow information estimating a first movement of the first feature to a first destination location in the current image; or depth information representing the depth of the first feature.

3. The apparatus of claim 1, wherein the one or more processors are configured to: determine the difference between the previous image and the estimated previous image.

4. The apparatus of claim 1, wherein the confidence map includes a first region corresponding to a first feature valid in the first disparity information.

5. The apparatus of claim 1, wherein the confidence map includes a first region corresponding to a first feature that is a false alarm in the first disparity information.

6. The apparatus of claim 5, wherein the one or more processors are configured to: Remove the first disparity information to generate the updated first disparity information.

7. The apparatus of claim 1, wherein the confidence map is determined based on a first threshold at a first time, and wherein the confidence map is determined based on a second threshold at a second time after the first time.

8. The apparatus of claim 7, wherein the second threshold includes a higher confidence level than the first threshold.

9. The apparatus of claim 1, wherein the one or more processors are configured to: Determine the sparsity of the region associated with a first feature in the current or previous image; and A threshold corresponding to the confidence level of the first feature in the first disparity information is determined based on the sparsity.

10. The apparatus of claim 1, wherein the one or more processors are configured to: Determine a first amount of movement associated with a first feature in the current image; Based on the first movement value, a first threshold corresponding to the confidence level of the first feature within the first disparity information is determined; and Based on the first threshold, it is determined whether the region associated with the first feature in the confidence map corresponds to the true disparity information.

11. The apparatus of claim 1, wherein the one or more processors are configured to: Determine the attention associated with a first feature, wherein the attention corresponds to the importance of the first feature associated with at least one other feature in the first disparity information; and A first threshold corresponding to the authentication of the first disparity information of the first feature is determined based on the attention.

12. The apparatus of claim 11, wherein the attention includes information identifying the importance of the first feature in the previous image and the current image compared to other features in the previous image and the current image.

13. The apparatus of claim 1, wherein the one or more processors are configured to: A second disparity information associated with a previous image is obtained, the second disparity information being used to estimate a second movement of a first feature within the current image or the previous image.

14. The apparatus of claim 13, wherein the one or more processors are configured to: Based on the first disparity information and the second disparity information, it is determined that the first feature is occluded in the current image or the previous image.

15. The apparatus of claim 13, wherein the one or more processors are configured to: The previous image is remapped based on the second disparity information to obtain an estimated current image; A second confidence map associated with the second disparity information is generated based on the difference associated with the estimated current image; and The second confidence map is applied to the second disparity information to generate updated second disparity information.

16. The apparatus of claim 1, further comprising one or more cameras configured to capture the one or more images.

17. The apparatus according to claim 1, wherein, In order to obtain the first disparity information associated with the current image, the one or more processors are configured to: Use one or more machine learning systems to generate features representing the current image; and The first disparity information is generated based on the features representing the current image.

18. The apparatus of claim 17, wherein the one or more machine learning systems comprise at least one of a deep neural network (DNN) or a convolutional neural network (CNN).

19. A method for processing one or more images by an image capture device, the method comprising: Obtain the first disparity information associated with the current image; The current image is remapped based on the first disparity information to obtain an estimated previous image; A confidence map associated with the confidence of the first disparity information is determined based on the difference associated with the previous image estimated thereto; as well as The confidence map is applied to the first disparity information to generate updated first disparity information.

20. The method of claim 19, wherein the first disparity information comprises at least one of: first optical flow information, the first optical flow information estimating a first movement of the first feature to a first destination location in the current image; or depth information representing the depth of the first feature.