Monitoring motion and lighting to implement modified stereo vision processing

A device with multiple cameras switches to modified stereo vision processing based on motion and lighting thresholds, improving accuracy and efficiency by using asymmetric downsampling operations.

WO2026060008A1PCT designated stage Publication Date: 2026-03-19QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing devices fail to select the optimal pair of cameras for stereo image capture based on motion and lighting conditions, leading to inefficient stereo vision processing.

Method used

A device with multiple cameras that monitors contextual data changes, switching to modified stereo vision processing when motion or lighting exceeds a threshold, utilizing asymmetric downsampling operations for improved accuracy and efficiency.

Benefits of technology

Enhances stereo vision processing accuracy and efficiency by selecting the most accurate stereo vision based on motion and lighting conditions, reducing computational complexity and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045767_19032026_PF_FP_ABST
    Figure US2025045767_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Aspects relate to monitoring motion and lighting to implement modified stereo vision processing. A device may include one or more memories configured to store one or more images. The device may further include one or more processors coupled to the one or memories, in which, the one or more processors are configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.
Need to check novelty before this filing date? Find Prior Art

Description

Qualcomm Ref. No. 2407692WO 1MONITORING MOTION AND LIGHTING TO IMPLEMENT MODIFIED STEREO VISION PROCESSINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present Application for Patent claims priority to pending Non-Provisional Application Serial No. 19 / 243,475 filed in the United States Patent and Trademark Office on June 19, 2025 and U.S. provisional patent application no. 63 / 693,596 filed on September 11, 2024, and assigned to the assignee hereof and hereby expressly incorporated by reference herein as if fully set forth below in their entireties and for all applicable purposes.TECHNICAL FIELD

[0002] The technology discussed below relates generally to monitoring motion and lighting for a device, and more particularly, to monitoring motion and lighting for a device to implement modified stereo vision processing.INTRODUCTION

[0003] Stereo vision may be defined as the ability to perceive depth and spatial information by using two images of the same scene from slightly different perspectives. It is based on the idea that humans have two eyes that see the world from slightly different positions, and the brain combines these views to create a three-dimensional sensation. Stereo video or pictures may be achieved using two views, e.g., a left view and a right view. In order to simulate a human vision system, which has depth perception, a device with two camera sensors may capture left eye and right eye views. In stereo vision, there is disparity in the distance between corresponding points in the two images taken from the slightly different positions of the two camera sensors having left and right views. A stereo image may be created by a device by combing the two images from the left and right camera sensors.

[0004] In many devices, various cameras are physically fixed at different locations in or on the device. Oftentimes, when a device captures an image with a pair of cameras, the device does not select the pair of cameras that provide the best stereo image based upon motion and / or lighting and does not implement modified and efficient stereo vision processing.Qualcomm Ref. No. 2407692WO 2BRIEF SUMMARY OF SOME EXAMPLES

[0005] The following presents a summary of one or more aspects of the present disclosure, in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated features of the disclosure and is intended neither to identify key or critical elements of all aspects of the disclosure nor to delineate the scope of any or all aspects of the disclosure. Its sole purpose is to present some concepts of one or more aspects of the disclosure in a form as a prelude to the more detailed description that is presented later.

[0006] In one example, a device is provided. The device may include one or more memories configured to store one or more images. The device may further include one or more processors coupled to the one or memories, in which, the one or more processors are configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

[0007] Another example is a method. The method includes: obtaining contextual data; determining if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

[0008] In yet another example, a non-transitory computer-readable data storage medium is provided that has stored thereon instructions that, when executed, cause one or more processors to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switching from stereo vision processing to a modified stereo vision processing.

[0009] These and other aspects will become more fully understood upon a review of the detailed description, which follows. Other aspects, features, and examples will become apparent to those of ordinary skill in the art, upon reviewing the following description of examples in conjunction with the accompanying figures. While features may be discussed relative to certain examples and figures below, all examples can include one or more of the advantageous features discussed herein. In other words, while one or more examples may be discussed as having certain advantageous features, one or more of such features may also be used in accordance with the various examples discussed herein. In similar fashion, while exemplary examples may be discussed below as device, system, or method examples such exemplary examples can be implemented in various devices.Qualcomm Ref. No. 2407692WO 3BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIG. 1 is a diagram illustrating an example of a device according to some aspects.

[0011] FIG. 2 is a flowchart related to contextual data according to some aspects.

[0012] FIG. 3 is a simplified diagram of FIG. 1 according to some aspects.

[0013] FIG. 4 is a diagram illustrating an example of cameras used in performing camera switching operations according to some aspects.

[0014] FIG. 5 is a diagram illustrating the use of a rectified epipolar line better aligned with dominant directions of local / global motion(s) used in performing camera switching operations according to some aspects.

[0015] FIG. 6 is a diagram illustrating different arrangements of multi-camera sets to support resolution enhancement for direction and speed according to some aspects.

[0016] FIG. 7 is a flowchart related to contextual data according to some aspects.

[0017] FIG. 8 is a diagram illustrating an example of a world point P(X,Y,Z), a left image plane of the left camera, and a right image plane of the right camera according to some aspects.

[0018] FIG. 9 is a diagram illustrating an example of disparities between pixels on the left and right image planes according to some aspects.

[0019] FIG. 10 is a flowchart illustrating an example operation for down-sampling according to some aspects.

[0020] FIG. 11 is a diagram illustrating an example operation for down-sampling images according to some aspects.

[0021] FIG. 12A is a diagram illustrating down-sampling utilizing multiple- aspect ratios according to some aspects.

[0022] FIG. 12B is a diagram illustrating down-sampling and up-sampling utilizing multiple-aspect ratios according to some aspects.

[0023] FIG. 13 is a diagram illustrating asymmetric S2D operations for down-sampling and asymmetric D2S for up- sampling according to some aspects.

[0024] FIG. 14 illustrates a proof-of-concept of the techniques of the disclosure related to encoding and decoding for stereo depth estimation utilizing asymmetric operations according to some aspects.DETAILED DESCRIPTIONQualcomm Ref. No. 2407692WO 4

[0025] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.

[0026] Aspects of the disclosure to be described relate to monitoring motion and lighting to implement modified stereo vision processing. The device may include one or more memories configured to store one or more images. The device may further include one or more processors coupled to the one or memories, in which, the one or more processors are configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing. In one example aspect, the device may further include a plurality of cameras configured to capture the plurality of images, in which the contextual data is based on the images. Further, as will be described, in some aspects, the contextual data is related to motion or lighting changes.

[0027] It should be appreciated that contextual data provides significant information about a particular scene and image - such as objects in an image, their arrangement, relative physical size to other objects, and location. Motion or lighting changes include significant contextual data that may be monitored according to aspects of the disclosure.

[0028] In one aspect, to be described in more detail hereafter, the processor of a device may be configured to: obtain motion and / or lighting changes (e.g., contextual data) from a sensor (e.g., a plurality of cameras) located on the device; determine if motion and / or lighting change exceeds a threshold; and when the motion and / or lighting change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

[0029] In one aspect, by having multiple cameras on a device and by providing multiple choices of stereo cameras, a device may select the most accurate stereo vision of an image or an object in an image based on motion and / or lighting changes. As will be described, various triggering and switching operations between cameras in response to observations and / or variations related to motion and / or lighting for scenes, images, or objects may be implemented. As an example, motion-related and / or lighting-related conditioning, triggering, and / or switching that affects stereo vision processing in order for betterQualcomm Ref. No. 2407692WO 5 accuracy or efficiency, or both, will be described. As will be described, a modified stereo vision processing implementation may be selected that utilizes an asymmetric downsampling operation for the height and width of the object. For example, the asymmetric down-sampling operation may include a higher resolution in width. A device utilizing stereo camera switching based upon motion and / or lighting changes and that implements a modified stereo vision processing operation that includes an asymmetric downsampling operation that includes a higher resolution in width, and that enables higher accuracy, consistency, diversity, and / or efficiency, will be described.

[0030] Fig. 1 illustrates a device 130 with multiple digital sensors (first, second ... N) 132, 134, 162 configured to capture and process 3-D stereo images and videos. It should be appreciated that digital sensors 132, 134, 162 may be camera sensors but that other sorts of sensors may be utilized. Also, device 130 may be a mobile device but also may be a fixed device or another sort of device. In general, device 130 may be configured to capture, create, process, modify, scale, encode, decode, transmit, store, and display digital images and / or video sequences. Device 130 may provide high-quality stereo image capturing, various sensor locations, view angle mismatch compensation, and an efficient solution to process and combine a stereo image.

[0031] In one aspect, device 130 may be: a mobile device, a mobile phone, a vehicle, a robot, a stationary Internet of Things (loT) device, a mobile loT device, or a security device. However, these devices are just examples and it should be appreciated that device 130 may be any suitable device.

[0032] Additionally device 130 may represent or be implemented in a wireless communication device, a personal digital assistant (PDA), a handheld device, a laptop computer, a desktop computer, a digital camera, a digital recording device, a network- enabled digital television, a mobile phone, a cellular phone, a satellite telephone, a camera phone, a terrestrial-based radiotelephone, a direct two-way communication device (sometimes referred to as a “walkie-talkie”), a camcorder, etc.

[0033] Device 130 may include a first camera sensor 132, a second camera sensor 134, a N-camera sensor 162, a first camera interface 136, a second camera interface 148, a N- camera interface 168, a first buffer 138, a second buffer 150, a N-buffer 170, a memory 146, a diversity combine module 140 (or engine), a camera process pipeline 142, a second memory 154, a diversity combine controller for 3-D image 152, a mobile display processor (MDP) 144, a processor 156, a user interface 120, a display device 122, a motion sensor 125, an image sensor 126, an audio device 127, and a transceiver or modemQualcomm Ref. No. 2407692WO 6129. It should be appreciated that motion sensor 125, image sensor 126, audio device 127, and transceiver 129 may also be coupled to processor 156. In addition to or instead of the components shown in Fig. 1, device 130 may include other components. The architecture in Fig. 1 is merely an example. The features and techniques described herein may be implemented with a variety of other architectures.

[0034] As will be described, device 130 may utilize processor 156 to interact with a plurality of different cameras (N-cameras) (e.g., camera 1 132, camera 2, camera N 162), in which, processor 156 may determine a subset of cameras to provide the best stereo vision of an image, scene, or object. In one example aspect, processor 156 may be configured to implement operations including: obtaining contextual data from a sensor (e.g., camera 1 132, camera 2 134, or camera N 162)); determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing. It should be appreciated that device 130 may include any number (N) of cameras.

[0035] The sensors 132, 134, 162 (N-sensors) may be digital camera sensors. The sensors 132, 134, 162 may have similar or different physical structures. The sensors 132, 134, 162 may have similar or different configured settings. The sensors 132, 134, 164 may capture still image snapshots and / or video sequences. Each sensor may include color filter arrays (CFAs) arranged on a surface of individual sensors or sensor elements.

[0036] The memories 146, 154 may be separate or integrated. The memories 146, 154 may store images or video sequences before and after processing. The memories 146, 154 may include volatile storage and / or non-volatile storage. The memories 146, 154 may comprise any type of data storage means, such as dynamic random access memory (DRAM), FLASH memory, NOR or NAND gate memory, or any other data storage technology.

[0037] The camera process pipeline 142 (also called engine, module, processing unit, video front end (VFE), etc.) may comprise a chip set for a mobile phone, which may include hardware, software, firmware, and / or one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or various combinations thereof. The pipeline 142 may perform one or more image processing techniques to improve quality of an image and / or video sequence.

[0038] Processor 156 may include one or more processors and may implement down- sampling / encoding functions and / or up-sampling / decoding functions. Processor 156 mayQualcomm Ref. No. 2407692WO 7 also implement other functions of device 130. Processor 156 may operate as a video encoder and may implement or comprise an encoder / decoder (CODEC) for encoding (or down-sample or compress, etc.) and decoding (or up-sample or decompress) digital video data. As an example, the processor operating to implement video encoder function may use one or more encoding / decoding standards or formats, such as MPEG or H.264. In other examples, separate video encoder and / or video decoder devices may be utilized.

[0039] The transceiver or modem 129 may receive and / or transmit coded images or video sequences to another device or a network. The transceiver or modem 129 may use a wireless communication standard, such as code division multiple access (CDMA). Examples of CDMA standards include CDMA lx Evolution Data Optimized (EV-DO) (3GPP2), Wideband CDMA (WCDMA) (3GPP), etc. In other examples, transceiver or modem 129 may utilize other cellular communication standards, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, 6G, or the like. In some examples, other wireless standards, such as IEEE 802.11 specification, IEEE 802.15 specification (e.g., ZigBee™), Bluetooth™ standard, or the like, may be utilized.

[0040] Device 130 may maintain a fixed horizontal distance between the sensors 132, 134, 162 such that 3-D stereo image and video can be generated efficiently. As shown in Fig. 1, the N-sensors 132, 134, 162 may be separated by a suitable fixed horizontal distance. The first sensor 132 may be a primary sensor, and the second sensor 134 and N-sensor 162 may be secondary sensors. The secondary sensors may be shut off for nonstereo mode to reduce power consumption. However, this is an optional sensor set-up.

[0041] The buffers 138, 150, 170 may store real time sensor input data, such as one row or line of pixel data from the sensors 132, 134, 162. Sensor pixel data may enter the small buffers 138, 150, 170 on-line (i.e., in real time) and be processed by the diversity combine module 140 and / or camera engine pipeline engine 142 offline with switching between the sensors 132, 134, 160 (or buffers 138, 150, 170) back and forth. The diversity combine module 140 and / or camera engine pipeline engine 142 may operate at about two times the speed of one sensor’s data rate. To reduce output data bandwidth and memory requirement, stereo image and video may be composed in the camera engine 142.

[0042] The diversity combine module 140 may first select data from the first buffer 138. At the end of one row of buffer 138, the diversity combine module 140 may switch to the second buffer 150 to obtain data from the second sensor 134 or likewise to the N-buffer 170 to obtain data from the N-sensor 162. The diversity combine module 140 may switchQualcomm Ref. No. 2407692WO 8 back to the first buffer 138 at the end of one row of data from the second buffer 150 or N-buffer 170.

[0043] In order to reduce processing power and data traffic bandwidth, the sensor image data in video mode may be sent directly through the buffers 138, 150, 170 (bypassing the first memory 146) to the diversity combine module 140. On the other hand, for a snapshot (image) processing mode, the sensor data may be saved in the memory 146 for offline processing. In addition, for low power consumption profiles, the second sensor 134 or N-sensor 162 may be turned off, and the camera pipeline driven clock may be reduced.

[0044] Aspects of the disclosure relate to monitoring motion and lighting to implement modified stereo vision processing. As previously described, device 130 may include one or more memories 146, 154 configured to store one or more images. In one embodiment, processor 156 coupled to the or more memories may be configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing. In one example aspect, device 130 may further include a plurality of cameras (first camera 132, second camera 134, and Nth camera 162), which act as sensors, configured to capture the plurality of images, in which, the contextual data is based on the images. Further, as will be described, in some aspects, the contextual data is related to motion or lighting changes.

[0045] With brief additional reference to FIG. 2, FIG. 2 is a flowchart related to contextual data according to some aspects. At block 202, contextual data (e.g., motion or lighting data) is obtained. At block 204, processor 156 determines if the contextual data change exceeds a threshold (e.g., as will be discussed in more detail hereafter). At block 206, when the contextual data change exceeds the threshold, processor 156 may command a switch from stereo vision processing to modified stereo vision processing. If the threshold is not exceeded regular stereo vision processing can proceed. As has been described, in one example aspect, device 130 may include a plurality of cameras (first camera 132, second camera 134, and Nth camera 162), which act as sensors, configured to capture the plurality of images, in which, the contextual data is based on the images. Further, as will be described, in some aspects, the contextual data is related to motion or lighting changes.

[0046] In one example aspect, contextual data obtained from a sensor (e.g., a camera) may be based upon contextual data change in a scene, image, an object in an image (e.g., a new object detected by cameras 132, 134, 162, etc.), etc. It should be appreciated thatQualcomm Ref. No. 2407692WO 9 contextual data provides significant information about a particular scene - such as images, objects in an image, their arrangement, relative physical size to other objects, and location. Motion or lighting changes is significant contextual data that may be monitored according to aspects of the disclosure.

[0047] In one aspect, processor 156 may be configured to: obtain motion and / or lighting changes (e.g., contextual data) from the plurality of cameras 132, 134, and 164; determine if motion and / or lighting change exceeds a threshold; and when the motion and / or lighting change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing. In one aspect, by utilizing multiple cameras 132, 134, and 164, and, by providing multiple choices of stereo cameras, processor 156 of device may select the most accurate stereo vision of an object based on motion and / or lighting changes. As will be described, various triggering and switching operations between cameras 132, 134, and 164 in response to observations and / or variations related to motion and / or lighting for scenes, images, or objects in images may be implemented. As an example, motion-related and / or lighting-related conditioning, triggering, and / or switching that affects the stereo vision processing in order for better accuracy or efficiency, or both, will be described.

[0048] As one example implementation, the plurality of cameras (e.g., cameras 132, 134, 162) may obtain contextual data including obtaining first contextual data at a first time and second contextual data at a second time. Processor 156 may determine if the contextual data change exceeds the threshold by comparing the image of the first contextual data at the first time and the image of the second data at the second time, and determine if the threshold is exceeded.

[0049] In one aspect, the contextual data changes exceeding the threshold may be based upon a motion change of first and second images. The contextual data change may be a motion change based upon speed and / or direction. As will be described, when the motion change is based upon a speed or direction, that exceeds a threshold, processor 156 may be configured to: determine a subset of cameras (e.g., 132, 134, 162) to account for the motion change; and based upon the determined subset of cameras (e.g., 132, 134, 162), utilize asymmetric down-sampling operations for height and width of the object to provide modified stereo vision. As will be described, a modified stereo vision processing implementation may be selected that utilizes an asymmetric down-sampling operation for the height and width of the object. For example, the asymmetric down-sampling operation may include a higher resolution in width. Therefore, as will be described, a device 130 utilizing stereo camera switching based upon motion and / or lighting changesQualcomm Ref. No. 2407692WO 10 and that implements a modified stereo vision processing operation that includes an asymmetric down- sampling operation that includes a higher resolution in width, and that enables higher accuracy, consistency, diversity, and / or efficiency, is disclosed herein.

[0050] Therefore, disclosed herein, is a motion-related solution for stereo vision processing (SVP) that is conditioned on the motions of scenes, images, and / or objects in images to trigger or switch cameras for better accuracy and / or efficiency. In particular, looking at particular aspects, the motion may include amplitude (e.g., speed) of the motion and / or direction (e.g., the orientation in a 2D / 3D coordinate space) of the motion. In some aspects, speed of motion may be referred to as being “faster” or “slower” motions. Moreover, the motion may apply to one or multiple objects in an image or scene or to the entire image or scene. In the former case, convenient reference to the local motion of one or multiple objects may be made. In the latter case, convenient reference may be to global motion, as such type of the motion applies to the global scope, such as due to the relative movement of the camera.

[0051] Typically, in the standard case of slow local / global motions, such as in an indoor scene with static objects, standard stereo vision processing (SVP) may be utilized. In the case of fast local motion for at least one object or fast global motion for the scene, however, it is beneficial to detect the occurrence of such scenario and apply conditions to trigger / switch among cameras (e.g., 132, 134, 162) for non-standard SVP, e.g., a modified stereo vision processing implementation that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, the asymmetric down-sampling operation may include a higher resolution in width. In order to detect a potential scenario of fast local / global motion, optical flow (2D) or scene flow (3D) estimation may be applied to detect the potential existence of any large 2D / 3D object / scene displacements. Optical flow may be referred to as 2D motion field that describes the apparent movement of pixels in an image, essentially showing how pixels seem to move between consecutive frames of a video, while scene flow is the 3D equivalent, representing the full 3D motion of points in a scene, essentially providing information about how points in the real world are moving relative to the camera. In order to reduce false positive (or false alarms) under noisy conditions, one or multiple fixed or dynamic thresholds may be applied such that only when the displacement quantity or the number of object / scene pixels becomes larger than the threshold(s) then it is conditioned as fast motion. It should be noted that in fast motion scenarios under a fixed frame rate (e.g., 30 FPS), there is unavoidable motionQualcomm Ref. No. 2407692WO 11 blurs given the duration in time for the image sensing. In such cases, the ability to better discriminative features in the motion directions would be desirable.

[0052] According to aspects of the disclosure, when the condition of fast local / global motion is determined by its amplitude (e.g., speed of motion between scenes, objects, images, etc.), further consideration of the direction(s) of the motion for further conditioning can be implemented. In general, it has been found that in prior art implementations, that standard stereo vision processing (SVP) can face a larger challenge, such as in accuracy or in computation complexity for depth estimation, under a fast motion condition.

[0053] With brief reference to FIG. 3, FIG. 3 is a simplified diagram of FIG. 1. Like FIG. 1, FIG. 3 illustrates device 130 including: processor 156, memories 146,154, and cameras 1-N (132, 134, 162), user interface 120, display 122, motion sensor 125, audio device 127, and transceiver or modem 129. FIG. 3 is a simplified diagram for ease of reference to aid in illustrating particular implementations described below. As previously described, device 130 may be: a mobile device, a mobile phone, a vehicle, a robot, a stationary Internet of Things (loT) device, a mobile loT device, or a security device. However, these devices are just examples and it should be appreciated that device 130 may be any suitable device.

[0054] Also, in one example aspect, in addition to utilizing optical flow (2D) and / or scene flow (3D) estimation to detect fast local / global motion based on the amplitude / speed of motion (and direction) of images, objects in images, and scenes, that exceed the threshold, to trigger switch among cameras (e.g., cameras 132, 134, 162), and for modified SVP (e.g., asymmetric down-sampling operations for height and width), input from motion sensor 125 may also be utilized. In one embodiment, input from motion sensor 125 may be utilized to determine if the contextual change (e.g., the difference between two images) exceeds the threshold. A motion sensor is a sensor device that measures the motion of a device. For example, a suitable motion sensor may be an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit (IMU), etc. The motion sensor 125 may measure the movement and acceleration of device 130 and therefore, the camera positions, along the x, y, and z-axis, that can be utilized to aid in measuring the amplitude / speed of motion and direction of images.

[0055] According to aspects of the disclosure, when the condition of fast local / global motion is determined by its amplitude, further consideration of the directions of the motion for further conditioning can be implemented. In general, it has been found that inQualcomm Ref. No. 2407692WO 12 prior art implementations, that standard stereo vision processing (SVP) can face a larger challenge, such as in accuracy or in computation complexity for depth estimation, under a fast motion condition.

[0056] With additional reference to FIG. 4, as will be described, based upon estimated local / global motions of the image, a subset of stereo cameras may be switched to from among the available sets of cameras. As examples of a device 130, with plurality of cameras, a variety of camera pairs can be switched to, e.g., BA, AC or AC, AB in either format 402, 406). These cameras being associated with cameras 132, 134, 164, etc., of device 130.

[0057] In one particular aspect example, with additional reference to FIG. 5, in a first local (e.g., fast local) and / or global motion detection in speed or direction that exceeds one or multiple corresponding thresholds, processor 156 may switch to stereo cameras (e.g., BA, AC or AC, AB in either format 402, 406) from the available candidate sets such that a rectified epipolar line 502 is better aligned with the dominant directions of local / global motion(s) of the object. As one aspect example, because there are limited candidate stereo cameras, such as, between vertical and horizontal candidates (e.g., BA, AC or AC, AB in either format 402, 406), the dominant directions of local / global motions can be directed to those directions of camera candidates and those candidate stereo cameras can be selected having the largest projected motion vector (in terms of motion amplitudes & directions).

[0058] In this way, by utilizing these techniques, selected stereo cameras are now better aligned with the dominant directions of local / global motions, so that a stereo depth estimation algorithm may be better utilized with the rectified features along the better aligned direction of motions for feature matching and cost derivation for a disparity hypotheses. In some cases, if different directions of fast local motions are present in different patches of the images / video, patchification of the images may be applied by partitioning the images into patches in the same way for both left and right images such that each patch matches better to one direction of candidates of stereo cameras In other words, each patch can have its best aligned stereo camera.

[0059] In one aspect example, when the motion change occurs relative to different patches of an image, processor 156 of device 130 may be configured to: determine a subset of cameras to account for the motion change relative to each patch; and switch to the determined subset of cameras (e.g., camera sets from 402 and / or 406) for each patch to account for the motion change to provide modified stereo vision.Qualcomm Ref. No. 2407692WO 13

[0060] It should be appreciated that a wide variety of camera set ups can be used to support multi-camera sets to support general resolution enhancement or patch-wise rectification resolution enhancement for direction and speed to be used in stereo depth estimation.

[0061] With additional reference to FIG. 6, different arrangements of multi-camera sets to support general resolution enhancement or patch-wise rectification resolution enhancement for direction and speed to be used in stereo depth estimation will be described. To begin with, in this example aspect, the full set of multi-switch camera stereo cameras may be M=4 (A, B, C, D) 602. As an example, device 130 under control of processor 156 may have multiple cameras N (e.g., N=4). It should be appreciated that these subsets of cameras may be considered to be positioned in a first axis along the device. In one example, the subset of cameras may include a horizontal candidate set of stereo cameras (e.g., C, B) 604. In one example, the first axis may be considered to be a horizontal axis aligned with an image displayed on a display of the device and in this example the subset of cameras may include a horizontal candidate set of stereo cameras (e.g., C, B) 604. In another example, the subset of cameras may include a vertical candidate set of stereo cameras (e.g., A, D) 606. In one example, the first axis may be considered to be a vertical axis aligned with an image displayed on a display of the device and in this example the subset of cameras may include a vertical candidate set of stereo cameras (e.g., A, D) 606. In a further example, the subset of cameras may include a slant candidate set of stereo cameras (e.g., A, B) 608. In this example, the first axis may be considered to be an off-axis with an image displayed on a display of the device with the subset of cameras including a slant candidate set of stereo cameras (e.g., A, B) 608. In an even further example, the subset of cameras may include another slant candidate set of stereo cameras (e.g., A, C) 610. In this example, the first axis may be considered to be an off-axis with an image displayed on a display of the device with the subset of cameras including a slant candidate set of stereo cameras (e.g., A, C) 610.

[0062] As to determining if a threshold as to motion is exceeded, the motion threshold may be defined as: 1) a non-directional motion quantity threshold or 2) a directional motion quantity threshold. In either example, a threshold may be derived statically (i.e., irrespective of the scenes or dynamics during inference) or dynamically (depending on scenes or dynamics in inference). More particularly, the threshold may be derived based on the range or distributions of motion amplitudes as estimated with a predefined motion estimation function. The threshold may be derived further depending on certain statisticsQualcomm Ref. No. 2407692WO 14 of the distribution of motions, such as mean, min, max, percentile, etc. The threshold may be derived through a heuristic function / logic, or through a learning-based function, or through other functions. In the example of a non-directional motion, the root- meansquare of all the 2D (x, y) or 3D (x, y, z) motion vector components may be utilized, depending on whether the task is to solve a 2D-projected feature SP. As previously described, the threshold may depend on a non-directional (e.g., RMS of the 2D / 3D motion vector) or a directional threshold (e.g., a single-directional x or y in 2D or a singledirectional x or y or z in 3D). In some cases, the motion threshold may be a global motion threshold that applies to all pixels / regions in the image(s), or a local / regional motion threshold that applies to only a local subset or region of the pixels. In some cases, the motion threshold may be applied to only a subset of pixels such as keypoints in the image, such as the eyes in a human face or the vertices of the bounding box of a car. As previously described, a variety of different subsets of cameras (e.g., FIGs. 4 and 6) may be selected based on the motion change.

[0063] Unlike conditions due to motions, conditions due to lighting do not typically involve directions in nature. However, lighting conditions play a significant role in the accuracy of stereo depth estimation. For example, in extreme lighting conditions (e.g., low light or dark color of the object / scene, or over-exposure such outdoor scenes (e.g., in / to wards sunlight), pixels or regions of severe lighting conditions may face higher challenges in stereo depth estimation. It may be desirable to properly mitigate such issues, such as due to day-time or night-time safety in auto driving, pictures and video from smart phones, surveillance cameras, etc.

[0064] As will be described, under lighting conditions that become relatively extreme, aspects of the disclosure describe techniques to trigger resolution enhancement such that the resolution of an image and / or video may be preserved or enhanced in the direction of rectification. In some aspects, enhanced stereo vision can be enabled or triggered or switched (ON) due to lighting conditions, in which, modified stereo vision may include asymmetric down-sampling operations for height and width, in which, the asymmetric down-sampling operations may include a higher resolution in width.

[0065] With brief additional reference to FIG. 7, FIG. 7 is a flowchart related to contextual data according to some aspects. At block 702, lighting data is obtained from a sensor. For example, the contextual data change may be a lighting condition change as measured by a sensor. In one example, the sensor may be an image sensor 126. An image sensor or imager is a sensor that detects and conveys information used to form an image.Qualcomm Ref. No. 2407692WO 15It does so by converting the variable attenuation of light waves (as they pass through or reflect off objects) into signals, small bursts of current that convey the information. As an example, an image sensor may be a semiconductor that converts light into electrical signals to create images or videos. The image sensor 126 may be a separate sensor of device 130 or may be part of one or more of the cameras (132, 134, 162, etc.). It is then determined whether a lighting condition change (e.g., measured by image sensor 126) exceeds a threshold by processor 156 (block 704), and when the lighting condition change exceeds the threshold, processor 156 may be configured to switch from stereo vision processing to modified stereo vision processing (block 705). If the threshold is not exceeded, regular stereo vision processing can proceed. In some aspects, modified stereo vision processing may provide enhanced resolution of the object in a direction of rectification. In one aspect example, the enhanced resolution of the object in the direction of rectification may include processor 156 utilizing asymmetric down-sampling operations for height and width of the object to provide modified stereo vision, in which, the asymmetric down-sampling operations may include a higher resolution in width.

[0066] As has been described, various lighting condition changes may be analyzed in view of thresholds. In one example aspect, the lighting condition change is determined to exceed the threshold when decreased lighting affects an object, and in another example, the lighting condition change is determined to exceed the threshold when increased sunlight affects the object. As has been described, under lighting conditions that become relatively extreme, aspects of the disclosure describe techniques to trigger resolution enhancement such that the resolution of an image and / or video may be preserved or enhanced in the direction of rectification. Additional descriptions of lighting thresholds will be hereafter described. As one example, the lighting threshold may be defined for the all-channel (e.g., 3-channel of RGB or YUV) or single-channel (e.g., G channel out of the RGB representation or Y out of the YUV representation). Furthermore, the lighting threshold may be further derived as a static / fixed- valued threshold or a dynamic threshold depending on the scenarios or task dynamics experienced in inference. Additionally, the threshold may be derived as a heuristic function or a learning-based function that may directly or indirectly depend on the range or statistics of the lighting metrics / conditions, such as the mean, min, max, percentile, density of global or local spatial regions, etc.

[0067] In one aspect example, when the lighting condition change exceeds the threshold for a first patch of an image, but not a second patch of the image, processor 156 may switch from stereo vision processing to modified stereo vision processing for the firstQualcomm Ref. No. 2407692WO 16 patch of the image but not the second patch of the image. Therefore, in some cases, severe lighting conditions may affect only a portion of an image or video. As described, an optional patch-wise resolution enhancement to enable or trigger or switch on an asymmetric stereo vision protocol (SVP) that favors higher / finer resolution in the direction of rectification can be applied only to needed patches in the image / video.

[0068] As has been described, asymmetric down-sampling operations for height and width of the object to provide modified stereo vision, in which, the asymmetric downsampling operations may include a higher resolution in width. The techniques to trigger or switch on asymmetric stereo vision protocol (SVP) (e.g., modified stereo vision protocol) can enable higher stereo depth accuracy only as necessary in terms of only the needed dimension and needed regions or patches to avoid waste in computation in unnecessary dimensions or unnecessary regions or patches.

[0069] As has been described, by having multiple cameras on a device and by providing multiple choices of stereo cameras, a device may select the most accurate stereo vision of an object based on motion and / or lighting changes. Various triggering and switching operations between cameras in response to observations and / or variations related to motion and / or lighting for scenes or objects have been described. As an example, motion- related and / or lighting-related conditioning, triggering, and / or switching that affects the stereo vision processing in order for better accuracy or efficiency, or both, have been described. The modified stereo vision processing implementation may be selected that utilizes an asymmetric down-sampling operation for the height and width of the object. For example, the asymmetric down-sampling operation may include a higher resolution in width. A device utilizing stereo camera switching based upon motion and / or lighting changes and that implements the modified stereo vision processing operation previously described enables higher accuracy, consistency, diversity, and / or efficiency, will be described.

[0070] The modified stereo vision processing implementation that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, a higher resolution is in width, will now be described in greater detail. Aspects of the disclosure relate to a description of modified stereo vision processing implementation that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, a higher resolution is in width. Aspects of the disclosure relate to a device or system that provides multi- aspect-ratio implementation in down-sampling for stereo disparity estimation. For example, multi- aspect-ratio encoding for stereo depth is presented thatQualcomm Ref. No. 2407692WO 17 provides for disparity preservation and width-centric processing for disparity handling. Further, as will be described, asymmetric space-to-depth encoding and depth-to- space decoding is provided for disparity estimation. For example, disparate-rate height-width space-to-depth encoding and disparate-rate height- width depth-to-space encoding will be described.

[0071] Aspects of the disclosure generally relate to down- sampling, and more particularly, to down-sampling in different directions. As will be described, aspects of the disclosure relate to down-sampling for stereo depth estimation utilizing asymmetric operations in different directions (e.g., in width and height). As shown in FIG. 1, device 130 may include one or more memories 146, 154 that are configured to store a plurality of images from cameras 132, 134. Cameras 132,134 may be configured to capture a left and right image, respectively, in which, each of the images includes one or more patches, each patch including plurality of pixels. Device 130 may further include one or more processors 156 that are coupled to the memories. Processor 156 may be configured to: down-sample in a first direction on a first set of pixels in a first patch of a first image to generate a first down-sample; and down-sample in a second direction on a second set of pixels in a second patch of a second image to generate a second down-sample, in which, the second down-sample includes a greater number of pixels. As has been described, processor 156 may implement the functions of an encoder and / or decoder or separate encoders and / or decoders may be utilized on the same device or different devices.

[0072] In one example, as be described in more detail hereafter, the first down-sample in the first direction is in height and the second down- sample in the second direction is in width, such that, the first and second down-sample is an asymmetric down-sample operation that includes a higher resolution in width.

[0073] Therefore, in one example, the left camera 132 is configured to capture the left image and the right camera 134 is configured to capture the right image, and the one or the processors 156 are configured to generate both the first and second down-samples in both height and width from each of the left and right images, respectively. In this way, device 130 may be configured to implement asymmetric operations for the width and height of the image captured by the plurality of cameras 132, 134 during down- sampling operations, in which, the asymmetric operations include higher resolution in width. Based upon the asymmetric operations, stereo vision of the image may be provided. Utilizing these techniques, stereo vision is provided that preserves disparity and enhanced resolution, while still being performed in an efficient manner. For example, stereo visionQualcomm Ref. No. 2407692WO 18 of the image may be displayed on a display device 122. In particular, by utilizing these techniques, stereo vision is provided that preserves disparity and enhanced resolution, by focusing more on width than height, while being done in a more efficient computational manner, which results in less computational tasks and less power than the conventional processes. It should be appreciated that terminology down-sampling and encoding and up-sampling and decoding are used interchangeably throughout the disclosure.

[0074] Aspects of the disclosure relate to a device or system that provides multi-aspect- ratio implementation in down-sampling for stereo disparity estimation. For example, multi-aspect-ratio down-sampling for stereo depth is presented that provides for disparity preservation and width-centric processing for disparity handling. Further, as will be described, asymmetric space-to -depth encoding and depth-to-space decoding is provided for disparity estimation. For example, disparate-rate height-width space-to-depth encoding and disparate-rate height-width depth-to-space encoding will be described.

[0075] The modified stereo vision processing implementation that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, a higher resolution is in width, will now be described in greater detail. Aspects of the disclosure relate to a description of modified stereo vision processing implementation that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, a higher resolution is in width.

[0076] It should be appreciated that system or device 130 is merely an example. Further, as has been described, processor 156 may implement the functions of an encoder and / or decoder or separate encoders and / or decoders may be utilized on the same device or different devices.

[0077] In one aspect, to address problems associated with the previously described common practice of an encoder performing feature extraction by down- sampling feature maps that results in the loss of critical details and depth estimation accuracy, aspects of the disclosure provide embodiments related to multi- aspect-ratio down-sampling for stereo depth that provide for disparity preservation and width-centric processing for disparity preservation. The multi- aspect-ratio down-sampling for stereo depth and widthcentric processing methods to be described estimate pixel-wise disparities between rectified stereo images in a manner that provides for disparity preservation. In one aspect, the disparity information per-pixel is carried by stereo inputs. As one example, processor 156 may operate as a feature extractor and / or encoder for down- sampling and may implement a machine-learning (ML) module to implicitly carry the disparity information.Qualcomm Ref. No. 2407692WO 19

[0078] As an example of implementation, with reference to FIG. 8, a world point P(X,Y,Z), a left image plane 802 of the left camera, and a right image plane 804 of the right camera are shown. Further, the left camera center 01 and right center camera Or are shown. Based upon these points, pl (xl,yl) and pr (xr,yr) on the left image plane and the right image plane are shown, respectively. It should be noted that in this horizontally rectified stereo set-up, the disparity information is carried between the stereo images for the world point P, which is projected on the stereo left and right images 802 and 804. In particular, the width disparity may be considered to be xr-xl.

[0079] With additional reference to FIG. 9, FIG. 9 illustrates disparities between pixels on the left and right image planes. As can be seen in FIG. 9, with respect to a top high resolution example 910, a left and right image plane 912 and 914 are shown, each having top and bottom pixels (the left and right image planes, having a y-axis in height (H) and x-axis in width (W)). In particular, as shown, the disparities between the top and bottom pixels on the left image plane 912 and the right image plane 914 are shown as dl and d2. The disparities may be considered equivalent to xr-xl, as previously described (e.g., in the width dimension).

[0080] Now considering the effect of the image down-sizing by a factor of r (e.g., r = 2, 4, 8, etc.) a lower resolution example 920 is shown, again with a left and right image plane 922 and 924, each having top and bottom pixels (the left and right image planes, having a y-axis in height (H / r) and x-axis in width (W / r)). As can be seen in this example, the down-sized (e.g., lower resolution) images are now shown with reduced disparities of dl’=dl / r and d2’=d2 / r, which are down-scaled by a factor of r. Accordingly, the model accuracy of disparity estimation is directly affected in width (e.g., horizontally), whereas height has not been found to be as an important of a factor. The utility of this disparity hypothesis will be further described hereafter in detail.

[0081] According to aspects of the disclosure, a technique for stereo depth estimation, in which, the disparity information as carried in the pixel- wise distance between the left and right image pairs (e.g., as previously shown in FIG. 9) and the encoded latent left and right (L and R) features is preserved. Further, convolution networks (e.g., neural networks) can further utilize these down-sized input images and latency encoded feature maps. It has been found in prior art implementations, that downsizing equally in height and width results in poor disparity estimation, whereas, aspects of the disclosure provide an approach to utilizing disparate down-sampling between height and width by keepingQualcomm Ref. No. 2407692WO 20 higher resolution in width (than in height) to better preserve disparity insight for stereo depth estimation.

[0082] With reference to FIG. 10, FIG. 10 is a flowchart illustrating down-sampling, in accordance with one or more techniques of this disclosure. At block 1002, one or more images are captured, in which each of the images includes one or more patches, each patch including a plurality of pixels. At block 1004, down- sampling occurs in a first direction on a first set of pixels in a first patch of a first image to generate a first downsample. At block 1006, down-sampling occurs in a second direction on a second set of pixels in a second patch of a second image to generate a second down- sample, wherein, the second down-sample includes a greater number of pixels.

[0083] With reference to FIG. 11, FIG. 11 is a diagram illustrating an example operation for down-sampling images, in accordance with one or more techniques of this disclosure. The operations presented in the flowcharts of this disclosure are provided merely as examples. At block 1100, left and right images are captured from left and right cameras (e.g., cameras 132 and 134). At block 1102, down-sampling occurs. As an example of down-sampling, down-sampling may occur asymmetrically with higher resolution in one direction (block 1104). As has been previously described, in one aspect, down-sampling occurs in a horizontal directional on a first set of pixels on each of the left and right images to generate a horizontal down-sample, and, down-sampling occurs in a width direction on a second set of pixels on each of the left and right images to generate a width downsample, in which, the width down-sample includes a greater number of pixels. In this way, the down-sampling is any asymmetric down-sample operation that includes a higher resolution in width. Stereo depth estimation may then be performed based upon the down-sampling operation (block 1106), as will be described in more detail hereafter. Further, as an example, processor 156 may render the output of the down-sampling for the left and right images and combine the left and right rendered images to generate a stereo image that is displayed on a display device 122 (block 1108), as will be described in more detail hereafter.

[0084] Therefore, down- sampling operations may be performed that are asymmetric (e.g., they may include higher resolution in width). In one example aspect, multiple asymmetric down- sample operations may be performed, in which, each asymmetric down- sample operation includes a pre-determined width-to-heigh aspect ratio.Qualcomm Ref. No. 2407692WO 21

[0085] In one example aspect, assuming the aspect ratio to a processing operation i is denoted as yt= — , i = 1,2, ... , 1V among a total of N operations of the model starting / i, with i = 1 for the first model operation in training or inference by a processor (e.g., processor 156 implementing a ML neural network), where ly and vy are the height and width for operation i, then the model architecture may include the property of multipleaspect ratios for encoding and decoding features:

[0086] Yt < Yj for i < j during encoding (down-sampling stages) among {1,2, ... , 1V}, and

[0087] YiYj f°ri < j during decoding (up-sampling stages) among {1,2, ... , 1V}.

[0088] With reference to FIG. 12A, FIG. 12A is a diagram illustrating encoding / down- sampling utilizing multiple- aspect ratios. As will be described in FIG. 12B, mirrored decoding / up-sampling will also be shown. As shown in FIG. 12A, encoding / down- sampling 1202 illustrates encoding / down-sampling of image data that is down-sized by a factor of yi (e.g., i= 1, 2, 3, 4, 5), such that the first encoded data image block has a downsize factor of i=l [yl] 1204 (stage 1), the second encoded image data block has downsize factor i=2 [y2] 1206 (stage 2), the third encoded image data block has down- size factor i=3 [y3] 1208 (stage 3), the fourth encoded image data block has down-size factor i=4 [y4] 1210 (stage 4), and the fifth encoded image data block has down- size factor i=5 [y5] 1212 (stage 5). Each of these image data blocks 1204, 1206, 1208, 1210, and 1212 (stages 1, 2, 3, 4, 5) is down-sized with an asymmetric aspect ratio yt=such that,horizontal width is weighted with more importance than height.

[0089] In this example of the encoding / down-sampling 1202, assuming the aspect ratio to this processing operation is set in a processing encoder (e.g., implementing a ML neural network (e.g., implemented by processor 156 or a particular encoder)), in which the aspect ratio, is defined as denoted as yt= — , i = 1,2, ... , N among a total of N operations (e.g. / i,N=5) of the model starting with i = 1 for the first model operation in training or inference and proceeding to i=5, where ly and vy are the height and width for each operation i, then the model architecture may include the property of multiple aspect ratios to encoded features - which can be seen as down-sized image data blocks 1204, 1206, 1208, 1210, and 1212 (stages 1, 2, 3, 4, 5).

[0090] It should be appreciated that in prior art implementations, down-sample factors in terms of width and height are be equally-weighted in terms of height and width. AnQualcomm Ref. No. 2407692WO 22 example of this would be equally down-sizing in both height and width by: 1 / 2, 1 / 4, 1 / 8, etc. For example, in prior down-sampling implementations R_h=R_w in each stage of down-sampling. For example, going from stages: I a 2 a 3 a 4 a 5 - the pair R_h=R_w may be (2,2)a(2,2) a(2,2) a(2,2) a(2,2). However, in the aspects of previously described disclosure,, i = 1,2, ... , N - R_h>R_w is implemented in each stage of down- / i, sampling. For example, going from stages: I a 2 a 3 a 4 a 5 (1204, 1206, 1208, 1210, 1212) - the pair (R_h, R_w) may be: (4,2)a(2,l) a(2,2) a(2,l) a(2,2). Other down-sizing implementations are also possible. However, because R_h>R_w is held true for each of the stages, implementing 5 stages in this example (1204, 1206, 1208, 1210, and 1212), resolution is preserved in the dimension of width better than in height. It should be appreciated that multiple width-to-height aspect ratios may be used during down- sampling / encoding. Also, the multiple width-to-height aspect ratios may be equal or increasing or decreasing during down-sampling / encoding operations.

[0091] With additional reference to FIG. 12B, in some example aspects, these down- sampled image data blocks 1204, 1206, 1208, 1210, and 1212 can be up-sampled by an automatic decoder 1215, in which, the up-sampled image data blocks are in the same feature / space domain and exactly match the down-sized image data blocks, as shown on the decoding / up-sampling side 1220 - as image data blocks 1222, 1224, 1226, 1228, and 1230. However, the use of decoder 1215 is completely optional. In general, decoded or up-sampled image data blocks 1222, 1224, 1226, 1228, and 1230 that may be utilized would exactly match the corresponding down-sampled image data blocks.

[0092] By utilizing the previously described multi-aspect-ratio down-sampling implementations that focus more on width than in height for stereo depth (e.g., widthcentric), pixel-wise disparities between rectified stereo images are processed in a manner that provides disparity preservation. In one aspect, the disparity information per-pixel is carried by the stereo inputs and is then down-sampled / encoded as previously illustrated. In one aspect, processor 156 may utilize an ML model to perform the previously described functions of down-sampling. Further, as example aspects, by utilizing encoder(s) that operate as ML modules the disparity information may be implicitly carried. The modules utilizing ML (e.g., encoder) can utilize learning and / or inference.

[0093] Therefore, as has been described, processor 156 may operate to perform down- sampling / encoding functions and can implement the ML functions for learning and / or inference. Also, it should be appreciated that variants in the model architecture mayQualcomm Ref. No. 2407692WO 23 include multiple encoders, multiple decoders, interleaved encoder-decoder module (e.g., hour-glass modules, etc.). Further, it should be appreciated that a wide variety of neural network models, neural processors, neural hardware and / or software accelerators, etc. may be utilized. In a broad aspect, processor 156 may implement ML models during down-sampling and / or up-sampling to perform down-sampling / encoding functions and / or up-sampling / decoding functions and can implement ML functions for learning and / or inference.

[0094] In one example aspect, an up-sampling process implemented by processor 156 (or a separate decoder) may be used for stereo depth. In this case, a “coarse-to-fine” feature may be used for stereo depth as an overall algorithm to start stereo estimation at the coarse level before continuing to the next finer level. One reason for such type of stereo depth algorithm is that local minimums can be effectively removed / reduced. In this example, both down-sampling in the encoding feature and up-sampling in the coarse-to-fine stereo depth may be used in order for the overall stereo depth algorithm to properly run. Also, an up-sampling process may be used to serve two purposes: 1) to support multi-resolution stereo matching algorithm with a mixture of respective fields; and 2) to recover the estimated stereo disparity / depth map back to the original or desirable (higher) resolution. Therefore, the stereo matching algorithm may be used to leverage the coarse-to-fine resolution levels to avoid local minimums in optimization.

[0095] Further, additional layers of 2D convolution functions and / or 3D convolution functions may be implemented that provide spatial filtering on top of the previously described asymmetric down-sampling operations. This allows processor 156 implementing ML functions for learning and / or inference (e.g., implementing a neural network) to obtain more opportunities for learning and inference. Based upon the ML- based stereo matching algorithm and filtering functions during the down-sampling by the processor 156, the stereo image output rendered by the down-sampling process is improved and includes stereo depth map resolution that closely replicates the original stereo depth map resolution associated with the original stereo image. An example of the stereo depth map resolution will be described with reference to FIG. 14. As previously described, processor 156 of device 130 may command the display on a display device 122 of the stereo image output (as will be described with reference to FIG. 14).

[0096] In one example aspect, based upon the implementation of the ML model during the down-sampling process 1202 by the processor 156, the stereo image output rendered by the down-sampling process is improved and includes stereo depth map resolution thatQualcomm Ref. No. 2407692WO 24 closely replicates the original stereo depth map resolution associated with the original stereo image. An example of the stereo depth map resolution will be described with reference to FIG. 14. As previously described, processor 156 of device 130 may command the display on a display device 122 of the stereo image output (as will be described with reference to FIG. 14).

[0097] It should be appreciated that artificial intelligence (Al) functionality and machine learning (ML) functionality may be utilized in these operations for learning, inference, etc., in the encoding, decoding, and other operations. Al generally is a field of research in computer science that develops and studies methods and software that enable machines to perceive their environment and use learning and intelligence to take actions that maximize their chances of achieving defined goals, such as, making predictions, recommendations or decisions influencing real or virtual environments. In particular, Al is a set of technologies that enable computers to perform a variety of advanced functions, including the ability to see, understand and translate spoken and written language, analyze data, make recommendations, and many other functions. ML may be considered a field of study in Al concerned with the development and study of statistical algorithms that can learn from data and generalize to unseen data and thus perform tasks without explicit instructions. The term AI / ML prediction, learning, inference, etc., referred to herein, may be any type of Al and / or ML related techniques, processes, algorithms, etc., that may be utilized herein to achieve the described functions. In other aspects other techniques that are not Al and / or ML related may be utilized to achieve the described functions.

[0098] According to aspects of the disclosure, the previously described techniques for stereo depth estimation, that utilize multi-aspect-ratio down- sampling 1202 implementations that focuses more on width than in height for stereo depth (e.g., widthcentric) results in pixel-wise disparities between rectified stereo images being processed in a manner that provides disparity preservation. In this way, the previously described down-sampling process 1202 that implements down-sampling operations provides stereo vision of images with improved disparity preservation. Also, by utilizing the previously described techniques of the disclosure, stereo vision is provided that preserves disparity and enhanced resolution, while being done in a more efficient computational manner by focusing more on width than height, which results in less computational tasks and less power than the conventional process.Qualcomm Ref. No. 2407692WO 25

[0099] In another aspect, width-centric or disparity-dimension-centric processing may be utilized in down-sampling and up-sampling operations to provide stereo vision of an image in order to provide improved disparity preservation. Width-centric or disparity- dimension-centric processing may be utilized in down-sampling and up-sampling operations to facilitate improved learning and / or inference in ML model implementations to provide improved disparity preservation. In one aspect, asymmetric operations during down-sampling and up-sampling operations to increase width dimension weighting may be utilized. In one example aspect, an asymmetric attention mechanism may be performed to focus more heavily on the width dimension. As an example type of asymmetric attention mechanism, variable width-to-height ratios for derivation of queries, keys, and / or values in favor of features in the width dimension may be utilized.

[0100] As one particular type of asymmetric operations, asymmetric tokenization rates in the width dimension may be utilized. As an example of an asymmetric operation, asymmetric operations that include the use of asymmetric patchification based upon asymmetric tokenization rates to increase width-to-height ratios of input images to allocate more patches in width than in height during asymmetric patchification may be utilized. As an example of asymmetric patchification, width-to-height ratios for 2D patches of features may be increased for encoding. For example, when provided an original R = W / H for an input, more patches may be allocated in width than in the height during patchification, such that, after patchification, the 2-D patches have an increased ratio of R’ = W7H’ > R = W / H. Therefore, patchification may be utilized as a special case of tokenization for 2D inputs in computer vision.

[0101] As another type of asymmetric operation, 1-D convolution for disparity-centric processing by focusing on the width dimension may be utilized in encoding and decoding operations to provide stereo vision of an image in order to provide improved disparity preservation. As to one type of asymmetric operation, asymmetric operations may include the use of variable-rate dilation for convolution in favor of the width dimension. As an example, when provided with 2-D inputs, dilated convolution that allows for asymmetric dilation rates between width and height dimensions may be utilized in favor of the width dimension. As another type of asymmetric operation, asymmetric operations may include the use of 1-D convolution for disparity-centric processing by focusing on the disparity dimension. For example, asymmetric separable convolution may be performed over the H and W dimensions. As one example, separable ID convolutions may be performed over the H and W dimension, but with different kernel sizes in favorQualcomm Ref. No. 2407692WO 26 of the width dimension. As one particular example, ConvlD of kernel Kh in height may be performed and another ConvlD of kernel Kw in width may be performed, where Kw > Kh so that the width dimension is favored. As another type of asymmetric operation, asymmetric kernels (or asummetric strids) for convolution to favor the width dimension may be utilized in encoding and decoding operations to provide stereo vision of an image in order to provide improved disparity preservation. For example, when provided with 2D inputs, a square K x K kernel for 2-D convolution may be utilized, such as 3 x 3. By utilizing asymmetric kernel convolution, Kh x Kw, may be utilized, where Kw > Kh, to favor the width dimension for more kernel weights to handle more details in the width dimension.

[0102] As yet another type of asymmetric operation according to another aspect, asymmetric Space-to-Depth (S2D) and Depth-to-Space (D2S) operations may be utilized. In current S2D / D2S operations, symmetric rates for Height (H) and Width (W) are utilized. According to another aspect, asymmetric S2D operations and asymmetric D2S operations for stereo depth estimation may be utilized in encoding-decoding implementations that focus more on width than in height for stereo depth results in pixelwise disparities between / among rectified stereo images being processed in a manner that provides disparity preservation.

[0103] As one example, asymmetric operations include the use of asymmetric S2D operations in the width dimension, in which, a smaller rate through division in the disparity width dimension is used than in other non-disparity dimensions. In particular, in order to preserve more feature information in the width dimension, a smaller rate through division in the width dimension than in the other non-disparity dimension is utilized when performing S2D operations.

[0104] As another example, asymmetric operations include the use of asymmetric D2S operations in the width dimension, in which, a larger rate through multiplication in the width dimension is used than in other non-disparity dimensions. In particular, in order to gain more feature information in the width dimension, a larger rate through multiplication in the width dimension is used than in the other non-disparity dimension.

[0105] In prior implementations, S2D operations and D2S operations were performed with symmetric rates, for down-sampling and up-sampling, in terms of [N, C, W, R].

[0106] In this implementation, N corresponds to batch, C corresponds to channel, H to height, W to width, and R to rate.Qualcomm Ref. No. 2407692WO 27

[0107] As can be seen with reference to FIG. 13, according to aspects of the disclosure, asymmetric operations include the use of asymmetric S2D operations in the width dimension for down-sampling 1302 (on the left side of the FIG. 13), in which, a smaller rate through division in the disparity width dimension is used than in other non-disparity dimensions. In particular, in order to preserve more feature information in the disparity dimension, a smaller rate “R” through division in the width dimension than in the other non-disparity dimension is utilized when performing S2D operations. This functionality is implemented by features below:

[0108] S2D: |N x C x H x WR |N x CRHRW x H / RH x W / RW] RH > RW

[0109] Instead of the standard symmetric operation [N x C x H x W], an asymmetric S2D operation may be utilized where [N x CRHRW x H / RH x W / RW] for down-sampling. RH may be considered a height rate factor and RW may be considered a width rate factor (in which RH is greater than RW) such that by utilizing a smaller rate factor through division in the width dimension than in the other non-disparity dimension in this formula more features in the width disparity dimension are preserved. Therefore, at each stage of S2D down- sampling, dimensionality changes in rates of RH and RW may be utilized.

[0110] As can be seen with reference to FIG. 13, D2S up-sampling rates of RH and RW for up-sampling 1304 can also be implemented, according to aspects of the disclosure, as shown on the right-side of FIG. 13. These asymmetric operations include the use of asymmetric D2S operations in the disparity dimension, in which, a larger rate through multiplication in the width dimension is used than in other non-disparity dimensions. In particular, in order to gain more feature information in the width dimension, a larger rate “R” through multiplication in the width dimension is used than in the other non-disparity dimension when performing D2S operations for up-sampling. This functionality is implemented by features below:

[0111] D2S: [N x C x H x W] [N x C / RHRW x HRH x W RW] RH < RW

[0112] In this aspect, an asymmetric D2S operation is utilized where [N x C / RHRW x HRH x W RW] . RH may be considered a height rate factor and RW may be considered a width rate factor (in which RH is less than RW) such that by utilizing a larger rate factor through multiplication in the width dimension than in the other non-disparity dimension in this formula more features in the width disparity dimension are preserved.

[0113] With brief reference to FIG. 14, FIG. 14 illustrates a proof-of-concept of the techniques of the disclosure related to down-sampling for stereo depth estimation utilizing asymmetric operations in width and height to provide a stereo view of an image thatQualcomm Ref. No. 2407692WO 28 preserves disparity and enhanced resolution, while still being performed in an efficient manner.

[0114] As has been described, processor 156 may operate to perform down- sampling / encoding functions and can implement ML functions for learning and / or inference. The encoding functions are based upon the asymmetric down-sizing operations for width and height of the image data captured by the cameras 132 and 134, as previously described, in which, the asymmetric operations include higher resolution in width. Based upon these implementation features during the down-sampling process by the processor 156, the stereo image output rendered by the up-sampling process is improved and includes stereo depth map resolution that closely replicates the original stereo depth map resolution associated with the original stereo image.

[0115] An example of the stereo depth map resolution can be seen with reference to FIG. 14. As can be seen in FIG. 14, in the upper-right, an image input of a man 1402 sitting at a table in front of kitchen with a plant 1404 in front of him is shown. The lower left image is a disparity map generated by a conventional process with down- sampling, in which height and width dimensions are equally weighted. The lower right is a disparity map generated by the previously described techniques to implement asymmetric operations for width and height of an image during down-sampling, in which, the asymmetric operations include higher resolution in width, in which, stereo vision is provided that preserves disparity and enhanced resolution, while still being performed in an efficient manner.

[0116] As can be seen in the lower right disparity map, performed with the previously described techniques of the disclosure, the disparity information is preserved. The disparity differences between the objects of the captured image - man 1402 sitting at the table in front of the kitchen with the plant 1404 in front of him - can be seen between the conventional process (left-hand side) and the previously described techniques of the disclosure (right-hand side), with few differences. However, by utilizing the previously described techniques of the disclosure, stereo vision is provided with preserved disparity and enhanced resolution, while being done more efficiently with less computational tasks and less power than the conventional process.

[0117] In particular, the quality of the lower right disparity map illustrates the improved features of the disclosure that utilize the previously described multi-aspect-ratio downsizing implementation that focuses more on width than in height for stereo depth. As has been described, the disparity information per-pixel is carried by the stereo inputs and isQualcomm Ref. No. 2407692WO 29 then down-sampled / encoded, as previously described. Further, by utilizing an encoder that utilizes ML functionality, the disparity information may be implicitly carried and encoder functions may utilize ML in learning and / or inference.

[0118] In order to support high-resolution input, down-sampling aggressively in order to meet real-time and power consumption requirements is currently needed. Aspects of the previously described disclosure describe multi-aspect ratio techniques related to downsampling for stereo depth estimation utilizing asymmetric operations in width and height, emphasizing width, to provide a stereo view that preserves disparity and enhanced resolution, while still being performed in an efficient manner. In one aspect, asymmetric down-sampling is implemented to better preserve disparity and to avoid low resolution in the width. Asymmetric super resolution may then be utilized to return desirable output as to the original input aspect ratio. For example, down-sampling may occur to as much as 32X in the height dimension, enabling a larger respective field, while keeping the disparity dimension down at 16X or even 8X. Further, disparity can be enhanced by allocating more computational power with asymmetric encoding and asymmetric super resolution.

[0119] As has been described, the previously described techniques for stereo depth estimation that utilize multi-aspect-ratio down-sizing implementations that focus more on width than in height for stereo depth (e.g., width-centric) results in pixel-wise disparities between rectified stereo images being processed in a manner that provides disparity preservation. Further, by the implementation of ML operations for learning and / or inference in encoding / down-sampling operations, stereo depth estimation for stereo images is further improved. In this way, down-sampling operations provide stereo vision of the image with improved disparity preservation. Also, by utilizing the previously described techniques of the disclosure, stereo vision is provided that preserves disparity and enhanced resolution, while being done in a more efficient computational manner by focusing more on width than height, which results in less computational tasks and less power than the conventional process. Therefore, as has been previously described in detail, a modified stereo vision processing implementation has been described that utilizes an asymmetric down-sampling operation for the height and width of the object, in which, a higher resolution is in width.

[0120] It should be appreciated that the features previously described for down-sampling for stereo depth estimation utilizing asymmetric operations in width and height may be utilized for a wide variety of different devices 130. In particular, these type of digitalQualcomm Ref. No. 2407692WO 30 video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, video teleconferencing devices, and the like. Also, such devices may implemented in scenarios related to vehicles, mobile devices, security, etc.

[0121] Various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as limitations.

[0122] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed with a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0123] Various modifications to the described aspects may be apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0124] The processes previously described may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.Qualcomm Ref. No. 2407692WO 31

[0125] Aspect 1: A device comprising: one or more memories configured to store one or more images; and one or more processors coupled to the one or memories, the one or more processors configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

[0126] Aspect 2: The device of aspect 1, further comprising a plurality of cameras configured to capture a plurality of images, the contextual data being based on the images.

[0127] Aspect 3: The de vice of aspect 2, wherein, obtaining contextual data from the plurality of images includes obtaining a first contextual data at a first time and a second contextual data at a second time.

[0128] Aspect 4: The device of aspect 3, wherein, determining if the contextual data change exceeds the threshold includes comparing the image of the first contextual data at the first time and the image of the second contextual data at the second time, and determining if the threshold is exceeded.

[0129] Aspect 5: The device of aspect 4, wherein, the contextual data change exceeding the threshold is based upon a motion change of the first and second images based upon at least one of a speed or direction.

[0130] Aspect 6: The device of aspect 5, further comprising a motion sensor, wherein the motion sensor contributes data to determining whether the contextual change exceeds the threshold.

[0131] Aspect 7: The device of aspect 6, wherein, when the motion change based upon the speed or direction, exceeds the threshold, proceed with the modified stereo vision processing procedure, wherein, the one or more processors are further configured to: determine a subset of cameras from the plurality of cameras to account for the motion change; and based upon the determined subset of cameras, utilize asymmetric downsampling operations for height and width to provide modified stereo vision.

[0132] Aspect 8: The device of aspect 7, wherein, the subset of cameras are determined based upon determining the subset of cameras that align a rectified epipolar line with a dominant direction of motion.

[0133] Aspect 9: The device of any aspects 1 through 8, wherein, the determined subset of cameras are positioned in a first axis along the device.

[0134] Aspect 10: The device of any aspects 1 through 9, wherein, the first axis is a horizontal axis aligned with an image displayed on a display of the device.Qualcomm Ref. No. 2407692WO 32

[0135] Aspect 11: The device of any aspects 1 through 10, wherein, the first axis is a vertical axis aligned with an image displayed on a display of the device.

[0136] Aspect 12: The device of any aspects 1 through 11, wherein, the first axis is off- axis with an image displayed on a display of the device.

[0137] Aspect 13: The device of any aspects 1 through 12, wherein, when motion change occurs relative to different patches of the image, the one or more processors are further configured to: determine a subset of cameras to account for the motion change relative to each patch; and switch to the determined subset of cameras for each patch to account for the motion change to provide modified stereo vision.

[0138] Aspect 14: The device of aspect 1, further comprising, an image sensor, wherein, the contextual data change is a change in a lighting condition measured by the image sensor, and when the lighting condition change exceeds the threshold, the one or more processors are configured to switch from stereo vision processing to modified stereo vision processing.

[0139] Aspect 15: The device of any aspects 1 through 14, wherein, the modified stereo vision processing utilizes asymmetric down-sampling operations for height and width.

[0140] Aspect 16: The device of any aspects 1 through 15, wherein, when the lighting condition change exceeds the threshold for a first patch of the image, but not a second patch of the image, the one or more processors are further configured to switch from stereo vision processing to modified stereo vision processing for the first patch of the image but not the second patch of the image.

[0141] Aspect 17: The device of any aspects 1 through 16, wherein, the lighting condition change is determined to exceed the threshold based on decreased lighting.

[0142] Aspect 18: The device of any aspects 1 through 17, wherein, the lighting condition change is determined to exceed the threshold based on increased sunlight.

[0143] Aspect 19: A method comprising: obtaining contextual data; determining if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switching from stereo vision processing to a modified stereo vision processing.

[0144] Aspect 20: A non-transitory computer-readable data storage medium having stored thereon instructions that, when executed, cause one or more processors to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.Qualcomm Ref. No. 2407692WO 33

[0145] This disclosure describes one or more examples that may be applied independently or in a combined way. It is to be recognized that depending on the example, certain acts or events of any of the techniques described herein can be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques). Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.

[0146] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware -based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0147] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer- readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but are instead directed to non-transitory, tangible storage media. Disk and disc, as used herein, includesQualcomm Ref. No. 2407692WO 34 compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0148] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the terms “processor” and “processing circuitry,” as used herein may refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0149] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0150] One or more of the components, steps, features and / or functions illustrated in FIGs. 1-14 may be rearranged and / or combined into a single component, step, feature or function or embodied in several components, steps, or functions. Additional elements, components, steps, and / or functions may also be added without departing from novel features disclosed herein. The apparatus, devices, and / or components illustrated in FIGs. 1-14 may be configured to perform one or more of the methods, features, or steps described herein. The novel algorithms described herein may also be efficiently implemented in software and / or embedded in hardware.

[0151] It is to be understood that the specific order or hierarchy of steps in the methods disclosed is an illustration of exemplary processes. Based upon design preferences, it is understood that the specific order or hierarchy of steps in the methods may be rearranged. The accompanying method claims present elements of the various steps in a sample orderQualcomm Ref. No. 2407692WO 35 and are not meant to be limited to the specific order or hierarchy presented unless specifically recited therein.

[0152] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” Unless specifically stated otherwise, the term “some” refers to one or more. A phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover: a; b; c; a and b; a and c; b and c; and a, b, and c. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

[0153] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. Qualcomm Ref. No. 2407692WO 36CLAIMSWhat is claimed is:

1. A device comprising: one or more memories configured to store one or more images; and one or more processors coupled to the one or memories, the one or more processors configured to: obtain contextual data; determine if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

2. The device of claim 1, further comprising a plurality of cameras configured to capture a plurality of images, the contextual data being based on the images.

3. The device of claim 2, wherein, obtaining contextual data from the plurality of images includes obtaining a first contextual data at a first time and a second contextual data at a second time.

4. The device of claim 3, wherein, determining if the contextual data change exceeds the threshold includes comparing the image of the first contextual data at the first time and the image of the second contextual data at the second time, and determining if the threshold is exceeded.

5. The device of claim 4, wherein, the contextual data change exceeding the threshold is based upon a motion change of the first and second images based upon at least one of a speed or direction.

6. The device of claim 5, further comprising a motion sensor, wherein the motion sensor contributes data to determining whether the contextual change exceeds the threshold.Qualcomm Ref. No. 2407692WO 377. The device of claim 6, wherein, when the motion change based upon the speed or direction, exceeds the threshold, proceed with the modified stereo vision processing procedure, wherein, the one or more processors are further configured to: determine a subset of cameras from the plurality of cameras to account for the motion change; and based upon the determined subset of cameras, utilize asymmetric downsampling operations for height and width to provide modified stereo vision.

8. The device of claim 7, wherein, the subset of cameras are determined based upon determining the subset of cameras that align a rectified epipolar line with a dominant direction of motion.

9. The device of claim 7, wherein, the determined subset of cameras are positioned in a first axis along the device.

10. The device of claim 9, wherein, the first axis is a horizontal axis aligned with an image displayed on a display of the device.

11. The device of claim 9, wherein, the first axis is a vertical axis aligned with an image displayed on a display of the device.

12. The device of claim 9, wherein, the first axis is off-axis with an image displayed on a display of the device.

13. The device of claim 5, wherein, when motion change occurs relative to different patches of the image, the one or more processors are further configured to: determine a subset of cameras to account for the motion change relative to each patch; and switch to the determined subset of cameras for each patch to account for the motion change to provide modified stereo vision.Qualcomm Ref. No. 2407692WO 3814. The device of claim 1, further comprising, an image sensor, wherein, the contextual data change is a change in a lighting condition measured by the image sensor, and when the lighting condition change exceeds the threshold, the one or more processors are configured to switch from stereo vision processing to modified stereo vision processing.

15. The device of claim 14, wherein, the modified stereo vision processing utilizes asymmetric down-sampling operations for height and width.

16. The device of claim 14, wherein, when the lighting condition change exceeds the threshold for a first patch of the image, but not a second patch of the image, the one or more processors are further configured to switch from stereo vision processing to modified stereo vision processing for the first patch of the image but not the second patch of the image.

17. The device of claim 14, wherein, the lighting condition change is determined to exceed the threshold based on decreased lighting.

18. The device of claim 14, wherein, the lighting condition change is determined to exceed the threshold based on increased sunlight.

19. A method comprising: obtaining contextual data; determining if a contextual data change exceeds a threshold; and when the contextual data change exceeds the threshold, switching from stereo vision processing to a modified stereo vision processing.

20. A non-transitory computer-readable data storage medium having stored thereon instructions that, when executed, cause one or more processors to: obtain contextual data; determine if a contextual data change exceeds a threshold; andQualcomm Ref. No. 2407692WO 39 when the contextual data change exceeds the threshold, switch from stereo vision processing to a modified stereo vision processing.

Citation Information

Patent Citations

  • Stereo vision system and stereo vision processing method

    US20090060280A1

  • Depth mapping with a head mounted display using stereo cameras and structured light

    US20190387218A1

  • Dynamic camera selection

    WO2023239877A1