Smart Sensor Implementation of Region of Interest Operation Mode

By operating in ROI modes and utilizing integrated circuit layers for parallel processing, the image sensor effectively addresses bandwidth and processing limitations, enhancing object detection and tracking in autonomous vehicles.

JP7737446B2Active Publication Date: 2025-09-10WAYMO LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023519430
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-16
Filing Date
2021-12-03
Publication Date
2025-09-10
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Autonomous vehicles face challenges in processing large amounts of image data generated by high-resolution image sensors due to limited data transfer bandwidth, latency, and processing power, which can hinder timely detection and analysis of objects in the environment.

Method used

The image sensor is configured to operate in multiple regions of interest (ROIs), allowing selective reading and processing of ROI images at a higher frame rate, reducing the amount of data generated and processed, and utilizing integrated circuit layers for parallel image processing.

Benefits of technology

This approach enables efficient processing and analysis of high-speed and distant objects with reduced bandwidth and latency, improving object detection and tracking capabilities in autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007737446000001
    Figure 0007737446000001
  • Figure 0007737446000002
    Figure 0007737446000002
  • Figure 0007737446000003
    Figure 0007737446000003
Patent Text Reader

Abstract

The system includes an image sensor having a plurality of pixels forming a plurality of regions of interest (ROIs) and configured to operate at a frame rate greater than a threshold rate. The system also includes image processing resources. The system further includes control circuitry configured to perform operations including acquiring a full-resolution image of an environment from the image sensor. The full-resolution image includes each respective ROI of the plurality of ROIs. The operations also include selecting a specific ROI based on the full-resolution image and detecting an object of interest within the specific ROI. The operations include determining an operating mode in which subsequent image data generated by the specific ROI will be processed. The operations further include processing the image data including the plurality of ROI images of the object of interest based on the operating mode and the frame rate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 17 / 123,440, filed December 16, 2020, which is incorporated herein by reference in its entirety. [Background technology]

[0002] The image sensor includes a plurality of light-sensing pixels that measure the intensity of light incident thereon, thereby collectively capturing an image of the environment. A frame rate may be applied to the image sensor to enable the image sensor to generate an image of the environment. Image sensors may be used in a number of applications, such as photography, robotics, and autonomous vehicles. Summary of the Invention

[0003] The image sensor may include multiple regions of interest (ROIs), each of which may be read out independently of the other ROIs. The image sensor may be used to capture a full-resolution image including each of the ROIs. The ROIs from which ROI images are acquired may be selected based on the content of the full-resolution image. These ROI images may be acquired at a higher frame rate than the full-resolution image. The image sensor may be configured to operate at a frame rate higher than a threshold rate, and a corresponding number of frames may be processed during a given duty cycle. Upon detection of an object of interest, a frame including the ROI image of the object of interest may be processed during one or more subsequent duty cycles, thereby increasing the image processing speed.

[0004] In a first example embodiment, a system is provided that includes an image sensor having a plurality of pixels forming a plurality of regions of interest (ROIs). The image sensor is configured to operate at a frame rate greater than a threshold rate. The system also includes image processing resources. The system further includes control circuitry configured to perform operations including acquiring a full-resolution image of an environment from the image sensor. The full-resolution image includes each respective ROI of the plurality of ROIs. The operations also include selecting a specific ROI based on the full-resolution image. The operations further include detecting at least one object of interest within the specific ROI. The operations also include determining an operating mode in which subsequent image data generated by the specific ROI will be processed. The operations further include processing the image data including the plurality of ROI images of the at least one object of interest based on the operating mode and the frame rate.

[0005] In a second example embodiment, a method is provided that includes acquiring, by a control circuit, a full-resolution image of an environment from an image sensor including a plurality of pixels forming a plurality of regions of interest (ROIs). The full-resolution image includes each respective ROI of the plurality of ROIs. The image sensor is configured to operate at a frame rate greater than a threshold rate. The method also includes selecting a specific ROI based on the full-resolution image. The method further includes detecting at least one object of interest within the specific ROI. The method also includes determining an operational mode in which subsequent image data generated by the specific ROI will be processed. The method further includes processing the image data including the plurality of ROI images of the at least one object of interest based on the operational mode and the frame rate.

[0006] In a third example embodiment, a non-transitory computer-readable storage medium is provided having stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations. The operations include acquiring a full-resolution image of an environment from an image sensor including a plurality of pixels forming a plurality of ROIs. The full-resolution image includes each respective ROI of the plurality of ROIs. The image sensor is configured to operate at a frame rate greater than a threshold rate. The operations also include selecting a specific ROI based on the full-resolution image. The operations further include detecting at least one object of interest within the specific ROI. The operations also include determining an operational mode in which subsequent image data generated by the specific ROI will be processed. The operations further include processing the image data including the plurality of ROI images of the at least one object of interest based on the operational mode and the frame rate.

[0007] These and other embodiments, aspects, advantages, and alternatives will become apparent to those skilled in the art upon reading the following detailed description, with reference to the accompanying drawings as appropriate. Moreover, this summary and other descriptions and figures provided herein are intended to illustrate embodiments by way of example only, and as such, numerous variations are possible. For example, structural elements and process steps may be rearranged, combined, distributed, eliminated, or otherwise modified while remaining within the scope of the claimed embodiments. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 shows a block diagram of an image sensor having three integrated circuit layers, according to an example embodiment. [Figure 2] FIG. 2 illustrates placement of regions of interest, according to an example embodiment. [Figure 3] FIG. 3 illustrates an example architecture of an ROI mode processor for an image sensor, according to an example embodiment. [Figure 4] FIG. 4 illustrates an example ROI mode processor, according to an example embodiment. [Figure 5]FIG. 5 is a conceptual diagram of wireless communication between various computing systems associated with an autonomous vehicle, according to an example embodiment. [Figure 6] FIG. 6 illustrates an example sensor architecture in which various types of sensors are configured to share one or more pieces of ROI information with each other via a peer-to-peer network, according to an example embodiment. [Figure 7] FIG. 7 shows a flow chart according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Example methods, apparatus, and systems are described herein. It should be understood that the terms "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example," "exemplary," and / or "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or features, unless stated otherwise. Accordingly, other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0010] Accordingly, the example embodiments described herein are not intended to be limiting, as it will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0011] Furthermore, unless the context indicates otherwise, the features shown in each of the figures may be used in combination with one another. Thus, the figures should be viewed generally as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are required for each embodiment.

[0012] Additionally, any listing of elements, blocks, or steps in the specification or claims is for clarity purposes. Thus, such listing should not be construed as requiring or implying that these elements, blocks, or steps adhere to a particular arrangement or be performed in a particular order. Unless otherwise noted, figures are not drawn to scale. I. Overview

[0013] Image sensors may be provided in autonomous vehicles to aid in perception of and navigation within an environment. In some cases, these image sensors may have the capability to generate more data than can be processed in a timely manner by the autonomous vehicle's control system. This may be the case, for example, when multiple image sensors are provided in the vehicle and / or when the resolution of each image sensor is high, which may result in the generation of large amounts of image data. In some cases, the amount of data transfer bandwidth available to the autonomous vehicle and / or the desired data transfer latency may limit the amount of image data that can be utilized by the control system. In other cases, the amount of processing power provided by the control system may limit the amount of image data that can be processed.

[0014] Generating and processing larger amounts of image data may generally be desirable because it may enable detection, tracking, classification, and other analysis of objects in an environment. For example, capturing images at a high frame rate may enable analysis of fast-moving objects by representing such objects without motion blur. In some implementations, the frame rate may be set to be higher than a threshold rate. For example, the threshold rate may be 50 frames per second (fps), and the frame rate may be 50 fps or greater. Similarly, capturing high-resolution images may enable analysis of objects that are far away or in low-light-intensity environments. Notably, in some cases, such objects may be represented in a portion of the overall image generated by the image sensor, rather than covering the entire image.

[0015] Thus, the image sensor and corresponding circuitry may be configured to operate in one or more modes of operation. For example, the image sensor may be configured to generate a full-resolution image (i.e., an image including each of the ROIs) and an ROI image that includes a portion of the object of interest therein, but not all of the ROI. Also, for example, the image sensor may be configured to generate an ROI image that includes a portion of the object of interest therein, but not a full-resolution image. In one example, one or more ROI images may be used to determine the velocity of at least one object of interest.

[0016] For example, the image sensor may be configured to operate at a high frame rate of 50 frames per second (fps). Generally, the frame rate may determine the number of ROI images of an object of interest. For example, a configuration with a frame rate of 50 fps may result in the generation of 5 frames per millisecond (msec). These five frames may be used to perform image processing tasks, such as noise reduction, image enhancement, image sharpening, object detection, object tracking, etc. Upon detecting an object of interest, an ROI image of the detected object may be generated and further processed. For example, one full-resolution image may be processed in 100 msec, followed by one full-resolution image and one or more ROI images for the next 100 msec. This may be followed by processing only the ROI image for the next 100 msec, and so on. This processing cycle may be repeated.

[0017] To this end, an image sensor and corresponding circuitry may be provided that divides the image sensor into multiple ROIs and enables selective reading of image data from individual ROIs. These ROIs may enable the image sensor to accommodate high frame rates and / or high-resolution imaging while reducing the total amount of image data generated, transmitted, and analyzed. That is, the image sensor may generate image data for some ROIs (e.g., those that include some objects of interest therein), while other ROIs (e.g., those that do not include at least one object of interest) are not used to generate image data. The full-resolution image may be analyzed to detect at least one object of interest therein, and the ROI in which at least one object of interest is represented may then be used to generate multiple ROI images for further analysis.

[0018] In another example, the full-resolution image may be used to determine the distance between the image sensor and at least one object of interest. When this distance straddles a threshold distance (e.g., exceeds or falls below the threshold distance), a ROI containing the object may be selected and used to acquire multiple ROI images of the object. For example, if a pedestrian is determined to be within a threshold distance of the autonomous vehicle based on the full-resolution image, multiple ROI images of the pedestrian may be acquired. The ROI may additionally or alternatively be selected based on a classification of the object in the full-resolution image or other characteristics of the object or the full-resolution image. For example, an ROI image of a traffic light or traffic signal may be acquired based on its classification. These and other attributes of at least one object of interest in the full-resolution image may be used independently or in combination to select a particular ROI.

[0019] In another example, the full resolution image may be used to determine the speed of at least one object of interest represented in the full resolution image. When this speed straddles a threshold speed (e.g., exceeds or falls below the threshold speed), a ROI containing the object may be selected and used to acquire multiple ROI images of the object. For example, if an oncoming vehicle is determined to be within the threshold speed based on the full resolution image, multiple ROI images of a pedestrian may be acquired. The ROI may additionally or alternatively be selected based on a classification of at least one object of interest in the full resolution image or other characteristics of the object or the full resolution image.

[0020] In some embodiments, detecting the at least one object of interest may include determining one or more of, among other possibilities, a geometric characteristic of the at least one object of interest, a position of the at least one object of interest within the environment, a change in the position of the at least one object of interest over time (e.g., speed, acceleration, trajectory), an optical flow associated with the at least one object of interest, a classification of the at least one object of interest (e.g., vehicle, pedestrian, vegetation, road, sidewalk, etc.), or one or more confidence values ​​associated with an analysis of the multiple ROI images (e.g., a confidence level that the object is another vehicle). The ROI used to generate the ROI image may change over time as new objects of interest are detected in different portions of the new full-resolution image.

[0021] Because of their smaller size, ROI images can be generated, transferred, and analyzed at a higher rate than full-resolution images, thereby enabling, for example, more accurate analysis of high-speed and / or distant objects (e.g., due to the absence of motion blur and the availability of redundant high-resolution data). Specifically, to capture ROI images at a higher frame rate, analog-to-digital converters (ADCs) may be reassigned by a multiplexer from other ROIs to a selected ROI. Thus, an ADC otherwise used to digitize pixels of other ROIs may be used to digitize pixels of the selected ROI in combination with the ADC assigned to the selected ROI. Each of these ADCs may operate to digitize multiple rows or columns of the ROI in parallel. For example, four different ADC banks may read out four different columns of the selected ROI in parallel.

[0022] At least a portion of the image sensor and associated circuitry may be implemented as layers of an integrated circuit, which may be communicatively connected to a central control system of the autonomous vehicle. For example, a first layer of the integrated circuit may implement pixels of the image sensor, a second layer of the integrated circuit may implement image processing circuitry (e.g., high dynamic range (HDR) algorithms, ADCs, pixel memory, etc.) configured to process signals generated by the pixels, and a third layer of the integrated circuit may implement neural network circuitry configured to analyze the signals generated by the image processing circuitry of the second layer for object detection, classification, and other attributes.

[0023] Thus, in some implementations, the generation and analysis of both the full-resolution image and the ROI image, as well as the selection of the ROI representing at least one object of interest, may be performed by the integrated circuit. Once the full-resolution image and the ROI image are processed and attributes of the at least one object of interest are determined, the attributes of the at least one object of interest may be transmitted to a control system of the autonomous vehicle. That is, the full-resolution image and / or the ROI image may not be transmitted to the control system, thereby reducing the amount of bandwidth utilized and required by communication between the integrated circuit and the control system. Notably, in some cases, a portion of the generated image data (e.g., the full-resolution image) may be transmitted along with the attributes of the at least one object of interest.

[0024] In other implementations, the analysis of the full-resolution image and the ROI image and the selection of the ROI representing at least one feature of interest may be performed by a control system. The control system may then cause the integrated circuit to generate the ROI image using the selected region of interest. While this approach may utilize more bandwidth, it may nevertheless allow the control system to acquire more images (i.e., ROI images) representing portions of the environment of interest, rather than acquiring fewer full-resolution images representing portions of the environment lacking the feature of interest.

[0025] The image sensor and its circuitry may additionally be configured to generate a stacked full-resolution image based on multiple full-resolution images. The stacked full-resolution image may be an HDR image, an image representing multiple objects that are in focus despite the objects being located at different depths, or an image representing some other combination or processing of multiple full-resolution images. In particular, the multiple full-resolution images may be combined by the circuitry rather than by a vehicle's control system. Thus, the multiple full-resolution images may be generated at a higher frame rate than these images would otherwise be generated if each were provided to the control system. The stacked image may contain more information than an individual full-resolution image due to the fact that information from multiple such full-resolution images is represented in the stacked image. II. Examples of Smart Sensors

[0026] 1 is a block diagram of an example image sensor 100 having three integrated circuit layers. Image sensor 100 may detect objects using the three integrated circuit layers. For example, image sensor 100 may capture an image including a person and output an indication that a "person detected." In another example, image sensor 100 may capture an image and output a portion of the image including a vehicle detected by image sensor 100.

[0027] The three integrated circuit layers include a first integrated circuit layer 110, a second integrated circuit layer 120, and a third integrated circuit layer 130. The first integrated circuit layer 110 is stacked on the second integrated circuit layer 120, which is stacked on the third integrated circuit layer 130. The first integrated circuit layer 110 may be in electrical communication with the second integrated circuit layer 120. For example, the first integrated circuit layer 110 and the second integrated circuit layer 120 may be physically connected to each other by an interconnect. The second integrated circuit layer 120 may be in electrical communication with the third integrated circuit layer 130. For example, the second integrated circuit layer 120 and the third integrated circuit layer 130 may be physically connected to each other by an interconnect.

[0028] The first integrated circuit layer 110 may have the same area as the second integrated circuit layer 120. For example, the first integrated circuit layer 110 and the second integrated circuit layer 120 may have the same length and width, but different heights. The third integrated circuit layer 130 may have a larger area than the first and second integrated circuit layers 110, 120. For example, the third integrated circuit layer 130 may have a length and width that are both 20 percent larger than the length and width of the first and second integrated circuit layers 110, 120.

[0029] The first integrated circuit layer 110 may include an array of pixel sensors grouped by location into pixel sensor groups 112A-112C (collectively 112) (each pixel sensor group is referred to as a "pixel group" in FIG. 1). For example, the first integrated circuit layer 110 may include an array of 6400x4800 pixel sensors grouped into 320x240 pixel sensor groups, each pixel sensor group including an array of 20x20 pixel sensors. The pixel sensor groups 112 may be further grouped to define an ROI.

[0030] Each of the pixel sensor groups 112 may include 2x2 pixel sensor subgroups. For example, each of the 20x20 pixel sensor pixel sensor groups may include 10x10 pixel sensor subgroups, each including a red pixel sensor at the top left, a green pixel sensor at the bottom right, a first clear pixel sensor at the bottom left, and a second clear pixel sensor at the top right, each subgroup also referred to as a red-clear-clear-green (RCCG) subgroup.

[0031] In some implementations, the size of the pixel sensor group can be selected to increase silicon utilization, for example, the size of the pixel sensor group can be such that more of the silicon is covered by pixel sensor groups having the same pattern of pixel sensors.

[0032] Second integrated circuit layer 120 may include image processing circuit groups 122A-122C (collectively 122) (each image processing circuit group is referred to as a "processing group" in FIG. 1). For example, second integrated circuit layer 120 may include 320×240 image processing circuit groups. Each image processing circuit group 122 may be configured to receive pixel information from a corresponding group of pixel sensors and may be further configured to perform image processing operations on the pixel information to provide processed pixel information during operation of image sensor 100.

[0033] In some implementations, each image processing circuit group 122 may receive pixel information from a single corresponding pixel sensor group 112. For example, image processing circuit group 122A may receive pixel information from pixel sensor group 112A but not from any other pixel group, and image processing circuit group 122B may receive pixel information from pixel sensor group 112B but not from any other pixel group.

[0034] In some implementations, each image processing circuit group 122 may receive pixel information from multiple corresponding pixel sensor groups 112. For example, image processing circuit group 122A may receive pixel information from both pixel sensor groups 112A and 112B, but not from other pixel groups, and image processing circuit group 112B may receive pixel information from pixel group 112C and another pixel group, but not from other pixel groups.

[0035] Because image processing circuit group 122 may be physically close to the corresponding pixel sensor group 112, having image processing circuit group 122 receive pixel information from the corresponding pixel group may result in fast transfer of pixel information from first integrated circuit layer 110 to second layer 120. The longer the distance the information is transferred, the longer the transfer may take. For example, pixel sensor group 112A may be directly above image processing circuit group 122A, but pixel sensor group 112A may not be directly above image processing circuit group 122C, so if there is an interconnection between pixel sensor group 112A and image processing circuit group 122C, the transfer of pixel information from pixel sensor group 112A to image processing circuit group 122A may be faster than the transfer of pixel information from pixel sensor group 112A to image processing circuit group 122C.

[0036] Image processing circuit group 122 may be configured to perform image processing operations on pixel information that image processing circuit group 122 receives from pixel groups. For example, image processing circuit group 122A may perform high dynamic range fusion on pixel information from pixel sensor group 112A, and image processing circuit group 122B may perform high dynamic range fusion on pixel information from pixel sensor group 112B. Other image processing operations may include, for example, analog-to-digital signal conversion and demosaicing.

[0037] Having image processing circuit groups 122 perform image processing operations on pixel information from corresponding pixel sensor groups 112 may allow image processing operations to be performed in a parallel and distributed manner by image processing circuit groups 122. For example, image processing circuit group 122A may perform image processing operations on pixel information from pixel sensor group 112A at the same time that image processing circuit group 122B performs image processing operations on pixel information from pixel group 122B.

[0038] The third integrated circuit layer 130 may include neural network circuit groups 132A-132C (each neural network circuit group is referred to as an "NN group" in FIG. 1) 132A-132C (collectively referred to as 132) and a full image neural network circuit 134. For example, the third integrated circuit layer 130 may include a 320×240 neural network circuit group.

[0039] Each neural network circuit group 132 may be configured to receive processed pixel information from a corresponding image processing circuit group and may be further configured to perform analysis on the processed pixel information for object detection during operation of image sensor 100. In some implementations, each neural network circuit group 132 may implement a convolutional neural network (CNN).

[0040] In some implementations, each neural network circuit group 132 may receive processed pixel information from a single corresponding image processing circuit group 122. For example, neural network circuit group 132A may receive processed pixel information from image processing circuit group 122A but not from any other image processing circuit group, and neural network circuit group 132B may receive processed pixel information from image processing circuit group 122B but not from any other image processing circuit group.

[0041] In some implementations, each neural network circuit group 132 may receive processed pixel information from multiple corresponding image processing circuit groups 122. For example, neural network circuit group 132A may receive processed pixel information from both image processing circuit groups 122A and 122B, but not from the other image processing circuit groups, and neural network circuit group 132B may receive processed pixel information from both image processing circuit group 122C and another pixel group, but not from the other pixel group.

[0042] Because neural network circuit group 132 may be physically close to the corresponding image processing circuit group 122, having neural network circuit group 132 receive processed pixel information from the corresponding image processing circuit group may result in a fast transfer of the processed pixel information from second integrated circuit layer 120 to third integrated circuit layer 130. Again, the longer the distance the information is transferred, the longer the transfer may take. For example, image processing circuit group 122A may be directly above neural network circuit group 132A, such that transferring processed pixel information from image processing circuit group 122A to neural network circuit group 132A may be faster than transferring processed pixel information from image processing circuit group 122A to neural network circuit group 132C if an interconnection existed between image processing circuit group 122A and neural network circuit group 132C.

[0043] Neural network circuit group 132 may be configured to detect objects from the processed pixel information that neural network circuit group 132 receives from image processing circuit group 122. For example, neural network circuit group 132A may detect objects from the processed pixel information from image processing circuit group 122A, and neural network circuit group 132B may detect objects from the processed pixel information from image processing circuit group 122B.

[0044] Having neural network circuit groups 132 detect objects from processed pixel information from corresponding image processing circuit groups 122 allows detection to be performed in a parallel and distributed manner by each of neural network circuit groups 132. For example, neural network circuit group 132A may detect objects from processed pixel information from image processing circuit group 122A at the same time that neural network circuit group 132B may detect objects from processed pixel information from image processing circuit group 122B.

[0045] In some implementations, neural network circuit group 132 may perform intermediate processing. Thus, image sensor 100 may use three integrated circuit layers 110, 120, and 130 to perform some intermediate processing and output only intermediate results. For example, image sensor 100 may capture an image containing a person and output an indication of "region of interest in some region of the image" without classifying at least one object of interest (person). Other processing performed outside image sensor 100 may classify the region of interest as a person.

[0046] Thus, the output from image sensor 100 may include data representing the output of a convolutional neural network. While this data itself may be difficult to interpret, if continued to be processed outside of image sensor 100, the data may be used to classify an area as containing a person. This hybrid approach may have the advantage of reducing bandwidth requirements. Thus, the output from neural network circuit group 132 may include one or more of a selected region of interest for pixels representing a detection, metadata including temporal and geometric location information, intermediate computation results prior to object detection, statistics regarding the network's confidence level, and a classification of the detected object.

[0047] In some implementations, neural network circuit group 132 may be configured to implement a CNN with high recall and low precision. Neural network circuit group 132 may each output a list of detected objects, the locations where the objects were detected, and the timing of the object detection.

[0048] The full-image neural network circuit 134 may be configured to receive data from each of the neural network circuit groups 132 indicative of objects detected by the neural network circuit groups 132 and detect the objects from the data. For example, the neural network circuit group 132 may not be able to detect objects captured by multiple pixel groups because each individual neural network circuit group may receive only a portion of the processed pixel information corresponding to the object. However, the full-image neural network circuit 134 may receive data from multiple neural network circuit groups 132 and therefore be able to detect objects sensed by multiple pixel groups. In some implementations, the full-image neural network circuit 134 may implement a recurrent neural network (RNN). The neural network may be configurable both in terms of its architecture (number and type of layers, activation function, etc.) and in terms of the actual values ​​of the neural network components (e.g., weights, biases, etc.).

[0049] In some implementations, offloading processing to image sensor 100 may simplify the processing pipeline architecture, provide higher bandwidth and lower latency, enable selective frame rate operation, reduce costs with a stacked architecture, provide higher system reliability as integrated circuits may have fewer potential points of failure, and provide significant cost and power savings on computing resources. III. ROI Placement and Hardware Examples

[0050] FIG. 2 illustrates an example of an arrangement of ROIs on an image sensor. That is, image sensor 200 may include pixels forming C columns and R rows. Image sensor 200 may correspond to first integrated circuit layer 110. The image sensor may be divided into eight ROIs, including ROI0, ROI1, ROI2, ROI3, ROI4, ROI5, ROI6, and ROI7 (i.e., ROIs 0-7), each of which includes m columns of pixels and n rows of pixels. Thus, C=2m and R=4n. In some implementations, each ROI may include multiple pixel groups 112 therein. Alternatively, pixel groups 112 may be sized and positioned such that each pixel group is also an ROI. In some implementations, ROIs may be positioned such that they do not collectively span the entire area of ​​image sensor 200. Thus, the union of ROIs may be smaller than the area of ​​image sensor 200. Thus, a full-resolution image may have a higher pixel count than the union of ROIs.

[0051] FIG. 2 shows the ROIs arranged in two columns, with even-numbered ROIs on the left and odd-numbered ROIs on the right. However, in other implementations, the ROIs and their numbering may be arranged differently. For example, ROIs 0-3 may be in the left column and ROIs 4-7 may be in the right column. In another example, the ROIs may divide the image sensor 200 into eight columns organized in a single row, with the ROIs numbered 0-7 arranged from left to right along the eight columns of the single row. In some implementations, the ROIs may be fixed in a given arrangement. Alternatively, the ROIs may be reconfigurable. That is, the number of ROIs, the location of each ROI, and the shape of each ROI may be reconfigurable. IV. Operating Mode Architecture Examples

[0052] 3 illustrates an example architecture of an ROI mode processor for image sensor 200. Specifically, image sensor 200 may include ROI mode processor 302, pixel groups 310 and 312-314 (i.e., pixel groups 310-314) that define ROIs 0-7, respectively, and multiple image processing resources. The image processing resources include pixel-level processing circuits 320, 322, and 324-330 (i.e., pixel-level processing circuits 320-330), machine learning circuits 340, 342, and 344-350 (i.e., machine learning circuits 340-350), and communication connections 316, 332, and 352. Image sensor 200 may be configured to provide image data to control system 360, which may also be considered part of the image processing resources. Control system 360 may represent a combination of hardware and software configured to generate actions for a robotic device or autonomous vehicle, among other possibilities.

[0053] Pixel groups 310-314 represent groupings of pixels that make up image sensor 200. In some implementations, each of pixel groups 310-314 may correspond to one or more of pixel sensor groups 112. Pixel groups 310-314 may represent circuitry disposed within first integrated circuit layer 110. The number of pixel sensor groups represented by each of pixel groups 310-314 may depend on the size of each of ROIs 0-7. In implementations in which the number, size, and / or shape of ROIs are reconfigurable, the subset of pixel sensor groups 112 that make up each of pixel groups 310-314 may change over time based on the number, size, and / or shape of the ROIs.

[0054] The pixel-level processing circuits 320-330 represent circuits configured to perform pixel-level image processing operations. The pixel-level processing circuits 320-330 may operate on outputs generated by the pixel groups 310-314. Pixel-level operations may include analog-to-digital conversion, demosaicing, high dynamic range fusion, image sharpening, filtering, edge detection, and / or binarization. Pixel-level operations may also include other types of operations not performed on the image sensor 200 or by machine learning models (e.g., neural networks) provided by the control system 360. In some implementations, each of the pixel-level processing circuits 320-330 may include one or more of the processing groups 122, among other circuits configured to perform pixel-level image processing operations. Thus, the pixel-level circuits 320-330 may represent circuits disposed on the second integrated circuit layer 120.

[0055] The machine learning circuits 340-350 may include circuits configured to perform operations associated with one or more machine learning models. The machine learning circuits 340-350 may operate on outputs generated by the pixel groups 310-314 and / or the pixel-level processing circuits 320-330. In some implementations, each of the machine learning circuits 340-350 may correspond to one or more of the neural network group 132 and / or the full-image neural network circuit 132, among many other circuits that implement machine learning models. Thus, the machine learning circuits 340-350 may represent circuits disposed within the third integrated circuit layer 130.

[0056] Communication connection 316 may represent an electrical interconnection between pixel-level processing circuitry 320-330 and pixel groups 310-314. Similarly, communication connection 332 may represent an electrical interconnection between (i) machine learning circuitry 340-350 and (ii) pixel-level processing circuitry 320-330 and / or pixel groups 310-314. Additionally, communication connection 352 may represent an electrical interconnection between machine learning circuitry 340-350 and control system 360. Communication connections 316, 332, and 352 may be considered subsets of image processing resources, at least because these connections (i) facilitate the transfer of data between circuitry configured to process image data and (ii) may be modified over time to transfer data between different combinations of circuitry configured to process image data.

[0057] In some implementations, communication connection 316 may represent an electrical interconnection between first integrated circuit layer 110 and second integrated circuit layer 120, and communication connection 332 may represent an electrical interconnection between second integrated circuit layer 120 and third integrated circuit layer 130. Communication connection 352 may represent an electrical interconnection between third integrated circuit layer 130 and one or more circuit boards by which image sensor 200 is connected to control system 360. Each of communication connections 316, 332, and 352 may be associated with a corresponding maximum bandwidth.

[0058] Processing a full-resolution image may require more computing resources than processing an ROI. At a higher frame rate, a higher number of frames are likely available for processing. Increasing the number of frames increases the ability to process an image. For example, increasing the number of frames available for processing increases the likelihood of detecting an ROI containing an object of interest. Thus, once these specific ROIs are determined, focus can be directed to the object of interest. For example, further ROIs containing these objects of interest may be determined, and additional ROIs containing these objects of interest may be acquired using a higher frame rate. In some implementations, these additional ROIs containing the object of interest may be acquired at 150 frames per second.

[0059] Processing images at higher frame rates can have several associated resource cost considerations. For example, processing a large number of frames can result in a larger allocation of memory resources for storing the frames and / or the results of the processing. This can result in higher latency and likely consumes more computational power. Focusing on additional ROIs containing objects of interest can enable the sensor to alleviate some of these resource allocation factors, enabling smart utilization of available resources.

[0060] Thus, the ROI mode processor 302 may be configured to dynamically utilize the image processing resources 316, 320-330, 340-350, 332, 352, and / or 360 (i.e., image processing resources 316-360) available to the image sensor 200. Specifically, some image sensors may be configured with a fixed ROI mode processor 302. For example, in some image sensors 200, the ROI mode processor 302 may be configured in a first mode of operation that processes an ROI containing an object of interest and does not process full-resolution images. Also, for example, in some other image sensors 200, the ROI mode processor 302 may be configured in a second mode of operation that processes an ROI containing an object of interest along with full-resolution images. However, in some image sensors 200, the ROI mode processor 302 may be configured to dynamically switch between one or more modes of operation (e.g., a first mode of operation and a second mode of operation).

[0061] The configuration of the ROI mode processor 302 may depend on several factors. For example, the configuration may depend on the type of image capture device of the image sensor 200. In some implementations, the configuration may depend on the field of view of the image capture device. For example, when the image sensor 200 is mounted on a vehicle, the image capture device may have a forward field of view, a side field of view, a rear field of view, etc. In some embodiments, for an image sensor 200 having an image capture device with a forward field of view, the ROI mode processor 302 may be configured to operate in a second mode of operation to continue capturing and processing full resolution images to detect ROIs and objects of interest. However, for an image sensor 200 having an image capture device with a side field of view and / or a rear field of view, the ROI mode processor 302 may be configured to operate in a first mode of operation to continue capturing and processing ROI images of detected objects to enable object tracking functionality.

[0062] In some implementations, the configuration of the ROI mode processor 302 may depend on the time of day and / or the intensity of ambient lighting. For example, during the day, when object detection may be less difficult, the ROI mode processor 302 may be configured to operate in a mode that processes multiple full-resolution images. However, at night, in foggy conditions, rainy weather, etc., when ambient lighting intensity may be low and detection of ROIs and / or objects of interest may be more difficult, the ROI mode processor 302 may be configured to operate in a mode that intermittently processes full-resolution images while processing multiple ROIs that contain objects of interest.

[0063] In some implementations, the configuration of the ROI mode processor 302 may depend on the type of object of interest. For example, the size, type, and / or speed of at least one object of interest may require processing different image sets. For example, different operational modes may be utilized to detect a sport utility vehicle traveling at a first speed and a motorcycle traveling at a second speed. In general, smaller objects and / or objects moving at faster speeds may require faster processing, and the ROI mode processor 302 may be configured to process multiple ROIs that include only the object of interest.

[0064] The ROI mode processor 302 may also be configured to determine the number and / or type of ROI images and / or full-resolution images. For example, the ROI mode processor 302 may be configured to determine whether it needs to process a single full-resolution image with 10 ROIs, or whether it needs to process three single full-resolution images and three ROIs containing objects of interest. For example, if at least one object of interest is a pedestrian crossing a street, the ROI mode processor 302 may intelligently determine to process a greater number of ROIs containing at least one object of interest to enable enhanced object tracking. Processing more images may result in more false positives, but this is desirable to minimize errors in detecting and / or tracking pedestrians.

[0065] As described herein, additional factors may cause the ROI mode processor 302 to dynamically switch between one or more operating modes, including traffic conditions, weather conditions, road conditions, road construction sites, speed limits, road type, type of landscape (e.g., urban, rural), number of sensors positioned on the vehicle, memory allocation availability, processing power of the image sensor, etc.

[0066] The ROI mode processor 302 may be communicatively coupled to each of the image processing resources 316-360 and may be aware of (e.g., receive, access, and / or store display of) the capabilities, workload assigned to, and / or features detected by each of the image processing resources 316-360. Thus, in some implementations, the ROI mode processor 302 may be configured to distribute image data among the image processing resources 316-360 in a manner that improves or minimizes latency between image data acquisition and processing, improves or maximizes utilization of the image processing resources 316-360, and / or improves or maximizes throughput of image data through the image processing resources 316-360. These objectives may be quantified by one or more objective functions, each of which may be minimized or maximized (e.g., globally or locally) to achieve the corresponding objective. V. Examples of Operation Modes

[0067] 4 shows an example of an ROI mode processor 402. As indicated by the header column 402, the first (top) row of diagram 400 indicates a first operating mode of the image processor, the second row indicates a second operating mode of the image processor, the third row indicates a third operating mode of the image processor, and the fourth (bottom) row indicates the amount of time dedicated to each operation in the duty cycle. In some implementations, the duty cycle may last 100 milliseconds (ms). The image processor may collectively represent operations performed by second integrated circuit layer 120, third integrated circuit layer 130, and any other control systems communicatively connected to image sensor 100 or 200.

[0068] During interval 404, which may last 100 milliseconds, the image processor operating in the first mode of operation may process multiple ROIs and full-resolution images. During intervals 406 and 408, each of which may last 100 milliseconds, the image processor may process only the ROI of the object of interest and may not process the full-resolution image. During intervals 410, 412, and 414, each of which may last 100 milliseconds, the image processor may repeat the operations performed in intervals 404, 406, and 408. Processing multiple ROIs may enable the image processor to determine various attributes of the content of the environment represented by these images.

[0069] During interval 404, which may last 100 milliseconds, the image processor operating in the second mode of operation may process multiple ROIs and full-resolution images. During interval 406, which may last 100 milliseconds, the image processor may process only the ROI of the object of interest and may not process the full-resolution image. During intervals 406 and 410, each of which may last 100 milliseconds, the image processor may repeat the operations performed in intervals 404 and 406. Similarly, during intervals 412 and 414, each of which may last 100 milliseconds, the image processor may repeat the operations performed in intervals 404 and 406, and then 408 and 410. Processing full-resolution images at frequent intervals may enable the image processor to determine various attributes of the content of the environment represented by these images, especially when object detection from full-resolution images may be somewhat difficult.

[0070] In interval 404, which may last 100 milliseconds, the image processor operating in the third mode of operation may again acquire a full resolution image. In interval 406, which may last 100 milliseconds, the image processor may process the full resolution image. In some embodiments, the full resolution image may be a stacked image including one or more detected objects. In interval 408, which may last 100 milliseconds, the image processor may acquire an ROI, and in interval 410, which may last 100 milliseconds, the image processor may detect an object of interest. In intervals 406 and 408, each of which may last 100 milliseconds, the image processor may process only the ROI of the object of interest and may not process the full resolution image.

[0071] In particular, during interval 404, the image processor may be configured to process ROI images captured during the previous cycle (not shown). In particular, the amount of time for each interval and the number of ROI images captured for each full-resolution image may vary. For example, some tasks may involve capturing more ROI images than shown (e.g., 16 ROI images per full-resolution image) or fewer ROI images (e.g., 4 ROI images per full-resolution image). Furthermore, the size of each ROI and / or the amount of ADC provided to the image sensor, among other factors, may be used to determine the length of the interval during which ROI images are captured.

[0072] Typically, multiple images are generated, and an image processor may be configured to perform temporal processing on the multiple images to track an object of interest. For example, the image processor may detect at least one object of interest moving along a trajectory, and the neural network circuit group 132A-132C may oversample another ROI on that trajectory. Because there may be multiple ways to predict where at least one object of interest may be at a given time, the image processor may focus processing along those expected trajectories. A larger number of images provides more information about the ROI and / or object of interest. In some implementations, the multiple images enable the neural network circuit group 132A-132C to correlate information and make more accurate predictions. For example, temporal events that may be outside of the expected normal range may be processed, increasing confidence that an object is present and subsequent identification of the object. VI. Vehicle System Examples

[0073] 5 is a conceptual diagram of wireless communication between various computing systems associated with an autonomous vehicle, according to an example embodiment. In particular, wireless communication may occur between a remote computing system 502 and a vehicle 508 over a network 504. Wireless communication may also occur between a server computing system 506 and the remote computing system 502, and between the server computing system 506 and the vehicle 508.

[0074] The example vehicle 508 includes an image sensor 510. The image sensor 510 is mounted on the vehicle 508 and includes one or more image capture devices configured to detect information about the environment surrounding the vehicle 508 and output an indication of the information. For example, the image sensor 510 may include one or more cameras. The image sensor 510 may include one or more movable stages that may be operable to adjust the orientation of one or more cameras in the image sensor 510. In one embodiment, the movable stage may include a rotating platform that can scan the cameras to obtain information from each direction around the vehicle 508. In another embodiment, the movable stage of the image sensor 510 may be movable in a scanning manner within a particular range of angles and / or orientations. The image sensor 510 may be mounted on the roof of the vehicle, although other mounting locations are possible.

[0075] Additionally, the cameras of image sensor 510 may be distributed at different locations and need not be co-located at a single location. Furthermore, each camera of image sensor 510 may be configured to be moved or scanned independently of the other cameras of image sensor 510.

[0076] The remote computing system 502 may represent any type of device associated with remote assistance technologies, including but not limited to those described herein. In an example, the remote computing system 502 may represent any type of device configured to (i) receive information related to the vehicle 508, (ii) provide an interface through which a human operator can then observe the information and enter a response related to the information, and (iii) transmit the response to the vehicle 508 or to another device. The remote computing system 502 may take various forms, such as a workstation, a desktop computer, a laptop, a tablet, a mobile phone (e.g., a smartphone), and / or a server. In some examples, the remote computing system 502 may include multiple computing devices operating together in a network configuration.

[0077] The remote computing system 502 may include a processor configured to perform the various operations described herein. In some embodiments, the remote computing system 502 may also include a user interface including input / output devices such as a touchscreen and speakers. Other examples are possible as well.

[0078] Network 504 represents an infrastructure that enables wireless communication between remote computing system 502 and vehicle 508. Network 504 also enables wireless communication between server computing system 506 and remote computing system 502, and between server computing system 506 and vehicle 508.

[0079] The location of the remote computing system 502 may vary within examples. For example, the remote computing system 502 may have a remote location from the vehicle 508 with wireless communication over the network 504. In another example, the remote computing system 502 may correspond to a computing device within the vehicle 508 that is separate from the vehicle 508 but that allows a human operator to simultaneously interact with a passenger or driver of the vehicle 508. In some examples, the remote computing system 502 may be a computing device with a touchscreen operable by a passenger of the vehicle 508.

[0080] In some embodiments, the operations described herein that are performed by the remote computing system 502 may additionally or alternatively be performed by the vehicle 508 (i.e., by any system(s) or subsystem(s) of the vehicle 508). In other words, the vehicle 508 may be configured to provide remote assistance mechanisms with which the driver or passengers of the vehicle can interact.

[0081] The server computing system 506 may be configured to wirelessly communicate with the remote computing system 502 and the vehicle 508 (or perhaps directly with the remote computing system 502 and / or the vehicle 508) over the network 504. The server computing system 506 may represent any computing device configured to receive, store, determine, and / or transmit information regarding the vehicle 508 and its remote assistance. As such, the server computing system 506 may be configured to perform any operation(s), or portions of such operation(s), described herein as being performed by the remote computing system 502 and / or the vehicle 508. Some embodiments of wireless communication related to remote assistance may utilize the server computing system 506, while other embodiments may not.

[0082] The server computing system 506 may include one or more subsystems and components similar to or identical to the subsystems and components of the remote computing system 502 and / or the vehicle 508, such as a processor configured to perform various operations described herein, and a wireless communication interface for receiving information from and providing information to the remote computing system 502 and the vehicle 508.

[0083] In keeping with the above discussion, a computing system (e.g., remote computing system 502, server computing system 506, or a computing system local to vehicle 508) may operate to capture images of the autonomous vehicle's environment using a camera. Generally, at least one computing system may analyze the images and possibly control the autonomous vehicle.

[0084] In some embodiments, to facilitate autonomous operation, a vehicle (e.g., vehicle 508) may receive data representing objects in the environment in which the vehicle operates (also referred to herein as "environmental data") in various ways. A sensor system on the vehicle may provide the environmental data representing objects in the environment. For example, the vehicle may have various sensors, including cameras. Each of these sensors may communicate environmental data to a processor within the vehicle regarding information each respective sensor receives.

[0085] While operating in autonomous mode, the vehicle may control its operation with little or no human input. For example, a human operator may input an address into the vehicle, and the vehicle may then be able to drive to the specified destination without further input from the human (e.g., the human does not need to turn the steering wheel or touch the brake / accelerator pedals). Additionally, while the vehicle is operating autonomously, the sensor system may receive environmental data. The vehicle's processing system may modify control of the vehicle based on the environmental data received from the various sensors. In some examples, the vehicle may modify the vehicle's speed in response to the environmental data from the various sensors. The vehicle may modify its speed to avoid obstacles, obey traffic laws, etc. When the processing system in the vehicle identifies an object near the vehicle, the vehicle may be able to modify its speed or otherwise alter its behavior.

[0086] To facilitate this, the vehicle may analyze environmental data representative of objects in the environment to determine at least one object having a detection confidence below a threshold. A processor in the vehicle may be configured to detect various objects in the environment based on the environmental data from various sensors. For example, in one embodiment, the processor may be configured to detect objects that may be important for the vehicle to recognize. Such objects may include pedestrians, road signs, other vehicles, indicator signals on other vehicles, distant activities and / or objects on highways, flashing school bus stop signs, flashing lights on emergency vehicles, and various other objects detected within the captured environmental data.

[0087] The processor may be configured to determine a detection confidence. The detection confidence may indicate a likelihood that a determined object is correctly identified or present in the environment. For example, the processor may perform object detection of objects in image data of the received environmental data and determine that at least one object has a detection confidence below a threshold based on failing to identify the object with a detection confidence above a threshold. If the object detection or object recognition results for the object are inconclusive, the detection confidence may be low or below a threshold.

[0088] The processor may be configured to determine an operating mode of the image sensor 510 based on the detection confidence. For example, if at least one object has a detection confidence below a threshold based on failing to identify the object with a detection confidence above the threshold, the image sensor may process an additional ROI including at least one object along with a full-resolution image of the environment. Also, for example, several successive duty cycles may change in detection confidence when the additional ROI including at least one object is processed. A higher detection confidence may be associated with processing fewer ROIs in successive duty cycles, and a lower detection confidence may be associated with processing more ROIs and more full-resolution images in successive duty cycles. VII. Example Peer-to-Peer Sensor Architecture

[0089] 6 illustrates an example sensor architecture in which two or more different sensors are configured to share one or more pieces of ROI information with each other via a peer-to-peer network. In some cases, such sharing via a peer-to-peer network may occur independently and / or without the involvement of a central control system. Specifically, FIG. 6 includes a vehicle 600, which may represent, among other possibilities, an autonomous vehicle (e.g., vehicle 500) or a robotic device. Vehicle 600 may include a vehicle control system 620, a first image sensor(s) 602, and a second image sensor(s) 604.

[0090] Vehicle control system 620 may represent hardware and / or software configured to control operation of vehicle 600 based on data from first image sensor(s) 602 and second image sensor(s) 604. Thus, vehicle control system 620 may be communicatively connected to first image sensor(s) 602 by connection 610 and to second image sensor(s) 604 by connection 612. Each of first image sensor(s) 602 and second image sensor(s) 604 may include a corresponding first control circuit 606 and a second control circuit 608, respectively, configured to process sensor data from the corresponding sensor and to handle communication with vehicle control system 620 and other sensors. In some implementations, first control circuit 606 and second control circuit 608 may be implemented as one or more layers of respective integrated circuits forming the respective sensors (e.g., as shown in FIG. 3 ).

[0091] Furthermore, each of the first image sensor(s) 602 and the second image sensor(s) 604 may be interconnected with each other by a peer-to-peer network connection. Specifically, the first image sensor(s) 602 may be communicatively connected to the second image sensor(s) 604 by a peer-to-peer connection 614.

[0092] Each of first image sensor(s) 602 and second image sensor(s) 604 may be configured to communicate with each other independently of vehicle control system 620, for example, via peer-to-peer connection 614. Thus, first image sensor(s) 602 and second image sensor(s) 604 may share ROI information with each other without involvement by vehicle control system 620. For example, first control circuitry 606 may be configured to (i) acquire full resolution images, (ii) determine information associated with at least one object of interest, (iii) select a second image sensor from the plurality of image sensors based on an attitude of the second image sensor relative to the environment and an expected position of the at least one object of interest in the environment at a second time, (iv) further determine a particular ROI based on the selection of the second image sensor, and (v) transmit an indication of the particular ROI to second control circuitry 608 via peer-to-peer connection 614 between first image sensor(s) 602 and second image sensor(s) 604. The second control circuit 608 may be configured to receive an indication of a particular ROI from the first control circuit 606 and acquire a plurality of ROI sensor data in response to receiving the indication.

[0093] As another example, the first control circuit 606 may be configured to (i) acquire a full-resolution image, (ii) determine information associated with the at least one object of interest, and (iii) broadcast an expected position of the at least one object of interest to the second image sensor(s) 604 via a peer-to-peer connection 614 between the first image sensor(s) 602 and the second image sensor(s) 604. In response to receiving the broadcast, the second control circuit 608 may be configured to determine a specific ROI for the second image sensor(s) 604 that is expected to view the at least one object of interest at a second time point based on an attitude of the second image sensor(s) 604 relative to the environment and the expected position of the at least one object of interest in the environment, and to acquire multiple ROI sensor data in response to determining the specific ROI. In some implementations, the first control circuit 606 may detect at least one object of interest moving in the environment. Thus, the first control circuitry 606 may generate a ROI that includes the detected object of interest and send the ROI to the second image sensor(s) 604 along with a request to capture a larger temporal image of the environment. In some implementations, one or more cameras may have a wider field of view, and such a wider field of view of the environment may not be desirable for image processing. Thus, the image from the camera with the wider field of view may be cropped to include the ROI.

[0094] Although the first image sensor(s) 602 and the second image sensor(s) 604 may communicate with a server (e.g., server computing system 506), it may be desirable to configure the first image sensor(s) 602 and the second image sensor(s) 604 to make certain decisions related to object detection and tracking. For example, if the first image sensor(s) 602 and the second image sensor(s) 604 send their respective full-resolution images and / or ROIs to the server, there would be latency in performing image processing tasks, and the position of the at least one object of interest may not be predicted with high accuracy. Thus, when it is determined that the at least one object of interest is moving in the environment, the first image sensor(s) 602 may be configured to detect the at least one object of interest in a first frame, sample the next frame, and generate a trajectory of the object's movement in the environment. This information, contained in the associated ROI image, may then be sent to the second image sensor(s) 604 for further tracking. For example, the second image sensor(s) 604 may receive information related to the trajectory, capture an image of the environment in which at least one object of interest is predicted to be located based on the trajectory, and process only pixels of the image corresponding to the ROI that includes the predicted location, instead of the full resolution image.

[0095] Such direct peer-to-peer communication of ROI images between sensors may enable the sensors to quickly react to changes in the environment and capture high-quality sensor data useful for determining how to operate the vehicle 600. By avoiding communication through the vehicle control system 620, the transmission path of the information may be shortened, thereby reducing communication delays. Furthermore, by selecting an ROI, detecting at least one object of interest, and / or determining one or more ROIs containing at least one object of interest using the first control circuitry 606 and the second control circuitry 608, the rate at which an ROI is selected, at least one object of interest is detected, additional ROIs are processed, and / or the rate at which an operational mode is determined may be independent of the processing load of the vehicle control system, thereby further reducing delays.

[0096] First image sensor(s) 602 and second image sensor(s) 604 may be configured to communicate directly with each other to share certain information directly, but these sensors may also share information with vehicle control system 620. For example, each sensor may share with vehicle control system 620 at least some of the sensor data itself, as well as the results of various sensor data processing operations performed on the captured sensor data. In general, first image sensor(s) 602 and second image sensor(s) 604 may share information useful for operating vehicle 600 with vehicle control system 620 and may share information with each other directly regarding where and how to capture sensor data. Thus, a peer-to-peer sensor network may enable vehicle 600 to avoid using vehicle control system 620 as a communications intermediary for certain types of communications between first image sensor(s) 602 and second image sensor(s) 604.

[0097] In some implementations, first image sensor(s) 602 may be configured to capture image data of the environment. Based on or in response to capturing the image data, first image sensor(s) 602 may be configured to transmit the image data to vehicle control system 620. This transmission may be performed over connection 610.

[0098] Based on or in response to receiving the image data, vehicle control system 620 may be configured to select an operational mode for first image sensor(s) 602. The selected operational mode may be based on the captured initial image data and may be selected to improve the quality of future images captured of the environment represented by the initial image data. Based on or in response to the selection of the operational mode, vehicle control system 620 may be configured to transmit the operational mode to first image sensor(s) 602. Again, this transmission may be performed over connection 610. In an alternative implementation, rather than relying on vehicle control system 620 to select the operational mode, this operation may be performed by first control circuitry 606 provided as part of first image sensor(s) 602.

[0099] Based on or in response to selecting and / or receiving an operational mode, first image sensor(s) 602 may be configured to adjust the number of ROIs and full resolution images processed. In doing so, first image sensor(s) 602 may capture additional image data that may have higher quality (e.g., have better exposure, magnification, white balance, etc.) than the initial image data captured. Thus, based on or in response to adjusting the operational mode, first image sensor(s) 602 may be configured to capture the additional image data.

[0100] Based on or in response to capturing the additional image data, the first image sensor(s) 602 may be configured to detect at least one object of interest in the additional image data. This detection may be performed by first control circuitry 606 provided as part of the first image sensor(s) 602. Based on or in response to detecting the at least one object of interest, the first image sensor(s) 602 may be configured to select another sensor and / or a ROI for that sensor that is expected to see the at least one object of interest in the future.

[0101] To that end, the first control circuitry 606 of the first image sensor(s) 602 may be configured to predict future positions relative to the vehicle 600 at which the at least one object of interest will be observed based on attributes of the object of interest. The control circuitry may also determine which of the sensors on the vehicle 600 will see (e.g., in the case of fixed sensors) and / or can be repositioned to see (e.g., in the case of sensors with adjustable pose) the at least one object of interest in the future. Additionally, the first control circuitry 606 may determine a particular ROI of the determined sensors in which the at least one object of interest will appear.

[0102] Based on the selection of the sensor and / or its ROI, the first image sensor(s) 602 may be configured to transmit the ROI selection, the at least one attribute of the object of interest, and / or the operational mode used by the first image sensor(s) 602 to the second image sensor(s) 604. Based on or in response to receiving the ROI selection, the at least one attribute of the object of interest, and / or the operational mode, the second image sensor(s) 604 may be configured to determine the operational mode used by the second image sensor(s) 604 to scan the at least one object of interest and adjust its operational mode accordingly. Based on or in response to adjusting the operational mode, the second image sensor(s) 604 may be configured to capture ROI image data of the selected ROI. VIII. Additional Examples of Actions

[0103] 7 shows a flowchart of operations associated with processing an ROI image. The operations may be performed by image sensor 100, image sensor 200, its components, and / or associated circuitry, among other possibilities. However, the operations may also be performed by other types of devices or device subsystems. For example, the processes may be performed by a server device, an autonomous vehicle, and / or a robotic device.

[0104] The embodiments of Figure 7 may be simplified by eliminating any one or more of the features shown therein. Additionally, these embodiments may be combined with any features, aspects, and / or implementations of previous figures or otherwise described herein.

[0105] Block 700 may involve acquiring, by control circuitry, a full resolution image of the environment from an image sensor including a plurality of pixels forming a plurality of regions of interest (ROIs), the image sensor being configured to operate at a frame rate greater than a threshold rate, the full resolution image including each respective ROI of the plurality of ROIs.

[0106] Block 702 may involve selecting, by the control circuitry, a particular ROI based on the full resolution image.

[0107] Block 704 may involve detecting, by the control circuitry, at least one object of interest within the particular ROI.

[0108] Block 706 may involve determining, by the control circuitry, the mode of operation in which subsequent image data generated by the particular ROI will be processed.

[0109] Block 708 may involve processing, by the control circuitry, the image data including multiple ROI images of the detected object of interest based on the operating mode and frame rate.

[0110] In some embodiments, based on the operating mode and frame rate, processing the image data may include generating multiple ROI images of at least one object of interest during one or more subsequent duty cycles instead of acquiring additional full resolution images.

[0111] In some embodiments, based on the operating mode and frame rate, processing the image data may include generating multiple ROIs and additional full resolution images including one or more objects of interest during one or more subsequent duty cycles.

[0112] In some embodiments, based on the operating mode and frame rate, processing the image data may include generating multiple additional full resolution images including one or more objects of interest during one or more subsequent duty cycles.

[0113] In some embodiments, the system may include a vehicle, and the operational mode may be based on the location of an image sensor on the vehicle.

[0114] In some embodiments, the number of the plurality of ROI images of the at least one object of interest may be based on the frame rate.

[0115] In some embodiments, the image sensor may be configured in a predetermined mode of operation.

[0116] In some embodiments, processing the image data may include performing one or more image processing tasks on multiple ROI images of at least one object of interest.

[0117] In some embodiments, the operations may further include comparing a distance between the image sensor and the at least one object of interest represented in the full resolution image to a threshold distance, and the operations also include acquiring multiple ROI images of the at least one object of interest instead of acquiring additional full resolution images based on a result of comparing the distance to the threshold distance.

[0118] In some embodiments, the operations may further include comparing a velocity of the at least one object of interest represented in the full resolution image to a threshold velocity, and the operations also include acquiring multiple ROI images of the at least one object of interest based on a result of comparing the velocity to the threshold velocity instead of acquiring additional full resolution images.

[0119] In some embodiments, the operations may further include providing to the server a plurality of ROI images of the at least one object of interest and a full resolution image including the at least one object of interest.

[0120] In some embodiments, a full-resolution image may be acquired at a first time point, and the operations may further include transmitting image data of at least one object of interest in the environment to a second image sensor. The operations also include determining a specific ROI from the multiple ROIs of the second image sensor. The specific ROI may correspond to an expected position of the at least one object of interest in the environment at a second time point, later than the first time point. The operations also include acquiring multiple ROI sensor data from the specific ROI instead of acquiring full-resolution sensor data including each respective ROI of the multiple ROIs from the second sensor.

[0121] In some embodiments, the first subset of control circuitry may form part of an image sensor. The second subset of control circuitry may form part of a second image sensor. The first subset of control circuitry may be configured to (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, (iii) select a second image sensor from the plurality of image sensors based on an orientation of the second image sensor relative to the environment and an expected location of the at least one object of interest in the environment at a second time, (iv) further determine a specific ROI based on the selection of the second image sensor, and (v) transmit an indication of the specific ROI to the second subset of control circuitry via a peer-to-peer connection between the first image sensor and the second image sensor. The second subset of control circuitry may be configured to receive the indication of the specific ROI from the first subset of control circuitry and acquire the plurality of ROI sensor data in response to receiving the indication.

[0122] In some embodiments, the first subset of control circuitry may form part of the image sensor. The second subset of control circuitry may form part of the second image sensor. The first subset of control circuitry may be configured to (i) acquire full-resolution images, (ii) determine information associated with the at least one object of interest, and (iii) broadcast an expected location of the at least one object of interest to the plurality of image sensors, including the second image sensor, via a peer-to-peer connection between the image sensor and the plurality of image sensors. The second subset of control circuitry may be configured, in response to receiving the broadcast, to determine a particular ROI of the second image sensor that is expected to view the at least one object of interest at a second time based on an attitude of the second image sensor relative to the environment and the expected location of the at least one object of interest within the environment, and, in response to determining the particular ROI, to acquire the plurality of ROI sensor data.

[0123] In some embodiments, the system may include a vehicle. The system may further include a control system configured to control the vehicle based on data generated by the image sensor. The transmission of image data, the determination of a particular ROI, and the acquisition of multiple ROI sensor data may occur independently of the control system.

[0124] In some embodiments, detecting the at least one object of interest may include determining one or more of: (i) a geometric characteristic of the at least one object of interest; (ii) an actual position of the at least one object of interest within the environment; (iii) a velocity of the at least one object of interest; (iv) an optical flow associated with the at least one object of interest; (v) a classification of the at least one object of interest; or (vi) one or more confidence values ​​associated with the results of processing the image data.

[0125] In some embodiments, the control circuitry may implement an artificial neural network configured to process the image data by (i) analyzing the image data using the neural network circuitry, and (ii) generating, by the neural network circuitry, neural network output data related to the results of processing the image data.

[0126] In some embodiments, the control circuitry may be configured to switch between operating in one of three different modes based on conditions in the environment. A first of the three different modes may involve acquiring a first plurality of full-resolution images at a first frame rate. A second of the three different modes may involve alternating between acquiring full-resolution images at the first frame rate and acquiring a plurality of ROI images at a second frame rate that is higher than the first frame rate. A third of the three different modes may involve (i) acquiring a second plurality of full-resolution images at a third frame rate that is higher than the first frame rate, and (ii) combining the second plurality of full-resolution images to generate a stacked full-resolution image. IX. Conclusion

[0127] The present disclosure is not limited with respect to the particular embodiments described in this application, which are intended as illustrations of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its scope. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims.

[0128] The above detailed description explains various features and operations of the disclosed systems, apparatus, and methods with reference to the accompanying drawings. In the figures, like symbols generally identify like components unless context dictates otherwise. The example embodiments described herein and in the figures are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0129] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the figures and discussed herein, each step, block, and / or communication may represent the processing of information and / or the transmission of information according to example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be executed in an order different from that shown or discussed, such as substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer blocks and / or operations may be used in any of the message flow diagrams, scenarios, and flowcharts discussed herein, and these message flow diagrams, scenarios, and flowcharts may be combined with each other, either partially or in whole.

[0130] Steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to perform specific logical operations or actions in the method or technique. The program code and / or associated data may be stored in any type of computer-readable medium, such as a storage device, including random access memory (RAM), a disk drive, a solid-state drive, or another storage medium.

[0131] Computer-readable media may also include non-transitory computer-readable media, such as computer-readable media that store data for a short period of time, such as register memory, processor cache, and RAM. Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for a longer period of time. Thus, computer-readable media may include, for example, secondary or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, solid-state drives, compact disc read-only memories (CD-ROMs), etc. Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, as computer-readable storage media or tangible storage devices.

[0132] Additionally, steps or blocks representing one or more information transmissions may correspond to information transmissions between software and / or hardware modules within the same physical device, although other information transmissions may be between software and / or hardware modules in different physical devices.

[0133] The particular arrangements shown in the figures should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in a given figure. Additionally, some of the illustrated elements may be combined or omitted. Still further, example embodiments may include elements not shown in the figures.

[0134] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

1. an image sensor including a plurality of pixels forming a plurality of regions of interest (ROIs); Image processing resources; A control circuit comprising: acquiring a full resolution image of an environment from the image sensor, the full resolution image including each respective ROI of the plurality of ROIs; detecting at least one object of interest within a particular ROI of the plurality of ROIs based on the full resolution image; determining a plurality of attributes associated with the object of interest based on processing the full resolution image using the image processing resources; selecting an operational mode from a plurality of predetermined operational modes based on the plurality of attributes, each respective operational mode of the plurality of predetermined operational modes indicating a different predetermined sequence of ROI image data and full resolution image data to be acquired as part of one or more subsequent processing cycles; and processing subsequent image data acquired as part of the one or more subsequent processing cycles using the image processing resource based on the operational mode.

2. 2. The system of claim 1, wherein the processing of the subsequent image data based on the operational mode includes generating multiple ROI images of the at least one object of interest instead of acquiring additional full resolution images during the one or more subsequent processing cycles.

3. 2. The system of claim 1, wherein the processing of the subsequent image data based on the operational mode includes generating, during the one or more subsequent processing cycles, multiple ROI images of the at least one object of interest and additional full resolution images including one or more objects of interest.

4. 2. The system of claim 1, wherein the processing of the subsequent image data based on the operational mode includes generating, during the one or more subsequent processing cycles, a plurality of additional full resolution images including one or more objects of interest.

5. The system of claim 1, further comprising a vehicle, wherein the operating mode is further based on the position of the image sensor on the vehicle.

6. The subsequent image data includes a plurality of ROI images of the at least one object of interest; The system of claim 1 , wherein the number of the plurality of ROI images of the at least one object of interest is based on a frame rate of the image sensor.

7. The method of claim 1, wherein the subsequent image data includes a plurality of ROI images of the at least one object of interest; The system of claim 1 , wherein the processing of the subsequent image data includes performing one or more image processing tasks on the plurality of ROI images of the at least one object of interest.

8. The control circuit further configured to compare a distance between the image sensor and the at least one object of interest represented in the full resolution image to a threshold distance; The system of claim 1 , wherein the operational mode is selected based on a comparison of the distance to a threshold distance.

9. The control circuit further configured to compare a velocity of the at least one object of interest represented in the full resolution image to a threshold velocity; The system of claim 1 , wherein the operating mode is selected based on a comparison of the speed to a threshold speed.

10. the full resolution image is acquired at a first time point and the subsequent image data is processed; transmitting image data of the at least one object of interest in the environment to a second image sensor; determining the specific ROI from a plurality of ROIs of the second image sensor, the specific ROI corresponding to an expected location of the at least one object of interest within the environment at a second time point that is later than the first time point; and acquiring multiple ROI sensor data from the particular ROI from the second image sensor based on the operation mode, instead of acquiring full resolution sensor data including each respective ROI of the plurality of ROIs.

11. 11. The system of claim 10, wherein the first subset of control circuitry forms a portion of the image sensor and the second subset of control circuitry forms a portion of the second image sensor, the first subset of control circuitry being configured to: (i) acquire the full resolution image; (ii) determine information associated with the at least one object of interest; (iii) select the second image sensor from a plurality of image sensors based on an attitude of the second image sensor relative to the environment and the expected position of the at least one object of interest within the environment at the second time; (iv) further determine the particular ROI based on the selection of the second image sensor; and (v) transmit an indication of the particular ROI to the second subset of control circuitry via a peer-to-peer connection between the image sensor and the second image sensor, the second subset of control circuitry being configured to receive the indication of the particular ROI from the first subset of control circuitry and acquire the plurality of ROI sensor data in response to receiving the indication.

12. 11. The system of claim 10, wherein the first subset of control circuits forms part of the image sensor, and the second subset of control circuits forms part of the second image sensor, the first subset of control circuits being configured to (i) acquire the full resolution image, (ii) determine information associated with the at least one object of interest, and (iii) broadcast the expected location of the at least one object of interest to a plurality of image sensors including the second image sensor over a plurality of peer-to-peer connections between the image sensor and the plurality of image sensors, the second subset of control circuits being configured, in response to receiving the broadcast, to determine the particular ROI of the second image sensor that is expected to see the at least one object of interest at the second time based on an attitude of the second image sensor relative to the environment and the expected location of the at least one object of interest within the environment, and to acquire the plurality of ROI sensor data in response to determining the particular ROI.

13. A vehicle; 11. The system of claim 10, further comprising: a control system configured to control the vehicle based on data generated by the image sensor, wherein the transmission of the image data, the determination of the particular ROI, and the acquisition of the multiple ROI sensor data are performed independently of the control system.

14. The system of claim 1, wherein the plurality of attributes include two or more of: (i) geometric properties of the at least one object of interest; (ii) actual position of the at least one object of interest within the environment; (iii) velocity of the at least one object of interest; (iv) optical flow associated with the at least one object of interest; (v) classification of the at least one object of interest; or (vi) one or more confidence values ​​associated with the results of the processing of the subsequent image data.

15. the control circuitry includes neural network circuitry, and the subsequent processing of the image data comprises: analyzing the subsequent image data using the neural network circuit; generating, by said neural network circuitry, neural network output data related to results of said processing of said subsequent image data.

16. The subsequent image data includes a plurality of ROI images of the at least one object of interest; 2. The system of claim 1, wherein the full resolution images are acquired at a first frame rate and the plurality of ROI images of the at least one object of interest are acquired at a second frame rate that is higher than the first frame rate.

17. acquiring, by control circuitry, a full resolution image of an environment from an image sensor including a plurality of pixels forming a plurality of regions of interest (ROIs), the full resolution image including each respective ROI of the plurality of ROIs; detecting, by the control circuitry, at least one object of interest within a particular ROI of the plurality of ROIs based on the full resolution image; determining, by the control circuitry, a plurality of attributes associated with the object of interest based on processing the full resolution image; selecting, by the control circuitry, an operational mode from a plurality of predetermined operational modes based on the plurality of attributes, each respective operational mode of the plurality of predetermined operational modes indicating a different predetermined sequence of ROI image data and full resolution image data to be acquired as part of one or more subsequent processing cycles; and processing, by the control circuitry, subsequent image data acquired as part of the one or more subsequent processing cycles based on the operational mode.

18. When executed by a computing device, the computing device: acquiring a full resolution image of an environment from an image sensor including a plurality of pixels forming a plurality of regions of interest (ROIs), the full resolution image including each respective ROI of the plurality of ROIs; detecting at least one object of interest within a particular ROI of the plurality of ROIs based on the full resolution image; determining a plurality of attributes associated with the object of interest based on processing the full resolution image; selecting an operational mode from a plurality of predetermined operational modes based on the plurality of attributes, each respective operational mode of the plurality of predetermined operational modes indicating a different predetermined sequence of ROI image data and full resolution image data to be acquired as part of one or more subsequent processing cycles; and processing subsequent image data acquired as part of the one or more subsequent processing cycles based on the operational mode.

19. 2. The system of claim 1, wherein the image sensor is provided on a first layer of an integrated circuit and the image processing resource is provided on a second layer of the integrated circuit communicatively connected to the first layer.

20. 2. The system of claim 1, wherein the operational mode is selected further based on at least one of traffic conditions, weather conditions, road conditions, road type, landscape type, processing capabilities of the image sensor, and number of other image sensors used in conjunction with the image sensor.

21. The system of claim 1 , wherein the operational mode is selected further based on a field of view of the image sensor relative to a vehicle in which the image sensor is mounted.

Citation Information

Patent Citations

  • Vehicle periphery monitoring device

    JP2010070127A

  • Recording device, recording method, and recording program

    JP2020145516A

  • Sensor device and signal processing method

    WO2020080140A1