Smart sensor implementation of region of interest operating mode
By configuring full-resolution and ROI image sensors on autonomous vehicles and combining them with distributed processing at the integrated circuit layer, the limitation of image sensor data processing capabilities has been solved, enabling efficient image data processing and analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WAYMO LLC
- Filing Date
- 2021-12-03
- Publication Date
- 2026-07-21
AI Technical Summary
Image sensors on autonomous vehicles generate a large amount of image data, which leads to limitations in data transmission bandwidth and processing capabilities, making it difficult to process high frame rate and high resolution images in a timely manner.
The image sensor is configured to generate full-resolution images and region of interest (ROI) images. Distributed image processing is achieved through ROI selective readout and high frame rate processing, combined with an integrated circuit layer, reducing data volume and improving processing efficiency.
By selectively processing ROI images, data transmission and processing requirements are reduced, image analysis speed and accuracy are improved, and system latency and resource consumption are reduced.
Smart Images

Figure CN116615911B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 123440, filed December 16, 2020, which is incorporated herein by reference in its entirety. Background Technology
[0003] An image sensor comprises multiple photosensitive pixels that measure the intensity of light incident upon them, thus collectively capturing an image of the environment. A frame rate can be applied to an image sensor to allow it to generate images of the environment. Image sensors can be used in a variety of applications, such as photography, robotics, and autonomous vehicles. Summary of the Invention
[0004] An image sensor can contain multiple regions of interest (ROIs), each of which can be read out independently of the others. The image sensor can be used to capture full-resolution images including each ROI. Based on the content of the full-resolution images, ROIs can be selected from the images to be acquired. These ROI images can be acquired at a higher frame rate than the full-resolution images. The image sensor can be configured to operate at a frame rate higher than a threshold rate and can process a corresponding number of frames during a given duty cycle. When an object of interest is detected, frames of the ROI image including the object of interest can be processed in one or more subsequent duty cycles, thereby improving image processing speed.
[0005] In a first example embodiment, a system is provided that includes an image sensor having a plurality of pixels forming a plurality of regions of interest (ROIs). The image sensor is configured to operate at a frame rate above a threshold rate. The system also includes image processing resources. The system further includes control circuitry configured to perform operations including acquiring a full-resolution image of the environment from the image sensor. The full-resolution image contains each of the plurality of ROIs. The operation also includes selecting a specific ROI based on the full-resolution image. The operation also includes detecting at least one object of interest within the specific ROI. The operation also includes determining an operation mode to process subsequent image data generated from the specific ROI. The operation also includes processing image data of the plurality of ROI images, including at least one object of interest, based on the operation mode and the frame rate.
[0006] In a second example embodiment, a method is provided that includes acquiring a full-resolution image of an environment from an image sensor comprising a plurality of pixels forming a plurality of regions of interest (ROIs) via control circuitry. The full-resolution image contains each of the plurality of ROIs. The image sensor is configured to operate at a frame rate above a threshold rate. The method further includes selecting a specific ROI based on the full-resolution image. The method further includes detecting at least one object of interest within the specific ROI. The method further includes determining an operating mode for processing subsequent image data generated from the specific ROI. The method further includes processing image data of the plurality of ROI images, each containing at least one object of interest, based on the operating mode and the frame rate.
[0007] In a third example embodiment, a non-transitory computer-readable storage medium is provided, having instructions stored thereon that, when executed by a computing device, cause the computing device to perform operations. The operations include acquiring a full-resolution image of an environment from an image sensor comprising a plurality of pixels forming a plurality of Regions of Interest (ROIs). The full-resolution image contains each corresponding ROI among the plurality of ROIs. The image sensor is configured to operate at a frame rate above a threshold rate. The operations also include selecting a specific ROI based on the full-resolution image. The operations further include detecting at least one object of interest (OI) within the specific ROI. The operations also include determining an operation mode for processing subsequent image data generated from the specific ROI. The operations further include processing image data of the plurality of ROI images, each containing at least one OI, based on the operation mode and the frame rate.
[0008] These and other embodiments, aspects, advantages, and alternatives will become clear to those skilled in the art upon reading the following detailed description and, where appropriate, referring to the accompanying drawings. Furthermore, the overview and other descriptions and drawings provided herein are intended to illustrate embodiments by way of example only, and many variations are possible accordingly. For example, structural elements and process steps may be rearranged, combined, distributed, eliminated, or otherwise altered while remaining within the scope of the claimed embodiments. Attached Figure Description
[0009] Figure 1 A block diagram of an image sensor with three integrated circuit layers according to an example embodiment is shown.
[0010] Figure 2 The arrangement of the region of interest according to an example embodiment is shown.
[0011] Figure 3 An example architecture of an ROI mode processor for an image sensor according to an example embodiment is shown.
[0012] Figure 4 An example ROI pattern processor according to an example embodiment is shown.
[0013] Figure 5 This is a conceptual diagram of wireless communication between various computing systems associated with an autonomous vehicle, based on an example embodiment.
[0014] Figure 6 An example sensor architecture according to an example embodiment is shown, wherein various types of sensors are configured to share one or more ROI information with each other via a peer-to-peer network.
[0015] Figure 7 A flowchart according to an example embodiment is shown. Detailed Implementation
[0016] This document describes example methods, devices, and systems. It should be understood that the terms "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example," "exemplary," and / or "illustrative" is not necessarily to be construed as superior to or more advantageous than other embodiments or features, unless so stated. Therefore, other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter set forth herein.
[0017] Therefore, the exemplary embodiments described herein are not limiting. It will be readily understood that the aspects of this disclosure, as generally described herein and illustrated in the accompanying drawings, can be arranged, replaced, combined, separated, and designed in a variety of different configurations.
[0018] Furthermore, unless the context otherwise requires, the features shown in the figures can be used in combination with each other. Therefore, the figures should generally be considered as aspects of one or more overall embodiments, and it should be understood that not all features shown are necessary for every embodiment.
[0019] Furthermore, the enumeration of any element, block, or step in this specification or claims is for clarity only. Therefore, such enumeration should not be construed as requiring or implying that these elements, blocks, or steps follow a particular arrangement or are performed in a particular order. Unless otherwise stated, the drawings are not to scale.
[0020] I. Overview
[0021] Image sensors can be provided on autonomous vehicles to assist in perceiving and navigating their environment. In some cases, these image sensors may be able to generate more data than the autonomous vehicle's control system can process in a timely manner. This may occur, for example, when a large number of image sensors are provided on the vehicle and / or when each image sensor has a high resolution, resulting in the generation of a large amount of image data. In some cases, the amount of data transmission bandwidth available on the autonomous vehicle and / or the expected data transmission latency may limit the amount of image data that the control system can utilize. In other cases, the amount of processing power provided by the control system may limit the amount of image data that can be processed.
[0022] Generating and processing large amounts of image data is often desirable because it allows for the detection, tracking, classification, and other analyses of objects within an environment. For example, capturing images at a high frame rate allows for the analysis of fast-moving objects by representing them without motion blur. In some implementations, the frame rate can be set above a threshold rate. For example, the threshold rate could be 50 frames per second (fps), and the frame rate could be 50 fps or higher. Similarly, capturing high-resolution images allows for the analysis of distant objects or objects in low-light environments. Notably, in some cases, such objects can be represented in a portion of the entire image generated by the image sensor, rather than occupying the entire image.
[0023] Therefore, the image sensor and corresponding circuitry can be configured to operate in one or more operating modes. For example, the image sensor can be configured to generate full-resolution images (i.e., images containing every ROI) and ROI images containing some objects of interest but not all ROIs. Furthermore, for example, the image sensor can be configured to generate ROI images containing some objects of interest and not generate full-resolution images. In one example, one or more ROI images can be used to determine the velocity of at least one object of interest.
[0024] For example, an image sensor can be configured to operate at a high frame rate of 50 frames per second (fps). Typically, the frame rate determines the number of multiple Regions of Interest (ROI) images for an object. For instance, a 50fps configuration might result in 5 frames being generated every millisecond (msec). These 5 frames can be used to perform image processing tasks such as noise reduction, image enhancement, image sharpening, object detection, object tracking, etc. When an object of interest is detected, an ROI image of the detected object can be generated and further processed. For example, a full-resolution image might be processed in 100 milliseconds, followed by processing of the full-resolution image and one or more ROI images over the next 100 msec. Subsequently, only ROI images might be processed over the next 100 msec, and so on. Such processing cycles can be repeated.
[0025] To this end, an image sensor and corresponding circuitry can be provided that divides the image sensor into multiple Regions of Interest (ROIs) and allows selective readout of image data from individual ROIs. These ROIs allow the image sensor to adapt to high frame rates and / or high-resolution imaging while reducing the total amount of image data generated, transmitted, and analyzed. In other words, the image sensor can generate image data in some ROIs (e.g., ROIs containing some objects of interest), while other ROIs (e.g., ROIs not containing at least one object of interest) are not used to generate image data. Full-resolution images can be analyzed to detect at least one object of interest, and the ROI representing at least one object of interest can then be used to generate multiple ROI images for further analysis.
[0026] In another example, a full-resolution image can be used to determine the distance between an image sensor and at least one object of interest. When this distance crosses a threshold distance (e.g., exceeding or falling below a threshold distance), a Region of Interest (ROI) containing the object can be selected and used to acquire multiple ROI images of the object. For example, when determining that a pedestrian is within a threshold distance of an autonomous vehicle based on a full-resolution image, multiple ROI images of the pedestrian can be acquired. ROIs can be additionally or alternatively selected based on the classification of the object in the full-resolution image or other attributes of the object or the full-resolution image. For example, ROI images of traffic lights or signals can be acquired based on their classification. These and other properties of at least one object of interest in the full-resolution image can be used independently or in combination to select a specific ROI.
[0027] In another example, a full-resolution image can be used to determine the velocity of at least one object of interest represented within the full-resolution image. When this velocity exceeds a threshold velocity (e.g., exceeding or falling below a threshold velocity), an ROI containing the object can be selected and used to acquire multiple ROI images of the object. For example, when determining that an oncoming vehicle is within a threshold velocity based on a full-resolution image, multiple ROI images of a pedestrian can be acquired. ROIs can be additionally or alternatively selected based on the classification of at least one object of interest in the full-resolution image or other attributes of the object or the full-resolution image.
[0028] In some aspects, the detection of at least one object of interest (ROI) may include determining the geometric properties of at least one ROI, the location of at least one ROI within the environment, the change in location of at least one ROI over time (e.g., velocity, acceleration, trajectory), the optical flow associated with at least one ROI, the classification of at least one ROI (e.g., car, pedestrian, vegetation, road, sidewalk, etc.), or one or more confidence values associated with the analysis results of multiple ROI images (e.g., confidence level that the object is another car), and one or more other possibilities. When a new ROI is detected within a different portion of a new full-resolution image, the ROI used to generate the ROI image may change over time.
[0029] Because of their smaller size, ROI images can be generated, transmitted, and analyzed at higher rates than full-resolution images, allowing for, for example, more accurate analysis of high-speed and / or distant objects (e.g., due to the absence of motion blur and the availability of redundant high-resolution data). Specifically, to capture ROI images at higher frame rates, analog-to-digital converters (ADCs) can be reassigned from other ROIs to a selected ROI using a multiplexer. Thus, an ADC originally used to digitize pixels of other ROIs can be combined with an ADC assigned to the selected ROI to digitize the pixels of that selected ROI. Each of these ADCs can operate to digitize multiple rows or columns of the ROI in parallel. For example, four different ADC banks can read out four different columns of the selected ROI in parallel.
[0030] The image sensor and at least some of its associated circuitry can be implemented as an integrated circuit layer. This integrated circuit can be communicatively connected to the central control system of the autonomous vehicle. For example, the first layer of the integrated circuit can implement the pixels of the image sensor, the second layer can implement image processing circuitry (e.g., high dynamic range (HDR) algorithms, ADCs, pixel memories, etc.) configured to process signals generated by the pixels, and the third layer can implement neural network circuitry configured to analyze the signals generated by the image processing circuitry in the second layer for object detection, classification, and other purposes.
[0031] Therefore, in some implementations, the generation and analysis of full-resolution images and ROI images, as well as the selection of ROIs representing at least one object of interest, can be performed by the integrated circuit. Once the full-resolution images and ROI images have been processed and the properties of at least one object of interest have been determined, the properties of at least one object of interest can be sent to the control system of the autonomous vehicle. That is, the full-resolution images and / or ROI images may not be sent to the control system, thereby reducing the amount of bandwidth used and required for communication between the integrated circuit and the control system. It is worth noting that in some cases, a portion of the generated image data (e.g., the full-resolution image) may be sent along with the properties of at least one object of interest.
[0032] In other implementations, the analysis of the full-resolution image and the ROI image, as well as the selection of the ROI representing at least one object of interest, can be performed by the control system. Thus, the control system can enable the integrated circuit to generate the ROI image using the selected region of interest. While this approach may utilize more bandwidth, it may still allow the control system to obtain more images representing the parts of the environment of interest (i.e., the ROI image) rather than fewer full-resolution images representing parts of the environment lacking features of interest.
[0033] The image sensor and its circuitry can be further configured to generate a stacked full-resolution image based on multiple full-resolution images. The stacked full-resolution image can be an HDR image, an image representing multiple focused objects (even if these objects are at different depths), or an image representing some other combination or processing of multiple full-resolution images. Notably, the multiple full-resolution images can be combined by the circuitry, rather than by the vehicle's control system. Therefore, the multiple full-resolution images can be generated at a higher frame rate than would be generated if these images were individually provided to the control system. Because information from multiple such full-resolution images is represented in the stacked image, the stacked image can contain more information than individual full-resolution images.
[0034] II. Example Smart Sensor
[0035] Figure 1 This is a block diagram of an example image sensor 100 with three integrated circuit layers. The image sensor 100 can use the three integrated circuit layers to detect objects. For example, the image sensor 100 can capture an image including a person and output an indication of "person detected". In another example, the image sensor 100 can capture an image and output a portion of the image that includes a vehicle detected by the image sensor 100.
[0036] The three integrated circuit layers include a first integrated circuit layer 110, a second integrated circuit layer 120, and a third integrated circuit layer 130. The first integrated circuit layer 110 is stacked on the second integrated circuit layer 120, and the second integrated circuit layer 120 is stacked on the third integrated circuit layer 130. The first integrated circuit layer 110 can electrically communicate with the second integrated circuit layer 120. For example, the first integrated circuit layer 110 and the second integrated circuit layer 120 can be physically connected to each other using interconnects. The second integrated circuit layer 120 can electrically communicate with the third integrated circuit layer 130. For example, the second integrated circuit layer 120 and the third integrated circuit layer 130 can be physically connected to each other using interconnects.
[0037] The first integrated circuit layer 110 may have the same area as the second integrated circuit layer 120. For example, the first integrated circuit layer 110 and the second integrated circuit layer 120 may have the same length and width, but different heights. The third integrated circuit layer 130 may have a larger area than the first and second integrated circuit layers 110 and 120. For example, the third integrated circuit layer 130 may have a length and width that are both 20% larger than those of the first and second integrated circuit layers 110 and 120.
[0038] The first integrated circuit layer 110 may include a pixel sensor array, which is grouped into pixel sensor groups according to their positions (each pixel sensor group is located in...). Figure 1 These are referred to as "pixel groups" 112A-112C (collectively referred to as 112). For example, the first integrated circuit layer 110 may include a 6400×4800 pixel sensor array grouped into 320×240 pixel sensor groups, wherein each pixel sensor group includes an array of 20×20 pixel sensors. Pixel sensor groups 112 may be further grouped to define ROIs.
[0039] Each pixel sensor group 112 may include 2×2 pixel sensor subgroups. For example, each pixel sensor group of a 20×20 pixel sensor may include 10×10 pixel sensor subgroups, wherein each pixel sensor subgroup includes a red pixel sensor in the upper left, a green pixel sensor in the lower right, a first transparent pixel sensor in the lower left, and a second transparent pixel sensor in the upper right. Each subgroup is also referred to as a red-transparent-transparent-green (RCCG) subgroup.
[0040] In some implementations, the size of the pixel sensor group can be selected to improve silicon utilization. For example, the size of the pixel sensor group can allow more silicon to be covered by pixel sensor groups having the same pixel sensor pattern.
[0041] The second integrated circuit layer 120 may include image processing circuit groups (each image processing circuit group in...) Figure 1The image processing circuits 122A-122C (collectively referred to as 122) are referred to as "processing groups". For example, the second integrated circuit layer 120 may include 320 by 240 image processing circuit groups. The image processing circuit groups 122 may be configured to each receive pixel information from a corresponding pixel sensor group and are also configured to perform image processing operations on the pixel information to provide processed pixel information during the operation of the image sensor 100.
[0042] In some implementations, each image processing circuit group 122 may receive pixel information from a single corresponding pixel sensor group 112. For example, image processing circuit group 122A may receive pixel information from pixel sensor group 112A instead of any other pixel group, and image processing circuit group 122B may receive pixel information from pixel sensor group 112B instead of any other pixel group.
[0043] In some implementations, each image processing circuit group 122 may receive pixel information from multiple corresponding pixel sensor groups 112. For example, image processing circuit group 122A may receive pixel information from both pixel sensor groups 112A and 112B, but not from other pixel groups, and image processing circuit group 122B may receive pixel information from pixel group 112C and another pixel group, but not from other pixel groups.
[0044] Enabling the image processing circuit group 122 to receive pixel information from the corresponding pixel group allows the pixel information to be quickly transmitted from the first integrated circuit layer 110 to the second layer 120, because the image processing circuit group 122 can be physically close to the corresponding pixel sensor group 112. The longer the information transmission distance, the longer the transmission time. For example, the pixel sensor group 112A can be directly above the image processing circuit group 122A, and the pixel sensor group 112A does not have to be directly above the image processing circuit group 122C. Therefore, if there is an interconnection between the pixel sensor group 112A and the image processing circuit group 122C, transmitting pixel information from the pixel sensor group 112A to the image processing circuit group 122A can be faster than transmitting pixel information from the pixel sensor group 112A to the image processing circuit group 122C.
[0045] Image processing circuit group 122 can be configured to perform image processing operations on pixel information received from pixel group 112A. For example, image processing circuit group 122A can perform high dynamic range fusion on pixel information from pixel sensor group 112A, and image processing circuit group 122B can perform high dynamic range fusion on pixel information from pixel sensor group 112B. Other image processing operations may include, for example, analog-to-digital conversion and de-mosaicing.
[0046] By enabling the image processing circuit group 122 to perform image processing operations on pixel information from the corresponding pixel sensor group 112, the image processing circuit group 122 can perform image processing operations in a distributed and parallel manner. For example, while the image processing circuit group 122B performs image processing operations on pixel information from the pixel group 122B, the image processing circuit group 122A can perform image processing operations on pixel information from the pixel sensor group 112A.
[0047] The third integrated circuit layer 130 may include neural network circuit groups 132A-132C (each neural network circuit group is...) Figure 1 The three integrated circuit layers 130 and 132A-132C (collectively referred to as 132) are called "NN groups" and 134 are full-image neural network circuits. For example, the third integrated circuit layer 130 may include 320 by 240 neural network circuit groups.
[0048] The neural network circuit groups 132 can be configured to each receive processed pixel information from a corresponding image processing circuit group, and are also configured to perform object detection analysis on the processed pixel information during operation of the image sensor 100. In some embodiments, the neural network circuit groups 132 can each implement a convolutional neural network (CNN).
[0049] In some implementations, each neural network circuit group 132 may receive processed pixel information from a single corresponding image processing circuit group 122. For example, neural network circuit group 132A may receive processed pixel information from image processing circuit group 122A instead of from any other image processing circuit group, and neural network circuit group 132B may receive processed pixel information from image processing circuit group 122B instead of from any other image processing circuit group.
[0050] In some implementations, each neural network circuit group 132 may receive processed pixel information from multiple corresponding image processing circuit groups 122. For example, neural network circuit group 132A may receive processed pixel information from both image processing circuit groups 122A and 122B, but not from other image processing circuit groups, and neural network circuit group 132B may receive processed pixel information from image processing circuit group 122C and another pixel group, but not from other pixel groups.
[0051] By enabling the neural network circuit group 132 to receive processed pixel information from the corresponding image processing circuit group, the processed pixel information can be quickly transmitted from the second integrated circuit layer 120 to the third integrated circuit layer 130, because the neural network circuit group 132 can be physically close to the corresponding image processing circuit group 122. Similarly, the longer the information transmission distance, the longer the transmission time. For example, the image processing circuit group 122A can be directly above the neural network circuit group 132A. Therefore, if there is an interconnection between the image processing circuit group 122A and the neural network circuit group 132C, transmitting processed pixel information from the image processing circuit group 122A to the neural network circuit group 132A can be faster than transmitting processed pixel information from the image processing circuit group 122A to the neural network circuit group 132C.
[0052] The neural network circuit group 132 can be configured to detect objects from processed pixel information received from the image processing circuit group 122. For example, the neural network circuit group 132A can detect objects from processed pixel information from the image processing circuit group 122A, and the neural network circuit group 132B can detect objects from processed pixel information from the image processing circuit group 122B.
[0053] The neural network circuit group 132 detects objects from the processed pixel information from the corresponding image processing circuit group 122, enabling each neural network circuit group 132 to perform detection in parallel in a distributed manner. For example, while neural network circuit group 132B can detect objects from the processed pixel information from image processing circuit group 122B, neural network circuit group 132A can detect objects from the processed pixel information from image processing circuit group 122A.
[0054] In some implementations, the neural network circuitry 132 can perform intermediate processing. Therefore, the image sensor 100 can use three integrated circuit layers 110, 120, and 130 to perform some intermediate processing and output only intermediate results. For example, the image sensor 100 can capture an image including a person and output an indication of a "region of interest" in a given area of the image, without classifying at least one object of interest (a person). Other processing performed outside the image sensor 100 can classify the region of interest as a person.
[0055] Therefore, the output from image sensor 100 may include some data representing the output of some convolutional neural network. This data may be difficult to decipher on its own, but once it continues to be processed outside of image sensor 100, it can be used to classify the region as including people. This hybrid approach may have the advantage of reducing the required bandwidth. Thus, the output from neural network circuitry 132 may include one or more of the following: a region of interest representing the selection of detected pixels, metadata containing temporal and geometrical location information, intermediate computation results prior to object detection, statistics regarding the network's deterministic level, and the classification of the detected objects.
[0056] In some implementations, the neural network circuit group 132 can be configured to implement a CNN with high recall and low precision. Each neural network circuit group 132 can output a list of detected objects, the location of the detected objects, and the timing of object detection.
[0057] The full-image neural network circuit 134 can be configured to receive data indicating objects detected by each neural network circuit group 132, and to detect objects from the data. For example, the neural network circuit group 132 may not be able to detect objects captured by multiple pixel groups because each individual neural network circuit group may only receive pixel information corresponding to a portion of the object being processed. However, the full-image neural network circuit 134 can receive data from multiple neural network circuit groups 132, and is therefore able to detect objects sensed by multiple pixel groups. In some embodiments, the full-image neural network circuit 134 can implement a recurrent neural network (RNN). The neural network can be configurable in terms of both its architecture (number and type of layers, activation functions, etc.) and the actual values of the neural network components (e.g., weights, biases, etc.).
[0058] In some implementations, having the image sensor 100 perform processing can simplify the processing pipeline architecture, provide higher bandwidth and lower latency, allow selective frame rate operation, reduce costs using a stacked architecture, provide higher system reliability because the integrated circuit can have fewer potential points of failure, and provide significant cost and power savings in computing resources.
[0059] III. Example ROI Layout and Hardware
[0060] Figure 2An example arrangement of Regions of Interest (ROIs) on an image sensor is shown. Specifically, image sensor 200 may include pixels forming C columns and R rows. Image sensor 200 may correspond to a first integrated circuit layer 110. The image sensor may be divided into eight ROIs, including ROI 0, ROI 1, ROI 2, ROI 3, ROI 4, ROI 5, ROI 6, and ROI 7 (i.e., ROI 0-7), each ROI comprising m columns of pixels and n rows of pixels. Therefore, C = 2m, R = 4n. In some embodiments, each ROI may include multiple pixel groups 112. Alternatively, pixel groups 112 may be resized and arranged such that each pixel group is also an ROI. In some embodiments, ROIs may be arranged such that they do not collectively span the entire area of image sensor 200. Therefore, the combined area of ROIs may be smaller than the area of image sensor 200. Thus, a full-resolution image may have a higher pixel count than the combined area of ROIs.
[0061] Figure 2 The diagram shows ROIs arranged in two columns, with even-numbered ROIs on the left and odd-numbered ROIs on the right. However, in other embodiments, ROIs and their numbers can be arranged differently. For example, ROIs 0-3 could be in the left column, while ROIs 4-7 could be in the right column. In another example, the image sensor 200 can be divided into eight columns arranged in a single row, with ROIs numbered 0-7 arranged from left to right along the eight columns. In some embodiments, ROIs can be fixed in a given arrangement. Alternatively, ROIs can be reconfigurable. That is, the number of ROIs, the position of each ROI, and the shape of each ROI can be reconfigured.
[0062] IV. Example Architecture of Operation Mode
[0063] Figure 3 An example architecture of the ROI pattern processor of image sensor 200 is shown. Specifically, image sensor 200 may include ROI pattern processor 302, pixel groups 310 and 312 to 314 defining ROIs 0-7 respectively (i.e., pixel groups 310-314), and multiple image processing resources. The image processing resources include pixel-level processing circuits 320, 322, and 324 to 330 (i.e., pixel-level processing circuits 320-330), machine learning circuits 340, 342, and 344 to 350 (i.e., machine learning circuits 340-350), and communication connections 316, 332, and 352. Image sensor 200 may be configured to provide image data to control system 360, which may also be considered part of the image processing resources. Control system 360 may represent a combination of hardware and software configured to generate operations for robotic devices or autonomous vehicles, etc.
[0064] Pixel groups 310-314 represent pixel groups constituting image sensor 200. In some embodiments, each of pixel groups 310-314 may correspond to one or more pixel sensor groups 112. Pixel groups 310-314 may represent circuitry disposed in the first integrated circuit layer 110. The number of pixel sensor groups represented by each of pixel groups 310-314 may depend on the size of each of ROIs 0-7. In embodiments where the number, size, and / or shape of ROIs can be reconfigured, a subset of pixel sensor groups 112 constituting each pixel group 310-314 may vary over time based on the number, size, and / or shape of ROIs.
[0065] Pixel-level processing circuits 320-330 represent circuitry configured to perform pixel-level image processing operations. Pixel-level processing circuits 320-330 can operate on the output generated by pixel groups 310-314. Pixel-level operations may include analog-to-digital conversion, de-mosaicing, high dynamic range fusion, image sharpening, filtering, edge detection, and / or thresholding. Pixel-level operations may also include other types of operations not performed by a machine learning model (e.g., a neural network) provided on the image sensor 200 or by the control system 360. In some embodiments, each of the pixel-level processing circuits 320-330 may include one or more processing groups 122, as well as other circuitry configured to perform pixel-level image processing operations. Therefore, pixel-level circuits 320-330 may represent circuitry disposed in a second integrated circuit layer 120.
[0066] Machine learning circuits 340-350 may include circuitry configured to perform operations associated with one or more machine learning models. Machine learning circuits 340-350 may operate on the outputs generated by pixel groups 310-314 and / or pixel-level processing circuits 320-330. In some embodiments, each of machine learning circuits 340-350 may correspond to one or more neural network groups 132 and / or full-image neural network circuits 132, as well as other circuitry implementing the machine learning models. Therefore, machine learning circuits 340-350 may represent circuitry disposed in a third integrated circuit layer 130.
[0067] Communication connection 316 can represent the electrical interconnection between pixel-level processing circuits 320-330 and pixel groups 310-314. Similarly, communication connection 332 can represent (i) the electrical interconnection between machine learning circuits 340-350 and (ii) the electrical interconnection between pixel-level processing circuits 320-330 and / or pixel groups 310-314. Furthermore, communication connection 352 can represent the electrical interconnection between machine learning circuits 340-350 and control system 360. Communication connections 316, 332, and 352 can be considered subsets of image processing resources, at least because these connections (i) facilitate data transfer between circuits configured to process image data, and (ii) can be modified over time to transfer data between different combinations of circuits configured to process image data.
[0068] In some implementations, communication connection 316 may represent an electrical interconnection between the first integrated circuit layer 110 and the second integrated circuit layer 120, and communication connection 332 may represent an electrical interconnection between the second integrated circuit layer 120 and the third integrated circuit layer 130. Communication connection 352 may represent an electrical interconnection between the third integrated circuit layer 130 and one or more circuit boards through which the image sensor 200 is connected to the control system 360. Each of communication connections 316, 332, and 352 may be associated with a corresponding maximum bandwidth.
[0069] Processing full-resolution images typically requires more computational resources than processing Regions of Interest (ROIs). At higher frame rates, the number of frames available for processing is likely to be high. This increase in the number of frames leads to increased image processing capabilities. For example, with an increased number of frames available for processing, the likelihood of ROIs including objects of interest (ROIs) is higher. Therefore, once these specific ROIs are identified, focus can be placed on the ROIs. For example, ROIs including these ROIs can also be identified, and additional ROIs including these ROIs can be obtained using a higher frame rate. In some implementations, these additional ROIs including ROIs can be obtained at a rate of 150 frames per second.
[0070] Processing images at higher frame rates can lead to several associated resource cost considerations. For example, processing a large number of frames may result in allocating more memory resources to store frames and / or processing results. This can lead to higher latency and potentially wasted computing power. Focusing on additional ROIs, including the object of interest, allows sensors to mitigate some of these resource allocation factors and enables intelligent utilization of available resources.
[0071] Therefore, the ROI mode processor 302 can be configured to dynamically utilize image processing resources 316, 320-330, 340-350, 332, 352 and / or 360 (i.e., image processing resources 316-360) available to the image sensor 200. Specifically, some image sensors can be configured to have a fixed ROI mode processor 302. For example, in some image sensor 200s, the ROI mode processor 302 can be configured to be in a first operating mode to process ROIs including objects of interest and not to process full-resolution images. Furthermore, for example, in some other image sensor 200s, the ROI mode processor 302 can be configured to be in a second operating mode to process ROIs including objects of interest along with full-resolution images. However, in some image sensor 200s, the ROI mode processor 302 can be configured to dynamically switch between one or more operating modes (e.g., a first operating mode and a second operating mode).
[0072] The configuration of the ROI mode processor 302 can depend on several factors. For example, the configuration can depend on the type of image capture device in the image sensor 200. In some embodiments, the configuration can depend on the field of view of the image capture device. For example, when the image sensor 200 is mounted on a vehicle, the image capture device can have a front view, a side view, a rear view, etc. In some aspects, for an image sensor 200 with an image capture device having a front view, the ROI mode processor 302 can be configured to operate in a second operating mode to continue capturing and processing full-resolution images to detect ROIs and objects of interest. However, for an image sensor 200 with an image capture device having a side view and / or a rear view, the ROI mode processor 302 can be configured to operate in a first operating mode to continue capturing and processing ROI images of detected objects and enable object tracking features.
[0073] In some implementations, the configuration of the ROI mode processor 302 can depend on the time of day and / or the intensity of ambient lighting. For example, during the daytime when object detection may be less challenging, the ROI mode processor 302 can be configured to operate in a mode that processes a large number of full-resolution images. However, at night, in foggy conditions, or during rainy conditions, when the intensity of ambient lighting may be low and detecting ROIs and / or objects of interest may be more challenging, the ROI mode processor 302 can be configured to operate in a mode that processes a large number of ROIs including objects of interest, and intermittently process full-resolution images.
[0074] In some implementations, the configuration of the ROI pattern processor 302 may depend on the type of object of interest. For example, the size, type, and / or speed of at least one object of interest may require processing different sets of images. For instance, different operating modes may be used to detect a SUV traveling at a first speed and a motorcycle traveling at a second speed. Typically, smaller objects and / or objects traveling at higher speeds may require faster processing, and the ROI pattern processor 302 may be configured to process a large number of ROIs that only include objects of interest.
[0075] The ROI pattern processor 302 can also be configured to determine the number and / or type of ROI images and / or full-resolution images. For example, the ROI pattern processor 302 can be configured to determine whether it is necessary to process a single full-resolution image with 10 ROIs, or 3 single full-resolution images and 3 ROIs including objects of interest. For example, when at least one object of interest is a pedestrian crossing the street, the ROI pattern processor 302 can intelligently decide to process a larger number of ROIs including at least one object of interest, thereby enhancing object tracking. Although processing more images may result in a higher number of false positives, this is desirable for minimizing errors in detecting and / or tracking pedestrians.
[0076] As described in this article, additional factors, including traffic conditions, weather conditions, road conditions, road construction sites, speed limits, road type, landscape type (e.g., urban, rural), the number of sensors on the vehicle, the availability of memory allocation, and the processing power of the image sensor, can enable the ROI mode processor 302 to dynamically switch between one or more operating modes.
[0077] The ROI pattern processor 302 can communicatively connect to each of the image processing resources 316-360 and can understand the capabilities of each of the image processing resources 316-360, the workload assigned to each of the image processing resources 316-360, and / or the features detected by each of the image processing resources 316-360 (e.g., can receive, access, and / or store their representations). Therefore, in some embodiments, the ROI pattern processor 302 can be configured to distribute image data among the image processing resources 316-360 in a manner that improves or minimizes the latency between acquiring and processing image data, improves or maximizes the utilization of the image processing resources 316-360, and / or improves or maximizes the throughput of image data through the image processing resources 316-360. These objectives can be quantified by one or more objective functions, each of which can be minimized or maximized (e.g., globally or locally) to achieve the corresponding objective.
[0078] V. Example Operation Mode
[0079] Figure 4 An example ROI mode processor 402 is illustrated. As shown in header column 402, the first row (topmost) of diagram 400 indicates a first operating mode of the image processor, the second row indicates a second operating mode of the image processor, the third row indicates a third operating mode of the image processor, and the fourth row (bottommost) indicates the amount of time dedicated to each operation in a duty cycle. In some embodiments, the duty cycle may last 100 milliseconds (ms). The image processor may collectively represent operations performed by a second integrated circuit layer 120, a third integrated circuit layer 130, and any other control system communicatively connected to the image sensor 100 or 200.
[0080] In interval 404, which can last for 100 ms, the image processor operating in the first operating mode can process multiple ROIs and full-resolution images. In intervals 406 and 408, each lasting for 100 ms, the image processor can process only the ROI of interest without processing the full-resolution image. In intervals 410, 412, and 414, each lasting for 100 ms, the image processor can repeat the operations performed in intervals 404, 406, and 408. Processing multiple ROIs allows the image processor to determine various properties of the environmental content represented by these images.
[0081] In interval 404, which can last for 100 ms, the image processor operating in the second operating mode can process multiple ROIs and full-resolution images. In interval 406, which can last for 100 ms, the image processor can process only the ROI of the object of interest without processing the full-resolution image. In intervals 406 and 410, which can each last for 100 ms, the image processor can repeat the operations performed in intervals 404 and 406. Similarly, in intervals 412 and 414, which can each last for 100 ms, the image processor can repeat the operations performed in intervals 404 and 406, and then in intervals 408 and 410. Processing full-resolution images at frequent intervals allows the image processor to determine various properties of the environmental content represented by these images, especially when detecting objects from full-resolution images can be challenging.
[0082] During an interval 404 that can last for 100 ms, the image processor operating in the third operating mode can again acquire the full-resolution image. During an interval 406 that can last for 100 ms, the image processor can process the full-resolution image. In some aspects, the full-resolution image can be a stacked image including one or more detected objects. During an interval 408 that can last for 100 ms, the image processor can acquire the ROI, and during an interval 410 that can last for 100 ms, the image processor can detect the object of interest. During intervals 406 and 408, each lasting for 100 ms, the image processor can process only the ROI of the object of interest without processing the full-resolution image.
[0083] It is worth noting that at interval 404, the image processor can be configured to process ROI images captured during the preceding period (not shown). It is also worth noting that the duration of each interval and the number of ROI images captured for each full-resolution image can vary. For example, some tasks may involve capturing more ROI images than shown (e.g., 16 ROI images per full-resolution image) or fewer ROI images (e.g., 4 ROI images per full-resolution image). Furthermore, among other factors, the size of each ROI and / or the number of ADCs provided to the image sensor can be used to determine the length of the interval for capturing ROI images.
[0084] Typically, multiple images are generated, and the image processor can be configured to perform temporal processing on these images to track objects of interest. For example, the image processor can detect at least one object of interest moving along a trajectory, and the neural network circuitry 132A-132C can oversample another ROI on that trajectory. There may be multiple ways to predict the possible location of at least one object of interest at a given time, so the image processor can focus processing along those expected trajectories. A larger number of images yields more information about the ROI and / or the object of interest. In some implementations, multiple images allow the neural network circuitry 132A-132C to correlate information to make more accurate predictions. For example, temporal events that may be outside the expected norm can be processed, and the certainty of the presence of objects and subsequent object identification can be enhanced.
[0085] VI. Example Vehicle System
[0086] Figure 5 This is a conceptual diagram of wireless communication between various computing systems associated with an autonomous vehicle according to an example embodiment. Specifically, wireless communication can occur between remote computing system 502 and vehicle 508 via network 504. Wireless communication can also occur between server computing system 506 and remote computing system 502, and between server computing system 506 and vehicle 508.
[0087] Example vehicle 508 includes an image sensor 510. The image sensor 510 is mounted on top of vehicle 508 and includes one or more image capturing devices configured to detect information about the environment surrounding vehicle 508 and output an indication of that information. For example, image sensor 510 may include one or more cameras. Image sensor 510 may include one or more movable bases operable to adjust the orientation of one or more cameras in image sensor 510. In one embodiment, the movable base may include a rotating platform that can scan the cameras to obtain information from every direction around vehicle 508. In another embodiment, the movable base of image sensor 510 may be movable in a scanning manner within a specific angular and / or azimuth range. Image sensor 510 may be mounted on top of top, although other mounting locations are also possible.
[0088] Furthermore, the cameras of the image sensor 510 can be distributed in different locations and do not need to be located in a single location. Additionally, each camera of the image sensor 510 can be configured to move or scan independently of the other cameras of the image sensor 510.
[0089] Remote computing system 502 can represent any type of device associated with remote assistance technology, including but not limited to the devices described herein. In the examples, remote computing system 502 can represent any type of device configured to (i) receive information related to vehicle 508, (ii) provide an interface through which a human operator can perceive the information and input a response related to the information, and (iii) send the response to vehicle 508 or other devices. Remote computing system 502 can take various forms, such as workstations, desktop computers, laptops, tablets, mobile phones (e.g., smartphones), and / or servers. In some examples, remote computing system 502 may include multiple computing devices operating together in a network configuration.
[0090] The remote computing system 502 may include a processor configured to perform the various operations described herein. In some embodiments, the remote computing system 502 may also include a user interface including input / output devices such as a touchscreen and speakers. Other examples are also possible.
[0091] Network 504 represents the infrastructure that enables wireless communication between remote computing system 502 and vehicle 508. Network 504 also enables wireless communication between server computing system 506 and remote computing system 502, as well as between server computing system 506 and vehicle 508.
[0092] The location of the remote computing system 502 can vary in the examples. For instance, the remote computing system 502 may be located remotely from the vehicle 508, which has wireless communication via network 504. In another example, the remote computing system 502 may correspond to a computing device within the vehicle 508, which is separate from the vehicle 508, but which can be used by an operator to interact with passengers or the driver of the vehicle 508. In some examples, the remote computing system 502 may be a computing device with a touchscreen that can be operated by passengers of the vehicle 508.
[0093] In some embodiments, the operations performed by the remote computing system 502 as described herein may be performed additionally or alternatively by the vehicle 508 (i.e., by any system or subsystem of the vehicle 508). In other words, the vehicle 508 may be configured to provide a remote assistance mechanism that the driver or passengers of the vehicle can use to interact.
[0094] Server computing system 506 can be configured to wirelessly communicate with remote computing system 502 and vehicle 508 via network 504 (or possibly directly with remote computing system 502 and / or vehicle 508). Server computing system 506 can represent any computing device configured to receive, store, determine, and / or transmit information related to vehicle 508 and its remote assistance. Thus, server computing system 506 can be configured to perform any operation or part of such operation, which is described herein as being performed by remote computing system 502 and / or vehicle 508. Some embodiments of wireless communication related to remote assistance may utilize server computing system 506, while others may not.
[0095] Server computing system 506 may include one or more subsystems and components similar to or the same as those of remote computing system 502 and / or vehicle 508, such as processors configured to perform the various operations described herein, and wireless communication interfaces for receiving and providing information to remote computing system 502 and vehicle 508.
[0096] Based on the above discussion, a computing system (e.g., a remote computing system 502, a server computing system 506, or a computing system local to the vehicle 508) can operate to use a camera to capture environmental images of the autonomous vehicle. Typically, at least one computing system will be able to analyze the images and potentially control the autonomous vehicle.
[0097] In some embodiments, to facilitate autonomous operation, a vehicle (e.g., vehicle 508) can receive data (also referred to herein as “environmental data”) representing objects in the environment in which the vehicle operates, in a variety of ways. Sensor systems on the vehicle can provide environmental data representing environmental objects. For example, the vehicle may have various sensors, including cameras. Each of these sensors can communicate environmental data to a processor within the vehicle regarding information received by each respective sensor.
[0098] When operating in autonomous mode, a vehicle can be controlled with little or no human input. For example, a human operator can input an address into the vehicle, which may then be able to drive to the designated destination without further human input (e.g., the person does not need to steer or touch the brake / accelerator pedal). Furthermore, when the vehicle is operating autonomously, sensor systems may be receiving environmental data. The vehicle's processing system can modify the vehicle's control based on the environmental data received from various sensors. In some examples, the vehicle may change its speed in response to environmental data from various sensors. The vehicle can change its speed to avoid obstacles, obey traffic regulations, etc. When the processing system in the vehicle identifies an object near the vehicle, the vehicle may be able to change its speed or modify its motion in another way.
[0099] To facilitate this, the vehicle can analyze environmental data representing objects in the environment to identify at least one object with a detection confidence level below a threshold. A processor within the vehicle can be configured to detect various objects in the environment based on environmental data from various sensors. For example, in one embodiment, the processor can be configured to detect objects that may be important for vehicle identification. Such objects may include pedestrians, street signs, other vehicles, indicator signals on other vehicles, long-distance activity and / or objects on highways, flashing school bus stop signs, hazard lights on emergency vehicles, and various other objects detected in the captured environmental data.
[0100] The processor can be configured to determine a detection confidence level. A detection confidence level indicates the likelihood that a identified object is correctly identified or exists in the environment. For example, the processor can perform object detection on objects within image data in received environmental data, and determine that at least one object has a detection confidence level below a threshold, based on the fact that no object with a detection confidence level above a threshold can be identified. If the result of object detection or object recognition is uncertain, the detection confidence level may be low or below the threshold.
[0101] The processor can be configured to determine the operating mode of the image sensor 510 based on the detection confidence level. For example, when an object with a detection confidence level above a threshold cannot be identified, and at least one object has a detection confidence level below a threshold, the image sensor can process an additional ROI including at least one object, as well as a full-resolution image of the environment. Furthermore, for example, when processing an additional ROI including at least one object, the number of consecutive duty cycles can vary with the detection confidence level. Higher detection confidence may be associated with processing fewer ROIs in consecutive duty cycles, while lower detection confidence may be associated with processing more ROIs and more full-resolution images in consecutive duty cycles.
[0102] VII. Example Peer-to-Peer Sensor Architecture
[0103] Figure 6 An example sensor architecture is shown, in which various two or more sensors are configured to share one or more ROI information with each other via a peer-to-peer network. In some cases, such sharing via a peer-to-peer network can be performed independently and / or without involving a central control system. Specifically, Figure 6 This includes a vehicle 600, which can represent an autonomous vehicle (e.g., vehicle 500) or a robotic device. Vehicle 600 may include a vehicle control system 620, a first image sensor 602, and a second image sensor 604.
[0104] The vehicle control system 620 can represent hardware and / or software configured to control the operation of the vehicle 600 based on data from the first image sensor 602 and the second image sensor 604. Therefore, the vehicle control system 620 can be communicatively connected to the first image sensor 602 via connection 610 and to the second image sensor 604 via connection 612. Each of the first image sensor 602 and the second image sensor 604 may respectively include a corresponding first control circuit 606 and a second control circuit 608, configured to process sensor data from the corresponding sensor and to handle communication with the vehicle control system 620 and other sensors. In some embodiments, the first control circuit 606 and the second control circuit 608 may be implemented as one or more layers forming the corresponding integrated circuits of the respective sensors (e.g., such as...). Figure 3 (As shown).
[0105] Furthermore, each of the first image sensor 602 and the second image sensor 604 can be interconnected via a peer-to-peer network. Specifically, the first image sensor 602 can be communicatively connected to the second image sensor 604 via a peer-to-peer connection 614.
[0106] Each of the first image sensor 602 and the second image sensor 604 can be configured to communicate with each other independently of the vehicle control system 620, for example, via a peer-to-peer connection 614. Therefore, the first image sensor 602 and the second image sensor 604 can share ROI information with each other without involving the vehicle control system 620. For example, the first control circuit 606 can be configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, (iii) select a second image sensor from a plurality of image sensors based on the pose of the second image sensor relative to the environment and the expected location of at least one object of interest within the environment at a second time, (iv) further determine a specific ROI based on the selection of the second image sensor, and (v) send an indication of the specific ROI to the second control circuit 608 via the peer-to-peer connection 614 between the first image sensor 602 and the second image sensor 604. The second control circuit 608 can be configured to receive the indication of the specific ROI from the first control circuit 606 and, in response to the receipt of the indication, acquire multiple ROI sensor data.
[0107] As another example, the first control circuit 606 may be configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, and (iii) broadcast the expected location of at least one object of interest to the second image sensor 604 via a peer-to-peer connection 614 between the first image sensor 602 and the second image sensor 604. The second control circuit 608 may be configured to, in response to receiving the broadcast, determine a specific ROI of the second image sensor 604 that is expected to view at least one object of interest at a second time, based on the pose of the second image sensor 604 relative to the environment and the expected location of at least one object of interest within the environment, and acquire multiple ROI sensor data in response to determining the specific ROI. In some embodiments, the first control circuit 606 may detect at least one object of interest moving within the environment. Therefore, the first control circuit 606 may generate an ROI including the detected object of interest and send the ROI, along with a request to capture a larger temporal image of the environment, to the second image sensor 604. In some embodiments, one or more cameras may have a wider field of view, and such a wide field of view of the environment may be undesirable for image processing. Therefore, images from cameras with a wider field of view can be cropped to include the ROI.
[0108] Although the first image sensor 602 and the second image sensor 604 can communicate with a server (e.g., server computing system 506), it may be desirable to configure the first image sensor 602 and the second image sensor 604 to make certain decisions related to object detection and object tracking. For example, if the first image sensor 602 and the second image sensor 604 send each full-resolution image and / or ROI to the server, there will be a delay in performing image processing tasks, and the location of at least one object of interest may not be predicted with high accuracy. Therefore, when at least one object of interest is determined to be moving within the environment, the first image sensor 602 can be configured to detect at least one object of interest in a first frame, sample the next frame, and generate a trajectory of the object's motion within the environment. This information is included in the relevant ROI image and can then be sent to the second image sensor 604 for further tracking. For example, the second image sensor 604 can receive information related to the trajectory, capture an image of the environment where the at least one object of interest is predicted to be located based on the trajectory, and process only the pixels in the image corresponding to the ROI including the predicted location, rather than the full-resolution image.
[0109] This direct peer-to-peer communication of ROI images among the sensors allows them to react quickly to environmental changes and capture high-quality sensor data for determining how to operate the vehicle 600. By avoiding communication through the vehicle control system 620, the information transmission path can be shortened, thereby reducing communication latency. Furthermore, by using the first control circuit 606 and the second control circuit 608 to select ROIs, detect at least one object of interest, and / or determine one or more ROIs including at least one object of interest, the speed of ROI selection, detection of at least one object of interest, processing of additional ROIs, and / or determination of the operating mode can be independent of the processing load on the vehicle control system, thereby further reducing latency.
[0110] While the first image sensor 602 and the second image sensor 604 can be configured to communicate directly with each other to share certain information, these sensors can also share information with the vehicle control system 620. For example, each sensor can share with the vehicle control system 620 the results of various sensor data processing operations performed on the captured sensor data, as well as at least a portion of the sensor data itself. Typically, the first image sensor 602 and the second image sensor 604 can share information useful in operating the vehicle 600 with the vehicle control system 620 and can directly share information with each other about where and how sensor data is captured. Therefore, a peer-to-peer sensor network can allow the vehicle 600 to avoid using the vehicle control system 620 as a communication intermediary for some type of communication between the first image sensor 602 and the second image sensor 604.
[0111] In some implementations, the first image sensor 602 may be configured to capture image data of the environment. Based on or in response to the captured image data, the first image sensor 602 may be configured to transmit the image data to the vehicle control system 620. This transmission may be performed via connection 610.
[0112] Based on or in response to the reception of image data, the vehicle control system 620 can be configured to select an operating mode for the first image sensor 602. The selected operating mode can be based on captured initial image data and can be selected to improve the quality of future images captured in the environment represented by the initial image data. Based on or in response to the selection of the operating mode, the vehicle control system 620 can be configured to transmit the operating mode to the first image sensor 602. This transmission can also be performed via connection 610. In an alternative embodiment, this operation can be performed by a first control circuit 606 provided as part of the first image sensor 602, rather than relying on the vehicle control system 620 to select the operating mode.
[0113] Based on or in response to the selection and / or reception of an operating mode, the first image sensor 602 can be configured to adjust the number of ROIs and full-resolution images to be processed. By doing so, the first image sensor 602 can capture additional image data, which can have a higher quality than the initial captured image data (e.g., better exposure, magnification, white balance, etc.). Therefore, based on or in response to adjusting the operating mode, the first image sensor 602 can be configured to capture additional image data.
[0114] Based on or in response to capturing additional image data, the first image sensor 602 can be configured to detect at least one object of interest within the additional image data. This detection can be performed by a first control circuitry 606, which is provided as part of the first image sensor 602. Based on or in response to detecting at least one object of interest, the first image sensor 602 can be configured to select another sensor and / or the ROI of that sensor that is expected to view at least one object of interest at a future time.
[0115] To this end, the first control circuit 606 of the first image sensor 602 can be configured to predict the future position of at least one object of interest relative to the vehicle 600 based on the properties of the object of interest. The control circuit can also determine which sensor on the vehicle 600 will view (e.g., in the case of a fixed sensor) and / or can be repositioned to view (e.g., in the case of a sensor with adjustable pose) at least one object of interest in the future. Furthermore, the first control circuit 606 can determine the specific ROI of the determined sensor where at least one object of interest will appear.
[0116] Based on the selection of the sensor and / or its ROI, the first image sensor 602 can be configured to send the ROI selection, the properties of at least one object of interest, and / or the operating mode used by the first image sensor 602 to the second image sensor 604. Based on or in response to receiving the ROI selection, the properties of at least one object of interest, and / or the operating mode, the second image sensor 604 can be configured to determine the operating mode used by the second image sensor 604 to scan at least one object of interest, and adjust its operating mode accordingly. Based on or in response to the adjustment of the operating mode, the second image sensor 604 can be configured to capture ROI image data of the selected ROI.
[0117] VIII. Additional Example Operations
[0118] Figure 7 A flowchart illustrating operations related to processing ROI images is shown. These operations can be performed by image sensor 100, image sensor 200, their components, and / or associated circuitry, etc. However, these operations can also be performed by other types of devices or device subsystems. For example, the process can be performed by server equipment, autonomous vehicles, and / or robotic devices.
[0119] Figure 7 The embodiments can be simplified by removing any one or more features shown in the figures. Furthermore, these embodiments can be combined with any features, aspects, and / or implementations described in the previous figures or herein.
[0120] Block 700 may involve obtaining a full-resolution image of the environment by means of control circuitry and from an image sensor comprising a plurality of pixels forming a plurality of regions of interest (ROIs), wherein the image sensor is configured to operate at a frame rate above a threshold rate, wherein the full-resolution image contains each of the plurality of ROIs.
[0121] Block 702 may involve selecting a specific ROI based on a full-resolution image via control circuitry.
[0122] Block 704 may involve detecting at least one object of interest in a specific ROI via control circuitry.
[0123] Block 706 may involve determining the operating mode for processing subsequent image data generated from a specific ROI via control circuitry.
[0124] Block 708 may involve processing image data, including multiple ROI images of detected objects of interest, based on the operating mode and frame rate via control circuitry.
[0125] In some embodiments, processing of image data based on operating mode and frame rate may include generating multiple ROI images of at least one object of interest during one or more next duty cycles, rather than obtaining additional full-resolution images.
[0126] In some embodiments, processing of image data based on operating mode and frame rate may include generating additional full-resolution images and multiple ROIs, including one or more objects of interest, during one or more next duty cycles.
[0127] In some embodiments, processing of image data based on operating mode and frame rate may include generating multiple additional full-resolution images, including one or more objects of interest, during one or more next duty cycles.
[0128] In some embodiments, the system may include a vehicle, and the operating mode may be based on the position of the image sensor on the vehicle.
[0129] In some embodiments, the number of multiple ROI images of at least one object of interest can be based on the frame rate.
[0130] In some embodiments, the image sensor may be configured to have a predetermined operating mode.
[0131] In some embodiments, image data processing may include performing one or more image processing tasks on multiple ROI images of at least one object of interest.
[0132] In some embodiments, the operation may further include comparing a distance between the image sensor and at least one object of interest represented within the full-resolution image with a threshold distance. These operations also include, based on the comparison result of the distance and the threshold distance, obtaining multiple ROI images of the at least one object of interest, rather than obtaining an additional full-resolution image.
[0133] In some embodiments, the operation may further include comparing the velocity of at least one object of interest represented within the full-resolution image with a threshold velocity. These operations also include, based on the comparison result, obtaining multiple ROI images of at least one object of interest, instead of obtaining an additional full-resolution image.
[0134] In some embodiments, the operation may further include providing the server with multiple ROI images of at least one object of interest and a full-resolution image including at least one object of interest.
[0135] In some embodiments, a full-resolution image can be acquired at a first time, and the operation may further include sending image data of at least one object of interest within the environment to a second image sensor. The operation also includes determining a specific ROI from a plurality of ROIs of the second image sensor. This specific ROI may correspond to the expected location of at least one object of interest within the environment at a second time later than the first time. The operation further includes acquiring multiple ROI sensor data from the specific ROI from the second sensor, rather than acquiring full-resolution sensor data containing each corresponding ROI among the plurality of ROIs.
[0136] In some embodiments, a first subset of the control circuitry may form part of an image sensor. A second subset of the control circuitry may form part of a second image sensor. The first subset of the control circuitry may be configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, (iii) select a second image sensor from a plurality of image sensors based on the pose of the second image sensor relative to the environment and the expected location of at least one object of interest within the environment at a second time, (iv) further determine a specific ROI based on the selection of the second image sensor, and (v) send an indication of the specific ROI to the second subset of the control circuitry via a peer-to-peer connection between the image sensor and the second image sensor. The second subset of the control circuitry may be configured to receive an indication of the specific ROI from the first subset of the control circuitry and, in response to receiving the indication, acquire multiple ROI sensor data.
[0137] In some embodiments, a first subset of the control circuitry may form part of an image sensor. A second subset of the control circuitry may form part of a second image sensor. The first subset of the control circuitry may be configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, and (iii) broadcast the expected location of at least one object of interest to the plurality of image sensors via a plurality of peer-to-peer connections between the image sensor and the plurality of image sensors including the second image sensor. The second subset of the control circuitry may be configured to, in response to receiving the broadcast and based on the pose of the second image sensor relative to the environment and the expected location of at least one object of interest within the environment, determine a specific ROI of the second image sensor that is expected to observe at least one object of interest at a second time, and acquire multiple ROI sensor data in response to determining the specific ROI.
[0138] In some embodiments, the system may include a vehicle. The system may also include a control system configured to control the vehicle based on data generated by the image sensors. Transmission of image data, determination of specific ROIs, and acquisition of sensor data for multiple ROIs can be performed independently of the control system.
[0139] In some embodiments, the detection of at least one object of interest may include determining one or more of the following: (i) the geometric properties of at least one object of interest, (ii) the actual location of at least one object of interest within the environment, (iii) the velocity of at least one object of interest, (iv) the optical flow associated with at least one object of interest, (v) the classification of at least one object of interest, or (vi) one or more confidence values associated with the processing results of the image data.
[0140] In some embodiments, the control circuit can implement an artificial neural network configured to process image data in the following ways: (i) analyzing the image data using a neural network circuit, and (ii) generating neural network output data related to the image data processing results using a neural network circuit.
[0141] In some embodiments, the control circuitry may be configured to switch between operating in one of three different modes based on conditions within the environment. The first mode of the three different modes may involve acquiring a first plurality of full-resolution images at a first frame rate. The second mode of the three different modes may involve alternating between acquiring full-resolution images at the first frame rate and acquiring a plurality of ROI images at a second frame rate higher than the first frame rate. The third mode of the three different modes may involve (i) acquiring a second plurality of full-resolution images at a third frame rate higher than the first frame rate, and (ii) combining the second plurality of full-resolution images to generate a stacked full-resolution image.
[0142] IX. Conclusion
[0143] This invention is not limited to the specific embodiments described herein, which are intended to illustrate various aspects. Many modifications and variations can be made without departing from its scope, as will be apparent to those skilled in the art. In addition to the methods and apparatus described herein, functionally equivalent methods and apparatus within the scope of this disclosure will be apparent to those skilled in the art based on the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0144] The detailed description above, with reference to the accompanying drawings, illustrates various features and operations of the disclosed systems, devices, and methods. In the drawings, similar symbols generally identify similar components unless the context indicates otherwise. The exemplary embodiments described herein and in the drawings are not intended to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter set forth herein. It will be readily understood that, as generally described herein and illustrated in the accompanying drawings, aspects of this disclosure can be arranged, replaced, combined, separated, and designed in a variety of different configurations.
[0145] Regarding any or all message flowcharts, scenarios, and processes described herein, each step, block, and / or communication may represent information processing and / or information transmission, according to the example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages may not be performed in the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved. Furthermore, more or fewer boxes and / or operations may be associated with any message flowcharts, scenarios, and processes discussed herein. Figure 1 They can be used together, and these message flow diagrams, scenarios, and flowcharts can be combined with each other, either partially or entirely.
[0146] A step or block representing information processing may correspond to a circuit that can be configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a block representing information processing may correspond to a module, segment, or portion of program code (including associated data). Program code may include one or more instructions executable by a processor for implementing a specific logical operation or action in the method or technique. Program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including random access memory (RAM), a disk drive, a solid-state drive, or other storage media.
[0147] Computer-readable media may also include non-transitory computer-readable media, such as short-term data storage media like register memory, processor cache, and RAM. Computer-readable media may also include long-term storage media for program code and / or data. Therefore, computer-readable media can include secondary or permanent long-term storage, such as read-only memory (ROM), optical discs or magnetic disks, solid-state drives, and optical disc read-only memory (CD-ROM). Computer-readable media can also be any other volatile or non-volatile storage system. Computer-readable media can be considered, for example, computer-readable storage media or tangible storage devices.
[0148] Furthermore, a step or block representing one or more information transfers may correspond to information transfers between software and / or hardware modules within the same physical device. However, other information transfers may occur between software and / or hardware modules in different physical devices.
[0149] The specific arrangement shown in the figures should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in the given figures. Furthermore, some of the elements shown may be combined or omitted. Additionally, the example embodiments may include elements not shown in the figures.
[0150] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and not for limitation, and the true scope is indicated by the appended claims.
Claims
1. A system for image sensing, comprising: An image sensor includes multiple regions of interest (ROIs), each ROI comprising a corresponding plurality of pixels, wherein the image sensor is configured to operate at a frame rate higher than a threshold rate; and The control circuit, configured to perform operations, includes: A full-resolution image of the environment is obtained from an image sensor, wherein the full-resolution image contains each of a plurality of ROIs. Select a specific ROI from multiple ROIs based on a full-resolution image; Detect at least one object of interest in a specific ROI; The operation mode to process subsequent image data generated by a specific ROI is determined from multiple operation modes, wherein each corresponding operation mode indicates the image data type of a corresponding predetermined sequence to be captured by the image sensor; and Based on the operating mode and frame rate, subsequent image data of multiple ROI images, including at least one object of interest, are processed.
2. The system according to claim 1, wherein, Processing of subsequent image data based on operating mode and frame rate involves generating multiple ROI images of at least one object of interest during one or more subsequent work cycles, rather than obtaining additional full-resolution images.
3. The system according to claim 1, wherein, Processing of subsequent image data based on operating mode and frame rate includes generating multiple ROI images and additional full-resolution images including one or more objects of interest during one or more subsequent work cycles.
4. The system according to claim 1, wherein, Processing of subsequent image data based on operating mode and frame rate involves generating multiple additional full-resolution images, including one or more objects of interest, during one or more subsequent work cycles.
5. The system of claim 1, further comprising a carrier, wherein, The operating mode is based on the position of the image sensor on the vehicle.
6. The system according to claim 1, wherein, The number of multiple ROI images of at least one object of interest is based on the frame rate.
7. The system according to claim 1, wherein, The processing of the subsequent image data includes performing one or more image processing tasks on multiple ROI images of at least one object of interest.
8. The system according to claim 1, wherein, The operation also includes: The distance between the image sensor and at least one object of interest represented within the full-resolution image is compared to a threshold distance; and Based on the comparison results of the distance and the threshold distance, multiple ROI images of at least one object of interest are obtained, instead of obtaining an additional full-resolution image.
9. The system according to claim 1, wherein, The operation also includes: The velocity of at least one object of interest represented within the full-resolution image is compared with a threshold velocity; and Based on the comparison between the stated velocity and the threshold velocity, multiple ROI images of at least one object of interest are obtained, instead of obtaining an additional full-resolution image.
10. The system according to claim 1, wherein, The operation also includes: Provide the server with multiple ROI images of at least one object of interest and a full-resolution image including at least one object of interest.
11. The system according to claim 1, wherein, The full-resolution image is acquired in a first-time event, and processing the subsequent image data includes: Send subsequent image data of at least one object of interest within the environment to the second image sensor; A second ROI is determined from a plurality of ROIs of a second image sensor, wherein the second ROI corresponds to the expected location of at least one object of interest within the environment at a second time later than the first time; and Instead of obtaining full-resolution sensor data containing each of the multiple ROIs, the sensor data is obtained from multiple ROIs from the second image sensor.
12. The system according to claim 11, wherein, A first subset of the control circuitry forms part of an image sensor, wherein a second subset of the control circuitry forms part of a second image sensor, wherein the first subset of the control circuitry is configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, (iii) select a second image sensor from a plurality of image sensors based on the pose of the second image sensor relative to the environment and the expected location of at least one object of interest in the environment at a second time, (iv) further determine a second ROI based on the selection of the second image sensor, and (v) send an indication of the second ROI to the second subset of the control circuitry via a peer-to-peer connection between the image sensor and the second image sensor, and wherein the second subset of the control circuitry is configured to receive the indication of the second ROI from the first subset of the control circuitry and acquire multiple ROI sensor data in response to the receipt of the indication.
13. The system according to claim 11, wherein, A first subset of the control circuitry forms part of an image sensor, wherein a second subset of the control circuitry forms part of a second image sensor, wherein the first subset of the control circuitry is configured to: (i) acquire a full-resolution image, (ii) determine information associated with at least one object of interest, and (iii) broadcast the expected location of at least one object of interest to the plurality of image sensors via a plurality of peer-to-peer connections between the image sensor and the plurality of image sensors including the second image sensor, wherein the second subset of the control circuitry is configured to, in response to receiving the broadcast and based on the pose of the second image sensor relative to the environment and the expected location of at least one object of interest within the environment, determine a second ROI of the second image sensor that is expected to view at least one object of interest at a second time, and acquire multiple ROI sensor data in response to determining the second ROI.
14. The system of claim 11, further comprising: Vehicle; as well as The control system is configured to control the vehicle based on data generated by the image sensor, wherein the transmission of subsequent image data, the determination of the second ROI, and the acquisition of multiple ROI sensor data are performed independently of the control system.
15. The system according to claim 1, wherein, The detection of at least one object of interest includes determining one or more of the following: (i) the geometric properties of at least one object of interest, (ii) the actual location of at least one object of interest within the environment, (iii) the velocity of at least one object of interest, (iv) the optical flow associated with at least one object of interest, (v) the classification of at least one object of interest, or (vi) one or more confidence values associated with the processing results of subsequent image data.
16. The system according to claim 1, wherein, The control circuit includes a neural network circuit, and the processing of subsequent image data includes: Using neural network circuits to analyze subsequent image data; and The neural network output data is generated using neural network circuits, which is related to the processing results of subsequent image data.
17. The system according to claim 1, wherein, The full-resolution image is acquired at a first frame rate, and wherein multiple ROI images of at least one object are acquired at a second frame rate higher than the first frame rate.
18. A method for image sensing, comprising: A full-resolution image of the environment is obtained by controlling the circuit and from an image sensor comprising multiple regions of interest (ROIs), wherein each corresponding ROI comprises a corresponding plurality of pixels, wherein the image sensor is configured to operate at a frame rate above a threshold rate, and wherein the full-resolution image contains each corresponding ROI among the plurality of ROIs. The control circuit selects a specific ROI from multiple ROIs based on a full-resolution image. The control circuit detects at least one object of interest in a specific ROI. The control circuit determines the operating mode from multiple operating modes to process subsequent image data generated from a specific ROI, wherein each corresponding operating mode indicates the image data type of a corresponding predetermined sequence to be captured by the image sensor; and Subsequent image data, including at least one object of interest, is processed by a control circuit based on the operating mode and frame rate.
19. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a computing device, cause the computing device to perform an operation comprising: A full-resolution image of the environment is obtained from an image sensor comprising multiple regions of interest (ROIs), wherein each corresponding ROI comprises a corresponding plurality of pixels, wherein the image sensor is configured to operate at a frame rate above a threshold rate, wherein the full-resolution image contains each corresponding ROI among the plurality of ROIs; Select a specific ROI from multiple ROIs based on a full-resolution image; Detect at least one object of interest in a specific ROI; The operation mode to process subsequent image data generated by a specific ROI is determined from multiple operation modes, wherein each corresponding operation mode indicates the image data type of a corresponding predetermined sequence to be captured by the image sensor; and Based on the operating mode and frame rate, subsequent image data of multiple ROI images, including at least one object of interest, are processed.