Methods and apparatus for performing computer vision-based tasks
Patent Information
- Application Number
- CN202580018134.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]此外,一般而言,尺度数量是预先确定且固定的,从而导致较高的计算成本
[0010]使用具有可变的尺度数量的动态尺度金字塔可以允许提供针对性能或计算效率进行优化的尺度。更具体地,可以增加或减少尺度数量,以改善任务性能或降低计算成本。此外,通过采用动态地缩放,可以在小分辨率输入上训练网络,并且在无需重新训练的情况下,使其能够在更高分辨率的输入上执行推理。这对于训练时间/资源或在高分辨率数据集不可用的情况下是有益的。
Smart Images

Figure CN122847729A_ABST
Abstract
Description
Technical Field
[0001] This technology relates to a method, neural network, and apparatus for performing computer vision-based tasks, and more particularly, to a method, neural network, and apparatus for providing improved performance for computer vision-based tasks. Background Technology
[0002] Computer vision is a subfield of deep learning and artificial intelligence (AI). It refers to the processing, analysis, and interpretation of image data to enable computers and other machines to handle complex real-world visual data, such as correctly identifying objects or people in digital images and taking appropriate actions when necessary. Computer vision-based tasks can include image classification, object detection and localization, semantic segmentation, instance segmentation, pose estimation, image generation and synthesis, and pattern recognition.
[0003] In this context, the image data to be processed can take many forms, such as video sequences, views from multiple cameras, multidimensional data from 3D scanners, 3D point clouds from LiDaR sensors, or medical scanning equipment. It can include scenes from around the world containing objects of various sizes, which may contain features of varying dimensions. Furthermore, objects can be located at different distances from the viewer. Therefore, appropriate methods for performing computer vision-based tasks may need to consider analyzing image data at various resolutions or scale levels. A common tool for performing multi-resolution analysis relates to pyramid representations of image data, which provide the same set of images at different resolutions. More specifically, an image pyramid is a data structure that supports scaled convolution through a reduced image representation. It consists of a series of copies of the original image, in which the sampling density and resolution decrease with a regular stride. This technique has been widely used in object detection methods. For example, feature pyramids built based on image pyramids form the basis of standard solution methods when applied to neural networks such as convolutional networks (ConvNet).
[0004] However, these common pyramid-based methods have some drawbacks. For example, feature-enhancing each level of an image pyramid increases the computation time of the neural network, making the method impractical for computer vision tasks such as object recognition. Furthermore, training the network end-to-end on the image pyramid implies a large memory footprint. Additionally, in common applications of pyramid implementations in neural networks, the maximum number of scales is determined by the input resolution of the input image. Therefore, to support multiple resolutions, multiple networks need to be trained using datasets of different resolutions.
[0005] Furthermore, the number of scales is generally predetermined and fixed, resulting in high computational costs.
[0006] Therefore, it is desirable to improve the methods used to perform computer vision-based tasks in order to provide optimized analysis of image data in terms of performance and / or computational efficiency. Summary of the Invention
[0007] To this end, a method for performing computer vision-based tasks is provided, comprising receiving an input image and processing the input image at different resolutions to manipulate image data of the input image at different scales by applying a dynamic scale pyramid, the dynamic scale pyramid providing a variable number of scales representing the input image at different resolution levels.
[0008] In addition, a neural network is provided that is configured to apply the above-described method.
[0009] Additionally, an apparatus is provided that includes an image sensor, an event-based vision sensor, and a processing unit configured to perform the method.
[0010] Using a dynamic scale pyramid with a variable number of scales allows for scales optimized for performance or computational efficiency. More specifically, the number of scales can be increased or decreased to improve task performance or reduce computational cost. Furthermore, by employing dynamic scaling, networks can be trained on low-resolution inputs and enabled to perform inference on higher-resolution inputs without retraining. This is beneficial for training time / resources or when high-resolution datasets are unavailable. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the sensor device.
[0012] Figure 2 This is a schematic block diagram of the sensor section.
[0013] Figure 3 This is a schematic block diagram of the pixel array section.
[0014] Figure 4 This is a schematic circuit diagram of a pixel block.
[0015] Figure 5 This is a schematic block diagram showing the event detection unit.
[0016] Figure 6 This is a schematic circuit diagram of the current-to-voltage conversion unit.
[0017] Figure 7 This is a schematic circuit diagram of the subtraction and quantization sections.
[0018] Figure 8 This is a schematic timing diagram illustrating an example operation of the sensor section.
[0019] Figure 9 This is a schematic diagram of a frame data generation method based on event data.
[0020] Figure 10 This is a schematic block diagram of another quantization section.
[0021] Figure 11 This is a schematic diagram of another event detection unit.
[0022] Figure 12 This is a schematic block diagram of another pixel array section.
[0023] Figure 13 This is a schematic circuit diagram of another pixel block.
[0024] Figure 14 This is a schematic block diagram of a scanning imaging device.
[0025] Figure 15 This is a schematic diagram illustrating an exemplary image processing method using a pyramid-based approach.
[0026] Figure 16 This is a schematic diagram illustrating an exemplary image processing method that applies another pyramid-based approach.
[0027] Figure 17 The flowchart of a method for performing computer vision-based tasks based on a dynamic scale pyramid is illustrated schematically.
[0028] Figure 18 This is a schematic diagram of exemplary input images at different resolutions processed using a fixed number of scales.
[0029] Figure 19A This is a schematic block diagram of an exemplary neural network.
[0030] Figure 19B This is a schematic block diagram of an exemplary neural network.
[0031] Figure 20 This is a schematic block diagram illustrating the application of an exemplary dynamic scale pyramid-based method for performing computer vision-based tasks.
[0032] Figure 21 This is a schematic block diagram illustrating the application of another exemplary method based on a dynamic scale pyramid for performing computer vision-based tasks.
[0033] Figure 22 This is a schematic block diagram illustrating the application of another exemplary method based on a dynamic scale pyramid for performing computer vision-based tasks.
[0034] Figure 23 This is a schematic block diagram of an exemplary device.
[0035] Figure 24 This is a schematic block diagram of a vehicle control system.
[0036] Figure 25 This is an explanatory diagram illustrating an example of the mounting positions of the vehicle exterior information detection unit and the imaging unit. Detailed Implementation
[0037] This disclosure aims to alleviate problems related to efficiency and computational cost in computer vision-based tasks. Specifically, the problem to be solved is how to improve task performance and / or reduce computational cost. Solutions to this problem include applying a dynamic scale pyramid that provides a variable number of scales, representing the input image at different resolution levels. According to a first aspect, the scale can be increased or decreased based on the resolution of the input image. Furthermore, according to a second aspect, the scale can be dynamically adapted based on information associated with the input image, such as event-based data. Therefore, this disclosure is based on the operation of conventional image sensors and conventional event / dynamic vision sensors (EVS / DVS).
[0038] Therefore, one possible implementation of EVS / DVS will be described first. This is, of course, purely exemplary. It should be understood that EVS / DVS can also be implemented in different ways.
[0039] Figure 1 This is a diagram showing an example configuration of the sensor device 10. Figure 1 In the example, sensor device 10 is composed of a sensor chip.
[0040] The sensor device 10 is a single-chip semiconductor chip and includes a sensor die (substrate) 11 serving as multiple dies (substrates) and stacked logic dies 12. Note that the sensor device 10 may also include only a single die or three or more stacked dies.
[0041] exist Figure 1 In the sensor device 10, the sensor die 11 includes a sensor section 21 (used as its circuitry), and the logic die 12 includes a logic section 22. Note that the sensor section 21 may be partially formed on the logic die 12. Furthermore, the logic section 22 may be partially formed on the sensor die 11.
[0042] Sensor unit 21 includes a plurality of pixels configured to perform photoelectric conversion on incident light to generate electrical signals, and generates event data indicating that an event has occurred, wherein the event is a change in the electrical signal of the pixel. Sensor unit 21 supplies this event data to logic unit 22. That is, sensor unit 21 performs imaging, which is similar to, for example, a synchronous image sensor, performing photoelectric conversion on incident light in pixels to generate electrical signals. However, sensor unit 21 generates event data indicating that an event has occurred (the event being a change in the electrical signal of the pixel), instead of generating image data (frame data) in frame format. Sensor unit 21 outputs the event data obtained through imaging to logic unit 22.
[0043] Here, the synchronous image sensor is an image sensor configured to perform imaging synchronously with the vertical synchronization signal and output frame data as frame-format image data. Since sensor unit 21 does not operate synchronously with the vertical synchronization signal when outputting event data, it can be considered asynchronous (an asynchronous image sensor) compared to the synchronous image sensor. Event detection can also be performed synchronously using a line scanner that determines which pixels generated events during past "frames" and assigns the same timestamp to these events. Scan-type event detection is useful for medium to high activity scenes because it incurs less readout overhead compared to arbitrator-type (i.e., asynchronous) event detection.
[0044] Note that, similar to the synchronous image sensor, the sensor unit 21 can generate and output frame data in addition to event data. Furthermore, the sensor unit 21 can output the electrical signal of the pixel where the event occurred as a pixel signal, along with the event data; the pixel signal is the pixel value of the pixel in the frame data.
[0045] The logic unit 22 controls the sensor unit 21 as needed. In addition, the logic unit 22 performs various types of data processing, such as data processing to generate frame data based on event data from the sensor unit 21, and image processing on frame data from the sensor unit 21 or frame data generated based on event data from the sensor unit 21, and outputs the data processing results obtained by performing various types of data processing to the event data and frame data.
[0046] Figure 2 It is shown Figure 1 A block diagram illustrating an example configuration of the sensor section 21.
[0047] The sensor unit 21 includes a pixel array unit 31, a driver unit 32, an arbitrator 33, an AD (analog-to-digital) converter unit 34, and an output unit 35.
[0048] The pixel array section 31 includes a plurality of pixels 51 arranged in a two-dimensional dot matrix pattern. Figure 3If the photocurrent (and its corresponding voltage) generated as an electrical signal by photoelectric conversion in pixel 51 changes by more than a predetermined threshold (including changes equal to or greater than the threshold, depending on the situation), the pixel array unit 31 detects this change in photocurrent as an event. Upon detecting an event, the pixel array unit 31 outputs a request to the arbitrator 33 to request the output of event data indicating that the event has occurred. Then, upon receiving a response from the arbitrator 33 authorizing the output of the event data, the pixel array unit 31 outputs the event data to the drive unit 32 and the output unit 35. Furthermore, the pixel array unit 31 outputs the electrical signal of the pixel 51 in which the event has been detected as a pixel signal to the AD conversion unit 34.
[0049] The driving unit 32 provides a control signal to the pixel array unit 31 to drive the pixel array unit 31. For example, the driving unit 32 drives the pixel array unit 31 to output the event data corresponding to the pixel 51, so that the pixel 51 provides (output) a pixel signal to the AD conversion unit 34.
[0050] Arbitrator 33 arbitrates the request from pixel array unit 31 for requesting the output of event data and returns a response to pixel array unit 31 indicating whether the output of event data is permitted or prohibited.
[0051] AD conversion unit 34, for example, is in pixel block 41 described later ( Figure 3 Each column of the image includes, for example, a single-slope ADC (analog-to-digital converter) (not shown). The AD conversion unit 34 uses the ADC in each column to perform AD conversion on the pixel signal of pixel 51 of pixel block 41 in that column and provides the result to the output unit 35. Note that the AD conversion unit 34 can perform CDS (correlated double sampling) while performing AD conversion on the pixel signal.
[0052] The output unit 35 performs necessary processing on the pixel signals from the AD conversion unit 34 and the event data from the pixel array unit 31, and provides the results to the logic unit 22. Figure 1 ).
[0053] Here, the change in the photocurrent generated in pixel 51 can be identified as the change in the amount of light entering pixel 51, so the event can also be said to be the change in the amount of light in pixel 51 (the change in the amount of light greater than the threshold).
[0054] Event data indicating the occurrence of an event includes at least location information (coordinates, etc.) indicating the location of the pixel block where the light intensity change has occurred as an event. Additionally, event data may include the polarity (positive or negative) of the light intensity change.
[0055] Regarding the series of event data output from the pixel array unit 31 in the event occurrence sequence, it can be said that as long as the event data interval is the same as the event occurrence interval, the event data implicitly includes time point information indicating the (relative) time point of the event occurrence. However, for example, when the event data is stored in memory and the event data interval is no longer the same as the event occurrence interval, the time point information implicitly included in the event data will be lost. Therefore, the output unit 35 includes time point information (e.g., a timestamp) indicating the (relative) time point of the event occurrence in the event data before the event data interval changes from the event occurrence interval. As long as this process is performed before the time point information implicitly included in the event data is lost, the process of including time point information in the event data can be performed in any block other than the output unit 35.
[0056] Figure 3 It is shown Figure 2 A block diagram illustrating an example configuration of the pixel array section 31.
[0057] The pixel array unit 31 includes a plurality of pixel blocks 41. Each pixel block 41 includes one or more pixels 51 arranged in I rows and J columns (I and J are integers), i.e., I × J pixels 51, an event detection unit 52, and a pixel signal generation unit 53. The aforementioned one or more pixels 51 in the pixel block 41 share the event detection unit 52 and the pixel signal generation unit 53. Furthermore, VSLs (vertical signal lines) for connecting the pixel block 41 to the ADC of the AD conversion unit 34 are wired in each column of the pixel block 41.
[0058] Pixel 51 receives light incident from an object and performs photoelectric conversion to generate a photocurrent as an electrical signal. Under the control of the driving unit 32, pixel 51 supplies the photocurrent to the event detection unit 52.
[0059] Under the control of the driving unit 32, the event detection unit 52 detects changes in photocurrent from each pixel in pixel 51 that are greater than a predetermined threshold as events. When an event is detected, the event detection unit 52 notifies the arbitrator 33 (…). Figure 2 The event detection unit 52 provides a request for the output of event data indicating that an event has occurred. Then, when the arbitrator 33 receives a response indicating permission to output the event data for the request, the event detection unit 52 outputs the event data to the drive unit 32 and the output unit 35.
[0060] When the event detection unit 52 detects an event, the pixel signal generation unit 53 generates a voltage, i.e., a pixel signal, corresponding to the photocurrent from the pixel 51 under the control of the driving unit 32, and supplies the voltage to the AD conversion unit 34 via VSL.
[0061] Here, a change in photocurrent larger than a predetermined threshold is detected as an event, or it can be identified as the absence of a change in photocurrent larger than a predetermined threshold being detected as an event. The pixel signal generation unit 53 can generate a pixel signal not only when a change in photocurrent larger than a predetermined threshold is detected as an event, but also when no change in photocurrent larger than a predetermined threshold is detected as an event.
[0062] Figure 4 This is a circuit diagram showing an example configuration of pixel block 41.
[0063] like Figure 3 As described, pixel block 41 includes pixel 51, event detection unit 52, and pixel signal generation unit 53.
[0064] Pixel 51 includes photoelectric conversion element 61 and transmission transistors 62 and 63.
[0065] The photoelectric conversion element 61 includes, for example, a PD (photodiode). The photoelectric conversion element 61 receives incident light and performs photoelectric conversion to generate charge.
[0066] The transmission transistor 62 includes, for example, an N-type MOS (metal-oxide-semiconductor) FET (field-effect transistor). The transmission transistor 62 of the nth pixel 51 among the I×J pixels 51 in pixel block 41 responds to the input from the driving unit 32 ( Figure 2 The control signal OFGn supplied by the transistor 62 turns the transistor on or off. When the transmission transistor 62 is turned on, the charge generated in the photoelectric conversion element 61 is transmitted (supplied) to the event detection unit 52 as a photocurrent.
[0067] The transmission transistor 63 includes, for example, an N-type MOSFET. The transmission transistor 63 of the nth pixel 51 among the I×J pixels 51 in the pixel block 41 is turned on or off in response to the control signal TRGn supplied from the driving unit 32. When the transmission transistor 63 is turned on, the charge generated in the photoelectric conversion element 61 is transferred to the FD74 of the pixel signal generation unit 53.
[0068] I×J pixels 51 in pixel block 41 are connected to event detection unit 52 of pixel block 41 via node 60. Thus, the photocurrent generated in pixel 51 (in photoelectric conversion element 61) is supplied to event detection unit 52 via node 60. As a result, event detection unit 52 receives the sum of photocurrents from all pixels 51 in pixel block 41. Therefore, event detection unit 52 detects changes in the sum of photocurrents supplied from the I×J pixels 51 in pixel block 41 as events.
[0069] The pixel signal generation unit 53 includes a reset transistor 71, an amplification transistor 72, a selection transistor 73, and an FD (floating diffusion) transistor 74.
[0070] The reset transistor 71, the amplification transistor 72, and the selection transistor 73 include, for example, an N-type MOSFET.
[0071] Reset transistor 71 responds to the drive section 32 ( Figure 2 The control signal RST supplied by the reset transistor 71 turns the FD74 on or off. When the reset transistor 71 is turned on, the FD74 is connected to the power supply VDD, and the charge stored in the FD74 is discharged to the power supply VDD. As a result, the FD74 is reset.
[0072] The gate of amplifying transistor 72 is connected to FD74, the drain is connected to the power supply VDD, and the source is connected to VSL through select transistor 73. Amplifying transistor 72 is a source follower that outputs a voltage (electrical signal) corresponding to the voltage supplied to the gate of FD74 to VSL via select transistor 73.
[0073] The selection transistor 73 is turned on or off in response to the control signal SEL supplied from the drive unit 32. When the selection transistor 73 is turned on, the voltage corresponding to the voltage of FD74 from the amplification transistor 72 is output to VSL.
[0074] FD74 stores the charge transferred from the photoelectric conversion element 61 of pixel 51 via the transfer transistor 63 and converts the charge into voltage.
[0075] Regarding the pixel 51 and pixel signal generation unit 53 configured as described above, the driving unit 32 uses the control signal OFGn to turn on the transmission transistor 62, thereby supplying the photocurrent generated in the photoelectric conversion element 61 based on the pixel 51 to the event detection unit 52. Thus, the event detection unit 52 receives the sum of the photocurrents from all pixels 51 in the pixel block 41, i.e., the current, or it may be the current of a single pixel.
[0076] When the event detection unit 52 detects a change in the photocurrent (sum of photocurrents) in pixel block 41 as an event, the driving unit 32 turns off the transmission transistors 62 of all pixels 51 in pixel block 41, thereby stopping the supply of photocurrent to the event detection unit 52. Then, the driving unit 32 uses the control signal TRGn to sequentially turn on the transmission transistors 63 of the pixels 51 in the pixel block 41 that detected the event, thereby transferring the charge generated in the photoelectric conversion element 61 to FD74. FD74 accumulates the charge transferred from the pixel 51 (in the photoelectric conversion element 61). The voltage corresponding to the charge accumulated in FD74 is output to VSL as the pixel signal of pixel 51 through the amplification transistor 72 and the selection transistor 73.
[0077] As described above, in sensor section 21 ( Figure 2In the process, only the pixel signals of pixel 51 in pixel block 41 where the event was detected are sequentially output to VSL. The pixel signals output to VSL are then supplied to AD conversion unit 34 for AD conversion.
[0078] Here, in pixel 51 of pixel block 41, the transmission transistor 63 can be turned on simultaneously instead of sequentially. In this case, the sum of the pixel signals of all pixels 51 in pixel block 41 can be output.
[0079] exist Figure 3 In the pixel array section 31, the pixel block 41 includes one or more pixels 51, and the one or more pixels 51 share the event detection section 52 and the pixel signal generation section 53. Therefore, when the pixel block 41 includes multiple pixels 51, compared with the case where an event detection section 52 and a pixel signal generation section 53 are provided for each pixel 51, the number of event detection sections 52 and pixel signal generation sections 53 can be reduced, and as a result, the size of the pixel array section 31 can be reduced.
[0080] Note that when pixel block 41 includes multiple pixels 51, an event detection unit 52 can also be provided for each pixel 51. When multiple pixels 51 in pixel block 41 share the event detection unit 52, events are detected on a per-pixel-block-41 basis. However, when an event detection unit 52 is provided for each pixel 51, events can be detected on a per-pixel-1 basis.
[0081] However, even when multiple pixels 51 in pixel block 41 share a single event detection unit 52, events can be detected on a pixel-by-pixel basis when the transmission transistors 62 of the multiple pixels 51 are temporarily turned on in a time-division manner.
[0082] Furthermore, if pixel signal output is not required, the pixel signal generation unit 53 can be omitted to form pixel block 41. When pixel block 41 is formed by omitting the pixel signal generation unit 53, the AD conversion unit 34 and the transmission transistor 63 can also be omitted to form sensor unit 21. In this case, the size of sensor unit 21 can be reduced. Then, the sensor outputs the address of the pixel (block) where the event occurred, and if necessary, also includes a timestamp.
[0083] Figure 5 It is shown Figure 3 A block diagram illustrating a configuration example of the event detection unit 52.
[0084] The event detection unit 52 includes a current-to-voltage conversion unit 81, a buffer 82, a subtraction unit 83, a quantization unit 84, and a transmission unit 85.
[0085] The current-to-voltage converter 81 converts the sum of photocurrents from the pixel 51 into a voltage corresponding to the logarithm of the photocurrents (hereinafter also referred to as "photovoltage"), and supplies the voltage to the buffer 82.
[0086] The buffer 82 buffers the photovoltage from the current-to-voltage conversion unit 81 and supplies the result to the subtraction unit 83.
[0087] The subtraction unit 83 calculates the difference between the current photovoltage and the photovoltage at a time slightly shifted from the current time when the line drive signal supplied from the drive unit 32 as a control signal is indicated, and supplies the differential signal corresponding to the difference to the quantization unit 84.
[0088] The quantization unit 84 quantizes the differential signal from the subtraction unit 83 into a digital signal and supplies the quantized value of the differential signal as event data to the transmission unit 85.
[0089] The transmission unit 85 transmits (outputs) event data to the output unit 35 based on the event data from the quantization unit 84. That is, the transmission unit 85 supplies a request to the arbitrator 33 for requesting the output of event data. Then, when the arbitrator 33 receives a response indicating permission to output the event data requested, the transmission unit 85 outputs the event data to the output unit 35.
[0090] Figure 6 It is shown Figure 5 A circuit diagram illustrating an example configuration of the current-to-voltage conversion unit 81.
[0091] The current-to-voltage conversion unit 81 includes transistors 91 to 93. For example, N-type MOSFETs can be used as transistors 91 and 93. For example, a P-type MOSFET can be used as transistor 92.
[0092] The source of transistor 91 is connected to the gate of transistor 93, and photocurrent is supplied from pixel 51 to the junction between the source of transistor 91 and the gate of transistor 93. The drain of transistor 91 is connected to the power supply VDD, and its gate is connected to the drain of transistor 93.
[0093] The source of transistor 92 is connected to the power supply VDD, and the drain is connected to the junction between the gate of transistor 91 and the drain of transistor 93. A predetermined bias voltage Vbias is applied to the gate of transistor 92. Using the bias voltage Vbias, transistor 92 is turned on or off, and the operation of the current-to-voltage conversion unit 81 is turned on or off accordingly.
[0094] The source of transistor 93 is grounded.
[0095] In the current-to-voltage conversion unit 81, the drain of transistor 91 is connected to the power supply VDD side, thus it is a source follower. The source of transistor 91, as a source follower, is connected to pixel 51 (…). Figure 4 Therefore, a photocurrent (flowing from the drain to the source) generated in the photoelectric conversion element 61 based on the charge of pixel 51 flows through transistor 91. Transistor 91 operates in the subthreshold region, generating a photovoltage at the gate of transistor 91 corresponding to the logarithm of the photocurrent flowing through transistor 91. As described above, in the current-to-voltage conversion unit 81, transistor 91 converts the photocurrent from pixel 51 into a photovoltage corresponding to the logarithm of the photocurrent.
[0096] In the current-to-voltage conversion unit 81, the gate of transistor 91 is connected to the connection point between the drain of transistor 92 and the drain of transistor 93, and the photovoltage is output from the connection point.
[0097] Figure 7 It is shown Figure 5 A circuit diagram illustrating the configuration of the subtraction unit 83 and the quantization unit 84.
[0098] The subtraction unit 83 includes a capacitor 101, an operational amplifier 102, a capacitor 103, and a switch 104. The quantization unit 84 includes a comparator 111.
[0099] One end of capacitor 101 is connected to buffer 82 ( Figure 5 The output terminal of the capacitor 101 is connected to the input terminal (inverting input terminal) of the operational amplifier 102. Thus, the photovoltage is input to the input terminal of the operational amplifier 102 through the capacitor 101.
[0100] The output terminal of operational amplifier 102 is connected to the non-inverting input terminal (+) of comparator 111.
[0101] One end of capacitor 103 is connected to the input terminal of operational amplifier 102, and the other end is connected to the output terminal of operational amplifier 102.
[0102] Switch 104 is connected to capacitor 103, switching the connection between the two ends of capacitor 103. Switch 104 is turned on or off in response to a row drive signal, which is a control signal supplied from drive unit 32, thereby switching the connection between the two ends of capacitor 103.
[0103] When switch 104 is turned on, the buffer 82 of capacitor 101 ( Figure 5 The photovoltage on one side is denoted as Vinit, and the capacitance (static capacitance) of capacitor 101 is denoted as C1. The input terminal of operational amplifier 102 serves as a virtual ground terminal, and the charge Qinit accumulated in capacitor 101 when switch 104 is turned on is represented by equation (1).
[0104] Qinit = C1 × Vinit (1)
[0105] Furthermore, when switch 104 is turned on, the connection across capacitor 103 is cut off (short-circuited), so no charge accumulates in capacitor 103.
[0106] When switch 104 is subsequently turned off, the buffer 82 of capacitor 101 ( Figure 5 When the photovoltage on one side is denoted as Vafter, the charge Qafter accumulated in capacitor 101 when switch 104 is open is represented by equation (2).
[0107] Qafter = C1 × Vafter (2)
[0108] When the capacitance of capacitor 103 is denoted as C2 and the output voltage of operational amplifier 102 is denoted as Vout, the charge Q2 accumulated in capacitor 103 is represented by equation (3).
[0109] Q2 = -C2 × Vout (3)
[0110] Since the total charge in capacitors 101 and 103 does not change before and after switch 104 is turned off, equation (4) holds true.
[0111] Qinit = Qafter + Q2 (4)
[0112] When equations (1) to (3) are substituted into equation (4), equation (5) is obtained.
[0113] Vout = -(C1 / C2) × (Vafter - Vinit) (5)
[0114] Using equation (5), the subtraction unit 83 subtracts the photovoltage Vinit from the photovoltage Vafter, thus calculating the differential signal (Vout) corresponding to the difference between the photovoltages Vafter and Vinit, Vafter - Vinit. Using equation (5), the subtraction gain of the subtraction unit 83 is C1 / C2. Since maximum gain is generally desired, C1 is preferably set to a large value, and C2 is preferably set to a small value. On the other hand, when C2 is too small, kTC noise increases, and there is a risk of noise characteristic degradation. Therefore, capacitor C2 can only be reduced within the range of acceptable noise. In addition, since each pixel block 41 is equipped with an event detection unit 52 containing the subtraction unit 83, there is a space constraint on capacitors C1 and C2. Taking these factors into consideration, the values of capacitors C1 and C2 are determined.
[0115] Comparator 111 compares the differential signal output from subtraction unit 83 with a predetermined threshold (voltage) Vth (>0) applied to the inverting input terminal (-), thereby quantizing the differential signal. Comparator 111 outputs the quantized value obtained by quantization to transmission unit 85 as event data.
[0116] For example, when the differential signal is greater than the threshold Vth, comparator 111 outputs a High level (H) indicating 1, as event data indicating that an event has occurred. When the differential signal is not greater than the threshold Vth, comparator 111 outputs a Low level (L) indicating 0, as event data indicating that no event has occurred.
[0117] When the transmission unit 85 confirms that a change in light intensity has occurred as an event based on the event data from the quantization unit 84, that is, when the differential signal (Vout) is greater than the threshold Vth, it supplies a request to the arbitrator 33. When a response indicating permission to output event data is received, the transmission unit 85 outputs event data (e.g., H level) indicating that an event has occurred to the output unit 35.
[0118] The output unit 35 includes in the event data received from the transmission unit 85 the position / address information of the pixel 51 (including the pixel 51) where the event shown in the event data occurred, and the time point information indicating the time point when the event occurred. Furthermore, if necessary, it also includes the polarity of the light intensity change during the event, i.e., whether the light intensity increases or decreases. The output unit 35 outputs the event data.
[0119] The data format, which includes the location information of pixel 51 where the event occurred, the time information indicating the time of the event, and the polarity of the light intensity change of the event, can be, for example, a data format called "AER (Address Event Representation)".
[0120] Furthermore, the gain A of the entire event detection unit 52 is expressed by the following formula, where the gain of the current-to-voltage conversion unit 81 is denoted as CGlog, and the gain of the buffer 82 is 1.
[0121] A = CGlogC1 / C2 (Σiphoto_n) (6)
[0122] Here, iphoto_n represents the photocurrent of the nth pixel 51 among the I×J pixels 51 in pixel block 41. In equation (6), Σ represents the summation of n over integers from 1 to I×J.
[0123] Furthermore, pixel 51 can receive arbitrary light as incident light through an optical filter (e.g., a color filter) that allows predetermined light to pass through. For example, when pixel 51 receives visible light as incident light, the event data represents an event in which the pixel value in an image containing a visible object changes. Additionally, for example, when pixel 51 receives infrared light, millimeter waves, or the like as incident light for distance measurement, the event data represents an event in which the distance to the object changes. Furthermore, for example, when pixel 51 receives infrared light as incident light for temperature measurement, the event data represents an event in which the temperature of the object changes. In this embodiment, pixel 51 is assumed to receive visible light as incident light.
[0124] Figure 8 It is shown Figure 2 Timing diagram of an example operation of sensor unit 21.
[0125] At time T0, the driving unit 32 changes all control signals OFGn from L level to H level, thereby turning on the transmission transistors 62 of all pixels 51 in pixel block 41. As a result, the sum of the photocurrents of all pixels 51 in pixel block 41 is supplied to the event detection unit 52. Here, all control signals TRGn are at L level, therefore the transmission transistors 63 of all pixels 51 are off.
[0126] For example, at time T1, when an event is detected, the event detection unit 52 outputs event data at level H in response to the detection of the event.
[0127] At time T2, the drive unit 32 sets all control signals OFGn to L level based on the H-level event data to stop supplying photocurrent from pixel 51 to event detection unit 52. Additionally, the drive unit 32 sets control signal SEL to H level and sets control signal RST to H level for a certain period to control FD 74 to discharge charge to power supply VDD, thereby resetting FD 74. The pixel signal generation unit 53 outputs the pixel signal corresponding to the voltage of FD 74 when FD 74 is reset as the reset level, and the AD conversion unit 34 performs AD conversion on the reset level.
[0128] At time T3 after the reset level AD conversion, the drive unit 32 sets the control signal TRG1 to level H for a certain period of time to control the first pixel 51 in the pixel block 41 that detected the event to transfer the charge generated by photoelectric conversion (in the photoelectric conversion element 61 of the first pixel 51) to FD 74. The pixel signal generation unit 53 outputs a pixel signal corresponding to the voltage of FD 74 from which charge has been transferred from pixel 51. As the signal level, the AD conversion unit 34 performs AD conversion on the signal level.
[0129] The AD conversion unit 34 outputs the difference between the signal level obtained after AD conversion and the reset level as a pixel signal used as the pixel value of the image (frame data) to the output unit 35.
[0130] Here, the process of obtaining the difference between the signal level and the reset level of the pixel signal used as the image pixel value is called "CDS". CDS can be performed after the AD conversion of the signal level and the reset level, or it can be performed simultaneously with the AD conversion of the signal level and the reset level when the AD conversion unit 34 performs single-slope AD conversion. In the latter case, the AD conversion result of the reset level is used as the initial value to perform AD conversion on the signal level.
[0131] At time T4 after the pixel signal of the first pixel 51 in pixel block 41 is converted by AD, the driving unit 32 sets the control signal TRG2 to H level for a certain period of time to control the second pixel 51 in pixel block 41 that has detected the event to output a pixel signal.
[0132] In the sensor unit 21, the same process is then performed, thereby sequentially outputting the pixel signal of pixel 51 in the pixel block 41 where the event was detected.
[0133] When the pixel signals of all pixels 51 in pixel block 41 are output, the driving unit 32 sets all control signals OFGn to H level to turn on the transmission transistors 62 of all pixels 51 in pixel block 41.
[0134] Figure 9 This is a diagram illustrating an example of a method for generating frame data based on event data.
[0135] Logic unit 22 sets the frame interval and frame width, for example, based on instructions from external input. Here, the frame interval refers to the interval between each frame of the frame data generated based on event data. The frame width refers to the time width of the event data used to generate a single frame of frame data. The frame interval and frame width set by logic unit 22 are also referred to as "set frame interval" and "set frame width".
[0136] The logic unit 22 generates frame data as image data in frame format based on the set frame interval, the set frame width and the event data from the sensor unit 21, thereby converting the event data into frame data.
[0137] That is, within each set frame interval, the logic unit 22 generates frame data based on event data within a set frame width starting from the set frame interval.
[0138] Here, it is assumed that the event data includes time point information ti (hereinafter also referred to as "event time point") representing the time when the event occurred, and coordinates (x, y) (hereinafter also referred to as "event position") representing the position information of the pixel 51 (in which the event occurred) (containing the pixel 51) in the pixel block 41.
[0139] exist Figure 9 In the three-dimensional space (time and space) consisting of the x-axis, y-axis and time axis t, points representing the event data are drawn based on the event time point t and the event position (coordinates) (x, y) contained in the event data.
[0140] That is, when the three-dimensional spatial position (x, y, t) represented by the event time point t and event position (x, y) contained in the event data is regarded as the spatial and temporal position of the event, in Figure 9 In this context, points representing event data are plotted at the spatial and temporal location (x, y, t) of the event.
[0141] The logic unit 22 uses a predetermined time point (e.g., the time point when the external instruction to start generating frame data is given, or the time point when the sensor device 10 is powered on) as the start time point for generating frame data, and begins to generate frame data based on event data.
[0142] Here, a cuboid with a set frame width in the time axis t direction, starting from the generation start time point and within each set frame interval, is called a "frame body". The size of the frame body in the x-axis or y-axis direction is, for example, equal to the number of pixel blocks 41 or pixels 51 in the x-axis or y-axis direction.
[0143] Within each set frame interval, the logic unit 22 generates frame data for a single frame based on event data within a frame body with a set frame width, starting from the set frame interval.
[0144] Frame data can be generated, for example, by setting the pixel (pixel value) at the event location (x, y) contained in the event data of the frame to white, and setting the pixels at other locations in the frame to gray, or by setting a predetermined color.
[0145] Furthermore, if the event data contains the polarity of the light intensity change as an event, the polarity contained in the event data can be taken into account when generating frame data. For example, a pixel can be set to white if the polarity is positive, and to black if the polarity is negative.
[0146] Additionally, as referenced Figure 3 and Figure 4As explained, when the pixel signal of pixel 51 is also output along with the event data, frame data can be generated based on the event data using the pixel signal of pixel 51. That is, frame data can be generated by setting the pixel at the event position (x, y) (in the block corresponding to pixel block 41) contained in the event data of the frame as the pixel signal of pixel 51 at position (x, y), and setting the pixels at other positions as a predetermined color such as gray.
[0147] Note that within a frame, there may sometimes be multiple event data points with different event timestamps (t) but the same event location (x, y). In such cases, for example, the most recent or oldest event data point at event timet (t) can be selected first. Additionally, if the event data contains polarity, the polarities of multiple event data points with different event timestamps but the same event location (x, y) can be added together, and the pixel value based on the summed value can be set to the pixel at the event location (x, y).
[0148] Here, when the frame width and frame interval are the same, the frames are adjacent to each other without any gaps. Alternatively, when the frame interval is greater than the frame width, the frames are arranged with gaps. And when the frame width is greater than the frame interval, the frames are arranged with partial overlap.
[0149] Figure 10 It is shown Figure 5 A block diagram of another configuration example of the quantization unit 84.
[0150] In addition, Figure 10 In, with Figure 7 The corresponding parts are marked with the same reference numerals, and their descriptions are omitted appropriately in the following text.
[0151] exist Figure 10 In the quantization section 84, comparators 111 and 112 are included, as well as an output section 113.
[0152] so, Figure 10 Quantitative division 84 and Figure 7 The situation is the same, including comparator 111. However, Figure 10 Quantitative division 84 and Figure 7 Unlike other cases, the addition includes a comparator 112 and an output section 113.
[0153] Include Figure 10 The quantization department 84 and the event detection department 52 ( Figure 5 In addition to detecting events, it also detects the polarity of changes in light intensity as events.
[0154] exist Figure 10In the quantization unit 84, when the differential signal is greater than the threshold Vth, the comparator 111 outputs an H level representing 1 as event data indicating that a positive polarity event has occurred. When the differential signal is not greater than the threshold Vth, the comparator 111 outputs an L level representing 0 as event data indicating that no positive polarity event has occurred.
[0155] Furthermore, in Figure 10 the quantization unit 84, a threshold Vth' (<Vth) is supplied to the non-inverting input terminal (+) of the comparator 112, and the differential signal from the subtraction unit 83 is supplied to the inverting input terminal (-) of the comparator 112. Here, for simplification of description, it is assumed that the threshold Vth' is equal to -Vth, but this is not necessary.
[0156] The comparator 112 quantizes the differential signal by comparing the differential signal output from the subtraction unit 83 with the threshold Vth' applied to the inverting input terminal (-). The comparator 112 outputs the quantized value obtained through the quantization as event data.
[0157] For example, when the differential signal is less than the threshold Vth' (the absolute value of the negative differential signal is greater than the threshold Vth), the comparator 112 outputs an H level representing 1 as event data indicating that a negative polarity event has occurred. Furthermore, when the differential signal is not less than the threshold Vth' (the absolute value of the negative differential signal is not greater than the threshold Vth), the comparator 112 outputs an L level representing 0 as event data indicating that no negative polarity event has occurred.
[0158] The output unit 113 outputs, to the transmission unit 85 according to the event data output from the comparators 111 and 112, event data indicating that a positive polarity event has occurred, event data indicating that a negative polarity event has occurred, or event data indicating that no event has occurred.
[0159] For example, when the event data from the comparator 111 is an H level representing 1, the output unit 113 outputs +V volts representing +1 to the transmission unit 85 as event data indicating that a positive polarity event has occurred. Furthermore, when the event data from the comparator 112 is an H level representing 1, the output unit 113 outputs -V volts representing -1 to the transmission unit 85 as event data indicating that a negative polarity event has occurred. In addition, when each event data from the comparators 111 and 112 is an L level representing 0, the output unit 113 outputs 0 volts (GND level) representing 0 to the transmission unit 85 as event data indicating that no event has occurred.
[0160] If, based on the event data from the output unit 113 of the quantization unit 84, it is confirmed that a change in light intensity has occurred as an event with positive or negative polarity, the transmission unit 85 supplies a request to the arbitrator 33. After receiving a response indicating permission to output event data, the transmission unit 85 outputs event data (representing +V volts of 1 or -V volts of -1) indicating that an event with positive or negative polarity has occurred to the output unit 35.
[0161] Preferably, the quantization unit 84 has the following characteristics: Figure 10 The configuration shown.
[0162] Figure 11 This is a diagram showing another configuration example of the event detection unit 52.
[0163] exist Figure 11 In this configuration, the event detection unit 52 includes a subtractor 430, a quantizer 440, a memory 451, and a control unit 452. The subtractor 430 and the quantizer 440 correspond to the subtraction unit 83 and the quantization unit 84, respectively.
[0164] Note that in Figure 11 In the event detection unit 52, there is also a block corresponding to the current-to-voltage conversion unit 81 and the buffer 82, but... Figure 11 The illustrations of these blocks are omitted in the text.
[0165] The subtractor 430 includes a capacitor 431, an operational amplifier 432, a capacitor 433, and a switch 434. The capacitors 431, 432, 433, and 434 correspond to the capacitors 101, 102, 103, and 104, respectively.
[0166] Quantizer 440 includes comparator 441. Comparator 441 corresponds to comparator 111.
[0167] Comparator 441 compares the voltage signal (differential signal) output from subtractor 430 with a predetermined threshold voltage Vth applied to the inverting input terminal (-). Comparator 441 outputs a signal representing the comparison result as a detection signal (quantized value).
[0168] The voltage signal from the subtractor 430 can be input to the input terminal (-) of the comparator 441, and the predetermined threshold voltage Vth can be input to the input terminal (+) of the comparator 441.
[0169] The control unit 452 supplies a predetermined threshold voltage Vth to the inverting input terminal (-) of the comparator 441. The supplied threshold voltage Vth can be changed in a time-division manner. For example, the control unit 452 supplies a threshold voltage Vth1 corresponding to an ON event (e.g., a positive change in photocurrent) and a threshold voltage Vth2 corresponding to an OFF event (e.g., a negative change in photocurrent) at different timings, thereby enabling the detection of multiple types of address events (events) using a single comparator.
[0170] The memory 451 accumulates the output of the comparator 441 based on the sampling signal supplied from the control unit 452. The memory 451 can be a sampling circuit such as a switch, transistor, or capacitor, or a digital storage circuit such as a latch or flip-flop. For example, the memory 451 can hold the comparison result performed by the comparator 441 using the threshold voltage Vth1 corresponding to the ON event while the inverting input terminal (-) of the comparator 441 is supplied with the threshold voltage Vth2 corresponding to the OFF event. Note that the memory 451 can be omitted, or it can be located inside the pixel (pixel block 41), or it can be located outside the pixel.
[0171] Figure 12 It is shown Figure 2 A block diagram of another configuration example of the pixel array section 31.
[0172] Note that in Figure 12 In, with Figure 3 The corresponding parts are indicated by the same reference numerals, and their descriptions are omitted below as appropriate.
[0173] exist Figure 12 In the pixel array unit 31, there are multiple pixel blocks 41. Each pixel block 41 includes I×J pixels 51, which are one or more pixels, and an event detection unit 52.
[0174] so, Figure 12 pixel array section 31 and Figure 3 The similarity lies in that the pixel array unit 31 includes a plurality of pixel blocks 41, and each pixel block 41 includes one or more pixels 51 and an event detection unit 52. However, Figure 12 pixel array section 31 and Figure 3 The difference is that pixel block 41 does not include pixel signal generation unit 53.
[0175] As mentioned above, in Figure 12 In the pixel array section 31, the pixel block 41 does not include the pixel signal generation section 53, so the sensor section 21 can be formed without the AD conversion section 34. Figure 2 ).
[0176] Figure 13 It is shown Figure 12 A circuit diagram of a configuration example for pixel block 41.
[0177] like Figure 12 The pixel block 41 includes a pixel 51 and an event detection unit 52, but does not include a pixel signal generation unit 53.
[0178] In this case, pixel 51 can only include photoelectric conversion element 61 without transmission transistor 62 and transmission transistor 63.
[0179] Note that pixel 51 has Figure 13 In the configuration shown, the event detection unit 52 can output a voltage corresponding to the photocurrent from the pixel 51 as a pixel signal.
[0180] The sensor device 10 has been described above as an asynchronous imaging device that reads out events via an asynchronous readout system. However, the event readout system is not limited to an asynchronous readout system; it can also be a synchronous readout system. The imaging device using a synchronous readout system is a scanning imaging device, the same as a general imaging device that captures images at a predetermined frame rate.
[0181] Figure 14 This is a block diagram illustrating an example configuration of a scanning imaging device.
[0182] like Figure 14 As shown, the imaging device 510 includes a pixel array unit 521, a driving unit 522, a signal processing unit 525, a readout area selection unit 527, and a signal generation unit 528.
[0183] The pixel array unit 521 includes a plurality of pixels 530. Each of the plurality of pixels 530 outputs an output signal in response to a selection signal from the readout region selection unit 527. Each of the plurality of pixels 530 may include, for example, Figure 11 The pixel-level quantizer is shown. Multiple pixels 530 output signals corresponding to changes in light intensity. Multiple pixels 530 can be used as follows... Figure 14 The diagram shows a two-dimensional matrix configuration.
[0184] The driving unit 522 drives multiple pixels 530, causing each pixel 530 to output its generated pixel signal to the signal processing unit 525 via the output line 514. Note that the driving unit 522 and the signal processing unit 525 are circuit units used to acquire grayscale information. Therefore, when only event information (event data) is acquired, the driving unit 522 and the signal processing unit 525 can be omitted.
[0185] The readout region selection unit 527 selects a portion of the pixels 530 included in the pixel array unit 521. For example, the readout region selection unit 527 selects one or more rows included in the two-dimensional matrix structure corresponding to the pixel array unit 521. The readout region selection unit 527 selects one or more rows sequentially based on a preset period. Furthermore, the readout region selection unit 527 can determine the selection region based on requests from the pixels 530 in the pixel array unit 521.
[0186] The signal generation unit 528 generates an event signal corresponding to an active pixel in the selected pixel 530 where an event has been detected, based on the output signal of the pixel 530 selected by the readout region selection unit 527. An event refers to a change in light intensity. An active pixel is a pixel 530 whose light intensity change corresponding to the output signal exceeds or falls below a preset threshold. For example, the signal generation unit 528 compares the output signal from the pixel 530 with a reference signal and detects pixels whose output signals are greater than or less than the reference signal as active pixels. The signal generation unit 528 generates an event signal (event data) corresponding to the active pixel.
[0187] The signal generation unit 528 may include, for example, a column selection circuit for arbitrating signals input to the signal generation unit 528. Furthermore, the signal generation unit 528 can output not only information about active pixels where an event was detected, but also information about inactive pixels where no event was detected.
[0188] The signal generation unit 528 outputs address information and timestamp information (e.g., (X, Y, T)) about the activated pixel of the detected event via the output line 515. However, the data output from the signal generation unit 528 can be not only address information and timestamp information, but also frame format information (e.g., (0, 0, 1, 0, ...)).
[0189] In the following description, for the sake of simplicity and to cover important application examples, reference will primarily be made to the EVS-type sensor device described above. However, the principles explained below are equally applicable to any event-based vision sensor capable of generating events based on the occurrence of time-varying intensity changes. The following describes how this basic concept of event-based vision sensors can be extended to color changes.
[0190] Figure 15 This illustrates the general principles of pyramid-based image processing methods. More specifically, in Figure 15The upper part of the diagram shows a simplified image pyramid GSP, which represents the input image I0 at different scales SC0, SC1, SC2, SC3, or resolution levels. At each scale SC1, SC2, SC3, the input image I0 is represented by copies of itself I1, I2, I3 with different, i.e., decreasing resolutions. This is achieved by smoothing the input image I0 using an appropriate smoothing filter or kernel, and then subsampling (downsampling) the smoothed image, typically by a factor of 2 along the x and y coordinate directions. The resulting image I1 is then subjected to the same processing, and this cycle is repeated multiple times. Each cycle of this process yields smaller images I2 and I3 with more smoothing but reduced spatial sampling density, i.e., reduced image resolution (width and height halved). The resulting multi-scale representation looks like... Figure 15 The pyramid shown above has the original input image I0 at the bottom, while the smaller images I1, I2, and I3 obtained from each loop are stacked one on top of the other. The number of loops and thus the number of duplicate images are not limited to... Figure 15 The examples shown are examples of those. More loops can be executed.
[0191] Figure 15 The lower part shows specific examples of the input image I0 and its copy sequences I1, I2, I3. It can be seen that starting with the original input image I0 with high resolution Hres, the resolution decreases as the scales SC0, SC1, SC2, SC3 increase, until the image I3 with the lowest resolution Lres.
[0192] The pyramid construction described above is equivalent to convolving the original input image I0 with a set of Gaussian-like weighted functions. This convolution acts as a low-pass filter, with its bandwidth limited by decreasing by one octave at each scale SC1, SC2, SC3, or level. Therefore, the multi-scale representation is also known as the low-pass pyramid or Gaussian pyramid GSP.
[0193] Besides low-pass pyramids or Gaussian pyramids (GSPs), bandpass pyramids can also be created as follows: In a GSP, the differences between images I1, I2, and I3 at adjacent scales SC1, SC2, and SC3 are calculated, and image interpolation is performed between the resolutions of adjacent scales SC1, SC2, and SC3 to enable the calculation of pixel-wise differences. This... Figure 16 The text is presented in a simplified manner.
[0194] Figure 16 This is a schematic diagram illustrating an exemplary image processing method that applies another pyramid-based approach. Figure 16The bandpass image shown on the right can be obtained by subtracting the images I3, I2, I1 corresponding to the corresponding scales SC3, SC2, SC1, or levels of the pyramid GSP from the images I2, I1, I0 at the next lower scales SC2, SC1, SC0 in the pyramid GSP. Only the smallest scale SC3 is not a difference image so that the input image I0 can be reconstructed. Since the sampling densities of these scales SC3, SC2, SC1 are different, it is necessary to interpolate new sample values between those sample values at a given scale SC3, SC2, SC1 before subtracting that scale SC3, SC2, SC1 from the next lower scale SC2, SC1, SC0. Interpolation can be achieved by expanding the corresponding images I3, I2, I1. Specifically, images I3, I2, I1 can be expanded by doubling the size of images I3, I2, I1 in each iteration, such as... Figure 16 As shown by the corresponding arrows in the image. Then, subtract the images to obtain the result as shown. Figure 16 The image shown is a bandpass image of the low-pass pyramid. Since each value of this bandpass pyramid can be obtained by convolving the difference between two Gaussian pyramids with the original image I0, similar to the Laplacian operator, the bandpass pyramid is also called the Laplacian pyramid (LPP). The Laplacian pyramid (LPP) provides a difference image of a blurred version of each scale (SC3, SC2, SC1) of the Gaussian pyramid (GSP).
[0195] The pyramid method described above can be applied to various computer vision-based tasks. Therefore, different combinations of pyramids GSP and LPP for multi-scale analysis are possible and implicitly included.
[0196] In the following text, based on Figure 15 and Figure 16 The principles and concepts of Gaussian / Laplace pyramids (GSP) and LPP, as shown and described in the preceding paragraphs, provide a detailed explanation of the dynamic scale pyramid DSP of this disclosure.
[0197] The method for performing computer vision-based tasks according to this disclosure includes the following operations.
[0198] In operation S201, an input image I0 is received. The input image I0 may include image data obtained by the image sensor 310, which will be referenced... Figure 23 To provide a more detailed explanation.
[0199] In operation S202, the input image I0 is processed at different resolutions. The image data of the input image I0 at different scales SC0, SC1, SC2, and SC3 are operated on by applying a dynamic scale pyramid DSP. The dynamic scale pyramid provides a variable number of scales SC0, SC1, SC2, and SC3 representing the input image I0 at different resolution levels.
[0200] Figure 17 The process of the above method is illustrated schematically.
[0201] Specifically, the dynamic scale pyramid DSP allows for the dynamic provision of additional scales SC1, SC2, SC3 (and their associated images) or skipping scales SC1, SC2, SC3 (and their associated images). This leads to more efficient performance of computer vision-based tasks such as object detection or pattern recognition. The dynamic scale pyramid method differs from conventional methods that use a fixed number of scales.
[0202] More specifically, in cases such as Figure 17 In operation S203 shown, the number of scales SC1, SC2, and SC3 can be dynamically adjusted based on information related to the input image I0. The number of scales SC1, SC2, and SC3 can be increased or decreased according to this information. An example of this dynamic adjustment is described below.
[0203] Generally, when training neural networks, different models with different predefined architectures and depths must be trained for input images with different input resolutions to achieve similar predictive performance. For example, "EfficientDet" (see, for example, arXiv:1911.09070v7, which is incorporated herein by reference) refers to a common object detector that includes a weighted bidirectional feature pyramid network (BiFPN) and a compound scaling method, as shown below, which requires training BiFPNs of different depths using datasets of different resolutions.
[0204]
[0205] In other words, each network can be trained using different numbers of fixed scales and depths.
[0206] Figure 18 The neural network is to be trained by applying a fixed number of scales SC1 and SC2 (e.g., two). Figure 18 A schematic diagram of an exemplary input image I0 with two different resolutions (not shown). The left portion shows an input image I0 with standard resolutions (e.g., 512 x 512, 640 x 360 pixels, 720 x 576 pixels, and / or 720 x 480 pixels) as the input resolution. SDThe right side shows an input image I with full HD resolution (such as, for example, 1024 x 720 pixels and / or 1920 x 1080 pixels). FHD The shaded areas represent kernels or filters with specific kernel sizes KS or filter sizes used to process (convolve) the input image I0 at scale SC0, thereby providing a reduced-resolution image I1 at a subsequent scale SC1. This may indicate an application as described above. Figure 15 and Figure 16 The first cycle of the pyramid. More specifically, the shaded area can also represent the input space of the neural network, called the receptive field. After processing scale SC0, for a standard resolution input image I... SD The receptive field can be approximately 75%, for full HD input images. FHD It can be approximately 37.5%.
[0207] Now, when inputting a full HD image I FHD Used for standard resolution input image I SD When training a fixed-scale network, the receptive field and consequently predictive performance decrease without dynamic pyramid scaling. In other words, it may be necessary to use a full HD input image. FHD Rescaling is performed to match the resolution the network was initially trained on.
[0208] Using a dynamic scale pyramid DSP, the network may only need to be trained once, and it can run on any (larger) input image I0 at any input resolution.
[0209] Figure 19A and Figure 19B This is a schematic diagram of an exemplary neural network 1000 applying a dynamic scale pyramid DSP. Non-limiting examples of such a neural network 1000 may include convolutional networks (ConvNet) or U-net.
[0210] Specifically, Figure 19A and Figure 19B The underlying architecture is shown, in which... Figure 19A In the middle, the low-resolution input image I0 (e.g., the standard-resolution input image I) SD ) as input, in Figure 19B In the middle, high-resolution input image I0 (e.g., full HD input image I) FHD The input is fed into network 1000 for training. The loop used to perform convolutions is provided by AvgPool2D. instruct. AvgPool2D Pooling layers, used in general convolutional network architectures, cover portions of the image (such as kernels or filters) from the image itself. Figure 19A and Figure 19B (As shown in the left part) Returns the average value. Network 1000 can also apply encoder / decoder ENC / DEC to use full weight sharing on all scales SC1, SC2, SC3, SC4.
[0211] More specifically, the information related to the input image I0 as described above can include metadata associated with the input image I0, such as the specific filter size or kernel size KS, and the receptive field required for the computer vision-based task. In other words, the dynamic scale pyramid DSP can be initialized with metadata such as the convolution kernel size KS and the maximum expected receptive field expressed as a percentage of the image size.
[0212] Then, for each input image I0, the number of scales SC1, SC2, SC3, and SC4 can be determined using the corresponding metadata associated with the corresponding input image I0.
[0213] In this context, the default number of scales SC1, SC2, SC3, and SC4 can be determined based on metadata related to accuracy or computational needs.
[0214] For a high-resolution input image I0, the number of scales SC1, SC2, SC3, and SC4 can be increased; for a low-resolution input image I0, the number of scales can be decreased.
[0215] Therefore, as Figure 19A and Figure 19B As illustrated, the number of scales SC1, SC2, SC3, and SC4 can be dynamically changed. If the input image I0 has high resolution, the number of scales SC1, SC2, and SC3 can be increased when processing the pixels of the input image I0 with a specific filter or kernel. Conversely, if the input image I0 has low resolution, the number of scales SC1, SC2, and SC3 can be decreased when processing the pixels of the input image I0 with a specific filter or kernel. This allows for flexible training. For example, a low-resolution input image I0... Figure 19A The middle indicator is I SD As can be seen, the scale SC4 generated by the corresponding convolution of image I3 may not be necessary and can be disregarded. This is indicated by the small cross. Furthermore, for Figure 19B The middle indicator is I FHD For high-resolution input images, scales SC1, SC2, and SC3 can be dynamically scaled up to accommodate large moving areas / objects. For example... Figure 19B As shown, the additional scale SC4 is considered.
[0216] The dynamic scale pyramid method can mean shorter training time, lower cost (energy consumption), and less required computer hardware.
[0217] For example, a low-resolution input image I0 with a resolution of 512 x 512 pixels may require approximately 5 hours of training time and 3GB of GPU memory.
[0218] According to another example, for a high-resolution input image I0 with a resolution of 1024 x 1024 pixels, it may require approximately 20 hours of training time and 12GB of GPU memory.
[0219] In summary, using Dynamic Scale Pyramid DSP, Network 1000 may not need to be trained on a fixed number of scales SC1, SC2, SC3, and SC4, and it does not need to be retrained when shifting to larger resolution inputs. Therefore, applying Dynamic Scale Pyramid DSP allows for greater flexibility during training and evaluation. It can be trained on smaller datasets with fewer scales SC1, SC2, and SC3, and dynamically scaled up to higher resolution inputs to obtain better receptive fields and large motion estimation.
[0220] In addition to the examples above, the Dynamic Scale Pyramid DSP can also improve the performance of computer vision-based tasks by applying information such as event data. Therefore, information related to the input image I0 can include event-based data E0 associated with the input image I0.
[0221] Figure 20 This is a schematic diagram of the exemplary dynamic scale pyramid-based method described above, used to perform computer vision-based tasks. The method considers information in the form of event data. The event-based data E0 can be obtained by an event-based vision sensor 330 (e.g., EVS), which will refer to... Figure 23 Please provide an explanation.
[0222] The underlying architecture includes encoder / decoder structures ENC and DEC. The input includes an input image I0 and corresponding event-based data E0 associated with the input image I0. In the first loop, the input image I0 and the corresponding event-based data E0 can be processed as follows: Figure 20 It is downsampled as indicated by DSAMP.
[0223] Subsequently, for a given scale (e.g., the first scale SC1 in this example), the occurrence of an event can be determined based on the event-based data E0 associated with the input image I0. This refers to... Figure 17 The operation S204 is shown.
[0224] If the event occurs above the predefined threshold THR, other scales SC2 can be provided, and subsequent loops can be executed. This means... Figure 17 Operation S205a in the middle.
[0225] If the event occurs below the predefined threshold THR, other scales SC2, SC3, and SC4 can be skipped. This is in Figure 20 The small cross indicates the second scale SC2, for which the event occurred below the threshold THR, therefore no further loops need to be executed, and the corresponding scales SC3 and SC4 can be disregarded. This indicates... Figure 17 Operation S205b in the middle.
[0226] According to the implementation, the input used to determine whether an event occurs above or below the threshold THR can be event-only (sparse) or event-stitched with image data (dense).
[0227] By utilizing sparse event information, it is possible to determine whether additional scales SC1, SC2, SC3, and SC4 are needed. This can be based on analyzing motion information from sparse events.
[0228] In this context, the threshold THR can represent a novelty heuristic, which may include a threshold for, for example, the amount of motion perceived in a scene, or another metric related to whether a downstream task needs to be triggered.
[0229] Classifier ( Figure 20 (Not shown in the text) Then it can be determined whether there is enough novelty or entropy available, and the next scale SC1, SC2, SC3, SC4 can be triggered.
[0230] When training end-to-end to infer the optimal performance-to-cost ratio, the classifier can also decide to skip the remaining scales SC1, SC2, SC3, and SC4. This also applies to neural networks such as U-Net.
[0231] In addition, this method can also apply Scale Invariant Feature Transform (SIFT), a computer vision algorithm used to detect, describe, and match local features in images.
[0232] More specifically, SIFT can detect robust features and local extrema at various scales SC0, SC1, SC2, SC3, and SC4. Assuming the default number of scales SC1, SC2, SC3, and SC4 is four, SIFT can utilize events from the event-based vision sensor 330 to predetermine the required number of scales SC1, SC2, SC3, and SC4 using the novelty heuristic provided by the threshold THR.
[0233] For example, in cases with fewer events (e.g., fewer edges or lower contrast), the chance of finding enough features may be lower, so several scales SC1, SC2, SC3, and SC4 can be skipped for computational efficiency.
[0234] When there are a large number of events, the chance of finding good key points may be higher, so more scales such as SC1, SC2, SC3, and SC4 can be used.
[0235] In summary, the scales SC0, SC1, and SC2 of the dynamic scale pyramid DSP can be dynamically adapted; for example, additional scales SC3 and SC4 can be added, or scales SC2, SC3, and SC4 can be skipped. It should be noted that... Figure 20 The number of scales shown is merely illustrative. Depending on the information provided, more or fewer scales SC1, SC2, SC3, and SC4 may be applied.
[0236] Figure 21 This is a schematic diagram of another exemplary dynamic scale pyramid-based method for performing computer vision-based tasks. In this example, the number of scales SC0, SC1, and SC2 to be processed can be determined based on event-based data E0 associated with the input image I0. Specifically, the order of the dynamic scale pyramid DSP can be reversed because the highest input resolution requires the greatest computational effort. Therefore, by first analyzing the event data E0, it can be determined whether applying the algorithm / encoder / task at the lowest scale (e.g., SC2 in this example) is sufficient, or whether it is necessary to move up to higher scales SC1 and SC0 to achieve higher accuracy.
[0237] For example, such as Figure 21 As shown, Figure 20 As explained, by applying the threshold THR, it is sufficient to determine whether image processing at scales SC2 and SC1 is adequate based on the frequency of event occurrences. Therefore, it is unnecessary to consider scale SC0, which uses the highest resolution input image I0 as input. This is in Figure 21 The middle part is represented by a small cross.
[0238] Figure 22 This is a schematic diagram for performing a computer vision-based task that applies yet another exemplary dynamic scale pyramid-based approach. In this example, the input image I0 can be divided into pixel blocks, and a scale pyramid DSP can be created for each block of the input image I0. More specifically, a given block can be divided into... Figure 22 The diagram shows additional blocks of different sizes, based on event-based data E0 associated with the input image I0. The order in which the blocks of different sizes are processed can be determined based on the event-based data E0.
[0239] Figure 22The example shown can illustrate a use case, such as the chunking mechanism provided in video coding (refer to the HEVC encoder). In this case, the order of chunk sizes can also be reversed, starting with the lowest resolution (maximum chunk size), since the highest input resolution requires the most computational effort. Therefore, by first analyzing the event-based data E0, it can be determined whether applying the algorithm / encoder / task at the lowest scale (e.g., SC2) is sufficient, or whether it is necessary to move up to a finer grid to achieve higher accuracy.
[0240] Figure 23 This is a schematic diagram of an exemplary device 300. The device may include an image sensor 320, an event-based vision sensor 330, and a processing unit 340 configured to perform the methods described above. Non-limiting examples of the image sensor 320 may include, but are not limited to, charge-coupled devices (CCDs) and active pixel sensors (CMOS sensors), which represent digital sensors and are used as color sensors such as high-quality RGB sensors.
[0241] In addition, the event-based vision sensor 330 can be a conventional event / motion-based vision sensor (EVS / DVS) as described herein.
[0242] The processing unit 340 can be configured as a central processing unit (CPU) commonly used in image processing.
[0243] The technology described above (i.e., this technology) can be applied to a variety of products. For example, the technology disclosed herein can be implemented as a device mounted on any type of mobile body, such as a vehicle, electric vehicle, hybrid electric vehicle, motorcycle, bicycle, personal mobility device, aircraft, drone, ship, and robot.
[0244] Figure 24 This is a block diagram illustrating a schematic configuration example of a vehicle control system (as an example of a technical mobile body control system to which one embodiment of the present disclosure may be applied).
[0245] The vehicle control system 12000 includes multiple electronic control units interconnected via a communication network 12001. Figure 24 In the example shown, the vehicle control system 12000 includes a drive system control unit 12010, a body system control unit 12020, an external information detection unit 12030, an internal information detection unit 12040, and an integrated control unit 12050. Furthermore, a microcomputer 12051, an audio / image output unit 12052, and an in-vehicle network interface (I / F) 12053 are shown as functional configurations of the integrated control unit 12050.
[0246] The drive system control unit 12010 controls the operation of equipment related to the vehicle's drive system according to various programs. For example, the drive system control unit 12010 is used as a control device for the following devices: a drive force generating device (such as an internal combustion engine, drive motor, etc.) for generating vehicle driving force, a drive force transmission mechanism for transmitting driving force to the wheels, a steering mechanism for adjusting the vehicle's steering angle, and a braking device for generating vehicle braking force.
[0247] The body system control unit 12020 controls the operation of various devices installed on the vehicle body according to various programs. For example, the body system control unit 12020 is used as a control device for keyless entry systems, smart key systems, power windows, or various lights (such as headlights, reversing lights, brake lights, turn signals, fog lights, etc.). In this case, radio waves emitted by mobile devices that serve as key substitutes or signals from various switches can be input to the body system control unit 12020. The body system control unit 12020 receives these input radio waves or signals and controls the vehicle's door locking devices, power windows, lights, etc.
[0248] The exterior information detection unit 12030 detects information about the exterior of the vehicle, including the vehicle control system 12000. For example, the exterior information detection unit 12030 is connected to the imaging unit 12031. The exterior information detection unit 12030 causes the imaging unit 12031 to image an image of the exterior of the vehicle and receives the image. Based on the received image, the exterior information detection unit 12030 can perform processing such as detecting objects like people, vehicles, obstacles, signs, characters on the road surface, or detecting the distance of such objects.
[0249] The imaging unit 12031 is an optical sensor that receives light and outputs an electrical signal corresponding to the amount of light received. The imaging unit 12031 can output the electrical signal as an image, or it can output the electrical signal as information about the measured distance. Furthermore, the light received by the imaging unit 12031 can be visible light or invisible light such as infrared light.
[0250] The in-vehicle information detection unit 12040 detects information about the interior of the vehicle. The in-vehicle information detection unit 12040 is connected, for example, to a driver state detection unit 12041 that detects the driver's state. The driver state detection unit 12041 includes, for example, a camera that images the driver. Based on the detection information input from the driver state detection unit 12041, the in-vehicle information detection unit 12040 can calculate the driver's level of fatigue or the driver's level of concentration, or it can determine whether the driver is dozing off.
[0251] The microcomputer 12051 can calculate control target values for the drive force generating device, steering mechanism, or braking device based on information about the vehicle's interior or exterior obtained by the external information detection unit 12030 or the internal information detection unit 12040, and output control commands to the drive system control unit 12010. For example, the microcomputer 12051 can perform cooperative control aimed at realizing advanced driver assistance system (ADAS) functions, including collision avoidance or impact mitigation, following distance-based driving, speed maintenance, collision warning, lane departure warning, etc.
[0252] Furthermore, the microcomputer 12051 can control the drive force generating device, steering mechanism, braking device, etc., based on information about the exterior or interior of the vehicle obtained by the external information detection unit 12030 or the internal information detection unit 12040, thereby performing cooperative control aimed at achieving autonomous driving such as autonomous driving without driver operation.
[0253] Furthermore, the microcomputer 12051 can output control commands to the body system control unit 12020 based on information about the exterior of the vehicle obtained by the vehicle exterior information detection unit 12030. For example, the microcomputer 12051 can perform cooperative control designed to prevent glare, such as controlling the headlights to switch from high beam to low beam based on the position of the preceding or oncoming vehicle detected by the vehicle exterior information detection unit 12030.
[0254] The sound / image output unit 12052 transmits the output signal of at least one of sound and image to an output device capable of visually or audibly notifying vehicle occupants or the outside of the vehicle. Figure 23 In the example, audio speaker 12061, display unit 12062, and dashboard 12063 are shown as output devices. Display unit 12062 may include, for example, at least one of an in-vehicle display and a head-up display.
[0255] Figure 25 This is a diagram depicting an example of the mounting position of the imaging unit 12031.
[0256] exist Figure 25 In the imaging unit 12031, there are imaging units 12101, 12102, 12103, 12104 and 12105.
[0257] Imaging units 12101, 12102, 12103, 12104, and 12105 are, for example, arranged on the front nose, side mirrors, rear bumper, and rear door of vehicle 12100, and on the upper part of the windshield inside the vehicle. Imaging unit 12101 on the front nose and imaging unit 12105 on the upper part of the windshield inside the vehicle primarily acquire images of the front of vehicle 12100. Imaging units 12102 and 12103 on the side mirrors primarily acquire images of the sides of vehicle 12100. Imaging unit 12104 on the rear bumper or rear door primarily acquires images of the rear of vehicle 12100. Imaging unit 12105 on the upper part of the windshield inside the vehicle is mainly used to detect vehicles ahead, pedestrians, obstacles, traffic lights, traffic signs, lanes, etc.
[0258] By the way, Figure 25 An example depicting the imaging range of imaging units 12101 to 12104 is shown. Imaging range 12111 represents the imaging range of imaging unit 12101 installed on the front nose. Imaging ranges 12112 and 12113 represent the imaging ranges of imaging units 12102 and 12103 installed on the side mirrors, respectively. Imaging range 12114 represents the imaging range of imaging unit 12104 installed on the rear bumper or rear door. For example, by superimposing the image data captured by imaging units 12101 to 12104, a bird's-eye view of the vehicle 12100 viewed from above can be obtained.
[0259] At least one of the imaging units 12101 to 12104 may have the function of obtaining distance information. For example, at least one of the imaging units 12101 to 12104 may be a stereo camera composed of multiple imaging elements, or may be an imaging element having pixels for phase difference detection.
[0260] For example, the microcomputer 12051 can determine the distance to each three-dimensional object within the imaging range 12111 to 12114 and the time change of that distance (relative speed relative to the vehicle 12100) based on distance information obtained from the imaging units 12101 to 12104. This allows it to specifically extract the nearest three-dimensional object existing on the driving path of the vehicle 12100 and traveling at a predetermined speed (e.g., equal to or greater than 0 km / h) in approximately the same direction as the vehicle 12100 as the preceding vehicle. Furthermore, the microcomputer 12051 can preset the following distance to be maintained in front of the preceding vehicle and execute automatic braking control (including stop-and-go control), automatic acceleration control (including start-and-go control), etc. This enables cooperative control aimed at achieving autonomous driving without relying on driver input.
[0261] For example, microcomputer 12051 can classify three-dimensional object data about three-dimensional objects into three-dimensional object data of two-wheeled vehicles, standard-sized vehicles, large vehicles, pedestrians, utility poles, and other three-dimensional objects based on distance information obtained from imaging units 12101 to 12104. It then extracts the classified three-dimensional object data and uses it for automatic obstacle avoidance. For example, microcomputer 12051 identifies obstacles around vehicle 12100 as obstacles that the driver of vehicle 12100 can visually recognize and obstacles that the driver cannot visually recognize. Then, microcomputer 12051 determines the collision risk indicating the possibility of a collision with each obstacle. If the collision risk is equal to or higher than a set value, and a collision is possible, microcomputer 12051 outputs a warning to the driver via audio speaker 12061 or display unit 12062, and executes forced deceleration or evasive steering via drive system control unit 12010. Thus, microcomputer 12051 can assist the driver in avoiding a collision.
[0262] At least one of the imaging units 12101 to 12104 can be an infrared camera that detects infrared light. For example, the microcomputer 12051 can identify a pedestrian by determining whether a pedestrian exists in the imaging images of the imaging units 12101 to 12104. This pedestrian identification is performed, for example, by extracting feature points from the imaging images of the imaging units 12101 to 12104, which are infrared cameras, and performing pattern matching processing on a series of feature points representing the outline of an object to determine whether the object is a pedestrian. When the microcomputer 12051 determines that a pedestrian exists in the imaging images of the imaging units 12101 to 12104 and thereby identifies the pedestrian, the sound / image output unit 12052 controls the display unit 12062 to overlay a square outline for emphasis onto the identified pedestrian. The sound / image output unit 12052 can also control the display unit 12062 to display an icon or the like representing a pedestrian at a desired location.
[0263] The above describes an example of a vehicle control system to which the technology of this disclosure can be applied. The technology of this disclosure can be applied to the imaging unit 12031 in the above configuration. Specifically, the sensor device 10 can be applied to the imaging unit 12031. The imaging unit 12031, to which the technology of this disclosure is applied, flexibly acquires event data and performs data processing on the event data, thereby enabling it to provide appropriate driving assistance.
[0264] Note that the implementation of this technology is not limited to the above-described implementation and various modifications can be made without departing from the spirit of this technology.
[0265] Furthermore, the effects described herein are merely illustrative and not restrictive, and other effects may also be provided.
[0266] Note that this technology can also be configured as follows.
[0267] [1] A method for performing a computer vision-based task, the method comprising: Receive input image (I0) (S201); and The input image (I0) is processed at different resolutions (S202) to operate the image data of the input image (I0) at different scales (SC0, SC1, SC2, SC3, SC4) by applying a dynamic scale pyramid (DSP). The dynamic scale pyramid provides a variable number of scales (SC0, SC1, SC2, SC3, SC4) representing the input image at different resolution levels.
[0268] [2] According to the method described in [1], the method further includes: The number of scales (SC1, SC2, SC3, SC4) is dynamically adapted (S203) based on information related to the input image (I0), wherein the number of scales (SC1, SC2, SC3, SC4) is increased or decreased according to this information.
[0269] [3] According to the method described in [1] or [2], the number of default scales (SC1, SC2, SC3, SC4) is determined based on metadata related to accuracy or computational requirements.
[0270] [4] According to the method of [2] or [3], the information related to the input image (I0) includes: metadata related to the input image (I0), which includes a specific filter size (KS) of the pixels used to process the input image (I0) and the receptive field required for the computer vision-based task.
[0271] [5] According to the method of [4], if the input image (I0) has high resolution, the number of scales (SC1, SC2, SC3, SC4) is increased to process the pixels of the input image (I0) using the particular filter, and if the input image (I0) has low resolution, the number of scales (SC1, SC2, SC3, SC4) is decreased to process the pixels of the input image (I0) using the particular filter.
[0272] [6] The method according to any one of [1] to [5], wherein the dynamic scale pyramid (DSP) comprises: a Gaussian scale pyramid.
[0273] [7] The method according to any one of [2] to [3], wherein the information related to the input image (I0) includes: event-based data (E0) related to the input image (I0).
[0274] [8] The method according to [7] further includes: For a given scale (SC1, SC2, SC3, SC4), the occurrence of an event is determined based on event-based data (E0) associated with the input image (I0) (S204). If the occurrence of an event exceeds a predefined threshold (THR), then provide (S205a) other scales (SC1, SC2, SC3, SC4); or If the occurrence of an event is below a predefined threshold (THR), then skip the other scales (SC1, SC2, SC3, SC4).
[0275] [9] According to the method of [7] or [8], the number of scales to be processed (SC1, SC2, SC3, SC4) is determined based on event-based data (E0) associated with the input image (I0).
[0276]
[10] The method according to any one of [7] to [9], wherein the input image (I0) is divided into pixel blocks and a dynamic scale pyramid (DSP) is created for each block of the input image (I0), wherein a given block is divided into additional blocks of different sizes based on event-based data (E0) associated with the input image (I0), and wherein the order of processing blocks of different sizes is determined based on event-based data (E0).
[0277]
[11] A neural network (1000) configured to apply the method according to any one of [1] to
[10] .
[0278]
[12] According to the neural network (1000) described in
[11] , for each input image (I0), the number of scales (SC1, SC2, SC3, SC4) is determined using metadata corresponding to the corresponding input image (I0).
[0279]
[13] According to the neural network (1000) described in
[11] or
[12] , the number of scales (SC1, SC2, SC3, SC4) is increased for high-resolution input images (I0), and the number of scales (SC1, SC2, SC3, SC4) is decreased for low-resolution input images (I0).
[0280]
[14] The neural network (1000) according to any one of
[11] to
[13] , wherein the neural network (1000) is a convolutional neural network (ConvNet) that provides weight sharing across all scales (SC0, SC1, SC2, SC3, SC4).
[0281]
[15] An apparatus (300) includes an image sensor (320), an event-based vision sensor (330), and a processing unit (340) configured to perform the method according to any one of [1] to
[10] .
Claims
1. A method for performing a computer vision-based task, the method comprising: Receive input image; as well as The input image is processed at different resolutions to operate on the image data of the input image at multiple scales by applying a dynamic scale pyramid, which provides a variable number of scales representing the input image at different resolution levels.
2. The method according to claim 1, wherein, The method further includes: The number of scales is dynamically adapted based on information related to the input image, wherein the number of scales is increased or decreased according to the information.
3. The method according to claim 1, wherein, The default number of scales is determined based on metadata related to accuracy or computational requirements.
4. The method according to claim 2, wherein, Information associated with the input image includes metadata related to the input image, which includes specific filter sizes for the pixels used to process the input image and the receptive field required for the computer vision-based task.
5. The method according to claim 4, wherein, If the input image has high resolution, the scale number is increased to process the pixels of the input image using the specific filter, and if the input image has low resolution, the scale number is decreased to process the pixels of the input image using the specific filter.
6. The method according to claim 1, wherein, The dynamic scale pyramid includes: the Gaussian scale pyramid.
7. The method according to claim 2, wherein, Information related to the input image includes event-based data associated with the input image.
8. The method according to claim 7, further comprising: For a given scale, the occurrence of an event is determined based on event-based data associated with the input image; If the occurrence of the event exceeds a predefined threshold, other scales are provided; or If the occurrence of the event is below the predefined threshold, then other scales are skipped.
9. The method according to claim 7, wherein, The number of scales to be processed is determined based on event-based data associated with the input image.
10. The method according to claim 7, wherein, The input image is divided into pixel blocks, and a dynamic scale pyramid is created for each pixel block of the input image, wherein a given block is divided into additional blocks of different sizes based on event-based data associated with the input image, and wherein the order in which the pixel blocks of different sizes are processed is determined based on the event-based data.
11. A neural network configured to apply the method of claim 1.
12. The neural network according to claim 11, wherein, For each input image, the number of scales is determined using the corresponding metadata associated with that input image.
13. The neural network according to claim 11, wherein, The number of scales is increased for high-resolution input images and decreased for low-resolution input images.
14. The neural network according to claim 11, wherein, The neural network is a convolutional neural network that provides weight sharing across all scales.
15. An apparatus comprising an image sensor, an event-based vision sensor, and a processing unit configured to perform the method of claim 1.