IMAGE PROCESSING DEVICE AND METHOD FOR FEATURE EXTRACTION

DE602018089483T2Active Publication Date: 2026-02-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602018089483
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2018-08-14
Publication Date
2026-02-25
Estimated Expiration
2038-08-14

AI Technical Summary

Technical Problem

Existing feature extraction methods from images and video frames are computationally intensive and time-consuming, limiting the speed and accuracy of analysis.

Method used

A dual-sensor setup combining an event-driven camera and a conventional CMOS camera is used to extract features, where motion statistics and intensity change information from the event-driven camera are utilized to identify objects of interest, allowing for efficient feature extraction directly from the CMOS camera output without additional computations.

Benefits of technology

This approach enables faster and more accurate feature extraction by focusing only on the objects of interest, reducing computational complexity and enhancing computer vision algorithm performance.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Generally, the present invention relates to the field of image processing. More specifically, the present invention relates to an image processing apparatus and method for extracting features from still images or images of video frames.BACKGROUND

[0002] The amount of video and image data has increased dramatically over recent years and continues to grow. Many vision applications based on video or image data require accurate and efficient analysis of the data. Much work has been recently devoted to automatic video / image analysis and recognition for different purposes. One of the important parts for automatic video / image processing is feature extraction, involving principal component analysis (PCA), deep learning or neural-network based methods, histograms of optical flow (HOF), histograms of 3D gradients, motion boundary histograms, local trinary patterns, motion estimation at dense grid and the like. However, most of these methods are time-consuming processes that limit the speed and the accuracy of feature extraction.

[0003] Event-driven cameras (also referred to as event cameras or event camera sensors) are triggered by objects or events that are interesting for video / image analysis and processing, and allow providing video characteristics in a more straightforward manner, which can spare the need of further analysis algorithms.

[0004] In conventional feature extraction approaches features are generally extracted on the basis of an output of a standard image sensor. For instance, the known feature extraction method SURF (speeded up robust features) uses a local feature detector and descriptor. Moreover, conventional feature extraction approaches normally compute the feature based on the whole image as the input.

[0005] In CENSI ANDREA ET AL: "Low-latency event-based visual odometry", 2014 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA), IEEE, 31 May 2014, pages 703 - 710,

[0006] XP32650286, DOI: 10.1109 / ICRA.2014.6906931, it is disclosed a visual odometry system based on a Dynamic Visions Sensor (DVS) plus a normal CMOS camera to provide the absolute brightness values. The two sources of data are automatically spatiotemporally calibrated from logs taken during normal operation. They design a visual odometry method that uses the DVS events estimates the relative displacement since the previous CMOS frame by processing each event individual.

[0007] In ELIAS MUEGGLER ET AL: "Continuous-Time Visual-Inertial Trajectory Estimation with Event Cameras", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 23 February 2017, XP80748368, it is disclosed a continuous time framework to perform trajectory estimation by fusing visual data from a moving event camera with inertial data from an inertial measurement unit (IMU). This framework allows direct integration of the asynchronous events with micro-second accuracy and the inertial measurements at high frequency.

[0008] In BARDOW PATRICK ET AL: "Simultaneous Optical Flow and Intensity Estimation from an Event Camera", 2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), IEEE, 27, June 2016, pages 884 - 892, XP33021219, DOI: 10.1109 / CVPR.2016.102 it is disclosed an algorithm to simultaneously recover the motion field and brightness image, while an event camera undergoes a generic motion through any scene. The approach employs minimisation of a cost function that contains the asynchronous event data as well as spatial and temporal regularisation within a sliding window time interval, and relies on GPU optimisation and runs in near real-time.SUMMARY

[0009] It is an object of the invention to provide an improved image processing apparatus and method allowing for a computationally more efficient extraction of features from still images or images of video frames.

[0010] The foregoing and other objects are achieved by the subject matter of the appended claims.

[0011] Generally, embodiments of the invention take advantage of image / video characteristics obtained from event-driven cameras for conducting feature extraction for video / image analysis and object / action recognition. Embodiments of the invention use a combination of an event-driven sensor and a conventional CMOS sensor. According to embodiments of the invention this could be a dual-sensor setup in a single camera or it could be a setup of two cameras very close to each other or a setup of two cameras very close to each other with a combination at the image pixel / sample level.

[0012] According to embodiments of the invention motion statistics and intensity change information can be obtained or extracted from the event-driven camera output and used for identifying the object whose features need to be extracted. In other words, the statistics extracted from the event-driven camera can be used to determine the object of the image data from a standard image sensor for feature extraction. According to embodiments of the invention the feature of the identified object of the image data from a standard image sensor can be extracted by a feature extraction method, such as a neural network-based feature extraction method or any conventional feature extraction method.

[0013] The combination of an event-driven camera and a conventional CMOS camera implemented in embodiments of the invention has two main advantages. On the one hand, the sensitivity of the event-driven camera to moving objects leads to fast and accurate detection, which facilitates the identification of the objects of interest form the output of the CMOS camera. On the other hand, the CMOS camera with an intrinsic high resolution and enhanced by the results of the event-driven camera can process a video which is more optimized for machine vision. According to embodiments of the invention the identification of the objects of interest by usage of the output of the event-driven camera could be used to extract the feature of the identified object directly without additional computations based on the image of CMOS sensor / camera.

[0014] Thus, an improved image processing apparatus is provided allowing for a computationally more efficient extraction of features of objects from still images or images of video frames, because generally only a part of the image which contains the object of interest is used for feature calculation and only the features of certain objects are extracted.

[0015] As used herein, a "feature" could be edge, color, texture, or the combination of edge, color, texture, or vector, or a layer or output from a neural network, etc., which can be mapped to detect or recognize or analyze an object. These features can then be compared to features in unknown images to detect and classify unknown objects therein.

[0016] The invention can be implemented in hardware and / or software.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Further embodiments of the invention will be described with respect to the following figures. It is noted that the embodiments of Figs. 1a, 2 4a, 4b and 6 are not part of the invention, but are illustrative examples helpful for understanding the invention. In the Figures: Fig. 1a shows a schematic diagram illustrating an example of an image processing apparatus according to an embodiment; Fig. 1b shows a schematic diagram illustrating an example of an image processing apparatus according to an embodiment; Fig. 1c shows a schematic diagram illustrating an example of an image processing apparatus according to an embodiment; Fig. 2 shows a schematic diagram illustrating an example of an image processing apparatus according to an embodiment; Fig. 3 shows a schematic diagram illustrating an example of a first image and a second image of a scene as processed by an image processing apparatus according to an embodiment; Fig. 4a shows a schematic diagram illustrating a further embodiment of the image processing apparatus of figure 1a; Fig. 4b shows a schematic diagram illustrating a further embodiment of the image processing apparatus of figure 1b; Fig. 5 shows a schematic diagram illustrating a further embodiment of the image processing apparatus of figure 1a; Fig. 6 shows a schematic diagram illustrating a further embodiment of the image processing apparatus of figure 2; and Fig. 7 shows a flow diagram illustrating an example of an image processing method according to an embodiment.

[0018] In the various figures, identical reference signs will be used for identical or at least functionally equivalent features.DETAILED DESCRIPTION OF EMBODIMENTS

[0019] In the following description, reference is made to the accompanying drawings, which form part of the disclosure, and in which are shown, by way of illustration, specific aspects in which the present invention may be placed. It is understood that other aspects may be utilized and structural or logical changes may be made without departing from the scope of the present invention. The following detailed description, therefore, is not to be taken in a limiting sense, as the scope of the present invention is defined by the appended claims.

[0020] For instance, it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if a specific method step is described, a corresponding device may include a unit to perform the described method step, even if such unit is not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary aspects described herein may be combined with each other, unless specifically noted otherwise.

[0021] As will be described in more detail further below, figures 1a and 2 show two embodiments of an image processing apparatus 100, 200 based on a combination of an event-driven sensor and a standard imaging sensor, such as a CMOS sensor. The image processing apparatus 100 of figure 1a is an example for the combination of a dual-sensor set up, where each sensor has its own optics and sensor. The image processing apparatus 200 of figure 2 is an example of the combination by single-sensor set up, where the single-sensor captures both standard image and event information. In both cases, the standard image and event information can be merged, i.e. at some stage to generate a single image, i.e. in processing block 121 of figure 1a and in processing block 221 of figure 2. A further embodiment of the image processing apparatus 100 is shown in figure 1b without merging of the standard image and event information in a fusion block 121.

[0022] In the embodiments shown in figures 1a and 1b the image processing apparatus 100 can comprise one or more optical elements, in particular lenses 103, 113 for providing two distinct optical paths for the standard imaging sensor 101 and the event driven sensor 111. In a further embodiment of the image processing apparatus 100 shown in figure 1c, the image processing apparatus 100 can comprise an optical splitter or a dedicated prism 106 for splitting a single optical path into two distinct optical paths, namely one for the standard imaging sensor 101 and one for the event driven sensor 111. Specifically for the embodiment in figure 1b, the output of the event information from the event sensor 111 can be used to adjust the output of the standard imaging, e.g. RGB sensor 101. For example, based on the event identified from the output of the event sensor 111, the frame rate, resolution or other parameters of the RGB sensor 101 could be adjusted accordingly.

[0023] More specifically, figures 1a,1b and 1c show a respective schematic diagram illustrating the image processing apparatus 100 according to an embodiment configured to extract a feature from an image of a scene. As will be described in more detail further below, the image processing apparatus 100 shown in figures 1a, 1b and 1c comprises processing circuitry, in particular one or more processors, configured, in response to a feature extraction event, to extract a feature from first image data representing a first image of a scene, wherein the feature extraction event is based on second image data representing a second image of the scene.

[0024] Figure 2 shows a schematic diagram illustrating the image processing apparatus 200 according to a further embodiment configured to extract a feature from an image of a scene. As will be described in more detail further below, the image processing apparatus 200 comprises processing circuitry, in particular one or more processors, configured, in response to a feature extraction event, to extract a feature from first image data representing a first image of a scene, wherein the feature extraction event is based on second image data representing a second image of the scene. The first image and the second image of the scene are generated from a single sensor.

[0025] Figure 3 shows a schematic diagram illustrating an example of first image data 301 provided by the standard image sensor and second image data 303 provided by the event-driven sensor based on a first image and a second image of a scene as processed by the image processing apparatus 100, 200. As illustrated in figure 3, an object from the output of the event-driven sensor could be mapped to the output of the standard image sensor, which allows reducing the complexity of the computation for finding the object and improving the computer vision algorithm performance, such as face recognition rate, car license plate number recognition rate, etc., based on the accurate identification of the object. Moreover, through the exaction of features, the computer vision (CV) algorithm performance can be improved in comparison to conventional methods.

[0026] As can be taken from figures 1a, 1b and 1c and as already mentioned above, the image processing apparatus 100 may optionally comprise at least one of a standard image sensor 101 and an event-driven sensor 111 in a dual-sensor set up and / or one or more capturing devices, such as a camera and / or an event-driven camera. Each sensor 101, 111 or camera may have its own optical elements, such as one or more respective lenses 103, 113, and preprocessing blocks 102, 112 (as illustrated in the further embodiment of figure 4a), such as a respective image signal processor (ISP) 102, 112. After preprocessing, the output of each sensor 101, 111 can be calibrated, corrected and registered, for instance, by the registration block 104 of the image processing apparatus 100 shown in figure 4a. As illustrated in figures 1a and 4a, fusion of the images can be performed in a further processing block 121. However, as illustrated in figure 1b, in embodiments of the invention the fusion block 121 can be omitted.

[0027] In an alternative realization, the standard image sensor 101 and an event-driven sensor 111 in a dual-sensor set up and / or one or more capturing devices may be not part of the image processing apparatus. As already described above, the processing circuitry of the image processing apparatus 100 is configured to receive and use the output of the event-driven sensor 111, i.e. the event signal and / or the second image data for identifying an object in the output, i.e. the first image data provided by the standard image sensor 101. As illustrated in figure 4a, the processing circuitry of the image processing apparatus 100 can comprise a further processing block 122 for extracting one or more features of the identified object and / or for coding for object detection or recognition from the first image or the fusion of the first and second image in a later stage.

[0028] In an embodiment, the event-driven sensor / camera 111 may have a smaller resolution than the standard image sensor 101. In an embodiment, the output of the standard image sensor 101 can be downsampled to perform sample / pixel registration for dual sensor output fusion.

[0029] The output of the event-driven sensor 111 can comprise motion information or other video characteristic information. This information can be obtained based on the output of the event-driven sensor 111. For instance, the second image or event signal data may include a video frame with a lower resolution than the first image data. Alternatively or in addition, the second image data or event signal data may include a positive / negative amount of the intensity change and the location of the intensity change. The term location refers to the location at which the event occurs, i.e. for instance the coordinate of the respective pixel where the intensity change exceeded a predetermined threshold. The term location may alternatively or in addition refer to the time (i.e. the time stamp) at which the intensity change occurred at said pixel coordinate. The identification of an object and the feature extraction implemented in the image processing apparatus 100, 200 can comprise the following processing stages illustrated in figure 4a. A pre-determined threshold can be used to detect an object. If the intensity change is bigger than the pre-determined threshold, the sample is marked as part of an object and an event signal can be generated. In block 104 of figure 4a, a mapping from the output of the event-driven sensor 111 can be made to match the sample position in the image of the standard image sensor 101 (e.g. using known pixel registration algorithms, geometric correction could be performed based on the sensor setup). An object is identified in the image of the standard image sensor 101. The feature of the identified object is extracted by a feature extraction method using, for instance, a neural network (block 122 of figure 4a). The extracted feature can be coded and transmitted to another entity, such as a cloud server.

[0030] Figure 4b illustrates a further embodiment of the image processing apparatus 100 shown in figure 4a. In the embodiment illustrated in figure 4b the image processing apparatus 100 further comprises the optical splitter already described in the context of figure 1c.

[0031] In an alternative embodiment, the output of the image sensor 101 and the output of the event sensor 111 are not merged into a single image. However, the event detected in the event sensor 111 can be mapped to identify the object in the output of the image sensor 101, as illustrated in the embodiments of the image processing apparatus 100 shown in figures 1b and 5. After the object is identified, the feature of that object can be further processed.

[0032] As can be taken from figure 2 and as already mentioned above, for the image processing apparatus 200 of figure 2 the combination of a standard image sensor and an event-driven sensor is implemented through a single sensor and optical arrangement, such as one or more lenses 203 and a sensor 201. The sensor 201 is capable of capturing the standard image and the event information. Thus, in comparison with the image processing apparatus 100 of figure 1a, the fusion (implemented in block 221 of figure 2) of the standard image data, i.e. the first image data and the event data, i.e. the second image data is more straightforward.

[0033] The identification of an object and the feature extraction can be done as illustrated in figure 6. A pre-determined threshold can be used to detect an object in the second image data. If the intensity change is bigger than the pre-determined threshold, the sample is marked as part of an object and an event signal is generated (block 202 of figure 6). A fusion of the image sensor data, i.e. the first image data 301 and the event sensor data, i.e. the second image data 303 can be performed (block 221 of figure 6). A mapping from the output of the event vision sensor 103 can be made in block 204 of figure 6 to match the sample position in the image 301 of the standard image sensor 201 using, for instance, known pixel registration algorithms. An object is identified in the image of the standard image sensor 201. The feature(s) of the identified object is extracted by a feature extraction method, such as a neural network (block 222 of figure 6). The extracted feature can be coded and transmitted to another entity, such as a cloud server.

[0034] Figure 7 shows a flow diagram illustrating an example of an image processing method 700 according to an embodiment, which summarizes some of the steps already described above. In a step 701, content is captured by the event-based sensor 103, 201 and a standard imaging sensor 101, 201. In step 703, an event is detected on the basis of the data captured by the event-based sensor 103, 201, i.e. the second image data and an event signal can be generated. In step 705, the object is detected, i.e. identified in the image captured by the standard imaging sensor 101, 201, i.e. the first image data, on the basis of the event signal. In step 707, one or more features of the object identified in the first image data 301 are extracted therefrom. In step 709, the extracted features can be coded.

[0035] To the extent that the terms "include", "have", "with", or other variants thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprise". Also, the terms "exemplary", "for example" and "e.g." are merely meant as an example, rather than the best or optimal. The terms "coupled" and "connected", along with derivatives may have been used. It should be understood that these terms may have been used to indicate that two elements cooperate or interact with each other regardless whether they are in direct physical or electrical contact, or they are not in direct contact with each other.

[0036] Although the elements in the following claims are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0037] Many alternatives, modifications, and variations will be apparent to those skilled in the art in light of the above teachings. Of course, those skilled in the art readily recognize that there are numerous applications of the invention beyond those described herein.

Claims

1. An image processing apparatus (100; 200) for extracting a feature from an image of a scene, wherein the apparatus (100; 200) comprises a standard imaging sensor (101, 201) configured to capture a first image of a scene and providing first image data (301); an event-driven sensor (111, 201) configured to capture a second image of the scene and providing second image data (303); and the apparatus (100; 200) further comprises processing circuitry configured, in response to a feature extraction event, to extract a feature from the first image data (301), the first image data (301) representing the first image of the scene, wherein the feature extraction event is based on the second image data (303), the second image data (303) representing the second image of the scene; wherein the second image is one of a plurality of images of a video stream and wherein the processing circuitry is further configured to determine motion statistics and sample value change information in the second image data (303) and to determine on the basis of the motion statistics and the sample value change information whether an event occurs.

2. The apparatus (100; 200) of claim 1, wherein the processing circuitry is configured to determine on the basis of the motion statistics and / or the sample value change information whether the event occurs by comparing the motion statistics and / or the sample value change information with one or more threshold values.

3. The apparatus (200) of any one of the preceding claims, wherein the standard imaging sensor (201) and the event-driven sensor (201) are both implemented in an imaging sensor (201) and an optical arrangement, the imaging sensor (201) being configured to capture an image of the scene comprising a plurality of sample values and to provide the first image data (301) as a first subset of the plurality of sample values and the second image data (303) as a second subset of the plurality of sample values.

4. The apparatus (200) of claim 3, wherein the apparatus (200) further comprises a spatial filter (205) configured to spatially filter the image before a sensor (201) for providing the first image data (301) as the first subset of the plurality of sample values and the second image data (303) as the second subset of the plurality of sample values.

5. The apparatus (100) of claim 1, wherein the first image captured by the standard imaging sensor (101) has a higher resolution than the second image captured by the event-driven sensor (111).

6. The apparatus (100) of claim 5, wherein the processing circuitry is further configured to downsample the first image to the lower resolution of the second image.

7. The apparatus (100) of any one of claims 5 to 6, wherein the standard imaging sensor (101) is a CMOS sensor and the second imaging sensor (111) is an event sensor.

8. The apparatus (100) of any one of claims 5 to 7, wherein the processing circuitry is further configured to identify in the first image data (301) the feature for which the feature extraction event occurs on the basis of the second image data.

9. The apparatus (100) of any one of claims 5 to 8, wherein the first image data (301) comprises a first plurality of sample values and the second image data (303) comprises a second plurality of sample values and wherein the processing circuitry is further configured to map the first plurality of sample values with the second plurality of sample values.

10. An image processing method for extracting a feature from an image of a scene, wherein the method comprises: capturing, with a standard imaging sensor (101, 201), a first image of a scene and providing first image data (301); capturing, with an event-driven sensor (111, 201), a second image of the scene and providing second image data (303); and extracting, with a processing circuitry, in response to a feature extraction event, a feature from the first image data (301), the first image data (301) representing the first image of a scene, wherein the feature extraction event is based on second image data (303), the second image data (303) representing the second image of the scene; wherein the second image is one of a plurality of images of a video stream and wherein the method further comprises determining, with the processing circuitry, motion statistics and sample value change information in the second image data (303) and to determine on the basis of the motion statistics and the sample value change information whether an event occurs.

11. A computer program product comprising instructions to cause the device of claim 1 to execute the steps of the method of claim 10.