Image Processing System
A multimodal CNN fuses frame-based and event cameras to overcome temporal resolution limitations and motion blur, ensuring accurate driver monitoring in DMS by leveraging the strengths of both camera types.
Patent Information
- Application Number
- JP2023542985
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-13
- Filing Date
- 2021-10-14
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-10-14
AI Technical Summary
Existing image processing systems using frame-based cameras suffer from limited temporal resolution and motion blur during high-speed movements, while event cameras are ineffective for stationary objects, limiting their effectiveness in applications like driver monitoring systems (DMS).
A multimodal convolutional neural network (CNN) fuses information from both frame-based and event cameras, generating sensor attention maps across intermediate layers to leverage the strengths of both modalities, allowing asynchronous inference based on event camera outputs to adapt to scene dynamics.
The system accurately analyzes facial features for DMS, enabling accurate driver state detection during normal driving and critical events like collisions, enhancing safety by allowing for injury estimation and autonomous system intervention.
Smart Images

Figure 0007770412000001 
Figure 0007770412000002 
Figure 0007770412000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing system. [Background technology]
[0002] Using a multimodal fusion architecture to fuse information from multiple different sensors not only improves performance compared to single-sensor-based architectures, as sensor fusion from different sensors can take advantage of the advantages of individual sensors in the system and minimize their disadvantages, but also provides a higher degree of redundancy than duplicating sensors of the same type.
[0003] C. Zhang, Z. Yang, X. He, and L. Deng, "Multimodal intelligence: Representation learning, information fusion, and applications," IEEE J. Sel. Top. Signal Process, 2020, discloses integrating information from different unimodal sensors into a single representation.
[0004] J.-M. Perez-Rua, V. Vielzeuf, S. Pateux, M. Baccouche, and F. Jurie, "MFAS: Multimodal fusion architecture search," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 6966-6975, describes a co-attention mechanism in which the network determines how to weight different modalities based on contextual information.
[0005] RA Jacobs, MI Jordan, SJ Nowlan and GE Hinton, "Adaptive mixtures of local experts," Neural Comput., Vol. 3, No. 1, pp. 79-87, 1991, discloses a core tension mechanism in which information is fused at the decision-level.
[0006] J. Arevalo, T. Solorio, M. Montes-y-Gomez, and F.A. Gonzalez, "Gated multimodal units for information fusion," arXiv Prepr.arXiv1702.01992, 2017, and J. Arevalo, T. Solorio, M. Montes-y-Gomez, and F.A. Gonzalez, "Gated multimodal networks," Neural Comput. Appl., pp. 1–20, 2020, propose a gated multimodal unit (GMU) that uses image and text inputs to enable feature-level fusion at any level within the network. The GMU can learn a latent variable that determines which modality contains useful information for a particular input.
[0007] A. Valada, A. Dhall, and W. Burgard, "Convoluted mixture of deep experts for robust semantic segmentation," IEEE / RSJ International conference on Intelligent Robots and Systems (IROS) workshop, State Estimation and Terrain Perception for All Terrain Mobile Robots, 2016, p. 23, proposes a network that includes an adaptive gating network that determines when and to what extent to rely on each "expert" (modality).
[0008] V. Vielzeuf, A. Lechervy, S. Pateux, and F. Jurie, "Centralnet: a multilayer approach for multimodal fusion," Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 575-589, discloses a multimodal network architecture that fuses information from individual networks for each modality at multiple layers.
[0009] R. Ranjan, S. Sankaranarayan, CD Castillo, and R. Chellappa, "An all-in-one convolutional neural network for face analysis," in 2017, 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), pp. 17-24; and R. Ranjan, VMPatel, and R. Chellappa, "Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition," in IEEE Trans. Pattern Anal. Mach. The All-in-One and Hyperface-ResNet network architectures disclosed in Intell., Vol. 41, No. 1, pp. 121-135, 2017, respectively apply the fusion of hidden layers in neural networks.
[0010] There is limited literature on multimodal fusion with event cameras. S. Pini, G. Borghi, and R. Vezzani, "Learn to see by events: Color frame synthesis from event and RGB cameras," International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2020, Vol. 4, pp. 37-47, feed the network with concatenated RGB and event as two input channels. Event frames are formed using a fixed time window, which eliminates many of the key properties of event cameras, namely temporal resolution and responsiveness to fast motion. [Prior art documents] [Patent documents]
[0011] [Patent Document 1] European Patent No. 3440833 [Patent Document 2] International Publication No. 2019 / 145516 [Patent Document 3] International Publication No. 2019 / 180033 [Patent Document 4] U.S. Patent Application Publication No. 16 / 904,122 [Patent Document 5] U.S. Patent Application Publication No. 16 / 941,799 [Patent Document 6] U.S. Patent Application Serial No. 17 / 037,420 [Patent Document 7] International Publication No. 2019 / 145578 [Patent Document 8] U.S. Patent Application Publication No. 16 / 544,238 [Non-patent literature]
[0012] [Non-Patent Document 1] C. Zhang, Z. Yang, X. He, and L. Deng, "Multimodal intelligence: Representation learning, information fusion, and applications," IEEE J. Sel. Top. Signal Process., 2020. [Non-patent document 2] J.-M. Perez-Rua, V. Vielzeuf, S. Pateux, M. Baccouche, and F. Jurie, "MFAS: Multimodal fusion architecture search," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 6966-6975. [Non-patent document 3] RA Jacobs, MI Jordan, SJ Nowlan, and GE Hinton, "Adaptive mixtures of local experts," Neural Comput., Vol. 3, No. 1, pp. 79-87, 1991. [Non-patent document 4] J. Arevalo, T. Solorio, M. Montes-y-Gomez, and F.A. Gonzalez, "Gated multimodal units for information fusion," arXiv Prepr.arXiv1702.01992, 2017. [Non-Patent Document 5] J. Arevalo, T. Solorio, M. Montes-y-Gomez, and F.A. Gonzalez, "Gated multimodal networks," Neural Comput. Appl., pp. 1–20, 2020. [Non-patent document 6] A. Valada, A. Dhall, and W. Burgard, "Convoluted mixture of deep experts for robust semantic segmentation," IEEE / RSJ International conference on Intelligent Robots and Systems (IROS) workshop, State Estimation and Terrain Perception for All Terrain Mobile Robots, 2016, p. 23 [Non-Patent Document 7] V. Vielzeuf, A. Lechervy, S. Pateux, and F. Jurie, "Centralnet: a multilayer approach for multimodal fusion," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 575-589. [Non-patent document 8] R. Ranjan, S. Sankaranarayan, CD Castillo, and R. Chellappa, "An all-in-one convolutional neural network for face analysis," in 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), pp. 17-24. [Non-Patent Document 9] R. Ranjan, VMPatel, and R. Chellappa, "Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition," IEEE Trans. Pattern Anal. Mach. Intell., Vol. 41, No. 1, pp. 121-135, 2017. [Non-Patent Document 10] S. Pini, G. Borghi, and R. Vezzani, "Learn to see by events: Color frame synthesis from event and RGB cameras," International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2020, vol. 4, pp. 37-47. [Non-Patent Document 11] Posch, C., Serrano-Gotarredona, T., Linares-Barranco, B., and Delbruck, T., "Retinomorphic event-based vision sensors: bioinspired cameras with spiking output," Proceedings of the IEEE, 102(10), 1470-1484, (2014) [Non-Patent Document 12] Scheerlinck, C., Rebecq, H., Gehrig, D., Barnes, N., Mahony, R., and Scaramuzza, D., 2020, "Fast image reconstruction with an event camera," IEEE Winter Conference on Applications of Computer Vision, pp. 156-163. Summary of the Invention [Means for solving the problem]
[0013] According to the present invention, there is provided an image processing system as set forth in claim 1.
[0014] In a second aspect, there is provided an image processing method according to claim 16, and a computer program product arranged to carry out this method.
[0015] Embodiments of the present invention may include a multimodal convolutional neural network (CNN) that fuses information from frame-based cameras, such as near-infrared (NIR) cameras and event cameras, to analyze facial features to generate classifications such as head pose or gaze.
[0016] Frame-based cameras have limited temporal resolution compared to event cameras and are therefore prone to blurring when objects in the camera's field of view are moving at high speeds, while event cameras are best suited to object motion but do not produce information when objects are stationary.
[0017] Embodiments of the present invention take advantage of the benefits of both worlds by fusing the intermediate layers of a CNN and assigning importance to each sensor based on the inputs provided.
[0018] Embodiments of the present invention generate sensor attention maps from multiple levels of intermediate layers throughout the network.
[0019] Embodiments are particularly well suited for driver monitoring systems (DMS). NIR is a standard camera often used in DMS. These standard frame-based cameras are prone to motion blur, which is particularly noticeable during vehicle crashes or other safety-critical high-speed events. In contrast, event cameras can adapt to scene dynamics and accurately track the driver with very high temporal resolution. However, event cameras are not particularly well suited for monitoring slow-moving or stationary objects, for example, to determine driver attentiveness.
[0020] Embodiments fuse both modalities into a unified CNN that can capture the advantages and minimize the disadvantages of each modality, allowing the network to accurately analyze normal driving and rare events such as collisions when implemented in a DMS.
[0021] Furthermore, inference can be performed asynchronously based on the output of the event camera, thus allowing the network to adapt to scene dynamics rather than running at a fixed rate.
[0022] This enables the DMS to sense and understand the driver's state during a vehicle collision, allowing for accurate injury estimation or autonomous system intervention.
[0023] The embodiments can be applied to other tasks, including external monitoring for autonomous driving purposes, such as vehicle / pedestrian detection and tracking, as well as DMS.
[0024] Embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0025] [Figure 1] FIG. 1 illustrates a system for fusing information provided by a frame-based NIR camera and an event camera, according to an embodiment of the present invention. [Figure 2] FIG. 2 illustrates a network for multimodal face analysis used in the system of FIG. 1. [Figure 3] FIG. 2 illustrates facial landmarks that can be detected in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] FIG. 1 illustrates an image processing system 10 according to an embodiment of the present invention. System 10 includes a frame-based camera 12, which in this case is a camera sensitive to near-infrared (NIR) wavelengths and generates frames of information at periodic intervals, such as at a rate of typically 30 frames per second (fps) to possibly up to 240 fps. It will be understood that the frame rate may vary over time depending on, for example, context or environmental conditions—higher frame rates may be impossible or inappropriate, for example, under low light conditions—and that the data acquired by camera 12 and provided to the rest of the system generally includes frames of information spanning the camera's entire field of view, regardless of any activity within the field of view. It will also be understood that in alternative implementations, the frame-based camera may be sensitive to other wavelengths, such as visible wavelengths, and provide either monochromatic intensity-only or multicolor frame information in any suitable format, including RGB, YUV, LCC, or LAB formats.
[0027] The system may also include an event camera 14 of the type disclosed, for example, in Posch, C., Serrano-Gotarredona, T., Linares-Barranco, B., and Delbruck, T., "Retinomorphic event-based vision sensors: bioinspired cameras with spiking output," Proceedings of the IEEE, 102(10), 1470-1484, (2014), EP 3440833, and WO 2019 / 145516 and WO 2019 / 180033 from Prophesee. Such cameras are based on asynchronously outputting image information from individual pixels whenever a change in pixel value exceeds a certain threshold. Thus, the pixels of the "event camera" report an asynchronous "event" stream of intensity changes characterized by the x, y position, timestamp, and polarity of the intensity change.
[0028] The event camera 14, like the frame camera 12, is sensitive to NIR or visible wavelengths and can provide monochromatic or multicolor RGB event information, and the like.
[0029] Events can occur asynchronously, possibly as frequently as the image sensor's clock period, and the minimum period over which an event can occur is referred to herein as the "event period."
[0030] When employed in a driver monitoring system (DMS), each of the cameras 12, 14 may be mounted on or near the rearview mirror facing forward of the vehicle cabin and facing rearward toward the cabin occupants.
[0031] The cameras 12, 14 may be spaced somewhat apart, and this stereoscopic perspective may assist in certain tasks such as detecting the occupant's head pose, as described in more detail below.
[0032] Nevertheless, it will be understood that the respective fields of view of cameras 12, 14 will generally overlap substantially to the extent that they can each capture the faces of one or more occupants of interest within the vehicle when in a range of typical positions within the cabin.
[0033] It should nevertheless be understood that the cameras 12, 14 need not be separate units, and that in some implementations a single integrated sensor, such as the Davis346 camera available at iniVation.com, may be used to provide the functionality of a frame camera and an event camera, which, of course, can reduce the need for a dual optical system.
[0034] As explained, the event camera 14 provides a stream of individual events as they occur, rather than frames of information as provided by the camera 12 .
[0035] In an embodiment of the present invention, this event information acquired and provided by the event camera 14 is accumulated by an event accumulator 16, which uses the accumulated event information to reconstruct texture-type image information that is provided in an image frame format 18 for further processing by the system.
[0036] Well-known neural network-based event camera reconstruction methods include E2VID and Firenet, which provide image frame information from event information, as described in Scheerlinck, C., Rebecq, H., Gehrig, D., Barnes, N., Mahony, R., and Scaramuzza, D., 2020, “Fast image reconstruction with an event camera,” IEEE Winter Conference on Applications of Computer Vision (pp. 156-163).
[0037] Further examples of methods and systems for accumulating event information and providing frame information are disclosed in U.S. Patent Application No. 17 / 037,420, entitled "Object Detection for Event Cameras," filed September 29, 2020, which is a continuation-in-part of U.S. Patent Application No. 16 / 941,799, filed July 29, 2020, which is a continuation-in-part of U.S. Patent Application No. 16 / 904,122 (Reference No. FN-662-US), filed June 17, 2020. These systems can identify regions of interest, such as face regions, within the field of view of the event camera and generate textured image frames of the face regions once a specified number of events, such as 20,000, have accumulated within the face regions.
[0038] Using one such method, the event accumulator 16 maintains a count of events occurring at each pixel location within the facial region within a time window during which the event accumulator 16 captures events to provide an image frame 18. The net polarity of the events occurring at each pixel location during this time window is determined to generate a decay factor for each pixel location as a function of the count. This decay factor is applied to a texture image generated for the facial region prior to the current time window, and the net polarity of the events occurring at each pixel location is added to the corresponding location in the decayed texture image to generate a texture image for the current time window. This allows the pixels in the frame 18 provided by the event accumulator 16 to maintain information over a time window as a form of motion memory, whereas the image frame generated by the camera 12 contains only relatively instantaneous information from the frame's exposure window.
[0039] In addition to or instead of using counts to generate decay factors, events can also be decayed as a function of time as image frames 18 are accumulated.
[0040] These methods are particularly useful for this application because in DMS systems, the location and characteristics of the vehicle occupant's facial region, such as facial landmarks, head pose, gaze and any occlusion, are of primary interest.
[0041] Nevertheless, in some embodiments of the present invention, event information for one or more facial regions in the field of view of the event camera 14 can alternatively or additionally be accumulated in each image frame by using a detector with frame information from the frame-based camera 12 to identify one or more regions of interest, such as facial regions, in the field of view of the camera 12 and mapping these regions of interest to corresponding regions in the field of view of the camera 14 taking into account the spatial relationship of the cameras 12, 14 and their respective camera models.
[0042] It will be understood that frames acquired from a frame-based camera 12 arrive periodically according to a frame rate configured for the camera 12. However, the event accumulator 16 can generate frames 18 asynchronously, theoretically with a time resolution as small as one event cycle.
[0043] If there is a large amount of motion within the region of interest within the field of view of the event camera 14, the event accumulator 16 can generate frames 18 very frequently, more frequently than would be generated by the frame-based camera 12 anyway.
[0044] In an embodiment of the present invention, the neural network 20 of FIG. 2 is applied to the most recent frame provided by the accumulator 16 and the most recent frame provided by the frame-based camera 12 .
[0045] This means that, given sufficient object motion, the neural network 20 can be run multiple times before a new NIR image frame is provided by the frame camera 12 .
[0046] Nevertheless, in some embodiments of the present invention, if an updated frame 18 is not provided by the event accumulator 16 within the interval between frames provided by the camera 12, the network 20 may be re-run using the latest NIR image and request the last available frame generated by the event accumulator 16, or request the event accumulator 16 to generate a frame of the required region of interest based on whatever events, if any, have been generated since that last frame.
[0047] This means that network 20 will operate at the lowest frame rate of camera 12, regardless of motion. Thus, for a camera 12 operating at 30 fps, network 20 will run if the elapsed time required to accumulate 20,000 events is greater than 0.033 seconds (equivalent to 30 fps).
[0048] In either case, the most recently acquired NIR imagery is used while the network 20 reacts to the scene dynamics provided by the event camera 14. As a result, the event image frames tend to be "ahead of time" and essentially represent NIR imagery plus motion.
[0049] While this time lag between the image frames provided by camera 12 and the image frames provided by event accumulator 16 may be considered problematic, it will be understood from the following discussion that this time lag does not adversely affect the techniques of the present application when determining the characteristics of any faces detected within the fields of view of cameras 12, 14.
[0050] Referring now in more detail to FIG. 2, network 20 includes two inputs: a first input for receiving image frames corresponding to facial regions detected by frame camera 12, and a second input for receiving image frames corresponding to facial regions detected within frames provided by event accumulator 16.
[0051] As can be seen, each input contains an intensity-only 224x224 image, and therefore each face region image frame needs to be upsampled / downsampled (normalized) as necessary before being provided to the network 20 so that it is provided at the correct resolution.
[0052] Each input image frame is processed by a network containing four successive blocks of two or three convolutional layers, with each of blocks i=1 to 3 followed by a maxpooling layer, similar to an incomplete VGG-16 network.
[0053] In Figure 2, where applicable, each block displays the kernel size (3x3), layer type (Conv / MaxPool), number of output filters (32, 64, 128, 256), and kernel stride (\2).
[0054] Note that while the structure of the network processing each input is the same, the weights used within the kernels of each convolutional layer may not match, and as will be appreciated, these are learned during training of the network.
[0055] The intermediate output of block i=1 (x v , x t ) are fused into a simple convolution 22 to produce the fused output hi.
[0056] The intermediate outputs of blocks i=2, 3, and 4 are the fused outputs of their previous blocks (h i-1 ), are fused using respective Gated Multimodal Units (GMUs) 24-2, 24-3, 24-4. Each GMU 24 contains an array of GMU cells of the type proposed in Areval et al., cited above, and shown in detail on the right side of Figure 2. Each GMU cell receives vectors xv, xt, and h i-1 is connected to each element of the cell, and the fusion output h i Generate h v =tanh(W v x v ) h t =tanh(W t x t ) z=σ(W z ·[x v ,x t ]) hi =h i-1 *(z*h v +(1-z)*h t ) and {W v ,W t ,W z} are the trained parameters, [·,·] represents the concatenation operator, σ is the cell h i feature x for the entire output of v ,x t represents the gate neuron that controls the contribution of
[0057] These GMUs 24 allow the network 20 to combine modalities and weight modalities that are likely to give better estimates.
[0058] Thus, for example, in a scene experiencing a lot of motion, it is expected that the image frames provided by camera 12 will tend to be blurry and exhibit low contrast. Any one or more frames provided by the event accumulator during or after the capture of such a blurry frame should be sharp, and network 20 can therefore be trained using a suitable training set to prioritize information from such frames from the event camera side of the network in these situations.
[0059] On the other hand, during times of low motion, the GMU 24 will tend to weight high contrast images from the frame camera 12 much more heavily than the last available image frame provided by the event accumulator 16, and thus will tend not to weight this image information as strongly, even if an image frame of unknown age is available when processing a sharp image from the frame camera 12. In any event, the less motion there is in the scene, the less likely any image information from the event camera 14 will corrupt the information available from the frame camera 12.
[0060] Furthermore, since the output of each convolution block xv, xt, as well as the fused output hi of the convolution layer 22 and the GMU 24, contain respective vectors whose individual elements are interconnected, this gives the network the possibility to respond differently to different spatial regions of each image provided by the camera 12 and the event accumulator 16.
[0061] 2, the sensor information is combined and weighted at four different levels within the network by the convolutional layers 22 and GMUs 24. This allows for emphasis on low-level features of one sensor 12, 14 and high-level features of another sensor, or vice versa. Nevertheless, it will be appreciated that the network architecture can be modified to enhance real-time performance and apply only one or two GMU fusions in earlier layers to reduce computational cost.
[0062] Note that convolutions 26-1 and 26-2 are performed between the outputs and inputs of GMUs 24-2 and 24-3, and between the outputs and inputs of GMUs 24-3 and 24-4 to match the downsampling of the intermediate outputs of blocks 3 and 4.
[0063] After the final feature fusion in GMU 24-4, a 1x1 convolution 28 is used to reduce the dimensionality of the feature vector provided by the final GMU.
[0064] In this embodiment, the feature vector provided by the convolution 28 may be fed into one or more separate task-specific channels.
[0065] An exemplary general structure of such a channel is shown in the top right corner of FIG.
[0066] In general, each such channel may include one or more further convolutional layers followed by one or more fully connected (fc) layers, with one or more nodes in the final fully connected layer providing the required output.
[0067] Exemplary facial features that can be determined using this structure include, but are not limited to, head pose, gaze, and occlusion.
[0068] Head pose and gaze can be represented using 3(x,y,z) output layer nodes for head pose, corresponding to the pitch, yaw and roll angles of the head, respectively, and 2(x,y) output layer nodes for gaze, corresponding to the yaw and pitch angles of the eyes.
[0069] Accurate estimation of head pose allows for the calculation of the angular velocity of the head, so for example knowing the initial orientation of the head during a collision provides context for the DMS to take more intelligent action in the event of a collision.
[0070] Gaze angle can provide information about whether a driver anticipated a collision. For example, a driver looking in the rearview mirror during a rear-end collision can indicate awareness of a potential collision. By tracking pupil saccades toward the impacting object, the system can calculate reaction time and whether advanced driver assistance systems (ADAS) intervention, such as autonomous emergency braking, is required.
[0071] Note that both head pose and gaze are determined with respect to the face as it appears in the face region image provided to the network and are therefore relative. Providing absolute head position or gaze angle requires knowledge of the relationship between the image plane and the cameras 12, 14.
[0072] The occlusion can be indicated by respective output nodes (x or y) corresponding to one indicating eye occlusion (if the occupant is wearing glasses) and one indicating mouth occlusion (if the occupant appears to be wearing a mask).
[0073] Other forms of facial features include facial landmarks, such as those described in International Publication No. WO 2019 / 145578 (Reference No. FN-630-PCT) and U.S. Patent Application Publication No. 16 / 544,238, filed August 19, 2019, entitled "Method of image processing using a neural network," which include the locations of a series of interest points around the facial region, as shown in FIG. 3, for example.
[0074] However, to generate such landmarks, the feature vectors generated by convolutional layer 24 may also be beneficially provided to a decoder network and a fully connected network of the type of network disclosed in U.S. Patent Application Publication No. 16 / 544,238, entitled "Image Processing Method Using Neural Networks," filed August 16, 2019 (Reference: FN-651-US), the disclosure of which is incorporated herein by reference.
[0075] In terms of training, the learning efficiency and performance of individual tasks, in this case gaze, head pose, and face occlusion, are enhanced by multi-task learning.
[0076] It is desirable to take into account the limitations of the NIR camera 12 (blur) and the limitations of the event camera 14 (no motion) so that the network 20 learns during training when to trust one camera over another. Therefore, the following augmentation methods can be employed: 1. To encourage reliance on the IR camera 12, the number of events in some parts of the training set can be limited to reflect limited motion. The attention mechanism should favor the NIR camera when the number of events is low. 2. To facilitate reliance on the vent camera, random object blurring can be applied to other parts of the NIR camera 12 training set to reflect very fast object motion, which should be the case during impacts where NIR is sensitive to blurring and lacks time resolution.
[0077] While the above embodiment has been described in terms of fusing two modalities, it will be appreciated that the network 20 can be extended to fuse more than two modalities by extending the cells of the convolutional layer 22 and GMU 24 to fuse more than two inputs. [Explanation of symbols]
[0078] 10 Image Processing System 12 Frame-based cameras 14 Event Camera 16 Event Accumulator 18 Image Frame Formats 20 Neural Networks
Claims
1. An image processing system (10), comprising: a frame-based camera (12) configured to periodically provide image frames covering the field of view of the camera; an event camera (14) having a field of view overlapping with the field of view of the frame-based camera and configured to provide event information in response to an event indicating that a change in light intensity detected at an x, y location within the field of view exceeds a threshold; a detector that identifies a region of interest within the overlapping fields of view; an accumulator (16) for accumulating event information from a plurality of events occurring during successive event cycles within the region of interest, each event indicating an x,y location within the region of interest, a polarity of a change in detected light intensity incident at said x,y location, and the event cycle in which the event occurred, and for generating image frames of the region of interest from the accumulated event information from within the region of interest in response to an event criterion for the region of interest being met, at a frequency that varies depending on the rate of occurrence of events as a function of an amount of movement within the region of interest; a neural network (20) configured to receive image frames from the frame-based camera and image frames of the region of interest from the accumulator, configured to process each image frame through a plurality of convolutional layers (Blocks 1-4) to provide a respective set of one or more intermediate images, and further configured to fuse at least one corresponding pair of intermediate images generated from each of the image frames through an array of fusion cells (22-28), each fusion cell connected to at least a respective element of each intermediate image and trained to weight each element from each intermediate image to provide a fusion output, the neural network further comprising at least one task network configured to receive the fusion output from a final set of intermediate images and generate one or more task outputs for the region of interest; Equipped with the detector is configured to identify a region of interest within the image frames provided by either the frame-based camera or the event camera; system.
2. each fusion cell of said intermediate image pair is further connected to a respective element of the fusion output (hi-1) from the previous intermediate image pair; The system of claim 1 .
3. Each fused cell is h v =tanh(W v ・x v ) h t =tanh(W t ・x t ) z=σ(W z ・[x v ,x t ]) h i =h i-1 *(z*h v +(1-z)*h t ) configured to generate a fusion output hi of the fusion cell according to a function of x v , x t is the element value of each intermediate image, {W v , W t , W z } are the trained parameters, h i-1 is the element value of the fusion output from the previous intermediate image pair, [・,・] indicates the concatenation operator, σ represents the gate neuron, The system of claim 1 .
4. The neural network is configured to fuse the first set of intermediate images through a convolutional layer (22). The system of claim 2 .
5. The neural network further includes one or more pooling layers between the plurality of convolutional layers. The system of claim 1 .
6. and further configured to match a resolution of the image frames from the frame-based camera and the image frames from the accumulator to a size required by the neural network. The system of claim 1 .
7. the region of interest includes a face region; The system of claim 1 .
8. a respective task network providing one of an indication of head pose, gaze, or facial occlusion; The system of claim 7.
9. Each task network includes one or more convolutional layers followed by one or more fully connected layers, the output layer of the head pose task network includes three output nodes, the output layer of the gaze task network includes two output nodes, and the output layer of the facial occlusion task network includes an output node for each type of occlusion. The system of claim 8.
10. a task network for providing a set of facial landmarks for the facial region; The system of claim 7.
11. the frame-based camera is sensitive to near-infrared (NIR) wavelengths; The system of claim 1 .
12. 10. A driver monitoring system comprising the image processing system of claim 1, wherein the image processing system is configured to provide the one or more task outputs to an advanced driver assistance system (ADAS). Driver monitoring system.
13. 1. A method of image processing operable by a system comprising a frame-based camera configured to periodically provide image frames covering a field of view of the camera, and an event camera having a field of view overlapping with the field of view of the frame-based camera and configured to provide event information in response to an event indicating a change in light intensity detected at an x, y location within the field of view of the event camera exceeds a threshold, the method comprising: identifying a region of interest within the overlapping fields of view; accumulating event information from a plurality of events occurring during successive event cycles within the region of interest, each event indicating an x,y location within the region of interest, a polarity of a change in detected light intensity incident at said x,y location, and an event cycle in which said event occurred; and generating image frames of the region of interest from the accumulated event information from within the region of interest in response to an event criterion for the region of interest being met, at a frequency that varies depending on a rate of occurrence of events depending on an amount of movement within the region of interest; receiving image frames from the frame-based camera and image frames of the region of interest from an accumulator, the region of interest being identified within the image frames provided by either the frame-based camera or the event camera; processing each image frame through multiple convolutional layers of a neural network to provide a respective set of one or more intermediate images; fusing at least one corresponding pair of intermediate images generated from each of the image frames through an array of fusion cells of the neural network, each connected to at least a respective element of each intermediate image, and trained to weight each element from each intermediate image to provide a fused output; receiving the fused output from a final pair of intermediate images; generating one or more task outputs of the neural network for the region of interest; A method comprising:
14. A computer readable medium comprising computer readable instructions stored on the computer readable medium, the computer readable instructions comprising instructions for performing the steps of claim 13 when executed on a computing device.
Citation Information
Patent Citations
Sample and hold based temporal contrast vision sensor
EP3440833A1
Event camera
JP2020182122A
Object detection for event cameras
US11164019B1
Event detector and method of generating textural image based on event count decay factor and net polarity
US11270137B2
Object detection for event cameras
US11301702B2