Multitasking perception network with applications for scene understanding and an advanced driver assistance system

A multitasking CNN addresses the computational inefficiencies and feature neglect in existing systems by extracting shared features and simultaneously performing multiple perception tasks, enhancing the efficiency and effectiveness of scene understanding and advanced driver assistance systems.

DE112020001103B4Active Publication Date: 2025-06-26NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112020001103
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-11
Filing Date
2020-02-12
Publication Date
2025-06-26
Estimated Expiration
2040-02-12

AI Technical Summary

Technical Problem

Existing scene understanding systems and advanced driver assistance systems require separate convolutional neural networks for each perception task, leading to high computational resource usage and neglect of mutual features between tasks.

Method used

A multitasking convolutional neural network (CNN) is used to extract shared features across different perception tasks, such as object detection and semantic segmentation, and simultaneously solve these tasks in a single pass, reducing computational requirements and leveraging shared features.

Benefits of technology

The proposed solution efficiently processes multiple perception tasks using a single GPU, explores shared features between tasks, and provides a parametric representation of driving scenes, enabling effective collision avoidance in advanced driver assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Computer-implemented method in an advanced driver assistance system (ADAS), comprising: Extracting (505), by a hardware processor, from an input video stream containing a plurality of images, using a multitasking convolutional neural network (CNN), shared features across different perceptual tasks, wherein the different perceptual tasks include object detection and other perceptual tasks; simultaneously solving (510), by the hardware processor, using the multitasking CNN, the different perception tasks in a single pass by simultaneously processing corresponding ones of the shared features by respectively different branches of the multitasking CNN to provide a plurality of outputs of different perception tasks, each of the respectively different branches corresponding to a respective one of the different perception tasks; Forming (530) a parametric representation of a driving scene as at least one top-view map in response to the plurality of outputs from different perception tasks; and Controlling operation of the vehicle to avoid collisions in response to at least one top-view map indicating an impending or imminent collision.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION ON RELATED APPLICATIONSThis application claims priority to U.S. Patent Application No. 16 / 787,727, filed February 11, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 814,886, filed March 7, 2019, the contents of which are incorporated herein by reference in their entirety.BACKGROUNDTechnical FieldThe present invention relates to machine learning and, more particularly, to a multi-tasking perception network having applications for scene understanding and an advanced driver assistance system.DESCRIPTION OF THE RELATED ARTMany scene understanding systems and advanced driver assistance systems require performing a variety of perception tasks, such as object recognition, semantic segmentation, and depth estimation, which are normally considered separate modules and implemented as independent convolutional neural networks (CNN). However, there are some disadvantages with the above approach. First, it requires many computing resources, e.g., a graphics processing unit (GPU) is needed for operation of a task-specific network. Second, it ignores mutual features between individual perception tasks, such as object detection and semantic segmentation. Therefore, there is a need for an improved approach for using a multi-tasking perception network for scene understanding and advanced driver assistance systems.SUMMARYAccording to an aspect of the present invention, a computer-implemented method is provided in an advanced driver assistance system (ADAS (= Adv Driver-Assistance System)). The method includes extracting, by a hardware processor, shared features across different perception tasks from an input video stream including a plurality of images using a convolutional neural network (CNN) with multi-tasking. The different perception tasks include object detection and other perception tasks. The method further includes concurrently resolving, by the hardware processor, the different perception tasks using the multi-tasking CNN in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks. Each of the respective different branches corresponds to a respective one of the different perception tasks. The method also includes forming a parametric representation of a driving scene as at least one top-view map in response to the plurality of outputs of different perception tasks. The method additionally includes controlling operation of the vehicle for collision avoidance in response to at least one top view map indicating an imminent collision.According to another aspect of the present invention, a computer program product for advanced driver assistance is provided. The computer program product includes a non-transitory computer readable storage medium having program instructions embodied therewith. The program instructions are executable by a computer to cause the computer to perform a method. The method includes extracting shared features across different perception tasks from an input video stream including a plurality of images using a convolutional neural network (CNN) with multi-tasking. The different perception tasks include object detection and other perception tasks. The method further includes concurrently resolving the different perception tasks using the multi-tasking CNN in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks. Each of the respective different branches corresponds to a respective one of the different perception tasks. The method also includes forming a parametric representation of a driving scene as at least one top view map in response to the plurality of outputs of different perception tasks by the hardware processor. The method additionally includes controlling, by the hardware processor, operation of the vehicle for collision avoidance in response to at least one top-view map indicating an imminent collision.According to yet another aspect of the present invention, a computer processing system for advanced driver assistance is provided. The computer processing system includes a storage device having program code stored thereon. The computer processing system further includes a hardware processor operatively coupled to the storage device and configured to run the program code stored on the storage device to extract shared features across different perception tasks from an input video stream including a plurality of images using a convolutional neural network (CNN) with multi-tasking. The different perception tasks include object detection and other perception tasks. The hardware processor continues to run the program code to concurrently solve the different perception tasks in a single pass using the multi-tasking CNN by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks. Each of the respective different branches corresponds to a respective one of the different perception tasks. The hardware processor continues to run the program code to form a parametric representation of a driving scene as at least one top-view map in response to the plurality of outputs of different perception tasks. The hardware processor also runs the program code to control operation of the vehicle for collision avoidance in response to the at least one top view map indicating an imminent collision.These and other features and advantages will become apparent from the following detailed description of their illustrative embodiments, which is to be read in conjunction with the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGSThe disclosure will provide details in the following description of preferred embodiments with reference to the following figures, in which: FIG. 1 is a block diagram illustrating an example processing system according to an embodiment of the present invention; FIG. 2 is a diagram showing an exemplary application overview according to an embodiment of the present invention; FIG. 3 is a block diagram illustrating an example multi-tasking perception network according to an embodiment of the present invention; FIG. 4 is a block diagram further illustrating the multi-tasking CNN of FIG. 3 according to an embodiment of the present invention; FIG. 5 is a flow diagram illustrating an example method for a multi-tasking perception network according to an embodiment of the present invention; and FIG. 6 is a block diagram illustrating an example advanced driver assistance system (ADAS) according to an embodiment of the present invention.DETAILED DESCRIPTION OF PREFERRED EMBODIMENTSEmbodiments of the present invention are directed to a multi-tasking perception network with applications in scene understanding and advanced driver assistance system (ADAS).One or more embodiments of the present invention propose a multi-tasking perception network that examines mutual characteristics between individual perception tasks and runs efficiently on a single GPU. In addition we demonstrate the applications of the proposed invention with alignment to scene understanding and advanced driver assistance systems.In one embodiment, the present invention proposes a new CNN architecture for simultaneously performing different perception tasks, such as object detection, semantic segmentation, depth estimation, occlusion augmentation, and 3D object localization, from a single input image. In particular, the input image is first passed through a feature extraction module that extracts features for sharing across different perception tasks. These shared features are then provided to task-specific branches, each performing one or more perception tasks. By sharing the feature extraction module, the network of the present invention is able to examine shared features between individual perception tasks and run efficiently on a single GPU. Additionally, the applications of the multi-tasking perception network towards scene understanding and advanced driver assistance systems are described. Of course, the present invention may be applied to other applications as will be readily appreciated by one of ordinary skill in the art in light of the teachings of the present invention provided herein.FIG. 1 is a block diagram illustrating an example processing system 100 according to an embodiment of the present invention. The processing system 100 includes a group of processing units (e.g., CPUs) 101, a group of GPUs 102, a group of memory devices 103, a group of communication devices 104, and a group of peripheral devices 105. The CPUs 101 may be single- or multi-core CPUs. The GPUs 102 may be single or multi-core GPUs. The one or more memory devices 103 may include caches, RAMs, ROMs, and other memory (flash, optical, magnetic, etc.). The communication devices 104 may include wireless and / or wired communication devices (e.g., network (e.g., WIFI, etc.) adapters, etc.). The peripheral devices 105 may include a display device, a user input device, a printer, an image capture device, and so forth. Elements of the processing system 100 are interconnected by one or more buses or networks (collectively referred to by the figure reference numeral 110).In one embodiment, storage devices 103 may store specially programmed software modules to transform the computer processing system into a special purpose computer configured to implement various aspects of the present invention. In one embodiment, special purpose hardware (e.g., application specific integrated circuits, field programmable gate arrays (FPGAs), and so forth) may be used to implement various aspects of the present invention. In one embodiment, the storage devices 103 include a multi-tasking perception network 103A for scene understanding and an advanced driver assistance system (ADAS).Of course, the processing system 100 may also include other elements (not shown) as readily contemplated by one of ordinary skill in the art, as well as omit certain elements. For example, various other input devices and / or output devices may be included in the processing system 100, depending on the particular implementation thereof, as will be readily understood by one of ordinary skill in the art. For example, various types of wireless and / or wired input and / or output devices may be used. In addition, additional processors, controllers, memories, and so forth may also be used in various configurations. These and other variations of the processing system 100 will be readily contemplated by one of ordinary skill in the art in light of the teachings of the present invention provided herein.Moreover, it is understood that various figures or configurations, as described below with respect to various elements and steps relating to the present invention, may be implemented in whole or in part by one or more of the elements of the system 100.As used herein, the term "hardware processor subsystem" or "hardware processor" may refer to a processor, memory, software, or combinations thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem may include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a separate processor- or computing- element-based controller (e.g., logic gates, etc.). The hardware processor subsystem may include one or more onboard memories (e.g., caches, dedicated memory arrays, read-only memory, etc.). In some embodiments, the hardware processor subsystem may include one or more memories, which may or may not be onboard or dedicated for use by the hardware processor subsystem (e.g., ROM, RAM, basic data exchange system (BIOS (= Basic Input / Output System)), etc.).In some embodiments, the hardware processor subsystem may include and execute one or more software elements. The one or more software elements may include an operating system and / or one or more applications and / or specific code to achieve a specified result.In other embodiments, the hardware processor subsystem may include dedicated, specialized circuitry that performs one or more electronic processing functions to achieve a specified result. Such a circuit may include one or more of application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).These and other variations of a hardware processor subsystem are also contemplated in accordance with embodiments of the present invention.FIG. 2 is a diagram illustrating an example application overview 200 according to an embodiment of the present invention.The application overview 200 includes input video 210, multi-tasking perception network 220, 2D object detection 231, 3D object detection 232, semantic segmentation 233, depth estimation 234, occlusion augmentation 235, motion and object tracking and 3D localization structure 240, top and top view maps 250, and scene understanding applications and advanced driver assistance systems 260.FIG. 2 illustrates example applications of a multi-tasking perception network that include scene understanding and ADAS 260. In particular, for a given input video 210, the perception network 220 processes each frame separately in a single forward pass and generates extensive per-frame outputs including 2D object detection 231, 3D object detection 232, semantic segmentation 233, depth estimation 234, and occlusion augmentation 235. These per-frame outputs may then be combined and fed into a structure of motion, object tracking, 3D localization 240, and top maps to produce a time and spatially consistent representation of a top map 250 of the captured scene, including the details of the scene layout, such as the number of lanes, road topology, and distance to an intersection, and the location of the objects consistent with the scene layout. The detailed top view map may then be useful for various applications such as scene understanding and ADAS 260 (e.g., blind spot reasoning, path planning, collision avoidance (through steering, brake inputs, etc.), etc.).FIG. 3 is a block diagram illustrating an example multi-tasking perception network 300 according to an embodiment of the present invention.The network 300 receives an input video 301.The network includes a convolutional neural network (CNN) with multi-tasking 310, a structure of a motion component 320, an object tracking component 330, a 3D localization component 340, top and top view maps 350, respectively, and applications 360.With respect to the input video 301, it may be a video stream of images (e.g., RGB or other type).With respect to the multi-tasking CNN 310, it takes as input (RGB) image and generates a plurality of outputs. The multi-tasking CNN 310 is configured to solve multiple tasks at a time.With respect to object tracking component 330, it receives 2D or 3D bounding ranges of object instances from multi-tasking CNN 310 for each frame of input video 301. The object tracking component 330 may operate with both 2D and 3D bounding ranges. The object tracking component 330 is to associate the 2D / 3D bounding areas across different frames, i.e., over time. An association between bounding areas indicates that these bounding areas capture exactly the same instance of an object.With respect to the structure of the motion component 320, it takes the video stream of RGB images 301 as input and outputs the relative camera pose to the first frame of the video. Thus, the structure from the motion component 320 measures how the camera itself moves through space and time. The input of 2D or 3D bounding regions helps the structure from the motion component 320 to improve its estimation because it can ignore dynamic parts of the scene that do not meet their internal assumptions about a static world.With respect to the 3D localization component 340, it integrates the estimated camera pose and the 3D bounding areas per frame to predict refined 3D bounding areas that are consistent over time.With respect to the top maps 350, they generate a consistent semantic top representation of the captured scene. The top and top view maps 350, respectively, integrate multiple outputs from the multi-tasking CNN 310, namely occlusion-based per-pixel semantics and depth estimates, as well as the refined 3D bounding regions from the 3D localization component 340. The output is a parametric representation of complex driving scenes including the number of lanes, the topology of line guides of a road, distances to intersections, a presence of zebra stripes and sidewalks, and some other attributes. It also provides a location of the object instances (given by the 3D location component 340) consistent with the scene layout.With respect to applications 360, the top semantic and parametric representation of top view maps 350 is a useful abstraction of the scene and may serve many different applications. Since it concludes from obscured areas, one application is the conclusion of a blind spot. Since it contains a metrically correct description of the road layout or the line guidance of a road, another application can be route planning. These are only two examples of possible applications that build on the output of the top view cards 350.FIG. 4 is a block diagram further illustrating the multi-tasking CNN 310 of FIG. 3 according to an embodiment of the present invention.The multi-tasking CNN 310 includes a shared feature extraction component 410, a task-specific CNN 420, and training data 430.The task-specific CNN 420 includes a 2D object detection component 421, a 3D object detection component 422, a depth estimation component 423, a semantic segmentation component 424, and an occlusion augmentation component 425.The training data 430 includes 2D object areas 431, 3D object areas 432, sparse 3D points 433, and semantic pixels 434. A sparse 3D point is a real point in 3D space relative to the camera that also captures the distance to the camera. Such sparse 3D points are typically collected with a laser scanner (lidar) and help the network estimate distances to objects.As described above, the multi-tasking CNN 310 takes as input an RGB image and generates a plurality of outputs (for the task specific CNN 420). The rough estimate of the calculation is still shared for all different outputs. The shared feature extraction component 410 and the task-specific CNN 420 are implemented as a shared convolutional neural network with multiple parameters that need to be estimated with the training data 430.With respect to the shared feature extraction component 410, it is represented as a convolutional neural network (CNN). The specific architecture of this CNN can be chosen arbitrarily as long as it generates a feature map of spatial dimensions that are proportional to the input image. The architecture may be adjusted depending on the available computing resources, enabling representations of heavy and strong features for offline applications as well as representations of weaker, but lighter, features for real-time applications.With respect to the task-specific CNN 420, CNN 420 applies multiple task-specific sub-CNNs thereto in view of the shared feature representation of block 210. These sub-CNNs are lightweight and require only a fraction of the runtime compared to the shared feature extraction component 410. This makes it possible to estimate the output of a plurality of tasks without significantly increasing the overall running time of the system. In one embodiment, the following outputs are estimated:Various components of the task specific CNN 420 will now be described in accordance with one or more embodiments of the present invention.With respect to the 2D object detection component 421, this output is a list of bounding areas (4 coordinates in image space, a confidence value, and a category label) that records the size of all instances of a predefined set of object categories, for example, cars, people, stop signs, traffic lights, etc.With respect to the 3D object detection component 422, for each detected object in 2D (from the 2D object detection component 421), the system estimates the 3D bounding area that encloses that object in actual 3D space (e.g., in meters or some other unit). This estimation provides the 3D location, orientation and dimension for each object, which is key information to fully understand the scene being captured.With respect to the depth estimation component 423, it assigns a distance (e.g., in meters or some other unit) to each pixel in the input image.With respect to the semantic segmentation component 424, it assigns a semantic category such as road, pavement, building, sky, car, or person to each pixel in the input image. The above listing is not limiting. The set of categories is different from the set of categories in the 2D object detection component 421, although some elements are the same. Importantly, the sentence in the semantic segmentation component 424 contains categories that cannot be easily drawn with a bounding area, such as the road.With respect to the occlusion augmentation component 425, it estimates the semantics and distances for all pixels that are obscured by foreground objects. A subset of the categories from the semantic segmentation component 424 is defined as foreground categories that can obscure the scene, such as cars, pedestrians, or line poles. The above listing is not limiting. All pixels associated with these categories in the output of the semantic segmentation component 424, which is also input to the occlusion augmentation component 425, are marked as occluded regions. The occlusion augmentation component 425 assigns each occluded area a category (from the set of background categories) as if it were not occluded. The occlusion augmentation component 425 essentially uses contextual information around the occluded pixels as well as from training data previously automatically learned to estimate semantic categories for occluded pixels. The same is done for distances in hidden areas. Importantly, as with all other components, the occlusion augmentation component 425 only operates on the feature representation provided by the shared feature extraction component 410 and also adds only a short amount of runtime to the overall system.With respect to the training data 430, they are required to estimate the convolutional neural network (CNN) parameters described with respect to the shared feature extraction component 410 and the task-specific CNN 420. Again, the CNN is a unified model that can be trained end-to-end, i.e., given an RGB input image and ground truth data for any of the tasks defined above. To better use data, in one embodiment we consider, without limitation, that each input image is annotated for all tasks. In one embodiment, we only require that an image for at least one task be annotated. In the case of an RGB input image and ground truth data for one (or more) task(s), the training algorithm then updates the parameters which are relevant for this task(s). Note that the shared feature representation from the shared feature extraction component 410 is always included. These updates are repeated with images and ground truth for all different tasks until the parameters are converged according to some loss functions for all tasks. The ground truth data required to train our multi-tasking CNN is as follows: 2D bounding areas 431; 3D bounding areas 432; sparse 3D points 433 (e.g., from a laser scanner) and semantic categories for each pixel (semantic pixels 434).It is important to note that the occlusion augmentation component 425 does not require annotations for occluded areas of the scene, which would be expensive and difficult to obtain.FIG. 5 is a flow diagram illustrating an example method 500 for a multi-tasking perception network according to an embodiment of the present invention. The method 500 may be applied to applications involving scene understanding and ADAS.At a block 505, extracting, from an input video stream containing a plurality of images using a convolutional neural network (CNN) for multi-tasking, shared features across different perception tasks, the different perception tasks including at least some of 2D and 3D object detection, depth estimation, semantic estimation, and occlusion augmentation.At a block 510, concurrent resolution, using the multi-tasking CNN, of the different perception tasks occurs in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks. Each of the respective different branches corresponds to a respective one of the different perception tasks.At a block 515, associating 2D and 3D bounding areas across different images is performed to obtain three-dimensional object tracks.At a block 520, processing of the 2D and 3D bounding areas is performed to determine a camera pose.At a block 525, locating objects encapsulated by the 2D and 3D bounding regions is responsive to the three-dimensional object tracks and the camera pose to provide refined 3D object tracks and 3D object tracks, respectively.At a block 530, forming a parametric representation of a driving scene as at least one top view map is performed in response to at least some of the plurality of outputs from different perception tasks (e.g., semantic segmentation, depth estimation, and occlusion augmentation) and the refined 3D object traces. It is noted that the remaining ones of the plurality of outputs from different perception tasks have been used to form the refined 3D object tracks.At a block 535, controlling operation of the vehicle for collision avoidance is responsive to at least one plan view map indicating an imminent collision.FIG. 6 illustrates an example advanced driver assistance system (ADAS) 600 based on tracking object detection, in accordance with an embodiment of the present invention.The ADAS 600 is used in an environment 601 where a user 2688 is located in a scene with multiple objects 699, each having their own locations and trajectories. The user 688 operates a vehicle 672 (e.g., a car, truck, motorcycle, etc.).The ADAS 600 includes a camera system 610. While a single camera system 610 is shown in FIG. 2 for purposes of illustration and brevity, it will be appreciated that multiple camera systems may also be used while retaining the spirit of the present invention. The ADAS 600 further includes a server 620 configured to perform object detection in accordance with the present invention. The server 620 may include a processor 621, a memory 622, and a wireless transceiver 623. The processor 621 and the memory 622 of the remote server 620 may be configured to perform driver assistance functions based on images received from the camera system 610 through the remote server 620 (the wireless transceiver 623 thereof). In this way, a corrective action may be taken by the user 688 and / or the vehicle 672.The ADAS 600 may interface with the user via one or more systems of the vehicle 672 that the user operates. For example, the ADAS 600 may provide the user information (e.g., detected objects, their locations, suggested actions, etc.) via a system 672A (e.g., a display system, a speaker system, and / or any other system) of the vehicle 672. Moreover, the ADAS 600 may interface with the vehicle 672 itself (e.g., via one or more systems of the vehicle 672 including, but not limited to, a steering system, a braking system, an accelerator system, a steering system, etc.) to control the vehicle or to cause the vehicle 672 to perform one or more actions. In this manner, the user or vehicle 672 can navigate around these objects 699 itself to avoid possible collisions between them.Embodiments described herein may include entirely hardware, entirely software, or both hardware and software elements. In a preferred embodiment, the present invention is implemented in software including, but not limited to, firmware, resident software, microcode, etc.Embodiments may include a computer program product accessible from a computer usable or computer readable medium that provides program code for use by or in connection with a computer or instruction execution system. A computer-usable or computer-readable medium may include any device that stores, communicates, propagates, or transports the program for use by, or in connection with, the instruction execution system, apparatus, or device. The medium may be a magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. The medium may include a computer readable storage medium such as a semiconductor or solid state memory, a magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a fixed magnetic disk, and an optical disk, etc.Each computer program may be tangiblely stored in a machine-readable storage medium or device (e.g., a program memory or a magnetic disk) readable by a general or special purpose programmable computer for configuring and controlling the operation of a computer when the storage medium or device is read by the computer to perform the procedures described herein. The inventive system may also be considered embodied in a computer readable storage medium configured with a computer program where the storage medium configured so causes a computer to operate in a specific and predefined manner to perform the functions described herein.A data processing system suitable for storing and / or executing program code may include at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements may include local memory used during actual execution of the program code, mass storage, and cache memories that provide temporary storage of at least some of program code to reduce the number of times code is retrieved from mass storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointing devices, etc.) may be coupled to the system either directly or via intervening I / O controllers.Network adapters may also be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices via intervening private or public networks. Modems, cable modem and Ethernet cards are only a few of the currently available types of network adapters.Reference in the specification to "a single embodiment" or "an embodiment" of the present invention, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrase "in a single embodiment" or "in an embodiment," as well as any other variations appearing at various places throughout the specification, are not necessarily all referring to the same embodiment. It is to be understood, however, that features of one or more embodiments may be combined in the teachings of the present invention provided herein.It will be appreciated that the use of any of the following " / ", "and / or" and "at least one of", such as in the cases of "A / B", "A and / or B" and "at least one of A and B", is intended to include only the selection of the first listed option (A) or the selection of the second listed option (B) or the selection of both options (A and B). As another example, such a phrase is intended to include, in the cases "A, B and / or C" and "at least one of A, B and C", only the selection of the first listed option (A) or only the selection of the second listed option (B), or only the selection of the third listed option (C), or only the selection of the first and second listed options (A and B), or only the selection of the first and third listed options (A and C), or only the selection of the second and third listed options (B and C), or the selection of all three options (A and B and C). This can be extended for as many elements as listed.The foregoing is to be considered in all respects as illustrative and exemplary, but not restrictive, and the scope of the invention disclosed herein is not to be determined from the detailed description, but from the claims as interpreted in accordance with the full breadth to which the patent laws are entitled. It is to be understood that the embodiments shown and described herein are merely illustrative of the present invention and that various modifications may be implemented by those skilled in the art without departing from the scope and spirit of the invention. Those skilled in the art could implement various other combinations of features without departing from the scope and spirit of the invention. Having thus described aspects of the invention with the details and particularities required by the patent laws, what is claimed and desired protected by the patent is set forth in the appended claims.

Claims

A computer-implemented method in an advanced driver assistance system (ADAS (= Adv Driver-Assistance System)), comprising: extracting (505), by a hardware processor, from an input video stream containing a plurality of images using a convolutional neural network (CNN) with multi-tasking, shared features across different perception tasks, wherein the different perception tasks include object detection and other perception tasks; A simultaneous solution (510), by the hardware processor, using the multi-tasking CNN, of the different perception tasks in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks, each of the respective different branches corresponding to a respective one of the different perception tasks; forming (530) a parametric representation of a driving scene as at least one plan view map in response to the plurality of outputs of different perception tasks; and controlling operation of the vehicle for collision avoidance in response to at least one plan view map indicating an imminent collision.The computer-implemented method of claim 1, wherein the other perception tasks comprise semantic segmentation, depth estimation, and occlusion augmentation.The computer-implemented method of claim 1, wherein the hardware processor is comprised of a single GPU.The computer-implemented method of claim 1, further comprising: associating bounding areas across different images to obtain object tracks; processing the bounding areas to determine a camera pose; and locating objects encapsulated by the bounding areas in response to the object tracks and the camera pose to provide refined object tracks for forming the at least one plan view map.The computer-implemented method of claim 4, wherein the refined object tracks are provided to be consistent over a given time period.The computer-implemented method of claim 4, further comprising generating a confidence value for each of the bounding areas, and wherein the confidence value is used to obtain the object tracks.The computer-implemented method of claim 1, wherein the multi-tasking CNN comprises a plurality of sub-CNNs, each of the plurality of sub-CNNs for processing a different one of the different perception tasks, respectively.The computer-implemented method of claim 1, further comprising training the multi-tasking CNN using training data comprising two-dimensional object boxes, three-dimensional object boxes, sparse three-dimensional points, and semantic pixels.The computer-implemented method of claim 8, wherein the training data is annotated for a respective one of the different perception tasks.The computer-implemented method of claim 1, wherein each of the semantic pixels is associated with one of a plurality of available semantic categories.The computer-implemented method of claim 1, further comprising forming the top view map using occlusion augmentation, wherein the occlusion augmentation estimates semantics and distances for any pixels in the input video stream that are occluded by foreground objects.The computer-implemented method of claim 1, wherein obscured areas of a scene in a frame of the input video stream are not annotated for occlusion augmentation.The computer-implemented method of claim 1, wherein the collision avoidance comprises controlling a vehicle input selected from the group consisting of braking and steering.The computer-implemented method of claim 1, further comprising performing a scene understanding task in response to the top view map to assist in collision avoidance.A computer program product for advanced driver assistance, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising: extracting (505), from an input video stream containing a plurality of images using a convolutional neural network (CNN) with multi-tasking, shared features across different perception tasks, the different perception tasks comprising object detection and other perception tasks; Simultaneously resolving (510), using the multi-tasking CNN, the different perception tasks in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks, each of the respective different branches corresponding to a respective one of the different perception tasks; forming (530) a parametric representation of a driving scene as at least one plan view map in response to the plurality of outputs of different perception tasks; and controlling operation of the vehicle for collision avoidance in response to at least one plan view map indicating an imminent collision.The computer program product of claim 15, wherein the other perception tasks comprise semantic segmentation, depth estimation, and occlusion augmentation.The computer program product of claim 15, wherein the hardware processor is comprised of a single GPU.The computer program product of claim 15, further comprising: associating bounding areas across different images to obtain object tracks; processing the bounding areas to determine a camera pose; and locating objects encapsulated by the bounding areas in response to the object tracks and the camera pose to provide refined object tracks for forming the at least one plan view map.The computer program product of claim 15, wherein the multi-tasking CNN comprises a plurality of sub-CNNs, each of the plurality of sub-CNNs for processing a different one of the different perception tasks, respectively.A computer processing system for advanced driver assistance, comprising: a storage device (103) including program code stored thereon; a hardware processor (102) operatively coupled to the storage device and configured to run the program code stored on the storage device to extract shared features across different perception tasks from an input video stream comprising a plurality of images using a convolutional neural network (CNN) with multi-tasking, wherein the different perception tasks comprise object detection and other perception tasks; using the multi-tasking CNN to concurrently solve the different perception tasks in a single pass by concurrently processing corresponding ones of the shared features through respective different branches of the multi-tasking CNN to provide a plurality of outputs of different perception tasks, each of the respective different branches corresponding to a respective one of the different perception tasks; forming a parametric representation of a driving scene as at least one plan view map in response to the plurality of outputs of different perception tasks; and controlling operation of the vehicle for collision avoidance in response to at least one plan view map indicating an imminent collision.

Citation Information

Patent Citations

  • US-PATENTANMELDUNGNR.62/814,886

  • US-PATENTANMELDUNGNR.16/787,727