Vision-based perception system
The vision-based perception system addresses the challenge of recognizing rare and unknown objects by using LIDAR and camera data to generate neural network representations for training, resulting in efficient and cost-effective object recognition.
Patent Information
- Application Number
- JP2024560584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-08-29
AI Technical Summary
Deep learning vision-based perception systems struggle to recognize rare objects and unknown objects without extensive training data, and current data augmentation methods are time-consuming and costly.
A vision-based perception system that uses a combination of LIDAR and camera data to identify unrecognized objects, with a second neural network generating neural network representations of these objects for training a first neural network, thereby enabling rapid acquisition of extended training datasets without human expert intervention.
Enables reliable recognition of rare objects and efficient learning of unknown objects, reducing the need for manual creation of synthetic assets and lowering costs, while improving the speed and accuracy of object recognition.
Smart Images

Figure 2025516460000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a vision-based perception system configured for object recognition by a neural network, the vision-based perception system comprising a LIDAR device and a camera device.
Background Art
[0002] A vision-based perception system comprising one or more light detection and ranging LIDAR devices configured to acquire a time series of 3D point cloud datasets of detected objects, one or more camera devices configured to capture a time series of 2D images of the objects, and a processing unit configured to recognize the objects is used in various applications. For example, vehicles such as automobiles, automated guided vehicles (AGVs), and autonomous mobile robots can be equipped with such a vision-based perception system to facilitate navigation, localization, and obstacle avoidance. In the context of an automobile, the vision-based perception system can be included by an advanced driver assistance system (ADAS).
[0003] In recent years, deep learning techniques have been used for automatic object recognition. Deep learning vision-based perception systems used for object recognition are trained with a vast amount of training data and have proven to reliably recognize objects belonging to classes of objects that include a large number of objects to be recognized. However, deep learning vision-based perception systems in this technical field tend not to be able to recognize objects with few examples (rare objects) in the training data set and are unable to recognize unknown (untrained) objects. Therefore, the training data set previously used for training is augmented with training samples representing unrecognized objects. Data augmentation has conventionally been achieved by synthetically creating computer graphics assets that are added to the training set for additional training of vision-based perception systems (U.S. Patent Application Publication No. 2020 / 0051291, U.S. Patent No. 9767565). However, computer graphics assets conventionally have to be created by experts in a time-consuming and cumbersome manner.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
[0005] In view of the above, the object underlying the present application is to provide a vision-based perception system that enables reliable recognition of rare objects and low-cost learning of the recognition of unknown objects. In the present application, by definition, it is explicitly stated that "recognition of an object" includes object recognition and / or object identification.
[0006] The foregoing and other objects are achieved by the subject matter of the independent claims. Further embodiments will become apparent from the dependent claims, the description, and the figures.
Means for Solving the Problems
[0007] According to a first aspect, there is provided a vision-based perception system (e.g., deep learning) for monitoring the physical environment of a vision-based perception system, comprising a light detection and ranging LIDAR device configured to acquire a time-series point cloud dataset representing the environment, a camera device configured to capture time-series images of the environment, and a processing unit including a first neural network and a second neural network different from the first neural network. The processing unit is configured to recognize objects (e.g., multiple objects) present in the time-series images by the first neural network and determine (detect) unrecognized objects (e.g., vehicles of unknown / untrained types) present in the time-series images. Further, the processing unit determines a bounding box of the determined unrecognized objects based on at least one of the point cloud datasets, and the second neural network is configured to obtain a neural network representation of the unrecognized objects based on the time-series images and the determined bounding box, and to train the first neural network for recognition of the unrecognized objects based on the obtained neural network representation of the unrecognized objects.
[0008] The processing unit may operate fully automatically or semi-automatically. The neural network representation of the unrecognized objects may be obtained without interaction by a human operator. The use of the second neural network to obtain the neural network representation of the unrecognized objects and to use that neural network representation of the unrecognized objects enables the rapid acquisition of multiple extended training datasets without the need for the time-consuming and costly manual creation of synthetic assets by human experts.
[0009] It goes without saying that in this specification, the term "neural network" refers to an artificial neural network. The first neural network may be a deep neural network, which may include a fully connected feedforward neural network (multi-layer perceptron, MLP) or a recurrent or convolutional neural network and / or a transformer. The second neural network is used to obtain a neural network representation of an unrecognized object, which may be or include an MLP, for example, a deep MLP.
[0010] According to one embodiment, a processing unit of a vision-based perception system clusters each point of a point cloud dataset to obtain each point cluster of the point cloud dataset, and for at least a part of the point cloud dataset, determines that one of the point clusters does not correspond to any object existing in a time-series image recognized by the first neural network, thereby being configured to determine an unrecognized object. Therefore, the unrecognized object can be determined reliably and quickly based on the data provided by the LIDAR device and the camera device.
[0011] According to another embodiment, the processing unit of the vision-based perception system performs a ground segmentation based on a point cloud dataset to determine the ground, and determines an unrecognized object by determining that the unrecognized object is located on the determined ground. In particular, the processing unit determines that none of the point clusters correspond to any object recognized by the first neural network for at least a portion of the point cloud dataset, and determines that the unrecognized object is located on the determined ground. The unrecognized object may be determined only by both of these. "Ground segmentation" means an automatic segmentation of the ground on which the object is located from the image and / or other parts of the point cloud. The ground segmentation may be performed based on the image or based on both the image and the LIDAR point cloud. An example of ground segmentation is road segmentation in the context of automotive applications. The adoption of ground segmentation facilitates reliably determining unrecognized / unrecognizable objects in the images captured by the camera device of the vision-based perception system.
[0012] According to another embodiment of the vision-based perception system of the first aspect, for each of a plurality of pre-stored images captured by a camera device and a corresponding pre-stored LIDAR point cloud dataset captured by a LIDAR device, based on at least one of the corresponding pre-stored LIDAR point cloud datasets having the same dimensions as the bounding box determined by the processing unit for the determined unrecognized object within a predetermined threshold, at least one other object having the corresponding bounding box determined by the processing unit is determined, and in each of the plurality of pre-stored images, at least one other object determined is replaced with a neural network representation of the unrecognized object in order to obtain a plurality of training images for training the first neural network for recognition of the unrecognized object. The processing unit is further configured to do so.
[0013] The pre-stored images may be captured by the same camera device of the vision-based perception system of the first aspect, and the pre-stored LIDAR point cloud dataset may be captured by the LIDAR device of the vision-based perception system of the first aspect. According to this embodiment, by means of a second neural network, in order to create a huge number of new training images (augmented training data) representing new training scenes that can be appropriately used to train the first neural network for object recognition, the objects existing in the pre-stored images are replaced with neural network representations of unrecognized objects.
[0014] To facilitate the training procedure of the first neural network that yields reliable recognition results, the lighting and shadow conditions can be appropriately taken into account when creating the training images. Thus, according to another embodiment, the processing unit of the vision-based perception system determines, for each of the plurality of pre-stored images captured by the camera device, a first light direction (direction of light incidence) with respect to at least one other object, determines a second light direction (direction of light incidence) with respect to the determined unrecognized object, and when the first light direction and the second light direction are offset from each other by less than a predetermined illumination (deviation) angle (for example, only at this time), it is configured to replace at least one other object determined in each of the plurality of pre-stored images with a neural network representation of an unrecognized object. According to this embodiment, an image having at least one other object illuminated from a light direction that is too different from the determined light direction for the unrecognized object is not considered to be sufficiently suitable to function as the basis for the training image.
[0015] According to another embodiment, the second neural network of the vision-based perception system includes a first multi-layer perceptron (MLP) trained based on neural radiance field technology, and a second MLP different from the first MLP and trained based on neural radiance field technology. The first MLP is configured to obtain a neural network representation of an unrecognized object, and the second MLP is configured to obtain a neural network representation of the background of the unrecognized object.
[0016] Neural Radiance Field (NERF) technology was proposed by B. Mildenhall et al. in the paper "Nerf: Representing scenes as neural radiance fields for view synthesis" in "Computer Vision - ECCV2020" (16 th European Conference, Glasgow, UK, August 23 - 28, 2020, Springer, Cham, 2020). NERF enables the efficient acquisition of an accurate neural network representation of the environment, particularly of unrecognized objects, based on color values and spatially dependent volume density values (see also the detailed description below).
[0017] The first MLP and the second MLP trained with NERF enable unrecognized objects to be automatically separated from the background without human interaction. According to one embodiment, some ray segmentation technique is used to obtain this separation. In this embodiment, the first MLP is configured to obtain a neural network representation of an unrecognized object based on the portion of the camera ray that crosses the region corresponding to the bounding box, and the second MLP is configured to obtain a neural network representation of the background of the unrecognized object based on the portion of the camera ray that does not cross the region corresponding to the bounding box. Based on the bounding box obtained from the LIDAR point cloud data, on the one hand, the neural network representation of the unrecognized object and, on the other hand, its background can be automatically and relatively easily obtained based on individually designed MLPs.
[0018] The training images can be obtained, for example, by rendering the neural network representation of an unrecognized object onto a pre-stored image, such as by replacing one or more recognized common objects present in the pre-stored image with an unrecognized object. For example, the neural network representation of a scene containing one or more recognized objects can be obtained by a neural network trained with a first NERF and a neural network trained with a second NERF, and at least one other determined object can be replaced with the neural network representation of a rare object. According to one embodiment, the first MLP is configured to obtain the neural network representation of an unrecognized object based on the camera pose (of the camera used to capture the image containing the unrecognized object), and the processing unit is configured to render the neural network representation of the unrecognized object onto each of a plurality of pre-stored images based on a rendering pose that deviates from the camera pose by less than a predetermined threshold (e.g., only by the rendering pose), so as to replace at least one other determined object with the neural network representation of the unrecognized object. As a result, in order to ensure that artifacts caused by the position / direction of the unrecognized object that do not exist in the time-series images of the environment containing the unrecognized object can be avoided, the position where the rendering is performed and the direction in which the rendering is performed can be controlled.
[0019] The vision-based perception system according to the first aspect and any of its embodiments can be suitably used in a vehicle. Thus, according to another embodiment, the vision-based perception system is configured to be installed in a vehicle, particularly an automobile, an autonomous mobile robot, or an automated guided vehicle (AGV), and the time-series point cloud dataset and the time-series images represent the driving scene of the vehicle. In automotive applications, the vision-based perception system according to the first aspect and any of its embodiments can be included by an advanced driver assistance system (ADAS).
[0020] According to a second aspect, a vehicle is provided that includes a vision-based perception system according to the first aspect or any of its embodiments. The vehicle may be an automobile, an autonomous mobile robot, or an automated guided vehicle (AGV).
[0021] According to a third aspect, a method is provided for training a first neural network of a vision-based perception system for object recognition. The method includes obtaining, by a light detection and ranging LIDAR device of the vision-based perception system, a time-series point cloud dataset representing the environment of the vision-based perception system; capturing, by a camera device of the vision-based perception system, time-series images of the environment; determining (detecting), by a processing unit of the vision-based perception system, unrecognized objects in the time-series images; determining, by the processing unit, a bounding box of the determined unrecognized objects based on at least one of the point cloud datasets; obtaining, by a second neural network of the vision-based perception system, a neural network representation of the unrecognized objects based on the time-series images and the determined bounding boxes; and training, based on the obtained neural network representation of the unrecognized objects, the first neural network for recognition of the unrecognized objects.
[0022] The unrecognized objects can be, for example, vehicles of unknown (untrained) brands or types, or unknown objects on streets or roads, or can include these.
[0023] According to one embodiment, the method of the third aspect further includes clustering, by the processing unit, each point of the point cloud dataset to obtain each point cluster of the point cloud dataset, and the unrecognized objects are determined by determining that, for at least a part of the point cloud dataset, none of the point clusters correspond to any object recognized by the first neural network.
[0024] According to another embodiment, the method further includes, by a processing unit, performing a ground classification (e.g., a road classification) based on a point cloud dataset to determine the ground, and an unrecognized object is determined by determining that the unrecognized object is located on the determined ground.
[0025] According to another embodiment, the method includes, by a processing unit, for each of a plurality of pre-stored images captured by a camera device (e.g., a camera device of a vision-based perception system) and a corresponding pre-stored LIDAR point cloud dataset captured by a LIDAR device (e.g., a LIDAR device of a vision-based perception system), determining at least one other object having a corresponding bounding box determined by the processing unit, the at least one other object having the same dimensions as the determined bounding box for the unrecognized object within a predetermined threshold, based on at least one of the corresponding pre-stored LIDAR point cloud datasets; and replacing, in each of the plurality of pre-stored images, at least one other object determined therein with a neural network representation of the unrecognized object to obtain a plurality of training images for training a first neural network for recognition of the unrecognized object.
[0026] According to another embodiment, the method further includes, by a processing unit, for each of a plurality of pre-stored images captured by a camera device, determining a first light direction for at least one other object and determining a second light direction for the determined unrecognized object, and the determined at least one other object is replaced with a neural network representation of the unrecognized object in each of the plurality of pre-stored images when the first light direction and the second light direction are offset from each other by less than a predetermined illumination (deviation) angle.
[0027] According to one embodiment, the second neural network includes a first multi-layer perceptron (MLP) trained based on neural radiance field technology and a second MLP trained based on neural radiance field technology, different from the first MLP. The neural network representation of an unrecognized object is obtained by the first MLP, and the method further includes obtaining, by the second MLP, the neural network representation of the background of the unrecognized object.
[0028] According to another embodiment, the neural network representation of an unrecognized object is obtained based on a portion of a camera ray crossing a region corresponding to a bounding box, and the neural network representation of the background of the unrecognized object is obtained based on a portion of a camera ray not crossing the region corresponding to the bounding box.
[0029] According to another embodiment, the neural network representation of an unrecognized object is obtained based on the camera pose (of the camera), and at least one other determined object is rendered onto each of a plurality of pre-stored images with a rendering pose that deviates from the camera pose by less than a predetermined threshold, replacing the neural network representation of the unrecognized object with the neural network representation of the unrecognized object.
[0030] According to a fourth aspect, a method of recognizing an object present in an image is provided, including the steps of training, for object recognition, a first neural network of a vision-based perception system according to the third aspect and its embodiments.
[0031] The method of training the first neural network of the vision-based perception system according to the third aspect and its embodiments for object recognition, and the method of recognizing an object existing in an image according to the fourth aspect, provide the same or similar advantages as those described above with reference to the vision-based perception system according to the first aspect and its embodiments. The vision-based perception system according to the first aspect and its embodiments may be configured to execute the methods according to the third aspect and its embodiments and the method according to the fourth aspect.
[0032] According to a fifth aspect, there is provided a computer program product including computer-readable instructions for executing or controlling the steps of the method according to the third aspect or the fourth aspect and its embodiments when executed on a computer.
[0033] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims.
[0034] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings and figures.
Brief Description of the Drawings
[0035]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0036] In this specification, a technique for data augmentation of a training dataset used in a deep - learning vision - based perception system configured for object recognition is provided. An embodiment of a vision - based perception system 100 for monitoring the physical environment of a vision - based perception system 100 in which such a technique is implemented is shown in FIG. 1. The vision - based perception system 100 includes a LIDAR device 110, for example, a plurality of LIDAR devices, and a camera device 120, for example, a plurality of camera devices. Further, the vision - based perception system 100 includes a processing unit 130 including a first neural network 132 and a second neural network 138. The processing unit 130 is configured to receive input data from both the LIDAR device 110 and the camera device 120.
[0037] The first neural network 132 may include a regression or convolutional neural network and / or a transducer, and is trained for object recognition. The first neural network 132 is configured to receive input data based on an image captured by the camera device 110. The input data may include image patches. The input data may include a tensor having a shape (number of images)×(image width)×(image height)×(image depth). Regarding the first layer of the neural network that processes the input data, the number of input channels may be equal to or greater than the number of channels of the data representation, for example, three channels for the RGB or YUV representation of the image. By passing through the neural network layers, the image is abstracted into a feature map having a shape (number of images)×(feature map width)×(feature map height)×(feature map channels) and can be further processed. The first neural network 132 may include a neural network that operates as a classifier. The first neural network 132 may also be configured to receive point cloud data acquired by the LIDAR device 110.
[0038] The second neural network 138 is configured to receive input data based on an image captured by the camera device 120. The second neural network 138 may include one or more MLPs and may be trained to generate an implicit neural network representation of the environment including information regarding the pose of the camera device. The second neural network 138 may include a neural network trained with one or more NERFs (see the following description).
[0039] The LIDAR device 110 is configured to acquire a time-series point cloud dataset representing the environment of the system, and the camera device 120 is configured to capture a time-series image of the environment. The processing unit 130 is configured to recognize objects (e.g., multiple objects) present in the time-series image by means of a first neural network and to determine (detect) unrecognized objects present in the time-series image. Unrecognized objects can be objects that were not trained to be recognized by the first neural network or objects that only existed in a relatively sparse subset of the training data.
[0040] Furthermore, the processing unit 130 is configured to determine a bounding box of the determined unrecognized objects based on at least one of the point cloud datasets, and to obtain a neural network representation of the unrecognized objects based on the time-series image and the determined bounding box by means of a second neural network.
[0041] Furthermore, the processing unit 130 of the vision-based perception system 100 is configured to train the first neural network for the recognition of unrecognized objects based on the obtained neural network representation of the unrecognized objects. The second neural network 138 can be used to create a plurality of new training images each containing a neural network representation of an unrecognized object, and the first neural network 132 can be trained based on the plurality of new training images thus created (see also the following description). By employing the second neural network 138, an extended training dataset suitable for training the first neural network 132 to recognize unrecognized objects can be generated without any interaction by human experts or with at least less interaction than required in the art.
[0042] A method 200 for training a first neural network of a vision-based perception system according to an embodiment for object recognition is shown by the flowchart shown in FIG. 2. This method can be implemented in the vision-based perception system 100 shown in FIG. 1, and the vision-based perception system 100 shown in FIG. 1 can be configured to execute one or more steps included by the method 200 shown by the flowchart shown in FIG. 2.
[0043] The method 200 includes a step S210 of obtaining a time-series point cloud dataset representing the environment of the vision-based perception system by a LIDAR device (one or more LIDAR devices) of the vision-based perception system, and a step S220 of capturing time-series images of the environment by a camera device (one or more camera devices) of the vision-based perception system. Further, the method 200 includes a step S230 of determining (detecting) unrecognized objects in the time-series images by a processing unit of the vision-based perception system, a step S240 of determining a bounding box of the determined unrecognized objects based on at least one of the point cloud datasets by the processing unit, and a step S250 of obtaining a neural network representation of the unrecognized objects based on the time-series images and the determined bounding boxes by a second neural network of the vision-based perception system.
[0044] Furthermore, the method 200 includes a step S260 of training the first neural network for recognition of unrecognized objects based on the neural network representation of the unrecognized objects obtained by the second neural network. In the training process S260, a plurality of new training images each including the neural network representation of an unrecognized object can be created and used.
[0045] Figure 3 shows the procedure of data augmentation 300 of training data used by the neural network of a vision-based perception system configured for object recognition. Data augmentation can be achieved by a fully automated method, or a semi-automated method that requires some human interaction but less human interaction than required for synthetic asset creation in the art. In the illustrated example, the vision-based perception system is installed in a vehicle. During the driving of the vehicle, the driving scene is captured by the LIDAR device and the camera device of the vision-based perception system installed in the vehicle. The captured driving scene is stored in the driving scene / sequence database (310) in the form of a time-series point cloud dataset acquired by the LIDAR device and a time-series image acquired by the camera device. In the rare object search procedure 320, a rare object that cannot be recognized (e.g., a rare vehicle) is automatically determined (detected), and a 3D bounding box corresponding to / including this object is automatically extracted from the Lidar point cloud. The determination of the 3D bounding box is based on the clustering of the points in the point cloud, which can be obtained by calculating the convex hull for the acquired 3D point cluster. The extracted 3D bounding box is defined by the center position (x, y, z), the direction angles (roll, pitch, yaw), and the dimensions (width, length, height).
[0046] The determination of an unrecognized object may include determining that a particular cluster of points in a consecutive point cloud of the time-series point cloud dataset cannot be associated with an object present in the image that is recognized by the neural network included in the vision-based perception system. Further, the determination of an unrecognized object may include determining that such a particular point cluster is located on the road determined by the road segmentation performed on the time-series point cloud dataset, and / or determining that the unrecognized object is located on the road determined by the road segmentation.
[0047] In Neural (Synthetic) Asset Creation Procedure 330, the neural asset is created by a neural network trained in NERF (see B. Mildenhall et al. in the paper "Nerf: Representing scenes as neural radiance fields for view synthesis" in "Computer Vision - ECCV2020" (16 th European Conference, Glasgow, UK, August 23 - 28, 2020, Springer, Cham, 2020)).
[0048] Data based on the captured images of the time - series images acquired by the camera device representing the driving scene is input into the neural network trained in NERF. The neural network trained in NERF learns the implicit neural network representation of rare objects. The input data represents the coordinates (x, y, z) of a set of sampled 3D points and the viewing direction (θ, φ) corresponding to the 3D points, and the neural network trained in NERF outputs a viewpoint - dependent color value (e.g., RGB) and a volume density value σ that can be interpreted as the differential probability of a camera ray ending at an infinitesimal particle at the position (x, y, z) (see the paper by B. Mildenhall et al. mentioned above). Thus, the MLP realizes F Θ :(x, y, z, θ, φ) → (R, G, B, σ). The output color value and volume density value are used for volume rendering to obtain the neural network representation of the driving scene captured by the camera device.
[0049] The extracted bounding box is used to "separate" the rendering of the rare object from the rendering of the background of the rare object. The rendered "separated" rare object is the neural network representation of the rare object. By "separating" the rare object based on the extracted bounding box, manual segmentation of the rare object in the image can be avoided, thereby significantly reducing time, cost, and human labor compared to the prior art.
[0050] According to one embodiment, the neural network trained with NERF includes a first neural network trained with NERF for obtaining the neural network representation of the rare object and a second neural network trained with NERF for obtaining the neural network representation of the background of the rare object. Training these two neural networks trained with different NERFs is based on the camera ray splitting technique shown in FIG. 4. Similar to the ray marching taught by B. Mildenhall et al., a set of camera rays for a set of camera poses is generated for each of the (training) images of the time series of images acquired by the camera device representing the driving scene. As shown in FIG. 4, the first portion of the camera ray crossing the determined rare object (bounding box) is distinguished from the second portion of the camera ray that does not cross the determined rare object (bounding box) but crosses its background. The first portion can be used to render the rare object, and the second portion can be used to render its background.
[0051] The neural network representation of the rare object is stored in the neural asset database (340). The procedure described above can be performed for a wide variety of rare objects that cannot be recognized based on the available training data used to train a neural network (different from the neural network 330 trained with NERF) used for object recognition. Therefore, a wide variety of neural network representations of various rare objects can be stored in the neural asset database (340).
[0052] In configuration procedure 350, the training images are obtained based on neural assets (neural network representations of rare objects) and previously stored images captured by one or more camera devices. A previously stored image is selected, and the neural network representation of the rare object is rendered onto that image, resulting in a new training image that includes the rare object. The new training neural network representation can be used to train the neural network used for object recognition. A vast number of previously stored images can be selected, and the rare object can be rendered onto each of the vast number of selected previously stored images (backgrounds) stored in the extended data database (360) to obtain a vast number of new training images, resulting in a new rich training dataset for training the neural network used for object recognition.
[0053] The selection of previously stored images from the vast amount of previously stored driving scene images during configuration procedure 350 can include a) determining whether another (e.g., recognized or recognizably common) object with a corresponding bounding box having the same dimensions as the rare object within a predetermined threshold exists within that image, and b) determining whether the direction of light with respect to the other object is similar to that with respect to the rare object within a predetermined limit (see Figure 5). When these conditions are met, i.e., in scenes with similar bounding boxes and light directions, the other object can be appropriately replaced with the neural asset created for the unrecognized rare object (a minivan in the example shown in Figure 5).
[0054] For example, the neural network representation of a scene containing other objects can be obtained by a neural network trained with a first NERF and a neural network trained with a second NERF, and the other objects can be replaced with the neural network representation of rare objects. Taking into account the direction of light enables the use of coherent / plausible lighting and shadows in the process of creating training images for the extended dataset by rendering the neural network representation onto the background of pre-stored images, thus improving / facilitating the training procedure based on the extended data.
[0055] Furthermore, it must be determined from which position and in which (line-of-sight) direction a synthetic image containing a rare object should be rendered. Creating a synthetic image depicting a rare object at a position and direction very different from those used to obtain the neural network representation of the rare object can result in artifacts that can significantly affect the training procedure based on the extended data. Thus, according to one embodiment, an appropriate rendering pose for rendering a desired synthetic image should be determined according to one embodiment by comparing the rendering pose with the pose used to obtain the neural network representation of the rare object (see FIGS. 5 and 6).
[0056] Figure 6 shows an example for conditioning the selection of possible rendering poses. Rare objects that are not recognized are detected by a camera device from a number of camera poses during the driving of an automobile in which a vision-based perception system is installed. In a synthetic image suitable for training the recognition of rare objects, the rare objects should be rendered in an acceptable rendering pose. The acceptable rendering pose is determined based on the ratio of valid rendering rays characterized by the angular deviation with respect to the corresponding camera rays used to obtain the neural representation of the rare object that does not exceed a predetermined angular deviation threshold. For example, a predetermined angular deviation threshold of 20° to 30° may be considered suitable for the valid rendering rays, and the rendering pose for rendering the rare object in a new synthetic image representing the extended data for training purposes may be allowed only if it is characterized by a ratio of valid rendering rays of at least 60% (see Figure 6). By selecting the acceptable pose and appropriate lighting conditions for the rendering process, realistic positioning, line-of-sight direction, and lighting of the rendered rare object can be achieved.
[0057] When a pre-stored image containing another object, for example a recognizable common object, is considered suitable for generating a new synthetic image containing the neural network representation of the rare object, the other object can be replaced by the neural network representation of the rare object. According to a particular example, the pre-stored image can be the same as the captured image containing the unrecognized rare object as shown in Figure 7. The stored actual image captured by one or more camera devices can be converted into the neural network representation of the captured scene.
[0058] The first neural network trained with NERF learns / acquires the implicit neural network representation of rare objects and selects the recognized common objects according to the similar lighting conditions also present in the scene for replacement with the neural network representation of rare objects in order to obtain twice the new synthetic training images containing rare objects using different rendering poses at different positions.
[0059] All of the above-described embodiments are not intended as limitations and serve as examples showing the features and advantages of the present invention. It should be understood that some or all of the features described above can also be combined in different ways.
Explanation of Signs
[0060] 100 Vision-based perception system 110 LIDAR device 120 Camera device 130 Processing unit 132 First neural network 138 Second neural network
Claims
1. A visual - based perception system (100) for monitoring a physical environment of the visual - based perception system (100), a light detection and ranging LIDAR device (110) configured to obtain a time - series point - cloud dataset representing the environment, a camera device (120) configured to capture time - series images of the environment, a processing unit (130) comprising a first neural network (132) and a second neural network (138) different from the first neural network, wherein the first neural network (132) recognizes objects present in the time - series images, determines unrecognized objects present in the time - series images, determines a bounding box of the determined unrecognized objects based on at least one of the point - cloud dataset, wherein the second neural network (138) obtains a neural - network representation of the unrecognized objects based on the time - series images and the determined bounding box, and trains the first neural network (132) for recognition of the unrecognized objects based on the obtained neural - network representation of the unrecognized objects a processing unit (130) configured as such, a visual - based perception system (100) comprising the same.
2. The processing unit (130) is configured to cluster each point of the point - cloud dataset to obtain each point cluster of the point - cloud dataset, and determine the unrecognized objects by determining that for at least a part of the point - cloud dataset, one of the point clusters does not correspond to any object recognized by the first neural network (132). The visual - based perception system (100) according to claim 1, configured as such.
3. The processing unit (130) is configured to perform a ground segmentation based on the point - cloud dataset to determine the ground, and determine the unrecognized objects by determining that the unrecognized objects are located on the determined ground. The visual - based perception system (100) according to claim 2, configured as such.
4. The processing unit (130) is configured to For each of the plurality of pre-stored images captured by the camera device (120) and the corresponding pre-stored LIDAR point cloud dataset captured by the LIDAR device (110), at least one of the corresponding pre-stored LIDAR point cloud datasets having the same dimensions as the bounding box determined by the processing unit (130) for the determined unrecognized object within a predetermined threshold is determined, and at least one other object having the corresponding bounding box determined by the processing unit (130) is determined, To obtain a plurality of training images for training the first neural network (132) for the recognition of the unrecognized object, replace at least one other object determined in each of the plurality of pre-stored images with the neural network representation of the unrecognized object Further configured as The visual-based perception system (100) according to any one of claims 1 to 3.
5. The processing unit (130) For each of the plurality of pre-stored images captured by the camera device (120), determine a first light direction for the at least one other object and a second light direction for the determined unrecognized object, When the first light direction and the second light direction are offset from each other by less than a predetermined illumination angle, replace at least one other object determined in each of the plurality of pre-stored images with the neural network representation of the unrecognized object The visual-based perception system (100) according to claim 4, configured as
6. The second neural network (138) includes a first multi-layer perceptron MLP trained based on neural radiance field technology and a second MLP trained based on the neural radiance field technology, different from the first MLP, The first MLP is configured to obtain the neural network representation of the unrecognized object, and the second MLP is configured to obtain the neural network representation of the background of the unrecognized object, The visual-based perception system (100) according to any one of claims 1 to 5.
7. The first MLP is configured to obtain the neural network representation of the unrecognized object based on a portion of the camera ray that crosses the region corresponding to the bounding box, The second MLP is configured to obtain the neural network representation of the background of the unrecognized object based on a portion of the camera ray that does not cross the region corresponding to the bounding box, The visual-based perception system (100) according to claim 6.
8. The first MLP is configured to obtain the neural network representation of the unrecognized object based on the camera pose, The processing unit (130) is configured to replace the determined at least one other object with the neural network representation of the unrecognized object by rendering the neural network representation of the unrecognized object onto each of the plurality of pre-stored images based on a rendering pose that deviates from the camera pose by less than a predetermined threshold, The visual-based perception system (100) according to claim 6 or 7.
9. The visual-based perception system (100) is configured to be installed in a vehicle, and the time-series point cloud dataset and the time-series images represent the driving scene of the vehicle. The visual-based perception system (100) according to any one of claims 1 to 8.
10. A method (200) for training a first neural network (132) of a visual-based perception system (100) for object recognition, Obtaining (S210) a time-series point cloud dataset representing the environment of the visual-based perception system (100) by an optical detection and ranging LIDAR device of the visual-based perception system (100); Capturing (S220) time-series images of the environment by a camera device (120) of the visual-based perception system (100); Determining (S230) unrecognized objects in the time-series images by a processing unit (130) of the visual-based perception system (100); Determining (S240) a bounding box of the determined unrecognized object based on at least one of the point cloud datasets by the processing unit (130); A step (S250) of obtaining a neural network representation of the unrecognized object based on the time-series images and the determined bounding boxes by a second neural network (138) of the visual-based perception system (100); A step (S260) of training the first neural network (132) for recognition of the unrecognized object based on the obtained neural network representation of the unrecognized object A method (200) including the above.
11. A step of clustering each point of the point cloud dataset by the processing unit (130) to obtain each point cluster of the point cloud dataset Further including The unrecognized object is determined (S230) by determining that, for at least a part of the point cloud dataset, none of the one of the point clusters corresponds to any object recognized by the first neural network (132). The method (200) according to claim 10.
12. A step of performing a ground classification based on the point cloud dataset by the processing unit (130) to determine the ground Further including The unrecognized object is determined (S230) by determining that the unrecognized object is located on the determined ground. The method (200) according to claim 11.
13. For each of a plurality of pre-stored images captured by a camera device (120) and a corresponding pre-stored LIDAR point cloud dataset captured by a LIDAR device (110) by the processing unit (130), at least one other object having a corresponding bounding box determined by the processing unit (130) based on at least one of the corresponding pre-stored LIDAR point cloud datasets having the same dimensions as the bounding box determined by the processing unit (130) for the determined unrecognized object within a predetermined threshold. A step of determining; To obtain a plurality of training images for training the first neural network (132) for the recognition of the unrecognized object, replacing, in each of the plurality of pre-stored images, the at least one other object determined with the neural network representation of the unrecognized object The method (200) according to any one of claims 10 to 12, further comprising.
14. Determining, by the processing unit (130), for each of the plurality of pre-stored images captured by the camera device (120), a first light direction for the at least one other object and determining a second light direction for the determined unrecognized object further comprising The method (200) according to claim 13, wherein the determined at least one other object is replaced with the neural network representation of the unrecognized object in each of the plurality of pre-stored images when the first light direction and the second light direction are offset from each other by less than a predetermined illumination angle.
15. The second neural network includes a first multi-layer perceptron MLP trained based on neural radiance field technology and a second MLP trained based on the neural radiance field technology, different from the first MLP, The neural network representation of the unrecognized object is obtained by the first MLP (S250), Obtaining, by the second MLP, a neural network representation of the background of the unrecognized object The method (200) according to any one of claims 10 to 14, further comprising.
16. The neural network representation of the unrecognized object is obtained based on a portion of the camera ray crossing the region corresponding to the bounding box (S250), The neural network representation of the background of the unrecognized object is obtained based on a portion of the camera ray that does not cross the region corresponding to the bounding box. The method (200) according to claim 15.
17. The neural network representation of the unrecognized object is obtained based on the camera pose (S250), The determined at least one other object replaces the neural network representation of the unrecognized object by rendering the neural network representation of the unrecognized object onto each of the plurality of pre-stored images based on a rendering pose that is offset from the camera pose by less than a predetermined threshold. The method (200) according to claim 15 or 16. **Claim 18** A computer program product including computer-readable instructions for performing or controlling the steps of the method (200) according to any one of claims 10 to 17 when executed on a computer.
Citation Information
Patent Citations
Classification of rare cases
JP2020524854A
Object Detection Training Based on Artificially Generated Images
US20200051291A1
Synthesizing training data for broad area geospatial object detection
US9767565B2