Method and system for object detection

By combining the alternating structure of the expanded convolutional layer and the pooling layer in the convolutional neural network, the problem of insufficient efficiency and accuracy in object detection is solved, and efficient processing and accurate detection of complex environmental data is achieved.

CN115082867BActive Publication Date: 2025-06-06APTIV TECHNOLOGIES AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210222984.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-10
Filing Date
2022-03-07
Publication Date
2025-06-06
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient and accurate results in object detection, especially when processing data in complex environments.

Method used

Object detection is achieved by combining the alternating structure of the expanded convolutional layer and the pooling layer in the convolutional neural network. The specific steps include: determining the output of the first pooling layer based on the input data, followed by the output of the expanded convolutional layer, then the output of the second pooling layer, and finally performing object detection based on these outputs.

Benefits of technology

This method improves the efficiency and accuracy of object detection, can effectively process data in complex environments, reduces the occurrence of grid artifacts, and maintains detailed local information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082867B_ABST
    Figure CN115082867B_ABST
Patent Text Reader

Abstract

Method and system for object detection. A computer-implemented method for object detection is provided, the method comprising the following steps performed by computer hardware components: determining an output of a first pooling layer based on input data; determining an output of a dilated convolutional layer disposed directly after the first pooling layer based on the output of the first pooling layer; determining an output of a second pooling layer disposed directly after the dilated convolutional layer based on the output of the dilated convolutional layer; and performing object detection based on at least the output of the dilated convolutional layer or the output of the second pooling layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods and systems for object detection. Background Art

[0002] Various sensors such as cameras, radar sensors, or LIDAR sensors may be used in automotive applications to monitor the environment of the vehicle. Driver assistance systems may utilize data captured from the sensors, for example, by analyzing the data to detect objects. For object detection, a convolutional neural network (CNN) may be used. However, object detection may be a cumbersome task.

[0003] Therefore, there is a need to provide a method and system for object detection that achieves efficient and accurate results. Summary of the invention

[0004] In one aspect, the present disclosure is directed to a computer-implemented method for object detection, the method comprising the following steps performed (in other words: executed) by computer hardware components: determining an output of a first pooling layer based on input data; determining an output of a dilated convolutional layer directly following the first pooling layer based on the output of the first pooling layer; determining an output of a second pooling layer directly following the dilated convolutional layer based on the output of the dilated convolutional layer; and performing object detection based on at least the output of the dilated convolutional layer or the output of the second pooling layer.

[0005] In other words, the pooling operation in the first pooling layer is based on the input data to determine the output of the first pooling layer. Before performing the pooling operation, the input data may be subjected to another layer (e.g., another dilated convolution layer). The output of the first pooling layer is the input of the dilated convolution layer. The dilated convolution layer directly follows the first pooling layer (in other words: no additional layer is set between the first pooling layer and the dilated convolution layer). The dilated convolution operation in the dilated convolution layer determines the output of the dilated convolution layer. The output of the dilated convolution layer is the input of the second pooling layer. The second pooling layer directly follows the dilated convolution layer (in other words: no additional layer is set between the dilated convolution layer and the second pooling layer). On the other hand, the dilated convolution layer is immediately preceding the second pooling layer, and the first pooling layer is immediately preceding the dilated convolution layer. Object detection is based at least on the output of the dilated convolution layer or the output of the second pooling layer.

[0006] It should be understood that one or more layers may be provided between the input data and the first pooling layer. Likewise, one or more other layers may be provided after the second pooling layer.

[0007] For example, a structure of a series of dilated convolutional layers may be provided, each of which is followed by a corresponding pooling layer. For example, an alternating structure of a single dilated convolutional layer followed by a corresponding single pooling layer may be provided.

[0008] Object detection can be understood as a combination of image classification and object localization. This means that object detection locates the presence and type of objects such as cars, buses, pedestrians, etc. by drawing bounding boxes around each object of interest in the image and assigning class labels to these objects. The input for object detection can be a data set with at least one object (e.g., from a radar sensor, or from a lidar sensor, or from a camera, such as an image). The output of object detection can be at least one bounding box and a class label for each bounding box, wherein a bounding box can be a rectangular box described by the coordinates of the center point of the bounding box and its width and height, and wherein a class label can be an integer value relating to a specific type of object. Object detection is widely used in many fields, for example, in autonomous driving technology by identifying the locations of vehicles, pedestrians, roads, obstacles, etc. in the captured images.

[0009] The input data may be data captured from a sensor such as a radar sensor, a camera, or a LIDAR sensor, or may include data from a system as described in detail below.

[0010] A convolutional layer may perform an operation known as a convolution operation or convolution. In the context of a convolutional neural network (CNN), a convolution may be a linear operation that processes the multiplication of a set of weights with an input data set. The set of weights may be referred to as a filter or a kernel. The multiplication may be a dot product, which means an element-by-element multiplication between a kernel element and a portion of the input data having the same size as the kernel. The multiplication results may be added and a single value may be produced. All single values ​​of all convolutions in a convolutional layer may define the output of the convolutional layer. The size of the kernel may be determined based on the number of weights.

[0011] The dilated convolutional layer can be, for example, Figure 4 The second dilated convolutional layer 428 is shown, which consists of a convolutional layer, wherein there is free space between kernel elements. In other words, the receptive field of the kernel is increased by adding zero weights between kernel elements. This will be described in detail below.

[0012] A pooling layer can perform an operation that can be referred to as a pooling operation or pooling. A pooling layer can aggregate the output of a dilated convolutional layer (i.e., focus on the most relevant information) and can reduce the number of parameters required to detect objects in given input data. Thus, the output of the pooling layer can be a narrowed down version of the output generated by the dilated convolutional layer, including summarized features rather than precisely located features. This can achieve better performance and lower memory requirements.

[0013] According to an embodiment, the method also includes the following steps performed by a computer hardware component: determining the output of another dilated convolutional layer directly disposed after the second pooling layer based on the output of the second pooling layer; and determining the output of another pooling layer directly disposed after the other dilated convolutional layer based on the output of the other dilated convolutional layer; wherein object detection is further performed based on at least the output of the other dilated convolutional layer or the output of the other pooling layer.

[0014] In other words, there is another dilated convolution layer and another pooling layer, wherein the another dilated convolution layer and the another pooling layer are arranged in the following manner: the another dilated convolution layer directly follows the second pooling layer (in other words: no additional layer is arranged between the second pooling layer and the another dilated convolution layer), and the another pooling layer directly follows the another dilated convolution layer (in other words: no additional layer is arranged between the another dilated convolution layer and the another pooling layer). The output of the second pooling layer is the input of the another dilated convolution layer, and the output of the another dilated convolution layer is the input of the another pooling layer. Object detection is further performed based on at least the output of the another dilated convolution layer or the output of the another pooling layer.

[0015] According to various embodiments, there may be a plurality of additional dilated convolutional layers and a plurality of additional pooling layers, wherein each of the plurality of additional dilated convolutional layers directly follows a corresponding pooling layer in the plurality of additional pooling layers, and the combination directly follows the already existing last pooling layer. Object detection may then be further performed based on at least one of the outputs of at least the respective layers (dilated convolutional layers and / or pooling layers).

[0016] Each dilated convolutional layer can identify different features that may be relevant to a specific task. By combining dilated convolutional layers with pooling layers, the dimension and therefore the number of parameters are reduced.

[0017] According to an embodiment, the method further comprises the following steps performed by a computer hardware component: upsampling the output of the dilated convolutional layer and the output of the another dilated convolutional layer to a predetermined resolution; and cascading the output of the dilated convolutional layer and the output of the another dilated convolutional layer, wherein object detection is performed based on the cascaded outputs.

[0018] As used herein, upsampling may refer to increasing the dimensionality of the output of a dilated convolutional layer and the output of the further dilated convolutional layer. This may be accomplished, for example, by repeating rows and columns of the output of the dilated convolutional layer and the output of the further dilated convolutional layer. Another possibility for upsampling may be a deconvolutional layer. A deconvolutional layer may perform a deconvolution operation, i.e., the forward pass and the reverse pass of a convolutional layer are reversed.

[0019] The predetermined resolution is the resolution of the input data, or the resolution of yet another dilated convolution layer or another convolution layer provided before the first pooling layer.

[0020] The output of yet another dilated convolutional layer can also be upsampled and concatenated.

[0021] The output of cascading the output of a dilated convolutional layer and the output of another dilated convolutional layer may be referred to as a cascaded output.

[0022] According to an embodiment, the dilation rate of the kernel of the dilated convolution layer may be different from the dilation rate of the kernel of the other dilated convolution layer.

[0023] The dilation rate of the kernel may define the spacing between kernel elements, wherein the spacing is filled with zero elements. A dilation rate of one means that there is no spacing between kernel elements, whereby the kernel elements are positioned in close proximity to each other. The size of a kernel with a dilation rate equal to or greater than two is increased by adding zero weights between kernel elements, so that, for example, a dilation rate of two may add one zero weight element between kernel elements, a dilation rate of three may add two zero weight elements between kernel elements, and so on. Different dilation rates, such as 2, 3, 4 or 5, may be used in the methods described herein. In another embodiment, at least two different dilation rates may be set among the kernels of the dilated convolutional layer, the kernel of the further dilated convolutional layer, and / or the kernels of other further dilated convolutional layers.

[0024] The stride of the kernel (of the corresponding layer) can define how the kernel or filter moves from left to right and from top to bottom across the input data (e.g., image) (of the corresponding layer). Using a higher stride (e.g., 2, 3, or 4) can have the effect of applying the filter in the following way: downsampling the output of the dilated convolutional layer, the further dilated convolutional layer, and / or other further dilated convolutional layers, where downsampling can refer to reducing the size of the output. Therefore, increasing the stride can reduce computation time and memory requirements. Different strides (e.g., 2, 3, or 4) of the kernel of the dilated convolutional layer, the further dilated convolutional layer, and / or other further dilated convolutional layers can be used in the methods described herein.

[0025] According to an embodiment, each of the first pooling layer and the second pooling layer includes mean pooling or max pooling.

[0026] The pooling operation may provide a downsampling of the image obtained from the previous layer, which may be understood as reducing the number of pixels of the image and thus reducing the number of parameters of the convolutional network. Max pooling may be a pooling operation that determines the maximum value (in other words: the largest value) in each considered region. Applying the max pooling operation to the downsampling may be computationally efficient. Mean pooling or average pooling may determine the average value of each considered region and may therefore retain information about elements that may be less important. This may be useful in cases where the location of an object in an image is important.

[0027] The further pooling layer and / or other further pooling layers may also include mean pooling or maximum pooling.

[0028] According to an embodiment, the size of the corresponding core of each pooling layer in the first pooling layer and the second pooling layer can be 2x2, 3x3, 4x4 or 5x5. The size of the core of the other pooling layer and / or the core of other additional pooling layers can also be 2x2, 3x3, 4x4 or 5x5.

[0029] Pooling in small local areas can smooth the mesh and thus reduce artifacts. It has been found that using smaller kernel sizes can achieve lower error rates.

[0030] It should be understood that reference to a core of a layer refers to the core used for operations of that layer.

[0031] According to an embodiment, the corresponding core of each pooling layer in the first pooling layer and the second pooling layer depends on the size of the core of the dilated convolution layer. The core of the other pooling layer may also depend on the size of the core of the other dilated convolution layer, and / or the core of other other pooling layers may also depend on the size of the core of other dilated convolution layers.

[0032] It has been found that performance can be enhanced if the kernel size of the pooling operation, the kernel size of the convolution operation, and the number of samples are selected in a corresponding relationship to each other.

[0033] According to an embodiment, the input data comprises a 2D grid having channels. The input data may be detected by a camera, wherein the input data is a two-dimensional image having RGB channels for color information.

[0034] According to an embodiment, the input data is determined based on data from a radar system.The input data may also be or may include data from said system.

[0035] Radar sensors are not affected by adverse or bad weather conditions and operate reliably in darkness, moisture or even fog. They are able to detect the distance, direction and relative speed of vehicles or hazards.

[0036] According to an embodiment, the input data is determined based on at least one of data from a LIDAR system or data from a camera system. The input data may also be or may include data from said system.

[0037] The input data recorded by the camera can be used to detect RGB information, for example to recognize traffic lights, road signs or red brake lights of other vehicles, with extremely high resolution.

[0038] The input data recorded from the LIDAR sensor can be very detailed and can include fine and accurate information about objects at long distances. Ambient light does not affect the quality of the information captured by the LIDAR, so day and night results can be obtained without any performance loss due to interference such as shadows, sunlight or headlight glare.

[0039] According to an embodiment, object detection is performed using a detection head. A corresponding detection head may be provided for each property of a detected object, such as the class of the object, the size of the object and / or the speed of the object.

[0040] According to an embodiment, the detection head comprises an artificial neural network. The neural network for the detection head can be trained together with the training of the artificial neural network, which comprises a first pooling layer, a dilated convolution layer, a second pooling layer (and, if applicable, the further pooling layer and the further dilated convolution layer).

[0041] In another aspect, the present disclosure is directed to a computer system configured to perform several or all of the steps of the computer-implemented methods described herein.

[0042] The computer system may include multiple computer hardware components (e.g., a processor (e.g., a processing unit or a processing network), at least one memory (e.g., a memory unit or a memory network), and at least one non-transitory data storage device). It should be understood that other computer hardware components may be provided and used to perform the steps of the computer-implemented method in the computer system. The non-transitory data storage device and / or the memory unit may include a computer program for instructing a computer, for example, using the processing unit and the at least one memory unit to perform some or all steps or aspects of the computer-implemented method described herein.

[0043] In another aspect, the present disclosure is directed to a vehicle comprising a computer system as described herein and a sensor, wherein the input data is determined based on an output of the sensor. The sensor may be a radar system, a camera, and / or a LIDAR system.

[0044] The vehicle may be a car or a truck, and the sensor may be mounted on the vehicle. The sensor may be directed toward an area in front of or behind or to the side of the vehicle. Images may be captured by the sensor while the vehicle is moving.

[0045] In another aspect, the present disclosure is directed to a non-transitory computer-readable medium including instructions for performing some or all steps or aspects of the computer-implemented methods described herein. The computer-readable medium may be configured as: an optical medium such as a compact disk or a digital versatile disk (DVD); a magnetic medium such as a hard disk drive (HDD); a solid-state drive (SSD); a read-only memory (ROM) such as a flash memory; and the like. Moreover, the computer-readable medium may be configured as a data storage device accessible via a data connection such as an Internet connection. The computer-readable medium may be, for example, an online data repository or cloud storage.

[0046] The present disclosure is also directed to a computer program for instructing a computer to perform several or all steps or aspects of the computer-implemented method described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Exemplary embodiments and functions of the present disclosure are described herein in conjunction with the following schematically illustrated drawings:

[0048] Figure 1 is a characteristic pyramid structure according to various embodiments;

[0049] Figure 2 are convolution kernels with different dilation rates according to various implementations;

[0050] Figure 3 is an example of a gridding artifact;

[0051] Figure 4 is a combination of dilated convolution and pooling according to various implementations;

[0052] Figure 5 is an implementation, wherein a pattern comprising a dilated convolutional layer and a second pooling layer is alternated multiple times in succession;

[0053] Figure 6 is a flow chart illustrating a method for object detection according to various embodiments; and

[0054] Figure 7The invention is a computer system having a plurality of computer hardware components, the computer system being configured to perform the steps of a computer-implemented method for object detection according to various embodiments. DETAILED DESCRIPTION

[0055] Convolutional Neural Networks (CNNs) can be used for many perception tasks, e.g., object detection, semantic segmentation. The input data for CNNs can be formulated as a 2D grid with channels, e.g., an image with RGB channels taken from a camera or a grid quantized from a LiDAR point cloud. Objects in perception tasks can have different sizes. For example, in an autonomous driving scenario, a bus may be much larger than a pedestrian. Furthermore, in the image domain, scale variations may be larger due to stereoscopic projections. To better recognize objects with multiple scales, a pyramid structure with different spatial resolutions can be provided.

[0056] Figure 1 A feature pyramid structure of a convolutional neural network (CNN) is shown 100. Perception tasks can be performed on each pyramid level (or layer) 102, 104, 106, 108, 110 or a fusion of them.

[0057] The pyramid structure starts at the first level 102 and refines features at each higher level 104, 106, 108, 110. In each step between the pyramid levels 112, 114, 116, 118, the spatial resolution is reduced. For example, a pooling operation (mean or max pooling) (followed by a downsampling operation) can be performed to reduce the spatial resolution of the grid so that the convolutions on it cover a larger area, also called the receptive field. If the spatial resolution is reduced and the size of the receptive field remains the same, the receptive field covers a larger area.

[0058] CNNs require large receptive fields to capture multi-scale objects in spatial context. However, pooling operations can discard detailed information embedded in local regions. Moreover, if upsampling is performed to restore spatial resolution, where the outputs 120, 122, 124, 126 of layers 104, 106, 108, 110 are used as input, smearing effects and smoothing effects may lead to further information loss.

[0059] Dilated convolution (meaning that the convolution kernel has at least two or higher dilation rates, such as Figure 2 ) can also have the effect of increasing the receptive field while still capturing local information well.

[0060] Figure 2An example diagram 200 of a convolution kernel of size 3x3 with different dilation rates is shown. If the dilation rate D is equal to one, the kernel of the dilated convolution can be consistent with the kernel of the normal convolution 202. As the dilation rate increases, the space between the original kernel elements 208 expands. For a dilation rate of two (D=2) 204, one space is left between each kernel element 208 of the original kernel. The kernel size remains 3x3, but the size of the receptive field is expanded to 5x5 (instead of 3x3). Therefore, a dilation rate of three (D=3) 206 means that there are two free spaces between two kernel elements 208, and the size of the receptive field grows to 7x7.

[0061] However, if multiple dilated convolutions are stacked (especially with the same dilation rate), spatial inconsistency problems, i.e., gridding artifacts, occur. In this case, neighboring positions in the grid come from different sets of positions in the preceding grid without any overlap. In other words, neighboring positions are no longer guaranteed to be smooth.

[0062] Figure 3 An example diagram 300 showing the spatial inconsistency problem of gridding artifacts. The operations between the layers are all dilated convolutions with a kernel size of 3x3 and a dilation rate of two. The four neighboring positions 310, 312, 314, 316 in layer i 302 are indicated by different shading. Their actual receptive fields in layer i-1 304 and layer i-2 306 are marked with the same shading, respectively. It should be understood that their actual receptive fields are separate sets of positions.

[0063] As mentioned above, pooling and dilated convolutions can be used to increase the receptive field. According to various embodiments, pooling and dilated convolutions can be combined. It has been found that when a pooling layer is inserted after each dilated convolution layer, gridding artifacts introduced by stacking multiple dilated convolutions can be reduced or avoided.

[0064] Figure 4An example diagram 400 of a combined structure of dilated convolution and pooling according to various embodiments is shown. Input data 402 may come from a digital device, such as a camera, radar, or LIDAR sensor. Input data 402 may also be or include data from the system. A first grid level 420 may include a 2D grid with channels, such as an image with RGB channels. A first dilated convolution layer 422 extracts features from the first grid level 420, wherein different dilation rates (e.g., 2, 3, 4, or 5) may be used to expand the receptive field. The input of the first dilated convolution layer 422 is the output 404 of the first grid level 420 (same as the input 402). The first dilated convolution layer 422 may be performed directly before the first pooling layer 424 to maintain local information at a limited resolution. The first pooling layer 424 may directly follow the first dilated convolution layer 422, wherein the input 406 of the first pooling layer 424 is the output 406 of the first dilated convolution layer 422. It is important that the first dilated convolutional layer 422 is directly followed by the first pooling layer 424 (without any further layers in between), otherwise gridding artifacts may occur.

[0065] The pooling operation within the first pooling layer 424 may be performed as mean pooling or max pooling and may be performed in small local areas (e.g., 2x2, 3x3, or 4x4) to smooth the grid and thus reduce artifacts. The output 408 of the first pooling layer 424 may specify a second grid level 426 and may be used as an input 410 to a second dilated convolutional layer 428 that directly follows the first pooling layer 424. Directly after the second dilated convolutional layer 428, a second pooling layer 430 is provided, wherein the input 412 of the second pooling layer 430 is the output 412 of the second dilated convolutional layer 428. Figure 4 As shown, outputs 408 and 412 with different spatial resolutions can be provided while maintaining detailed local information at each resolution. Finally, object detection is performed based on at least the output 412 of the second dilated convolutional layer 428 or the output 414 of the second pooling layer 430.

[0066] Figure 4 The method shown can be adopted as the backbone of a radar-based object detection network that detects vehicles in autonomous driving scenarios. The method provides high performance, especially by detecting large moving vehicles. This is likely due to the larger receptive field given by the combination of dilated convolutions and pooling. At the same time, detailed local information can be maintained to provide good detection of stationary vehicles that may rely more on local spatial structure.

[0067] Figure 5 An example diagram 500 showing alternating application of dilated convolutions and pooling in accordance with various embodiments is shown. Figure 5 The various elements shown can be used with Figure 4 The elements shown are similar or identical, so the same reference numerals may be used and repeated descriptions may be omitted.

[0068] Additional convolutional layers 502, 504, 512, additional dilated convolutional layers 506, 510, and another pooling layer 508 are added to Figure 4 In the combined structure of the dilated convolution and pooling shown. Immediately before the first dilated convolution layer 422, another convolution layer 504 can be set, which can be a dilated convolution layer. The input data 402 of the first dilated convolution layer 422 can be the output of another convolution layer 504, so the input data 402 is data from the (sensor) system. Before the other convolution layer 504, another convolution layer 502 can be set, which can be a dilated convolution layer. The convolution layer 502 can be directly connected to the other convolution layer 504. The output 414 of the second pooling layer 430 can be the input of another dilated convolution layer 506 (which can directly follow the second pooling layer 430). Directly after the other dilated convolution layer 506, another pooling layer 508 can be set, wherein the input of the other pooling layer 508 can be the output of the other dilated convolution layer 506. Directly after the further pooling layer 508, a further dilated convolutional layer 510 may be provided, wherein the input of the further dilated convolutional layer 510 may be the output of the further pooling layer 508. Directly after the further dilated convolutional layer 510, a further convolutional layer 512 may be provided, which may be a dilated convolutional layer. The input of the further convolutional layer 512 may be the output of the further dilated convolutional layer 510. Thus, a wide range of receptive fields may be embedded in the multi-level outputs 406, 412, 514, 516, 518, 530, 532. In 520, 522, 524, 526, respectively, the output 412 of the dilated convolutional layer 428, the outputs 514, 516 of the further dilated convolutional layers 506, 510, and the output 518 of the further convolutional layer 512 may be upsampled to a predetermined resolution. The outputs 534, 536, 538, 540 of the upsampling 520, 522, 524, 526 and the output 532 of another convolutional layer 504 and the output 530 of yet another convolutional layer 502 may be cascaded in a cascade block 528. The output 542 of the cascade block 528 may be an input to a detection head 546, and the output 544 of the cascade block 528 may be an input to another detection head 548. Object detection is performed using the outputs 550, 552 of the detection heads 546, 548.

[0069] In another embodiment, the plurality of outputs 406, 412, 514, 516, 518, 530, 532 may be directly connected to a plurality of task-specific detection heads (not shown).

[0070] Figure 6A flow chart 600 of a method for object detection according to various embodiments is shown. At 602, an output of a first pooling layer can be determined based on input data. At 604, an output of a dilated convolutional layer disposed directly after the first pooling layer can be determined based on the output of the first pooling layer. At 606, an output of a second pooling layer disposed directly after the dilated convolutional layer can be determined based on the output of the dilated convolutional layer. At 608, object detection can be performed based on at least the output of the dilated convolutional layer or the output of the second pooling layer.

[0071] According to various embodiments, the method may further include the following steps: determining the output of another dilated convolutional layer directly disposed after the second pooling layer based on the output of the second pooling layer; and determining the output of another pooling layer directly disposed after the other dilated convolutional layer based on the output of the other dilated convolutional layer; wherein object detection is also performed based at least on the output of the other dilated convolutional layer or the output of the other pooling layer.

[0072] According to various embodiments, the method may further include the steps of: upsampling the output of the dilated convolution layer and the output of another dilated convolution layer to a predetermined resolution; and cascading the output of the dilated convolution layer and the output of another dilated convolution layer; wherein object detection is performed based on the cascaded output.

[0073] According to various embodiments, a dilation rate of a kernel of a dilated convolutional layer may be different from a dilation rate of a kernel of another dilated convolutional layer.

[0074] According to various embodiments, each of the first pooling layer and the second pooling layer may include or may be mean pooling or maximum pooling.

[0075] According to various embodiments, the size of the corresponding kernel of each of the first pooling layer and the second pooling layer may be 2x2, 3x3, 4x4, or 5x5.

[0076] According to various embodiments, the corresponding kernel of each of the first pooling layer and the second pooling layer may depend on (rely on) the size of the kernel of the dilated convolution layer.

[0077] According to various embodiments, the input data may include or may be a 2D grid having channels.

[0078] According to various embodiments, the input data may be determined based on data from a radar system.

[0079] According to various embodiments, the input data may be determined based on at least one of data from a LIDAR system and data from a camera system.

[0080] According to various implementations, a detection head may be used to perform object detection.

[0081] According to various embodiments, the detection head may include or may be an artificial neural network.

[0082] Each of the above steps 602, 604, 606, 608 and another step may be performed by computer hardware components.

[0083] Using the methods and systems described herein, deep learning based object detection can be provided.

[0084] Figure 7 A computer system 700 having a plurality of computer hardware components is shown, the computer system being configured to perform the steps of a computer-implemented method for object detection according to various embodiments. The system 700 may include: a processor 702, a memory 704, and a non-transitory data storage device 706. A camera 708 and / or a distance sensor 710 (e.g., a radar sensor or a LIDAR sensor) may be provided as part of the computer system 700 (e.g., Figure 7 exemplified), or may be located outside the computer system 700.

[0085] The processor 702 may execute instructions provided in the memory 704. The non-transitory data storage device 706 may store a computer program including instructions that may be transferred to the memory 704 and then executed by the processor 702. The camera 708 and / or the range sensor 710 may be used to determine input data, for example, input data provided to a convolutional layer or a dilated convolutional layer or a pooling layer as described herein.

[0086] The processor 702, the memory 704 and the non-transitory data storage device 706 may be connected to each other for exchanging electrical signals, for example, via an electrical connection 712 (e.g., such as a cable or a computer bus) or via any other suitable electrical connection. The camera 708 and / or the distance sensor 710 may be connected to the computer system 700, for example, via an external interface, or may be arranged as part of the computer system (in other words: inside the computer system, for example, connected via the electrical connection 712).

[0087] The terms "coupled" or "connected" are intended to include a direct "coupled" (eg, via a physical link) or direct "connected" as well as an indirect "coupled" or indirect "connected" (eg, via a logical link), respectively.

[0088] It should be understood that what is described above with respect to one of the methods may similarly be applied to the computer system 700 .

[0089] Label list

[0090] 100 Feature pyramid structure

[0091] 102, 104, 106, 108, 110 Pyramid level (or layer)

[0092] 112, 114, 116, 118 Steps between pyramid levels

[0093] Output of pyramid levels 120, 122, 124, and 126

[0094] 200 Example of a 3x3 convolution kernel with different dilation rates 200

[0095] 202 Kernel of normal convolution with dilation rate D = 1

[0096] 204 Convolution kernel with dilation rate D = 2

[0097] 206 Convolution kernel with dilation rate D = 3

[0098] 208 Nuclear Elements

[0099] 300 Illustration of spatial inconsistency problem of gridding artifacts

[0100] 302 layeri

[0101] 304 layer i-1

[0102] 306 layer i-2

[0103] 310, 312, 314, 316 nearby locations

[0104] 400 An example diagram of a combined structure of dilated convolution and pooling according to various embodiments

[0105] 402 Input Data

[0106] 404 Output of the first grid level

[0107] 406 Input of the first pooling layer, output of the first dilated convolutional layer

[0108] 408 Output of the first pooling layer

[0109] 410 Input of the second dilated convolutional layer

[0110] 412 Input of the second pooling layer, output of the second dilated convolutional layer

[0111] 414 Output of the second pooling layer

[0112] 420 First grid level

[0113] 422 First dilated convolutional layer

[0114] 424 First Pooling Layer

[0115] 426 Second Grid Level

[0116] 428 Second dilated convolutional layer

[0117] 430 Second Pooling Layer

[0118] 500 An example diagram of alternating application of dilated convolution and pooling according to various embodiments

[0119] 502 Another convolutional layer

[0120] 504 Another convolutional layer

[0121] 506 Another dilated convolutional layer

[0122] 508 Another pooling layer

[0123] 510 Another dilated convolutional layer

[0124] 512 Another convolutional layer

[0125] 514, 516, 518 output

[0126] 520, 522, 524, 526 upsampling

[0127] 528 Cascade Frame

[0128] 530, 532 Output of another convolutional layer

[0129] 534, 536, 538, 540 upsampled output

[0130] 542, 544 Cascade frame output, detection terminal input

[0131] 546, 548 detection terminal

[0132] 550, 552 detection terminal output

[0133] 600 Flowchart of a method for object detection according to various embodiments

[0134] 602 Step of determining the output of the first pooling layer

[0135] 604 Steps to determine the output of the dilated convolutional layer

[0136] 606 Step of determining the output of the second pooling layer

[0137] 608 Steps to perform object detection

[0138] 700 Computer system according to various embodiments

[0139] 702 processor

[0140] 704 Memory

[0141] 706 Non-transitory data storage device

[0142] 708 Camera

[0143] 710 Distance Sensor

[0144] 712 connections

Claims

1. A computer-implemented method for object detection, the method comprising the following steps performed by computer hardware components: determining an output (408) of a first pooling layer (424) based on the input data (402), in, The input data is an image; Based on the output (408) of the first pooling layer (424), determining the output (412) of a dilated convolutional layer (428) disposed directly after the first pooling layer (424); Based on the output (412) of the dilated convolutional layer (428), determining the output (414) of a second pooling layer (430) disposed directly after the dilated convolutional layer (428); Performing the object detection based on at least an output (412) of the dilated convolutional layer (428) or an output (414) of the second pooling layer (430); Based on the output (414) of the second pooling layer (430), determining the output of another dilated convolutional layer (506) disposed directly after the second pooling layer (430); Based on the output of the further dilated convolutional layer (506), determining the output of a further pooling layer (508) arranged directly after the further dilated convolutional layer (506), wherein the object detection is also performed based on at least an output of the further dilated convolutional layer (506) or an output of the further pooling layer (508); Upsampling the output of the dilated convolutional layer (428) and the output of the further dilated convolutional layer (506) to a predetermined resolution; and cascading (528) the output of the dilated convolutional layer (428) and the output of the further dilated convolutional layer (506), The object detection is performed based on the concatenated outputs (542, 544).

2. The method according to claim 1, in, The dilation rate of the kernel of the dilated convolution layer (428) is different from the dilation rate of the kernel of the other dilated convolution layer (506).

3. The method according to claim 1, in, Each pooling layer in the first pooling layer (424) and the second pooling layer (430) includes mean pooling or maximum pooling.

4. The method according to claim 1, in, The size of the corresponding kernel of each pooling layer in the first pooling layer (424) and the second pooling layer (430) is 2x2, 3x3, 4x4 or 5x5.

5. The method according to claim 1, in, The corresponding kernel of each pooling layer in the first pooling layer (424) and the second pooling layer (430) depends on the size of the kernel of the dilated convolution layer (428).

6. The method according to claim 1, in, The input data (402) includes a 2D grid having channels.

7. The method according to claim 1, in, The input data (402) is determined based on data from a radar system.

8. The method according to claim 1, in, The input data (402) is determined based on at least one of data from a LIDAR system and data from a camera system.

9. The method according to claim 1, in, The object detection is performed using a detection head (546, 548).

10. The method according to claim 9, in, The detection end (546, 548) includes an artificial neural network.

11. A computer system (700), comprising a plurality of computer hardware components configured to perform the steps of the computer-implemented method according to any one of claims 1 to 10.

12. A vehicle comprising sensors (708, 710) and a computer system (700) according to claim 11, in, The input data (402) is determined based on outputs of the sensors (708, 710).

13. A non-transitory computer-readable medium comprising instructions for executing the computer-implemented method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • System and method for semantic segmentation using hybrid dilated convolution (HDC)

    US20180260956A1