Method, artificial neural network and device for semantic segmentation of image data
By introducing a splicing function between the encoder and decoder paths of the convolutional neural network, the problem of excessive computing and storage resources demands in the semantic segmentation of image data is solved, and more efficient image data processing and classification are achieved.
Patent Information
- Application Number
- CN201910950033.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-05
- Filing Date
- 2019-10-08
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2039-10-08
AI Technical Summary
The prior art uses convolutional neural networks to perform semantic segmentation of image data, and the demand for computing resources and storage resources is too high, especially when running on embedded computing units, resulting in performance bottlenecks.
By introducing a splicing function between the encoder path and the decoder path, the splicing of the input tensor and the jump tensor is realized, the merged tensor is generated, and the convolution function is applied to generate a review tensor. The review tensor and the input tensor are then spliced for the second time to generate an output tensor for further processing in the decoder path.
This method not only allows processing image data more precisely and improves classification accuracy, but also applies to resource-constrained embedded systems by reducing computing and storage requirements.
Smart Images

Figure CN111008973B_ABST
Abstract
Description
Technical Field
[0001] The invention proceeds from a method, an artificial neural network and a device for semantic segmentation of image data. Background Art
[0002] "Fully Convolutional Models for Semantic Segmentation by Evan Shelhamer, Jonathan Long, Trevor Darrell (PAMI: Transactions on Pattern Analysis and Machine Intelligence, 2016)" discloses an extension of convolutional neural networks. Convolutional neural networks are powerful artificial neural networks that process visual data that can generate a feature hierarchy of the semantics of the visual data. The document discloses the following scheme: using a "fully convolutional network" that can receive input data of any size and can output an output corresponding in size by efficient derivation of features.
[0003] "Convolutional Networks for Biomedical Image Segmentation by Olaf Ronneberger, Philipp Fischer, Thomas Brox. U-Net (Medical Image Computing and Computer-Assisted Intervention (MICCAI), Springer, LNCS, Vol. 9351)" discloses a training strategy for the network, which is based on the use of extended (enhanced) training data in order to use existing annotated examples more effectively. The architecture of the network includes a "Contracting Path" (encoder path) for detecting the context of the input data, and symmetrically includes an "Expanding Path" (decoder path) that enables precise positioning of the detected context. The artificial neural network can be trained using a relatively small amount of training data. Summary of the invention
[0004] Artificial neural networks, in particular so-called convolutional neural networks (CNNs), for semantic segmentation of image data, in particular for localization and classification of objects in image data, have a high demand for computing resources. The image data are reconstructed to the original resolution after semantic analysis in the encoder component by means of a decoding component or an upsampling component and a splicing component (jump component), which further significantly increases the demand for computing resources due to the addition of a decoding component or an upsampling component and a splicing component (jump component). In some implementations, this may lead to an exponential increase in the demand for computing resources.
[0005] In addition to the increased demand for computing resources when using artificial neural networks, especially convolutional neural networks, pixel-based semantic segmentation of image data also requires more storage resources, i.e., more memory bandwidth, memory accesses, and memory space during the training phase and when applying the network.
[0006] This disadvantage of the additional requirement for computing and storage resources is exacerbated if the application is not executed on special computing units with high memory and distributed computing, such as graphics processing unit clusters (GPU clusters), but on embedded computing units, such as embedded hardware or the like.
[0007] In this context, the present invention provides a method, an artificial neural network, a device, a computer program and a machine-readable storage medium for semantic segmentation of image data of an imaging sensor.
[0008] Image data may be understood here as data of an imaging sensor. Image data are to be understood primarily as data of a video sensor, thus a camera. Similarly, due to the similarity of the data, data of a radar sensor, an ultrasonic sensor, a lidar sensor or the like may be processed as image data with the aid of the present invention. Therefore, in relation to the present invention, a radar sensor, an ultrasonic sensor, a lidar sensor or the like may be understood as an imaging sensor.
[0009] Of particular importance for the present invention are image data of imaging sensors or the like which are suitable for use in vehicles, thus automotive image sensors.
[0010] In this context, semantic segmentation is to be understood as the processing of image data, the purpose of which is to determine the semantic class of objects contained in the image and the location of these objects in the image. It should be taken into account that global information in the image allows conclusions to be drawn about the semantic class of the object, whereas local information in the image allows conclusions to be drawn about the location of the object in the image.
[0011] One aspect of the present invention is a method for semantic segmentation of image data by means of an artificial neural network, in particular a convolutional neural network (CNN). The artificial neural network has an encoder path for ascertaining semantic categories in the image data and a decoder path for locating the ascertained categories in the image data. The method comprises the following steps:
[0012] By means of a first concatenation function, the input tensor and the jump tensor are first concatenated or merged to obtain a merged tensor.
[0013] Here, the input tensor and the jump tensor may be associated with the image data.
[0014] A function of the neural network, in particular a convolution, is applied to the combined tensor in order to obtain a review tensor.
[0015] The review tensor is secondly concatenated or merged with the input tensor by means of a second concatenation function to obtain an output tensor.
[0016] Feed the output tensor into the decoder path of the artificial neural network.
[0017] In this case, an artificial neural network is to be understood as a network of artificial neurons for processing information, for example for processing image data, in particular for localizing and classifying objects in image data.
[0018] Here, a convolutional neural network (CNN) is to be understood as a type of artificial neural network that is considered "state of the art" in the field of classification. The basic structure of a CNN consists of an arbitrary sequence of convolutional and pooling layers, which are terminated by one or more fully connected layers. The corresponding layers are composed of artificial neurons.
[0019] In this case, an encoder path is to be understood as meaning the path from the processing of the image data to the classification of the objects in the image data.
[0020] In this context, a decoder path is to be understood as a path which is connected to the encoder path and which reconstructs the original image data based on the classification in order to localize the classified objects.
[0021] In this context, a tensor is understood to be a data representation during processing in an artificial neural network. The data items include the processed state of the image data and the associated feature maps. The tensor at level l of the i-th step in an artificial neural network is usually represented as It has n rows, m columns, and f feature maps.
[0022] The input tensor is a data representation before being processed by the method of the invention. According to the invention, the input tensor is based on the up-converted output tensor of the previous stage l-1 of the artificial neural network. Here, the up-conversion is performed as follows: the dimension of the up-converted tensor corresponds to the dimension of the jump tensor in the first splicing step.
[0023] The skip tensor is a representation of the data at level l in the j-th step in the neural network. The skip tensor can be provided by a concatenation component (skip component) and thus provides information from level l of the artificial neural network in the encoder path to the decoder path of the artificial neural network directly, i.e. without further processing in the encoder path.
[0024] Here, a splicing component is to be understood as an architectural component in an artificial neural network for semantic segmentation, which provides information from the encoder path at the corresponding position of the decoder path. The splicing component can appear as a skip splicing or as a skip module.
[0025] The merged tensor is a data representation after the first concatenation step of the method according to the present invention. The merged tensor is the result of the first concatenation function. As concatenation function, function concatenation, addition, multiplication or the like can be considered.
[0026] The review tensor is a data representation after the application of the function of the artificial neural network, especially the convolutional neural network (CNN) according to the method of the present invention. As artificial neural networks, especially CNNs, the function convolution (Convolution) can be considered, and also the configuration of convolution blocks, that is, multiple applications of convolution, depthwise convolution, compression, residual (Residual), density (Dense), Inception, activation (Activation, Act), normalization, pooling (Pooling) or the like.
[0027] Inception is to be understood here as an architectural variant of an artificial neural network, in particular a convolutional neural network, which was first described in “Going deeper with convolutions” by Szegedy et al. (Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-9, 2015).
[0028] The output tensor is a data representation after the second concatenation step of the method according to the invention, which is used for further processing in the decoder path of the artificial neural network. The output tensor is the result of the second concatenation function. As concatenation functions, concatenation, addition, multiplication or the like can be considered.
[0029] In this context, a feature map is to be understood as the output of a layer of an artificial neural network. In a CNN, this is usually the result of processing by a convolutional layer followed by an associated pooling layer and can be used as input data for subsequent layers or, if provided, for a full concatenation layer.
[0030] Here, a function of an artificial neural network can be understood as any function of a neuron layer of an artificial neural network. The function can be a convolution - also a configuration of a convolution block - that is, multiple applications of convolution, depthwise convolution, compression, residual, dense, inception, activation, normalization, pooling or the like.
[0031] Inception is to be understood here as an architectural variant of an artificial neural network, in particular a convolutional neural network, which was first described in “Going deeper with convolutions” by Szegedy et al. (Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-9, 2015).
[0032] An advantage of the method of the present invention lies in the step of applying the function of the artificial neural network to obtain the verification tensor and the subsequent step of splicing the verification tensor with the input tensor. In the step of applying, not only the coarse-grained features from the encoder path but also the fine-grained features from the decoder path are spliced with each other. In the step of splicing, the input tensor is refined with the help of the verification tensor to generate the output tensor for the next layer.
[0033] According to one embodiment of the method according to the invention, in the step of applying, the function of the artificial neural network is associated with one or more feature maps of the input tensor, i.e. the function is selected such that, despite the application of the function to the combined sensor, the function is still suitable for the one or more feature maps.
[0034] This embodiment of the method has the advantage that the classification performed in the deeper layers of the artificial neural network is thereby refined, ie performed more accurately.
[0035] According to one specific embodiment of the method according to the invention, the first concatenation function and the second concatenation function are designed in such a way that the dimensionality of the input tensor is preserved.
[0036] According to one embodiment of the method according to the invention, the steps of the method are performed in a decoder path of the artificial neural network.
[0037] Another aspect of the invention is an artificial neural network for localizing and classifying image data, wherein the artificial neural network has an encoder path for classifying image data, a decoder path for localizing image data and is configured to carry out the steps of the method according to the invention.
[0038] The artificial neural network configured in this way is preferably used in a technical system, in particular a robot, a vehicle, a tool or a machine tool, in order to determine an output variable as a function of an input variable. Sensor data or a variable associated with the sensor data may be considered as an input variable of the artificial neural network. The sensor data may originate from a sensor of the technical system or may be received from the outside by the technical system. Based on the output variable of the artificial neural network, at least one actuator of the technical system is controlled by a control device of the technical system with the aid of a control signal. In this way, for example, the movement of a robot or a vehicle may be controlled, or a tool or a machine tool may be controlled.
[0039] In one embodiment of the artificial neural network according to the present invention, the artificial neural network can be configured as a convolutional neural network.
[0040] A further aspect of the invention is a device which is configured to carry out the steps of the method according to the invention.
[0041] A further aspect of the invention is a computer program which is configured to carry out the steps of the method according to the invention.
[0042] Another aspect of the invention is a machine-readable storage medium on which an artificial neural network according to the invention or a computer program according to the invention is stored.
[0043] Details and embodiments of the present invention are explained in more detail below with reference to a number of figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A block diagram of a fully convolutional network in the prior art is shown;
[0045] Figure 2 A block diagram showing a U-Net architecture of a convolutional network in the prior art;
[0046] Figure 3 A diagram showing an artificial neural network according to a fully convolutional network architecture;
[0047] Figure 4 A diagram showing an artificial neural network according to the "U-Net" architecture;
[0048] Figure 5 A diagram showing a concatenation function in a decoder block of an artificial neural network according to a fully convolutional network architecture;
[0049] Figure 6 A diagram showing a concatenation function in an artificial neural network according to a U-Net architecture;
[0050] Figure 7 A diagram showing a concatenation function in an artificial neural network according to the U-Net architecture according to the present invention;
[0051] Figure 8 A diagram showing the application of the present invention in an artificial neural network according to a fully convolutional network architecture;
[0052] Fig. 9 A flow chart of one specific embodiment of a method 900 according to the present invention is shown. DETAILED DESCRIPTION
[0053] Figure 1 Block diagram showing a fully convolutional network from "Fully Convolutional Models for Semantic Segmentation by Evan Shelhamer, Jonathan Long, Trevor Darrell, PAMI, 2016".
[0054] The figure summarizes part of the process shown in the artificial neural network into modules.
[0055] Block coder 110 shows the processing steps through a plurality of layers of a convolutional neural network (CNN) starting from image data as input data 111. The convolution layer 112a and the pooling layer 112b are clearly evident from the figure.
[0056] In the block decoder 120, the "deconvolution" results 121, 122, 123 of the CNN are shown. In this case, the deconvolution can be achieved by the inversion of the convolution step. In this case, it is possible to map the image of the coarse-grained classification result onto the original image data in order to thus localize the classified objects.
[0057] In the box jump module 130, the concatenation of the higher-level classified intermediate results of the CNN to the "deconvolution" results is shown. Thus, in the second row, the intermediate results of the 4th pool have been associated with the final result 122, and in the third row, the intermediate results of the 3rd and 4th pools are associated with the final result 123.
[0058] The advantage of these associations lies in the possibility to determine finer details and at the same time preserve higher level semantic information.
[0059] Figure 2 Block diagram showing the U-Net architecture of a convolutional network from "Olaf Ronneberger, Philipp Fischer, Thomas Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation, Medical Image Computing and Computer-Assisted Intervention (MICCAI), Springer, LNCS, Vol. 9351".
[0060] Block encoder 210 shows the processing steps starting from image data as input data 211 through a plurality of layers of a convolutional neural network (CNN) in order to classify the input data 211 .
[0061] The “deconvolution step (up-convolution)” is shown in the block decoder 220 , which proceeds from the deepest classification level through a corresponding number of deconvolution layers (layers) to a semantic segmentation map 221 of the localized and classified objects with the input data 211 .
[0062] In block 230, the splicing (jump splicing) between the classification layer (Layer) and the corresponding localization layer (Layer) is shown. These splicings show the information flow between the classification task and the localization task in the artificial neural network. It is thus possible to make the coarse-grained semantic segmentation consistent with the higher degree of reconstruction of the input data.
[0063] Figure 3 An artificial neural network in the form of a tensor-based "fully convolutional network" (FCN) is shown. Here, image data is provided to the network as an input tensor 301. A feature map 302 is generated based on the input tensor 301 by a convolution function, the so-called convolution. The information in the feature map obtained by the convolution function is compressed by the so-called pooling and mapped to the encoder tensor 303. The encoder tensor 303 is used as an input tensor for the next deeper layer (layer) of the fully convolutional network. The convolution function, the so-called convolution, is reapplied here in order to obtain semantically richer information mapped into the corresponding feature map 302. The result tensor 310 of the encoder block of the fully convolutional network is used as an input tensor for the decoder block of the fully convolutional network, wherein the semantically rich information is enriched with the help of the localization information of the higher layers (layers) of the fully convolutional network.
[0064] For this purpose, the result tensor 310 is first upconverted (upsampling: upsampling) into an upsampled tensor 304 and concatenated with a skip tensor 306, which has been derived, for example, from a higher-level encoder tensor 303. It is also conceivable that, instead of the encoder tensor 303, one or more feature map tensors 302 are conveyed from the encoder block of the fully convolutional network to the decoder block of the fully convolutional network by means of a skip module.
[0065] The result of this process is a decoder tensor 315 which is up-converted and used as an up-sampled tensor 304 for the next higher level layer of the decoder block of the fully convolutional network.
[0066] At the end of the decoder block, the decoder tensor 315 may be up-converted to the original parameters of the input tensor 301 .
[0067] The result is semantically segmented image data 320 having class and localization information about objects or features contained in the image data.
[0068] Because in fully convolutional networks, no semantic information is transferred between deeper and finer representations (i.e., at deeper layers of the network), the finer representations are less unique. Thus, these layers contribute more significantly to the determination error.
[0069] In addition, deeper layers are less susceptible to the so-called “Gradient Vanishing.” The fewer layers are removed from the input tensor 301, the more significantly the effect of “Gradient Vanishing” affects these layers.
[0070] "Gradient vanishing" is understood to be an effect that may occur during the training of artificial neural networks, where the change in parameters vanishes slightly. In the worst case, this effect leads to a stagnation of the change or improvement of the trained parameters.
[0071] The introduction of skip module 130 or skip splicing 230 helps to overcome this effect.
[0072] For these reasons, fully convolutional networks are particularly suitable for a large number of semantic categories (i.e., more than 3 categories) and for flat networks, since the semantic features of finer layers can no longer be distinguished.
[0073] Figure 4 A diagram showing an artificial neural network according to the "U-Net" architecture.
[0074] According to the diagram, the processing of image data is performed from left to right. The image data to be processed will be introduced into the artificial neural network as input tensor 410. Input tensor 401 represents the image data to be processed. By applying functions of the neural network, such as convolution - also the configuration of convolution blocks - that is - multiple applications of convolution, depth convolution, compression, residual, density, inception, activation, normalization, pooling or the like, a feature map 402 can be generated from the input tensor 401 and the feature map 402 can be further processed as a tensor in the network.
[0075] Artificial neural networks are usually constructed in layers. In a layer, the following functions of the artificial neural network are usually applied: the functions do not lead to a change in the resolution of the tensor.
[0076] In the case of layer transformations, the following functions of artificial neural networks are usually applied: As a result, the resolution of the tensor is changed. In the direction of deeper layers, the resolution is reduced (pooling, downsampling), while in the direction of higher layers, the resolution is upconverted (upsampling).
[0077] In order to perform downsampling, a so-called pooling function can be applied to the tensor. As a result of the pooling function, a pooled tensor 403 exists, which serves as an input tensor for deeper layers in the artificial neural network. Figure 4 As shown in the diagram in , the function of the artificial neural network can be applied to the pooling tensor 403 to obtain the feature map 402 based on the pooling tensor 403.
[0078] In a U-Net architecture, the deepest layer is reached if the image data has been processed to such an extent that there is (sought or desired) classification information. Typically, information about the presence of a certain semantic class in the image data lacks information about the localization of the identified semantic class, i.e., information about where the identified class is located in the image data).
[0079] The U-Net architecture provides a decoder path for this purpose, in which the tensors (pooling tensor 403 and feature map 402) are up-converted (up-sampled). Depending on the application, the up-conversion can be performed up to the original resolution of the image data.
[0080] With the addition of information from the corresponding level of the encoder path 210, up-conversion by the deepest layer of the artificial neural network can be achieved. Figure 4 In the illustration of , this is illustrated by the splicing arrows between the corresponding layers of the encoder path 210 and the corresponding layers of the decoder path 220.
[0081] This addition is achieved by concatenating the tensor of one layer of up-conversion in the decoder path 220 with the skip tensor from the encoder into a cascaded tensor 411 in the decoder 220 .
[0082] Functions of artificial neural networks such as convolution - as well as configurations of convolution blocks - i.e. multiple applications of convolution, depthwise convolution, compression, residual, dense, inception, activation (Act), normalization, pooling or the like can be applied to the cascaded tensor 411 to obtain a feature map 412 in the decoder path 220.
[0083] The result of the decoder path 220 is a result tensor 420 in which a representation of the processed image data is shown, and in addition to the image data, the identified semantic categories are also shown, as well as the location of these semantic categories in the image data.
[0084] By concatenating features of encoder paths 210 with each other and by subsequently concatenating (merging) with knowledge about deeper and finer levels of the network, the U-Net architecture allows image data to be accurately defined up to the original resolution.
[0085] This architecture aims to address the shortcomings of the fully convolutional network architecture by using more resources.
[0086] Such resource use may lead to increased costs. The increased costs can be counteracted by keeping the number of output classes, ie the number of objects to be differentiated in the image data, small, for example in the order of magnitude of 2 to 3 classes.
[0087] The biggest drawback of the U-Net architecture is the strong effect of "vanishing gradients" in the deeper layers of the network. This effect is caused by the multiple layers arranged between the "loss function" and the discriminative layer.
[0088] Therefore, the U-Net architecture is particularly suitable for tasks that require only a small number of categories and therefore high localization accuracy.
[0089] Figure 5 A diagram showing the concatenation function in the decoder block of an artificial neural network according to a fully convolutional network architecture.
[0090] In the encoder block of the network, an encoder tensor 501 is formed by using the functions of the artificial neural network.
[0091] By means of the skip module, the encoder tensor 501 can be directly provided to the decoder module as a skip tensor 502 without further processing in the encoder block and, if necessary, in the decoder block.
[0092] The result tensor is provided by a deeper layer of the decoder block or by the deepest layer of the encoder block at the beginning of the decoder as decoder tensor 503. The decoder tensor 503 is first up-converted (up-sampled) into an up-sampled tensor 504 when entering the next higher-level layer. The up-sampled tensor 504 and the skip tensor 502 are concatenated (merged) with each other by means of a concatenation function 520 and thus form the result tensor 515 of the layer shown.
[0093] Figure 6A diagram showing the concatenation function (merging) in an artificial neural network according to the U-Net architecture.
[0094] As shown in the above figure, the up-converted decoder tensor 502 is concatenated (merged) as an up-sampled tensor 504 with the skip tensor 502 into a concatenated tensor 605 by means of a concatenation function 520. In this figure, concatenation is applied as the concatenation function 520. Similarly, other concatenation functions, such as addition, multiplication or the like, can be considered.
[0095] Subsequently, a convolution function 620 (Convolution) of the artificial neural network is applied to the concatenated tensor 605 to form a result tensor 615 of the layer shown.
[0096] With the help of the convolution function (Convolution) 620, coarse and fine semantic features are concatenated with each other without direct relation to the target output category.
[0097] Figure 7 A diagram showing a concatenation function in an artificial neural network extended with the present invention according to a U-Net architecture.
[0098] As shown in the above diagram, the up-converted decoder tensor 503 is concatenated (merged) as the up-sampled tensor 704 and the skip tensor 502 together into the concatenated tensor 705 by means of the first concatenation function 520 .
[0099] A series of functions 620 of the artificial neural network are applied to the concatenated tensor 705 in order to obtain a review tensor 706. The applied functions 620 of the artificial neural network should concatenate the coarse features and fine features represented by the corresponding tensors with each other and should be matched to the feature maps of the lower layers. For example, convolutions - also the configuration of convolution blocks - i.e. - multiple applications of convolutions, depthwise convolutions, compression, residual, density, inception, activation, normalization, pooling or the like.
[0100] Subsequently, the review tensor 706 is concatenated (merged) with the upsampled tensor 704 by means of a second concatenation function 720 to form a result tensor 715 of the layer shown.
[0101] By rejoining (merging) 720 the review tensor 706 with the upsampled tensor 704 by means of a stitching function 720 , it is possible to correct the positioning of features within a certain level. Features identified in the image data are thereby improved in that they become more precise.
[0102] By means of the concatenation function (merge) 720 , not only the verification tensor 706 can be concatenated with the upconverted decoder tensor 704 . It is also conceivable to additionally concatenate a further tensor 707 with the help of the concatenation function (merge) 720 to form the result tensor 715 .
[0103] Different functions are applied to the sampling tensor 704 , the concatenated tensor 705 and the review tensor 706 to form the so-called correction module (review module) 700 .
[0104] Figure 8 A diagram showing the application of the present invention in an artificial neural network according to a fully convolutional network architecture.
[0105] In this case, the application of the present invention enables enhanced knowledge transfer between layers, wherein in particular the “vanishing gradient” effect can be prevented at the discriminative layer.
[0106] The application of the present invention to an artificial neural network based on a fully convolutional network architecture can be achieved in the following manner: in the decoder module 120, the individual skip tensors 802 and the upsampling tensor 304 are concatenated (merged) together into a review tensor 806 by means of a concatenation function.
[0107] Subsequently, the review tensor 806 is again concatenated (merged) with the upsampled tensor 304 into the decoder tensor 815 by means of a concatenation function.
[0108] As a result of the last layer of the decoder module 120 of the artificial neural network according to the fully convolutional network architecture, there is a result tensor 320 with an optimized semantic segmentation and a resolution up to the original resolution of the processed image data.
[0109] Fig. 9 A flow chart showing one specific embodiment of a method 900 of the present invention is shown.
[0110] In step 910, with the help of a first splicing function, the input tensors 304, 504, 704 and the jump tensors 502, 802 are spliced (merged) to obtain merged tensors 605, 705, wherein the input tensors 304, 504, 704 and the jump tensors 502, 802 are respectively related to the image data 111, 211.
[0111] In step 920 , a function of a neural network, in particular a convolution, is applied to the merged tensors 605 , 705 in order to obtain the review tensors 706 , 806 .
[0112] In step 930 , with the help of a second concatenation function, a second concatenation (merging) of the review tensors 706 , 806 and the input tensors 304 , 504 , 704 is performed to obtain output tensors 715 , 815 .
[0113] In step 904, the output tensor 715, 815 is output to the decoder path 120, 220 of the artificial neural network.
[0114] The present invention is preferably suitable for use in automotive systems, in particular in combination with driver assistance systems, up to partially automated driving or fully automated driving.
[0115] Of particular interest here is the processing of image data or image streams representing the vehicle's surroundings.
[0116] These image data or image streams can be detected by an imaging sensor of the vehicle. In this case, the detection is performed with the aid of a single sensor. It is also conceivable to fuse the image data of multiple sensors, if necessary with different detection sensors, such as video sensors, radar sensors, ultrasonic sensors, lidar sensors.
[0117] In this context, the detection of free space (free space detection) and the semantic distinction between foreground and background in the image data or image stream are of particular importance.
[0118] These features can be obtained by processing the image data or image stream using an artificial neural network according to the invention. Based on this information, the control system for the longitudinal or lateral control of the vehicle can be controlled accordingly, so that the vehicle reacts appropriately to the detection of these features in the image data.
[0119] Another area of application of the invention can be seen as the precise pre-labeling of image data or image data streams for camera-based vehicle control systems.
[0120] In this case, the label to be assigned indicates the class of the object to be detected in the image data or the image stream.
[0121] The invention can also be applied in all fields where accurate pixel-wise prediction with the aid of artificial neural networks is required (e.g. automotive, robotics, health, surveillance, etc.). Here, for example: optical flow, depth of monochrome image data, digital, edge recognition, key cards, object detection, etc. can be mentioned.
Claims
1. A method (900) for performing computationally efficient semantic segmentation of image data (111, 211) of an imaging sensor, the method being performed with the aid of an artificial neural network, wherein: The artificial neural network has an encoder path (110, 210) and a decoder path (120, 220), and the method comprises the following steps: Performing a first concatenation (910) of an input tensor (304, 704) and a jump tensor (502, 802) with the aid of a first concatenation function (520) to obtain a merged tensor (705), wherein the input tensor (304, 704) and the jump tensor (502, 802) are related to the image data (111, 211); applying (920) a neural network function (620) to the merged tensor (705) to obtain a review tensor (706, 806); By means of a second concatenation function (720), the review tensor (706, 806) and the input tensor (304, 704) are second concatenated (930) to obtain an output tensor (715, 815); The output tensor (715, 815) is output (940) to the decoder path (210, 220) of the artificial neural network.
2. The method (900) of claim 1, wherein: The input tensor (304, 704) has a feature map (302, 402), and in the step of applying (920), the function (620) of the neural network is associated with the feature map (302, 402).
3. The method (900) of claim 1, wherein: The first concatenation function (520) and / or the second concatenation function (720) are configured in such a way that the dimensionality of the input tensor (304, 704) is preserved.
4. The method (900) according to any one of claims 1 to 3, wherein: The steps of the method are performed in the decoder path (120, 22) of the artificial neural network.
5. The method (900) of claim 1, wherein: The artificial neural network is a convolutional neural network.
6. The method (900) of claim 5, wherein: A convolution is applied (920) to the merged tensor (705) to obtain a review tensor (706, 806).
7. A device for semantically segmenting image data (111, 211) of an imaging sensor, comprising an artificial neural network having an encoder path (110, 210) for classifying the image data (111, 211) and a decoder path (120, 220) for locating the image data (111, 211), and wherein the artificial neural network is configured to implement the steps of the method (900) according to any one of claims 1 to 6.
8. The device for semantically segmenting image data (111, 211) of an imaging sensor according to claim 7, wherein: The artificial neural network is a convolutional neural network.
9. A computer program product comprising instructions which, when executed by a computer, cause the computer to implement the method (900) according to any one of claims 1 to 6.
10. A machine-readable storage medium having instructions stored thereon, which, when executed by a computer, cause the computer to implement the method (900) according to any one of claims 1 to 6.
Citation Information
Patent Citations
RGBD image semantic segmentation method
CN107403430A
Image semantic segmentation method based on pyramid pooled coding-decoding structure
CN107644426A