Method for training a neural convolutional network, method for determining a positioning pose, device and storage medium
By training a neural convolutional network and combining aerial and ground images, the problem of insufficient uniqueness of localization within a large geographical area is solved, realizing an economical and efficient localization method with good scalability and constant query time.
Patent Information
- Application Number
- CN202011154501.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-24
- Filing Date
- 2020-10-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2040-10-26
AI Technical Summary
Existing positioning methods based on high-resolution maps suffer from insufficient uniqueness of positioning within large geographical areas and are also costly.
By training a neural convolutional network and combining aerial and ground images of the mobile platform's surrounding environment, an end-to-end learning approach is adopted, utilizing a two-stage method of pre-training with aerial images and training with ground images to determine the localization pose.
It achieves accurate positioning within large geographical areas, reduces economic costs, does not require high-resolution maps, and has good scalability and constant query time.
Smart Images

Figure CN112712556B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method for training a neural convolutional network for determining a localization pose of a mobile platform by means of the neural convolutional network by means of ground images. The present invention further relates to a method for determining a localization pose of a mobile platform, a device, and a machine-readable storage medium. BACKGROUND
[0002] Precise localization is a prerequisite for the travel of at least partially automated platforms, such as autonomous vehicles.
[0003] In order to localize such mobile platforms by means of ground images of the surroundings of such mobile platforms, a number of different approaches have been made, which are usually based on features of the surroundings of the mobile platform, wherein these features are then assigned to a pose of the mobile platform by means of a high-resolution map. SUMMARY
[0004] However, the use of such high-resolution maps is accompanied by economic disadvantages. In contrast, deep learning-based methods for determining a pose by means of a regression based on ground images have the advantage of a determined size of the respective map; constant query time. By means of monocular images, video image sequences and depth images from a direct camera position, a localization can be determined by means of such methods. Here, a localization in very large geographical areas is a challenge in terms of the uniqueness of the determined pose.
[0005] The present invention discloses a method for training a neural convolutional network for determining a localization pose of a mobile platform by means of ground images, a method for determining a localization pose, a method for maneuvering a mobile platform, a computer program and a machine-readable storage medium according to the features of the present invention. Advantageous configurations are preferred embodiments of the present invention and the subject matter described below.
[0006] The present invention is based on the knowledge that by means of aerial images centered on an estimated position of the mobile platform, the spatial context and the perspective of the surroundings of the mobile platform can be used to train a neural network in combination with ground images in order to determine a pose of the mobile platform. This enables, inter alia, the correct assignment of non-unique features from ground images over larger geographical areas.
[0007] According to one aspect, a method for training a neural convolutional network for determining a localization pose of a mobile platform by means of the neural convolutional network by means of ground images is proposed.
[0008] Here, the method has a first plurality of aerial image training cycles, wherein each aerial image training cycle has the following steps:
[0009] In one step of the aerial image training cycle, a reference pose of the mobile platform is provided. In another step, an aerial image of the surrounding environment of the mobile platform in the reference pose is provided. In another step, the aerial image is used as an input signal for the neural convolutional network. In another step, a corresponding localization pose is determined by means of an output signal of the neural convolutional network. In another step, the neural convolutional network is adapted in order to minimize a deviation of the corresponding localization pose determined by means of the corresponding aerial image from the corresponding reference pose.
[0010] In a further step, the method trains the neural convolutional network trained by means of the first plurality of aerial image training cycles by means of a second plurality of ground image training cycles, wherein each ground image training cycle has the following steps:
[0011] In one step, a reference pose of the mobile platform is provided. In another step, a ground image of the surrounding environment of the mobile platform in the reference pose is provided. In another step, the ground image is used as an input signal for the neural convolutional network trained by means of the first plurality of aerial image training cycles. In another step, a localization pose is determined by means of an output signal of the neural convolutional network. In another step, the neural convolutional network is adapted in order to minimize a deviation of the corresponding localization pose determined by means of the corresponding ground image from the corresponding reference pose in order to provide a trained neural convolutional network for determining localization poses by means of ground images.
[0012] For the method, an untrained convolutional neural network can be provided for the first aerial image training cycle, as described below.
[0013] Advantageously, in the method, for each aerial image training cycle of the first plurality of aerial image training cycles, a different reference pose of a different surrounding environment of the mobile platform and a corresponding different aerial image are provided.
[0014] Advantageously, by means of the method, localization poses can be determined by means of visual ground images and visual aerial images of the surrounding environment of the mobile platform without the need to use high-resolution maps. Thus, aerial images are used for pre-training a neural convolutional network for determining localization poses of the mobile platform. Since the method is not based on artificially completed features, it can be well scaled with respect to larger geographical areas.
[0015] Here, using a ground image or an aerial image as an input signal for a neural network means that the ground image or the aerial image is passed to an input layer of the neural network.
[0016] Here, the ground image is typically generated by a front camera of the mobile platform by means of a digital camera system at a corresponding viewing angle.
[0017] In this method, the neural convolutional network is provided with ground images (for example, RGB images from a mobile platform's front-facing camera) and aerial images (for example, satellite images).
[0018] By pre-training the neural convolutional network by means of a first plurality of aerial image training cycles, blurring of the following ground images is eliminated when training by means of a second plurality of ground image training cycles: these ground images look very similar, but are spatially far apart. Thus, the convolutional network is first trained by means of a first plurality of aerial image training cycles, and then successively trained by means of a second plurality of ground image training cycles. Here, the aerial images of the first plurality of aerial image training cycles provided can correspond to the ground images of the second plurality of ground image training cycles in the sense that the geographical information contained in the entirety of the aerial images or ground images mutually complement and / or jointly act for improving the determination of the pose of the mobile platform. This joint action and / or complementation can in particular relate to aerial images and ground images representing similar geographical regions. However, a discriminative effect can also be achieved by different geographical regions of the aerial images and ground images.
[0019] By considering (by means of the pre-training) aerial images of the surroundings of the mobile platform, by means of the significant spatial arrangement of the aerial image features, the neural convolutional network is trained to learn discriminative features and, in addition, to be able to determine the localization pose more precisely.
[0020] In order to be able to determine the vehicle position or vehicle localization pose with high precision, the similarity of the ground images and the aerial images (for example, at least local parts of satellite images) is not compared, but the pose of the mobile platform is derived from the ground images provided in combination with the corresponding local aerial images or local satellite images.
[0021] Thus, an end-to-end learning based on ground images and aerial images is carried out in order to achieve good scalability. Thus, the position precursor The advantages in terms of good scalability are combined with the advantages of applying a neural convolutional network.
[0022] The neural convolutional network essentially has alternately repeated filter (Convolutional Layer: convolutional layer) and aggregation layers (Pooling Layer: pooling layer) and can contain one or more layers of "normal" fully connected neurons (Dense / Fully Connected Layer: dense layer / full connection layer) at the end of the network.
[0023] Here, the first or second trained neural encoder convolutional network part can be configured as part of the neural convolutional network, or these network parts can each be implemented in the form of a single neural convolutional network.
[0024] Both ground images and aerial images can exist as digital images of different angles on the surroundings of the mobile platform and can be generated, for example, by means of a digital camera system. The perspective of an aerial image on the surroundings of the mobile platform is a top down view. Such an aerial image can be generated, for example, by a camera system of a satellite, an airplane or a drone. Here, such an aerial image can not only be a single generated aerial image of the surroundings of the mobile platform, but can also be a partial image, for example, from a larger aerial image, wherein the partial image is centered, in particular, on the estimated pose of the mobile platform. Such an aerial image can be, in particular, a satellite image tile that can be called up for a determined satellite navigation position, for example, a GPS position.
[0025] The localization pose of the mobile platform is a pose that is defined by a position in three spatial dimensions and an orientation of the mobile platform in space, for example, which can be described by three Euler angles, by means of which the pose is determined.
[0026] The reference pose of the mobile platform is a pose that provides a very accurate description of the training for determining the localization pose of the method, for example, by means of a reference system pose for determining the pose of the mobile platform.
[0027] Feedforward neural networks provide a framework for many different algorithms for machine learning, for cooperation and for processing complex data inputs. Such neural networks learn a task based on examples, often without being programmed with specific rules for the task.
[0028] Such a neural network is based on a collection of connected units or nodes, called artificial neurons. Each connection is capable of transmitting a signal from an artificial neuron to another. An artificial neuron that receives a signal can process it and then activate other artificial neurons connected to it.
[0029] In a conventional implementation of a neural network, the signals at the synapses of an artificial neuron are real numbers, and the output of the artificial neuron is computed by a non-linear function of the sum of its inputs. The synapses of an artificial neuron typically have weights matched to the strength of their incoming connections. The weights are adjusted during learning processes. Artificial neurons can have a threshold such that the signal is only output if the total input exceeds the threshold.
[0030] Typically, multiple artificial neurons are combined in layers. Different layers can perform different kinds of transformation on their inputs. It is possible that the signals propagate from the first layer (input layer) to the last layer (output layer) after multiple passes through the layers.
[0031] As a supplement to the above implementation of the feedforward neural network, the structure of an artificial neural convolutional network (Convolutional Neural Network: CNN) also comprises one or more convolutional layers (convolutional layers), if necessary followed by pooling layers. The order of the layers can be used with or without a normalization layer (e.g. batch normalization), a zero padding layer, an exit layer and an activation function (e.g. rectified linear unit (ReLU), sigmoid function, tanh function or softmax function). In principle, these units can be repeated any number of times, and in the case of sufficient repetition, we speak of a deep convolutional neural network.
[0032] In order to train the thus defined structure of the neural encoder / decoder convolutional network, each neuron obtains initial weights, for example randomly. The input data is then given to the network, and each neuron weights the input signal with its weights and continues to give the result to the neurons of the next layer. The result is then provided in the output layer. The size of the error can be calculated, as well as the proportion of each neuron in this error, and then the weights of each neuron are changed in the direction of minimizing the error. The traversal is then carried out recursively, the error is re-measured, the weights are matched, until the error is below a predefined limit.
[0033] The order of the method steps is shown in the overall description of the application in such a way that the method is easy to understand. However, the person skilled in the art will recognize that many of the method steps can also be traversed in a different order and lead to the same result. In this sense, the order of the method steps can be changed accordingly and is therefore also disclosed.
[0034] A mobile platform can be understood as a driver assistance system of a mobile system and / or a vehicle that is at least partially automated. One example can be an at least partially automated vehicle or a vehicle with a driver assistance system. This means that in this context the at least partially automated system contains the mobile platform in terms of at least partially automated functions, but the mobile platform also contains vehicles and other mobile machines that include a driver assistance system. Other examples of mobile platforms can be a driver assistance system with multiple sensors, a mobile multi-sensor robot (e.g. a robotic vacuum cleaner or a lawnmower), a multi-sensor monitoring system, a manufacturing machine, a personal assistant, a shuttle, a robot, a ship, an aircraft, a commercial vehicle or an access control system. Each of these systems can be a fully or partially automated system.
[0035] It is proposed according to one aspect that the first plurality of aerial images training period is determined by the deviation of the respective determined positioning pose from the respective reference pose being smaller than a predefined first value.
[0036] Thus, the accuracy sought for the determination of the localization poses can be determined in a first part of the method having the first plurality of aerial image training cycles, and / or termination criteria for the first plurality of aerial image training cycles can be defined.
[0037] According to one aspect it is proposed to determine the second plurality of ground image training cycles by the deviation of the respective determined localization pose from the respective reference pose being smaller than a predefined second value.
[0038] Thus, the accuracy sought for the determination of the localization poses can be determined in a second part of the method having the second plurality of ground image training cycles, and / or termination criteria for the second plurality of ground image training cycles can be defined.
[0039] According to one aspect it is proposed that the neural convolutional network to be trained is a neural encoder convolutional network or encoder network.
[0040] According to one aspect it is proposed to generate the aerial images for the method of training and for the method of determining a localization pose of a surrounding environment of a mobile platform by means of a satellite, an airplane or a drone.
[0041] According to one aspect it is proposed to select the aerial images by means of a pose determined by means of a navigation system supported by a global navigation system and / or a mobile radio.
[0042] By means of the position predefinition of the navigation system, the search space for the features can be reduced and the pose determination by means of the ground images can be estimated more precisely by means of a reduced amount of data.
[0043] According to one aspect it is proposed that, when adapting the neural convolutional network, the weights of the neural convolutional network are changed in at least some training cycles in order to minimize the deviation of the respective localization pose from the respective reference pose.
[0044] According to one aspect it is proposed that, when adapting the neural convolutional network trained by means of the first plurality of aerial image training cycles, the weights of the neural convolutional network trained by means of the first plurality of aerial image training cycles are changed in at least some training cycles in order to minimize the deviation of the respective localization pose from the respective reference pose.
[0045] A method for determining a localization pose of a mobile platform is proposed, wherein the mobile platform is provided for generating a ground image of the surrounding environment of the mobile platform. In the method, a ground image of the surrounding environment of the mobile platform is provided in one step. In another step, a localization pose of the mobile platform is generated by means of a neural convolutional network which has been trained in succession by means of aerial images and corresponding ground images of the respective surrounding environment of the mobile platform and the provided ground image as an input signal for the neural convolutional network which has been trained in succession.
[0046] The method is based on a neural convolutional network which has been trained in succession by means of aerial images and corresponding ground images of the respective surrounding environment of the mobile platform. Thereby, in a two-stage training of the neural network with a successively connected aerial image training cycle and a ground image training cycle, features from a larger spatial context, for example from aerial images, can advantageously be incorporated into the training of the neural convolutional network in order to achieve a higher accuracy of the determination of the localization pose of the mobile platform.
[0047] The method for determining a localization pose of a mobile platform can be combined with different existing methods for improving the pose determination. This is in particular for example the integration of sequential information and the consideration of geometric constraints, which can lead to further performance improvements.
[0048] A main advantage of the method is the scalability of the application of the method, since both contextual information and large-area localization information are incorporated into the method.
[0049] Furthermore, a constant query time for the pose determination is derived by means of the method, which is not the case in conventional feature-based methods. For example, in 3D-3D / 2D-3D feature matching, a good scaling cannot be achieved in the case of large map sizes.
[0050] In the method, a fixed "map size" is derived, since the map is implicitly represented by the weights of the provided and stored network.
[0051] Additionally, publicly accessible information is used for the first estimated pose by means of the method, and for example satellite images can be used for the aerial images, which are economically advantageous and do not require a manual labeling.
[0052] It is proposed according to one aspect to train a neural convolutional network which has been trained in succession by means of aerial images and corresponding ground images of the respective surrounding environment of the mobile platform according to one of the above-described methods for training a neural convolutional network.
[0053] It is proposed according to one aspect that the digital ground image is provided by the mobile platform.
[0054] It is proposed according to an aspect that the output signal is generated by a neural convolutional network when determining the pose of the mobile platform and that the output signal has a value for determining the localization pose.
[0055] It is proposed according to an aspect that the neural convolutional network has a fully connected network layer.
[0056] Herein, in a fully connected layer, the neurons of one layer are connected with all neurons of the next layer and are therefore called "fully connected layer" (also called "dense layer"). The neural layer then has a high weight like the number of connections.
[0057] It is proposed according to an aspect that the neural convolutional network is a neural encoder convolutional network.
[0058] It is proposed according to an aspect that the ground image of the surrounding environment of the mobile platform is a digital ground image.
[0059] It is proposed according to an aspect that the ground image of the surrounding environment of the mobile platform is generated by means of a digital camera system.
[0060] The use of a digital camera system has the advantage that the digital image generated here can be further processed simply.
[0061] It is proposed according to an aspect that the ground image of the surrounding environment of the mobile platform is generated from the perspective of the mobile platform by means of a front camera of the mobile platform.
[0062] It is proposed according to an aspect that a control signal for maneuvering the at least partially automated mobile platform is provided on the basis of the localization pose; and / or a warning signal for warning an occupant of the at least partially automated mobile platform is provided on the basis of the localization pose.
[0063] With regard to the feature "providing a control signal on the basis of the localization pose", the term "on the basis of" is to be understood broadly. It is to be understood that the localization pose is considered for each determination or calculation of the control signal, wherein this does not exclude that further input variables are also considered for the determination of the control signal. Similarly, this applies likewise to the provision of a warning signal.
[0064] An apparatus is specified which is set up for carrying out one of the above-mentioned methods. By means of this apparatus, the method can be integrated simply into different systems.
[0065] A computer program is specified which comprises instructions which, when the program is implemented by a computer, cause the computer to implement one of the above-mentioned methods. Such a computer program makes it possible to use the described method in different systems.
[0066] A machine-readable storage medium is specified on which the above-mentioned computer program is stored. BRIEF DESCRIPTION OF DRAWINGS
[0067] REFERENCE Figure 1 AND Figure 2 Embodiments of the present application are illustrated and hereinafter further described. The accompanying drawings illustrate:
[0068] Figure 1 a flow chart illustrating a method for training a neural convolutional network for determining a localization pose;
[0069] Figure 2 a flow chart illustrating a method for determining a localization pose of an at least partially automated mobile platform. DETAILED DESCRIPTION
[0070] Figure 1 A method 100 for training a neural convolutional network 110 for determining a localization pose 150 of a mobile platform by means of the neural convolutional network 110 by means of a ground image 140 is schematically depicted in a data flow diagram.
[0071] Herein, the method 100 has a first plurality of aerial image training cycles, wherein each aerial image training cycle has the following steps:
[0072] In a step S1 of the aerial image training cycle, a reference pose 120 of the mobile platform is provided. In a further step S2, an aerial image 130 of the surroundings of the mobile platform in the reference pose 120 is provided. In a further step S3, the aerial image 130 is used as an input signal for the neural convolutional network 110. In a further step S4, a corresponding localization pose 150 is determined by means of an output signal of the neural convolutional network 110. In a further step S5, the neural convolutional network 110 is adapted in order to minimize a deviation of the corresponding localization pose 150 determined by means of the corresponding aerial image 130 from the corresponding reference pose 120.
[0073] In a further step, the method 100 trains the neural convolutional network 110 trained by means of the first plurality of aerial image training cycles by means of a second plurality of ground image training cycles, wherein each ground image training cycle has the following steps:
[0074] In a step S6, a reference pose 120 of the mobile platform is provided. In a further step S7, a ground image 140 of the surrounding environment of the mobile platform in the reference pose 120 is provided. In a further step S8, the ground image 140 is used as an input signal for the neural convolutional network 110. In a further step S9, a localization pose 150 is determined by means of an output signal of the neural convolutional network 110. In a further step S10, the neural convolutional network 110 trained by means of the first plurality of training cycles of aerial images is adapted in order to minimize a deviation of a respective localization pose 150 determined by means of a respective ground image 140 from a respective reference pose 120 in order to provide a trained convolutional neural network 110 for determining a localization pose 150 by means of a ground image 130. Here, the neural convolutional network 110 can have a first number of convolutional layers 112 and a second number of fully connected layers 114. Therein, the second number of fully connected layers 114 can be connected on the first number of convolutional layers 112 in a layer sequence of the neural convolutional network 110.
[0075] Figure 2 The method 200 for determining a localization pose 150 of a mobile platform is schematically depicted in a data flow diagram, wherein the mobile platform is provided for generating a ground image 140 of a surrounding environment of the mobile platform. In the method, a ground image 140 of a surrounding environment of the mobile platform is provided in a step S21. In a further step S22, a localization pose 150 of the mobile platform is generated by means of the provided ground image 140 as an input signal for a sequentially trained neural convolutional network 110 trained by means of a respective aerial image 130 and a respective ground image 140 of a respective surrounding environment of the mobile platform in sequence.
[0076] Here, the sequentially trained neural convolutional network 110 by means of a respective aerial image 130 and a respective ground image 140 of a respective surrounding environment of the mobile platform can be trained according to the method 100 described in Figure 1 In this way, the sequentially trained neural convolutional network 110 by means of a respective aerial image 130 and a respective ground image 140 of a respective surrounding environment of the mobile platform can be trained according to the method 100 described in
Claims
1. A method (100) for training a neural convolutional network (110) for determining a localization pose (150) of a mobile platform by means of the neural convolutional network (110) by means of ground images (140), the method having a first plurality of aerial image training cycles, wherein, Each aerial image training cycle has the following steps: providing a reference pose (120) of the mobile platform (S1); providing an aerial image (130) of the surrounding environment of the mobile platform in the reference pose (120) (S2); using the aerial image (130) as an input signal for the neural convolutional network (110) (S3); determining a corresponding localization pose (150) by means of an output signal of the neural convolutional network (110) (S4); adapting the neural convolutional network (110) in order to minimize a deviation of the corresponding localization pose (150) determined by means of the corresponding aerial image (130) from the corresponding reference pose (120) (S5); training the neural convolutional network (110) trained by means of the first plurality of aerial image training cycles by means of a second plurality of ground image training cycles, wherein each ground image training cycle has the following steps: providing a reference pose (120) of the mobile platform (S6); providing a ground image (140) of the surrounding environment of the mobile platform in the reference pose (120) (S7); using the ground image (140) as an input signal for the neural convolutional network (110) trained by means of the first plurality of aerial image training cycles (S8); determining the localization pose (150) by means of an output signal of the neural convolutional network (110) (S9); adapting the neural convolutional network (110) in order to minimize a deviation of the corresponding localization pose (150) determined by means of the corresponding ground image (140) from the corresponding reference pose (120) in order to provide a trained neural convolutional network (110) for determining localization poses (150) by means of ground images (140) (S10), wherein the localization pose of the mobile platform is a pose comprising a position definition having three spatial dimensions and an orientation of the mobile platform in space, wherein the aerial image (130) is selected by means of a pose determined by means of a navigation system supported by a global navigation system and / or a mobile radio, wherein the aerial image (130) is an aerial image centered on an estimated position of the mobile platform.
2. The method (100) of claim 1, wherein The first plurality of aerial image training cycles is determined in such a way that a deviation of the corresponding determined localization pose (150) from the corresponding reference pose (120) is smaller than a predetermined first value.
3. The method (100) according to claim 1 or 2, wherein The second plurality of ground image training cycles is determined in such a way that a deviation of the corresponding determined localization pose (150) from the corresponding reference pose (120) is smaller than a predetermined second value.
4. The method (100) according to any one of claims 1 to 3, wherein The aerial image (130) of the surrounding environment of the mobile platform is generated by means of a satellite, an airplane or a drone.
5. A method (200) for determining a positioning pose (150) of a mobile platform, wherein The mobile platform is provided for generating ground images (140) of the surrounding environment of the mobile platform, the method having the following steps: providing a ground image (140) of the surrounding environment of the mobile platform (S21); providing a ground image (140) of the surrounding environment of the mobile platform in the reference pose (120) (S7); generating a localization pose (150) of the mobile platform (S22) by means of a neural convolutional network (110) and a provided ground image (140), the neural convolutional network being successively trained by means of aerial images (130) and corresponding ground images (140) of respective surroundings of the mobile platform, the provided ground image (140) being an input signal for the successively trained neural convolutional network (110), wherein the neural convolutional network (110) successively trained by means of aerial images (130) and corresponding ground images (140) of respective surroundings of the mobile platform is trained according to any one of claims 1 to 4.
6. The method (200) of claim 5, wherein providing a digital ground image (140) by means of the mobile platform.
7. The method (200) according to any one of claims 5 to 6, wherein, generating an output signal by means of the neural convolutional network (110), the output signal having a value for determining the localization pose (150).
8. The method (200) according to any of the preceding claims, wherein the neural convolutional network (110) is a neural encoder convolutional network.
9. The method (100) (200) according to any of the preceding claims, wherein generating a ground image (140) of a surrounding of the mobile platform from a perspective of the mobile platform by means of a front-facing camera of the mobile platform.
10. The method (200) according to any one of claims 5 to 9, wherein providing a control signal for maneuvering the at least partially automated mobile platform based on the localization pose (150); and / or providing a warning signal for warning an occupant of the at least partially automated mobile platform based on the localization pose (150).
11. A device configured to perform the method according to any one of claims 1 to 10.
12. A computer program comprising instructions which, when the computer program is implemented by a computer, cause the computer to carry out the method according to any one of claims 1 to 10.
13. A machine-readable storage medium having stored thereon the computer program according to claim 12.
Citation Information
Patent Citations
Aerial video saliency region detection method and apparatus
CN109543561A
Scene modeling method, system and device integrating aerial photography and ground visual angle images
CN110223380A