Positioning with point-to-line matching

Through the neural network matching the 3D feature points of the vehicle sensor image and satellite image, combined with the depth estimation technology, the problem of insufficient vehicle attitude resolution in the prior art is solved, the determination of high-definition attitude is achieved, and the operation accuracy of the vehicle on the road is improved.

CN120014022APending Publication Date: 2025-05-16FORD GLOBAL TECH LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411579269.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-11-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art uses satellite image-guided geolocation, and it is difficult to determine the three-degree-of-freedom attitude of the vehicle at a higher resolution, especially when operating the vehicle on the road, the resolution of the position and orientation is insufficient.

Method used

By inputting the images acquired by the vehicle sensor and satellite images into the neural network, the features are extracted and matched, the high-definition attitude of the vehicle is determined. This method does not require a predetermined high-definition map, and uses 3D feature point matching and depth estimation techniques to iteratively adjust the pose until the threshold of the global loss function is reached.

Benefits of technology

The determination of vehicle attitude within the +/-1 meter position resolution and +/-1 degree orientation resolution is achieved, which is sufficient to operate the vehicle safely on the road and improve the navigation accuracy of the vehicle in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014022A_ABST
    Figure CN120014022A_ABST
Patent Text Reader

Abstract

The invention provides positioning with point-to-line matching. A computer includes a processor and a memory including instructions executable by the processor to: determine a top keypoint from one of an aerial feature map or one or more ground feature maps, and projecting the top keypoint as a corresponding line on the other of the aerial feature map or the one or more ground feature maps. The memory includes instructions to determine a depth estimate of the top keypoint on the corresponding line. A high definition estimated three degree of freedom pose of the ground view camera is determined in global coordinates by iteratively determining a geometric correspondence between the top keypoint and the corresponding line until a global loss function is less than a user determined threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to determining a position of a vehicle based on a system pose. Background Art

[0002] The computer can be used to operate a system, including a vehicle, a robot, a drone, and / or an object tracking system. Data, including images, can be acquired by sensors and processed using a computer to determine the position of the system relative to objects in the environment surrounding the system. The computer can use the position to determine a trajectory for moving the system in the environment. The computer can then determine control data to be transmitted to system components to control the system components to move the components according to the determined trajectory. Summary of the invention

[0003] Systems (including vehicles, robots, drones, etc.) can operate by acquiring sensor data about the environment surrounding the system and processing the sensor data to determine a path on which to operate the system or part of the system. Sensor data can be processed to determine the location of objects in the environment. Objects can include roads, buildings, conveyors, vehicles, manufactured parts, etc. Sensor data can be processed to determine the attitude of the system, where the "attitude" specifies the position and orientation of an object (such as a system and / or its components). The system attitude can be determined based on a full six-degree-of-freedom (DoF) attitude, which includes x, y, and z position coordinates, and roll, pitch, and yaw rotation coordinates relative to the x, y, and z axes, respectively. The six-DoF attitude can be determined relative to a global coordinate system such as a Cartesian coordinate system, where points can be specified based on latitude, longitude, and altitude or some other x, y, and z axes.

[0004] A vehicle is used herein as a non-limiting example of a system. A simpler three-DoF pose can be used to position the vehicle relative to the environment surrounding the vehicle, which assumes that the vehicle is supported on a planar surface such as a road that fixes the z, pitch, and roll coordinates of the vehicle to match the road. The vehicle pose can be described by x and y position coordinates and yaw rotation coordinates to provide a three-DoF pose that defines the position and orientation of the vehicle relative to the supporting surface.

[0005] Vehicle sensors can provide data that can be used to determine vehicle attitude and can then be used to locate the vehicle relative to an aerial image that includes position data in global coordinates. For example, vehicle sensors can provide data to determine position and / or attitude based on a satellite-based global positioning system (GPS) and / or an accelerometer-based inertial measurement unit (IMU). For example, the position data included in the aerial image can be used to determine the position of any pixel address position in the aerial image in global coordinates. Aerial images can be obtained by satellites, aircraft, drones or other aerial platforms. Satellite data will be used as a non-limiting example of aerial image data in this article. For example, satellite images can be obtained by downloading GOOGLE™ maps from the Internet, etc.

[0006] The use of global coordinate data included in or with satellite images to determine the attitude of an object (such as a vehicle) relative to satellite image data can generally provide attitude data within + / -3 meter position and + / -3 degree orientation resolution. Operating a vehicle can rely on attitude data, which includes a position resolution of one meter or less and an orientation resolution of one degree or less. For example, + / -3 meter position data may not be enough to determine the position of a vehicle relative to a lane on a road. The technology for geolocation guided by satellite images as discussed herein can determine the vehicle attitude within a specified resolution, which is generally within a position resolution of one meter or less and an orientation resolution of one degree or less, for example, a resolution sufficient to operate a vehicle on a road. The vehicle attitude data determined within a specified resolution (that is, exceeding one or more specified resolution thresholds, for example, a position resolution of one meter or less and an orientation resolution of one degree or less in an exemplary implementation) is referred to as high-definition attitude data in this article.

[0007] The technology described herein uses satellite image guided geolocation to enhance the determination of the high definition pose of the vehicle. Satellite image guided geolocation uses images acquired by sensors included in the vehicle to determine the high definition pose relative to the satellite image without the need for a predetermined high definition (HD) map. The vehicle sensor image and the satellite image are input to one or more neural networks that extract features and confidence and / or attention maps from the image. In some examples, one or more neural networks can be the same neural network. 3D feature points from the vehicle image are matched with 3D feature points from the satellite image to determine the high definition pose of the vehicle relative to the satellite image. The high definition pose of the vehicle can be used to operate the vehicle by determining the vehicle path based on the high definition pose.

[0008] A system is disclosed herein, the system comprising a computer, the computer comprising a processor and a memory. The memory comprises instructions executable by the processor to determine a top keypoint from one of an aerial feature map or one or more ground feature maps. The top keypoint is projected as a corresponding line on the other of the aerial feature map or one or more ground feature maps. The processor determines a depth estimate of the top keypoint on the corresponding line, and determines a high-definition estimated three-degree-of-freedom pose of a ground view camera in global coordinates by iteratively determining a geometric correspondence between the top keypoint and the corresponding line until a global loss function is less than a user-determined threshold.

[0009] The global loss function may be determined by summing: 1) a pose-aware branch loss function determined by computing a triplet loss between the top keypoint and the corresponding line and 2) a recursive pose refinement branch loss function determined by computing a residual between the top keypoint and the corresponding line using a Levenberg-Marquardt algorithm. The pose-aware branch loss function may determine feature residuals based on the determined high-definition estimated three-degree-of-freedom pose of the ground-truth three-degree-of-freedom pose of the ground-view camera.

[0010] The instructions may also include instructions for determining the one or more ground feature maps and the one or more ground attention maps from one or more ground perspective images using one or more neural networks, and determining the aerial feature map and the aerial attention map from the aerial perspective images using the one or more neural networks.

[0011] The instructions may include instructions for weighting the feature map with the attention map. The instructions for determining the top keypoint may include instructions for determining the top keypoint from the one or more ground feature maps. The aerial view image may be a satellite image.

[0012] The instructions for determining the high-definition estimated three-degree-of-freedom pose of the ground view camera may include instructions for determining the high-definition estimated three-degree-of-freedom pose based on an initial estimate of the three-degree-of-freedom pose of the ground view camera. The instructions may also include instructions for outputting the high-definition estimated three-degree-of-freedom pose of the ground view camera to operate a vehicle. The system may include a vehicle computer configured to determine a vehicle path on which to operate the vehicle based on the high-definition estimated three-degree-of-freedom pose of the ground view camera and the aerial view image.

[0013] A method is disclosed herein that includes determining a top keypoint from one of an aerial feature map or one or more ground feature maps. The top keypoint is projected as a corresponding line on the other of the aerial feature map or one or more ground feature maps. The method includes determining a depth estimate of the top keypoint on the corresponding line, and determining a high-resolution estimated three-degree-of-freedom pose of a ground view camera in global coordinates by iteratively determining a geometric correspondence between the top keypoint and the corresponding line until a global loss function is less than a user-determined threshold.

[0014] The global loss function may be determined by summing: 1) a pose-aware branch loss function determined by computing a triplet loss between the top keypoint and the corresponding line and 2) a recursive pose refinement branch loss function determined by computing a residual between the top keypoint and the corresponding line using a Levenberg-Marquardt algorithm. The pose-aware branch loss function may determine feature residuals based on the determined high-definition estimated three-degree-of-freedom pose of the ground-truth three-degree-of-freedom pose of the ground-view camera.

[0015] The method may include determining the one or more ground feature maps and the one or more ground attention maps from one or more ground perspective images using one or more neural networks, and determining the aerial feature map and the aerial attention map from the aerial perspective images using the one or more neural networks.

[0016] The method may include weighting the feature map with the attention map. Top keypoints may be determined from the ground feature map. One or more neural networks may have a U-Net architecture.

[0017] The high-definition estimated three-degree-of-freedom pose of the vehicle camera may be determined based on an initial estimate of the three-degree-of-freedom pose of the ground-view camera. The method may include outputting the high-definition estimated three-degree-of-freedom pose of the ground-view camera to operate a vehicle. The method may include determining a vehicle path on which to operate the vehicle based on the high-definition estimated three-degree-of-freedom pose of the ground-view camera and the aerial-view image. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a block diagram of an exemplary vehicle sensing system.

[0019] Figure 2 is a diagram of an exemplary satellite image including a vehicle.

[0020] Figure 3 is a diagram of another exemplary satellite image including a vehicle.

[0021] Figure 4is a diagram of an exemplary system for determining a high-resolution three-degree-of-freedom vehicle position in global coordinates.

[0022] Figure 5A is a diagram of an exemplary image of a traffic scene including features.

[0023] Figure 5B is a diagram of an exemplary satellite image including line features.

[0024] Fig. 6A is a diagram of an exemplary satellite image including features.

[0025] Figure 6B is a diagram of an exemplary image of a traffic scene including line features.

[0026] Figure 7 is a flow chart of an exemplary process for determining a high-resolution three-degree-of-freedom vehicle position in global coordinates.

[0027] Figure 8 is a flow chart of an exemplary process for operating a vehicle based on a high-resolution vehicle position in global coordinates. DETAILED DESCRIPTION

[0028] Figure 1 is a diagram of a sensing system 100. The sensing system 100 includes a vehicle 110 that can be operated by a user and / or under the control of a computing device 115, which can include one or more vehicle electronic control units (ECUs) or computers such as are known, which ECUs or computers may include additional hardware, software and / or programming as described herein. The computing device 115 can receive data about the operation of the vehicle 110 from sensors 116. The computing device 115 can operate the vehicle 110 or components thereof in lieu of or in combination with the control of a human user. The system 100 can also include a server computer 120 that can communicate with the vehicle 110 via a network 130.

[0029] The computing device 115 may include one or more processors and one or more memory devices such as are known. In addition, the memory includes one or more forms of computer-readable media and stores instructions that can be executed by the processor to perform various operations (including operations as disclosed herein). For example, the computing device 115 may include programming to operate one or more of vehicle braking, propulsion (e.g., controlling the acceleration of the vehicle 110 by controlling one or more of an internal combustion engine, an electric motor, a hybrid engine, etc.), steering, climate control, interior lights and / or exterior lights, etc., and determine whether and when the computing device 115 (rather than a human operator) controls such operations.

[0030] The computing device 115 may include more than one computing device, for example, a controller, ECU, etc. included in the vehicle 110 for monitoring and / or controlling various vehicle subsystems (e.g., propulsion subsystem 112, braking subsystem 113, steering subsystem 114, etc.), or may be communicatively coupled to the more than one computing device, for example, via a vehicle communication bus as further described below. The computing device 115 is typically arranged to communicate on a vehicle communication network (e.g., including a bus in the vehicle 110, such as a controller area network (CAN), etc.); the vehicle network may additionally or alternatively include, for example, known wired or wireless communication mechanisms, such as Ethernet or other communication protocols.

[0031] The computing device 115 may transmit messages to and / or receive messages from various devices in the vehicle (e.g., controllers, actuators, sensors (including sensor 116), etc.) via the vehicle network. Alternatively or additionally, where the computing device 115 actually includes multiple devices, the vehicle communication network may be used for communication between devices represented as computing device 115 in this disclosure. In addition, as mentioned below, various controllers or sensing elements (such as sensor 116) may provide data to the computing device 115 via the vehicle communication network.

[0032] Additionally, computing device 115 may be configured to communicate with remote server computer 120 (e.g., a cloud server) via network 130 through vehicle-to-infrastructure (V2X) interface 111, as described below, which includes hardware, firmware, and software that permit computing device 115 to communicate with remote server computer 120 (e.g., a cloud server) via a vehicle-to-infrastructure (V2X) interface 111, such as a wireless Internet connection. The V2X interface 111 may thus include a computer configured to utilize various wired and / or wireless networking technologies (e.g., cellular, The computing device 115 may be configured to communicate with other vehicles 110 through a V2X (vehicle to the outside world) interface 111 using, for example, a vehicle-to-vehicle (V2V) network formed between nearby vehicles 110 on the basis of a mobile ad hoc network or through an infrastructure-based network (e.g., based on cellular communication (C-V2X) wireless communication cellular, dedicated short-range communication (DSRC) and / or similar communications). The computing device 115 also includes a non-volatile memory such as known. The computing device 115 can record data by storing the data in the non-volatile memory for later retrieval and transmission to the server computer 120 or the user mobile device 160 via the vehicle communication network and the vehicle to infrastructure (V2X) interface 111.

[0033] As already mentioned, the instructions stored in the memory and executable by the processor of the computing device 115 generally include programming for operating one or more vehicle 110 components (e.g., braking, steering, propulsion, etc.). Using data received in the computing device 115 (e.g., sensor data from sensors 116, server computer 120, etc.), the computing device 115 can make various determinations and / or control various vehicle 110 components and / or operations. For example, the computing device 115 can include programming to regulate or control vehicle 110 operating behavior (e.g., the physical manifestation of vehicle 110 operation), such as speed, acceleration, deceleration, steering, etc., as well as strategic behavior (e.g., generally controlling operating behavior in a manner intended to achieve efficient traversal of a route), such as the distance between vehicles and / or the amount of time between vehicles, lane changes, minimum gaps between vehicles, left turn crossing path minimums, arrival times at specific locations, and the minimum time from arrival at an intersection to crossing the intersection (without a signal light).

[0034] Each of the subsystems 112, 113, 114 may include a corresponding processor and memory and / or one or more actuators. The subsystems 112, 113, 114 may be programmed and connected to a vehicle 110 communication bus, such as a controller area network (CAN) bus or a local interconnect network (LIN) bus, to receive instructions from a computing device 115 and control the actuators based on the instructions.

[0035] Sensors 116 may include a variety of devices such as are known to provide data via a vehicle communication bus. For example, a radar fixed to a front bumper (not shown) of vehicle 110 may provide a distance from vehicle 110 to the next vehicle in front of vehicle 110, or a global positioning system (GPS) sensor disposed in vehicle 110 may provide geographic coordinates of vehicle 110. The distances provided by radar and / or other sensors 116 and / or the geographic coordinates provided by a GPS sensor may be used by computing device 115 to operate vehicle 110.

[0036] The vehicle 110 is typically a land-based vehicle 110 with three or more wheels, such as a passenger car, a light truck, etc. The vehicle 110 includes one or more sensors 116, a V2X interface 111, a computing device 115, and one or more subsystems 112, 113, 114. The sensor 116 can collect data related to the vehicle 110 and the operating environment of the vehicle 110. By way of example and not limitation, the sensor 116 can include, for example, an altimeter, a camera, a lidar, a radar, an ultrasonic sensor, an infrared sensor, a pressure sensor, an accelerometer, a gyroscope, a temperature sensor, a pressure sensor, a Hall sensor, an optical sensor, a voltage sensor, a current sensor, a mechanical sensor (such as a switch), etc. The sensor 116 can be used to sense the operating environment of the vehicle 110, for example, the sensor 116 can detect phenomena such as weather conditions (rainfall, external ambient temperature, etc.), road slope, road position (e.g., using road edges, lane markings, etc.) or the position of a target object (such as a neighboring vehicle 110). Sensors 116 may also be used to collect data, including dynamic vehicle data related to the operation of vehicle 110 , such as speed, yaw rate, steering angle, engine speed, brake pressure, oil pressure, power levels applied to subsystems 112 , 113 , 114 in vehicle 110 , connectivity between components, and accurate and timely performance of components of vehicle 110 .

[0037] The server computer 120 generally has features in common with the V2X interface 111 and computing device 115 of the vehicle 110, such as a computer processor and memory and a configuration for communicating via the network 130, and therefore these features will not be described further. The server computer 120 can be used to develop and train software that can be transferred to the computing device 115 in the vehicle 110.

[0038] Figure 2is a diagram of a satellite image 200. The satellite image 200 may be a map downloaded to a computing device 115 in a vehicle 110 via a network 130, for example, from a source such as GOOGLE maps. The satellite image 200 includes a road 202 indicated by a straight shape, a building 204, and foliage 206 indicated by an irregular shape. The version of the satellite image 200 used herein is a version of a photographic portrait including objects such as the road 202, the building 204, and the foliage 206. A vehicle, such as the vehicle 110, is included in the satellite image 200. The vehicle 110 includes a sensor 116, including a camera. The satellite image 200 includes four fields of view 208, 210, 212, 214 (e.g., spatial areas within which the corresponding cameras can capture images) for four cameras respectively included in the front, right, rear, and left sides of the vehicle 110.

[0039] Figure 3 is a diagram of satellite image 200 including estimated three DoF pose 302 of vehicle 110. For example, initial estimated three DoF pose 302 of vehicle 110 relative to satellite image 200 may be based on vehicle sensor data including a GPS sensor included in vehicle 110. Due to the limited resolution of the GPS sensor and the limited resolution of satellite image 200, estimated three DoF pose 302 of vehicle 110 generally does not represent a sufficiently accurate pose of vehicle 110. Due to the limited resolution of the GPS sensor and satellite image 200, estimated three DoF pose 302 is generally not used to operate vehicle 110.

[0040] One way to obtain high-definition data for operating the vehicle 110 may be to generate an HD map of all areas over which the vehicle 110 operates. High-definition maps typically require a significant amount of cartographic work and a significant amount of computer resources to generate and store the HD maps, and typically require a significant amount of network bandwidth to download the HD maps to the vehicle 110, not to mention the significant amount of computer memory typically required to store the maps in a computing device 115 included in the vehicle. The satellite image-guided geolocation techniques described herein use 3D feature points determined based on video images acquired by a camera included in the vehicle 110 to determine a high-definition estimated three-DoF pose of the vehicle 110 based on satellite images, without requiring the significant amount of computer processing, networking, and / or memory resources typically required to generate, transmit, and store HD maps.

[0041] The keypoints detected by the disclosed systems and methods can circumvent the use of a flat ground homography. A flat ground homography assumes that the world lies on a flat plane and maps all pixels from a given viewpoint onto that flat plane via homography projection. Thus, the techniques described herein remove the constraint that keypoints are restricted to the ground plane. This enhancement enables the disclosed systems and methods to determine useful poses in a wider range of scenes, generally resulting in better performance.

[0042] Figure 4 4 is a diagram of an exemplary system 400 for determining a high-resolution estimated three-DoF vehicle pose in global coordinates. For example, the system 400 may be implemented by operating software instructions on a computing device 115 included in the vehicle 110. The system 400 may be trained on a server computer 120 and downloaded or otherwise installed to the computing device 115 in the vehicle 110. Starting from a rough pose, the system 400 may estimate an accurate three-DoF pose 402 of the vehicle, including lateral shift, longitudinal shift, and yaw angle in a satellite image 404, using a ground-view image 406 taken at the same location. The system 400 includes a feature and confidence map extractor (FCE) 408 having one or more convolutional neural networks (CNNs) to extract a satellite feature map 410 and a ground-view feature map 412 from the satellite image 404 and the ground-view image 406, respectively. The CNN may be a U-Net structure to obtain feature maps with original resolution, which is conducive to accurate pose estimation.

[0043] The CNN for extracting satellite feature map 410 and ground view feature map 412 may include a convolution layer followed by a fully connected layer. The convolution layer extracts latent variables indicating the location of feature points by convolving an input image (such as images 404, 406) with a series of convolution kernels. The latent variables are input to a fully connected layer, which determines feature points by combining the latent variables using linear and nonlinear functions. The convolution kernels and the linear and nonlinear functions are programmed using weights determined by training the feature extractor.

[0044] The spatial attention maps (satellite attention map 414 and ground view attention map 416) are computed and used to weight the feature maps 410 and 412 to identify pixels with potential correspondence (i.e., co-visibility) between the two sets of images. The feature maps 410 and 412 are used to calculate the weight of the feature maps using the equation r[p] P =F s [p] P -F g [p] to calculate the point residual 426. The spatial attention map (satellite attention map 414 and ground view attention map 416) uses the equation W[p] P =A s [p] P*A g [p] is generated and used as point weights 428. Here, 'P' represents the vehicle pose and 'p' represents the keypoints (top K points) detected by the conventional keypoint detection (KPD) method 418. The attention map is used as the weight of the pixel residuals. During the training process, the attention to moving and temporal objects is reduced. Therefore, pixels with high values ​​in the attention map indicate potential common visibility between two views.

[0045] The KPD method 418 can be a conventional method such as ORB (Oriented FAST and Rotated BRIEF) or SIFT (Sorted Intolerant from Tolerant) and can be applied to identify the top keypoints, such as the top K points 420, from each query image to create a keypoint map 420. The top K points refer to the top K keypoints with the highest scores; their scores are provided by KPD 418. High scores indicate more unique points, such as corner points. ORB is built on the well-known FAST keypoint detector and BRIEF detector.

[0046] The top K points from each query confidence map are detected and then projected as lines to the satellite image 410 to create a projection map 422. Since the depth of the top K points in the ground view image is unknown (see Figure 5A ), so the projection on satellite images is depicted as lines (see Figure 5B ) and depends on the vehicle pose. The point-to-line depth estimation (DE) module 424 uses the transformer attention mechanism to accurately estimate the corresponding depth of the first K points. The depth is presented as the distance from the camera along the camera's z-axis (facing direction). Using this depth information, the system calculates the coordinates of the key points in the satellite image. Subsequently, the residual 426 and point weights 428 are calculated based on these sparse representations across two views.

[0047] The global loss function is determined by summing the pose awareness branch (PAB) loss function 430 and the recursive pose refinement branch (RPRB) loss function 434. The PAB 430 uses a triplet loss 432 to distinguish the residual between two views conditioned on the correct (ground truth) and incorrect (initial) poses. The triplet loss is a function in which a reference input (i.e., an anchor point) is compared with a matching input (i.e., positive) and a non-matching input (i.e., negative). The triplet loss function minimizes the distance from the anchor point to the positive and maximizes the distance from the anchor point to the negative. The PAB 430 can be determined by calculating the triplet loss between the first K points and the corresponding line. The PAB 430 is enabled only when the initial pose (incorrect pose) is far from the ground truth pose. The RPRB 434 is deployed to iteratively optimize the initial pose toward the ground truth pose using the Levenberg-Marquardt (LM) algorithm 436. The RPRB can be determined by calculating the residual between the top key point and the corresponding line using the LM algorithm 436. In addition to the triplet loss 432, the reprojection error 438 is also minimized when optimizing the vehicle pose. It should be noted that both the PAB 430 and RPRB 434 objective branches supervise feature extraction, but they have different focuses. The PAB 430 encourages correct pose estimation and penalizes incorrect estimates. The RPRB 434 forces the most correct predicted pose to be close to the ground truth.

[0048] The two values ​​outputted from PAB 430 and RPRB 434, respectively, are added to form a global loss function and compared to a predetermined threshold to determine whether the system has converged to a solution. If the global loss function is greater than the threshold, the process loops back to reduce the loss function at the next iteration. When the global loss function is less than the threshold, the system has converged to a high-definition estimated three-DoF pose 402.

[0049] Figure 5A 1 is a graph of four images 500, 502, 504, 506 acquired by a camera included in the vehicle 110 corresponding to different fields of view (similar to the fields of view 208, 210, 212, 214, respectively). For example, the images 500, 502, 504, 506 can be red, green and blue (RGB) color images acquired at a standard video resolution of about 2K x 1K pixels. The images 500, 502, 504, 506 have been processed to determine the top K features, such as point groups 508, 510, 512, 514, respectively. The images 500, 502, 504, 506 including the feature points 508, 510, 512, 514 are referred to as key point graphs. The key feature points 508, 510, 512, 514 are indicated by the hatched areas in the images 500, 502, 504, 506.

[0050] Figure 5B 506. The key feature points 508, 510, 512, 514 are projected onto the satellite reference image 520 as lines (i.e., line groups 522, 524) using the initial pose (depicted as the cross-hatched area 522) and the ground truth pose (depicted in the hatched area 524). Due to the unknown depth of the key points in the ground view images 500, 502, 504, 506, the projections 522 and 524 on the satellite image 520 are lines. In other words, it is unknown how far the key points are from the ground view camera in the ground view image. The projections 522 and 524 are lines extending horizontally from the camera position.

[0051] Although this paper is about detecting key points in the ground view and projecting them as lines on the satellite view (e.g. Figure 5A and Figure 5B However, key points in the satellite view can be detected and projected as lines on the ground view image (as shown in FIG. Fig. 6A and Figure 6B shown). Fig. 6A corresponds to the ground perspective images 604, 606, 608, 610 ( Figure 6B ). The satellite image 600 has been processed to determine the top K features, such as point groups 602. Figure 6B , the projection of the key point 602 can be visualized as a line on the ground perspective images 604, 606, 608, 610. These lines are depicted as a group of vertical shadow lines 612, 614, 616, 618. Due to the unknown depth (height) of the key point in the satellite image 600, the projections 612, 614, 616, 618 on the ground perspective images 604, 606, 608, 610 are lines. The ground perspective images 604, 606, 608, 610 can be acquired by a camera included in the vehicle 110 corresponding to different fields of view (similar to the fields of view 208, 210, 212, 214, respectively).

[0052] Figure 7 About Figure 1 6 is a flow chart of a process 700 for determining a high-definition estimated three-DoF pose based on satellite image-guided geolocation. The process 700 may be implemented in a computing device 115 included in the vehicle 110. The process 700 includes a plurality of blocks that may be executed in the order shown. Alternatively or in addition, the process 700 may include fewer blocks, or may include blocks that are executed in a different order.

[0053] The process 700 begins at block 702, where a computing device 115 in a vehicle 110 receives an image 406 from, for example, one or more cameras included in the vehicle 110. The one or more images 406 include image data regarding an environment surrounding the vehicle 110, and may include any portion of the environment surrounding the vehicle, including, for example, overlapping fields of view 208, 210, 212, 214.

[0054] At block 704, computing device 115 receives an aerial view image (e.g., satellite image 404). For example, satellite image 404 may be obtained by downloading satellite image 404 from the Internet via network 130. Satellite image 404 may also be retrieved from a memory included in computing device 115. Satellite image 404 includes position data in global coordinates, which may be used to determine the position of any point in satellite image 404 in global coordinates. Satellite image 404 may be selected to include estimated three DoF pose 302. Estimated three DoF pose 302 may be determined by acquiring data from vehicle sensors 116 (e.g., GPS).

[0055] At block 706, the computing device 115 inputs the received ground perspective image 406 to one or more trained neural networks, such as the FCE 408. The one or more neural networks may be trained on the server computer 120 and downloaded or otherwise installed to the computing device 115 in the vehicle 110. The one or more neural networks determine a ground feature map 412 and a ground attention map 416 corresponding to the received ground perspective image 406.

[0056] At block 708, the computing device 115 also inputs the received aerial perspective image (e.g., satellite image 404) to the one or more neural networks, e.g., FCE 408. The one or more neural networks determine an aerial feature map 410 and an aerial attention map 414 corresponding to the received aerial perspective image 404.

[0057] At block 710, the computing device 115 determines the top K feature points. The attention maps (satellite attention map 414 and ground view attention map 416) are used to weight the feature maps 410 and 412 to identify pixels with potential correspondence between the two sets of images. The KPD method 418 can be applied to identify the top K points 420 from each query image to create a keypoint map 420.

[0058] At block 712, the computing device 115 projects the first K points (eg, keypoint map 420) as lines onto the satellite image 410 to create a projection map 422. Because the depths of the first K points in the ground view image are unknown (see Figure 5A ), so the projection on satellite images is depicted as lines (see Figure 5B ) and depends on the vehicle pose. An initial iteration of block 712 may use the estimated three DoF pose 302 from the vehicle sensor 116 data. Subsequent iterations of process 700 enhance the estimated three DoF pose 302 by reducing the global loss function as described above.

[0059] At block 714, the computing device 115 estimates the corresponding depths of the first K points using the point-to-line DE module 424. The point-to-line DE module 424 may use a transformer attention mechanism to accurately estimate the corresponding depths. The depth is presented as the distance from the camera along the projection line (z-axis facing direction).

[0060] At block 716, the computing device 115 determines a high-resolution estimated three-degree-of-freedom pose 402 of a ground view camera (e.g., a camera on the vehicle 110) in global coordinates by iteratively determining geometric correspondences between the top keypoints and the correspondence lines and / or depths until a global loss function is less than a user-determined threshold. Geometric correspondence is the process of pairing data points in the projection map 422 and the satellite map 410 and iteratively reprojecting the entire projection map 422 to minimize the pairwise position error or difference for each pair of data points.

[0061] At box 718, the computing device 115 determines the PAB loss function 430 and the RPRB loss function 434. The two values ​​output from these functions are added to form a global loss function and compared with a predetermined threshold at box 720 to determine whether the process 700 has converged to a solution. If the global loss function is greater than the threshold, the process 700 loops back to box 712 to reduce the loss function at the next iteration. The key point map 420 is reprojected using the new estimated three DoF pose to form a new projection map 422 and a new geometric correspondence between the new projection map 422 and the satellite map 410 to determine a new global loss function. When the global loss function is less than the threshold, the process 700 stops iterating and outputs the current estimated three DoF pose as the high-definition estimated three DoF pose 402 at box 722.

[0062] At box 722, the computing device 115 outputs the high-definition estimated three-DoF pose 402 from box 718 for use in operating the vehicle 110, as described below with respect to Figure 8 The high-resolution estimated three-DoF pose 402 may be described by x and y position coordinates and yaw rotation coordinates to provide a three-DoF pose defining the vehicle's position and orientation. After block 722, process 700 ends.

[0063] Figure 8 About Figures 1 to 7A flow chart of a process 800 for operating a vehicle 110 based on a high-resolution estimated three-DoF pose determined by a satellite image-based geo-positioning system 400 is described. The process 800 may be implemented by a computing device 115 included in the vehicle 110. The process 800 includes a plurality of blocks that may be performed in the order shown. Alternatively or additionally, the process 800 may include fewer blocks, or may include blocks that are performed in a different order.

[0064] Process 800 begins at block 802, where computing device 115 in vehicle 110 acquires one or more images, such as images 500, 502, 504, 506, from one or more cameras included in vehicle 110, and acquires a satellite image, such as satellite image 520, by downloading via network 130 or retrieving from a memory included in computing device 115. An estimated three-DoF pose 302 of vehicle 110 is determined based on the data acquired by vehicle sensors 116.

[0065] At block 804, computing device 115 Figure 4 The depicted satellite image guided geo-positioning system 400 processes one or more images 500 , 502 , 504 , 506 and a satellite image 520 to enhance the estimated three DoF pose 302 into a high resolution estimated three DoF pose 402 .

[0066] At box 806, the computing device 115 determines a vehicle path for the vehicle 110 using the high-definition estimated three-DoF pose 402. The vehicle can operate on a road based on the vehicle path by determining commands that direct the vehicle's propulsion (e.g., powertrain), braking, and steering components to operate the vehicle so as to travel along the path. The vehicle path is generally a polynomial function on which a vehicle such as the vehicle 110 can operate. Sometimes referred to as a path polynomial, the polynomial function can specify the vehicle position (e.g., in terms of x, y, and z coordinates) and / or attitude (e.g., roll, pitch, and yaw) over time. That is, the path polynomial can be a polynomial function of third degree or less that describes the movement of the vehicle on the ground. The movement of the vehicle on the road is described by a multi-dimensional state vector, which includes the vehicle position, orientation, velocity, and acceleration. Specifically, the vehicle motion vector may include position in x, v, z, yaw, pitch, roll, yaw rate, pitch rate, roll rate, heading velocity, and heading acceleration, which may be determined, for example, by fitting a polynomial function to a continuous 2D position relative to the ground including the vehicle motion vector. In addition, for example, the path polynomial p(x) is a model that predicts the path as a line described by a polynomial equation. The path polynomial p(x) predicts a predetermined upcoming distance x (e.g., measured in meters) of the path by determining a lateral coordinate p:

[0067] p(x)=a0+a1x+a2x 2 +a3x 3 (1)

[0068] Where a0 is the offset, eg, the lateral distance between the path and the centerline of the vehicle 110 at the upcoming distance x, a1 is the heading angle of the path, a2 is the curvature of the path, and a3 is the rate of change of the curvature of the path.

[0069] The polynomial function may be used to guide the vehicle 110 from a current position indicated by the high-definition estimated three-DoF pose to another position in the environment surrounding the vehicle while maintaining minimum and maximum limits on lateral and longitudinal accelerations. The vehicle 110 may be operated along the vehicle path by transmitting commands to the subsystems 112, 113, 114 to control vehicle propulsion, steering, and braking. After block 806, the process 800 ends.

[0070] Computing devices such as those described herein typically each include commands that can be executed by one or more computing devices such as those identified above and used to implement the blocks or steps of the processes described above. For example, the process blocks described above can be embodied as computer executable commands.

[0071] The computer executable instructions may be compiled or interpreted by a computer program created using a variety of programming languages ​​and / or techniques, including but not limited to the following in single or combined form: Java TM , C, C++, Python, Julia, SCALA, Visual Basic, Java Script, Perl, HTML, etc. Typically, a processor (e.g., a microprocessor) receives commands, such as from a memory, a computer-readable medium, etc., and executes these commands, thereby performing one or more processes including one or more of the processes described herein. A variety of computer-readable media can be used to store such commands and other data in files and to transmit such commands and other data. Files in a computing device are typically collections of data stored on computer-readable media such as storage media, random access memory, etc.

[0072] Computer-readable media (also referred to as processor-readable media) include any non-transitory (i.e., tangible) media that participate in providing data (e.g., instructions) that can be read by a computer (e.g., by a processor of a computer). Such media can take many forms, including, but not limited to, non-volatile media and volatile media. Instructions can be transmitted via one or more transmission media, including optical fibers, wires, wireless communications, including internals that make up a system bus coupled to a processor of a computer. Common forms of computer-readable media include, for example, RAM, PROM, EPROM, FLASH-EEPROM, any other memory chip or cassette, or any other medium from which a computer can read.

[0073] Unless otherwise expressly indicated herein, all terms used in the claims are intended to be given their ordinary and customary meanings as understood by those skilled in the art. Specifically, unless a claim recites an express limitation to the contrary, the use of singular articles such as "a," "an," "the," and "said" should be interpreted as reciting one or more of the indicated elements.

[0074] The term "exemplary" is used herein in the sense of referring to an example; for example, reference to "exemplary widget" should be interpreted to refer merely to an example of a widget.

[0075] The adverb “approximately” modifying a value or result means that the shape, structure, measurement, value, determination, calculation, etc. may vary from the exactly described geometry, distance, measurement, value, determination, calculation, etc. due to imperfections in materials, machining, manufacturing, sensor measurement, calculation, processing time, communication time, etc.

[0076] In the accompanying drawings, the same candidate mark indicates the same element. In addition, some or all of these elements can be changed. With respect to the media, processes, systems, methods, etc. described herein, it should be understood that although the steps or frames of such processes, etc. have been described as occurring according to a sequence of a specific order, such processes can be practiced by the described steps performed in an order other than the order described herein. It should be understood that certain steps can be performed simultaneously, other steps can be added, or certain steps described herein can be omitted. In other words, the description of the process herein is provided for the purpose of illustrating certain embodiments and should never be interpreted as limiting the claimed invention. Any use of "based on" and "in response to" herein (including reference to the media, processes, systems, methods, etc. described herein) indicates a causal relationship, not just a temporal relationship.

[0077] According to the present invention, a system is provided, the system having: a computer, the computer including a processor and a memory, the memory including instructions, the instructions executable by the processor to: determine a top key point from one of an aerial feature map or one or more ground feature maps; project the top key point as a corresponding line on the other of the aerial feature map or the one or more ground feature maps; determine a depth estimate of the top key point on the corresponding line; and determine a high-definition estimated three-degree-of-freedom pose of a ground-view camera in global coordinates by iteratively determining a geometric correspondence between the top key point and the corresponding line until a global loss function is less than a user-determined threshold.

[0078] According to an embodiment, the global loss function is determined by summing the following items: 1) a posture-aware branch loss function determined by calculating the triplet loss between the top key point and the corresponding line and 2) a recursive posture refinement branch loss function determined by calculating the residual between the top key point and the corresponding line using the Levenberg-Marquardt algorithm.

[0079] According to an embodiment, the pose-aware branch loss function determines feature residuals based on the determined high-resolution estimated three-degree-of-freedom pose and the ground-truth three-degree-of-freedom pose of the ground-view camera.

[0080] According to an embodiment, the instructions also include instructions for performing the following operations: using one or more neural networks to determine the one or more ground feature maps and one or more ground attention maps from one or more ground perspective images, and using the one or more neural networks to determine the aerial feature map and aerial attention map from aerial perspective images.

[0081] According to an embodiment, the instructions further comprise instructions for weighting the feature map using the attention map.

[0082] According to an embodiment, the instructions for determining the top keypoint include instructions for determining the top keypoint from the one or more ground feature maps.

[0083] According to an embodiment, the instructions for determining the high-resolution estimated three-degree-of-freedom pose of the ground view camera include instructions for determining the high-resolution estimated three-degree-of-freedom pose based on an initial estimate of the three-degree-of-freedom pose of the ground view camera.

[0084] According to an embodiment, the aerial view image is a satellite image.

[0085] According to an embodiment, the instructions further include instructions for outputting the high-resolution estimated three-degree-of-freedom pose of the ground-view camera to operate a vehicle.

[0086] According to an embodiment, the invention also features a vehicle computer configured to determine a vehicle path on which to operate the vehicle based on the high-resolution estimated three-degree-of-freedom pose of the ground-view camera and the aerial-view image.

[0087] According to the present invention, a method includes: determining a top key point from one of an aerial feature map or one or more ground feature maps; projecting the top key point as a corresponding line on the other of the aerial feature map or the one or more ground feature maps; determining a depth estimate of the top key point on the corresponding line; and determining a high-definition estimated three-degree-of-freedom pose of a ground-view camera in global coordinates by iteratively determining a geometric correspondence between the top key point and the corresponding line until a global loss function is less than a user-determined threshold.

[0088] In one aspect of the invention, the global loss function is determined by summing: 1) a pose-aware branch loss function determined by calculating a triplet loss between the top keypoint and the corresponding line and 2) a recursive pose refinement branch loss function determined by calculating a residual between the top keypoint and the corresponding line using a Levenberg-Marquardt algorithm.

[0089] In one aspect of the invention, the pose-aware branch loss function determines feature residuals based on the determined high-resolution estimated three-degree-of-freedom pose and the ground-truth three-degree-of-freedom pose of the ground-view camera.

[0090] In one aspect of the present invention, the method includes using one or more neural networks to determine the one or more ground feature maps and one or more ground attention maps from one or more ground perspective images, and using the one or more neural networks to determine the aerial feature map and aerial attention map from aerial perspective images.

[0091] In one aspect of the invention, the method comprises weighting the feature map with the attention map.

[0092] In one aspect of the invention, top keypoints are determined from a ground feature map.

[0093] In one aspect of the invention, the determined high-resolution estimated three-degree-of-freedom pose of the vehicle camera is determined based on an initial estimate of the three-degree-of-freedom pose of the ground-view camera.

[0094] In one aspect of the invention, one or more neural networks have a U-Net architecture.

[0095] In one aspect of the invention, the method includes outputting the high-resolution estimated three-degree-of-freedom pose of the ground-view camera to operate a vehicle.

[0096] In one aspect of the present invention, the method includes determining a vehicle path on which to operate the vehicle based on the high-resolution estimated three-degree-of-freedom pose of the ground-view camera and the aerial-view image.

Claims

1. A system comprising: A computer comprising a processor and a memory, the memory comprising instructions executable by the processor to: Determining a top keypoint from one of the aerial feature map or one or more ground feature maps; projecting the top keypoint as a corresponding line on the other of the aerial feature map or the one or more ground feature maps; determining a depth estimate of the top keypoint on the corresponding line; as well as A high-resolution estimated three-degree-of-freedom pose of the ground view camera is determined in global coordinates by iteratively determining geometric correspondences between the top keypoints and the correspondence lines until a global loss function is less than a user-determined threshold.

2. The system of claim 1 , wherein the global loss function is determined by summing: 1) a pose-aware branch loss function determined by computing a triplet loss between the top keypoint and the corresponding line and 2) a recursive pose refinement branch loss function determined by computing a residual between the top keypoint and the corresponding line using a Levenberg-Mar quardt algorithm.

3. The system of claim 2, wherein the pose-aware branch loss function determines feature residuals based on the determined high-resolution estimated three-degree-of-freedom pose of the ground-truth three-degree-of-freedom pose of the ground-view camera.

4. The system of claim 1 , wherein the instructions further comprise instructions for determining the one or more ground feature maps and the one or more ground attention maps from one or more ground perspective images using one or more neural networks, and determining the aerial feature map and the aerial attention map from aerial perspective images using the one or more neural networks.

5. The system of claim 4, wherein the instructions further comprise instructions for weighting the feature map with the attention map.

6. The system of claim 1, wherein the instructions for determining the top keypoint include instructions for determining the top keypoint from the one or more ground feature maps.

7. The system of claim 1 , wherein the instructions for determining the high-definition estimated three-degree-of-freedom pose of the ground-view camera include instructions for determining the high-definition estimated three-degree-of-freedom pose based on an initial estimate of the three-degree-of-freedom pose of the ground-view camera.

8. The system of any one of claims 1 to 7, wherein the instructions further comprise instructions for outputting the high-resolution estimated three-degree-of-freedom pose of the ground-view camera to operate a vehicle.

9. The system of claim 8, further comprising a vehicle computer configured to determine a vehicle path on which to operate the vehicle based on the high-resolution estimated three-degree-of-freedom pose of the ground-view camera and the aerial-view image.

10. A method comprising: Determining a top keypoint from one of the aerial feature map or one or more ground feature maps; projecting the top keypoint as a corresponding line on the other of the aerial feature map or the one or more ground feature maps; determining a depth estimate of the top keypoint on the corresponding line; as well as A high-resolution estimated three-degree-of-freedom pose of the ground view camera is determined in global coordinates by iteratively determining geometric correspondences between the top keypoints and the correspondence lines until a global loss function is less than a user-determined threshold.

11. The method of claim 10, wherein the global loss function is determined by summing: 1) a pose-aware branch loss function determined by calculating a triplet loss between the top keypoint and the corresponding line and 2) a recursive pose refinement branch loss function determined by calculating a residual between the top keypoint and the corresponding line using a Levenberg-Marquardt algorithm.

12. The method of claim 11, wherein the pose-aware branch loss function determines feature residuals based on the determined high-resolution estimated three-degree-of-freedom pose and the ground-truth three-degree-of-freedom pose of the ground-view camera.

13. The method of claim 10, further comprising determining the one or more ground feature maps and the one or more ground attention maps from one or more ground perspective images using one or more neural networks, and determining the aerial feature map and the aerial attention map from aerial perspective images using the one or more neural networks.

14. The method of claim 13, further comprising weighting the feature map with the attention map.

15. The method of any one of claims 10 to 14, wherein the top keypoint is determined from the ground feature map.