Method and apparatus for positioning based on image and map data
By generating the first image of the object and the second image based on map data, using feature pooling and score calculation of candidate positioning information, the problem of inaccurate positioning in the prior art is solved, and high-accuracy positioning and augmented reality services are realized in complex environments.
Patent Information
- Application Number
- CN201910648410.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-24
- Filing Date
- 2019-07-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-07-17
AI Technical Summary
In the prior art, it is difficult to accurately locate the location and orientation of objects when providing augmented reality (AR) services, especially in complex environments.
By generating a first image of the object and a second image based on map data, the positioning information of the device is determined by using feature pooling and score calculations of candidate positioning information.
Improves the accuracy and stability of positioning and can effectively provide augmented reality services in complex environments.
Smart Images

Figure CN111089597B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of Korean Patent Application No. 10 - 2018 - 0127589, filed on October 24, 2018, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical field
[0003] The following description relates to a method and apparatus for positioning based on image and map data. Background art
[0004] Various augmented reality (AR) services are provided in fields such as driving assistance for vehicles and other transportation means, games, or entertainment. To provide more accurate and realistic AR, many positioning methods are used. For example, sensor - based positioning methods use a combination of sensors such as global positioning system (GPS) sensors and inertial measurement unit (IMU) sensors to determine the position and orientation of an object. In addition, vision - based positioning methods use camera information. Summary of the invention
[0005] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0006] In one general aspect, a positioning method is disclosed, including: generating a first image of an object according to an input image; generating a second image based on map data including the position of the object, the second image projecting the object's candidate positioning information relative to a device; pooling eigenvalue corresponding to vertices in the second image according to the first image; and determining a score of the candidate positioning information based on the pooled eigenvalue.
[0007] Generating the first image may include: generating a feature map corresponding to a plurality of features.
[0008] Generating the second image may include: extracting a region corresponding to a field of view in the candidate positioning information from the map data; and projecting vertices included in the region into projection points corresponding to the candidate positioning information.
[0009] The pooling may include: selecting pixels in the first image based on the coordinates of the vertices; and obtaining the eigenvalue of the selected pixels.
[0010] The determining may include determining the sum of the pooled eigenvalue.
[0011] Determining the sum may include: determining a weighted sum of the feature values based on weights determined for the feature in response to the first image including a feature map corresponding to the feature.
[0012] The positioning method may include: determining the positioning information of the device based on the scores of the candidate positioning information.
[0013] Determining the positioning information of the device may include: determining, as the positioning information of the device, the candidate positioning information corresponding to the highest score among the scores of multiple pieces of candidate positioning information.
[0014] Determining the positioning information of the device may include: segmenting the second image into regions; and sequentially determining multiple degrees of freedom "DOF" values included in the candidate positioning information using the scores calculated in the regions.
[0015] The multiple DOF values may include: three translational DOF values; and three rotational DOF values.
[0016] The segmentation may include: segmenting the second image into a long-distance region and a short-distance region based on a first criterion associated with distance; and segmenting the short-distance region into a short-distance region towards the vanishing point and a short-distance region not towards the vanishing point based on a second criterion associated with the vanishing point.
[0017] The sequential determination may include: determining the rotational DOF based on the long-distance region; determining the left-right translational DOF based on the short-distance region towards the vanishing point; and determining the front-back translational DOF based on the short-distance region not towards the vanishing point.
[0018] The determination may include: determining the rotational DOF based on long-distance vertices in the vertices included in the second image that are less affected by the translational DOF than a first threshold; determining the left-right translational DOF based on short-distance vertices towards the vanishing point among the short-distance vertices other than the long-distance vertices in the second image that are less affected by the front-back translational DOF than a second threshold; and determining the front-back translational DOF based on short-distance vertices not towards the vanishing point among the short-distance vertices other than the short-distance vertices towards the vanishing point.
[0019] Determining the positioning information of the device may include: determining a direction for improving the score based on the distribution of the pooled feature values; and correcting the candidate positioning information based on the direction.
[0020] The first image may include a probability distribution indicating the degree of proximity to the object, wherein determining the direction includes determining the direction based on the probability distribution.
[0021] Determining the positioning information of the device includes: generating a corrected second image in which the object is projected relative to the corrected candidate positioning information; and determining a corrected score of the corrected candidate positioning information by pooling eigenvalue corresponding to vertices in the corrected second image according to the first image, wherein determining the direction, correcting the candidate positioning information, generating the corrected second image, and calculating the corrected score are iteratively performed until the corrected score meets the condition.
[0022] The positioning method may include: determining a virtual object on the map data to provide an augmented reality (AR) service; and displaying the virtual object based on the determined positioning information.
[0023] The input image may include a driving image of a vehicle, and the virtual object indicates driving route information.
[0024] In another general aspect, a positioning method is disclosed, including: generating a first image of an object according to an input image; generating a second image based on map data including the position of the object, the second image projecting the object relative to candidate positioning information of the device; segmenting the second image into regions; and determining a degree of freedom "DOF" value included in the candidate positioning information through matching between the first image and the regions.
[0025] The determination may include: determining the DOF value included in the candidate positioning information by sequentially using scores calculated through the matching in the regions.
[0026] The determination may include: while changing the DOF value determined for the region, calculating a score corresponding to the changed DOF value by pooling eigenvalue corresponding to vertices in the region according to the first image; and selecting the DOF value corresponding to the highest score.
[0027] The plurality of DOF values may include: three translational DOF values; and three rotational DOF values.
[0028] The segmentation may include: segmenting the second image into a long-distance region and a short-distance region based on a first criterion associated with distance; and segmenting the short-distance region into a short-distance region towards the vanishing point and a short-distance region not towards the vanishing point based on a second criterion associated with the vanishing point.
[0029] The determination may include: determining a rotational DOF based on the long-distance region; determining a left-right translation DOF based on the short-distance region towards the vanishing point; and determining a front-back translation DOF based on the short-distance region not towards the vanishing point.
[0030] The determination may include: determining a rotational DOF based on long-distance vertices among the vertices included in the second image that are less affected by the translation DOF than a first threshold; determining a left-right translation DOF based on short-distance vertices towards the vanishing point among the short-distance vertices other than the long-distance vertices in the second image that are less affected by the front-back translation DOF than a second threshold; and determining the front-back translation DOF based on short-distance vertices not towards the vanishing point among the short-distance vertices other than the short-distance vertices towards the vanishing point.
[0031] The positioning method may include: determining a virtual object on the map data to provide an augmented reality "AR" service; and displaying the virtual object based on the determined DOF values.
[0032] The input image may include a driving image of a vehicle, and the virtual object indicates driving route information.
[0033] In another general aspect, a positioning device is disclosed, including: a processor configured to: generate a first image of an object according to an input image, generate a second image based on map data including the position of the object, the second image projecting candidate positioning information of the object relative to the device, pool eigenvalue corresponding to vertices in the second image according to the first image, and determine a score of the candidate positioning information based on the pooled eigenvalue.
[0034] In another general aspect, a positioning device is disclosed, including: a processor configured to: generate a first image of an object according to an input image, generate a second image based on map data including the position of the object, the second image projecting candidate positioning information of the object relative to the device, segment the second image into regions, and determine a degree of freedom "DOF" value included in the candidate positioning information by matching between the first image and the regions.
[0035] In another general aspect, a positioning device is disclosed, comprising: a sensor disposed on a device and configured to sense one or more of an image of the device and candidate positioning information; a processor configured to: generate a first image of an object based on the image, generate a second image based on map data including the position of the object, the second image projecting the object relative to the candidate positioning information, determine a score of the candidate positioning information based on pooling eigenvalue corresponding to vertices in the first image and the second image, and determine positioning information of the device based on the score; and a head-up display (HUD) configured to visualize a virtual object on the map data based on the determined positioning information.
[0036] The processor may be configured to: segment the second image into a long-distance region and a short-distance region based on distance; and segment the short-distance region into a short-distance region towards a vanishing point and a short-distance region not towards the vanishing point based on the vanishing point.
[0037] The processor may be configured to: determine a rotational degree of freedom (DOF) based on the long-distance region; determine a left-right translation DOF based on the short-distance region towards the vanishing point; and determine a front-back translation DOF based on the short-distance region not towards the vanishing point.
[0038] The processor may be configured to: generate the first image including a feature map corresponding to a plurality of features using a neural network.
[0039] The second image may include a projection of two-dimensional (2D) vertices corresponding to the object.
[0040] The positioning device may include: a memory configured to store the map data, the image, the first image, the second image, the score, and instructions which, when executed, configure the processor to determine any one or any combination of the determined positioning information and the virtual object.
[0041] Other features and aspects will be apparent from the following detailed description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figures 1A to 1C An example demonstrating the importance of positioning accuracy in an augmented reality (AR) application is shown.
[0043] Figure 2 An example of calculating a positioning score is shown.
[0044] Figure 3 An example of calculating a positioning score is shown.
[0045] Figure 4 An example of determining the location information of a device by using a location score is shown.
[0046] Figure 5 An example of the scores of each piece of candidate location information is shown.
[0047] Figure 6 An example of a location method is shown.
[0048] Figure 7 An example of determining the location information of a device by using an optimization technique is shown.
[0049] Figure 8A and Figure 8B An example of an optimization technique is shown.
[0050] Figures 9A to 9E An example of the result of applying the optimization technique is shown.
[0051] Figure 10 An example of a location method by parameter update is shown.
[0052] Figure 11 An example of the result of applying the location method by gradually performing parameter update is shown.
[0053] Figure 12 An example of a neural network for generating a feature map is shown.
[0054] Figure 13 An example of a location device is shown.
[0055] Throughout the drawings and the detailed description, unless otherwise described or provided, the same reference numerals will be understood to refer to the same elements, features, and structures. The drawings may not be drawn to scale, and for clarity, illustration, and convenience, the relative dimensions, scales, and depictions of the elements in the drawings may be exaggerated. Detailed Description
[0056] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent. For example, the operation sequences described herein are merely examples and are not limited to those set forth herein, but may be changed, which becomes apparent after understanding the disclosure of the present application, except for operations that must be performed in a certain order. Moreover, descriptions of features known in the art may be omitted for increased clarity and conciseness.
[0057] The features described herein may be embodied in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0058] Terms such as first, second, etc. may be used herein to describe components. Each of these terms is not used to define the essence, order, or sequence of the corresponding component, but is only used to distinguish the corresponding component from other components. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0059] If the specification states that a first component is "connected", "coupled", or "joined" to a second component, the first component may be directly "connected", "coupled", or "joined" to the second component, or a third component may be "connected", "coupled", or "joined" between the first component and the second component. However, if the specification states that the first component is "directly connected" or "directly joined" to the second component, then no third component may be "connected" or "joined" between the first component and the second component. Similar expressions, such as "between" and "immediately between" and "adjacent to" and "immediately adjacent to", should also be interpreted in this way.
[0060] The terms used herein are for the purpose of describing particular examples only and are not limiting of the examples. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should also be understood that the terms "comprises" and / or "comprising", when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0061] The term "may" is used herein with respect to examples or embodiments, such as what an example or embodiment may include or implement, meaning that there is at least one example or embodiment in which this feature is included or implemented, but not all examples and embodiments are limited thereto.
[0062] The examples set forth below may be implemented in hardware suitable for positioning technologies based on image and map data. For example, these examples can be used to improve the accuracy of positioning in an augmented reality head-up display (AR HUD). In addition, many location-based services, in addition to HUDs, also require positioning, and these examples can be used to estimate position and orientation in an environment that provides high-density (HD) map data for high-precision positioning.
[0063] In the following, examples will be described in detail with reference to the accompanying drawings. In the drawings, like reference numerals are used for like elements.
[0064] Figures 1A to 1C An example showing the importance of positioning accuracy in an AR application is shown.
[0065] Reference Figures 1A to 1C , in the example, AR adds or enhances information based on reality and provides the added or enhanced information. For example, AR adds a virtual object corresponding to a virtual image to a background image or an image of the real world and presents an image with the added object. AR appropriately combines the virtual world and the real world so that the user experiences an immersive experience when interacting with the virtual world in real time without recognizing the separation between the real environment and the virtual environment. To match the virtual object with the real image, it is necessary to determine the position and orientation (i.e., positioning information) of the user device or the user providing AR.
[0066] The positioning information used to provide AR is used to arrange the virtual object at a desired position in the image. In the following, for ease of description, an example of a driving guidance lane corresponding to a virtual object is shown on a road surface. However, the example is not limited thereto.
[0067] Figure 1A An AR image 120 with a relatively small positioning error is shown. Figure 1B An AR image 140 with a relatively large positioning error is shown.
[0068] For example, a reference route of a vehicle is displayed on a road image based on the positioning information of an object 110. In the example, the object corresponds to the vehicle and / or the user terminal performing the positioning. When the positioning information of the object 110 includes an error within a small tolerance range, the driving guidance lane 115, which is a virtual object to be displayed by the device, is visually properly aligned with the real road image, as shown in the image 120. When the positioning information of the object 130 includes a relatively large error, that is, outside the tolerance range, the driving guidance lane 135, which is a virtual object to be displayed by the device, is not visually properly aligned with the real road image, as shown in the image 140.
[0069] Reference Figure 1C , the positioning information includes the position and orientation of the device. The position corresponds to three-dimensional (3D) coordinates, such as lateral (t x ), vertical (t y ) and longitudinal (t z ), that is, (x, y, z), as translational degrees of freedom (DOF). In addition, the orientation corresponds to pitch (r x ), yaw (r y ) and roll (r z) as the rotational DOF. This position is obtained by, for example, a Global Positioning System (GPS) sensor and a Light Detection and Ranging (LiDAR), and the orientation is obtained by, for example, an Inertial Measurement Unit (IMU) sensor and a gyroscope sensor. The positioning information is interpreted as having 6 DOFs, including the position and the orientation.
[0070] The vehicle described herein refers to any transportation, delivery, or conveyance tool, such as a car, truck, tractor, scooter, motorcycle, bicycle, amphibious vehicle, snowmobile, boat, bus vehicle, bus, monorail, train, tram, autonomous or self-driving vehicle, intelligent vehicle, self-driving vehicle, unmanned aerial vehicle, electric vehicle (EV), hybrid vehicle, intelligent mobile device, intelligent vehicle with an Advanced Driving Assistance System (ADAS), or drone. In an example, the intelligent mobile device includes mobile devices such as electric wheels, electric kickboards, and electric bicycles. In an example, the vehicle includes motor vehicles and non-motor vehicles, such as vehicles with a power engine (e.g., a cultivator or a motorcycle), a bicycle, or a handcart.
[0071] In addition to the vehicles described herein, the methods and apparatuses described herein may be included in various other devices, such as smart phones, walking assistance devices, wearable devices, safety devices, robots, mobile terminals, and various Internet of Things (IoT) devices.
[0072] The term "road" is a passageway, route, or connection between two places that has been improved to allow travel by foot or some form of conveyance tool (e.g., a vehicle). A road may include various types of roads, which refer to the path on which a vehicle travels, and includes various types of roads, such as highways, national roads, local roads, expressways, rural roads, local roads, high-speed national roads, and motor vehicle lanes. A road includes one or more lanes.
[0073] The term "lane" refers to the road space distinguished by lines marked on the road surface. A lane is distinguished by its left line and right line or its lane boundary lines. In addition, the lines are various lines, such as solid lines, dashed lines, curves, and zigzag lines marked with colors (e.g., white, blue, and yellow) on the road surface. A line corresponds to a single line separating a single lane, or corresponds to a pair of lines separating a single lane, i.e., the left line and the right line corresponding to the lane boundary lines. The term "lane boundary" may be used interchangeably with the term "lane marking".
[0074] The methods and apparatuses described herein are used for road guidance information in navigation devices of vehicles (e.g., augmented reality head-up displays (AR 3D HUDs) and autonomous vehicles). The examples set forth below can be used to display lines in an AR navigation system of an intelligent vehicle, generate visual information to assist in the maneuvering of an autonomous vehicle, or provide various control information related to the driving of a vehicle. Additionally, these examples are used to assist in safe and enjoyable driving by providing visual information to devices including intelligent systems (e.g., HUDs installed on a vehicle for driving assistance or full autonomous driving). In an example, the examples described herein can also be used to interpret visual information of an intelligent system installed in a vehicle for full autonomous driving or driving assistance and for assisting in safe and comfortable driving. The examples described herein can be applicable to vehicles and vehicle management systems, such as autonomous vehicles, automated or autonomous driving systems, intelligent vehicles, advanced driver assistance systems (ADASs), navigation systems that assist a vehicle in safely staying in the lane in which the vehicle is traveling, smart phones, or mobile devices. The examples related to displaying road guidance information of a vehicle are provided only as examples, and other examples (e.g., training, gaming, applications in healthcare, public safety, tourism, and marketing) are considered to be within the scope of the present disclosure.
[0075] Figure 2 An example of calculating a localization score is shown.
[0076] Reference Figure 2 , the localization device calculates a score s(θ) of a localization parameter θ based on map data Q and an image I. The localization device can be implemented by one or more hardware modules.
[0077] In an example, the map data is a point cloud including a plurality of 3D vertices corresponding to objects such as lines. The 3D vertices of the map data are projected onto two-dimensional (2D) vertices based on the localization parameter. The features of the image include feature values extracted based on the pixels included in the image. Thus, for the examples described herein, a correspondence between the vertices of the map data and the features of the image may not be required.
[0078] For the examples described herein, information related to the correspondence or matching between the vertices of the map data (e.g., 2D vertices) and the features or pixels of the image may not be required. Additionally, since the features extracted from the image may not be parameterized, a separate analysis of the relationships between the features or a search of the map data may not be required.
[0079] The localization parameter θ is a position / orientation information parameter and is defined as Figure 1C the 6-DOF variable described in. The localization parameter θ corresponds to approximate position / orientation information. The localization device improves the localization accuracy by using a scoring technique based on the image and the map data to correct the localization parameter θ.
[0080] In an example, the positioning device configures a feature map by extracting features from an image I. The positioning device calculates a matching score related to positioning parameters θ. Specifically, the positioning device calculates the matching score by projecting vertices from map data Q based on the positioning parameters and pooling the feature values of pixels in the feature map corresponding to the 2D coordinates of the projected vertices. The positioning device updates the positioning parameters θ to increase the matching score.
[0081] In an example, the device is any device that executes the positioning method and includes devices such as vehicles, navigation systems, or user devices (e.g., smart phones). As described above, the positioning information has 6 DOFs, including the position and orientation of the device. The positioning information is obtained based on the outputs of sensors such as IMU sensors, GPS sensors, lidar sensors, and radio detection and ranging (radar).
[0082] The input image is a background image or other image to be displayed together with a virtual object to provide an AR service. The input image includes, for example, a driving image of a vehicle. In an example, the driving image is a driving image acquired using a capturing device mounted on the vehicle and includes one or more frames.
[0083] The positioning device obtains the input image based on the output of the capturing device. The capturing device is fixed to a position on the vehicle, such as the windshield, dashboard, or rearview mirror of the vehicle, to capture a driving image of the view in front of the vehicle. The capturing device includes, for example, a vision sensor, an image sensor, or a device performing a similar function. According to an example, the capturing device captures a single image per frame or captures multiple images. In an example, an image captured by a device other than the capturing device fixed to the vehicle is also used as the driving image. The object includes, for example, lines, road markings, traffic lights, traffic signs, curbs, pedestrians, and structures. The lines include lines such as lane boundary lines, road centerlines, and stop lines. The road markings include markings such as no parking markings, crosswalk markings, parking trailer area markings, and speed limit markings.
[0084] In an example, the map data is high-definition (HD) map data. The HD map is a 3D map with high density (e.g., centimeter-level density) and can be used for autonomous driving. The HD map includes, for example, line information related to road centerlines and boundary lines in the form of 3D digital data, and information related to traffic lights, traffic signs, curbs, road markings, and various structures. The HD map is established by, for example, a mobile mapping system (MMS). The MMS (a 3D spatial information surveying system equipped with various sensors) uses a moving object equipped with sensors such as cameras, lidar, and GPS to measure positions and geographical features to obtain minute position information.
[0085] Figure 3 An example of calculating a localization score is shown.
[0086] Reference Figure 3 , the localization device 200 includes transformation devices 210 and 220, a feature extractor 230, and a pooler 240. In the example, the transformation device 210 receives the parameter θ and the map data Q, and applies the 3D position and 3D orientation corresponding to the parameter θ of the device to the map data Q through the 3D transformation T. For example, the transformation device 210 extracts from the map data a region corresponding to the field of view at the position and orientation corresponding to the parameter θ. In the example, the transformation device 220 generates a projection image at the viewpoint of the device through the perspective transformation P. For example, the transformation device 220 projects the 3D vertices included in the region extracted by the transformation device 210 onto a 2D projection plane corresponding to the parameter θ. In this example, the 3D vertex q included in the map data Q i k is transformed into the 2D vertex p in the projection image based on the parameter θ through the transformations T and P i k . Here, k represents an index indicating different features or classes, and i represents an index indicating the vertex within the corresponding feature or class.
[0087] The feature extractor 230 extracts features from the image I. Depending on the type or class of the object, the features include one or more feature maps F1 and F2 235. For example, the feature map F1 includes features related to lines in the image, and the feature map F2 includes features related to traffic signs in the image. For ease of description, an example of extracting two feature maps F1 and F2 235 is described. However, the example is not limited to this.
[0088] In the example, the localization device includes separate feature extractors to extract multiple feature maps. In another example, the localization device includes a single feature extractor such as a deep neural network (DNN) to output multiple feature maps for each lane.
[0089] In some examples, the extracted feature maps F1 and F2 235 may include errors and thus may not accurately specify the values of the corresponding features pixel by pixel. In this example, each feature map has a value between "0" and "1" for each pixel. The feature value of a pixel indicates the intensity of the pixel related to the feature.
[0090] The 2D vertex p of the projection image i k refers to the pixel in the image I that is mapped to the 3D vertex q of the map data i k corresponding to. Referring to Equation 1, the scores of the features of the pixels mapped to the image I are summed.
[0091] [Equation 1]
[0092]
[0093] In Equation 1, T() represents transformation T, and P() represents transformation P. In this example, represents a mapped point, and F k () represents the eigenvalue or score of the mapped point corresponding to the k-th feature or class in the feature map. In this example, if the mapped point is not an integer, operations such as rounding or interpolation are performed. Referring to Equation 2, the final score is calculated by computing the weighted sum of the scores of the features.
[0094] [Equation 2]
[0095]
[0096] In this example, an arbitrary scheme is used to set the weight w k . For example, the weight w k is set to a weight equally distributed at once or a value adjusted by training data.
[0097] Figure 4 Shows an example of determining the location information of a device by utilizing the localization score. The operations in Figure 4 can be performed in the order and manner shown. However, the order of some operations can be changed or some operations can be omitted without departing from the spirit and scope of the described illustrative example. Many of the operations shown in Figure 4 can be performed in parallel or concurrently. One or more boxes and combinations of boxes in Figure 4 can be implemented by a computer (e.g., a processor) based on dedicated hardware that performs the specified functions or a combination of dedicated hardware and computer instructions. Except for the description of Figure 4 below, the descriptions of FIGS. 1-3 also apply to Figure 4 and are incorporated herein by reference. Therefore, the above description may not be repeated here.
[0098] Referring to Figure 4 , in operation 410, an input image is received. In operation 430, features are extracted. In operation 420, map data is obtained. In operation 440, candidate location information is obtained. In the example, the candidate location information includes multiple pieces of candidate location information.
[0099] In operation 450, a projected image of the map data with respect to the candidate location information is generated. In the example, when multiple pieces of candidate location information are provided, multiple projected images with respect to the multiple pieces of candidate location information are generated.
[0100] In operation 460, eigenvalue corresponding to the 2D vertex in the projection image is pooled according to the feature map. In addition, in operation 460, a score of candidate localization information is calculated based on the pooled eigenvalues. When multiple pieces of candidate localization information are provided, scores of the multiple pieces of candidate localization information are calculated.
[0101] In operation 470, the best score, such as the highest score, is determined. In operation 480, the candidate localization information with the determined best score is determined as the localization information of the device.
[0102] Although not shown in the drawings, the positioning device 200 determines a virtual object on the map data Q to provide an AR service. For example, the virtual object indicates driving route information and is represented in the form of an arrow or a road marking indicating the traveling direction. The positioning device displays the virtual object together with the input image on a display of the user equipment, a navigation system, or an HUD based on the localization information determined in operation 480.
[0103] Figure 5 An example of scores of multiple pieces of candidate localization information is shown.
[0104] Reference Figure 5 , the degree of visual alignment between the image 510 and the projection image with respect to the first candidate localization information is lower than the degree of visual alignment between the image 520 and the projection image with respect to the second candidate localization information. Therefore, the score pooled based on the first candidate localization information according to the feature map of the image 510 is calculated to be lower than the score pooled based on the second candidate localization information according to the feature map of the image 520.
[0105] Figure 6 An example of the positioning method is shown. The operations in Figure 6 can be performed in the shown order and manner. However, the order of some operations can be changed or some operations can be omitted without departing from the spirit and scope of the described illustrative example. Many operations shown in Figure 6 can be performed in parallel or concurrently. One or more boxes and combinations of boxes in Figure 6 can be implemented by a computer (e.g., a processor) based on dedicated hardware that performs the specified functions or a combination of dedicated hardware and computer instructions. In addition to the description of Figure 6 below, the descriptions of FIGS. 1-5 also apply to Figure 6 and are incorporated herein by reference. Therefore, the above description may not be repeated here.
[0106] Reference Figure 6, in operation 610, at least one feature map is extracted from the input image. In operation 620, a candidate set of positioning parameters (e.g., position / orientation parameters) is selected. In operation 630, it is determined whether additional candidate positioning information is to be evaluated. In operation 640, when it is determined that there is candidate positioning information to be evaluated, a projection image corresponding to the candidate positioning information is generated. In operation 650, the score of the feature is calculated. In operation 660, the final score is calculated by the weighted sum of the scores of the features. In operation 670, it is determined whether the best candidate positioning information is to be updated. In the example, the best candidate positioning information is determined by comparing the prior best candidate positioning information among the evaluated candidate positioning information with the final score.
[0107] When there is no other candidate positioning information to be evaluated, in operation 680, the best candidate positioning information among the evaluated candidate positioning information is determined as the positioning information of the device. In this example, the parameters of the best candidate positioning information are determined as the position / orientation parameters of the device.
[0108] Figure 7 An example of determining the positioning information of the device by an optimization technique is shown. The operations in Figure 7 can be performed in the order and manner shown. However, the order of some operations can be changed or some operations can be omitted without departing from the spirit and scope of the described illustrative example. Many of the operations shown in Figure 7 can be performed in parallel or concurrently. One or more of the boxes and combinations of boxes in Figure 7 can be implemented by a computer (e.g., a processor) based on dedicated hardware that performs the specified functions or a combination of dedicated hardware and computer instructions. In addition to the description of Figure 7 below, the descriptions of FIGS. 1-6 also apply to Figure 7 and are incorporated herein by reference. Therefore, the above description may not be repeated here.
[0109] Referring to Figure 7 , in operation 710, an input image is received. In operation 730, a feature map is extracted. In operation 720, map data is received. In operation 740, initial positioning information is received. In operation 750, a projection image with respect to the initial positioning information is generated.
[0110] In operation 760, the initial positioning information is updated by an optimization technique. In operation 770, the positioning information of the device is determined as the optimized positioning information.
[0111] Hereinafter, the optimization technique of operation 760 will be described in detail.
[0112] Figure 8A and Figure 8BAn example of an optimization technique is shown. The operations in Figure 8A can be performed in the order and manner shown. However, the order of some operations can be changed or some operations can be omitted without departing from the spirit and scope of the illustrative examples described. Many of the operations shown in Figure 8A can be performed in parallel or concurrently. One or more boxes and combinations of boxes of Figure 8A can be implemented by a computer (e.g., a processor) based on dedicated hardware that performs a specified function or a combination of dedicated hardware and computer instructions. In addition to the description of Figure 8A below, the descriptions of FIGS. 1-7 also apply to Figure 8A and are incorporated herein by reference. Therefore, the above description may not be repeated here.
[0113] The positioning device supports the global optimization process. The positioning device classifies 2D vertices projected from map data by criteria other than features (e.g., whether the distance or area is oriented towards the vanishing point) and uses the classified 2D vertices to estimate different DOFs of the positioning parameters.
[0114] In an example, the positioning device divides the projected image into multiple regions and determines the positioning information of the device by matching between the feature map and the regions. Specifically, the positioning device sequentially determines multiple DOF values included in the positioning information by sequentially using the scores calculated by matching in the regions. For example, the positioning device pools the feature values corresponding to the 2D vertices included in the region according to the feature map while changing the DOF values determined for the region. The positioning device calculates the scores corresponding to the changed DOF values based on the pooled feature values. In the example, the positioning device determines the DOF as the value corresponding to the highest score.
[0115] Distant vertices in the projected image have the characteristic of being almost invariant to changes in the position parameters. Based on such a characteristic, the positioning device separately performs a process of determining the orientation parameter by using the long-distance vertices to calculate the scores and a process of determining the position parameter by using the short-distance vertices to calculate the scores. This reduces the DOFs to be estimated in each process, and thus the search complexity or the possibility of local convergence during optimization is reduced.
[0116] In an example, the positioning device divides the projected image into a long-distance region and a short-distance region based on a first criterion associated with the distance, and divides the short-distance region into a short-distance region oriented towards the vanishing point and a short-distance region not oriented towards the vanishing point based on a second criterion associated with the vanishing point, which will be further described below. Here, the long-distance region includes 2D vertices whose influence by the translational DOF is below a threshold. The short-distance region oriented towards the vanishing point includes 2D vertices whose influence caused by the DOF related to the movement in the front-back direction (or the front-back translational DOF) is less than the threshold.
[0117] In an example, the positioning device uses a portion of the DOF of the positioning parameter as a value determined by prior calibration. In addition, in an example, the r of the positioning parameter is determined by prior calibration. z and t y This is because the camera is mounted at a height t y and roll z is fixed.
[0118] refer to Figure 8A and Figure 8B In operation 810, a feature map is extracted from the input image. In operation 820, a long-distance vertex Q1 is selected from the projected image of the map data. The long-distance vertex Q1 is substantially unaffected by the translation DOF t in the DOF of the positioning parameter. x and t z Therefore, in operation 830, the rotation DOF r is determined based on the long-distance vertex Q1. x and r y . Rotation DOFr x and r y is called the orientation parameter.
[0119] In the example, the positioning device changes r x While performing a parallel translation of the long-distance vertex Q1 in the longitudinal direction, and changing r y The positioning device searches for r x and r y The value of r x and r y The value of makes the score calculated for the long-distance vertex Q1 greater than or equal to the target value.
[0120] In operation 840, short-distance vertices are selected from the map data. z and t y and r determined by the long-distance vertex Q1 x and r y , select short-distance vertices.
[0121] The positioning device selects the vertex Q2 corresponding to the line toward the vanishing point among the short-distance vertices, and selects another vertex Q3. The vertex Q2 is substantially not affected by the front-back translation DOF t in the DOF of the positioning parameter. z (Translation DOF t z ). Therefore, in operation 850, the translation DOF t is determined based on vertex Q2. x In addition, the translation DOFt is determined based on vertex Q3. z . Translation DOF tx and t z are called positional parameters.
[0122] Figures 9A to 9E An example showing the result of applying the optimization technique is shown.
[0123] In Figure 9A the feature 911 extracted from the feature map 910 and the vertex 921 projected from the map data 920 are shown. The initial positioning information related to the map data 920 is inaccurate, and thus the feature 911 and the vertex 921 do not match. Hereinafter, a process of sequentially determining the DOF of the positioning parameters to match the feature 911 and the vertex 921 will be described.
[0124] Considering that the camera mounted on the vehicle has a relatively constant height and a relatively constant roll with respect to the road surface, t y and r z are pre-calibrated.
[0125] Referring to Figure 9B , the positioning device removes the roll effect from the feature map 910 and the map data 920. In 930, the positioning device removes the roll effect from the feature map 910 by rotating the feature map 910 based on the pre-calibrated r z . In addition, the positioning device detects the vertices near the initial positioning information from the map data 920, approximates the detected vertices to a plane, and rotates the map data 920 to remove the roll of the plane.
[0126] In addition, the positioning device uses the pre-calibrated t y to correct the height of the vertices of the map data 920.
[0127] Referring to Figure 9C , the positioning device infers r x and r y corresponding to the parallel translation based on the long-distance vertices 940 in the projected image of the map data using the initial positioning information. In the example, the positioning device performs a parallel translation of the long-distance vertices 940 such that the correlation between the long-distance vertices 940 and the feature map 910 is greater than or equal to the target value. For example, by adjusting r x as in 945 to rotate the vertices on the map data 920, the long-distance vertices 940 in the projected image can be well matched with the features of the feature map 910.
[0128] Referring to Figure 9D, the positioning device obtains the vanishing point of the adjacent line 950 by analyzing the vertices in the projection image. The positioning device aligns the vanishing point of the adjacent line 950 at a position in the feature map 910, such as the center of the feature map. In this example, the adjacent line has the property of being invariant to translation in the z direction. The positioning device moves the vertex corresponding to the adjacent line 950 in the x direction such that the correlation between the adjacent line 950 and the feature map 910 is greater than or equal to a target value. For example, by adjusting t as in 955 x to move the vertex on the map data 920, the vertex corresponding to the adjacent line 950 in the projection image can be well matched with the features of the feature map 910.
[0129] Reference Figure 9E , the positioning device uses the remaining vertices 960 other than the lanes in the direction of the vanishing point among the short-distance vertices in the projection image to detect translation in the z direction. The positioning device moves the remaining vertices 960 in the z direction such that the correlation between the remaining vertices 960 and the feature map 910 is greater than or equal to a target value. For example, by adjusting t as in 965 z to move the vertex on the map data 920, the remaining vertices 960 in the projection image can be well matched with the features of the feature map 910.
[0130] Figure 10 An example of the positioning method by parameter update is shown. The operations in Figure 10 can be performed in the order and manner shown. However, the order of some operations can be changed or some operations can be omitted without departing from the spirit and scope of the described illustrative example. Many of the operations shown in Figure 10 can be performed in parallel or concurrently. One or more boxes and combinations of boxes in Figure 10 can be implemented by a computer based on dedicated hardware that performs the specified functions (e.g., a processor) or a combination of dedicated hardware and computer instructions. In addition to the description of Figure 10 below, the descriptions of FIGS. 1-9 also apply to Figure 10 and are incorporated herein by reference. Therefore, the above description may not be repeated here.
[0131] Reference Figure 10 , in operation 1010, a feature map is extracted from the input image. In operation 1020, initial positioning information, such as initial values of position / orientation parameters, is selected. In operation 1030, the score of the current parameter is calculated, and the direction to improve the score of the current parameter is calculated.
[0132] In the example, the feature map includes a probability distribution indicating the degree of proximity to an object. For example, the features included in the feature map include information related to the distance to the nearest object, which is represented using a normalized value between "0" and "1". In this example, the feature map provides information related to the direction towards the object. The positioning device pools the feature values of the feature map corresponding to the 2D vertices projected from the map data by the current parameters. The positioning device determines the direction for increasing the score of the current parameters based on the pooled feature values.
[0133] In operation 1040, it is determined whether an iteration termination condition is satisfied. When it is determined that the iteration termination condition is not satisfied, the parameters are updated in operation 1050. The positioning device updates the parameters based on the direction calculated in operation 1030. Operations 1050, 1030, and 1040 are iteratively executed until the iteration termination condition is satisfied. The iteration termination condition includes whether the score of the parameters is greater than or equal to a target value. In the example, the iteration termination condition also includes whether the iteration count exceeds a threshold for system stability.
[0134] When it is determined that the iteration termination condition is satisfied, the current parameters are selected as the final positioning information, such as the final position / orientation parameters, in operation 1060.
[0135] In Figure 10 the example, a process of gradually searching starting from an initial value θ0 is performed to determine better parameters. For better performance, such a local optimization scheme requires a good initial value. Therefore, Figure 10 the example selectively includes operations for respectively estimating the initial value. For example, while considering the parameters obtained through the example of FIG. 8 as the initial value, Figure 10 the example is performed. In this example, the effect of correcting values (such as the camera height and roll fixed by pre-calibration) to adapt to changes occurring in the real driving environment is achieved.
[0136] Figure 11 An example showing the step-by-step results of applying the positioning method through parameter update is shown.
[0137] Referring to Figure 11 , an input image 1105, a first image 1110, and a second image 1120 are shown. The first image 1110 is generated to correspond to the input image 1105. In addition, the second image 1120 is an image generated by projecting an object based on the map data with respect to the positioning information corresponding to the initial positioning information The second image 1120 is a projection image including a plurality of 2D vertices corresponding to the object.
[0138] As shown in Image 1130, the positioning device calculates a score by matching the first image 1110 and the second image 1120. The positioning device calculates the score by summing the values of the pixels in the first image 1110 that correspond to the object included in the second image 1120 among the multiple pixels included in the first image 1110.
[0139] For example, based on the distance to adjacent objects, the multiple pixels included in the first image 1110 have values between "0" and "1". Each pixel has a value close to "1" when close to an adjacent object and a value close to "0" when far from an adjacent object. The positioning device extracts the pixels that match the second image 1120 from the multiple pixels included in the first image 1110 and calculates the score by summing the values of the extracted pixels.
[0140] The positioning device corrects the positioning information based on the directionality of the first image 1110 to increase the degree of visual alignment, that is, the score. The positioning device calculates a positioning correction value such that the positioning information of the object included in the second image 1120 conforms to the directionality of the first image 1110. In operation 1140, the positioning device applies the positioning correction value to the positioning information corresponding to the initial positioning information, thereby according to to update the positioning information. For example, the positioning device determines the direction in which the object of the second image 1120 should move to increase the score based on the directionality of the first image 1110. When the positioning information is updated, the object of the second image 1120 is moved, and thus the positioning device updates the positioning information based on the directionality included in the first image 1110.
[0141] The positioning device generates an updated second image 1150 based on the updated positioning information The positioning device calculates the score by matching the updated second image 1150 and the first image 1110.
[0142] The positioning device finally outputs the optimized positioning information by calculating the positioning correction value that makes the score greater than or equal to the standard through the above process
[0143] Figure 12 An example of a neural network that generates a feature map is shown.
[0144] Refer to Figure 12 , a process of generating a distance field map 1250 corresponding to the first image by applying an input image 1210 to a neural network 1230 is shown.
[0145] In an example, a neural network 1230 is trained to generate a first image based on an input image 1210 that includes directionality corresponding to an object included in the input image 1210. The neural network 1230 is implemented on a hardware-based model that includes a framework or structure having multiple layers or operations to provide many different machine learning algorithms to work together, process complex data inputs, and identify patterns. The neural network 1230 is implemented in various structures, such as a convolutional neural network (CNN), a deep neural network (DNN), an n-layer neural network, a recurrent neural network (RNN), or a bidirectional long short-term memory network (BLSTM). The DNN includes, for example, a fully connected network, a CNN, a deep convolutional network, or a recurrent neural network (RNN), a deep belief network, a bidirectional neural network, a restricted Boltzman machine, or may include different or overlapping neural network portions having fully connected, convolutional, recurrent, and / or bidirectional connections, respectively. The neural network 1230 is based on deep learning and maps input data and output data in a non-linear relationship to perform, for example, object classification, object recognition, speech recognition, or image recognition.
[0146] A neural network can be implemented as an architecture having multiple layers, which includes an input image, a feature map, and an output. In the neural network, a convolution operation is performed between the input image and a filter called a kernel, and as a result of the convolution operation, a feature map is output. Here, the output feature map is the input feature map, and again, a convolution operation is performed between the output feature map and the kernel, and a new feature map is output as a result. Based on this repeatedly performed convolution operation, a result of identifying the characteristics of the input image via the neural network can be output.
[0147] In an example, the neural network 1230 estimates an object included in the input image 1210 in the form of a distance field map 1250. For example, when the first image includes directionality information toward a nearby object as in the distance field map 1250, an optimized directionality can be determined by using gradient descent. In addition, when there is a probability distribution indicating the degree of proximity to an object over the entire image as in the distance field map 1250, the amount of data for training increases, and thus the performance of the neural network is improved compared to the case of training using sparse data.
[0148] Figure 13 An example of a positioning device is shown.
[0149] Reference Figure 13 , the positioning device 1300 includes a sensor 1310 and a processor 1330. The positioning device 1300 further includes a memory 1350, a communication interface 1370, and a display device 1390. The sensor 1310, the processor 1330, the memory 1350, the communication interface 1370, and the display device 1390 are connected to each other via a communication bus 1305.
[0150] The sensor 1310 includes, for example, an image sensor, a vision sensor, an acceleration sensor, a gyroscope sensor, a GPS sensor, an IMU sensor, a radar, and a lidar. The sensor 1310 acquires or captures an input image including a driving image of the vehicle. In addition to sensing positioning information (e.g., GPS coordinates, position, and orientation of the vehicle), the sensor 1310 also senses information such as the speed, acceleration, traveling direction, and steering angle of the vehicle.
[0151] In the example, the positioning device 1300 obtains sensing information of various sensors including the input image through the communication interface 1370. The communication interface 1370 receives sensing information including the driving image from other sensors existing outside the positioning device 1300.
[0152] The processor 1330 outputs the corrected positioning information through the communication interface 1370 and / or the display device 1390, or displays a virtual object and the input image on the map data based on the corrected positioning information, thereby providing an AR service. In addition, the processor 1330 executes at least one of the methods described above with reference to FIGS. 1 to Figure 13 the algorithm corresponding to the at least one method.
[0153] The processor 1330 is a data processing device implemented by hardware, and the hardware includes a circuit having a physical structure for performing desired operations. For example, the desired operations include instructions or codes included in a program. For example, the data processing device implemented by hardware includes a microprocessor, a central processing unit (CPU), a processor core, a multi-core processor, a multi-processor, an application specific integrated circuit (ASIC), and a field programmable gate array (FPGA). In the example, the processor 1330 may be a graphics processing unit (GPU), a reconfigurable processor, or have any other form of multi-processor or single-processor configuration. The processor 1330 executes a program and controls the positioning device 1300. In the example, the processor 1330 executes a program and controls the neural network 1230. The program code executed by the processor 1330 is stored in the memory 1350. Further details about the processor 1330 are provided below.
[0154] The memory 1350 stores the positioning information, the first image, the second image, and / or the corrected positioning information of the positioning device 1300. The memory 1350 stores various information generated during the processing executed by the processor 1330. In addition, the memory 1350 stores various data and programs. The memory 1350 includes a volatile memory or a non-volatile memory. The memory 1350 includes a large-capacity storage medium such as a hard disk to store various data. Further details about the memory 1120 are provided below.
[0155] The display device 1390 outputs the corrected positioning information by the processor 1330, or displays a virtual object together with the input image on the map data based on the corrected positioning information. The display device 1390 is a physical structure including one or more hardware components that provide the ability to present a user interface, present a display, and / or receive user input. However, the display device 1390 is not limited to the above examples, and any other display (e.g., a smart phone and an eyewear display (EGD)) effectively connected to the positioning device 1300 may be used without departing from the spirit and scope of the described illustrative examples.
[0156] According to an example, even when the viewpoints of the capture device and the positioning device do not match like those of a HUD or AR glasses, the positioning device performs positioning in a viewpoint-independent manner by using the result of performing the above-described positioning method based on the capture device to update the 3D positioning information of the positioning device. In addition, when the viewpoints of the capture device and the positioning device match like those of a mobile terminal or a smart phone, the positioning device updates the 3D positioning information and is also used to directly correct the 2D position in the image.
[0157] The examples described herein provide techniques for performing positioning without establishing a correspondence between the vertices of an image and the vertices of map data. In addition, these examples provide techniques for performing positioning without parameterizing the features of an image, without extracting relationships invariant to three-dimensional (3D) transformations and perspective transformations, or without easily specifying such invariant relationships during the search for map data.
[0158] The positioning devices 200 and 1300, the transformation devices 210 and 220, the feature extractor 230, the pooling device 240, and herein with respect to FIGS. 1 to Figure 13The other apparatuses, units, modules, devices, and other components described are implemented by hardware components. Examples of hardware components that can be used to perform the operations described in the present application, where appropriate, include controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in the present application. In other examples, one or more hardware components that perform the operations described in the present application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements, such as a logic gate array, a controller and an arithmetic logic unit, a digital signal processor, a microcomputer, a programmable logic controller, a field programmable gate array, a programmable logic array, a microprocessor, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result. In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software, such as an operating system (OS) and one or more software applications running on the OS, to perform the operations described in the present application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For the sake of brevity, the singular terms "processor" or "computer" can be used in the description of the examples described in the present application, but in other examples, multiple processors or computers can be used, or a processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors or another processor and another controller. One or more processors or a processor and a controller can implement a single hardware component, or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of which include single processor, independent processor, parallel processor, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.
[0159] FIG. 1 to perform the operations described in the present application Figure 13The method shown is performed by computing hardware, e.g., by one or more processors or computers that implement instructions or software as described above to perform the operations described in this application (the operations performed by the method). For example, a single operation or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors or a processor and a controller, and one or more other operations can be performed by one or more other processors or another processor and another controller. One or more processors or a processor and a controller can perform a single operation or two or more operations.
[0160] Instructions or software for controlling a processor or computer to implement the hardware components and perform the method as described above are written as a computer program, code segment, instruction, or any combination thereof, for individually or jointly instructing or configuring the processor or computer to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and the method described above. In an example, the instructions or software include at least one of the following: applet, dynamic link library (DLL), middleware, firmware, device driver, application that stores output status information. In one example, the instructions or software include machine code directly executable by the processor or computer, e.g., machine code generated by a compiler. In another example, the instructions or software include higher-level code executable by the processor or computer using an interpreter. An ordinary programmer in the art can easily write the instructions or software based on the block diagrams and flowcharts shown in the drawings and the corresponding descriptions in the specification, which disclose algorithms for performing the operations performed by the hardware components and the method described above.
[0161] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and execute the methods as described above, as well as any associated data, data files, and data structures, can be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), flash memory, card-type memory (such as, multimedia card, secure digital (SD) card, or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that one or more processors or computers store, access, and execute the instructions and software and any associated data, data files, and data structures in a distributed manner.
[0162] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein are considered to be illustrative only and not for the purpose of limitation. The description of the features or aspects in each example is considered to be applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner and / or replaced or supplemented by other components or their equivalents. Accordingly, the scope of the present disclosure is not defined by the specific embodiments, but by the claims and their equivalents, and all variations within the scope of the claims and their equivalents are construed as being included in the present disclosure.
Claims
1. A positioning method, comprising: Generating a first image of an object based on an input image; Generating a second image based on map data including the position of the object to project candidate positioning information of the object relative to a device; Pooling eigenvalue corresponding to vertices in the second image according to the first image; Determining a score of the candidate positioning information based on the pooled eigenvalues, and Determining positioning information of the device based on the score of the candidate positioning information, wherein determining the positioning information of the device includes: Determining a rotational degree of freedom (DOF) based on long-distance vertices among the vertices included in the second image that are affected by a translational DOF less than a first threshold; Determining a left-right translational DOF based on short-distance vertices toward a vanishing point among short-distance vertices other than the long-distance vertices in the second image that are affected by a front-back translational DOF less than a second threshold; and Determining the front-back translational DOF based on short-distance vertices other than the short-distance vertices toward the vanishing point among the short-distance vertices.
2. The positioning method according to claim 1, wherein, Generating the first image includes: generating a feature map corresponding to multiple features.
3. The positioning method according to claim 1, wherein Generating the second image includes: Extracting a region corresponding to a field of view in the candidate positioning information from the map data; and Projecting vertices included in the region to projection points corresponding to the candidate positioning information.
4. The positioning method according to claim 1, wherein, The pooling includes: Selecting pixels in the first image based on coordinates of the vertices; and Obtaining eigenvalues of the selected pixels.
5. The positioning method according to claim 1, wherein Determining the score of the candidate positioning information includes determining a sum of the pooled eigenvalues.
6. The positioning method according to claim 5, wherein, Determining the sum includes: determining a weighted sum of the eigenvalues based on weights determined for the features in response to the first image including a feature map corresponding to the features.
7. The positioning method according to claim 1, wherein, Determining the positioning information of the device includes: determining, as the positioning information of the device, the candidate positioning information corresponding to the highest score among scores of multiple pieces of candidate positioning information.
8. The positioning method according to claim 1, wherein Determining the positioning information of the device includes: Segmenting the second image into regions; and Sequentially determining multiple DOF values included in the candidate positioning information using scores calculated in the regions.
9. The positioning method according to claim 8, wherein, The multiple DOF values include: Three translational DOF values; and Three rotational DOF values.
10. The positioning method according to claim 8, wherein, The segmentation includes: Segmenting the second image into a long-distance region and a short-distance region based on a first criterion associated with distance; and Segmenting the short-distance region into a short-distance region toward the vanishing point and a short-distance region not toward the vanishing point based on a second criterion associated with the vanishing point.
11. The positioning method according to claim 10, wherein, The sequential determination includes: Determining the rotational DOF based on the long-distance region; Determining the left-right translational DOF based on the short-distance region toward the vanishing point; and Determining the front-back translational DOF based on the short-distance region not toward the vanishing point.
12. The positioning method according to claim 1, wherein, Determining the positioning information of the device includes: Determining a direction for increasing the score based on a distribution of the pooled eigenvalues; and Correcting the candidate positioning information based on the direction.
13. The positioning method according to claim 12, wherein, The first image includes a probability distribution indicating a degree of proximity to the object, Wherein, determining the direction includes determining the direction based on the probability distribution.
14. The positioning method according to claim 12, wherein, Determining the positioning information of the device includes: Generating a corrected second image, in which the object is projected relative to the corrected candidate positioning information; and Determining a corrected score of the corrected candidate positioning information by pooling eigenvalue corresponding to vertices in the corrected second image according to the first image, Wherein, determining the direction, correcting the candidate positioning information, generating the corrected second image, and calculating the corrected score are iteratively performed until the corrected score meets the condition.
15. The positioning method according to claim 1, further comprising: Determining a virtual object on the map data to provide an augmented reality "AR" service; And Displaying the virtual object based on the determined positioning information.
16. The positioning method according to claim 15, wherein, The input image includes a driving image of a vehicle, and The virtual object indicates driving route information.
17. A non-transitory computer-readable storage medium storing instructions, which when executed by a processor cause the processor to execute the positioning method according to claim 1.
18. A positioning method, comprising: Generating a first image of an object according to an input image; Generating a second image based on map data including the position of the object to project the object relative to candidate positioning information of the device; Segmenting the second image into regions; And Determining a degree of freedom "DOF" value included in the candidate positioning information through matching between the first image and the regions, Wherein, the determination includes: Determining the rotational DOF based on long-distance vertices in the vertices included in the second image that are less affected by translational DOF than a first threshold; Determining the left-right translational DOF based on short-distance vertices facing the vanishing point among short-distance vertices other than the long-distance vertices in the second image that are less affected by the front-back translational DOF than a second threshold; and Determining the front-back translational DOF based on short-distance vertices other than the short-distance vertices facing the vanishing point among the short-distance vertices.
19. The positioning method according to claim 18, wherein, The determination includes: determining the DOF value included in the candidate positioning information by sequentially using scores calculated through the matching in the regions.
20. The positioning method according to claim 18, wherein, The determination includes: While changing the DOF value determined for the region, calculating a score corresponding to the changed DOF value by pooling eigenvalue corresponding to vertices in the region according to the first image; and Selecting the DOF value corresponding to the highest score.
21. The positioning method according to claim 18, wherein, Multiple DOF values include: Three translational DOF values; and Three rotational DOF values.
22. The positioning method according to claim 18, wherein, The segmentation includes: Segmenting the second image into a long-distance region and a short-distance region based on a first criterion associated with distance; and Segmenting the short-distance region into a short-distance region facing the vanishing point and a short-distance region not facing the vanishing point based on a second criterion associated with the vanishing point.
23. The positioning method according to claim 22, wherein, The determination includes: Determining the rotational DOF based on the long-distance region; Determine the left - right translation DOF based on the short - distance region towards the vanishing point; and Determine the front - back translation DOF based on the short - distance region not towards the vanishing point.
24. The positioning method according to claim 18, further comprising: Determine a virtual object on the map data to provide an augmented reality "AR" service; And Display the virtual object based on the determined DOF values.
25. The positioning method according to claim 24, wherein, The input image includes a driving image of a vehicle, and The virtual object indicates driving route information.
26. A non - transitory computer - readable storage medium storing instructions that, when executed by a processor, cause the processor to execute the positioning method according to claim 18.
27. A positioning device, comprising: A processor configured to: Generate a first image of an object according to an input image, Generate a second image based on map data including the position of the object to project candidate positioning information of the object relative to the device, Pool feature values corresponding to vertices in the second image according to the first image, Determine a score of the candidate positioning information based on the pooled feature values, and Determine the positioning information of the device based on the score of the candidate positioning information, Wherein, the processor is further configured to: Determine the rotation DOF based on long - distance vertices in the vertices included in the second image, whose influence by the translational degree of freedom DOF is lower than a first threshold; Determine the left - right translation DOF based on short - distance vertices towards the vanishing point in the short - distance vertices other than the long - distance vertices in the second image, whose influence by the front - back translation DOF is lower than a second threshold; and Determine the front - back translation DOF based on short - distance vertices not towards the vanishing point in the short - distance vertices other than the short - distance vertices towards the vanishing point.
28. A positioning device, comprising: A processor configured to: Generate a first image of an object according to an input image, Generate a second image based on map data including the position of the object to project candidate positioning information of the object relative to the device, Segment the second image into regions, and Determine the degree of freedom "DOF" values included in the candidate positioning information through the matching between the first image and the regions, Wherein, the processor is further configured to: Determine the rotation DOF based on long - distance vertices in the vertices included in the second image, whose influence by the translational DOF is lower than a first threshold; Determine the left - right translation DOF based on short - distance vertices towards the vanishing point in the short - distance vertices other than the long - distance vertices in the second image, whose influence by the front - back translation DOF is lower than a second threshold; and Determine the front - back translation DOF based on short - distance vertices not towards the vanishing point in the short - distance vertices other than the short - distance vertices towards the vanishing point.
29. A positioning device, comprising: A sensor disposed on the device and configured to sense one or more of the image of the device and candidate positioning information; A processor configured to: Generate a first image of an object according to the image, Generate a second image based on map data including the position of the object to project the object relative to the candidate positioning information. Determine a score of the candidate positioning information based on pooling eigenvalues corresponding to vertices in the second image according to the first image, and Determine the positioning information of the device based on the score; And A head-up display "HUD", configured to visualize a virtual object on the map data based on the determined positioning information, Wherein the processor is further configured to: Determine the rotational degrees of freedom "DOF" based on long-distance vertices in the vertices included in the second image that are less affected by the translational degrees of freedom "DOF" than a first threshold; Determine the left-right translational DOF based on short-distance vertices towards the vanishing point in the short-distance vertices other than the long-distance vertices in the second image that are less affected by the front-back translational DOF than a second threshold; And Determine the front-back translational DOF based on short-distance vertices other than the short-distance vertices towards the vanishing point in the short-distance vertices.
30. The positioning device according to claim 29, wherein, The processor is further configured to: Segment the second image into a long-distance region and a short-distance region based on distance; and Segment the short-distance region into a short-distance region towards the vanishing point and a short-distance region not towards the vanishing point based on the vanishing point.
31. The positioning device according to claim 30, wherein, The processor is further configured to: Determine the rotational degrees of freedom "DOF" based on the long-distance region; Determine the left-right translational DOF based on the short-distance region towards the vanishing point; And Determine the front-back translational DOF based on the short-distance region not towards the vanishing point.
32. The positioning device according to claim 29, wherein, The processor is further configured to: use a neural network to generate the first image including a feature map corresponding to multiple features.
33. The positioning device according to claim 29, wherein, The second image includes a projection of two-dimensional "2D" vertices corresponding to the object.
34. The positioning device according to claim 29 further comprises: A memory, configured to store the map data, the image, the first image, the second image, the score, and instructions that, when executed, configure the processor to determine any one or any combination of: the determined positioning information and the virtual object.
Citation Information
Patent Citations
Non-slip pavement forming device of grooved structure
KR1020180127589A
Method and system for vehicle localization from camera image
CN108256411A
System and method for image based vehicle localization
CN108450034A