Method and apparatus for position estimation
By calculating edge direction information and determining inner point pixels in the vehicle navigation system, estimating the rotation reference point and correcting the positioning information, the problem of difficulty in correcting the positioning information error in the prior art is solved, and the accuracy of the navigation system is improved.
Patent Information
- Application Number
- CN202010168378.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-03-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-03-11
AI Technical Summary
The prior art is difficult to accurately estimate the rotation reference point of a vehicle in a vehicle navigation system, resulting in difficult correction of position errors and rotation errors of positioning information.
By calculating edge direction information of edge component pixels in the input image, the inner point pixel is determined, and the rotation reference point is estimated based on the inner point pixel, and the positioning information is finally corrected.
It improves the accuracy of vehicle positioning information, reduces position error and rotation error, and enhances the accuracy of the navigation system.
Smart Images

Figure CN112539753B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of Korean Patent Application No. 10 - 2019 - 0115939, filed on September 20, 2019, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical field
[0003] The following description relates to a technology for position estimation. Background art
[0004] When a vehicle or other object is moving, the navigation system of the vehicle receives radio waves from satellites belonging to multiple Global Positioning Systems (GPS), and verifies the current position and speed of the vehicle. The navigation system can calculate the three - dimensional (3D) current position of the vehicle including latitude, longitude, and altitude information based on the information received from the GPS receiver. Summary of the invention
[0005] This summary of the invention is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.
[0006] In one general aspect, a processor - implemented method includes: calculating edge - direction information of edge - component pixels extracted from an input image; determining in - point pixels among the edge - component pixels based on the edge - direction information and virtual horizontal - line information; estimating a rotation reference point based on the in - point pixels; and correcting the positioning information of the device based on the estimated rotation reference point.
[0007] Calculating the edge - direction information may include: extracting edge - component pixels by pre - processing the input image.
[0008] The extraction of the edge - component pixels may include: calculating gradient information for each pixel included in the input image; and extracting selected pixels having gradient information exceeding a threshold among the pixels included in the input image as the edge - component pixels.
[0009] Calculating the edge - direction information may include: masking a part of the edge - component pixels.
[0010] Masking a part of the edge - component pixels may include: excluding pixels on the boundary from the edge - component pixels.
[0011] Excluding pixels on the boundary may include: excluding pixels corresponding to the area above the virtual horizontal line in the input image.
[0012] Excluding pixels on the boundary may also include: determining a virtual horizontal line in the input image based on the roll parameter and the pitch parameter of the camera sensor that captures the input image.
[0013] Calculating edge direction information may include: determining the edge direction of each edge component pixel based on an estimation of the edge direction.
[0014] Determining the edge direction may include: using a convolutional neural network (CNN) to calculate the convolution operation value of each edge component pixel in the edge component pixels, the CNN including corresponding kernels for each predetermined angle; and determining the angle value of the kernel corresponding to the highest convolution operation value among the convolution operation values calculated for each edge component pixel in the edge component pixels as the edge direction information.
[0015] Determining inlier pixels may include: retaining inlier pixels by removing outlier pixels in the edge component pixels based on the intersection points between the virtual horizontal line and the edge line corresponding to the edge direction information of the edge component pixels.
[0016] Retaining inlier pixels may include: generating inlier data with outlier pixels removed based on an attention neural network.
[0017] Retaining outlier pixels may include: removing outlier pixels based on the proximity level between intersection points.
[0018] Removing outlier pixels may include: for each edge component pixel in the edge component pixels, calculating a reference vector from an intersection point between an edge line on one side of the edge line and the virtual horizontal line in the edge line to a reference point in the input image; and calculating the similarity between the reference vectors based on the proximity level between the intersection points.
[0019] Removing outlier pixels may also include: removing outlier pixels by applying the calculated similarity to a vector value obtained by transforming the edge direction information and the image coordinates of an edge component pixel in the edge component pixels.
[0020] Estimating a rotation reference point may include: performing line fitting on the inlier pixels; and estimating the rotation reference point based on the edge line according to the line fitting.
[0021] Performing line fitting may include: converting an image including a plurality of inlier pixels into a bird's-eye view image; clustering the plurality of inlier pixels into at least one straight line group based on lines corresponding to the plurality of inlier pixels in the bird's-eye view image; and determining the edge line as the edge line representing each straight line group in the at least one straight line group.
[0022] Estimating a rotation reference point may include: determining a point with the shortest distance from the edge line corresponding to the inlier pixels in the input image as the rotation reference point.
[0023] The method may further include: performing visualization in a display for a user by mapping a virtual content object to display coordinates determined based on the final positioning information of the device. Calibrating the positioning information may include: calibrating any one or both of the position error and the rotation error of the positioning information based on the estimated rotation reference point; and generating the final positioning information based on the result of calibrating any one or both of the position error and the rotation error.
[0024] The method may further include: obtaining positioning information using a Global Navigation Satellite System (GNSS) module and an Inertial Measurement Unit (IMU); and acquiring an input image using an image sensor.
[0025] In another general aspect, there is provided a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the above method.
[0026] In another general aspect, a device includes: a sensor configured to acquire an input image; and a processor configured to: calculate edge direction information of edge component pixels extracted from the input image; determine inlier pixels among the edge component pixels based on the edge direction information and virtual horizontal line information; estimate a rotation reference point based on the inlier pixels; and calibrate the positioning information of the device based on the estimated rotation reference point.
[0027] The edge component pixels may include selected pixels that are determined to have gradient information exceeding a threshold among the pixels included in the input image.
[0028] The processor may further be configured to: mask a part of the edge component pixels by excluding pixels corresponding to a region above the virtual horizontal line in the input image among the edge component pixels.
[0029] The processor may further be configured to: determine the virtual horizontal line based on the roll parameter and the pitch parameter of the sensor.
[0030] The processor may further be configured to calculate the edge direction information by: implementing a neural network to perform a convolution operation on each edge component pixel among the edge component pixels; and determining, based on the result of the convolution operation, an angle corresponding to each edge component pixel among the edge component pixels as the edge direction information.
[0031] The processor may further be configured to: determine the inlier pixels by removing outlier pixels among the edge component pixels based on the proximity of intersections between the virtual horizontal line and edge lines corresponding to the edge direction information of the edge component pixels.
[0032] Removing outlier pixels may include: for each edge component pixel among the edge component pixels, calculating a reference vector from an intersection point between an edge line on one side of the edge line in the intersection points and a virtual horizontal line to a reference point in the input image; calculating a similarity between the reference vectors based on the proximity level between the intersection points; and removing outlier pixels by applying the calculated similarity to a vector value obtained by transforming the edge direction information and the image coordinates of an edge component pixel among the edge component pixels.
[0033] The processor may also be configured to: determine inlier pixels by removing outlier pixels in the edge component pixels based on the intersection points between the virtual horizontal line and the edge lines corresponding to the edge direction information of the edge component pixels.
[0034] The processor may also be configured to determine inlier pixels by: removing outlier pixels in the edge component pixels based on the intersection points between the virtual horizontal line and the edge lines corresponding to the edge direction information of the edge component pixels; and generating inlier pixels with outlier pixels removed based on an attention neural network.
[0035] The processor may also be configured to estimate a rotation reference point by: performing line fitting on the inlier pixels in the bird's-eye view image based on the input image; and estimating the rotation reference point based on the edge line according to the line fitting.
[0036] The processor may also be configured to estimate a rotation reference point by determining a point with the shortest distance to the edge line corresponding to the inlier pixels in the input image as the rotation reference point.
[0037] The processor may also be configured to: perform visualization in the user's display for transmitting an external scene by mapping a virtual content object to display coordinates determined based on the final positioning information of the device. The processor may also be configured to: correct the positioning information by correcting any one or both of the position error and the rotation error of the positioning information based on the estimated rotation reference point, and generate final positioning information based on the result of correcting any one or both of the position error and the rotation error.
[0038] The processor may also be configured to: obtain positioning information from a Global Navigation Satellite System (GNSS) module and an Inertial Measurement Unit (IMU).
[0039] In another general aspect, an augmented reality (AR) device includes: a display configured to provide a view of an external environment; a sensor configured to acquire an input image; and a processor configured to: calculate edge direction information of edge component pixels extracted from the input image; determine inlier pixels among the edge component pixels based on the edge direction information and virtual horizontal line information; estimate a rotation reference point based on the inlier pixels; generate final positioning information by correcting initial positioning information of the device based on the estimated rotation reference point; and perform visualization by mapping a virtual content object to coordinates determined based on the final positioning information in the display.
[0040] The sensor may include a global navigation satellite system (GNSS) module and an inertial measurement unit (IMU). The initial positioning information may include position information and attitude information obtained from the GNSS module and the IMU.
[0041] The processor may also be configured to: determine an initial rotation reference point based on map data and the initial positioning information. Correcting the initial positioning information may include: applying a rotation parameter such that the initial rotation reference point matches the estimated rotation reference point.
[0042] The processor may also be configured to calculate the edge direction information by: implementing a neural network to perform a convolution operation on each edge component pixel among the edge component pixels; and based on the result of the convolution operation, determining an angle corresponding to each edge component pixel among the edge component pixels as the edge direction information.
[0043] The processor may also be configured to: estimate the rotation reference point by determining a point with the shortest distance to an edge line corresponding to an inlier pixel in the input image as the rotation reference point.
[0044] The processor may also be configured to estimate the rotation reference point by: performing line fitting on inlier pixels in a bird's-eye view image based on the input image; and based on the line fitting, estimating the rotation reference point based on the edge line.
[0045] The processor may also be configured to: determine inlier pixels by removing outlier pixels among the edge component pixels based on proximity between intersections between the virtual horizontal line and the edge line corresponding to the edge direction information of the edge component pixels.
[0046] The processor may also be configured to determine inlier pixels by: removing outlier pixels among the edge component pixels based on intersections between the virtual horizontal line and the edge line corresponding to the edge direction information of the edge component pixels; and generating inlier pixels with outlier pixels removed based on an attention neural network.
[0047] The display may include a windshield of a vehicle.
[0048] Other features and aspects will become apparent from the detailed description, the drawings, and the claims. Description of the Drawings
[0049] Figure 1 is a block diagram showing an example of a position estimation device.
[0050] Figures 2 to 4 shows an example of a position estimation error.
[0051] Figure 5 is a flowchart showing an example of a position estimation method.
[0052] Figure 6 shows an example of a position estimation process.
[0053] Figure 7 shows an example of a preprocessing operation in the position estimation process.
[0054] Figure 8 shows an example of virtual horizontal line information.
[0055] Figure 9 shows an example of local filtering.
[0056] Figure 10 shows an example of non-local filtering.
[0057] Figure 11 shows an example of an attention neural network for excluding outlier pixels.
[0058] Figure 12 shows an example of line fitting.
[0059] Figure 13 shows an example of a rotational reference point error between the initial positioning information estimated by the position estimation device and the actual attitude of the position estimation device.
[0060] Figure 14 shows an example of the operation of a position estimation device mounted on a vehicle.
[0061] In all of the drawings and the detailed description, like reference numerals refer to like elements, features, and structures. The drawings may not be drawn to scale, and the relative dimensions, proportions, and depictions of elements in the drawings may be enlarged for clarity, illustration, and convenience. Detailed Description
[0062] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely illustrative and is not limited to those set forth herein, but can be significantly changed after understanding the disclosure of this application, except for operations that must be performed in a certain order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0063] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of this application.
[0064] In this document, it should be noted that the use of the term "may" with respect to an example or embodiment (e.g., what may be included or implemented with respect to the example or embodiment) means that there is at least one example or embodiment in which such a feature is included or implemented, but not all examples and embodiments are limited thereto. Throughout the specification, when an element such as a layer, region, or substrate is described as being "on," "connected to," or "coupled to" another element, it may be directly "on," "connected to," or "coupled to" the other element, or there may be one or more other elements therebetween. Conversely, when an element is described as being "directly on," "directly connected to," or "directly coupled to" another element, there may be no other elements therebetween.
[0065] As used herein, the term "and / or" includes any one of the associated listed items, or any combination of any two or more of the listed items.
[0066] Although terms such as "first," "second," and "third" may be used herein to describe various components, elements, regions, layers, or portions, these components, elements, regions, layers, or portions are not limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or portion from another. Thus, a first component, element, region, layer, or portion referred to in the examples described herein may also be referred to as a second component, element, region, layer, or portion without departing from the teachings of the examples.
[0067] The terms used herein are for the purpose of describing various examples only and are not intended to limit the disclosure. The articles "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. The terms "comprising", "including", and "having" specify the presence of the stated features, numbers, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, components, elements, and / or combinations thereof.
[0068] Unless otherwise defined, all terms, including technical and scientific terms, used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding the disclosure of this application. Terms (such as those defined in a common dictionary) should be interpreted as having a meaning consistent with their meaning in the relevant art and the disclosure of this application, and should not be interpreted in an idealized or overly formal sense unless clearly so defined herein.
[0069] After understanding the disclosure of this application, it will be apparent that the features of the examples described herein can be combined in various ways. In addition, although the examples described herein have various configurations, other configurations can become apparent after understanding the disclosure of this application.
[0070] For example, a global positioning system (GPS) signal for estimating the position of a target (e.g., a moving object such as a vehicle) may include a GPS position error of about 10 meters (m) to 100 m. According to the example embodiments disclosed herein, such a position error can be corrected based on other information.
[0071] Figure 1 is a block diagram showing an example of a position estimation device.
[0072] Referring Figure 1 , the position estimation device 100 may include, for example, a sensor 110, a processor 120, and a memory 130.
[0073] Sensor 110 generates sensed data. For example, sensor 110 generates sensed data by sensing information for estimating the position of a target. The target is an object whose attitude and position are estimated by the position estimation device 100, for example. In an example, the target may be the position estimation device 100. In another example, the position estimation device 100 may be mounted on a vehicle, and the target may be the vehicle. The information for estimating the position includes various signals, such as global navigation satellite system (GNSS) signals (e.g., global positioning system (GPS) signals), acceleration signals, speed signals, and image signals. Processor 120 measures initial positioning information mainly based on the sensed data. The initial positioning information includes information indicating the position and attitude of the target measured roughly. For example, sensor 110 may include an inertial measurement unit (IMU) and a GNSS module. In such an example, sensor 110 acquires IMU signals indicating the acceleration and angular velocity of the target and GNSS signals as the sensed data.
[0074] The IMU is also referred to as an "inertial measurer". The IMU measures changes in attitude, the rate of change of position movement, and displacement. The IMU includes a three-axis accelerometer that senses translational motion (e.g., acceleration) and a three-axis gyroscope that senses rotational motion (e.g., angular velocity). Since the IMU does not rely on external information, acceleration signals and angular velocity signals can often be collected stably. The GNSS module can receive signals sent from at least three artificial satellites to calculate the positions of the satellites and the position of the position estimation device 100, and can also be referred to as GNSS. The GNSS module can measure the absolute position at 1 hertz (Hz) and can operate stably due to the relatively low noise of the GNSS module. The IMU can measure the relative position at 100 Hz and perform measurements at high speed.
[0075] In addition, sensor 110 senses image data. For example, sensor 110 may include a camera sensor. The camera sensor receives rays in the visible light range and senses the intensity of light. The camera sensor generates a color image by sensing the intensity of light corresponding to each of the red channel, blue channel, and green channel. However, the wavelength band that can be sensed by sensor 110 is not limited to the visible light band. Sensor 110 may also or alternatively include, for example, an infrared sensor that senses rays in the infrared band, a lidar sensor, a RADAR sensor that senses signals in the electromagnetic band, a thermal image sensor, or a depth sensor. The camera sensor generates image data by capturing an external view (e.g., a front view) from the position estimation device 100.
[0076] The processor 120 can generate the final positioning information of the target by correcting the initial positioning information of the target. The processor 120 can estimate the initial positioning information and the rotation reference point to correct the positioning error (e.g., rotation error and position error) between the actual position and attitude of the target. Herein, the rotation reference point is a point used as a reference for correcting the rotation error in the positioning information estimated for the target. The positioning information can include information about the position, attitude, and movement (e.g., speed and acceleration) of the target. The rotation reference point can be used in positioning, and the rotation reference point can be used to correct positioning based on map data, for example.
[0077] For example, the first rotation reference point can be a three-dimensional (3D) point corresponding to the vanishing point on a two-dimensional (2D) image plane, and can indicate a point on the lane boundary line that is at a sufficient distance from the image sensor. The processor 120 can determine the first rotation reference point based on the map data and positioning information estimated for the target. The second rotation reference point can be a rotation reference point estimated based on the image data sensed by the sensor 110. As a type of vanishing point, the second rotation reference point can be a point where at least two parallel lines (e.g., lane boundary lines) converge in 3D space when projected onto the 2D image plane. The processor 120 can estimate the second rotation reference point according to the image data associated with the front view of the target.
[0078] In Figures 5 to 11 the following description, the first rotation reference point is referred to as the "initial reference point", and the second rotation reference point is referred to as the "rotation reference point".
[0079] Reference Figure 1 , the processor 120 corrects the positioning information (e.g., initial positioning information) based on the rotation parameter, which is calculated based on the first rotation reference point and the second rotation reference point. The rotation parameter is a parameter used to compensate for the error between the initially estimated initial positioning information of the target and the actual position and attitude, and can be, for example, a parameter that matches the first rotation reference point and the second rotation reference point. For example, the camera sensor included in the sensor 110 can be attached to the target (e.g., a vehicle) such that the optical axis of the camera sensor is parallel to the longitudinal axis of the target. In this example, since the camera sensor moves the same as the target, the image data sensed by the camera sensor includes visual information that matches the position and attitude of the target. Therefore, the processor 120 can apply the rotation parameter so that the first rotation reference point corresponding to the initial positioning information matches the second rotation reference point detected from the image data.
[0080] Memory 130 stores data for performing the position estimation method temporarily or permanently. For example, memory 130 stores a convolutional neural network (CNN) trained for local filtering and an attention neural network trained for non-local filtering based on image data. However, the present disclosure is not limited to memory 130 storing a CNN and an attention neural network. For example, other types of neural networks can be stored on memory 130 to perform local filtering and non-local filtering, such as a recurrent neural network (RNN), a deep belief network, a fully connected network, a bidirectional neural network, a restricted Boltzmann machine, or a neural network including different or overlapping neural network parts having fully connected, convolutional, recurrent, and / or bidirectional connections respectively.
[0081] As described above, the position estimation device 100 estimates initial positioning information including an initial position and an initial attitude of a target based on the sensed data sensed by the sensor 110. The initial positioning information may include an attitude error and a position error. Errors in the initial positioning information will be described in more detail with reference to Figures 2 to 4 Errors in the initial positioning information will be described in more detail.
[0082] Figures 2 to 4 An example of the position estimation error is shown.
[0083] The position estimation device according to an example embodiment (e.g., the position estimation device 100) measures the initial positioning information based on, for example, GNSS signals and acceleration signals. Figure 2 The case where the initial attitude of the initial positioning information has a rotational error Δθ with respect to the actual attitude is shown. The initial positioning information indicates that the target including the position estimation device has an arbitrary initial position and an arbitrary initial attitude. Based on the map data, for example, a high-definition (HD) map, the map lane boundary line 210 corresponding to the initial positioning information is as Figure 2 shown. The image lane boundary line 220 appearing in the image data actually captured by the image sensor may have a different angle from the map lane boundary line 210. Therefore, the position estimation device can compensate for the above rotational error based on the rotation parameter.
[0084] The map data includes information associated with the map and includes, for example, information about landmarks. A landmark may be, for example, an object fixed at a predetermined geographical location to provide information for a driver to drive a vehicle on a road. For example, road signs and traffic lights may belong to landmarks. For example, according to the Road Traffic Act of Korea, landmarks are classified into six categories, such as warning signs, regulatory signs, guiding signs, auxiliary signs, road signs, and signal signs. However, the classification of landmarks is not limited to the above description. For example, for each country, the categories of landmarks may be different.
[0085] In addition, in this document, a lane boundary line (e.g., a line corresponding to the map lane boundary line 210 or the image lane boundary line 220) is a line for defining a lane. The lane boundary line can be, for example, a solid line or a dashed line marked on the road surface, or can be a curb arranged along the outer edge of the road.
[0086] Figure 3 An example situation is shown where the initial position in the initial positioning information has a horizontal position error Δx with respect to the actual position. Based on the map data, the map lane boundary line 310 corresponding to the initial positioning information is as Figure 3 shown. The image lane boundary line 320 appearing in the image data actually captured by the image sensor may be located at a different position from the map lane boundary line 310.
[0087] Figure 4 An example situation is shown where the initial attitude in the initial positioning information has a rotational error 420Δθ with respect to the actual attitude about the transverse axis y . Based on the map data, the first rotational reference point 412 corresponding to the map lane boundary line is as Figure 4 shown, where the map lane boundary line corresponds to the initial positioning information. The second rotational reference point 411 corresponding to the image lane boundary line appearing in the image data actually captured by the image sensor corresponds to an attitude different from the attitude of the first rotational reference point 412. Therefore, the position estimation device (e.g., the position estimation device 100) can compensate for the rotational error 420Δθ based on the rotational parameters y .
[0088] Figure 5 is a flowchart showing an example of a position estimation method.
[0089] Referring to Figure 5 , in operation 510, a position estimation device (e.g., the position estimation device 100) according to an example embodiment calculates edge direction information of a plurality of edge component pixels extracted from an input image. An edge component pixel is a pixel indicating an edge component in the input image. Edge component pixels will be further described below with reference to Figure 7 . The edge direction information is information indicating the direction (e.g., a 2D direction in the image) faced by the edge represented by the edge component pixels. Edge direction information will be further described below with reference to Figure 9Further describe the example edge direction information. Although the input image described in this document is a grayscale image obtained by converting a color image, the present disclosure is not limited to such an example. The input image can be, for example, a color image. When the position estimation device is included in a vehicle, the input image can be a driving image representing a front view from the vehicle while the vehicle is driving. For example, when the input image is an RGB image with 256 horizontal pixels and 256 vertical pixels, the total value of 256×256×3 pixels can be used as input data.
[0090] In operation 520, the position estimation device determines inlier pixels among a plurality of edge component pixels based on the edge direction information and the virtual horizontal line information. The virtual horizontal line information is line information corresponding to a horizontal line within the angle of the field (or, field of view) of the camera sensor as prior information. The virtual horizontal line can be a virtual horizon, a virtual horizontal line including or depending on the determined estimated vanishing point, or otherwise determined or preset (e.g., depending on the attitude of the captured image sensor) virtual line. This will be further described below with reference to Figure 8 Further describe the virtual horizontal line information. Inlier pixels are pixels among the plurality of edge component pixels for estimating the rotation reference point, and outlier pixels are one of the pixels other than the inlier pixels among the plurality of edge component pixels. For example, the position estimation device can calculate the intersection point between the line based on the edge direction information and the line based on the virtual horizontal line information at each position among the positions of the plurality of edge component pixels, and can determine the pixels forming the intersection point at adjacent positions as inlier pixels. This will be further described below with reference to Figure 10 and Figure 11 Further describe an example of determining inlier pixels.
[0091] In operation 530, the position estimation device estimates the rotation reference point based on the inlier pixels. In an example, the position estimation device can determine the point in the image with the shortest straight-line distance from the line corresponding to the inlier pixels as the above-mentioned rotation reference point. In another example, the position estimation device can determine the average point of the intersection points between the virtual horizontal line and the line based on the edge direction information of the inlier pixels as the rotation reference point. In another example, the position estimation device can determine the average point of the intersection points between the lines based on the edge direction information of the inlier pixels as the rotation reference point. However, the present disclosure is not limited to the foregoing examples, and the rotation reference point can be determined by various schemes. The rotation reference point can be represented by 2D coordinates (r x , r y ), and the values of r x and r y can be within the range of [0, 1] as relative coordinate values with respect to the image size.
[0092] At operation 540, the position estimation device corrects the positioning information of the position estimation device based on the estimated rotation reference point. For example, the position estimation device may estimate initial positioning information based on GNSS signals, IMU signals, and map data, and may also estimate corresponding initial reference points. The position estimation device may estimate the positioning error (e.g., rotation error and position error) between the initial reference point and the rotation reference point. The error between the initial reference point and the rotation reference point may correspond to the error between the initial attitude and the actual attitude. Therefore, the position estimation device may correct the initial positioning information by rotating and / or moving the initial positioning information according to the positioning error, and may generate final positioning information. Therefore, the position estimation device may determine a more accurate attitude based on the estimated initial attitude that is distorted from the actual attitude.
[0093] Figure 6 An example of the position estimation process is shown.
[0094] Refer to Figure 6 , in operation 610, the position estimation device (e.g., the position estimation device 100) according to an example embodiment preprocesses the input image. The position estimation device may preprocess the input image to extract edge component pixels. In operation 620, the position estimation device performs local filtering on the edge component pixels to determine edge direction information for each of the edge component pixels among the edge component pixels. In operation 630, the position estimation device performs non-local filtering on the edge component pixels to exclude outlier pixels and retain inlier pixels. Although not limited thereto, the examples of preprocessing, local filtering, and non-local filtering will be further described with reference to Figure 7 , Figure 9 and Figure 10 respectively.
[0095] Figure 7 An example of the preprocessing operation in the position estimation process is shown.
[0096] Refer to Figure 7 , the position estimation device (e.g., the position estimation device 100) according to an example embodiment acquires the input image 701. For example, the position estimation device may be installed on a vehicle and may include a camera sensor installed to have a field of view angle to show the front view of the vehicle. The camera sensor may capture a scene corresponding to the field of view angle and acquire the input image 701.
[0097] As Figure 7 shown, Figure 6 the operation 610 in Figure 6 may include operation 711 and operation 712. In Figure 6 's operation 610, the position estimation device extracts a plurality of edge component pixels by preprocessing the input image 701. For example, as Figure 7As shown in the figure, in operation 711, the position estimation device calculates the gradient information of each pixel based on the input image 701. The gradient information is a value indicating, for example, the intensity change or direction change of the color. The gradient information at any point in the image can be represented as a 2D vector having differences in the vertical and horizontal directions. The position estimation device can calculate the gradient value in the horizontal direction suitable for estimating the lane boundary line. The calculated gradient value can be a predetermined width value regardless of the width of the lane boundary line.
[0098] The position estimation device extracts edge component pixels having the calculated gradient information exceeding a predefined threshold from the pixels included in the input image 701. For example, the position estimation device calculates the magnitude of the gradient vector (e.g., a 2D vector having differences) as the gradient information of each pixel in the input image 701, and determines the pixels having a gradient vector exceeding the threshold as edge component pixels. Thus, the position estimation device can extract the pixels whose intensity value or color value changes rapidly compared to adjacent pixels as edge component pixels. The image including a plurality of edge component pixels is the edge component image 703.
[0099] In operation 712, the position estimation device masks a part of the plurality of edge component pixels. For example, the position estimation device can exclude the pixels on the boundary from the plurality of edge component pixels. The position estimation device can exclude the pixels corresponding to the area above the virtual horizontal line 702 in the input image 701. Thus, the position estimation device can obtain the information focusing on the road area. The virtual horizontal line 702 will be further described below. Although the virtual horizontal line 702 is shown as a boundary in Figure 8 the figure, the present disclosure is not limited to this example. Figure 7 the figure, the present disclosure is not limited to this example.
[0100] The plurality of edge component pixels extracted in operation 711 may indicate the lane boundary line and the lines along the lane (e.g., the curb), or may indicate the lines unrelated to the lane (e.g., the outline of the buildings around the road). Since the rotation reference point 790 as the vanishing point is the point where the lines along the lane (e.g., the lane boundary line) converge in the image, the lines perpendicular to the road (such as the outline of the buildings) may not be highly relevant to the vanishing point. Therefore, as described above, the position estimation device can remove the pixels having no determined or estimated correlation with the vanishing point or having a selectively reduced correlation by excluding the edge component pixels above the road area in operation 712. The image from which the pixels corresponding to such an area above the road area are removed is the masked image 704.
[0101] The position estimation device generates an edge direction image 705 by local filtering in operation 620 and generates an inlier image 706 by non-local filtering in operation 630. The position estimation device estimates a rotation reference point 790 based on the inlier image 706 and obtains a result 707 mapped with the rotation reference point 790.
[0102] Unless otherwise described, the above operations are not necessarily executed in the order shown in the drawings, but may be executed in parallel or in a different order. In an example, the position estimation device may calculate a virtual horizontal line while calculating a gradient value. In another example, the position estimation device may extract edge component pixels after performing masking based on the virtual horizontal line.
[0103] Figure 8 An example of virtual horizontal line information is shown.
[0104] Reference Figure 8 above, the Figure 7 virtual horizontal line information indicates a straight line where a horizontal line virtually exists in the image captured by the camera sensor 810. The virtual horizontal line information may be prior information and may be determined in advance by camera parameters (e.g., camera calibration information). For example, the position estimation device may determine the virtual horizontal line in the input image based on the roll parameter and pitch parameter of the camera sensor 810 that captures the input image.
[0105] In Figure 8 , the x-axis and z-axis are parallel to the ground, and the y-axis is defined as the axis perpendicular to the ground. In the image 801 captured when the central axis of the camera sensor 810 is parallel to the z-axis, the virtual horizontal line appears as a straight line intersecting the central part of the image 801, as shown in Figure 8 . The central axis of the camera sensor 810 may be referred to as the "optical axis". When assuming that the central axis of the camera sensor 810 is parallel to the z-axis, the roll axis 811 and pitch axis 812 of the camera sensor 810 may be the z-axis and x-axis respectively. The roll parameter of the camera sensor 810 may be determined by the roll rotation 820, and the pitch parameter of the camera sensor 810 may be determined by the pitch rotation 830. For example, the roll parameter may indicate the angle by which the camera sensor 810 rotates around the roll axis 811 relative to the ground, and the virtual horizontal line may appear as a straight line tilted according to the roll parameter in the image 802. The pitch parameter may indicate the angle by which the camera sensor 810 rotates around the pitch axis 812 relative to the ground, and the virtual horizontal line may appear as a straight line at a position vertically shifted according to the pitch parameter in the image 803.
[0106] Figure 9 An example of local filtering is shown.
[0107] Reference Figure 9, the position estimation device (e.g., the position estimation device 100) according to the example embodiment calculates edge direction information for each edge component pixel among a plurality of edge component pixels in the masked image 704 by performing local filtering in operation 620 of Figure 6 . The edge direction information indicates the edge direction indicated by the edge component pixel and represents the direction of the edge in the image. For example, the edge direction information can be distinguished by the value of the angle formed with an arbitrary reference axis (e.g., the horizontal axis) in the input image, however, the present disclosure is not limited to this example.
[0108] The position estimation device determines the edge direction for each edge component pixel among the plurality of edge component pixels based on the edge direction estimation model. The edge direction estimation model is a model for outputting the edge direction of each edge component pixel from the edge component image 703, can have, for example, a machine learning structure, and can include, for example, a neural network. In Figure 9 , the example edge direction estimation model is implemented as a CNN 920 including a plurality of kernels 925. For example, the CNN 920 can include a convolutional layer connected by the plurality of kernels 925. The neural network can extract the edge direction from the image by mapping the input data and the output data in a non-linear relationship based on deep learning. This example of deep learning can be a machine learning scheme configured to solve the edge direction extraction problem from a large dataset and can be used to map the input data and the output data by supervised learning or unsupervised learning. Therefore, the CNN 920 can be a network pre-trained based on training data to output the edge direction information corresponding to each edge component pixel from each edge component pixel.
[0109] For example, the position estimation device calculates the convolution operation value of each edge component pixel among the plurality of edge component pixels based on the kernel for each preset angle in the CNN 920. Each kernel among the plurality of kernels 925 can be a filter in which element values are set, and the element values are used to extract image information along a predetermined edge direction or with respect to a predetermined edge direction through the corresponding convolution operation and input information. For example, the kernel 931 among the plurality of kernels 925 can be a filter including element values configured to extract image information along an edge direction at 45 degrees with respect to the horizontal axis of the image or with respect to this edge direction. The position estimation device can perform a convolution operation on the pixel at an arbitrary position in the image and the adjacent pixels based on the element values of the filter. For example, the convolution operation value of an arbitrary pixel calculated based on the kernel for extracting image information along a predetermined edge direction (e.g., the direction corresponding to 45 degrees) or with respect to a predetermined edge direction can be a value (e.g., probability) indicating the possibility that the pixel is the predetermined edge direction.
[0110] The position estimation device can calculate the convolution operation value for each pixel included in the pixels swept for one core while sweeping each pixel in the pixels. In addition, the position estimation device can calculate the convolution operation value of the entire image for each of the multiple cores 925. For example, the edge component image 703 can be an image including 256 horizontal pixels and 256 vertical pixels, and the CNN 920 can include 180 cores. Although the 180 cores are filters for extracting the edge direction corresponding to the angle samples (such as 1 degree, 2 degrees, 179 degrees, or 180 degrees) degree by degree from 0 degrees to 180 degrees, the angles that can be extracted by the cores are not limited to these examples. The position estimation device can calculate the convolution operation value for each pixel among 256×256 pixels for each core. In addition, the position estimation device can calculate the convolution operation value for each pixel among 256×256 pixels for each of the 180 cores. Therefore, the output data of the CNN 920 can include tensor data of size 256×256×180. In the described example, each element of the tensor data indicates the edge probability information of 180 angle samples at the sample position of the 2D image with 256×256 pixels.
[0111] For example, the value represented by the tensor coordinates (x, y, θ) in the tensor data can be a value obtained by applying the core for extracting the edge direction of the angle θ to the coordinates (x, y) of the edge component image 703 through convolution operation. Therefore, when the edge direction of the angle indicated by the core is detected at each 2D position of the image, the element of the tensor data can be a non-zero value (for example, as the probability of the corresponding angle), otherwise it can be a "0" value.
[0112] The position estimation device determines the angle value indicated by the core corresponding to the highest convolution operation value among the multiple convolution operation values calculated for each edge component pixel as the edge direction information. As described above, each element value of the tensor data can indicate the probability that the pixel at the predetermined coordinate in the image is the edge direction of the predetermined angle. Therefore, the highest convolution operation value among the multiple convolution operation values calculated for each edge component pixel can correspond to the convolution operation value based on the core with the highest probability as the angle of the corresponding edge component pixel. Therefore, the position estimation device can estimate the edge direction with the highest possibility as the angle of the corresponding edge component pixel for each edge component pixel.
[0113] Figure 10 An example of non-local filtering is shown.
[0114] Reference Figure 10, the position estimation device according to the embodiment (e.g., the position estimation device 100) determines the inlier pixels in the edge direction image 705 by performing non-local filtering in operation 630. For example, the position estimation device can retain the inlier pixels by removing the outlier pixels from the multiple edge component pixels based on the intersections between the virtual horizontal line and the edge lines corresponding to the edge direction information of the multiple edge component pixels. The position estimation device can remove the outlier pixels based on the proximity level between the intersections. The proximity level between the intersections can be the degree or extent to which the intersections are adjacent to each other in the image and can correspond to, for example, the determined and / or compared distance between the intersections.
[0115] For example, in Figure 10 , a first straight line 1091 corresponding to the edge direction of the first edge component pixel 1021 and a virtual horizontal line 1002 can be determined to form a first intersection 1031. A second straight line 1092 corresponding to the edge direction of the second edge component pixel 1022 and a virtual horizontal line 1002 can be determined to form a second intersection 1032. A third straight line 1093 corresponding to the edge direction of the third edge component pixel 1023 and a virtual horizontal line 1002 form a third intersection 1033. As Figure 10 shown in, the first edge component pixel 1021 and the second edge component pixel 1022 correspond to points on the lane boundary line of the road. The third edge component pixel 1023 corresponds to a point on the bridge rail. The positions of the first intersection 1031 and the second intersection 1032 are detected to be adjacent to each other, and the third intersection 1033 is detected at a relatively far position. Therefore, the position estimation device can exclude the third edge component pixel 1023 forming the third intersection 1033 as an outlier pixel. The position estimation device can retain the first edge component pixel 1021 forming the first intersection 1031 and the second edge component pixel 1022 forming the second intersection 1032 as inlier pixels.
[0116] The following refers to Figure 11 for a more detailed description of an example of using an attention neural network to determine inlier pixels and exclude outlier pixels.
[0117] Figure 11 An example of an attention neural network 1130 for excluding outlier pixels is shown.
[0118] Referring to Figure 11 , the position estimation device according to the example embodiment (e.g., the position estimation device 100) generates inlier data with outlier pixels removed from the multiple edge component pixels based on the attention neural network 1130. The attention neural network 1130 can be a network pre-trained based on training data to output the remaining inlier pixels by removing the outlier pixels from the multiple edge component pixels.
[0119] For example, the position estimation device can determine inlier pixels among a plurality of edge component pixels based on the similarity between the intersection points formed by single edge component pixels and the intersection points formed by another single edge component pixel.
[0120] The position estimation device can calculate, for each edge component pixel among the plurality of edge component pixels, a reference vector from the intersection point between the edge line and the virtual horizontal line 1102 to the reference point 1101 on the input image through the first reference vector transformation 1131 and the second reference vector transformation 1132. The reference point 1101 is designated as an arbitrary point on the input image. For example, a reference vector for expressing similarity can be defined. As Figure 11 shown, the intersection points between the virtual horizontal line 1102 and the lines corresponding to the edge directions of each edge component pixel can be formed by each edge component pixel among the edge component pixels. The reference vector corresponding to any edge component pixel can be a vector (e.g., a unit vector) from the reference point 1101 to the intersection point between the virtual horizontal line 1102 and the line corresponding to the edge direction of each edge component pixel.
[0121] When the distance between the intersection points corresponding to the plurality of edge component pixels decreases, the similarity between the reference vectors corresponding to the plurality of edge component pixels increases. Therefore, the inner product between the reference vectors can be implemented to determine the proximity level of the intersection points. The intersection point between the virtual horizontal line 1102 and the line corresponding to the edge direction of the i-th edge component pixel is represented by φ i The reference vector from the reference point 1101 to the intersection point φ i corresponding to the i-th edge component pixel is represented by In Figure 11 , i, j, and k are integers greater than or equal to "1". The position estimation device can transform the intersection point φ i corresponding to the i-th edge component pixel into the i-th reference vector through the first reference vector transformation 1131. In addition, the position estimation device can transform the intersection point φ j corresponding to the j-th edge component pixel into the j-th reference vector
[0122] The position estimation device calculates the similarity between the reference vectors of the plurality of edge component pixels as the proximity level between the intersection points. For example, the position estimation device can input the i-th reference vector as a query value into the attention neural network 1130 and input the j-th reference vector as a key value. The position estimation device can calculate the similarity between the reference vectors as an attention value through matrix multiplication (MatMul) 1134, as shown in Equation 1 below.
[0123] [Equation 1]
[0124]
[0125] In Equation 1, a j is the inner product between the transposed vector of the i-th reference vector and the j-th reference vector , and A() is the inner product operation. Additionally, q is the query value, and k j is the j-th key value. As described above, the inner product between reference vectors indicates the proximity level between intersections, so a j is the proximity level between the intersection point φ i corresponding to the i-th edge component pixel and the intersection point φ j corresponding to the j-th edge component pixel.
[0126] The position estimation device can remove outlier pixels by applying the calculated similarity to the vector values converted from the image coordinates and edge direction information of the edge component pixels. For example, the position estimation device can apply the inner product value between reference vectors to the SoftMax operation 1135 to convert the similarity between reference vectors into weights based on Equation 2 and Equation 3 shown below.
[0127] [Equation 2]
[0128] w 1 ,..., w n = softmax(a 1 ,..., a n )
[0129] [Equation 3]
[0130]
[0131] In Equation 2, n is an integer greater than or equal to "1" and is the number of intersection points, and the weights w 1 to weight w n are the results of the softmax function of the attention values based on Equation 1. A single weight is as shown in Equation 3. In Equation 3, a is the index indicating the reference vector and is an integer greater than or equal to "1". In Equation 3, e is the Euler number, and w j is the weight corresponding to the proximity level between the intersection point φ i corresponding to the i-th edge component pixel and the intersection point φ j corresponding to the j-th edge component pixel. The position estimation device can obtain the weight vector [w 1 ,..., w n for all intersection points from Equation 3.
[0132] In addition, the position estimation device can obtain a vector value by transforming the image coordinates and edge direction information of the edge component pixels. For example, the position estimation device can generate value data v through the g transformation 1133 according to the j-th edge component pixel j = g j (x j , y j , θ j ). The position estimation device can input the vector value v generated through the g transformation 1133 j into the attention neural network 1130.
[0133] The position estimation device can perform matrix multiplication 1136 between the vector value v j and the weight vector of Equation 3, and can generate an output value, as shown in Equation 4 below.
[0134] [Equation 4]
[0135]
[0136] In Equation 4, o i is the value of the i-th edge component pixel. In the example, when the intersection point corresponding to the i-th edge component pixel is far from the intersection points corresponding to other edge component pixels, the value corresponding to the i-th edge component pixel can converge to "0", so the outlier pixels can be excluded. In another example, when the intersection point corresponding to the i-th edge component pixel is adjacent to the intersection points corresponding to other edge component pixels, the value corresponding to the i-th edge component pixel can be a non-zero value due to its large weight, so the inlier pixels can be retained. Therefore, the position estimation device can retain the inlier pixels by excluding the outlier pixels from multiple edge component pixels through the attention neural network 1130 and Equations 1 to 4.
[0137] As described above, the position estimation device can determine the edge component pixels that form intersection points with the virtual horizontal line 1102 and are adjacent to each other as inlier pixels. The position estimation device can estimate the rotation reference point 790 based on the inlier image including the inlier pixels. For example, the position estimation device can determine the point with the shortest distance from the edge line corresponding to the inlier pixels in the input image as the rotation reference point 790. The edge line is a straight line passing through the inlier pixels, and can be, for example, a straight line with a slope corresponding to the edge direction indicated by the inlier pixels. The position estimation device can determine the point with the minimum distance between the edge line and the horizontal line as the rotation reference point 790. However, the determination of the rotation reference point 790 is not limited to the above description, but various schemes can be used to determine the rotation reference point 790. Although Figure 11 an example of directly estimating the rotation reference point according to the inlier pixels has been described, the present disclosure is not limited to this example.
[0138] Refer to the followingFigure 12 Describe an example of performing line fitting and estimating a rotation reference point.
[0139] Figure 12 An example of line fitting is shown. A position estimation device (e.g., position estimation device 100) according to an example embodiment performs line fitting on inlier pixels among a plurality of edge component pixels. The position estimation device clusters the plurality of edge component pixels into at least one straight line group, and determines a representative edge line for each of the at least one straight line group to perform line fitting.
[0140] According to an example, Figure 12 A first edge component pixel 1211 and a second edge component pixel 1212 in a part 1210 of an inlier image 706 are shown. A first edge line 1221 of the first edge component pixel 1211 and a second edge line 1222 of the second edge component pixel 1212 form angles 1231 and 1232 that are different from each other with respect to an arbitrary reference axis. Since the edge directions of the first edge line 1221 and the second edge line 1222 are similar to each other, the position estimation device clusters the first edge line 1221 and the second edge line 1222 into the same straight line group, and determines a representative edge line 1290 of the straight line group to which the first edge line 1221 and the second edge line 1222 belong. The representative edge line 1290 has an average edge direction of the edge lines included in the straight line group. However, the present disclosure is not limited to such an example.
[0141] For the above clustering, the position estimation device converts an image including a plurality of inlier pixels into a bird's-eye view image. Compared with edge lines having different edge directions, edge lines having similar edge directions among the edge lines projected onto the bird's-eye view image can be more clearly distinguished from each other. The position estimation device clusters the plurality of inlier pixels into at least one straight line group based on lines corresponding to the plurality of inlier pixels in the bird's-eye view image. The position estimation device determines an edge line representing the at least one straight line group.
[0142] The position estimation device estimates a rotation reference point based on the edge line obtained by line fitting. For example, the position estimation device can obtain a representative edge line of each individual straight line group by line fitting, and determine a point with the shortest distance from the representative edge line as the rotation reference point. Refer to the following Figure 13 Describe an example of correcting positioning information based on a rotation reference point determined by the method of Figures 1 to 12 .
[0143] Figure 13 A diagram showing an example of a rotation reference point error between the initial positioning information estimated by the position estimation device and the actual attitude of the position estimation device is shown.
[0144] In Figure 13In the example, it is assumed that the target 1310 in the attitude based on the initial positioning information is positioned obliquely with respect to the lane boundary. The attitude of the target 1310 is virtually shown based on the initial positioning information. For example, a position estimation device that may be set on or within the target determines points outside the threshold distance from the target 1310 among the points defining the boundary line of the lane on which the target 1310 travels as the initial reference point 1350. When the initial reference point 1350 is projected onto the image plane corresponding to the field of view angle of the sensor 1311 attached to the target 1310 in the attitude based on the initial positioning information, the point 1351 onto which the initial reference point 1350 is projected is as Figure 13 shown in
[0145] In Figure 13 the example, different from the initial positioning information, the actual target 1320 is oriented parallel to the lane boundary. The sensor 1321 of the actual target 1320 captures a forward view image from the actual target 1320. As described with reference to Figures 1 to 12 , the position estimation device detects the rotation reference point 1361 from the image data 1360.
[0146] As described above, due to the error in the initial positioning information, the point where the initial reference point 1350 is directly projected onto the image plane and the rotation reference point 1361 may appear at different points. Therefore, the difference 1356 between the point where the initial reference point 1350 is projected onto the image plane and the rotation reference point 1361 may correspond to the error in the initial positioning information, such as the rotation error and the position error.
[0147] By using the rotation parameter to correct the positioning error, the position estimation device can correct the initial positioning information. The position estimation device can generate the final positioning information of the position estimation device by correcting at least one of the position error and the rotation error in the initial positioning information based on the rotation reference point.
[0148] Figure 14 shows an example of the operation of the position estimation device 1400 installed on a vehicle according to an embodiment.
[0149] Referring to Figure 14 , the position estimation device 1400 includes, for example, a sensor 1410, a processor 1420, a memory 1430, and a display 1440.
[0150] The sensor 1410 acquires an input image. The sensor 1410 may be a camera sensor 1410 configured to capture the external scene 1401. The camera sensor 1410 generates a color image as the input image. The color image includes images corresponding to each of the red channel, the blue channel, and the green channel, however, the color space is not limited to the RGB color space. The camera sensor 1410 is positioned to capture the external environment of the position estimation device 1400.
[0151] The processor 1420 calculates edge direction information of a plurality of edge component pixels extracted from an input image, determines inlier pixels among the plurality of edge component pixels based on the edge direction information and virtual horizontal line information, estimates a rotation reference point based on the inlier pixels, and corrects the positioning information of the position estimation device 1400 based on the estimated rotation reference point.
[0152] The memory 1430 temporarily or permanently stores data required to execute the position estimation method.
[0153] The display 1440 performs visualization by mapping the virtual content object 1490 to coordinates determined based on the final positioning information of the position estimation device 1400 in the display 1440. The display 1440 transmits the external scene 1401 to the user. For example, the display 1440 may provide the external scene 1401 captured by the camera sensor 1410 to the user. However, the present disclosure is not limited to such an example. In an example, the display 1440 may be implemented as a transparent display 1440 to transmit the external scene 1401 to the user by allowing the external scene 1401 to be seen through the transparent display 1440. In another example, when the position estimation device 1400 is installed on a vehicle, the transparent display 1440 may be integrated with the vehicle's windshield.
[0154] The processor 1420 corrects the initial positioning information of the position estimation device 1400 based on the rotation reference point and generates the final positioning information of the position estimation device 1400. Thus, the processor 1420 can match the estimated position with the actual position 1480 of the position estimation device 1400. For example, the position estimation device 1400 may control the operation of the device based on the positioning information that is corrected based on the rotation reference point.
[0155] For example, the processor 1420 may determine the position of the virtual content object 1490 on the display 1440 based on the final positioning information of the position estimation device 1400. For example, the display 1440 may be as Figure 14Visualization is performed by overlaying a virtual content object 1490 on an external scene 1401 as shown. When the display 1440 visualizes the external scene 1401 captured by the camera sensor 1410, the processor 1420 may perform visualization through the display 1440 by overlaying the virtual content object 1490 on the rendered external scene 1401. When the display 1440 is implemented as a transparent display 1440, the processor 1420 may map the position of the virtual content object 1490 to the actual physical position coordinates of the external scene 1401, and may visualize the virtual content object 1490 at a point corresponding to the mapped position on the display 1440. The virtual content object 1490 may be, for example, an object that guides information associated with the travel of a vehicle, and may indicate various information associated with the travel direction and direction indicators (e.g., left turn points and right turn points), speed limits, and the current speed of the vehicle.
[0156] For example, when the position estimation device 1400 is included in a vehicle, the position estimation device 1400 may control the operation of the vehicle based on the final positioning information. The position estimation device 1400 may determine the lane in which the vehicle is located among a plurality of lanes defined by lane boundary lines 1470 on the road on which the vehicle travels based on the final positioning information, and may detect the distance between the position estimation device 1400 and an adjacent obstacle (e.g., an object ahead). The position estimation device 1400 may control any one or any combination of the acceleration, speed, and steering of the vehicle based on the distance detected based on the final positioning information. For example, when the distance between the position estimation device 1400 and an object ahead decreases, the position estimation device 1400 may decrease the speed of the vehicle. When the distance between the position estimation device 1400 and an object ahead increases, the position estimation device 1400 may increase the speed of the vehicle. In addition, the position estimation device 1400 may change the steering of the vehicle (e.g., turn right and / or turn left) based on the final positioning information of the vehicle along a preset travel route.
[0157] has been referred to Figure 14 Examples of using a rotational reference point in positioning based on map data and travel images in an augmented reality (AR) environment have been described. However, the present disclosure is not limited to the application of AR in a travel environment. For example, even in various environments, such as AR navigation during walking, the rotational reference point estimation scheme according to the examples may be used to enhance positioning performance.
[0158] For example, the position estimation device can also make the estimation performance of the input data have characteristics different from those of the data used for training the neural network. Therefore, the position estimation device can exhibit rotation reference point estimation performance that is robust to various environments. For example, the position estimation device can accurately estimate the rotation reference point regardless of the unique characteristics of the camera sensor, the characteristics according to the driving area (e.g., urban area or rural area), the image characteristics of the driving city and / or country, and the illuminance characteristics. This is because the edge direction information used to estimate the rotation reference point is less affected by the characteristics of the camera sensor, the area, or the illuminance. The position estimation device can estimate the rotation reference point based on a single image instead of performing a separate operation of the motion vector. The above position estimation method can be implemented by, for example, software, network services, and chips (e.g., processors and memories).
[0159] The parameters of the machine learning module (e.g., any, any combination, or all in the neural network) are stored in one or more memories and implemented by one or more processors by reading and implementing the read parameters.
[0160] Figures 1 to 14The processors, processor 120, and processor 1420, sensors, GPS receivers, GNSS modules, IMUs, memories, and memories 130 and 1430 that perform the operations described in this application are implemented by hardware components configured to perform the operations described in this application, and this application is executed by hardware components. Examples of hardware components that can be used to perform the operations described in this application, where appropriate, include controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtracters, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more hardware components for performing the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (e.g., logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes (or is connected to) one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (e.g., an operating system (OS) and one or more software applications running on the OS) to perform the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to executing the instructions or software. For the sake of brevity, the singular terms "processor" or "computer" may be used in the description of the examples described in this application, but in other examples, multiple processors or computers may be used, or a processor or computer may include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, can implement a single hardware component, or two or more hardware components. The hardware components can have one or more of different processing configurations, examples of which include single-processor, independent-processor, parallel-processor, single-instruction single-data (SISD) multiprocessing, single-instruction multiple-data (SIMD) multiprocessing, multiple-instruction single-data (MISD) multiprocessing, and multiple-instruction multiple-data (MIMD) multiprocessing.
[0161] Figures 1 to 14The method of performing the operations described in this application, as shown, is executed by computing hardware, such as by one or more processors or computers that implement the instructions or software as described above to perform the operations performed by the method described in this application. For example, a single operation or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors, or a processor and a controller, and one or more other operations can be performed by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller, can perform a single operation, or two or more operations.
[0162] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method as described above can be written as a computer program, code segment, instruction, or any combination thereof, for individually or jointly instructing or configuring one or more processors or computers to operate as a machine or special-purpose computer to perform the operations performed by the above-described hardware components and method. In one example, the instructions or software include machine code directly executable by one or more processors or computers, such as machine code generated by a compiler. In another example, the instructions or software include higher-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions in the specification, which disclose algorithms for performing the operations performed by the hardware components and the method as described above.
[0163] Instructions or software for controlling a computing hardware (e.g., one or more processors or computers) to implement the hardware components and execute the methods as described above, as well as any associated data, data files, and data structures, can be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), flash memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the one or more processors or computers store, access, and execute the instructions and software and any associated data, data files, and data structures in a distributed manner.
[0164] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made to these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered only as descriptive and not for purposes of limitation. The description of a feature or aspect in each example is considered applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner and / or replaced or supplemented by other components or their equivalents. Accordingly, the scope of the present disclosure is defined not by the specific embodiments but by the claims and their equivalents, and all variations within the scope of the claims and their equivalents are construed as being included in the present disclosure.
Claims
1. A processor-implemented method, comprising: calculating edge direction information of edge component pixels extracted from an input image; determining the inlier pixels among the edge component pixels by removing outlier pixels among the edge component pixels based on intersections between a virtual horizontal line and edge lines, while retaining the inlier pixels among the edge component pixels, wherein the virtual horizontal line corresponds to a virtual horizon in the input image, and the edge lines correspond to the edge direction information of the edge component pixels; estimating a rotation reference point based on the inlier pixels; and correcting the positioning information of a device by applying rotation parameters to make an initial rotation reference point match the estimated rotation reference point, wherein the initial rotation reference point is determined based on map data and the positioning information.
2. The method according to claim 1, wherein calculating the edge direction information comprises: extracting the edge component pixels by preprocessing the input image.
3. The method according to claim 2, wherein extracting the edge component pixels comprises: calculating gradient information for each pixel among the pixels included in the input image; and extracting selected pixels having gradient information exceeding a threshold as the edge component pixels among the pixels included in the input image.
4. The method according to claim 1, wherein calculating the edge direction information comprises: masking a part of the edge component pixels.
5. The method according to claim 4, wherein masking a part of the edge component pixels comprises: excluding pixels on the boundary from the edge component pixels.
6. The method according to claim 5, wherein excluding the pixels on the boundary comprises: excluding pixels corresponding to a region above the virtual horizontal line in the input image.
7. The method according to claim 6, wherein excluding the pixels on the boundary further comprises: determining the virtual horizontal line in the input image based on roll parameters and pitch parameters of a camera sensor that captured the input image.
8. The method according to claim 1, wherein calculating the edge direction information comprises: determining the edge direction of each edge component pixel among the edge component pixels based on an estimation of the edge direction.
9. The method according to claim 8, wherein determining the edge direction comprises: using a convolutional neural network (CNN) to calculate a convolution operation value for each edge component pixel among the edge component pixels, the convolutional neural network (CNN) including corresponding kernels for each predetermined angle; and determining the angle value of the kernel corresponding to the highest convolution operation value among the convolution operation values calculated for each edge component pixel among the edge component pixels as the edge direction information.
10. The method according to claim 1, wherein retaining the inlier pixels comprises: generating inlier data with the outlier pixels removed based on an attention neural network.
11. The method according to claim 1, wherein retaining the inlier pixels comprises: removing the outlier pixels based on the proximity level between the intersections.
12. The method according to claim 11, wherein removing the outlier pixels comprises: For each edge component pixel among the edge component pixels, calculating a reference vector from an intersection point between an edge line in the edge lines and a virtual horizontal line among the intersection points to a reference point in the input image; and Calculating a similarity between the reference vectors based on the proximity level between the intersection points.
13. The method according to claim 12, wherein removing the outlier pixels further comprises: Removing the outlier pixels by applying the calculated similarity to a vector value obtained by transforming the edge direction information and the image coordinates of an edge component pixel among the edge component pixels.
14. The method according to claim 1, wherein estimating the rotation reference point comprises: Performing line fitting on the inlier pixels; and Based on the line fitting, estimating the rotation reference point based on the edge line.
15. The method according to claim 14, wherein performing the line fitting comprises: Converting an image including a plurality of inlier pixels into a bird's-eye view image; Clustering the plurality of inlier pixels into at least one straight line group based on lines corresponding to the plurality of inlier pixels in the bird's-eye view image; and Determining the edge line as the edge line representing each straight line group among the at least one straight line group.
16. The method according to claim 1, wherein estimating the rotation reference point comprises: Determining a point with the shortest distance from an edge line corresponding to the inlier pixels in the input image as the rotation reference point.
17. The method according to claim 1, further comprises: Performing visualization in a display for transmitting an external scene to a user by mapping a virtual content object to display coordinates determined based on the final positioning information of the device, wherein correcting the positioning information comprises: correcting any one or both of a position error and a rotation error of the positioning information based on the estimated rotation reference point; and generating the final positioning information based on a result of correcting any one or both of the position error and the rotation error.
18. The method according to claim 1, further comprises: Obtaining the positioning information using a global navigation satellite system GNSS module and an inertial measurement unit IMU; and Obtaining the input image using an image sensor.
19. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method according to claim 1.
20. A device for position estimation, comprises: A sensor configured to obtain an input image; and A processor configured to: Calculate edge direction information of edge component pixels extracted from the input image; Determine the inlier pixels by removing outlier pixels among the edge component pixels and retaining inlier pixels among the edge component pixels based on intersection points between a virtual horizontal line and an edge line, wherein the virtual horizontal line corresponds to a virtual horizon in the input image, and the edge line corresponds to the edge direction information of the edge component pixels; Estimate a rotation reference point based on the inlier pixels; and correct the positioning information of the device by applying rotation parameters to match the initial rotation reference point with the estimated rotation reference point, where the initial rotation reference point is determined based on map data and the positioning information.
21. The device according to claim 20, wherein the edge component pixels include selected pixels that are determined to have gradient information exceeding a threshold among the pixels included in the input image.
22. The device according to claim 20, wherein the processor is further configured to: mask a part of the edge component pixels by excluding pixels corresponding to a region above a virtual horizontal line in the input image from the edge component pixels.
23. The device according to claim 22, wherein the processor is further configured to determine the virtual horizontal line based on the roll parameter and the pitch parameter of the sensor.
24. The device according to claim 20, wherein the processor is further configured to calculate the edge direction information by: implementing a neural network to perform a convolution operation on each edge component pixel in the edge component pixels; and based on the result of the convolution operation, determining an angle corresponding to each edge component pixel in the edge component pixels as the edge direction information.
25. The device according to claim 20, wherein the processor is further configured to: remove outlier pixels in the edge component pixels based on the proximity level between intersections between the virtual horizontal line and the edge lines.
26. The device according to claim 25, wherein removing the outlier pixels includes: for each edge component pixel in the edge component pixels, calculating a reference vector from an intersection between one of the edge lines in the edge lines and the virtual horizontal line among the intersections to a reference point in the input image; calculating the similarity between the reference vectors based on the proximity level between the intersections; and removing the outlier pixels by applying the calculated similarity to a vector value obtained by transforming the edge direction information and the image coordinates of one of the edge component pixels in the edge component pixels.
27. The device according to claim 20, wherein the processor is further configured to determine the inlier pixels by: generating inlier data with the outlier pixels removed based on an attention neural network.
28. The device according to claim 20, wherein the processor is further configured to estimate the rotation reference point by: performing line fitting on the inlier pixels in the bird's-eye view image based on the input image; and estimating the rotation reference point based on the edge lines according to the line fitting.
29. The device according to claim 20, wherein the processor is further configured to: estimate the rotation reference point by determining a point with the shortest distance from the edge line corresponding to the inlier pixels in the input image as the rotation reference point.
30. The apparatus according to claim 20, wherein the processor is further configured to: perform visualization in a display for a user of an external scene by mapping a virtual content object to display coordinates determined based on the final positioning information of the apparatus, and wherein the processor is further configured to: correct the positioning information by correcting any one or both of a position error and a rotation error of the positioning information based on an estimated rotation reference point, and generate the final positioning information based on a result of correcting any one or both of the position error and the rotation error.
31. The apparatus according to claim 20, wherein the processor is further configured to obtain the positioning information from a Global Navigation Satellite System (GNSS) module and an Inertial Measurement Unit (IMU).
32. An Augmented Reality (AR) apparatus comprising: a display configured to provide a view of an external environment; a sensor configured to acquire an input image; and a processor configured to: calculate edge direction information of edge component pixels extracted from the input image; determine inlier pixels among the edge component pixels by removing outlier pixels among the edge component pixels based on intersections between a virtual horizontal line and edge lines, wherein the virtual horizontal line corresponds to a virtual horizon in the input image, and the edge lines correspond to the edge direction information of the edge component pixels; estimate a rotation reference point based on the inlier pixels; correct initial positioning information of the apparatus by applying rotation parameters to match an initial rotation reference point with the estimated rotation reference point to generate final positioning information, wherein the initial rotation reference point is determined based on map data and the initial positioning information; and perform visualization by mapping a virtual content object to coordinates determined based on the final positioning information in the display.
33. The AR apparatus according to claim 32, wherein the sensor includes a Global Navigation Satellite System (GNSS) module and an Inertial Measurement Unit (IMU), and the initial positioning information includes position information and attitude information obtained from the GNSS module and the IMU.
34. The AR apparatus according to claim 32, wherein the processor is further configured to calculate the edge direction information by: implementing a neural network to perform a convolution operation on each of the edge component pixels in the edge component pixels; and based on a result of the convolution operation, determining an angle corresponding to each of the edge component pixels as the edge direction information.
35. The AR apparatus according to claim 32, wherein the processor is further configured to estimate the rotation reference point by determining a point having the shortest distance from an edge line corresponding to the inlier pixels in the input image as the rotation reference point.
36. The AR apparatus according to claim 32, wherein the processor is further configured to estimate the rotation reference point by: performing line fitting on inlier pixels in a bird's-eye view image based on the input image; and Based on the line fitting, estimate the rotation reference point based on the edge line.
37. The AR device according to claim 32, wherein the processor is further configured to determine the inlier pixels by: Generate inlier data with the outlier pixels removed based on an attention neural network.
38. The AR device according to claim 32, wherein the display includes a windshield of a vehicle.
Citation Information
Patent Citations
A pressing jig for pressing a bus bar and a battery module manufacturing system comprising the same
KR1020190115939A
Method and apparatus for estimating position
CN111060946A