Vehicle attitude determination
By detecting static environmental features in the vehicle's camera image and generating distance transformed images, combined with map data, the complex environmental problem of vehicle attitude determination is solved, and high-precision vehicle attitude recognition is achieved.
Patent Information
- Application Number
- CN202411830497.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively utilize static environmental features and map data to determine the attitude of a vehicle, especially in complex environments.
By detecting static environment features in the camera image of the vehicle, a distance transformed image is generated, and the value of the cost function is calculated to determine the attitude of the vehicle based on the comparison of map data indicating static environment features with the distance transformed image.
The ability to accurately determine vehicle attitude in complex environments is achieved, improving the accuracy and reliability of Advanced Driver Assistance Systems (ADAS).
Smart Images

Figure CN120182364A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure describes techniques for determining the pose of a vehicle based on static environmental features detected in a camera image and map data of the static environmental features. Background Art
[0002] Advanced driver assistance systems (ADAS) are electronic technologies that assist a driver in performing driving functions and parking functions. Examples of ADAS include forward proximity detection, lane departure detection, blind spot detection, brake actuation, adaptive cruise control, and lane keeping assistance systems. Summary of the Invention
[0003] The pose of a vehicle includes the position and / or orientation of the vehicle, and some advanced driver assistance systems (ADAS) can use the pose as an input. Static environmental features can include lane lines, street light poles, and the like. To determine the pose, a computer of the vehicle uses what will be referred to as a "distance transform image". A distance transform image is a two-dimensional matrix of pixels, and each pixel can have a scalar pixel value that indicates the pixel distance of that pixel from the nearest location of one of the static environmental features in the distance transform image. Pixels that coincide with the location of one of the static environmental features can have a pixel value of zero, and the pixel value can increase at locations away from the coinciding pixels. The computer of the vehicle is programmed to detect static environmental features in a camera image, generate a distance transform image of the static environmental features as detected in the camera image, and determine the pose of the vehicle based on a comparison of map data indicating the static environmental features with the distance transform image. For example, the computer can use the distance transform image to calculate a cost function that optimizes the projection of map data of the static environmental features onto the image plane of the distance transform image. The projection can indicate the pose of the vehicle.
[0004] A computer includes a processor and a memory, and the memory stores instructions that are executable by the processor to detect static environmental features in a camera image from a camera of a vehicle, generate a distance transform image of the static environmental features as detected in the camera image, and determine the pose of the vehicle based on a comparison of map data indicating the static environmental features with the distance transform image. The pixel value of a corresponding pixel in the distance transform image indicates the corresponding pixel distance of the corresponding pixel from the static environmental features in the distance transform image.
[0005] In one example, the instructions may further include instructions for performing the following operations: calculating a value of a cost function based on the map data indicating the static environmental features and the distance transform image, and determining an attitude of the vehicle that minimizes the value of the cost function. In another example, the instructions may further include instructions for performing the following operations: projecting the map data indicating the static environmental features onto the distance transform image, and calculating the value of the cost function based on pixel values of the pixels onto which the map data is projected.
[0006] In yet another example, the attitude may be a first attitude, and the instructions may further include instructions for performing the following operations: determining a GNSS attitude based on Global Navigation Satellite System (GNSS) data, and initializing the first attitude at the GNSS attitude to minimize the value of the cost function.
[0007] In one example, the static environmental features may include lane lines.
[0008] In one example, the distance transform image may be an overhead distance transform image from an overhead perspective. In another example, the attitude may be a first attitude, and the instructions may further include instructions for performing the following operations: generating an image plane distance transform image from the perspective of the camera, and determining a second attitude based on a comparison of the first attitude and the map data indicating the static environmental features with the image plane distance transform image. In yet another example, the instructions may further include instructions for performing the following operations: calculating a value of a cost function based on the map data indicating the static environmental features and the image plane distance transform image, and determining a second attitude of the vehicle that minimizes the value of the cost function. In still yet another example, the instructions may further include instructions for performing the following operations: initializing the second attitude at the first attitude to minimize the value of the cost function.
[0009] In still yet another example, the first attitude may include only two horizontal spatial dimensions and a heading, and the second attitude may include three spatial dimensions and three angular dimensions.
[0010] In one example, the distance transform image may be an image plane distance transform image from the perspective of the camera. In another example, the static environmental features may include linear vertical features.
[0011] In one example, the instructions may further include instructions for generating a binary image depicting the static environmental features and generating the distance transform image from the same perspective as the binary image. In additional examples, the binary image may depict only the static environmental features.
[0012] In one example, the pose may include two horizontal spatial dimensions and a heading.
[0013] In one example, the instructions may further include instructions for actuating components of the vehicle based on the pose of the vehicle.
[0014] A method includes detecting static environmental features in a camera image from a camera of a vehicle, generating a distance transform image of the static environmental features as detected in the camera image, and determining a pose of the vehicle based on a comparison of map data indicative of the static environmental features and the distance transform image. A pixel value of a corresponding pixel in the distance transform image indicates a distance of the corresponding pixel from a corresponding pixel of the static environmental features in the distance transform image.
[0015] In one example, the method further includes calculating a value of a cost function based on the map data indicative of the static environmental features and the distance transform image, and determining the pose of the vehicle that minimizes the value of the cost function. In additional examples, the method further includes projecting the map data indicative of the static environmental features onto the distance transform image and calculating the value of the cost function based on the pixel values of the pixels onto which the map data is projected.
[0016] In one example, the method further includes generating a binary image depicting the static environmental features and generating the distance transform image from the same perspective as the binary image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram of an exemplary vehicle.
[0018] Figure 2 is a block diagram of an exemplary algorithm for determining a pose of a vehicle.
[0019] Figure 3A is an exemplary camera image in which static environmental features are detected.
[0020] Figure 3B is depicted in Figure 3A is an exemplary binary image depicting the static environmental features detected in the camera image.
[0021] Figure 3C is based onFigure 3B Exemplary distance transformation image generated from a binary image
[0022] Figure 3D is Figure 3C Three-dimensional plot of the pixel values of the distance transformation image
[0023] Figure 3E is Figure 3D Exemplary cross-section of the three-dimensional plot
[0024] Figure 4A Another exemplary camera image detecting static environmental features
[0025] Figure 4B is depicting Figure 4A Exemplary binary image depicting static environmental features detected in the camera image
[0026] Figure 4C is based on Figure 4B Exemplary distance transformation image generated from the binary image
[0027] Figure 5A Another exemplary camera image detecting static environmental features
[0028] Figure 5B is based on Figure 5A Exemplary distance transformation image generated from the camera image
[0029] Figure 6 is Figure 3B Illustration of map data projected and transformed on the binary image
[0030] Figure 7 Flowchart of an exemplary process for determining the pose of a vehicle DETAILED DESCRIPTION
[0031] Referring to the accompanying drawings, in which like numerals indicate like parts throughout the several views, computer 105 includes a processor and a memory, and the memory stores instructions that can be executed by the processor to detect static environmental features 310 in a camera image 305 from a camera 110 of a vehicle 100, generate a distance transformation image 315 of the static environmental features 310 detected in the camera image 305, and determine the poses 205, 210 of the vehicle 100 based on a comparison of map data 245 indicating the static environmental features 310 and the distance transformation image 315. The pixel value of a corresponding pixel in the distance transformation image 315 indicates the corresponding pixel distance of the corresponding pixel from the static environmental features 310 in the distance transformation image 315.
[0032] Referring Figure 1, Vehicle 100 can be any passenger or commercial vehicle, such as a sedan, truck, sport utility vehicle, crossover vehicle, van, minivan, taxi, bus, etc. Vehicle 100 can include a computer 105, a communication network 115, a camera 110, a global navigation satellite system (GNSS) receiver 120, other sensors 125, a propulsion system 130, a braking system 135, a steering system 140, and a user interface 145.
[0033] The computer 105 is a microprocessor-based computing device, such as a general-purpose computing device (including a processor and a memory, an electronic controller, etc.), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a combination of the foregoing, etc. Generally, hardware description languages such as VHDL (VHSIC (very high speed integrated circuit) hardware description language) are used in electronic design automation to describe digital and mixed-signal systems such as FPGAs and ASICs. For example, an ASIC is manufactured based on the VHDL programming provided before manufacturing, and the logic components inside an FPGA can be configured based on, for example, the VHDL programming stored in a memory electrically connected to the FPGA circuit. Thus, the computer 105 can include a processor, a memory, etc. The memory of the computer 105 can include a medium for storing instructions executable by the processor and for electronically storing data and / or databases, and / or the computer 105 can include a structure such as the foregoing structures that provide programming. The computer 105 can be a plurality of computers connected together.
[0034] The computer 105 can transmit and receive data through the communication network 115. The communication network 115 can be, for example, a controller area network (CAN) bus, Ethernet, WiFi, local interconnect network (LIN), on-board diagnostic connector (OBD-II), and / or any other wired or wireless communication network. The computer 105 can be communicatively coupled to the camera 110, the GNSS receiver 120, other sensors 125, the propulsion system 130, the braking system 135, the steering system 140, the user interface 145, and other components via the communication network 115.
[0035] Vehicle 100 includes at least one camera 110, such as a plurality of cameras 110. The camera 110 can detect electromagnetic radiation in a certain wavelength range. For example, the camera 110 can detect visible light, infrared radiation, ultraviolet light, or a certain range of wavelengths including visible light, infrared light, and / or ultraviolet light. For example, the camera 110 may be a charge-coupled device (CCD), a complementary metal-oxide semiconductor (CMOS), or any other suitable type. The camera 110 can form a surround-view camera system, where the cameras 110 are oriented in different directions away from the vehicle 100. For example, at least one camera 110 faces forward, at least one camera 110 faces right, at least one camera 110 faces left, and at least one camera 110 faces rearward. In addition to the features discussed herein, the camera 110 can also support other features, such as ADAS features, such as parking assistance.
[0036] The GNSS receiver 120 receives data from GNSS satellites. Systems for GNSS include the Global Positioning System (GPS), GLONASS, Beidou, Galileo, etc. The GNSS satellites broadcast time and geographical location data. The GNSS receiver 120 or the computer 105 can determine the GNSS attitude 215 of the vehicle 100, such as latitude and longitude, based on the GNSS receiver 120 receiving time and geographical location data from multiple GNSS satellites simultaneously and using the trilateration principle.
[0037] Other sensors 125 can provide data about the operation of the vehicle 100, such as wheel speed, wheel orientation, and engine and transmission data (e.g., temperature, fuel consumption, etc.). The other sensors 125 can detect the position and / or orientation of the vehicle 100. For example, the other sensors 125 can include accelerometers, such as piezoelectric systems or microelectromechanical systems (MEMS); gyroscopes, such as rate gyroscopes, ring laser gyroscopes, or fiber optic gyroscopes; inertial measurement units (IMU); and magnetometers. The other sensors 125 can detect the external world, such as objects and / or characteristics of the surrounding environment of the vehicle 100, such as other vehicles, road lane markings, traffic lights and / or signs, road users, etc. For example, the other sensors 125 can include radar sensors, ultrasonic sensors, scanning lidar, and optical detection and ranging (lidar) devices.
[0038] The propulsion system 130 of the vehicle 100 generates energy and converts the energy into the motion of the vehicle 100. The propulsion system 130 can be a conventional vehicle propulsion subsystem, for example, a conventional powertrain that includes an internal combustion engine coupled to a transmission that transfers rotational motion to the wheels; an electric powertrain that includes a battery, an electric motor, and a transmission that transfers rotational motion to the wheels; a hybrid powertrain that includes elements of a conventional powertrain and an electric powertrain; or any other type of propulsion device. The propulsion system 130 can include an electronic control unit (ECU) etc. that communicates with and receives input from the computer 105 and / or a human operator. The human operator can control the propulsion system 130 via, for example, an accelerator pedal and / or a gear shift lever.
[0039] The braking system 135 is typically a conventional vehicle braking subsystem and inhibits the motion of the vehicle 100, thereby slowing down and / or stopping the vehicle 100. The braking system 135 can include friction brakes, such as disc brakes, drum brakes, band brakes, etc.; regenerative brakes; any other suitable type of brakes; or a combination thereof. The braking system 135 can include an electronic control unit (ECU) etc. that communicates with and receives input from the computer 105 and / or a human operator. The human operator can control the braking system 135 via, for example, a brake pedal.
[0040] The steering system 140 is typically a conventional vehicle steering subsystem and controls the turning of the wheels. The steering system 140 can be a rack and pinion system with electric power steering, a steer-by-wire system (both of which are known), or any other suitable system. The steering system 140 can include an electronic control unit (ECU) etc. that communicates with and receives input from the computer 105 and / or a human operator. The human operator can control the steering system 140 via, for example, a steering wheel.
[0041] The user interface 145 presents information to and receives information from the operator of the vehicle 100. The user interface 145 can be located, for example, on the dashboard in the passenger compartment of the vehicle 100, or anywhere easily visible to the operator. The user interface 145 can include dials, digital readouts, screens, speakers, etc. for providing information to the operator, for example, such as known human-machine interface (HMI) elements. The user interface 145 can include buttons, knobs, keypads, microphones, etc. for receiving information from the operator.
[0042] Reference Figure 2, the computer 105 is programmed to determine the poses 205, 210 of the vehicle 100. The poses 205, 210 describe the position and / or orientation of the vehicle 100. The poses 205, 210 can include at least two horizontal spatial dimensions and one angular dimension such as heading (also known as yaw). For example, the poses 205, 210 can include only two horizontal spatial dimensions and heading, or can include three spatial dimensions and three angular dimensions. In the specific example below, the computer 105 determines two poses 205, 210 at two stages of the process, which will be referred to as the first pose 205 and the second pose 210. The first pose 205 can include only two horizontal spatial dimensions and heading, and the first pose 205 can be used as an input for determining the second pose 210. The second pose 210 can include three spatial dimensions and three angular dimensions.
[0043] The computer 105 determines the poses 205, 210 based on the static environmental features 310. The static environmental features 310 are aspects of the environment around the vehicle 100 that remain constant over time. The static environmental features 310 can be selected to be both described in the map data 245 and detectable in the camera image 305. For example, the static environmental features 310 can include generally horizontal edges (such as lane lines) and linear vertical features (such as street light poles, traffic signal poles, power line poles, etc.). The linear vertical features are selected to be a type of feature that linearly extends along the vertical dimension, i.e., straight up and down or having a main sub-component that is straight up and down. The computer 105 can use the lane lines when determining the first pose 205 because the first pose 205 is determined from an overhead perspective (as will be described). The computer 105 can use both the lane lines and the linear vertical features when determining the second pose 210 because the second pose 210 is determined from the perspective of the camera 110 (as will also be described).
[0044] As an overall overview, computer 105 determines the GNSS attitude 215 at the GNSS box 220, and computer 105 determines the ranging attitude 225 at the ranging box 230 based on data from other sensors 125. Computer 105 detects some static environmental features 310, such as lane lines, in the camera image 305 from camera 110 in the lane detection box 235. Computer 105 detects other static environmental features 310, such as linear vertical features, in the camera image 305 in the vertical detection box 240. Computer 105 extracts map data 245 indicating the static environmental features 310 in the map box 250. In the lane positioning box 255, computer 105 determines the first attitude 205 based on the GNSS attitude 215, the ranging attitude 225, the map data 245, and the detected static environmental features 310 from the lane detection box 235. Computer 105 initializes the first attitude 205 to the GNSS attitude 215 or the ranging attitude 225, and starting from the initialized first attitude 205, determines the final first attitude 205 based on a comparison of the map data 245 indicating the static environmental features 310 with an overhead distance transform image 315a generated from the static environmental features 310 detected in the camera image 305. In the full positioning box 260, computer 105 determines the second attitude 210 based on the first attitude 205, the map data 245, and the detected static environmental features 310 from the lane detection box 235 and the vertical detection box 240. Computer 105 initializes the second attitude 210 to the first attitude 205, and starting from the initialized second attitude 210, determines the final second attitude 210 based on a comparison of the map data 245 indicating the static environmental features 310 with an image plane distance transform image 315b generated from the static environmental features 310 detected in the camera image 305.
[0045] In the GNSS box 220, computer 105 or GNSS receiver 120 determines the GNSS attitude 215 of vehicle 100 based on GNSS data received by GNSS receiver 120. The GNSS attitude 215 describes the position and / or orientation of vehicle 100, e.g., two horizontal spatial dimensions (such as latitude and longitude) and one angular dimension (such as heading), or three spatial dimensions and three angular dimensions. As is known, computer 105 or GNSS receiver 120 uses trilateration to determine the GNSS attitude 215. The GNSS attitude 215 can be specified in an absolute coordinate system (i.e., a coordinate system fixed relative to the Earth). The GNSS attitude 215 is provided to the lane positioning box 255.
[0046] In the map frame 250, the computer 105 extracts, for example, map data 245 indicating static environmental features 310 from a map database stored in the memory of the computer 105. The map data 245 can be specified as positions in an absolute coordinate system paired with specific static environmental features.
[0047] In the odometry frame 230, the computer 105 estimates the odometry pose 225 of the vehicle 100 based on data from other sensors 125 (e.g., proprioceptive measurements, i.e., self-sensing measurements of movement). For example, the odometry pose 225 can be based on data from wheel speed sensors and an inertial measurement unit (IMU) of the other sensors 125. The odometry pose 225 can also be based on changes in radar measurements of stationary objects in the environment. The computer 105 can estimate the odometry pose 225 by integrating the velocity vector inferred from data from the other sensors 125 from a previous time to the current time, resulting in changes in position and orientation, and adding the changes in position and orientation to the previous pose from the previous time, i.e., inertial navigation. The previous pose can be, for example, the previously determined second pose 210, or if the most recent second pose 210 is not available, the previous GNSS pose 215, or if the most recent GNSS pose 215 is not available, the previous odometry pose 225.
[0048] The lane localization frame 255 and the full localization frame 260 employ a similar process, an overview of which is given here and a more comprehensive description is provided below. In both the lane localization frame 255 and the full localization frame 260, the computer 105 generates a binary image 320 depicting static environmental features 310, generates a distance transform image 315 from the binary image 320, projects the map data 245 indicating static environmental features 310 onto the distance transform image 315, and determines the pose 205, 210 of the vehicle 100 that minimizes the value of a cost function. The value of the cost function is based on the map data 245 indicating static environmental features 310 and the distance transform image 315. The minimization of the cost function starts with the pose 205, 210 at an initial pose. The difference between the lane localization frame 255 and the full localization frame 260 is the initial pose and the perspective. The lane localization frame 255 uses the GNSS pose 215 or the odometry pose 225 as the initial pose, and the full localization frame 260 uses the first pose 205 output by the lane localization frame 255 as the initial pose. In the lane localization frame 255, the binary image 320 and the distance transform image 315 are from a top-down perspective (also known as a bird's-eye view) and will be referred to as the top-down binary image 320a and the top-down distance transform image 315a, respectively. Figure 3B The top-down binary image 320 is shown, and Figure 3CShows a top-down distance-transformed image 315a. In the full localization box 260, the binary image 320 and the distance-transformed image 315 are from the perspective of the camera 110, and will be referred to as the image-plane binary image 320b and the image-plane distance-transformed image 315b, respectively; that is, the image-plane binary image 320b and the image-plane distance-transformed image 315b have the same image plane as the camera image 305. Figure 4B Shows the image-plane binary image 320b, and Figure 4C and Figure 5B Shows the image-plane distance-transformed image 315b. Using different perspectives in the lane localization box 255 and the full localization box 260 results in high accuracy in the three-dimensional space by covering different dimensions in the three-dimensional space.
[0049] Reference Figure 3A 、 Figure 4A and Figure 5A , the computer 105 can be programmed to receive the camera image 305 from one of the cameras 110 via the communication network 115. The camera image 305 is a two-dimensional pixel matrix. The brightness or color of each pixel is represented as one or more numerical values, for example, a scalar unitless value of photometric light intensity between 0 (black) and 1 (white), or the values of each of red, green, and blue, for example, each on an 8-bit scale (0 to 255) or a 12-bit or 16-bit scale. The pixels can be a mixed representation, such as a repeating pattern of three pixels and a scalar value of the intensity of a fourth pixel with three numerical color values, or some other pattern. The position in the camera image 305 (i.e., the position in the field of view of the corresponding camera 110 when recording the camera image 305) can be specified by pixel dimensions or coordinates, for example, a pair of ordered pixel distances, such as the number of pixels from the top edge of the camera image 305 and the number of pixels from the left edge of the camera image.
[0050] The computer 105 is programmed to detect static environmental features 310 in the camera image 305 for Figure 2 the lane detection box 235 and the vertical detection box 240 in. The computer 105 can use any conventional object recognition technique suitable for identifying the type of static environmental features 310 of interest (such as lane lines, power line poles, etc.). For example, the computer 105 can execute a machine learning model, such as a deep neural network, such as PersFormer for detecting lane lines or YoloV5 for detecting linear vertical features. The machine learning model can be trained on camera images from the same perspective as the cameras 110 on the vehicle 100, and the camera images are annotated with the identification of the static environmental features 310 to be used as ground truth.
[0051] The computer 105 can be programmed to determine a geometric description of at least some of the static environmental features 310 (e.g., lane lines). The geometric description specifies the shape of the static environmental feature 310 in two - dimensional or three - dimensional space, e.g., a series of coordinate points in space or the formula of a spline curve. For example, the computer 105 can execute a machine - learning model, e.g., the same machine - learning model used to detect the static environmental feature 310, e.g., PersFormer.
[0052] Reference Figure 3B , for the lane localization box 255, the computer 105 can be programmed to project the static environmental features 310 (e.g., lane lines) detected in the camera image 305 to a bird's - eye view. For example, as is known, the computer 105 can use inverse perspective mapping. Again, for example, the computer 105 can use the horizontal coordinates from the geometric description of the static environmental feature 310, e.g., the coordinates (x, y) from the position (x, y, z).
[0053] Reference Figure 3B and Figure 4B , for both the lane localization box 255 and the full - localization box 260, the computer 105 is programmed to generate a binary image 320 depicting the static environmental feature 310. The binary image 320 is a two - dimensional pixel matrix. Each pixel is a binary variable, i.e., takes on one of two values, e.g., 0 or 1. The position in the binary image 320 can be specified in pixel dimensions or coordinates, e.g., a pair of ordered pixel distances, such as the number of pixels from the top edge of the binary image 320 and the number of pixels from the left edge of the binary image. The image - plane binary image 320b can have the same pixel dimensions as the camera image 305. The binary image 320 can depict only the static environmental feature 310. For example, the pixels occupied by the static environmental feature 310 can have the value 1 (depicted as black in Figure 3B and Figure 4B ), and the remaining pixels can have the value 0 (depicted as white in Figure 3B and Figure 4B ).
[0054] Reference Figures 3C to 3E , Figure 4C and Figure 5B, the distance transform image 315 is a two-dimensional pixel matrix. Each pixel has a pixel value, such as a scalar pixel value. Positions in the distance transform image 315 can be specified in pixel dimensions or coordinates, e.g., a pair of ordered pixel distances such as the number of pixels from the top edge of the distance transform image 315 and the number of pixels from the left edge of the distance transform image. The distance transform image 315 can have the same perspective as the corresponding binary image 320. The distance transform image 315 can have the same pixel dimensions as the corresponding binary image 320; i.e., the top-down view of the distance transform image 315a can have the same pixel dimensions as the top-down view of the binary image 320a, and the image plane top-down view of the distance transform image 315a can have the same pixel dimensions as the image plane binary image 320b. Each pixel value indicates the pixel distance of that pixel from the static environment feature 310 in the distance transform image 315. For example, each pixel can have a pixel value equal to the pixel distance from that pixel to the nearest pixel that is part of the static environment feature 310. Figure 3C , Figure 4C and Figure 5B Pixels that coincide with the static environment feature 310 (i.e., pixel distance is zero) are represented as black, and increasing pixel distances are represented as gradually lighter shades. Figure 3D shows a three-dimensional plot of the pixel values from Figure 3C , where the horizontal axis corresponds to pixel coordinates and the vertical axis corresponds to pixel values. Figure 3E is Figure 3D a cross-section of the plot of x . The pixel distance can be the Euclidean distance. For example, if (p y ) are the coordinates of a pixel and (p i , p j ) are the coordinates of the nearest pixel that is part of the static environment feature 310, then the pixel value D at (p x , p y ) can be given by the following expression:
[0055]
[0056] The computer 105 is programmed to generate a distance transform image 315 of the static environment feature 310 based on the corresponding binary image 320. The computer 105 can determine the pixel value D of each pixel in the distance transform image 315 as the pixel distance from the nearest pixel (e.g., whose value is 1) indicated by the binary image 320 to be occupied.
[0057] Refer to Figure 6, the computer 105 is programmed to determine the poses 205, 210 of the vehicle 100 based on a comparison of map data 245 indicating static environmental features 310 with a distance transform image 315. The computer 105 can determine the poses 205, 210 that minimize the value of a cost function. The computer 105 can start from an initial pose and execute an optimization algorithm that iteratively refines the poses 205, 210. At each iteration, the computer 105 projects the map data 245 indicating static environmental features 310 onto the distance transform image 315 based on the current poses 205, 210, calculates the value of the cost function, and adjusts the current poses 205, 210 for use in the next iteration. The computer 105 can execute the optimization algorithm until a convergence condition is met. Figure 6 The static environmental features 310 detected in the camera image 305 are represented by black lines, the map data 245 projected with the initial pose is represented by gray lines, and the map data 245 projected with the poses 205, 210 after convergence is represented by white lines.
[0058] The computer 105 can be programmed to project the map data 245 indicating static environmental features 310 onto the distance transform image 315 based on the current values of the poses 205, 210. For example, the map data 245 can be represented as a set of three-dimensional map points, and the poses 205, 210 can be represented as geometric transformation matrices. For each map point, the computer 105 can calculate the product of the inverse matrix of the pose 205, 210 and the map point, thereby generating a point in three-dimensional space, and apply a projection model that transforms the point into pixel coordinates in the distance transform image 315. For example, as shown in the following expression:
[0059]
[0060] where π is the projection model, W Tc is the pose 205, 210, and i X W is the i-th map point. The projection model is selected to transform a point in three-dimensional space into corresponding pixel coordinates in the same perspective (top-down perspective or the perspective of the camera 110) as the distance transform image 315.
[0061] The computer 105 can be programmed to calculate the value of a cost function based on the map data 245 indicating the static environmental features 310 and the distance transform image 315. For example, the pixel values of the pixels in the distance transform image 315 onto which the map data 245 is projected. For example, the cost function can be the sum of the pixel values of each pixel onto which one of the map points is projected, for example, the square of the pixel value, as shown in the following expression:
[0062]
[0063] where N is the number of map points, and D() is the pixel value of the pixel coordinates provided as an argument.
[0064] The computer 105 can be programmed to determine the poses 205, 210 of the vehicle 100 that minimize the value of the cost function, e.g., as shown in the following expression:
[0065]
[0066] The computer 105 can determine the poses 205, 210 that minimize the value of the cost function by executing an optimization algorithm. The computer 105 can use any suitable optimization algorithm, such as iterative non-linear least squares optimization.
[0067] Regarding determining the first pose 205, the computer 105 can detect generally horizontal static environmental features 310, such as lane lines, in the camera image 305; project those static environmental features 310 to a top-down view; generate a top-down binary image 320a from the projected static environmental features 310; generate a top-down distance transform image 315a from the top-down binary image 320a; initialize the first pose 205 at the GNSS pose 215, i.e., use the GNSS pose 215 as the initial pose; and determine the first pose 205 that minimizes the cost function, where the cost function includes a projection model in the top-down view.
[0068] Regarding determining the second pose 210, the computer 105 can detect static environmental features 310 in the camera image 305; generate an image plane binary image 320b from the detected static environmental features 310; generate an image plane distance transform image 315b from the planar image binary image 320b; initialize the second pose 210 at the first pose 205, i.e., use the first pose 205 as the initial pose; and determine the second pose 210 that minimizes the cost function, where the cost function includes a projection model in the image plane view. The computer 105 can detect the static environmental features 310 by detecting linear vertical features and using the generally horizontal static environmental features 310 that have already been detected for determining the first pose 205. The computer 105 does not need to project the detected static environmental features 310 to the view of the camera 110 because the camera image 305 is already in the view of the camera 110.
[0069] Figure 7FIG. is a flowchart showing an exemplary process 700 for determining the attitude 205, 210 of a vehicle 100. The memory of the computer 105 stores executable instructions for performing the steps of process 700, and / or can be programmed in a structure such as that mentioned above. As an overall overview of process 700, the computer 105 receives camera images 305 from the camera 110, GNSS attitude 215 from the GNSS receiver 120, and sensor data from the other sensors 125; estimates the ranging attitude 225; initializes the first attitude 205; detects the substantially horizontal features, such as lane lines; projects the detected static environment features 310 into a top-down view; generates a top-down binary image 320a; generates a top-down distance transform image 315a; determines the first attitude 205; detects the remaining static environment features 310; generates an image plane binary image 320b; generates an image plane distance transform image 315b; determines the second attitude 210; and actuates components of the vehicle 100 based on the second attitude 210.
[0070] Process 700 begins at block 705, where the computer 105 receives the camera image 305 from the camera 110, the GNSS attitude 215 from the GNSS receiver 120, and the sensor data from the other sensors 125, as described above.
[0071] Next, at block 710, the computer 105 determines the ranging attitude 225 based on the sensor data from the other sensors 125, as described above.
[0072] Next, at block 715, the computer 105 initializes the first attitude 205 to the GNSS attitude 215 from block 705 or the ranging attitude 225 from block 710, as described above.
[0073] Next, at block 720, the computer 105 detects substantially horizontal static environment features 310, such as lane lines, in the camera image 305 from block 705, as described above.
[0074] Next, at block 725, the computer 105 projects the static environment features 310 detected at block 720 into a top-down view, as described above.
[0075] Next, at block 730, the computer 105 generates a top-down binary image 320a based on the projected static environment features 310 from block 725, as described above.
[0076] Next, at block 735, the computer 105 generates a top-down distance transform image 315a based on the top-down binary image 320a from block 730, as described above.
[0077] Next, in block 740, computer 105 determines a first pose 205 based on a comparison of map data 245 indicating static environmental features 310 with the top-down distance transform image 315a from block 735. Computer 105 may determine the first pose 205 that minimizes the value of a cost function, starting from the first pose 205 set to the initial pose from block 715, as described above.
[0078] Next, in block 745, computer 105 detects linear vertical features in the camera image 305 from block 705, as described above.
[0079] Next, in block 750, computer 105 generates an image-plane binary image 320b based on the static environmental features 310 from blocks 725 and 745, as described above.
[0080] Next, in block 755, computer 105 generates an image-plane distance transform image 315b based on the image-plane binary image 320b from block 750, as described above.
[0081] Next, in block 760, computer 105 determines a second pose 210 based on a comparison of map data 245 indicating static environmental features 310 with the image-plane distance transform image 315b from block 755. Computer 105 may determine the second pose 210 that minimizes the value of a cost function, starting from the second pose 210 set to the first pose 205 from block 740, as described above.
[0082] Next, in block 765, computer 105 actuates components of vehicle 100 based on the postures 205, 210 of vehicle 100. The components can include, for example, propulsion system 130, braking system 135, steering system 140, and / or user interface 145. Computer 105 can actuate the components based on the second posture 210, which means actuating the components indirectly based on the first posture 205 because the first posture 205 is an input for determining the second posture 210. For example, computer 105 can actuate the components when executing an Advanced Driver Assistance System (ADAS). ADAS is an electronic technology that assists a driver in achieving driving functions and parking functions. Examples of ADAS include forward proximity detection, lane departure detection, blind spot detection, braking actuation, adaptive cruise control, and lane keeping assist systems. For example, as part of a lane centering feature, computer 105 can actuate steering system 140 based on the distance from lane lines, e.g., steering to prevent vehicle 100 from driving too close to a lane line. Computer 105 can use the detections and / or map data 245 from block 720 to identify lane lines. Computer 105 can determine the position of vehicle 100 relative to the lane lines based on the second posture 210 of vehicle 100. If the position of vehicle 100 is within a distance threshold of one of the lane lines, computer 105 can instruct steering system 140 to actuate to steer vehicle 100 towards the center of the lane. For another example, computer 105 can operate vehicle 100 autonomously (i.e., actuate propulsion system 130, braking system 135, and steering system 140 based on the second posture 210 of vehicle 100) to, for example, navigate vehicle 100 through a certain area.
[0083] Generally, the described computing systems and / or devices can employ any of a variety of computer operating systems, including but not limited to the following versions and / or varieties: Ford Applications; AppLink / Smart Device Link middleware; Microsoft Operating System; Microsoft Operating System; Unix operating system (e.g., the operating system released by Oracle Corporation of Redwood Shores, California); AIX UNIX operating system released by International Business Machines Corporation of Armonk, New York; Linux operating system; Mac OSX and iOS operating systems released by Apple Inc. of Cupertino, California; BlackBerry operating system released by BlackBerry Limited of Waterloo, Canada; and Android operating system developed by Google Inc. and the Open Handset Alliance; or provided by QNX Software Systems Limited CAR information entertainment platform. Examples of computing devices include, but are not limited to, in-vehicle computers, computer workstations, servers, desktops, notebooks, laptop computers, or handheld computers, or some other computing system and / or device.
[0084] The computing device generally includes computer-executable instructions, where the instructions can be executed by one or more computing devices such as those listed above. The computer-executable instructions can be compiled or interpreted from computer programs created using a variety of programming languages and / or technologies, which alone or in combination include, but are not limited to, Java TM , C, C++, Matlab, Simulink, Stateflow, Visual Basic, Java Script, Python, Perl, HTML, etc. Some of these applications can be compiled and executed on virtual machines such as the Java virtual machine, the Dalvik virtual machine, etc. Generally, a processor (e.g., a microprocessor) receives instructions from, for example, a memory, a computer-readable medium, etc., and executes these instructions, thereby performing one or more processes, including one or more of the processes described herein. Such instructions and other data can be stored and transmitted using a variety of computer-readable media. Files in a computing device are typically a collection of data stored on a computer-readable medium such as a storage medium, random access memory, etc.
[0085] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory (e.g., tangible) medium that participates in providing data (e.g., instructions) that can be read by a computer (e.g., by a processor of the computer). Such media can take many forms, including but not limited to non-volatile media and volatile media. Instructions can be transmitted through one or more transmission media, which include optical fibers, wires, wireless communication, including internal components that make up the system bus coupled to the processor of the computer. Common forms of computer-readable media include, for example, RAM, PROM, EPROM, FLASH-EEPROM, any other memory chip or cartridge, or any other medium from which a computer can read.
[0086] The databases, data repositories, or other data stores described herein may include various mechanisms for storing, accessing, and retrieving various data, including hierarchical databases, sets of files in a file system, application databases in a proprietary format, relational database management systems (RDBMSs), non-relational databases (NoSQL), graph databases (GDBs), etc. Each such data store is typically included within a computing device employing a computer operating system such as one of those mentioned above, and is accessed via a network in any one or more of various ways. A file system can be accessed from a computer operating system and can include files stored in various formats. In addition to languages for creating, storing, editing, and executing stored programs such as the PL / SQL language mentioned above, an RDBMS typically also employs the Structured Query Language (SQL).
[0087] In some examples, system components may be implemented as computer-readable instructions (e.g., software) on one or more computing devices (e.g., servers, personal computers, etc.) and stored on a computer-readable medium associated therewith (e.g., disk, memory, etc.). A computer program product may include such instructions stored on a computer-readable medium for performing the functions described herein.
[0088] In the drawings, like reference numerals indicate like elements. Additionally, some or all of these elements may be varied. Regarding the media, processes, systems, methods, heuristics, etc. described herein, it should be understood that while the steps of such processes, etc. have been described as occurring in a certain ordered sequence, such processes may be practiced by performing the steps in an order different from that described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. The operations, systems, and methods described herein should always be implemented and / or performed in accordance with applicable owner / user manuals and / or safety guidelines.
[0089] The present disclosure has been described in an illustrative manner, and it should be understood that the terms used are of a descriptive nature and not restrictive. The adjectives "first" and "second" are used throughout this document as identifiers and are not intended to indicate importance, order, or quantity. The use of "in response to," "after determining...," etc. indicates a causal relationship and not just a temporal relationship. Given the above teachings, many modifications and variations of the present disclosure are possible, and the present disclosure may be practiced in other ways than specifically described.
[0090] According to the present invention, there is provided a computer having a processor and a memory, the memory storing instructions executable by the processor to: detect static environmental features in a camera image from a vehicle camera; generate a distance transform image of the static environmental features as detected in the camera image, wherein a pixel value of a corresponding pixel in the distance transform image indicates a distance of the corresponding pixel from a corresponding pixel of the static environmental features in the distance transform image; and determine a pose of the vehicle based on a comparison of map data indicating the static environmental features with the distance transform image.
[0091] According to an embodiment, the instructions further include instructions for: calculating a value of a cost function based on the map data indicating the static environmental features and the distance transform image; and determining the pose of the vehicle that minimizes the value of the cost function.
[0092] According to an embodiment, the instructions further include instructions for: projecting the map data indicating the static environmental features onto the distance transform image; and calculating the value of the cost function based on the pixel values of the pixels onto which the map data is projected.
[0093] According to an embodiment, the pose is a first pose, and the instructions further include instructions for: determining a GNSS pose based on Global Navigation Satellite System (GNSS) data; and initializing the first pose at the GNSS pose to minimize the value of the cost function.
[0094] According to an embodiment, the static environmental features include lane lines.
[0095] According to an embodiment, the distance transform image is an overhead distance transform image from an overhead perspective.
[0096] According to an embodiment, the pose is a first pose, and the instructions further include instructions for: generating an image plane distance transform image from the perspective of the camera; and determining a second pose based on the first pose and a comparison of the map data indicating the static environmental features with the image plane distance transform image.
[0097] According to an embodiment, the instructions further include instructions for: calculating a value of a cost function based on the map data indicating the static environmental features and the image plane distance transform image; and determining the second pose of the vehicle that minimizes the value of the cost function.
[0098] According to an embodiment, the instructions further include instructions for performing the following operation: initializing the second pose at the first pose to minimize the value of the cost function.
[0099] According to an embodiment, the first pose includes only two horizontal spatial dimensions and a heading; and the second pose includes three spatial dimensions and three angular dimensions.
[0100] According to an embodiment, the distance transform image is an image plane distance transform image from the perspective of the camera.
[0101] According to an embodiment, the static environment features include linear vertical features.
[0102] According to an embodiment, the instructions further include instructions for performing the following operations: generating a binary image depicting the static environment features; and generating the distance transform image from the same perspective as the binary image based on the binary image.
[0103] According to an embodiment, the binary image depicts only the static environment features.
[0104] According to an embodiment, the pose includes two horizontal spatial dimensions and a heading.
[0105] According to an embodiment, the instructions further include instructions for actuating components of the vehicle based on the pose of the vehicle.
[0106] According to the present invention, a method includes: detecting static environment features in a camera image from a camera of a vehicle; generating a distance transform image of the static environment features detected in the camera image, wherein the pixel value of a corresponding pixel in the distance transform image indicates the distance of the corresponding pixel from the corresponding pixel of the static environment features in the distance transform image; and determining the pose of the vehicle based on a comparison of map data indicating the static environment features with the distance transform image.
[0107] In one aspect of the present invention, the method includes: calculating a value of a cost function based on the map data indicating the static environment features and the distance transform image; and determining the pose of the vehicle that minimizes the value of the cost function.
[0108] In one aspect of the present invention, the method includes: projecting the map data indicating the static environment features onto the distance transform image; and calculating the value of the cost function based on the pixel values of the pixels onto which the map data is projected.
[0109] In one aspect of the present invention, the method includes: generating a binary image depicting the characteristics of the static environment; and generating the distance transformation image from the same perspective as the binary image based on the binary image.
Claims
1. A method comprising: detecting static environmental features in a camera image from a camera of the vehicle; generating a range transformed image of the static environmental feature as detected in the camera image, wherein a pixel value of a corresponding pixel in the range transformed image indicates a corresponding pixel distance of the corresponding pixel from the static environmental feature in the range transformed image; and A pose of the vehicle is determined based on a comparison of map data indicative of features of the static environment and the range transform image.
2. The method of claim 1, further comprising: calculating a value of a cost function based on the map data indicating the static environment features and the range transform image; as well as The pose of the vehicle that minimizes the value of the cost function is determined.
3. The method of claim 2, further comprising: projecting the map data indicative of the static environment features onto the range transformed image; as well as The value of the cost function is calculated based on the pixel value of the pixel onto which the map data is projected.
4. The method of claim 2, wherein the gesture is a first gesture, the method further comprising: Determine GNSS attitude based on Global Navigation Satellite System (GNSS) data; as well as The first attitude is initialized at the GNSS attitude to minimize the value of the cost function. The method of claim 1 , wherein the static environmental features include lane lines. The method of claim 1 , wherein the distance transformed image is a top-down distance transformed image from a top-down perspective.
7. The method of claim 6, wherein the gesture is a first gesture, the method further comprising: generating an image plane range transform image from the viewpoint of the camera; as well as A second pose is determined based on the first pose and based on a comparison of the map data indicative of the static environment feature and the image plane distance transform image.
8. The method of claim 7, further comprising: calculating a value of a cost function based on the map data indicating the static environment features and the image plane distance transform image; as well as The second pose of the vehicle that minimizes the value of the cost function is determined.
9. The method of claim 8, further comprising initializing the second pose at the first pose to minimize the value of the cost function.
10. The method of claim 7, wherein The first posture includes only two horizontal spatial dimensions and heading; and The second posture includes three spatial dimensions and three angular dimensions.
11. The method of claim 1, wherein the range transform image is an image plane range transform image from the perspective of the camera.
12. The method of claim 1, further comprising: generating a binary image characterizing the static environment; as well as The distance transformed image is generated based on the binary image from the same viewing angle as the binary image. The method of claim 12 , wherein the binary image depicts only the static environmental features.
14. The method of claim 1 further comprising actuating a component of the vehicle based on the posture of the vehicle.
15. A computer comprising a processor and a memory, the memory storing instructions executable by the processor to perform the method of one of claims 1 to 14.