Positioning method, positioning system and vehicle

By encoding and storing laser features and visual semantic information in the same positioning map, and integrating laser positioning and visual positioning technology, the problem of excessive data and computing overhead in the existing technology is solved, efficient positioning effect is achieved, and the needs of autonomous driving in all scenarios is met.

CN114200481BActive Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010884916.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-28
Publication Date
2025-05-16
Estimated Expiration
2040-08-28

AI Technical Summary

Technical Problem

The existing laser positioning technology based on laser feature matching and visual positioning technology based on visual feature matching have too much data storage and calculation overhead, which is inefficient and cannot meet the needs of autonomous driving in all scenarios.

Method used

The laser feature and visual semantic information encoding are stored in the same positioning map, and the positioning efficiency is improved and data and calculation overhead is reduced by integrating laser positioning and visual positioning technology.

Benefits of technology

It realizes the efficient integration of laser positioning and visual positioning technology, reduces data and calculation overhead during positioning, improves positioning efficiency, and meets the needs of autonomous driving in all scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114200481B_ABST
    Figure CN114200481B_ABST
Patent Text Reader

Abstract

The present application provides a positioning method, a positioning system, and a vehicle. Among them, the method includes generating N sampling points C1 to C around the initial pose N , respectively determining first weights of the current estimated poses corresponding to the sampling points based on the matching of laser features, determining second weights of the current estimated poses corresponding to the sampling points based on the matching of visual semantic information and road surface semantic information (i.e., the matching of visual features), and then weighted averaging the current estimated poses according to the first weights and the second weights to obtain the actual pose of the vehicle, thereby realizing the fusion of the laser positioning technology based on laser features and the visual positioning technology based on features, and improving the positioning efficiency. Moreover, the technical solution provided by the embodiments of the present application encodes and stores the laser features and the visual semantic information in the same positioning map, realizing the fusion of map data and reducing the data overhead and computational overhead generated during the positioning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a positioning method, a positioning system and a vehicle. Background Art

[0002] Autonomous driving needs to solve three core driving problems: where is it? (vehicle positioning); where is it going? (determine the destination); how to get there? (route planning). Among them, positioning technology is mainly used to solve the problem of "where is it?" and is one of the key technologies necessary to realize autonomous driving.

[0003] Signal-based positioning technology uses satellite signals or 5G signals to locate the vehicle, so it has the ability of global positioning and is currently the most widely used positioning technology. However, since GNSS satellite signals are easily blocked by high buildings, mountains, etc., signal-based positioning technology cannot provide accurate positioning when the vehicle is driving in cities, tunnels and other road conditions, and therefore cannot meet the needs of achieving full-scenario autonomous driving. To solve this problem, the current research direction in the field of autonomous driving is to integrate signal-based positioning technology, dead reckoning-based positioning technology, and feature matching-based positioning technology to make up for the shortcomings of signal-based positioning technology.

[0004] At present, for positioning technology based on environmental feature matching, laser positioning technology based on laser feature matching supplements or replaces signal-based positioning technology, or visual positioning technology based on visual feature matching supplements or replaces signal-based positioning technology. These are two independent technical routes. The map data relied on by laser positioning and visual positioning are stored separately, and the algorithms run independently, resulting in excessive data storage and computing overheads, high requirements on hardware systems such as the vehicle's control unit ECU, and low efficiency. Summary of the invention

[0005] The embodiments of the present application provide a positioning method, a positioning system and a vehicle, which can integrate laser positioning technology based on laser feature matching and visual positioning technology based on visual feature matching to improve positioning efficiency and reduce data and computing overhead generated during the positioning process.

[0006] In a first aspect, an embodiment of the present application provides a positioning method, the method comprising: determining an initial position of a vehicle, generating N sampling points C1 to C2 around the initial position, N , N is a positive integer; extract the first laser feature and at least one visual semantic information from the positioning map according to the current predicted posture P of the vehicle; for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n, the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of , the second laser feature is extracted from the point cloud data collected by the laser radar; and, for any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of at least one road semantic information is extracted from the image data collected by the camera; according to N sampling points C1~C N The current estimated pose P1-P N and its first weight and second weight, calculate the current estimated pose P 1- P N The weighted average of the weighted average is used as the current posture of the vehicle; wherein the positioning map includes a plurality of positioning pictures spliced ​​together, the positioning picture includes a color channel, and the encoding of the first laser feature and the encoding of the visual semantic information are stored in the color channel.

[0007] The technical solution provided in the embodiment of the present application can determine the first weight of the current estimated posture based on the matching of laser features, determine the second weight of the current estimated posture based on the matching of visual semantic information and road semantic information (i.e., the matching of visual features), and then calculate the weighted average of the current estimated posture based on the first weight and the second weight, and use the weighted average as the current posture of the vehicle, thereby realizing the fusion of the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching, and improving the positioning efficiency. In addition, the technical solution provided in the embodiment of the present application encodes and stores the laser features and visual semantic information in the same positioning map, realizing the fusion of map data and reducing the data overhead and computing overhead generated during the positioning process.

[0008] In an optional implementation, the positioning image includes a first color channel, and the first color channel is used to store the encoding of the visual semantic information. In this way, the method provided in the embodiment of the present application can decode the first color channel to obtain the visual semantic information.

[0009] In an optional implementation, the positioning image further includes a second color channel, and the second color channel is used to store the encoding of the first laser feature. In this way, the method provided in the embodiment of the present application can decode the second color channel to obtain the first laser feature.

[0010] In an optional implementation, the encoding of the visual semantic information includes at least one of a flag bit, a type code, and a brightness code; the flag bit is used to indicate the type of the road sign, the type code is used to indicate the content of the road sign, and the brightness information code is used to indicate the brightness information of the image. In this way, the encoding of the visual semantic information can determine what kind of road sign the visual semantic information contains, such as a white dotted line, a white solid line, a straight sign, etc., so as to facilitate matching with the road semantic information.

[0011] In an optional implementation, at least one visual semantic information is extracted by the following steps: extracting a local positioning map from the positioning map according to the current predicted posture, the local positioning map comprising M positioning pictures, the M positioning pictures comprising a first picture where the current predicted posture is located, and M-1 second pictures near the first picture, where M is a positive integer greater than 1; extracting at least one visual semantic information from the local positioning map. In this way, the visual semantic information can include information such as road signs around the current predicted posture, so as to facilitate matching with the road semantic information in the image data around the vehicle collected by the camera to determine the second weight.

[0012] In an alternative implementation, for any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of comprises: determining at least one valid road semantic information from at least one road semantic information, wherein the number of pixels of each valid road semantic information is within a preset range; and determining at least one valid road semantic information from at least one road semantic information according to the current estimated posture P. n Project at least one valid road semantic information into the coordinate system of the local positioning map; determine the semantic association relationship between at least one valid road semantic information and at least one visual semantic information; perform semantic matching on each pair of semantically associated valid road semantic information and visual semantic information, and determine a second weight value according to the semantic matching result. In this way, by determining the valid road semantic information, some incomplete or oversized misdetected visual semantic information can be filtered out, thereby improving the semantic matching efficiency and reducing the computational complexity of the semantic matching.

[0013] In an optional implementation, determining the semantic association relationship between at least one valid road semantic information and at least one visual semantic information includes: calculating any valid road semantic information a i The semantic weights and any visual semantic information b j The semantic weight of the effective road semantic information a i The semantic weight and visual semantic information b j The difference of the semantic weights of the effective road semantic information a iand visual semantic information b j When the semantic association is less than the preset first threshold, the effective road semantic information a is determined. i and visual semantic information b j Thus, when the semantic relevance is less than the first threshold, it indicates that the effective road semantic information a i The semantic weight and visual semantic information b j The semantic weights of are close, which indicates that the effective road semantic information a i and visual semantic information b j Have semantic associations.

[0014] In an optional implementation, semantic matching is performed on each pair of semantically associated effective road semantic information and visual semantic information, and a second weight is determined based on the semantic matching result, including: calculating the matching distance of each pair of semantically associated effective road semantic information and visual semantic information respectively; performing a weighted summation of the calculated matching distances to obtain a total matching distance; and determining the second weight based on the total matching distance. In this way, the second weight can reflect the current estimated posture P n The closeness to the vehicle's true pose.

[0015] In an optional implementation, according to N sampling points C1-C N The current estimated pose P1-P N and its first and second weights to calculate the current estimated pose P 1- P N The weighted average of the weighted average is used as the current posture of the vehicle, including: using the current estimated posture P1-P N The first weight of the current estimated pose P1-P N Weighted averaging is performed to obtain the first weighted average value; the current estimated pose P1-P N The second weight of the current estimated pose P1-P N The weighted average is obtained to obtain a second weighted average; the first weighted average and the second weighted average are weighted averaged to obtain a weighted average, and the weighted average is used as the current posture of the vehicle. In this way, by using the first weight and the second weight to weight the current estimated posture, the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching are integrated.

[0016] In an optional implementation, the road semantic information includes: a pixel block containing at least one road sign, the number of pixels in the pixel block, and the type of road sign to which each pixel belongs. In this way, the type and size of the road sign can be determined through the road semantic information to facilitate matching with the road semantic information.

[0017] In an optional implementation, the current predicted posture is determined by the following steps: determining the relative posture of the vehicle between the current time t and the first historical time t-1 according to the odometer data; n The predicted posture corresponding to the first historical moment t-1 is added to the relative posture to obtain the current predicted posture.

[0018] In an optional implementation, when a preset ratio of the first weight or the second weight is lower than the second threshold, N sampling points C1 to C N In this way, the divergence of sampling points as the number of positioning times increases can be eliminated.

[0019] In a second aspect, an embodiment of the present application provides a positioning system, including: a GNSS / INS combination module, a control unit, a memory, a laser radar and a camera installed on a vehicle; the GNSS / INS combination module is used to determine the initial position and posture of the vehicle; the memory is used to store a positioning map, the positioning map includes a plurality of positioning pictures spliced ​​together, the positioning picture includes a color channel, and the color channel stores the encoding of the first laser feature and the encoding of the visual semantic information; the laser radar is used to collect point cloud data, the point cloud data includes the second laser feature; the camera is used to collect image data, the image data includes at least one road semantic information; the control unit is used to generate N sampling points C1~C around the initial position and posture. N , N is a positive integer; the control unit is further used to extract the first laser feature and at least one visual semantic information from the positioning map according to the current predicted position of the vehicle; the control unit is also used to extract the first laser feature and at least one visual semantic information from the positioning map for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of the second laser feature is extracted from the point cloud data collected by the laser radar; the control unit is also used for any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of at least one road semantic information is extracted from the image data collected by the camera; the control unit is also used to extract the road semantic information according to the N sampling points C1~C N The current estimated pose P1-P N and its first and second weights to calculate the current estimated pose P 1- P N The weighted average of is taken as the current posture of the vehicle.

[0020] The technical solution provided in the embodiment of the present application can determine the first weight of the current estimated posture based on the matching of laser features, determine the second weight of the current estimated posture based on the matching of visual semantic information and road semantic information (i.e., the matching of visual features), and then calculate the weighted average of the current estimated posture based on the first weight and the second weight, and use the weighted average as the current posture of the vehicle, thereby realizing the fusion of the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching, and improving the positioning efficiency. In addition, the technical solution provided in the embodiment of the present application encodes and stores the laser features and visual semantic information in the same positioning map, realizing the fusion of map data and reducing the data overhead and computing overhead generated during the positioning process.

[0021] In an optional implementation, the positioning picture includes a first color channel, and the first color channel is used to store the encoding of the visual semantic information. In this way, the positioning system provided in the embodiment of the present application can decode the first color channel to obtain the visual semantic information.

[0022] In an optional implementation, the positioning image further includes a second color channel, and the second color channel is used to store the encoding of the first laser feature. In this way, the positioning system provided in the embodiment of the present application can decode the second color channel to obtain the first laser feature.

[0023] In an optional implementation, the encoding of the visual semantic information includes at least one of a flag bit, a type code, and a brightness code; the flag bit is used to indicate the type of the road sign, the type code is used to indicate the content of the road sign, and the brightness information code is used to indicate the brightness information of the image. In this way, the encoding of the visual semantic information can determine what kind of road sign the visual semantic information contains, such as a white dotted line, a white solid line, a straight sign, etc., so as to facilitate matching with the road semantic information.

[0024] In an optional implementation, when the control unit is used to extract at least one visual semantic information from the positioning map according to the current predicted position of the vehicle: the control unit is specifically used to extract a local positioning map from the positioning map according to the current predicted position, the local positioning map includes M positioning pictures, the M positioning pictures include a first picture where the current predicted position is located, and M-1 second pictures near the first picture, M is a positive integer greater than 1; the control unit is also used to extract at least one visual semantic information from the local positioning map. In this way, the visual semantic information can include information such as road signs around the current predicted position, so as to facilitate matching with the road semantic information in the image data around the vehicle collected by the camera, so as to determine the semantic association relationship and the second weight.

[0025] In an optional implementation, when the control unit is used for any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n When the second weight is: a control unit, specifically used to determine at least one valid road semantic information from at least one road semantic information, and the number of pixels of each valid road semantic information is within a preset range; the control unit is also used to determine at least one valid road semantic information from at least one road semantic information according to the current estimated posture P n At least one valid road semantic information is projected into the coordinate system of the local positioning map; the control unit is further used to determine the semantic association relationship between at least one valid road semantic information and at least one visual semantic information; the control unit is further used to perform semantic matching on each pair of semantically associated valid road semantic information and visual semantic information, and determine the second weight according to the semantic matching result. In this way, by determining the valid road semantic information, some incomplete or oversized misdetected visual semantic information can be filtered out, thereby improving the efficiency of semantic matching and reducing the computational complexity of semantic matching.

[0026] In an optional implementation, when the control unit is used to determine the semantic association relationship between at least one valid road semantic information and at least one visual semantic information: the control unit is specifically used to calculate any valid road semantic information a i The semantic weights and any visual semantic information b j The control unit is also used to calculate the semantic weight of the effective road surface semantic information a i The semantic weight and visual semantic information b j The difference of the semantic weights of the effective road semantic information a i and visual semantic information b j The control unit is also used to determine the effective road semantic information a when the semantic association is less than a preset first threshold value. i and visual semantic information b j Thus, when the semantic relevance is less than the first threshold, it indicates that the effective road semantic information a i The semantic weight and visual semantic information b j The semantic weights of are close, which indicates that the effective road semantic information a i and visual semantic information b j Have semantic associations.

[0027] In an optional implementation, when the control unit is used to perform semantic matching on each pair of semantically associated effective road semantic information and visual semantic information, and determine the second weight according to the semantic matching result: the control unit is specifically used to calculate the matching distance of each pair of semantically associated effective road semantic information and visual semantic information respectively; the control unit is also used to perform the weighted sum of the calculated matching distances to obtain the total matching distance; the control unit is also used to determine the second weight according to the total matching distance. In this way, the second weight can reflect the current predicted posture P n The closeness to the vehicle's true pose.

[0028] In an optional implementation, when the control unit is used to select the N sampling points C1 to C N The current estimated pose P1-P N and its first weight and second weight, calculate the current estimated pose P 1- P N The weighted average of the weighted average is used as the current posture of the vehicle: a control unit, specifically for using the current estimated posture P1-P N The first weight of the current estimated pose P1-P N Weighted averaging is performed to obtain a first weighted average value; a control unit is also used to use the current estimated posture P1-P N The second weight of the current estimated pose P1-P N The weighted average is used to obtain a second weighted average; the control unit is also used to weight the first weighted average and the second weighted average to obtain a weighted average, and the weighted average is used as the current posture of the vehicle. In this way, by using the first weight and the second weight to weight the current estimated posture, the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching are integrated.

[0029] In an optional implementation, the road semantic information includes: a pixel block containing at least one road sign, the number of pixels in the pixel block, and the type of road sign to which each pixel belongs. In this way, the type and size of the road sign can be determined through the road semantic information to facilitate matching with the road semantic information.

[0030] In an optional implementation, the positioning system further includes an odometer; a control unit further configured to determine the relative position of the vehicle between the current time t and the first historical time t-1 according to the odometer data; and a control unit further configured to convert the sampling point C n The predicted posture corresponding to the first historical moment t-1 is added to the relative posture to obtain the current predicted posture.

[0031] In an optional implementation, the control unit is further configured to regenerate N sampling points C1 to C2 when a preset ratio of the first weight or the second weight is lower than a second threshold. N In this way, the divergence of sampling points as the number of positioning times increases can be eliminated.

[0032] In a third aspect, an embodiment of the present application provides a vehicle, which includes the positioning system provided by the second aspect of the embodiment of the present application and its various implementation methods.

[0033] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned aspects and methods of each implementation thereof.

[0034] In a fifth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the above-mentioned aspects and methods of each of its implementation methods.

[0035] In a sixth aspect, an embodiment of the present application further provides a chip system, which includes a processor for supporting the above-mentioned device or system to implement the functions involved in the above-mentioned aspects, for example, generating or processing the information involved in the above-mentioned method. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of a point cloud positioning map;

[0037] Figure 2 This is a schematic diagram of a picture-based location map.

[0038] Figure 3 It is a schematic diagram of the pixel quantity of sparse feature maps, semi-dense maps and dense maps;

[0039] Figure 4 It is a module configuration diagram of the positioning system of the autonomous driving vehicle;

[0040] Figure 5 It is a schematic diagram of the laser map;

[0041] Figure 6 is a diagram of the encoding format of visual semantic information provided by an embodiment of the present application;

[0042] Figure 7 It is a schematic diagram of the 8-bit encoding of the visual semantic information corresponding to the pixels in different areas on the positioning map;

[0043] Figure 8 is a hardware framework diagram of a positioning system for implementing a positioning method provided in an embodiment of the present application;

[0044] Fig. 9 is a flowchart of a positioning method provided in an embodiment of the present application;

[0045] Fig.10 It is a data flow block diagram involved in the positioning method provided in the embodiment of the present application;

[0046] Fig.11 A scheme for generating sampling points is exemplarily provided;

[0047] Fig.12 It is a schematic diagram of the odometer coordinate system;

[0048] Fig.13 is a schematic diagram of a process for determining a first weight provided in an embodiment of the present application;

[0049] Fig.14 is a schematic diagram of a process for determining a second weight provided in an embodiment of the present application;

[0050] Fig.15 is a flowchart of semantic matching provided by an embodiment of the present application;

[0051] Fig.16 is a flowchart of step S201 of the positioning method provided in an embodiment of the present application;

[0052] Fig.17 is a schematic diagram of calculating the pixel area of ​​road semantic information provided by an embodiment of the present application;

[0053] Fig.18 is a flowchart of step S203 of the positioning method provided in an embodiment of the present application;

[0054] Fig.19 is a flowchart of step S204 of the positioning method provided in an embodiment of the present application;

[0055] Fig. 20 is a schematic diagram of the matching distance provided in an embodiment of the present application;

[0056] Fig.21 is a flowchart of step S502 of the positioning method provided in an embodiment of the present application;

[0057] Fig. 22 is a flowchart of step S105 of the positioning method provided in an embodiment of the present application;

[0058] Fig.23 It is a software module block diagram of the positioning system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] Autonomous vehicles or self-driving automobiles are also called unmanned driving or computer driving. Autonomous vehicles can sense their surroundings with sensors, global navigation satellite systems (GNSS), and machine vision technologies, and determine their own positions, plan navigation routes, update map information, avoid obstacles, etc. based on the sensing data, and ultimately achieve automatic driving of vehicles without or with little human active operation.

[0060] Generally speaking, autonomous driving needs to solve three core problems of driving: where is it? (vehicle positioning); where is it going? (determine the destination); how to get there? (route planning). Among them, positioning technology is mainly used to solve the problem of "where is it?" and is one of the key technologies essential to realize autonomous driving.

[0061] At present, depending on the sensors relied on, the positioning technologies for autonomous driving can mainly include the following three types:

[0062] 1. Signal-based positioning technology.

[0063] This technology mainly realizes the positioning of the vehicle based on satellite signals or 5G signals. The current mainstream solution is to install a global navigation satellite system (GNSS) receiver in the vehicle to receive satellite signals from multiple GNSSs, and use the received satellite signals to calculate the global position of the vehicle in the space environment. GNSS ground stations can also be used in conjunction with GNSS satellites to improve positioning accuracy. Common GNSS systems include: Beidou Navigation Satellite System (BDS), Global Positioning System (GPS), etc.

[0064] 2. Positioning technology based on dead reckoning.

[0065] This technology requires the vehicle to be equipped with sensors such as an inertial measurement unit (IMU) and a wheel speed meter. Among them, the IMU can measure the angular velocity, acceleration and other information of the vehicle, and the wheel speed meter can measure the rotation speed of the wheel. Based on the sensor data, after determining the initial position of the vehicle, the current position (position and attitude) of the vehicle can be estimated according to the vehicle's dynamic equations.

[0066] 3. Positioning technology based on environmental feature matching.

[0067] At present, this technology mainly includes two forms: laser positioning and visual positioning. Based on laser sensors and visual sensors, respectively, the environmental information around the vehicle is obtained in real time, and the obtained environmental information is processed and matched with the pre-stored positioning map to determine the position of the vehicle. It can be understood that the implementation of this technology requires the pre-construction of the positioning map, and the positioning map has different construction methods depending on the different sensors used.

[0068] When laser positioning is used, positioning maps are mainly in point cloud and picture types.

[0069] ① Point cloud type: Use laser sensors to collect point cloud data, then filter the point cloud data to remove noise, and finally splice and superimpose the processed point cloud data together to form a Figure 1 The point cloud map shown.

[0070] ② Image type: The point cloud data after splicing in ① is processed by rasterization, and each raster is encoded as a pixel to convert the spliced ​​point cloud data into Figure 2 The positioning map in the picture format shown. Among them, the laser map in the picture format can be a single-channel grayscale picture or a three-channel color picture. The picture-based positioning map occupies a small storage space and can solve the problem of high storage resource overhead of the point cloud map.

[0071] When visual positioning is used, the positioning map (visual map) is constructed based on the way and number of pixels selected from the original image collected by the visual sensor. Figure 3 As shown, it can include sparse feature maps, semi-dense maps and dense maps. Figure 3 Currently, it is mainly saved in the form of point cloud.

[0072] ① Sparse feature map: The pixels corresponding to the feature points in the original image are saved in the map, so the number of pixels is minimal.

[0073] ② Semi-dense map: Some pixels in the original image, such as pixels with gradients, are saved in the map, so the number of pixels is medium.

[0074] ③Dense map: All pixels in the original image are saved in the map, so the number of pixels is the largest.

[0075] Signal-based positioning technology is currently the most widely used positioning technology because of its global positioning capabilities. However, since GNSS satellite signals are easily blocked by tall buildings, mountains, etc., signal-based positioning technology cannot provide accurate positioning when vehicles are driving in cities, tunnels, and other road conditions, and therefore cannot meet the needs of achieving full-scenario autonomous driving. To solve this problem, the current research direction in the field of autonomous driving is to integrate signal-based positioning technology, dead reckoning-based positioning technology, and feature matching-based positioning technology to make up for the shortcomings of signal-based positioning technology.

[0076] At present, for positioning technology based on environmental feature matching, laser positioning technology using laser feature matching to supplement or replace signal-based positioning technology, and visual positioning technology using visual feature matching to supplement or replace signal-based positioning technology, are two independent technical routes for integrated positioning. Since the technical routes are independent of each other, the map data of laser positioning and visual positioning are stored separately, and the algorithms are independent of each other, resulting in excessive data storage and computing overheads, high requirements for the vehicle's control unit ECU and other hardware systems, and low efficiency.

[0077] In order to solve the problems existing in the prior art, an embodiment of the present application provides a positioning method.

[0078] The technical solutions of the embodiments of the present application can be applied to various vehicles that adopt automatic driving technology or positioning technology, including but not limited to various means of transportation: such as vehicles (cars), ships, trains, subways, airplanes, etc., and various robots, such as: service robots, transport robots, automated guided vehicles (AGV), unmanned ground vehicles (UGV), etc., and various engineering machinery, such as: tunnel boring machines, etc.

[0079] The following uses a vehicle as an example to illustrate the hardware environment in which the technical solution of the embodiment of the present application is implemented.

[0080] like Figure 4 As shown, the vehicle is equipped with the following modules: a LiDAR 110, a camera 120, a GNSS / INS combined module 130, and a control unit 140. Among them:

[0081] The laser radar 110 is used to collect information about the distance between the elements in the environment (e.g., vehicles, obstacles, pedestrians, road signs, etc.) and the vehicle. The laser radar 110 can perform a 360-degree omnidirectional scan of the environment, or it can only scan the environment information within a partial range (e.g., 180 degrees) in front of the vehicle.

[0082] The camera 120 is used to collect image information around the vehicle. The camera 120 may have one (i.e., a monocular camera) or multiple (i.e., multi-cameras), and may collect 360-degree panoramic images of the surroundings, or may only collect images of a portion of the front of the vehicle.

[0083] The GNSS / INS combination module 130 may include devices such as a GNSS receiver and an inertial measurement unit (IMU) to achieve fusion positioning based on satellite signals and IMU.

[0084] The control unit 140 may be a core computing unit of the entire electronic system of the autonomous driving vehicle, such as a mobile data center (MDC), an electronic control unit (ECU), etc., which is used to process data generated by other modules and generate vehicle control information based on the processing results.

[0085] In addition, to realize other assisted driving functions of the vehicle, the vehicle can also be configured with the following modules:

[0086] The ultrasonic sensor 150 is used for short-distance distance measurement. For example, it is turned on during parking assistance to provide short-distance warning information.

[0087] The millimeter wave radar 160 is used for long-distance ranging. Since the millimeter wave has strong anti-interference ability and strong ability to penetrate fog, smoke and dust, the millimeter wave radar can work all day long, for example, it can be used to assist in obstacle ranging in bad weather conditions.

[0088] It is to be understood that the hardware environment illustrated in the embodiment of the present application does not constitute a specific limitation on the technical solution of the embodiment of the present application. In other embodiments of the present application, the hardware environment implemented by the technical solution of the embodiment of the present application may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0089] In the embodiment of the present application, in order to realize the integration of laser positioning and visual positioning technology and reduce storage overhead and computing overhead, the basic idea is: to integrate the map data in the positioning map of laser positioning (hereinafter referred to as laser map) and the positioning map of visual positioning (hereinafter referred to as visual map) into a positioning map, so that when laser positioning and visual positioning technology are applied at the same time, the data required for these two positioning methods can be obtained from the integrated positioning map, thereby reducing data storage overhead and computing overhead and improving positioning efficiency.

[0090] The following describes in detail how the positioning map is obtained in the embodiments of the present application with reference to some examples.

[0091] The basic idea of ​​obtaining a positioning map in the embodiment of the present application is to integrate the visual semantic information of the road surface markings into the laser map to obtain a positioning map. Among them, the road surface markings include information used to assist in determining the location of the vehicle and the lane it is in, such as: single yellow line, double yellow line, white dashed line, white solid line, straight sign, left turn sign, right turn sign, U-turn sign, etc. In order to reduce storage overhead, the embodiment of the present application preferably uses a laser map in a picture format. The laser map can be composed of a grayscale image with a single color channel, or a color image with multiple color channels.

[0092] In one embodiment, Figure 5 As shown, the laser map is implemented as a three-channel color image, for example, including an R channel (red channel), a G channel (green channel) and a B channel (blue channel), and some characteristic information of the laser map can be added to each channel. As an optional implementation method, laser features can be added to the R channel, such as line features, corner features, gradient features, height features, etc. of various elements in the map; the G channel can be added with brightness information of the laser map; the B channel can be added with relative height information, such as the height of the pixel relative to the ground, etc.

[0093] In addition, in addition to the above-mentioned RGB channels, the laser map may also include more channels, such as an alpha channel, etc., which is not limited in the embodiments of the present application. In some other implementations, the laser map may also be implemented in other color modes, such as the RYYB format, etc. In this case, the laser map may include four channels accordingly.

[0094] In one embodiment, each channel in a pixel of a laser map may contain a certain number of bits of information, such as 8 bits (8 bits of information), 16 bits, 32 bits, etc. Therefore, characteristic information such as the above-mentioned laser characteristics, brightness information, and relative height information may be encoded and represented in the bit position of the pixel. For example, brightness information may be represented by 1 bit encoding, such as a bit value of 1 indicating brightness information, and 0 indicating no brightness information.

[0095] In order to merge the visual map and the laser map into one map, the embodiment of the present application merges the visual semantic information usually located in the visual map into the laser map. In a specific implementation, the pixels in the laser map that will contain the visual semantic information can be determined based on the corresponding positions of the road surface signs and road traffic signs in the laser map, and then the visual semantic information can be encoded and stored in the bit information of one or some channels of these pixels.

[0096] In one embodiment, the visual semantic information may be represented in an 8-bit binary encoding format, occupying 8 bits of the channel in which it is located. Figure 6 is an example of an 8-bit encoding format for visual semantic information. Figure 6 As shown, the 8-bit encoding format can be composed of at least one of the three parts: a flag bit, a type code, and a brightness code, wherein the flag bit can be used to indicate the type of the road sign; the type code is used to indicate the content of the road sign; and the brightness information code can be used to indicate the brightness information of the image.

[0097] In one implementation, the flag bit is as follows Figure 6 As shown, it occupies 1 bit, for example, the first bit of an 8-bit code, or other bits. At this time, the flag bit can have two values ​​0 and 1, representing at most two categories. For example, if the visual semantic information is divided into text elements and graphic marking elements, then 0 can represent text elements and 1 can represent graphic marking elements, or, 1 can represent text elements and 0 can represent graphic marking elements. For example, if the visual semantic information is divided into road sign elements and other sign elements, then 0 can represent road sign elements and 1 can represent other sign elements, or, 1 can represent road sign elements and 0 can represent other sign elements.

[0098] In some other implementations, the flag bit may occupy more than 1 bit, such as 2 bits, 3 bits, etc., so that more types of visual semantic information can be represented, for example, 2 bits can represent up to 4 types, and 3 bits can represent up to 8 types. In specific practice, those skilled in the art can determine the length of the flag bit according to the actual classification requirements, and the embodiment of the present application does not specifically limit the length of the flag bit.

[0099] In one implementation, the type encoding is as follows Figure 6 As shown, it occupies 6 bits, such as the 6 consecutive bits after the flag bit, or other bits. In this case, the value range of the type code can be from 000000 to 111111, which can represent up to 2 6 For example, 000000 can represent a road surface, 000001 can represent a road sign, etc.; more specifically, 000001 can represent a white dotted line, 000010 can represent a white solid line, 000011 can represent a straight sign, 000100 can represent a left turn sign, 000101 can represent a warning sign, and 000110 can represent a road sign, etc. In the embodiment of the present application, the type code is a necessary part of the 8-bit code of the visual semantic information.

[0100] In some other implementations, the type code may occupy more than 6 bits, such as 7 bits, so that more content types can be represented; it may also be less than 6 bits, such as 5 bits, 4 bits, etc., so that the bit length of the type code can be reduced to reduce data overhead while being able to represent all required road surface signs and / or road traffic signs, and more bits of information can be reserved in the 8-bit code to represent other information. In specific practice, those skilled in the art can determine the length of the type code according to the number of road surface signs and / or road traffic signs that need to be distinguished, and the embodiment of the present application does not specifically limit the length of the type code.

[0101] In one implementation, the brightness is encoded as Figure 6 As shown, it occupies 1 bit, such as the last bit of the 8-bit code, or other bits. At this time, the flag bit can have two values ​​0 and 1, 0 means no brightness information, and 1 means brightness information. Taking the lane line as an example, its brightness is higher than that of the road surface, so its brightness information can be 1; taking the road surface marking as an example, it may include the painted white line part and the road surface part, so the brightness information of the white line part can be 1, and the brightness information of the road surface part can be 0.

[0102] In some other implementations, the brightness information may occupy more than 1 bit, such as 2 bits, 3 bits, etc., so that the brightness can be represented in a more detailed manner. In specific practice, those skilled in the art can determine the length of the brightness information according to the actual classification requirements, and the embodiment of the present application does not specifically limit the length of the brightness information.

[0103] In one embodiment, the binary encoding of the visual semantic information may include only a part of the flag bit, the type code, and the brightness code, for example, only the type code and the brightness code, or, only the flag bit and the type code, or, only the type code. In addition, the binary encoding of the visual semantic information may also be other encoding formats other than the 8-bit encoding format, such as a code greater than 8 bits, such as a 16-bit code, or a code less than 8 bits, such as a 4-bit code, etc., which is not limited in the embodiments of the present application.

[0104] In one embodiment, the encoded visual semantic information may be stored in the G channel of the pixel, so that when the G channel of the pixel contains 8 bits of information, the 8 bits of information may include a flag bit, a type code, and a brightness code in sequence from front to back.

[0105] Figure 7 It is a schematic diagram of the 8-bit encoding of the visual semantic information corresponding to the pixels in different areas on the positioning map. Figure 6 As shown, according to the encoding rules of the above example, the pixels in area ① correspond to the road surface, its flag bit is 0, the type code is 000000, and the brightness code is 0, so the visual semantic information is 00000000; the pixels in area ② correspond to the straight sign, its flag bit is 1, the type code is 000011, and the brightness code is 1, so the visual semantic information is 10000111; the pixels in area ③ correspond to the white solid line sign, its flag bit is 1, the type code is 000010, and the brightness code is 1, so the visual semantic information is 10000101.

[0106] It can be understood that the embodiment of the present application encodes and stores the visual semantic information originally located in the visual map into the pixel channels of the laser map, thereby realizing the fusion of the laser map and the visual map, and obtaining a positioning map that contains both visual map features and laser map features, thereby reducing the storage overhead of the map.

[0107] The technical solution of the positioning method provided in the embodiment of the present application is described in detail below.

[0108] Figure 8 1 is a hardware framework diagram of a positioning system for implementing a positioning method provided in an embodiment of the present application. Figure 8As shown, the positioning system may include a control unit 140, a GNSS / INS combination module 130, a wheel speed meter 170, an odometer 180, a laser radar 110, a camera 120, and a memory 190, etc. Among them, the GNSS / INS combination module, the wheel speed meter, the odometer, the laser radar, the camera and other modules are used to collect data respectively and send the data to the control unit for processing, and the memory can be used to store positioning maps, store data collected by the above modules, store program instructions for execution by the control unit, and store data generated by the control unit during data processing, etc.

[0109] The following is based on Figure 8 The hardware structure shown takes the vehicle as an example of the target to be positioned to specifically illustrate the step flow of the positioning method provided in the embodiment of the present application. It can be understood that, in addition to vehicles, the positioning target of the method in the embodiment of the present application can also be other means of transportation such as ships and trains, various robots, and construction machinery, etc.

[0110] Fig. 9 is a flowchart of a positioning method provided in an embodiment of the present application, Fig.10 This is the data flow diagram involved in the positioning method. Fig. 9 and Fig.10 As shown, the positioning method can be implemented by the following steps S101 to S105:

[0111] Step S101: determine the initial posture of the vehicle and generate N sampling points C1 to C2 around the initial posture. N , N is a positive integer.

[0112] The initial position and posture may include the initial position and initial posture of the vehicle.

[0113] In a specific implementation, the control unit can obtain data collected by the GNSS / INS combination module. Then, the control unit can determine the initial position of the vehicle based on the antenna signal of the GNSS. Generally speaking, the initial position of the vehicle can be a global position. In addition, the control unit can also determine the initial attitude of the vehicle based on information such as the angular velocity and acceleration of the vehicle measured by the inertial measurement unit IMU of the INS module. Generally speaking, the initial attitude of the vehicle can be composed of one or more parameters of the initial heading angle, pitch angle and roll angle of the vehicle. Since the heading angle is mainly used in vehicle positioning and navigation, the initial attitude of the vehicle can also only include the heading angle.

[0114] In one embodiment, after determining the initial position of the vehicle, the control unit may also generate N sampling points C1 to C2 within a certain range near the vehicle and within a certain range near the heading angle of the vehicle, with the initial position of the vehicle as the center. N .

[0115] Fig.11 The scheme for generating sampling points is provided as an example. Fig.11 As shown, the control unit can determine a circular range with a radius of R = 5 meters with the initial position of the vehicle as the center, and select a sector range with a deviation of 2° to the left and right, i.e. yaw ± 2°, with the direction of the heading angle yaw (the vehicle's forward direction) as the center, and then the overlapping area of ​​the circular range and the sector range (i.e. Fig.11 Generate N=1000 discrete sampling points in the gray shaded area in ). In some implementations, these 1000 sampling points can be generated in a uniform distribution manner, so that the 1000 sampling points are relatively evenly distributed within their distribution area. It is understandable that after the sampling points are selected, the initial posture of the sampling points is also determined based on the initial posture of the vehicle. In other implementations, these 1000 sampling points can also be generated in a non-uniform manner, such as a normal distribution, etc., which is not limited in the embodiments of the present application.

[0116] It should be noted that the embodiment of the present application selects a large number of sampling points around the initial posture of the vehicle, which can realize multiple sampling and multiple calculation of the vehicle posture, and improve the positioning accuracy by combining the sampling point filtering technology. In the specific implementation, the control unit can perform steps S102 to S105 for each sampling point respectively.

[0117] Step S102: extracting a first laser feature and at least one visual semantic information from a positioning map according to a current predicted position of the vehicle.

[0118] It should be noted here that, since the position and posture of the vehicle is constantly changing during driving, various positioning methods are required to be able to locate the vehicle in real time, so as to facilitate the automatic driving system to realize real-time path planning and navigation functions. In order to achieve the purpose of real-time positioning, the control unit can periodically locate the vehicle, and the positioning behavior of each cycle can be called a positioning frame. In the specific description, in order to distinguish the positioning frames at different times, if the current time is set to t, then the positioning frame corresponding to the current time t can be called the current frame, the previous positioning frame of the current frame is called the first historical frame, and the time of the first historical frame is recorded as the first historical moment t-1.

[0119] It is understandable that during the driving process of the vehicle, the control unit will obtain a predicted posture of the vehicle each time it locates the vehicle. For ease of description, the embodiment of the present application refers to the posture of the vehicle obtained at the first historical moment t-1 as the first historical posture. Then, based on the first historical posture and the odometer parameters, the current predicted posture P of the vehicle at the current moment t can be predicted. t It should be noted that if the first historical moment t-1 is the initial moment, then the first historical posture is the initial posture of the vehicle.

[0120] Based on the above definition, the embodiment of the present application can determine the current predicted position of the vehicle in the following manner:

[0121] Step a: Get the initial posture of the vehicle.

[0122] As mentioned above, the initial position and posture can be the position and posture determined according to the data collected by the GNSS / INS combination module when the positioning method is initially executed. Step a is only used to be executed when the method is initialized.

[0123] Step b: determining the relative position of the vehicle between the current time t and the first historical time t-1 based on the odometer data.

[0124] In one implementation, the embodiment of the present application uses a mileage coordinate system to obtain relative posture.

[0125] Fig.12 is a schematic diagram of the odometer coordinate system. Fig.12 As shown, the odometer coordinate system can take the initial posture of the vehicle as the origin Odom, the front direction of the vehicle in the initial posture as the X-axis direction, and the direction perpendicular to the front direction of the vehicle and pointing to the left side of the vehicle as the Y-axis direction.

[0126] The local pose of the vehicle in the odometer coordinate system can be calculated by the odometer based on the measurement data of the wheel speed meter and the inertial measurement unit IMU. When calculating the local pose, the odometer can use the following kinematic model:

[0127] S=V*Δt (1)

[0128]

[0129] Among them, S represents the movement mileage of the vehicle relative to the initial position, V represents the movement speed of the vehicle, and Δt represents the movement time of the vehicle; It represents the change of the vehicle's heading angle relative to the initial position, and ω represents the angular velocity of the vehicle.

[0130] According to the above motion model (1) (2), we can get:

[0131] x0=V*cos(yaw)*Δt (3)

[0132] y0=V*sin(yaw)*Δt (4)

[0133] Among them, x0 is the X-axis coordinate value of the vehicle when it is moving, y0 is the Y-axis coordinate value of the vehicle when it is moving, and yaw represents the heading angle of the vehicle. The value of yaw in the odometer coordinate system is

[0134] According to the above motion models (1)(2) and formulas (3)(4), the local pose expressed by the parameters of the first local coordinate system at any time when the vehicle is moving can be obtained. For example, the local pose may include (x0, y0, yaw).

[0135] It should be noted here that the local pose only represents the pose of the vehicle in the odometer coordinate system, and does not represent the absolute pose of the vehicle in the spatial environment.

[0136] It should be noted that in addition to using the odometer coordinate system, the embodiments of the present application can also use other coordinate systems to obtain relative posture, such as the GNSS coordinate system, the IMU coordinate system, the vehicle's rear axle ground projection coordinate system, etc., and the embodiments of the present application are not limited to this.

[0137] Based on the odometer coordinate system, the odometer can send the local posture of the vehicle at the current time t and the local posture of the first historical time t-1 to the control unit. Then, the control unit can calculate the relative posture of the vehicle between the current time t and the first historical time t-1 based on the local posture of the vehicle at the current time t and the local posture of the first historical time t-1. The specific calculation method is as shown in formula (5):

[0138] ΔP=o t -o t-1 (5)

[0139] Among them, ΔP is the relative position of the vehicle between the current time t and the first historical time t-1, o t is the local position of the vehicle at the current time t, o t-1 is the local position of the vehicle at the first historical moment t-1.

[0140] Step c: Add the first historical posture corresponding to the vehicle at the first historical moment t-1 to the relative posture to obtain the current predicted posture P of the vehicle t .

[0141] As shown in the following formula (6)

[0142] P t =P t-1 +ΔP (6)

[0143] Among them, P t-1 is the first historical position of the vehicle at the first historical moment t-1.

[0144] It is understandable that the sampling point C n The current estimated pose P n It is also estimated through the above step c, namely:

[0145] P n =P n(t-1)+ΔP

[0146] Among them, P n(t-1) is the sampling point C n The estimated pose corresponding to the first historical moment t-1.

[0147] Furthermore, the control unit can calculate the vehicle's current predicted position P according to the vehicle's current predicted position P. t , get the current predicted pose P from the positioning map t A nearby area, for the convenience of description, may be referred to as a local positioning map, and then the first laser feature is extracted from the local positioning map.

[0148] In one embodiment, the positioning map may be composed of a large number of images of preset sizes, each of which corresponds to a range of a specified size in the spatial environment. For example, each image of the positioning map is a square image with equal length and width, and each image corresponds to a square range of 100 meters in length and 100 meters in width.

[0149] In one embodiment, when the positioning map is composed of a large number of images, the control unit can obtain the current predicted pose P from the positioning map. t The image where the image is and at least one nearby image are used as a local positioning map. For example, Fig.13 As shown in 14, the control unit can obtain the current predicted posture P t There are 9 pictures of the 3×3 image of the current position and the surrounding area. If the range of each picture is 100m×100m, then the local positioning map includes the current predicted position P t A nearby area of ​​300m x 300m.

[0150] Based on the above-extracted image of the positioning map, the control unit may extract the first laser feature from a channel of the image storing the laser feature, for example, extract the first laser feature from the R channel.

[0151] In addition, the control unit may extract at least one visual semantic information from a channel storing the visual semantic information of the local positioning map, for example, by decoding the G channel data of the image to extract at least one visual semantic information in the G channel.

[0152] Step S103: for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of .

[0153] The second laser feature can be extracted from the point cloud data collected by the laser radar.

[0154] In a specific implementation, the control unit can sample the point cloud data collected by the laser radar at the current time t to obtain the second laser feature, and the sampling method can be determined according to the specific form of the positioning map. For example: when the positioning map is a sparse feature map, the control unit can perform sparse feature sampling on the point cloud data; when the positioning map is a semi-dense map, the control unit can perform semi-dense feature sampling on the point cloud data; when the positioning map is a dense map, the control unit can perform dense feature sampling on the point cloud data. In this way, it is convenient to match the first laser feature with the second laser feature.

[0155] Next, the control unit calculates the current estimated pose P n Project the laser feature into the coordinate system of the local positioning map. For different sampling points C n For example, due to its current estimated pose P n Different, so the above laser characteristics according to different sampling points C n The current estimated pose P n After projection, different coordinate distributions will correspond in the local positioning map.

[0156] Next, the control unit can match the first laser feature with the second laser feature based on the coordinate distribution of the first laser feature and the second laser feature in the local positioning map, calculate the matching distance between the first laser feature and the second laser feature, and determine the current estimated pose P according to the matching distance. n The first weight Among them, the matching distance represents the distance between the actual position of the vehicle determined based on the laser feature and the sampling point C n The current estimated pose P n The closer the degree of closeness is, the greater the first weight is. The larger the value, the lower the degree of closeness. The smaller it is.

[0157] In some embodiments, the matching distance can be a cosine distance or an Euler distance. The present application embodiment does not limit the algorithm used to obtain the matching distance. For example, when the matching distance is a cosine distance, the numerical range of the matching distance can be [0, 1]. The larger the value, the closer the actual vehicle posture determined based on the laser feature is to the sampling point C. n The current estimated pose P n The lower the degree of closeness, the smaller the value, indicating that the actual position of the vehicle determined based on the laser feature is close to the sampling point C. n The current estimated pose P n The closer the distance is, the closer the distance is to each other.

[0158] Step S104: for any sampling point C n, according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of .

[0159] The at least one road semantic information is extracted from image data collected by a camera.

[0160] In the specific implementation, Fig.14 As shown, the control unit may first pre-process the image data collected by the camera, such as removing noise, cropping, grayscale processing, etc. Next, the control unit may use a pre-trained deep neural network to perform pixel-level semantic segmentation on the pre-processed image to extract at least one road surface semantic information from the image. The road surface semantic information may be pixel-level information, and each road surface semantic information may include: a pixel block containing at least one road surface sign, the number of pixels in the pixel block, the type and probability of the road surface sign to which each pixel belongs, etc. Among them, the pixel block should at least contain all the pixels of the road surface sign. In some embodiments, the pixel block may be a regular shape such as a rectangle, a circle, or other shapes, preferably a regular shape, to facilitate data processing. In addition, while ensuring that the pixel block contains all the pixels of the road surface sign, the pixel block preferably contains as few pixels as possible that are not road surface signs.

[0161] The deep neural network used in the embodiments of the present application may be, for example: a convolutional neural network (CNN), a long short-term memory network (LSTM), a recurrent neural network (RNN) or other neural networks, or a combination of multiple neural networks. The deep neural network uses training corpus as input in the training phase, and the training corpus may be a road surface picture collected in advance, and the road surface signs in the road surface picture are annotated at the pixel level; the output of the deep neural network in the training phase is the annotation result of the training corpus, such as the type of annotated road surface signs, etc. The input of the neural network in the use phase is the image collected by the camera, and the output is the pixel-level information of the road surface signs contained in the image. The specific method of using a deep neural network for information extraction is not the focus of discussion in the embodiments of the present application, and due to space limitations, it will not be repeated here.

[0162] In addition, it should be noted that road semantic information actually also belongs to visual semantic information. The difference is that it is extracted from images collected by the camera instead of being stored in the positioning map.

[0163] Based on the above-extracted road surface semantic information and visual semantic information, step S104 can be implemented through the following steps S201 - S204 as shown in Fig.15 :

[0164] Step S201: Determine at least one valid road surface semantic information from at least one road surface semantic information, and establish a set of valid road surface semantic information. Among them, the number of pixels of each valid road surface semantic information is within a preset range.

[0165] In one embodiment, step S201 can be specifically implemented through the following steps S301 - S303 as shown in Fig.16 :

[0166] Step S301: Calculate the pixel area of each road surface semantic information respectively.

[0167] In specific implementation, the pixel area of the road surface semantic information can be the number of pixels of the pixel block. Taking a rectangular pixel block as an example, assuming the resolution of the pixel block is W pixels × H pixels, where W and H are the number of pixels in the horizontal and vertical directions of the pixel block respectively, then the pixel area S of this pixel block = W × H.

[0168] Exemplarily, as shown in Fig.17 : In step S105, the control unit obtains multiple road surface semantic information from the images collected by the camera, such as including: road surface semantic information L0, road surface semantic information L1, and road surface semantic information L2. Then, in step S301, the number of pixels of the pixel blocks of L0, L1, and L2 can be calculated respectively to obtain the pixel area of L0 as S0 = W0 × H0, the pixel area of L1 as S1 = W1 × H1, and the pixel area of L2 as S2 = W2 × H2.

[0169] Step S302: Determine the valid road surface semantic information according to the pixel area.

[0170] In specific implementation, in the embodiments of the present application, a lower pixel area threshold T1 and an upper pixel area threshold T2 for determining the valid road surface semantic information can be set. The control unit uses the lower threshold T1 and the upper threshold T2 to compare with the pixel area S of each road surface semantic information (for example: S0, S1, S2, etc.) respectively. When T1 < S < T2, the road surface semantic information is the valid semantic information. When S < T1 or S > T2, the road surface semantic information is the invalid semantic information. Among them, the lower threshold T1 and the upper threshold T2 can be pre-set values or dynamically generated values.

[0171] In one embodiment, when the lower threshold T1 and the upper threshold T2 are dynamically generated values, the control unit can count the pixel areas of all road surface semantic information within a period of time to obtain the distribution range of the pixel area, and then select a certain range in the distribution range of the pixel area as the range of valid road surface semantic information, and then determine the lower threshold T1 and the upper threshold T2.

[0172] Step S303: establishing an effective road surface semantic information set for the effective road surface semantic information.

[0173] According to the different results of the effective semantic road surface information determined in step S302, the number of effective road surface semantic information included in the effective road surface semantic information set is also different. For example, when step S302 determines that the road surface semantic information does not include the effective road surface semantic information set, the effective road surface semantic information set is an empty set; when step S302 determines that a part of the road surface semantic information is the effective road surface semantic information, the effective road surface semantic information set is a subset of the road surface semantic information set; when step S302 determines that all the road surface semantic information is the effective road surface semantic information, the effective road surface semantic information set is the same as the road surface semantic information set.

[0174] For example, any valid road semantic information in the valid road semantic information set may be in the following form:

[0175] a i =[M, p1~p M ,Ka i ]

[0176] Among them, a i represents the i-th valid road semantic information in the valid road semantic information set, and M represents a i The number of pixels contained, Ka i Indicates a i The corresponding semantic type value (such as the type value of the road sign), different semantic types have different type values, p1~p M Respectively represent a i The 1st to Mth pixels belong to Ka i The probability of p1~p M It can be obtained from the output of the deep neural network.

[0177] The above steps S301 to S303 are exemplary implementation methods of step S201.

[0178] Step S202: unify the coordinate systems of the effective road semantic information and the visual semantic information.

[0179] In a specific implementation, the control unit can calculate the estimated pose P according to the current nProject the above-mentioned at least one valid road semantic information into the coordinate system of the local positioning map. n For example, due to its current estimated pose P n Therefore, at least one valid road semantic information is obtained according to different sampling points C n The current estimated pose P n After projection, different coordinate distributions will correspond in the local positioning map.

[0180] In one embodiment, the local positioning map can use a known coordinate system such as the GNSS coordinate system, or it can have its own coordinate system. For example, the control unit can establish the coordinate system of the local positioning map with the center point of the local positioning map as the origin and the horizontal and vertical directions as the X-axis and Y-axis.

[0181] Furthermore, after determining the coordinate system of the local positioning map, the control unit can project the effective road semantic information from the camera coordinate system to the coordinate system of the local positioning map by means of matrix transformation. The transformation matrix can be, for example, a 4×4 matrix, and its mathematical meaning represents a spatial transformation process of one translation and one rotation, that is, any pixel point in the effective road semantic information can be projected into the coordinate system of the local positioning map through one translation and one rotation. The projection transformation of spatial points between different coordinate systems is a common method in the field of navigation and positioning, which will not be elaborated here.

[0182] Step S203: determining a semantic association relationship between at least one valid road surface semantic information and at least one visual semantic information.

[0183] The purpose of step S203 is to: for any valid road semantic information a in the valid road semantic information set i (a i ∈A), find a corresponding to a from the visual semantic information set B of the local positioning map i Visual semantic information of semantic association j (b j ∈B).

[0184] In one embodiment, step S203 is as follows: Fig.18 The following steps S401 to S404 can be used to implement the above:

[0185] Step S401, calculate effective road semantic information a i The semantic weight of .

[0186] In the specific implementation, the following formula can be used:

[0187]

[0188] in, Indicates a i The semantic weight of a i The number of pixels contained, p m Indicates a i The mth pixel in belongs to Ka i The probability of Ka i Indicates a i The corresponding semantic type value.

[0189] Step S402, calculating visual semantic information b j The semantic weight of .

[0190] In the specific implementation, the following formula can be used:

[0191]

[0192] in, Indicates b j The semantic weight of b j The number of pixels contained, p g Indicates b j The g-th pixel belongs to The probability of Indicates b j The corresponding semantic type value.

[0193] It should be noted that in the local map, whether a pixel belongs to a road sign or a road traffic sign is known, so for b j For the g-th pixel, its p g There are only two possible values: 0 and 1. If the pixel belongs to a road sign or a road traffic sign, then its p g =1, otherwise p g =0.

[0194] Step S403: Based on the effective road semantic information a i The semantic weight and visual semantic information b j The difference of the semantic weights of the effective road semantic information a i and visual semantic information b j semantic relevance.

[0195] In the specific implementation, for any valid road semantic information a i and any visual semantic information b j , the semantic association Δw is the effective road semantic information a i and any visual semantic information b j The absolute value of the difference is the following formula:

[0196]

[0197] Step S404: when the semantic relevance is less than a preset first threshold, determining the effective road semantic information a i and visual semantic information b j Have semantic associations.

[0198] In the specific implementation, for any valid road semantic information a i and any visual semantic information b j If its semantic association Δw is less than the preset first threshold σ, then the effective road semantic information a i and any visual semantic information b j If the semantic association degree Δw is greater than or equal to the preset first threshold σ, then the effective road semantic information a i and any visual semantic information b j No semantic association.

[0199] The embodiment of the present application can eventually obtain at least one pair of valid road surface semantic information and visual semantic information with semantic association by repeatedly executing the operations of step S401 to step S404. For the sake of ease of description, the valid road surface semantic information and visual semantic information with semantic association can be referred to as a semantic association group.

[0200] The above steps S401 to S404 are exemplary implementation methods of step S203.

[0201] Step S204: semantically match each pair of semantically associated valid road surface semantic information and visual semantic information, and determine a second weight according to the semantic matching result.

[0202] In the specific implementation, step S204 is as follows: Fig.19 The following steps S501 to S503 can be used to implement the above:

[0203] Step S501 : calculating the matching distance of each pair of semantically associated valid road semantic information and visual semantic information respectively.

[0204] Among them, for each pair of valid road semantic information and visual semantic information with semantic association, the matching distance can be the Euclidean distance (i.e., Euclidean distance) of the valid road semantic information and the visual semantic information, or the cosine distance, etc., which is not specifically limited in the embodiments of the present application.

[0205] Combine the following Fig. 20 The mathematical meaning of Euclidean distance and cosine distance is explained. Fig. 20 The effective road semantic information a is shown in the coordinate system of the local positioning map iand visual semantic information b j The Euclidean distance is the effective road semantic information a i and visual semantic information b j The straight-line distance in the coordinate system of the local positioning map. The cosine distance is the effective road semantic information a i and visual semantic information b j Angle α with the line connecting the origin of the coordinate system ij The cosine value of .

[0206] Step S502: The weighted sum of each matching distance is obtained to obtain the total matching distance.

[0207] To obtain the total matching distance, step S502 is as follows Fig.21 The above can be achieved through the following steps S601-S602.

[0208] Step S601: Calculate the weight of each semantic association group.

[0209] Specifically, the following formula can be used:

[0210]

[0211] Where K represents the number of semantically related groups, w′ k represents the weight of the kth semantic association group, Δw k represents the semantic association degree of the kth semantic association group.

[0212] Step S602: determining a total matching distance according to the weights and matching distances of all semantic association groups.

[0213] Among them, the weight of each semantic association group and the matching distance can be multiplied respectively, and then all the multiplication results are summed to obtain the total matching distance E, that is, the following formula is used:

[0214]

[0215] Where K represents the number of semantically related groups, w′ k represents the weight of the kth semantic association group, E k represents the matching distance of the kth semantic association group.

[0216] Step S503: determine the current estimated pose P according to the total matching distance n The second weight

[0217] The total matching distance represents the distance between the actual position of the vehicle determined based on visual features and the sampling point C n The current estimated pose P n The closer the degree of closeness is, the greater the second weight is. The higher the value, the lower the degree of closeness. The lower it is.

[0218] In one embodiment, when the matching distance is the Euclidean distance, the smaller the value of the total matching distance is, the closer the actual position of the vehicle determined based on the visual features is to the sampling point C. n The current estimated pose P n The closer the degree of proximity is, the higher the corresponding second weight is. The larger the total matching distance is, the closer the actual vehicle posture determined based on the visual features is to the sampling point C. n The current estimated pose P n The lower the proximity between them, the corresponding second weight The smaller.

[0219] In one embodiment, when the matching distance is the cosine distance, the smaller the value of the total matching distance is, the closer the sampling point C determined based on the visual feature is. n The actual pose and the current predicted pose P n The closer the degree of proximity is, the higher the corresponding second weight is. The larger the total matching distance is, the closer the sampling point C is to the one determined based on the visual features. n The actual pose and the current predicted pose P n The lower the proximity between them, the corresponding second weight The smaller.

[0220] Based on the above total matching distance and the second weight The control unit can use any algorithm to determine the second weight. This embodiment of the present application does not limit this. The value range of can be, for example, [0, 1], or other ranges, which are not limited in the present embodiment. It can be obtained by complementing the total matching distance, taking the inverse, normalizing the numerical range, etc.

[0221] The above steps S501 to S503 are exemplary implementation methods of step S204.

[0222] Step S105: according to the N sampling points C1-C N The current estimated pose P1-P N and the first weight and the second weight, calculate the current estimated pose P 1- P N The weighted average of the weighted average is used as the current posture of the vehicle.

[0223] In the specific implementation, step S105 is as follows: Fig. 22 The following steps S701 to S703 can be used to implement the above:

[0224] Step S701, using the current estimated pose P1-P N The first weight of the current estimated pose P1-P N Weighted averaging is performed to obtain the first weighted average.

[0225] In one embodiment, the first weighted average value may be obtained by the following formula:

[0226]

[0227] Among them, P l represents the first weighted average, N is the number of sampling points; P n Represents the current estimated pose of the nth sampling point; is the first weight corresponding to the current estimated pose of the nth sampling point, For sampling points C1~C N The current estimated pose P 1- P N The corresponding first weight sum.

[0228] Step S702, using the current estimated pose P1-P N The second weight of the current estimated pose P1-P N Weighted averaging is performed to obtain a second weighted average.

[0229] In one embodiment, the second weighted average value may be obtained by the following formula:

[0230]

[0231] Among them, P c represents the second weighted average, N is the number of sampling points; P n Represents the current estimated pose of the nth sampling point; is the second weight corresponding to the current estimated pose of the nth sampling point, For sampling points C1~C N The current estimated pose P 1- P N The corresponding second weight sum.

[0232] Among them, step S701 and step S702 are only used to describe the method steps, and do not represent the order of the steps. Generally speaking, the control unit can execute step S701 and step S702 in parallel, or can execute S701 and step S702 in sequence. In addition, in the embodiment of the present application, the first weighted average and the second weighted average are both postures.

[0233] Step S703: weighted average the first weighted average value and the second weighted average value to obtain a weighted average value, and use the weighted average value as the current posture of the vehicle.

[0234] In the specific implementation, the positioning result of the vehicle can be obtained by the following formula:

[0235] P=α l ·P l +α c ·P c (14)

[0236] Among them, P is the current position of the vehicle, that is, the actual position of the vehicle output by the positioning system this time; P l represents the first weighted average, P c represents the second weighted average; α l represents the weight of the first weighted average, α c represents the weight of the second weighted average, α l +α c =1. Among them, α l and α c The value of can be determined according to actual needs. l or α c The larger the value of , the higher the weight of the first weighted average or the second weighted average. In practice, if the technician wants to use the laser feature to dominate the positioning result, he can increase α l The value of, for example, α l The value can be 0.7, 0.8, etc. If the technician wants to use visual features to dominate the positioning results, α can be increased c The value of, for example, α c The value is 0.7, 0.8, etc. If the technician hopes that the laser feature and the visual feature will play an equal role in the positioning result, then α l and α c The value of can be 0.5.

[0237] It is understandable that after the execution of step S101 to step S105, the positioning system has completed a complete positioning process. During the driving process of the vehicle, since the position and posture of the vehicle changes all the time, the positioning process is also continuously performed.

[0238] It is understandable that, affected by the odometer error or other factors, as the number of positioning times increases, the distribution of sampling points may diverge. Among them, the divergence conditions of the sampling points can be set by those skilled in the art, for example: when the second weight of the current estimated posture corresponding to a preset proportion of sampling points is lower than a preset threshold, the sampling points are considered to be divergent; or, when the first weight of the current estimated posture corresponding to a preset proportion of sampling points is lower than a preset threshold, the sampling points are considered to be divergent. When the sampling points diverge, the control unit can reselect the same number of sampling points as before, and the embodiment of the present application does not limit the method of reselecting the sampling points, for example: select a certain proportion (for example: 90%) of sampling points around one or more current estimated postures with higher first weights and / or second weights, and select the remaining proportion (for example: 10%) of sampling points around the posture output by the GNSS / INS combination module. In addition, the control unit can also periodically reselect the sampling points, for example, every 100 positioning frames as a reselection cycle, and reselect the sampling points.

[0239] The positioning method provided in the embodiment of the present application can determine the first weight of the current estimated posture based on the matching of laser features, determine the second weight of the current estimated posture based on the matching of visual semantic information and road semantic information (i.e., the matching of visual features), and then weight the current estimated posture according to the first weight and the second weight to obtain the positioning result of the vehicle, thereby realizing the fusion of the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching, and improving the positioning efficiency. In addition, the positioning method provided in the embodiment of the present application encodes and stores the laser features and visual semantic information in the same positioning map, realizing the fusion of map data and reducing the data overhead and computing overhead generated during the positioning process.

[0240] The embodiments provided in the present application above introduce various schemes of the positioning method. It is understandable that, in order to realize the above functions, the positioning system may include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0241] In one embodiment, the positioning system can be Figure 8 The hardware structure shown in the figure realizes the corresponding functions. The positioning system can be installed in a vehicle. The installation method of some components can be as follows: Figure 4As shown. Among them, the GNSS / INS combination module 130 is used to determine the initial posture of the vehicle; the memory 190 is used to store the positioning map, the positioning map includes multiple positioning pictures spliced ​​together, the positioning picture includes a color channel, and the color channel stores the encoding of the first laser feature and the encoding of the visual semantic information; the laser radar 110 is used to collect point cloud data, and the point cloud data includes the second laser feature; the camera 120 is used to collect image data, and the image data includes at least one road semantic information; the control unit 140 is used to generate N sampling points C1~C around the initial posture N , N is a positive integer; the control unit 140 is further used to extract the first laser feature and at least one visual semantic information from the positioning map according to the current predicted position of the vehicle; the control unit 140 is also used to extract the first laser feature and at least one visual semantic information from the positioning map for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of the second laser feature is extracted from the point cloud data collected by the laser radar; the control unit 140 is also used for any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of at least one road semantic information is extracted from the image data collected by the camera; the control unit 140 is also used to extract the road semantic information according to the N sampling points C1~C N The current estimated pose P1-P N and its first and second weights to calculate the current estimated pose P 1- P N The weighted average of is taken as the current posture of the vehicle.

[0242] In one embodiment, the positioning picture includes a first color channel, and the first color channel is used to store an encoding of visual semantic information.

[0243] In one embodiment, the positioning image further includes a second color channel, and the second color channel is used to store the code of the first laser feature.

[0244] In one embodiment, the encoding of visual semantic information includes at least one of a flag bit, a type code, and a brightness code; the flag bit is used to indicate the type of the road sign, the type code is used to indicate the content of the road sign, and the brightness information code is used to indicate the brightness information of the image.

[0245] In one embodiment, when the control unit 140 is used to extract at least one visual semantic information from the positioning map according to the current predicted position of the vehicle: the control unit 140 is specifically used to extract a local positioning map from the positioning map according to the current predicted position, the local positioning map includes M positioning pictures, the M positioning pictures include a first picture where the current predicted position is located, and M-1 second pictures near the first picture, M is a positive integer greater than 1; the control unit 140 is also used to extract at least one visual semantic information from the local positioning map.

[0246] In one embodiment, when the control unit 140 is used to n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The control unit 140 is specifically configured to determine at least one valid road surface semantic information from at least one road surface semantic information, and the number of pixels of each valid road surface semantic information is within a preset range; the control unit 140 is also configured to determine at least one valid road surface semantic information from at least one road surface semantic information according to the current estimated posture P n At least one valid road surface semantic information is projected into the coordinate system of the local positioning map; the control unit 140 is also used to determine the semantic association relationship between at least one valid road surface semantic information and at least one visual semantic information; the control unit 140 is also used to perform semantic matching on each pair of semantically associated valid road surface semantic information and visual semantic information, and determine the second weight according to the semantic matching result.

[0247] In one embodiment, when the control unit 140 is used to determine the semantic association relationship between at least one valid road semantic information and at least one visual semantic information: the control unit 140 is specifically used to calculate any valid road semantic information a i The semantic weights and any visual semantic information b j The control unit 140 is also used to calculate the semantic weight of the effective road surface semantic information a i The semantic weight and visual semantic information b j The difference of the semantic weights of the effective road semantic information a i and visual semantic information b j The control unit 140 is also used to determine the effective road semantic information a when the semantic association is less than the preset first threshold value i and visual semantic information b j Have semantic associations.

[0248] In one embodiment, when the control unit 140 is used to perform semantic matching on each pair of semantically associated effective road surface semantic information and visual semantic information, and determine the second weight based on the semantic matching result: the control unit 140 is specifically used to respectively calculate the matching distance of each pair of semantically associated effective road surface semantic information and visual semantic information; the control unit 140 is also used to obtain the total matching distance by weighted summation of the calculated matching distances; the control unit 140 is also used to determine the second weight based on the total matching distance.

[0249] In one embodiment, when the control unit 140 is used to N The current estimated pose P1-P N and its first weight and second weight, calculate the current estimated pose P 1- P N When the weighted average value is used as the current posture of the vehicle: the control unit 140 is specifically configured to use the current estimated posture P1-P N The first weight of the current estimated pose P1-P N The weighted average is obtained to obtain a first weighted average; the control unit 140 is also used to use the current estimated posture P1-P N The second weight of the current estimated pose P1-P N The weighted average is obtained to obtain a second weighted average value; the control unit 140 is further used to weighted average the first weighted average value and the second weighted average value to obtain a weighted average value, and the weighted average value is used as the current posture of the vehicle.

[0250] In one embodiment, the road semantic information includes: a pixel block including at least one road sign, the number of pixels in the pixel block, and the type of road sign to which each pixel belongs.

[0251] In one embodiment, the positioning system further includes an odometer 180; the control unit 140 is further configured to determine the relative position of the vehicle between the current time t and the first historical time t-1 according to the data of the odometer 180; the control unit 140 is further configured to calculate ... n The predicted posture corresponding to the first historical moment t-1 is added to the relative posture to obtain the current predicted posture.

[0252] In one embodiment, the control unit 140 is further configured to regenerate N sampling points C1 to C2 when a preset ratio of the first weight or the second weight is lower than a second threshold. N .

[0253] In another embodiment, the positioning system can be Fig.23 The software modules shown in the figure implement the corresponding functions. Fig.23As shown, the positioning system may include a sampling point generation module 810, an extraction module 820, a first matching module 830, a second matching module 840, and a solution module 850. The functions of the above modules are described in detail below:

[0254] The sampling point generation module 810 is used to determine the initial posture of the vehicle and generate N sampling points C1 to C2 around the initial posture. N , N is a positive integer;

[0255] The extraction module 820 is used to extract the first laser feature and at least one visual semantic information from the positioning map according to the current predicted position of the vehicle.

[0256] The first matching module 830 is used to match any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight is, and the second laser feature is extracted from the point cloud data collected by the lidar.

[0257] The second matching module 840 is used to match any sampling point C n , according to the current estimated pose P n , matching at least one visual semantic information with at least one road semantic information to determine the current estimated pose P n The second weight of at least one road semantic information is extracted from the image data collected by the camera.

[0258] The solution module 850 is used to solve the problem according to the N sampling points C1-C N The current estimated pose P1-P N and its first and second weights to calculate the current estimated pose P 1- P N The weighted average of is taken as the current posture of the vehicle.

[0259] The positioning map includes a plurality of positioning pictures spliced ​​together, the positioning picture includes a color channel, and the encoding of the first laser feature and the encoding of the visual semantic information are stored in the color channel.

[0260] The positioning system provided in the embodiment of the present application can determine the first weight of the current estimated posture based on the matching of laser features, determine the second weight of the current estimated posture based on the matching of visual semantic information and road semantic information (i.e., the matching of visual features), and then calculate the weighted average of the current estimated posture based on the first weight and the second weight, and use the weighted average as the current posture of the vehicle, thereby realizing the fusion of the laser positioning technology based on laser feature matching and the visual positioning technology based on visual feature matching, and improving the positioning efficiency. In addition, the technical solution provided in the embodiment of the present application encodes and stores the laser features and visual semantic information in the same positioning map, realizing the fusion of map data and reducing the data overhead and computing overhead generated during the positioning process.

[0261] An embodiment of the present application also provides a vehicle, which may include the positioning system provided by the aforementioned embodiments, and the user executes the positioning method provided by the aforementioned embodiments.

[0262] The embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned methods.

[0263] The embodiment of the present application also provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the above-mentioned methods.

[0264] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above are only specific implementation methods of the embodiments of the present application and are not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the protection scope of the embodiments of the present application.

Claims

1. A positioning method, characterized in that: include: Determine the initial position of the vehicle and generate N sampling points C1 to C2 around the initial position. N , N is a positive integer; Extract the first laser feature and at least one visual semantic information from the positioning map according to the current predicted position of the vehicle; for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of the laser feature is extracted from the point cloud data collected by the laser radar; as well as, For any sampling point C n , according to the current estimated pose P n , matching the at least one visual semantic information with the at least one road semantic information to determine the current estimated pose P n A second weight of , wherein the at least one road semantic information is extracted from image data collected by a camera; According to the N sampling points C1-C N The current estimated pose P1-P N and the first weight and the second weight, calculate the current estimated pose P 1- P N The weighted average of , taking the weighted average as the current posture of the vehicle; The positioning map includes a plurality of positioning pictures spliced ​​together, the positioning picture includes a color channel, and the encoding of the first laser feature and the encoding of the visual semantic information are stored in the color channel.

2. The method according to claim 1, characterized in that The positioning picture includes a first color channel, and the first color channel is used to store the encoding of the visual semantic information.

3. The method according to claim 2, characterized in that The positioning picture also includes a second color channel, and the second color channel is used to store the code of the first laser feature.

4. The method according to any one of claims 1 to 3, characterized in that: The encoding of the visual semantic information includes at least one of a flag bit, a type code and a brightness code; the flag bit is used to indicate the type of the road sign, the type code is used to indicate the content of the road sign, and the brightness code is used to indicate the brightness information of the picture.

5. The method according to claim 1, characterized in that The at least one visual semantic information is extracted by the following steps: Extracting a local positioning map from the positioning map according to the current predicted posture, the local positioning map comprising M positioning pictures, the M positioning pictures comprising a first picture where the current predicted posture is located, and M-1 second pictures near the first picture, where M is a positive integer greater than 1; The at least one visual semantic information is extracted from the local positioning map.

6. The method according to claim 5, characterized in that For any sampling point C n , according to the current estimated pose P n , matching the at least one visual semantic information with the at least one road semantic information to determine the current estimated pose P n The second weight of includes: Determine at least one valid road surface semantic information from the at least one road surface semantic information, wherein the number of pixels of each valid road surface semantic information is within a preset range; According to the current estimated pose P n Projecting the at least one valid road semantic information into the coordinate system of the local positioning map; Determining a semantic association relationship between the at least one valid road surface semantic information and the at least one visual semantic information; Semantic matching is performed on each pair of semantically associated valid road surface semantic information and visual semantic information, and the second weight is determined according to the semantic matching result.

7. The method according to claim 6, characterized in that The determining of the semantic association relationship between the at least one valid road surface semantic information and the at least one visual semantic information includes: Calculate any valid road semantic information a i The semantic weights and any visual semantic information b j The semantic weight of According to the effective road semantic information a i The semantic weight and the visual semantic information b j The difference of the semantic weights of the effective road semantic information a is determined i and the visual semantic information b j The semantic relevance of When the semantic association is less than a preset first threshold, the valid road semantic information a is determined. i and the visual semantic information b j Have semantic associations.

8. The method according to claim 6 or 7, characterized in that: The performing semantic matching on each pair of the semantically associated valid road surface semantic information and the visual semantic information, and determining the second weight according to the semantic matching result, comprises: respectively calculating the matching distance between each pair of semantically associated effective road surface semantic information and visual semantic information; The weighted sum of the calculated matching distances is used to obtain a total matching distance; The second weight is determined according to the total matching distance.

9. The method according to claim 1, characterized in that: According to the N sampling points C1-C N The current estimated pose P1-P N and the first weight and the second weight, calculate the current estimated pose P 1- P N The weighted average of , taking the weighted average as the current posture of the vehicle, includes: Using the current estimated pose P1-P N The first weight of the current estimated pose P1-P N Weighted averaging is performed to obtain the first weighted average value; Using the current estimated pose P1-P N The second weight of the current estimated pose P1-P N Weighted averaging is performed to obtain a second weighted average value; The first weighted average value and the second weighted average value are weighted and averaged to obtain the weighted average value, and the weighted average value is used as the current posture of the vehicle.

10. The method according to claim 1, characterized in that The road semantic information includes: a pixel block including at least one road sign, the number of pixels in the pixel block, and the type of road sign to which each pixel belongs.

11. The method according to claim 1, characterized in that: The current predicted pose is determined by the following steps: Determine the relative position of the vehicle between the current time t and the first historical time t-1 according to the odometer data; The sampling point C n The predicted posture corresponding to the first historical moment t-1 is added to the relative posture to obtain the current predicted posture.

12. The method according to claim 1, characterized in that Also includes: When a preset ratio of the first weight or the second weight is lower than a second threshold, the N sampling points C1 to C N .

13. A positioning system, characterized in that: include: The combined GNSS / INS module, control unit, memory, lidar and camera installed in the vehicle; The GNSS / INS combination module is used to determine the initial position and posture of the vehicle; The memory is used to store a positioning map, the positioning map includes a plurality of positioning pictures spliced ​​together, the positioning picture includes a color channel, and the color channel stores a code of the first laser feature and a code of visual semantic information; The laser radar is used to collect point cloud data, wherein the point cloud data includes a second laser feature; The camera is used to collect image data, wherein the image data includes at least one piece of road semantic information; The control unit is used to generate N sampling points C1 to C2 around the initial posture. N , N is a positive integer; The control unit is further used to extract a first laser feature and at least one visual semantic information from a positioning map according to a current predicted position of the vehicle; The control unit is also used for any sampling point C n , n is a positive integer, n≤N, according to its corresponding current estimated pose P n , the first laser feature extracted from the positioning map is matched with the second laser feature to determine the current estimated pose P n The first weight of the laser feature is extracted from the point cloud data collected by the laser radar; The control unit is also used for any sampling point C n , according to the current estimated pose P n , matching the at least one visual semantic information with the at least one road semantic information to determine the current estimated pose P n A second weight of , wherein the at least one road semantic information is extracted from image data collected by a camera; The control unit is further configured to: N The current estimated pose P1-P N and the first weight and the second weight, calculate the current estimated pose P 1- P N The weighted average of is taken as the current posture of the vehicle.

14. The positioning system according to claim 13, characterized in that: The positioning picture includes a first color channel, and the first color channel is used to store the encoding of the visual semantic information.

15. The positioning system according to claim 13, characterized in that: The positioning picture also includes a second color channel, and the second color channel is used to store the code of the first laser feature.

16. The positioning system according to any one of claims 13 to 15, characterized in that: The encoding of the visual semantic information includes at least one of a flag bit, a type code and a brightness code; the flag bit is used to indicate the type of the road sign, the type code is used to indicate the content of the road sign, and the brightness code is used to indicate the brightness information of the picture.

17. The positioning system according to claim 13, characterized in that: When the control unit is used to extract the at least one visual semantic information from the positioning map according to the current predicted position of the vehicle: The control unit is specifically configured to extract a local positioning map from the positioning map according to the current predicted posture, wherein the local positioning map includes M positioning pictures, and the M positioning pictures include a first picture where the current predicted posture is located, and M-1 second pictures near the first picture, where M is a positive integer greater than 1; The control unit is further used to extract the at least one visual semantic information from the local positioning map.

18. The positioning system according to claim 17, characterized in that: When the control unit is used for any sampling point C n , according to the current estimated pose P n , matching the at least one visual semantic information with the at least one road semantic information to determine the current estimated pose P n The second weight of is: The control unit is specifically used to determine at least one valid road surface semantic information from the at least one road surface semantic information, and the number of pixels of each valid road surface semantic information is within a preset range; The control unit is further configured to: n Projecting the at least one valid road semantic information into the coordinate system of the local positioning map; The control unit is further used to determine a semantic association relationship between the at least one valid road surface semantic information and the at least one visual semantic information; The control unit is further used to perform semantic matching on each pair of semantically associated valid road surface semantic information and visual semantic information, and determine the second weight according to the semantic matching result.

19. The positioning system according to claim 18, characterized in that When the control unit is used to determine the semantic association relationship between the at least one valid road semantic information and the at least one visual semantic information: The control unit is specifically used to calculate any valid road semantic information a i The semantic weights and any visual semantic information b j The semantic weight of The control unit is further configured to: i The semantic weight and the visual semantic information b j The difference of the semantic weights of the effective road semantic information a is determined i and the visual semantic information b j The semantic relevance of The control unit is further configured to determine the effective road semantic information a when the semantic association is less than a preset first threshold. i and the visual semantic information b j Have semantic associations.

20. The positioning system according to claim 18 or 19, characterized in that: When the control unit is used to perform semantic matching on each pair of semantically associated valid road semantic information and visual semantic information, and determine the second weight according to the semantic matching result: The control unit is specifically used to calculate the matching distance of each pair of semantically associated effective road semantic information and visual semantic information respectively; The control unit is further used to obtain a total matching distance by performing a weighted summation of the calculated matching distances; The control unit is further configured to determine the second weight according to the total matching distance.

21. The positioning system according to claim 13, characterized in that When the control unit is used to N The current estimated pose P1-P N and the first weight and the second weight, calculate the current estimated pose P 1- P N The weighted average of , when the weighted average is used as the current posture of the vehicle: The control unit is specifically configured to use the current estimated posture P1-P N The first weight of the current estimated pose P1-P N Weighted averaging is performed to obtain the first weighted average value; The control unit is further configured to use the current estimated posture P1-P N The second weight of the current estimated pose P1-P N Weighted averaging is performed to obtain a second weighted average value; The control unit is further used to perform weighted averaging on the first weighted average value and the second weighted average value to obtain the weighted average value, and use the weighted average value as the current posture of the vehicle.

22. The positioning system according to claim 13, characterized in that The road semantic information includes: a pixel block including at least one road sign, the number of pixels in the pixel block, and the type of road sign to which each pixel belongs.

23. The positioning system according to claim 13, characterized in that Also includes: Odometer; The control unit is further used to determine the relative position of the vehicle between the current time t and the first historical time t-1 according to the odometer data; The control unit is also used to set the sampling point C n The predicted posture corresponding to the first historical moment t-1 is added to the relative posture to obtain the current predicted posture.

24. The positioning system according to claim 13, characterized in that The control unit is further configured to regenerate the N sampling points C1 to C2 when a preset proportion of the first weight or the second weight is lower than a second threshold. N .

25. A vehicle, characterized in that: Comprising a positioning system as described in any one of claims 13-24.

26. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 12.

27. A computer program product, characterized in that When the computer program product is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Automatic guided vehicle (AGV) positioning method based on multi-sensor data fusion

    CN110108269A

  • Pose estimation method and device, electronic equipment and computer readable storage medium

    CN111174782A