Vision ai-based fusion positioning method, system, device and storage medium

By using a visual AI-based fusion positioning method and combining real-time images with high-precision maps for verification, the problems of high computational load, high cost, and low recognition accuracy in existing technologies have been solved, achieving fast and accurate vehicle positioning.

CN115164924BActive Publication Date: 2025-11-25SHANGHAI WESTWELL INFORMATION & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210789254.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-11-25
Estimated Expiration
2042-07-06

AI Technical Summary

Technical Problem

Existing vehicle positioning and road surface detection solutions suffer from problems such as high computational load, high cost, low recognition accuracy, and slow positioning speed, especially in complex environments lacking lane markings where positioning accuracy is insufficient.

Method used

A vision-based AI-based fusion positioning method is adopted. Real-time image detection and image segmentation results are used as input, and high-precision maps are combined for hybrid verification. End-to-end detection results and image segmentation results are used for self-verification, filtering out false lane lines, interpolating and expanding the point set, and improving positioning accuracy through skeletonization and Kdtree matching.

Benefits of technology

It reduces the amount of computation, speeds up positioning, improves positioning accuracy, avoids longitudinal positioning deviations of vehicles in environments lacking markers, and reduces the demand for hardware and computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115164924B_ABST
    Figure CN115164924B_ABST
Patent Text Reader

Abstract

The application provides a vision AI-based fusion positioning method, system, device and storage medium, which comprises the following steps: a vehicle collects real-time images of a road surface and GPS positioning information; image recognition is performed on the real-time images to determine whether lane lines exist in the real-time images; if yes, map positioning is performed on a high-precision map based on GPS positioning and real-time images, and a point set in a local area range is interpolated and expanded; if no, the point set of the high-precision map is interpolated and expanded based on GPS positioning; the interpolated point set is classified and matched with a partitioned image to obtain a corresponding area of the point set of the high-precision map in the image. The application can use the detection result and the image segmentation result of the real-time image as input to realize hybrid verification based on the high-precision map, reduce the calculation amount, speed up the positioning speed, and improve the positioning precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of AI vision, and particularly relates to a fusion positioning method and system based on visual AI, equipment and a storage medium. BACKGROUND

[0002] In the field of automatic driving, in recent years, solutions through high-precision maps and visual positioning have been paid more and more attention. At present, the mainstream solution on the market is still in the form of single lane line feature, single model support and multi-sensor loose coupling. This results in that the current high-precision map visual positioning solution still lacks self-checking, so that a single visual positioning solution cannot complete a relatively complex task alone.

[0003] Most of the existing vehicle positioning and road surface detection use the fusion of multiple devices such as visual sensors, laser radars and ultrasonic radars, and fuse the multi-party data into the three-dimensional space on the spot. The steps of projection to the overhead view are complicated, although the accuracy is high, but the cost is extremely high, and it is difficult to promote. Part of the road surface detection is only performed by using a graphic sensor, but the recognition accuracy is low, and the large amount of calculation also prolongs the processing time of the algorithm, and improves the requirements for vehicle-mounted computing power or bandwidth.

[0004] Therefore, the present application provides a fusion positioning method and system based on visual AI, equipment and a storage medium.

[0005] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present application, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] In view of the problems in the prior art, the present application aims to provide a fusion positioning method and system based on visual AI, equipment and a storage medium, which overcomes the difficulties of the prior art, can use the detection result and image segmentation result of the real-time image as input to realize hybrid checking based on high-precision maps, reduce the amount of calculation, speed up the positioning speed, and improve the positioning accuracy.

[0007] The embodiment of the present application provides a fusion positioning method based on visual AI, comprising the following steps:

[0008] The vehicle collects real-time images of the road surface and GPS positioning information;

[0009] The real-time images are subjected to image recognition, and a mapping relationship between the pixel region after partitioning and the label category corresponding to the partitioned image is obtained;

[0010] It is judged whether lane lines exist in the real-time images;

[0011] If the lane line exists, the point set in the local area range in the high-definition map is interpolated and expanded based on the current GPS positioning information of the vehicle and the map positioning based on the position of the lane line in the real-time image;

[0012] If the lane line does not exist, the point set in the local area range in the high-definition map is interpolated and expanded based on the current GPS positioning information of the vehicle;

[0013] Based on the classified matching of the interpolated point set and the partitioned image based on the label category, the corresponding area of the point set of the high-definition map in the image is obtained.

[0014] Preferably, the vehicle collects the real-time image of the road surface and the GPS positioning information, comprising:

[0015] The vehicle collects the real-time image of the road surface through the image sensor;

[0016] The vehicle obtains the current GPS positioning information of the vehicle through the GPS.

[0017] Preferably, the image recognition of the real-time image is performed, and the mapping relationship between the pixel area after partitioning and the label category corresponding to the partitioned image is obtained, comprising:

[0018] The real-time image is subjected to image partition recognition to obtain the label category corresponding to each partition;

[0019] The mapping relationship between the pixel corresponding to each partition and the label category of the partition is established.

[0020] Preferably, the judgment of whether the lane line exists in the real-time image includes:

[0021] Judging whether each partition in the real-time image includes at least one image partition with the label category of the lane line.

[0022] Preferably, the map positioning based on the current GPS positioning information of the vehicle and the position of the lane line in the real-time image based on the high-definition map, and the interpolation and expansion of the point set in the local area range based on the map positioning, comprise:

[0023] Based on the image, the curve fitting of the extension direction of each lane line in the image is performed, and the curve trajectory is obtained;

[0024] Based on the current GPS positioning information of the vehicle as the center and the preset length as the radius, a local area in the high-definition map is obtained;

[0025] filtering, by the curve trajectory, a point set of points within a range of the local area, and filtering out points with a distance greater than a preset distance threshold from the curve trajectory; and

[0026] interpolating and expanding the filtered point set.

[0027] Preferably, the filtering, by the curve trajectory, a point set of points within a range of the local area, and filtering out points with a distance greater than a preset distance threshold from the curve trajectory, further comprises:

[0028] The preset distance threshold is 1 / 15 to 1 / 30 of a width pixel value of the real-time image.

[0029] Preferably, the interpolating and expanding the point set within the local area of the high-definition map based on the current GPS positioning information of the vehicle comprises:

[0030] obtaining a local area in the high-definition map based on the current GPS positioning information of the vehicle as a center and a preset length as a radius;

[0031] obtaining a point set of points within a range of the local area;

[0032] interpolating and expanding the point set.

[0033] Preferably, the high-definition map comprises a plurality of sub-point sets, each point in each sub-point set has a label category, the label category corresponds to an obstacle or a road surface mark, when interpolating and expanding the point set, an interpolation point is added between adjacent points based on spatial positions of a plurality of adjacent points, and a label category of the interpolation point is determined according to a label category with a maximum number of adjacent points around the interpolation point.

[0034] Preferably, the interpolating and expanding is performed on the filtered point set and a sub-point set in the high-definition map with a distance from the point set less than or equal to a preset distance threshold.

[0035] Preferably, the classification and matching based on the label category between the interpolated point set and the partitioned image to obtain a corresponding area of the point set of the high-definition map in the image comprises:

[0036] skeletonizing each of the partitioned images;

[0037] performing classification and matching based on the label category between the interpolated point set and the skeletonized partitioned image to obtain a corresponding area of the point set of the high-definition map in the image.

[0038] The embodiment of the present application also provides a fusion positioning system based on visual AI, which is used for realizing the fusion positioning method based on visual AI.

[0039] An image acquisition module is configured to acquire real-time images of a road surface and GPS positioning information of the vehicle;

[0040] An image recognition module is configured to perform image recognition on the real-time images and obtain a mapping relationship between a pixel region after partition and a label category corresponding to a partitioned image;

[0041] A lane detection module is configured to determine whether lane lines exist in the real-time images;

[0042] A filtering and interpolation module is configured to, when the lane lines exist, perform map positioning based on the current GPS positioning information of the vehicle and positions of the lane lines in the real-time images based on a high-definition map, and perform interpolation expansion on a point set in a local region range based on the map positioning;

[0043] A positioning and interpolation module is configured to, when the lane lines do not exist, perform interpolation expansion on the point set in the local region range in the high-definition map based on the current GPS positioning information of the vehicle;

[0044] A matching and positioning module is configured to, based on the point set after interpolation and the partitioned image, perform classification matching based on a label category, and obtain a corresponding region of the point set in the high-definition map in the image.

[0045] The embodiment of the present application also provides a fusion positioning device based on visual AI, which comprises:

[0046] A processor;

[0047] A memory in which executable instructions of the processor are stored;

[0048] The processor is configured to execute the steps of the fusion positioning method based on visual AI by executing the executable instructions.

[0049] The embodiment of the present application also provides a computer readable storage medium for storing a program, the program being executed to realize the steps of the fusion positioning method based on visual AI.

[0050] The fusion positioning method, system, device and storage medium based on visual AI can realize hybrid verification based on a high-definition map by taking a detection result of real-time images and an image segmentation result as input, reduce a calculation amount, accelerate a positioning speed, and improve positioning accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0051] Other features, objects, and advantages of the application will become more apparent from the following detailed description when read in connection with the following drawings.

[0052] Figure 1 is a flowchart of the fusion positioning method based on visual AI of the present application.

[0053] Figures 2 to 6 is a schematic diagram of one implementation scenario of the fusion positioning method based on visual AI of the present application.

[0054] Figures 7 to 8 is a schematic diagram of another implementation scenario of the fusion positioning method based on visual AI of the present application.

[0055] Figure 9 is a structural schematic diagram of the fusion positioning system based on visual AI of the present application.

[0056] Figure 10 is a structural schematic diagram of the fusion positioning device based on visual AI of the present application. And

[0057] Figure 11 is a structural schematic diagram of the computer readable storage medium of an embodiment of the present application. DETAILED DESCRIPTION

[0058] The present application is described below by way of specific examples, and other advantages and effects of the present application can be easily understood by those skilled in the art from the disclosure of the present application. The present application can also be implemented or applied in different specific embodiments or systems, and various modifications or changes can be made to the details of the present application without departing from the spirit of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0059] The embodiments of the present application are described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present application. The present application can be embodied in various different forms, and is not limited to the embodiments described herein.

[0060] In the description of the present application, the expressions of "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" mean that the specific features, structures, materials or characteristics expressed in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics expressed can be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples expressed in the present application and the features of the different embodiments or examples without conflict.

[0061] Furthermore, the terms "first", "second", etc. are used herein only to distinguish one element from another and do not necessarily indicate any relative importance or imply the presence of two or more elements. Thus, a feature defined with "first", "second", etc. can include at least one of the features, explicitly or implicitly.

[0062] For the sake of clearness of the present application, parts irrelevant to the description are omitted, and the same or similar components are designated by the same reference numerals throughout the several drawings.

[0063] Throughout the specification, when it is said that an element is "connected" to another element, this includes not only a case where it is "directly connected", but also a case where it is "indirectly connected" with other elements interposed therebetween. In addition, when it is said that an element "includes" a certain component, unless specifically noted to the contrary, other components are not excluded, but it means that other components can also be included.

[0064] When it is said that an element is "on" another element, it can be directly on the other element, but can also be on the other element with other elements interposed therebetween. When it is said in contrast that an element is "directly on" another element, there are no other elements interposed therebetween.

[0065] Although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first interface and a second interface, etc. are distinguished from each other. Also, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, steps, operations, elements, components, items, kinds and / or groups but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, items, kinds and / or groups thereof. As used herein, the terms "or" and "and / or" are construed to be inclusive, or mean any one or any combination of items in the alternative. Thus, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C". Exceptions to this definition apply only when the combination of elements, functions, steps or acts are mutually exclusive between themselves as a matter of some fact.

[0066] The specific terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, in the specification and the claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The term "comprising" as used herein is intended to mean that the compositions and methods include the recited elements, but not excluding others.

[0067] Although not all defined differently, the technical terms and scientific terms used herein include the meanings commonly understood by one of ordinary skill in the art to which this application belongs. Terms defined in commonly used dictionaries add to the interpretation that the context of the relevant art and the contemporaneous descriptions provided herein bear, unless otherwise defined specifically herein. In implementing the present application, one of ordinary skill in the art will not feel bound by any specific listed

[0068] Figure 1 is a flowchart of a fusion positioning method based on visual AI of the present application. As shown in Figure 1 The embodiment of the present application provides a fusion positioning method based on visual AI, comprising the following steps:

[0069] S110, the vehicle collects real-time images of the road surface and GPS positioning information.

[0070] S120, image recognition is performed on the real-time images, and a mapping relationship between the pixel regions after partitioning and the label categories corresponding to the partitioned images is obtained.

[0071] S130, it is judged whether there is a lane line in the real-time image, if yes, step S140 is executed, if no, step S150 is executed.

[0072] S140, map positioning is performed based on the high-definition map by using the current GPS positioning information of the vehicle and the position of the lane line in the real-time image, and the point set in the local area range is interpolated and expanded based on the map positioning, and step S160 is executed.

[0073] S150, the point set in the local area range in the high-definition map is interpolated and expanded based on the current GPS positioning information of the vehicle.

[0074] S160, classification matching is performed based on the label categories by using the interpolated point set and the partitioned image, and the corresponding area of the point set in the high-definition map in the image is obtained.

[0075] To solve the problems in the prior art, the application provides a multi-input visual auxiliary positioning scheme, which takes end-to-end detection results and image segmentation results as inputs and completes self-checking through a series of operations. After obtaining the matching results, the matching results, the odometer and the GPS can be further involved in fusion, and further smoothing between visual positioning frames can be completed. Thus, the problem of longitudinal positioning deviation of the vehicle in the straight-ahead case without signs can be effectively avoided.

[0076] In a preferred embodiment, step S110 comprises:

[0077] S110, the vehicle collects real-time images of the road surface through an image sensor.

[0078] S120, the vehicle obtains the current GPS positioning information of the vehicle through the GPS.

[0079] In a preferred embodiment, step S120 comprises:

[0080] S121, image partition recognition is performed on the real-time images to obtain the label category corresponding to each partition.

[0081] S122, a mapping relationship between the pixels corresponding to each partition and the label category of the partition is established.

[0082] In a preferred embodiment, step S130 comprises:

[0083] It is judged whether at least one image partition with the label category of lane line is included in each partition of the real-time images.

[0084] In a preferred embodiment, step S140 comprises:

[0085] S141, a planar coordinate is established based on the images, each lane line is curve-fitted in the extension direction in the images, and a curve trajectory is obtained.

[0086] S142, based on the current GPS positioning information of the vehicle as the center and a preset length as the radius, a local area is obtained in the high-precision map.

[0087] S143, the set of points with the label category of lane line in the range of the local area is filtered through the curve trajectory, and points with a distance greater than a preset distance threshold from the curve trajectory are filtered out.

[0088] and

[0089] S144, the filtered point set is interpolated and expanded.

[0090] In a preferred embodiment, in step S143, the preset distance threshold is 1 / 15 to 1 / 30 of the width pixel value of the real-time images.

[0091] In a preferred embodiment, the high-definition map comprises a plurality of sub-point sets, each point in each sub-point set has a label category and accurate positioning information, the label category corresponds to an obstacle or a road marking, when the point set is interpolated and expanded, an interpolation point is added between adjacent points based on the spatial positions of the adjacent points, and the label category of the interpolation point is determined according to the label category with the maximum number of adjacent points around the interpolation point, thereby realizing automatic definition of the interpolation point.

[0092] In a preferred embodiment, in step S144, the filtered point set and the sub-point set in the high-definition map with a distance less than or equal to the preset distance threshold are interpolated and expanded.

[0093] In a preferred embodiment, the filtered point set is interpolated and expanded.

[0094] In a preferred embodiment, step S150 comprises:

[0095] S151, based on the current GPS positioning information of the vehicle as the center and a preset length as the radius, a local area is obtained in the high-definition map.

[0096] S152, a point set of points located within the range of the local area is obtained.

[0097] S153, the point set is interpolated and expanded.

[0098] In a preferred embodiment, step S160 comprises:

[0099] S161, each partition image is skeletonized.

[0100] S162, the interpolated point set and the skeletonized partition image are classified and matched based on the label category, and the corresponding area of the point set of the high-definition map in the image is obtained.

[0101] The fusion positioning method based on visual AI can realize hybrid verification based on the high-definition map by taking the detection result and the image segmentation result of the real-time image as input, reduce the calculation amount, speed up the positioning speed, and improve the positioning accuracy.

[0102] In order to solve the problems of the prior art, the present application proposes a multi-input visual auxiliary positioning scheme, which takes the end-to-end detection result and the image segmentation result as input and completes self-checking through a series of operations. After obtaining the matching result, the matching result, the odometer and the GPS can be further involved in fusion, and further smoothing between visual positioning frames can be completed. Thus, the problem of longitudinal positioning deviation of the vehicle in the straight line without marking can be effectively avoided.

[0103] FIG. 1 is a schematic diagram of an embodiment of the fusion positioning method based on visual AI of the present application. As shown in the figure, in this embodiment, the vehicle collects real-time images 1 of the road surface through an image sensor. The vehicle obtains the current GPS positioning information of the vehicle through GPS. Figures 2 to 6 Figure 2

[0104] As shown in the figure, the real-time images 1 are subjected to image partition recognition through the vehicle-mounted visual chip to obtain the label category corresponding to each partition. A mapping relationship between the pixels corresponding to each partition and the label category of the partition is established. In this embodiment, the label category of partition 11 is lane line, the label category of partition 12 is lane line, and the label category of partition 13 is street lamp (the lane line of the other lane at the upper left corner of the picture should be located at the edge and is severely deformed, so it is not recognized, which will not affect the positioning accuracy of the present application, but rather removes irrelevant data, which is conducive to reducing the calculation amount and speeding up the recognition speed). It is judged whether at least one image partition with the label category of lane line is included in each partition of the real-time images 1. The AI image recognition in this embodiment refers to the technology of using a computer to process, analyze and understand images to identify various different patterns of targets and objects, which is a practical application of deep learning algorithm. The traditional recognition process of images is divided into four steps: image acquisition → image preprocessing → feature extraction → image recognition. An image recognition neural network is trained based on a large amount of road condition images, so as to segment the identified target objects in the real-time images, and establish a mapping relationship between the segmented images and the target object labels. Figure 3

[0105] As shown in the figure, based on the image, a plane coordinate is established, each lane line is subjected to curve fitting in the extension direction in the image, and curve trajectories 14 and 15 are obtained. Based on the current GPS positioning information of the vehicle as the center and a pre-set length as the radius, a local area is obtained in the pre-stored high-precision map in the vehicle-mounted system. The high-precision map in this embodiment includes a large number of point sets with high-precision positioning information, wherein each point in each sub-point set has a label category, and the label category corresponds to an obstacle or a road surface mark. An end-to-end method can directly extract the equation information of the lane line of the image, if the point set of the high-precision map projected into the image does not deviate too much from the end-to-end equation information, it will be considered that the position passes the verification, and it will be considered that the point is in the effective area. A large number of map point sets not in the image can be quickly screened through such a method. Figure 4

[0106] ​​​​In a preferred example, the specific pattern of the point set of the high-definition map projected to the real-time image 1 can be simulated by the on-board chip and the prior information of the image sensor (sensor size, position mounted on the vehicle body, optical parameters of the lens, etc.), so that the actual captured lane line is compared with the simulated lane line, and the area where the two overlap or are close is the area close to the vehicle, which needs high-precision identification; the area where the actual captured lane line and the simulated lane line are far away from each other is far away from the vehicle, and high-precision identification is not needed.

[0107] As shown in Figure 5 , the point set in the local area whose label category is lane line is filtered by the curve trajectory, and the points with a distance greater than a preset distance threshold from the curve trajectory are filtered out. For example, the minimum distance of each point in the point set from the curve trajectory is calculated, and since each curve trajectory corresponds to a lane line, one lane line is left in the point set by filtering (the other points in the point set corresponding to the same lane line may be filtered out because they are too far away or too close to the curve trajectory). The preset distance threshold is 1 / 15 to 1 / 30 of the width pixel value of the real-time image 1, and in this embodiment, the image width is 960 pixels, so the preset distance threshold is 50 pixels. By filtering, the points 16 close to the curve trajectory are retained, and the points 17 with a distance greater than the preset distance threshold from the curve trajectory are filtered out (since the distance is too far, the area represented by the point 17 does not pose a threat to the current driving, so removing these points will not affect the current road identification, and it is also beneficial to reduce the calculation amount and speed up the identification speed). Since the image sensor can capture the road and obstacles hundreds of meters away, this part does not need to be used, and the objects far away that are not identified are not included in the subsequent point cloud calculation, thereby reducing the calculation amount.

[0108] As shown in Figure 6 , the high-definition map includes a plurality of sub-point sets, each point in each sub-point set has a label category, and the label category corresponds to an obstacle or a road identifier. When the point set is interpolated and expanded, an interpolation point is added between adjacent points based on the spatial positions of the adjacent points, and the label category of the interpolation point is determined according to the label category with the maximum number of adjacent points around the interpolation point. The filtered point set 18 and the sub-point set 19 in the high-definition map with a distance less than or equal to a preset distance threshold from the point set 18 are interpolated and expanded. By interpolating the point set 18 representing the filtered lane line close to the curve trajectory and the other sub-point set 19 adjacent in the high-definition map, more effective information of the local high-definition map of the relevant area of the road is obtained through interpolation.

[0109] Finally, each partition image is skeletonized. Existing image skeletonization algorithms can be used to classify and match the interpolated point set with the skeletonized partition image based on the label category, to obtain the corresponding area of the point set 18, 19, etc. of the high-precision map in the image, thereby strengthening the vehicle's accurate information about the road surface conditions, obstacle positions, etc. of the surrounding environment. The vehicle can quickly and accurately obtain the position of itself and the surrounding signs or obstacles in the high-precision map, thereby achieving accurate road surface recognition. The skeletonization in the present embodiment includes the following steps: eroding the image, the eroded object becomes narrower; performing an open operation on the eroded image, the pixels deleted during the open operation are part of the skeleton, which are added to the skeleton image; repeating the above process until the image is completely eroded, thereby obtaining the skeleton of each partition image after image thinning, but this is not limiting.

[0110] The present application does not require the use of sensors with depth information or laser radars, etc., saving hardware devices, and at the algorithm level, only through the implementation of image 1 planar coordinates for recognition and positioning, without constructing a three-dimensional space requiring huge computing power, greatly saving computing power. The fusion positioning method based on visual AI of the present application can realize hybrid verification based on high-precision maps by taking the detection results and image segmentation results of real-time images 1 as inputs, reducing the amount of calculation, speeding up the positioning speed, and improving the positioning accuracy.

[0111] Figures 7 to 8 is a schematic diagram of one implementation process of the fusion positioning method based on visual AI of the present application. As shown in Figure 7 , in the present embodiment, the visual matching verification system of the present project is as shown in Figure 7 , the method has three feature inputs. And there are three self-checking strategies inside, so as to ensure the stability of the visual feature matching result.

[0112] Firstly, through the end-to-end model input, the situation where there is no lane line can be effectively considered, so as to avoid the interference problem caused by the misidentification of the segmentation result.

[0113] The first self-checking strategy is to identify whether there is a lane line in the image. If there is a lane line, it is to match the end-to-end result and the information of the high-precision map. The operation at this place can effectively deal with complex scenes such as dense lane lines. By matching the projection result of the high-precision map and the end-to-end output result with the current position of the vehicle, the appropriate lane line point set on the high-precision map is checked from the vehicle, so as to determine the position of the vehicle in the horizontal direction.

[0114] If there is no lane line, the lane line calibration step will be omitted, and the linear difference will be used to expand the high-precision map point set for the second self-checking strategy. This step is mainly used to expand the lane line and lamp post map point information, thereby providing more effective information for lane line and lamp post fitting.

[0115] At the segmentation result, the skeletonization needs to be performed according to the type of the segmentation result. Through the form of skeletonization, the center of the lane line, the lamp post and the sign can be effectively extracted, thereby facilitating the subsequent feature matching.

[0116] Subsequently, the high-precision map point set and the point set after skeletonization are matched using Kdtree one by one, to obtain the corresponding position of the high-precision map point set in the skeletonized graph. Although the interference of the segmentation result will cause errors during skeletonization, the third self-checking is completed through the skeletonization and Kdtree, which can effectively filter out the areas with poor segmentation results. Meanwhile, because the linear interpolation obtains more feature points, the interference of the poor segmentation area can be effectively reduced. Kd-tree (short for k-dimensional tree) is a tree data structure that stores instance points in k-dimensional space for fast retrieval. It is mainly used for searching of key data in multi-dimensional space (such as range search and nearest neighbor search). K-D tree is a special case of binary space partitioning tree. In computer science, k-d tree (short for k-dimensional tree) is a data structure for organizing points in k-dimensional Euclidean space. K-d tree can be used in various applications, such as multi-dimensional key value search (e.g., range search and nearest neighbor search). K-d tree is a special case of binary space partitioning. K-d tree is a binary tree with k-dimensional points. All non-leaf nodes can be regarded as a hyperplane that divides the space into two half spaces. The left subtree of the node represents the points on the left side of the hyperplane, and the right subtree of the node represents the points on the right side of the hyperplane. The method of selecting the hyperplane is as follows: each node is related to the k-dimensional dimension perpendicular to the hyperplane. Therefore, if the x-axis is selected for division, all nodes with x values less than the specified value will be in the left subtree, and all nodes with x values greater than the specified value will be in the right subtree. In this way, the hyperplane can be determined by the x value, and the normal vector of the x-axis is the unit vector. After obtaining the skeletonized points matched with the high-precision map, the three-order curve equation of the lane line, the straight line equation of the lamp post and the point coordinate information of the lane sign are obtained through the point set correspondence. For example, the kd-tree nearest neighbor search is performed by taking all the points after skeletonization as a tree. Thus, the nearest skeletonized point corresponding to each expanded point is obtained. The matching value is calculated based on the Euclidean distance between the pixels in the image, and the search range of the kd-tree is limited. At present, the pixel value is set to 20 pixels. That is, the search range of the kd-tree is limited to 20 pixels around the expanded point (xex , y ex ) and the Euclidean distance between the skeletonized point (x sk , y sk ) is less than 20 pixels, it is considered that the bone point is matched. Finally, a cubic curve equation is fitted again according to these bone points to optimize, but not limited to.

[0117] As shown in the embodiment, when the visual verification matching result is taken as the output, the optimized input of the current frame is obtained, and in this case, the feature information of the last 10 frames is selected. Figure 8 At this time, because the frequency of the odometer is high and the error is small in a short time, it is involved to realize inter-frame optimization.

[0118] The GPS is used in a loose coupling manner to participate in the fusion positioning. Because the update frequency of the GPS is low, and there is drift when the high-rise is blocked, whether the GPS is added to the fusion positioning is judged according to the current GPS state and the difference between the two positioning. If the GPS is added, the positioning results of the two are further fused through Kalman filtering, and the result is output as the final positioning result to the unmanned vehicle. If the GPS is not added, the result of the inter-frame optimization is directly input.

[0119] In the present application, a visual auxiliary high-precision fusion positioning scheme is proposed, and the main technical advantages include:

[0120] Based on the end-to-end detection result and the image segmentation result as input, a hybrid verification based on a high-precision map is realized.

[0121] Through the end-to-end detection result, the situation where there is no lane line can be effectively filtered, so as to avoid the positioning error caused by lane line misdetection

[0122] The image segmentation result is used in the present application, and the skeletonization+kdtree method is used to effectively extract the lane line cubic curve equation suitable for three-dimensional space. And through the lamp pole, the signboard and the lane marking, the accuracy of vehicle positioning is comprehensively constrained.

[0123] In the optimization process, the odometer information is added to the matching result to complete the inter-frame constraint of vision, and the fusion is realized through the judgment of the GPS positioning information. Thus, the problem of longitudinal positioning deviation of the vehicle in the straight line without marking is avoided.

[0124] Figure 9 is a structure diagram of the fusion positioning system based on visual AI of the present application. Figure 9As shown, the visual AI-based fusion positioning system 5 of the present application comprises:

[0125] An image acquisition module 51 acquires real-time images of the road surface and GPS positioning information of the vehicle;

[0126] An image recognition module 52 performs image recognition on the real-time images and obtains a mapping relationship between the partitioned pixel regions and the label categories corresponding to the partitioned images;

[0127] A lane detection module 53 determines whether lane lines exist in the real-time images;

[0128] A filtering and interpolation module 54, when lane lines exist, performs map positioning based on the current GPS positioning information of the vehicle and the positions of the lane lines in the real-time images based on the high-definition map, and performs interpolation expansion on the point set within the local region range based on the map positioning;

[0129] A positioning and interpolation module 55, when lane lines do not exist, performs interpolation expansion on the point set within the local region range in the high-definition map based on the current GPS positioning information of the vehicle;

[0130] A matching positioning module 56, based on the interpolated point set and the partitioned images, classifies and matches based on the label categories to obtain the corresponding region of the point set of the high-definition map in the images.

[0131] In a preferred embodiment, the image acquisition module 51 is configured to cause the vehicle to acquire real-time images of the road surface through an image sensor, and the vehicle to obtain the current GPS positioning information of the vehicle through GPS.

[0132] In a preferred embodiment, the image recognition module 52 is configured to perform image partition recognition on the real-time images to obtain the label category corresponding to each partition and establish a mapping relationship between the pixels corresponding to each partition and the label category of the partition.

[0133] In a preferred embodiment, the lane detection module 53 is configured to determine whether each partition of the real-time images includes at least one image partition with a label category of lane lines.

[0134] In a preferred embodiment, the filtering and interpolation module 54 is configured to establish a planar coordinate based on the image, perform curve fitting on the extension direction of each lane line in the image to obtain a curve trajectory, obtain a local region in the high-definition map with the current GPS positioning information of the vehicle as the center and a preset length as the radius, filter the point set of the points within the range of the local region through the curve trajectory, filter out points with a distance greater than a preset distance threshold from the curve trajectory, and perform interpolation expansion on the filtered point set.

[0135] In a preferred embodiment, the preset distance threshold in the filtering interpolation module 54 is 1 / 15 to 1 / 30 of the width pixel value of the real-time image.

[0136] In a preferred embodiment, the positioning interpolation module 55 is configured to obtain a local area in the high-definition map based on the current GPS positioning information of the vehicle as the center and with a preset length as the radius, obtain a point set of points located within the range of the local area, and interpolate and expand the point set.

[0137] In a preferred embodiment, the high-definition map includes a plurality of sub-point sets, each point in each sub-point set has a label category, the label category corresponds to an obstacle or a road surface mark, and when the point set is interpolated and expanded, an interpolation point is added between adjacent points based on the spatial positions of the adjacent points, and the label category of the interpolation point is determined according to the label category with the maximum number of adjacent points around the interpolation point.

[0138] In a preferred embodiment, the matching positioning module 56 is configured to skeletonize each partition image, and classify and match the interpolated point set and the skeletonized partition image based on the label category to obtain the corresponding area of the point set of the high-definition map in the image.

[0139] The fusion positioning system based on visual AI can realize hybrid verification based on the high-definition map by taking the detection result and the image segmentation result of the real-time image as input, thereby reducing the amount of calculation, accelerating the positioning speed, and improving the positioning accuracy.

[0140] The embodiment of the present application also provides a fusion positioning device based on visual AI, which comprises a processor and a memory having executable instructions of the processor stored therein. The processor is configured to execute the steps of the fusion positioning method based on visual AI by executing the executable instructions.

[0141] As described above, the fusion positioning device based on visual AI can realize hybrid verification based on the high-definition map by taking the detection result and the image segmentation result of the real-time image as input, thereby reducing the amount of calculation, accelerating the positioning speed, and improving the positioning accuracy.

[0142] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a program product. Therefore, each aspect of the present application can be specifically implemented as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "platform" here.

[0143] Figure 10 is a structural schematic diagram of the fusion positioning device based on visual AI of the present application. Hereinafter, the fusion positioning device based on visual AI of the present application will be described in detail with reference to the accompanying drawings. Figure 10The electronic device 600 according to this embodiment of the present application will be described. Figure 10 The electronic device 600 shown is merely one example and should not be taken as limiting the scope of the present application embodiments.

[0144] As shown in Figure 10 The electronic device 600 is in the form of a general computing device. Components of the electronic device 600 can include, but are not limited to, at least one processing unit 610, at least one memory unit 620, a bus 630 that connects the various platform components including the memory unit 620 and the processing unit 610, a display unit 640, etc.

[0145] The memory unit stores program code that can be executed by the processing unit 610 such that the processing unit 610 performs the steps described in the above electronic prescription flow processing method section of this specification according to various exemplary embodiments of the present application. For example, the processing unit 610 can perform the steps as shown in Figure 1

[0146] The memory unit 620 can include a readable medium in the form of volatile memory units such as a random access memory (RAM) 6201 and / or a cache memory unit 6202, and can further include a read-only memory (ROM) 6203.

[0147] The memory unit 620 can further include a program / utility 6204 having a set of program modules 6205 such as an operating system, one or more application programs, other program modules, and program data, and can include an implementation of a network environment, each or a combination of these examples.

[0148] The bus 630 can be representative of one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus structures, and the like.

[0149] ​The electronic device 600 can also communicate with one or more external devices 700 such as a keyboard or a pointing device, a Bluetooth device, etc.; other devices that enable a user to interact with the electronic device 600; and / or any devices (e.g., a router, a modem, a peer device or other computing device) that enable the electronic device 600 to communicate with one or more other computing devices. Such communication can occur via an input / output (I / O) interface 650. Still yet, the electronic device 600 can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network such as the Internet, via a network adapter 660. The network adapter 660 can be any of a plurality of different types of adapters suitable for interfacing the electronic device 600 to various networking schemes. The network adapter 660 can include an integrated services digital network (ISDN) adapter, a digital subscriber line (DSL) adapter, a telephone modem, or a wireless network adapter, for example. It should be appreciated that for purposes of clarity, not all of the hardware of a computer is shown in the figure, including, but not limited to, a main frame computer, a multiprocessor, a networked computer, a microcomputer, a minicomputer, or a computer system that includes a plurality of computers.

[0150] The embodiment of the present application also provides a computer readable storage medium for storing a program, the program being executed to implement the steps of the fusion positioning method based on visual AI. In some possible implementation manners, various aspects of the present application can also be implemented in the form of a program product, which includes program codes for causing a terminal device to perform the steps of the various exemplary embodiments of the present application described in the above electronic prescription flow processing method part of the specification when the program product is run on the terminal device.

[0151] As shown above, the program of the computer readable storage medium of the embodiment, when executed, can implement the hybrid verification based on the high-definition map by taking the detection result of the real-time image and the image segmentation result as inputs, thereby reducing the amount of calculation, accelerating the positioning speed, and improving the positioning accuracy.

[0152] Figure 11 is a structural schematic diagram of the computer readable storage medium of the present application. Referring to Figure 11 As shown in the figure, the program product 800 for implementing the above method according to the embodiment of the present application can be in the form of a portable compact disc read-only memory (CD-ROM) and includes program codes, and can be run on a terminal device such as a personal computer. However, the program product of the present application is not limited to this, and in this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device or apparatus.

[0153] The program product can employ any combination of one or more computer-readable media. The computer-readable media can be a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0154] The computer-readable storage medium can include a data signal traveling in a baseband or a carrier wave traveling in a propagation medium, in which the computer-readable program code embodied in the computer-readable storage medium is contained. The computer-readable storage medium can also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate or transport the program for use by or in connection with an instruction execution system, apparatus or device. The computer-readable program code embodied on the computer-readable storage medium can be transmitted or transported by a computer-readable medium including, but not limited to, a wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0155] The program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.

[0156] In summary, the fusion positioning method, system, device and storage medium based on visual AI of the present application can use the detection result and image segmentation result of the real-time image as input to realize hybrid verification based on high-precision map, reduce the amount of calculation, speed up the positioning speed, and improve the positioning accuracy.

[0157] The above is further detailed description of the present application in combination with specific preferred embodiments, and cannot be deemed as limitation of the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, and all of them shall be deemed as falling within the protection scope of the present application.

Claims

1. A visual AI-based fusion positioning method, characterized in that, The method comprises the following steps: The vehicle collects real-time images of the road surface and GPS positioning information; Image recognition is performed on the real-time images, and a mapping relationship between the pixel regions after partitioning and the label categories corresponding to the partitioned images is obtained; It is judged whether lane lines exist in the real-time images; If lane lines exist, the current GPS positioning information of the vehicle is taken as the center, a local area in the high-precision map is obtained with a preset length as the radius, the curve trajectory of each lane line in the real-time images is obtained by curve fitting in the direction of extension, the set of points in the local area whose label categories are lane lines is filtered through the curve trajectory, and points with a distance greater than a preset distance threshold from the curve trajectory are filtered out, the preset distance threshold being 1 / 15 to 1 / 30 of the width pixel value of the real-time images, and the filtered point set is interpolated and expanded; If lane lines do not exist, the current GPS positioning information of the vehicle is taken as the center, a local area in the high-precision map is obtained with a preset length as the radius, a point set of points located in the local area is obtained, and the point set is interpolated and expanded, the high-precision map comprising a plurality of sub-point sets, each point in each sub-point set having a label category, the label category corresponding to an obstacle or a road surface mark, and when the point set is interpolated and expanded, an interpolation point is added between adjacent points based on the spatial positions of the adjacent points, and the label category of the interpolation point is determined according to the label category with the maximum number of adjacent points around the interpolation point; Each partitioned image is skeletonized, the interpolated point set is classified and matched with the skeletonized partitioned image based on the label categories of the Kdtree, the matching value being obtained according to the Euclidean distance between pixels in the image, and when the Euclidean distance between the extension point of the kd-tree and the skeletonized point is less than the Euclidean distance between pixels in the image, it is considered that a skeleton point is matched, so as to obtain the corresponding area of the point set of the high-precision map in the partitioned image. 2.The vision AI-based fusion positioning method according to claim 1, wherein, The vehicle collects real-time images of the road surface and GPS positioning information, comprising: The vehicle collects real-time images of the road surface through an image sensor; The vehicle obtains the current GPS positioning information of the vehicle through GPS. 3.The vision AI-based fusion positioning method of claim 1, wherein, The image recognition of the real-time images and the obtaining of the mapping relationship between the pixel regions after partitioning and the label categories corresponding to the partitioned images comprise: Image partition recognition is performed on the real-time images to obtain the label categories corresponding to each partition; A mapping relationship between the pixels corresponding to each partition and the label categories of the partition is established. 4.The vision AI-based fusion positioning method of claim 1, wherein, The judgment of whether lane lines exist in the real-time images comprises: It is judged whether at least one image partition with a label category of lane lines is included in each partition of the real-time images.

5. A vision AI based fusion positioning system, characterized in that, The system comprises: An image collection module, which collects real-time images of the road surface and GPS positioning information; An image recognition module, which performs image recognition on the real-time images and obtains a mapping relationship between the pixel regions after partitioning and the label categories corresponding to the partitioned images; lane detection module, judging whether lane lines exist in the real-time image; filtering and interpolating module, when the lane lines exist, establishing a plane coordinate based on the real-time image, performing curve fitting on each lane line in the image, and obtaining a curve track; based on the current GPS positioning information of the vehicle as the center and a preset length as the radius, a local area is obtained in the high-definition map; filtering a point set of points in the local area whose label category is lane line through the curve track, filtering out points whose distance from the curve track is greater than a preset distance threshold, and the preset distance threshold is 1 / 15 to 1 / 30 of the width pixel value of the real-time image; interpolating and expanding the filtered point set; positioning and interpolating module, when the lane lines do not exist, based on the current GPS positioning information of the vehicle as the center and a preset length as the radius, a local area is obtained in the high-definition map, a point set of points in the local area is obtained, and the point set is interpolated and expanded, the high-definition map includes a plurality of sub-point sets, each point in each sub-point set has a label category, the label category corresponds to an obstacle or a road surface mark, when interpolating and expanding the point set, an interpolation point is added between adjacent points based on the spatial positions of the adjacent points, and the label category of the interpolation point is determined according to the label category of the adjacent points around the interpolation point with the largest number; matching and positioning module, skeletonizing each of the partition images, classifying and matching the interpolated point set and the skeletonized partition images based on the label category of the kd-tree, the matching value is obtained according to the Euclidean distance between pixels in the image, when the Euclidean distance between the extension point of the kd-tree and the skeleton point is less than the Euclidean distance between pixels in the image, it is considered that the skeleton point is matched, so as to obtain the corresponding area of the point set of the high-definition map in the partition image.

6. A visual AI-based fusion positioning device, characterized in that, comprise: a processor; a memory having executable instructions of the processor stored therein; wherein the processor is configured to perform the steps of the fusion positioning method based on visual AI of any one of claims 1-4 by executing the executable instructions.

7. A computer readable storage medium for storing a program, characterized in that, The program is executed to realize the steps of the fusion positioning method based on visual AI of any one of claims 1-4. The program is executed to realize the steps of the fusion positioning method based on visual AI of any one of claims 1-4.

Citation Information

Patent Citations

  • Coarse-to-fine multi-sensor fusion positioning method based on semantic edge alignment

    CN113920198A

  • Positioning pose calibration method based on vision and map lane line matching and automobile

    CN114396957A