Positioning method and device of vehicle, vehicle and storage medium

By performing semantic segmentation and depth estimation on panoramic images, a 3D projection plane is constructed and semantic features are fused to generate a semantic map. This solves the problems of insensitive environmental perception and unstable pose estimation in automatic parking, and achieves more accurate positioning and a more intelligent automatic parking solution.

CN116612190BActive Publication Date: 2025-12-16CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310577629.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-12-16
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

In existing technologies, automatic parking suffers from poor environmental perception, unstable pose estimation, and inaccurate global positioning, resulting in a poor user experience.

Method used

By acquiring panoramic images around the vehicle, semantic segmentation is performed to obtain ground and spatial semantic features. Depth estimation is performed by combining the vehicle's pose state, a 3D projection plane is constructed and semantic features are projected, and features from multiple frames are tracked and fused to generate a semantic map for localization.

Benefits of technology

It improves the accuracy and stability of vehicle positioning, enhances the intelligence and reliability of automatic parking solutions, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612190B_ABST
    Figure CN116612190B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of vehicles, in particular to a positioning method and device of a vehicle, a vehicle and a storage medium, wherein the method comprises the following steps: obtaining ground semantic features and space semantic features by performing semantic segmentation on a panoramic image; projecting the space semantic features to a three-dimensional projection plane to obtain space features, and projecting the ground semantic features to a ground projection plane to obtain ground features; tracking and fusing the ground features and the space features of multiple panoramic images to obtain ground semantic feature objects and space semantic feature objects; constructing a semantic map based on the ground semantic feature objects and the space semantic feature objects; and positioning the current position of the vehicle by using the semantic map and the ground semantic feature objects and the space semantic feature objects of the region where the vehicle is currently located. Thus, the problems in the prior art that automatic parking environment perception is not sensitive, pose estimation is unstable, global positioning is not accurate, and the user experience is poor when the user uses the vehicle are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a positioning method and device of a vehicle, a vehicle and a storage medium. BACKGROUND

[0002] The positioning function is an important part of automatic parking. Since the parking scene is usually narrow and complex, the accuracy of the vehicle perceiving its own pose affects the accuracy and rationality of the route calculation of the control module.

[0003] In the related art, a method using a combination of cameras, millimeter waves, laser radars, ultrasonic waves, IMUs (Inertial Measurement Units) and wheel speed meters as sensors is used, but the odometer part has cumulative errors, the laser radar matching algorithm has large calculation amount, slow repositioning speed, high cost and poor popularity.

[0004] In the related art, a visual positioning algorithm can also be used, which is usually divided into a feature point-based positioning algorithm and a visual semantic-based positioning algorithm. The feature point-based algorithm uses a front-view or rear-view camera, which is susceptible to interference and has poor stability. The visual semantic-based positioning algorithm uses a surround-view camera to collect images, and obtains scene semantic features through surround-view splicing and semantic segmentation. The features are less and the repositioning range is limited. SUMMARY

[0005] The present application provides a positioning method and device of a vehicle, a vehicle and a storage medium to solve the problems in the related art that automatic parking environment perception is not sensitive, pose estimation is unstable, global positioning is not accurate, and user experience is poor when using.

[0006] The first aspect of the present application provides a positioning method of a vehicle, comprising the following steps: acquiring a panoramic image around the vehicle; performing semantic segmentation on the panoramic image to obtain ground semantic features and spatial semantic features, performing depth estimation on a space point according to a pose state of the vehicle and the panoramic image to obtain depth estimation information; determining a three-dimensional projection plane according to the depth estimation information, projecting the spatial semantic features to the three-dimensional projection plane to obtain spatial features, and projecting the ground semantic features to a ground projection plane to obtain ground features; tracking and fusing the ground features and the spatial features of multiple frames of panoramic images to obtain ground semantic feature objects and spatial semantic feature objects, constructing a semantic map based on the ground semantic feature objects and the spatial semantic feature objects, and positioning a current position of the vehicle using the semantic map and ground semantic feature objects and spatial semantic feature objects in a region where the vehicle is currently located.

[0007] According to the technical means, the embodiment of the present application can rely on the collection device around the vehicle body to perceive the panoramic image, extract key semantic information from the panoramic image, construct a scene stereoscopic projection plane by depth information, associate and project the semantic information to the predicted plane to form a semantic object, and generate a semantic feature object of the global map by tracking and optimizing the semantic object. The semantic feature is more stable, the feature element is more complete, the feature object matching speed is faster, and the success rate is higher. Therefore, the semantic map is constructed by the spatial and ground semantic features for positioning. Since the spatial objects are fully considered for the help of positioning, the accuracy of vehicle positioning can be improved. The positioning in a closed space can be used. When applied to the scene of automatic parking, the automatic parking scheme can be more intelligent and reliable, and the user experience can be improved.

[0008] Optionally, the depth estimation information of the space point according to the pose state of the vehicle and the panoramic image comprises: identifying a ground line of an entity around the vehicle in the panoramic image and the ground; projecting the ground line to the ground projection plane to obtain the distribution coordinates of the ground line on the ground projection plane, constructing a straight line in a three-dimensional space according to the distribution coordinates to obtain a plurality of straight lines along the ground line; fitting one or more fitting planes by using randomly sampled points on the plurality of straight lines, retaining a fitting plane with an area greater than a preset area in the one or more fitting planes, and obtaining the depth estimation information of the space point in combination with the depth estimation result of the discrete points in the space.

[0009] According to the technical means, the embodiment of the present application can obtain the ground semantic feature and the spatial semantic feature by semantic segmentation of the panoramic image, and then estimate the depth of the space point according to the pose state of the vehicle and the panoramic image to obtain the depth estimation information. The influence of noise on depth estimation is reduced, the stability of depth estimation is improved, and then the stability of the fitting plane is improved, and the error is reduced.

[0010] Optionally, the projecting the spatial semantic feature to the three-dimensional projection plane to obtain a spatial feature comprises: identifying a moving object feature in the spatial semantic feature; filtering the moving object feature in the spatial semantic feature, and associating a space point in the spatial semantic feature to the nearest fitting plane in combination with the depth estimation information; taking the distance between the space point and the fitting plane and the category of the space point as the spatial feature, and obtaining the spatial features of all space points after traversing all space points.

[0011] According to the technical means, the embodiment of the present application can effectively retain the semantic special elements in the space, improve the stability of the semantic feature, improve the matching speed and success rate, improve the intelligence of the scheme, and improve the user experience.

[0012] Optionally, the tracking and fusing of the ground features and the space features of the multiple panoramic images respectively to obtain the ground semantic feature object and the space semantic feature object comprises: obtaining a fitting plane based on panoramic images in a moving direction of the vehicle; constructing a binding box outside a ground line of the fitting plane, taking the fitting plane of the first frame as a search starting point, searching for a fitting plane of an adjacent frame with a binding box similarity greater than a preset threshold to obtain an associated plane; fusing the same space features on the associated plane, and updating new space features, searching for a boundary range of the associated plane, and cropping along the boundary range, and obtaining the space semantic feature object after the number of iterations is greater than a preset number.

[0013] According to the above technical means, the embodiment of the application can introduce a feature object, that is, the projected semantic points are clustered to construct a single observation object, and then the objects observed multiple times are tracked and fused to obtain a semantic feature object that can be stably represented, thereby the embodiment of the application can improve stability, improve feature matching speed and success rate, improve scheme intelligence, and improve user experience.

[0014] Optionally, the projecting the ground semantic features to a ground projection plane to obtain ground features comprises: projecting the ground semantic features to the ground projection plane to obtain semantic information of the ground; identifying ground elements of the semantic information, constructing a feature area according to the ground elements, constructing a feature area at an interval of a preset distance, and calculating the ground features by using the ground elements of the feature area.

[0015] According to the above technical means, the embodiment of the application can enhance the anti-interference and stability of the semantic features, reduce the complexity of semantic processing, and quickly and stably realize the calculation and extraction of the semantic features.

[0016] Optionally, the tracking and fusing of the ground features and the space features of the multiple panoramic images respectively to obtain the ground semantic feature object and the space semantic feature object comprises: matching a sampling box of each ground feature; merging the sampling boxes at the same position, and fusing the same ground elements to obtain the ground semantic feature object.

[0017] According to the above technical means, the embodiment of the application can track and fuse the features by projecting or associating the features to a plane and clustering the features, which provides a possibility for creating an omnidirectional feature object in a scene.

[0018] Optionally, the constructing the semantic map based on the ground semantic feature object and the space semantic feature object comprises: arranging the space semantic feature object according to position height, compressively encoding the arranged feature to obtain a space feature vector; splicing the ground semantic feature object according to region, compressively encoding the spliced feature to obtain a ground feature vector; identifying the space feature vector and the ground feature vector of the same road section according to a pre-constructed topological road, and arranging the space feature vector and the ground feature vector according to road section order to obtain a feature sequence of the road section, and constructing the semantic map based on the feature sequence of all road sections.

[0019] According to the technical means described above, the embodiments of the present application can use feature objects for positioning, construct a semantic map, obtain a more accurate and perfect map, improve matching speed, improve matching success rate, and make the parking scheme more accurate and reliable.

[0020] Optionally, the positioning the current position of the vehicle using the semantic map and the ground semantic feature object and the space semantic feature object of the region where the vehicle is currently located comprises: determining one or more candidate regions from the semantic map according to the ground semantic feature object and the space semantic feature object of the region where the vehicle is currently located; performing feature matching between the current region and the candidate regions, taking the candidate region with the most matching points as a matching region, and determining the current position of the vehicle according to the position of the matching region on the semantic map.

[0021] According to the technical means described above, the embodiments of the present application can use feature objects for positioning, make positioning more accurate, improve positioning and matching efficiency, save scheme execution time, improve parking intelligence, and improve user experience.

[0022] Optionally, the acquiring the panoramic image around the vehicle comprises: simultaneously triggering a plurality of collection devices on the vehicle, and generating the panoramic image according to images collected by the plurality of collection devices, wherein the plurality of collection devices have consistent installation heights.

[0023] According to the technical means described above, the embodiments of the present application can reduce collection errors by uniformly controlling the installation heights of the plurality of collection devices, fully utilize the collection range of the collection devices, reduce costs, and ensure the richness of features and the reliability of positioning.

[0024] Optionally, the three-dimensional projection plane is a solid plane around the vehicle.

[0025] According to the technical means described above, the embodiments of the present application can preferentially select a stand column and a wall surface in a scene as a projection plane instead of a virtual plane designed by a human being, improve the consistency of plane creation at different times, effectively improve the environment coverage and comprehensiveness of semantic object construction during mapping and positioning, and help track and fuse semantic features in multiple frames.

[0026] The second aspect embodiment of the present application provides a positioning device of a vehicle, comprising: an acquisition module configured to acquire a panoramic image around the vehicle; an estimation module configured to perform semantic segmentation on the panoramic image to obtain ground semantic features and space semantic features, and perform depth estimation on a space point based on a pose state of the vehicle and the panoramic image to obtain depth estimation information; a projection module configured to determine a three-dimensional projection plane based on the depth estimation information, project the space semantic features to the three-dimensional projection plane to obtain space features, and project the ground semantic features to a ground projection plane to obtain ground features; and a positioning module configured to track and fuse the ground features and the space features of multiple frames of panoramic images to obtain a ground semantic feature object and a space semantic feature object, construct a semantic map based on the ground semantic feature object and the space semantic feature object, and locate a current position of the vehicle by using the semantic map and the ground semantic feature object and the space semantic feature object of a region where the vehicle is currently located.

[0027] Optionally, the estimation module is further configured to: identify a ground line of an entity around the vehicle in the panoramic image; project the ground line to the ground projection plane to obtain distribution coordinates of the ground line on the ground projection plane, and construct a space straight line in a three-dimensional space based on the distribution coordinates to obtain a plurality of straight lines along the ground line; and fit one or more fitting planes by using randomly sampled points on the plurality of straight lines, retain a fitting plane with an area greater than a preset area from the one or more fitting planes, and obtain the depth estimation information of the space point in combination with a depth estimation result of a discrete point in the space.

[0028] Optionally, the projection module is further configured to: identify a moving object feature in the space semantic features; filter the moving object feature in the space semantic features, and associate a space point in the space semantic features to a nearest fitting plane in combination with the depth estimation information; and take a distance of the space point from the fitting plane and a category of the space point as the space feature, and obtain the space features of all the space points after traversing all the space points.

[0029] Optionally, the positioning module is further configured to: acquire a fitting plane created based on a panoramic image in a motion direction of the vehicle; construct a binding box outside a ground line of the fitting plane, search for a fitting plane of a neighboring frame with a similarity greater than a preset threshold to a binding box of the fitting plane of a first frame to obtain an associated plane; fuse the same space features on the associated plane, and update new space features, search for a boundary range of the associated plane, and crop along the boundary range, and obtain the space semantic feature object after a number of iterations of optimization is greater than a preset number.

[0030] Optionally, the projection module is further configured to project the ground semantic features to the ground projection plane to obtain semantic information of the ground, identify ground elements of the semantic information, construct a feature region according to the ground elements, construct a feature region at a preset interval, and calculate the ground features by using the ground elements of the feature region.

[0031] Optionally, the positioning module is further configured to match the sampling frame of each ground feature, merge the sampling frames at the same position, and fuse the same ground elements to obtain the ground semantic feature object.

[0032] Optionally, the positioning module is further configured to arrange the spatial semantic feature object according to the position height, compress and encode the arranged features to obtain a spatial feature vector, splice the ground semantic feature object according to the region, compress and encode the spliced features to obtain a ground feature vector, identify the spatial feature vector and the ground feature vector of the same road section according to a pre-constructed topological road, arrange the feature vectors according to the road section order to obtain a feature sequence of the road section, and construct the semantic map based on the feature sequences of all road sections.

[0033] Optionally, the positioning module is further configured to determine one or more candidate regions from the semantic map according to the ground semantic feature object and the spatial semantic feature object of a current region where the vehicle is located, perform feature matching on the current region and the candidate regions, select a candidate region with the most matching points as a matching region, and determine the current position of the vehicle according to the position of the matching region on the semantic map.

[0034] Optionally, the acquisition module is further configured to trigger multiple acquisition devices on the vehicle at the same time, and generate the panoramic image according to images collected by the multiple acquisition devices, wherein the multiple acquisition devices are installed at the same height.

[0035] Optionally, the three-dimensional projection plane is a solid plane around the vehicle.

[0036] The third aspect of the embodiments of the present application provides a vehicle, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the positioning method of the vehicle as described in the above embodiments.

[0037] The fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program executable by a processor to implement the positioning method of the vehicle as described in the above embodiments.

[0038] Therefore, the present application has at least the following beneficial effects:

[0039] (1) The embodiment of the present application can rely on the collection equipment around the vehicle body to perceive panoramic images, extract key semantic information from the panoramic image, construct a scene stereoscopic projection plane by depth information, associate and project the semantic information to the predicted plane to form a semantic object, and generate a semantic feature object of the global map through tracking and optimization of the semantic object. The semantic features are more stable, the feature elements are more complete, the feature object matching speed is faster and the success rate is higher, so as to construct a semantic map for positioning through the semantic features of space and ground. Since the help of spatial objects for positioning is fully considered, the accuracy of vehicle positioning can be improved, and the positioning in a closed space can be used. When applied to the scene of automatic parking, the automatic parking scheme can be more intelligent and reliable, and the user experience can be improved;

[0040] (2) The embodiment of the present application can obtain ground semantic features and spatial semantic features through semantic segmentation of the panoramic image, and then estimate the depth of the spatial point according to the pose state of the vehicle and the obtained panoramic image to obtain depth estimation information, reduce the influence of noise on depth estimation, improve the stability of depth estimation, and then improve the stability of the fitting plane and reduce errors;

[0041] (3) The embodiment of the present application can effectively retain semantic special elements in space, improve the stability of semantic features, improve the matching speed and success rate, improve the intelligence of the scheme, and improve the user experience;

[0042] (4) The embodiment of the present application can introduce a feature object, that is, the projected semantic points are clustered to construct an object observed at a single time, and then the objects observed at multiple times are tracked and fused to obtain a semantic feature object that can be stably represented. Therefore, the embodiment of the present application can improve stability, improve feature matching speed and success rate, improve the intelligence of the scheme, and improve the user experience;

[0043] (5) The embodiment of the present application can enhance the anti-interference and stability of semantic features, reduce the complexity of semantic processing, and quickly and stably realize the calculation and extraction of semantic features;

[0044] (6) The embodiment of the present application can project or associate the features to the plane and cluster the features, which is beneficial to the tracking and fusion of the features and provides the possibility of creating an omnidirectional feature object in the scene;

[0045] (7) The embodiment of the present application can use the feature object for positioning, construct a semantic map, obtain a more accurate and perfect map, improve the matching speed, improve the matching success rate, make the parking scheme more accurate and reliable, and improve the user experience;

[0046] (8) The embodiment of the present application can use feature objects for positioning, so that the positioning is more accurate, the positioning matching efficiency is improved, the scheme execution time is saved, the parking intelligence is improved, and the user experience is improved;

[0047] (9) The embodiment of the present application can reduce the collection error by uniformly controlling the installation height of the plurality of collection devices, fully utilize the collection range of the collection device, reduce the cost, and ensure the richness of the features and the reliability of the positioning;

[0048] (10) The embodiment of the present application can preferentially select a column and a wall surface in a scene as a projection surface instead of a virtual plane artificially designed, improve the consistency of plane creation at different times, effectively improve the environment coverage and comprehensiveness of semantic object construction during mapping and positioning, and help the semantic features to be tracked and fused in multiple frames.

[0049] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 a flowchart of a positioning method of a vehicle according to an embodiment of the present application;

[0051] Figure 2 a panoramic projection schematic diagram according to an embodiment of the present application;

[0052] Figure 3 a detailed flowchart of panoramic mapping according to an embodiment of the present application;

[0053] Figure 4 an example diagram of a positioning device of a vehicle according to an embodiment of the present application;

[0054] Figure 5 a structural schematic diagram of a vehicle according to an embodiment of the present application. DETAILED DESCRIPTION

[0055] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0056] In the related art, underground parking lot visual positioning algorithms are generally based on features and semantics.

[0057] The feature-based algorithm can adopt a forward-looking or rear-view camera, and the mapping can be performed by a monocular camera. The monocular camera-based visual positioning algorithm has the following problems: (1) when the vehicle enters a new scene in a certain direction, the light changes suddenly, and the camera will have a dynamic adjustment process, which causes the feature points of the front and rear frames to be unable to be correctly matched, and the pose estimation will have a large deviation; (2) in a dark environment, the monocular camera can not match enough feature points, resulting in positioning failure; (3) the monocular camera can only observe the features in one direction, so the camera needs to have the same orientation during mapping and positioning. If the vehicle passes through the same road section in different driving directions, two trajectories containing feature points in different directions need to be constructed during mapping. Since each trajectory has features in only one direction, due to the cumulative error, the two trajectories can have a large deviation.

[0058] The semantic positioning algorithm can adopt a surround-view camera to collect images, and obtain the semantic features of the scene through surround-view splicing and semantic segmentation. In related algorithms, the ground features only include lane lines, speed bump lines, and warehouse lines, and the ground semantics are relatively sparse, and each observation contains fewer semantic features. Generally, the warehouse lines and lane lines on the ground have high similarity. During positioning, searching for matching features on the map through semantic features will obtain many results, and the positioning initialization cannot be completed.

[0059] The positioning method, device, vehicle, and storage medium of the vehicle according to the embodiments of the present application are described below with reference to the accompanying drawings.

[0060] Specifically, Figure 1 A flowchart of a positioning method of a vehicle provided by an embodiment of the present application is shown in FIG. 1.

[0061] As Figure 1 shown, the positioning method of the vehicle includes the following steps:

[0062] In step S101, a panoramic image around the vehicle is acquired.

[0063] It can be understood that the embodiment of the present application can first acquire the panoramic image around the vehicle to facilitate the segmentation of the image in the subsequent steps. The panoramic projection diagram of the embodiment of the present application can be as shown in FIG. 2; and the embodiment of the present application can acquire the panoramic image around the vehicle in at least one way, which is not limited specifically. Figure 2

[0064] As a possible implementation, acquiring the panoramic image around the vehicle includes simultaneously triggering a plurality of collection devices on the vehicle, and generating the panoramic image according to the images collected by the plurality of collection devices.

[0065] ​The installation heights of the plurality of collection devices are consistent. The plurality of collection devices can be set and selected according to actual conditions, and the acquisition ranges of the plurality of collection devices can cover all scenes around the vehicle body, for example, front-view cameras, rear-view cameras, left-view and right-view cameras can be set, and no specific limitation is made in this regard.

[0066] It can be understood that the embodiments of the present application can reduce collection errors, fully utilize the collection range of the collection device, and reduce costs through unified control of the installation heights of the plurality of collection devices. For example, the embodiments of the present application can install four fisheye cameras with consistent heights around the vehicle body as collection devices, wherein the four fisheye cameras can be used as front-view cameras, rear-view cameras, and left and right side cameras of the vehicle. The embodiments of the present application can trigger the four fisheye cameras to take pictures covering the scenes around the vehicle body at the same time, thereby achieving acquisition of the scenes around the vehicle.

[0067] In step S102, the semantic segmentation panoramic image obtains ground semantic features and spatial semantic features, and depth estimation information is obtained by depth estimation of the spatial points according to the pose state of the vehicle and the panoramic image.

[0068] The ground semantic features can include semantic information of ground artificial groups extracted from the original image by the deep learning model, for example, can include traffic lanes, speed bumps, and storage location information. The spatial semantic features can include image features around the vehicle body, for example, can include walls, green belts, and fences around the vehicle body. The spatial semantic features in the embodiments of the present application preferentially select a center column or a wall as a projection surface, rather than a virtual plane designed by humans, which will be described in detail in the following embodiments, and will not be described here. The depth estimation information can be an estimated value of the distance between the semantic features and the actual pose of the vehicle body, which can be calculated by at least one way, and no specific limitation is made in this regard.

[0069] It can be understood that the embodiments of the present application can implement semantic segmentation of the panoramic image by at least one way, and no specific limitation is made in this regard. For example, the embodiments of the present application can use a model segmentation method to implement semantic segmentation of the panoramic image: the embodiments of the present application can sequentially input the images collected by each camera into a network model improved based on UNet (a full convolutional network containing 4 layers of down-sampling, 4 layers of up-sampling, and a similar skip connection structure), to obtain the semantic segmentation result of the panoramic image acquired in step S101, and combine the device information of the plurality of collection devices in the above steps, to separate the semantic segmentation result of the non-ground element from the semantic features of the ground, and make them as the features of the space and the ground, respectively.

[0070] In the embodiment of the present application, the depth estimation information of the space point is obtained according to the pose state of the vehicle and the panoramic image, including: identifying the ground line of the entity around the vehicle in the panoramic image and the ground; projecting the ground line to the ground projection plane to obtain the distribution coordinates of the ground line on the ground projection plane, and constructing a space straight line in a three-dimensional space according to the distribution coordinates to obtain a plurality of straight lines along the ground line; fitting one or more fitting planes by using randomly sampled points on the plurality of straight lines, retaining the fitting plane with an area greater than a preset area in the one or more fitting planes, and obtaining the depth estimation information of the space point in combination with the depth estimation result of the discrete points in the space.

[0071] It can be understood that, the embodiment of the present application can obtain the ground semantic features and the space semantic features of the panoramic image through semantic segmentation of the panoramic image, and then obtain the depth estimation information of the space point according to the pose state of the vehicle and the panoramic image, so as to reduce the influence of noise on the depth estimation, improve the stability of the depth estimation, and further improve the stability of the fitting plane and reduce the error.

[0072] Specifically, the embodiment of the present application can perform IPM (Inverse Perspective Mapping) inverse projection on the semantic region of the ground according to the above-mentioned calibrated external reference to the ground, to obtain the depth information of the ground point, and fit the ground plane by projecting the ground point; the ground line of the wall surface and the ground is detected by semantic segmentation, the ground line is projected to the ground plane by inverse projection to obtain the distribution coordinates of the ground line on the ground; a space straight line represented by a plucker coordinate system is constructed in a 3D (3 Dimension) space through these coordinate points to obtain a plurality of straight lines along the ground line; then some points are randomly taken on the above-mentioned straight lines, and then a plurality of optimal planes are fitted by using ransc (Random Sample Consensus) to retain the plane with the largest area, and the remaining straight lines are retained; the depth estimation of the discrete points in the space is obtained through inter-frame optical flow tracking of the motion of the vehicle, and the depth points appearing behind the plane are filtered out in combination with the plane fitted by the ground line, and the remaining points are retained. Thus, the single-frame depth estimation and plane fitting are completed.

[0073] It should be noted that, in the pose state of the vehicle and the panoramic image, in order to constrain and optimize the scale in the depth estimation, reduce the error, the application embodiment can also perform track prediction, wherein the track prediction can be realized by at least one way, which is not limited. For example, the application embodiment can use a track prediction method: first, use imu and wheel speed meter as input, then fuse and integrate the acceleration and angular velocity measurement values of the imu after removing the initial zero offset; construct an error state space of the vehicle body pose, including pose error and zero offset error term; use gravity and wheel speed to predict and update the error state of the pose, and obtain the error estimation of the pose and zero offset through iterative update; integrate the predicted pose and wheel speed to obtain the translation of the vehicle body; use the translation to constrain and optimize the scale in the depth estimation.

[0074] In step S103, a three-dimensional projection plane is determined according to the depth estimation information, a space semantic feature is projected to the three-dimensional projection plane to obtain a space feature, and a ground semantic feature is projected to a ground projection plane to obtain a ground feature.

[0075] In the application embodiment, the three-dimensional projection plane is a solid plane around the vehicle.

[0076] It can be understood that the application embodiment can construct a scene stereoscopic projection plane from the depth information, associate and project the semantic information to the predicted plane to form a semantic object, generate a semantic feature object of the global map through tracking and optimization of the semantic object, enhance the anti-interference and stability of the semantic feature, reduce the complexity of semantic processing, and quickly and stably realize the calculation and extraction of the semantic feature.

[0077] In the application embodiment, projecting the space semantic feature to the three-dimensional projection plane to obtain the space feature includes: identifying moving object features in the space semantic feature; filtering the moving object features in the space semantic feature, and associating the space points in the space semantic feature to the nearest fitting plane in combination with the depth estimation information; taking the distance of the space point to the fitting plane and the category of the space point as the space feature, and obtaining the space features of all space points after traversing all space points.

[0078] It can be understood that, after the application embodiment excludes the elements corresponding to the moving object from the non-ground semantic elements as in step S102, the point is associated to the nearest fitting plane in combination with the depth estimation information of the point; the distance of the point to the plane and the category of the point are taken as the space feature of the point; after traversing all space points, the feature vector of the effective point about the associated plane is obtained; the above feature vector is saved to the descriptor queue of the plane, which can be used as the feature matrix of the plane. Therefore, the application embodiment can effectively retain the semantic special elements in the space, improve the stability of the semantic feature, and improve the matching speed and success rate.

[0079] In the embodiment of the present application, the ground semantic feature is projected to the ground projection plane to obtain the ground feature, including: projecting the ground semantic feature to the ground projection plane to obtain the semantic information of the ground; identifying the ground elements of the semantic information, constructing a feature area according to the ground elements, constructing a feature area at a preset distance interval, and calculating the ground feature by using the ground elements of the feature area.

[0080] The preset distance can be calibrated according to the actual situation, such as 1 meter, etc., which is not limited here; for example, the embodiment of the present application can directly project the semantic elements of the ground to the ground according to the calibrated extrinsic parameters through IPM inverse projection to obtain the semantic information of the ground; then taking the direction of the road as the main direction, a feature area is constructed at each fixed distance, and the ground feature is calculated by taking the ground elements in the area.

[0081] It can be understood that the embodiment of the present application can continue to construct the ground semantic feature area containing the ground semantic elements according to the ground semantic information, calculate the ground elements by using the area to obtain the ground feature, enhance the anti-interference and stability of the semantic feature, reduce the complexity of semantic processing, and quickly and stably realize the calculation and extraction of the semantic feature.

[0082] In step S104, the ground features and spatial features of the multiple panoramic images are tracked and fused respectively to obtain ground semantic feature objects and spatial semantic feature objects, a semantic map is constructed based on the ground semantic feature objects and the spatial semantic feature objects, and the current position of the vehicle is located by using the semantic map and the ground semantic feature objects and the spatial semantic feature objects of the current area where the vehicle is located.

[0083] It can be understood that after obtaining the spatial features and the ground features, the embodiment of the present application can generate the semantic feature objects of the global map by tracking and optimizing the semantic objects, so that the semantic features are more stable, the feature elements are more complete, the feature object matching speed is faster and the success rate is higher, so that the semantic map is constructed by the spatial and ground semantic features for positioning. Since the help of the spatial object for positioning is fully considered, the accuracy of vehicle positioning can be improved, and the positioning in a closed space can be used. When applied to the scene of automatic parking, the automatic parking scheme can be more intelligent and reliable, and the user experience can be improved.

[0084] In the embodiment of the present application, the ground features and spatial features of the multiple panoramic images are tracked and fused respectively to obtain the content of the ground semantic feature objects and the spatial semantic feature objects as follows:

[0085] (1) Obtain the spatial semantic feature object:

[0086] In the embodiment of the present application, the ground features and the space features of the multiple panoramic images are tracked and fused respectively to obtain the ground semantic feature object and the space semantic feature object, including: obtaining a fitting plane created based on the panoramic image in the moving direction of the vehicle; constructing a binding box outside the ground line of the fitting plane, taking the fitting plane of the first frame as the search starting point, searching for the fitting plane of the adjacent frame with a binding box similarity greater than a preset threshold to obtain an associated plane; fusing the same space features on the associated plane, and updating the new space features, searching for the boundary range of the associated plane, and after the number of iterative optimizations is greater than a preset number, obtaining the space semantic feature object.

[0087] In the embodiment of the present application, the acquisition of the fitting plane can be as shown in step S103; the preset threshold and the preset number of the embodiment of the present application can be set according to the actual situation, and no specific limitation is made.

[0088] It should be noted that the search and matching of the binding box in the embodiment of the present application can be implemented in at least one way, such as constructing a similarity measurement function and using the Hungarian matching algorithm for search and matching, and no specific limitation is made.

[0089] Specifically, the embodiment of the present application can take the feature plane created by the first panoramic image as a reference, then search for the corresponding feature plane in the adjacent area of the subsequent frame in combination with the vehicle movement, construct a binding box outside the ground line of the feature plane, and search and match the binding boxes of the front and rear frames. When the feature plane that has completed matching is obtained, the plane equation is optimized, and then the associated feature queue is matched, the features with consistent descriptors on the same plane are fused and updated, and if there are new feature points, they are added to the feature queue. Then search the boundary range of the plane, and crop the plane into a rectangular area. After multiple iterative optimizations, the acquisition of the space semantic feature object is completed.

[0090] It can be understood that the embodiment of the present application can introduce a feature object, that is, the projected semantic points are first clustered to construct a single observation object, and then the objects observed multiple times are tracked and fused to obtain a semantic feature object that can be stably represented. Therefore, the embodiment of the present application can improve stability, improve feature matching speed and success rate, improve the intelligence of the scheme, and improve the user experience.

[0091] (2) Obtain the ground semantic feature object:

[0092] In the embodiment of the present application, the ground features and the space features of the multiple panoramic images are tracked and fused respectively to obtain the ground semantic feature object and the space semantic feature object, including: matching the sampling box of each ground feature; merging the sampling boxes at the same position, and fusing the same ground elements to obtain the ground semantic feature object.

[0093] It should be understood that, for the fusion of the same ground elements, at least one manner can be used, such as the fusion by using a calculation manner, and the like, which is not specifically limited.

[0094] It should be understood that, for the ground features, the corresponding sampling frames of the ground features can be matched, the sampling frames at the same position can be merged, the semantic elements can be fused, the semantic feature objects are obtained, and then the features are projected or associated to the plane and the features are clustered, which is beneficial to the tracking and fusion of the features, and provides a possibility for creating omnidirectional feature objects in a scene. When applied to the scene of automatic parking, the automatic parking can be made more reliable.

[0095] In the embodiment of the present application, the semantic map is constructed based on the ground semantic feature objects and the space semantic feature objects, including: arranging the space semantic feature objects according to the position height, compressing and encoding the arranged features to obtain a space feature vector; splicing the ground semantic feature objects according to the region, compressing and encoding the spliced features to obtain a ground feature vector; identifying the space feature vector and the ground feature vector of the same road section according to the pre-constructed topological road, and arranging the feature vectors according to the road section sequence to obtain a feature sequence of the road section, and constructing the semantic map based on the feature sequences of all road sections.

[0096] It should be understood that, in the embodiment of the present application, the features can be encoded by using at least one manner, such as a model based on a variational autoencoder design, and the like, which is not specifically limited.

[0097] It should be understood that, in the embodiment of the present application, the feature objects can be used for positioning, the semantic map can be constructed, a more accurate and perfect map can be obtained, the matching speed can be improved, the matching success rate can be improved, and the parking scheme can be more accurate and reliable.

[0098] Specifically, in the embodiment of the present application, a map containing semantic feature objects can be constructed after the vehicle passes through a complete mapping area. First, the semantic vectors projected in the feature plane are arranged from low to high according to the position, and the semantic features are further compressed and encoded to obtain a feature vector; for the ground features, the compression and encoding are performed. The features in a certain region are spliced together to obtain a regional feature object. According to the constructed topological road, the features of the same road section are arranged in sequence to obtain a feature sequence of the road section, and the feature sequences of all road sections constitute a feature map of the scene.

[0099] In the embodiment of the present application, the current position of the vehicle is located by using the semantic map and the ground semantic feature object and the space semantic feature object in the region where the vehicle is currently located, including: determining one or more candidate regions from the semantic map according to the ground semantic feature object and the space semantic feature object in the region where the vehicle is currently located; performing feature matching on the current region and the candidate regions, taking the candidate region with the most matching points as a matching region, and determining the current position of the vehicle according to the position of the matching region on the semantic map.

[0100] In the embodiment of the present application, the current position of the vehicle is located by using the semantic map and the ground semantic feature object and the space semantic feature object in the region where the vehicle is currently located, including: determining one or more candidate regions from the semantic map according to the ground semantic feature object and the space semantic feature object in the region where the vehicle is currently located; performing feature matching on the current region and the candidate regions, taking the candidate region with the most matching points as a matching region, and determining the current position of the vehicle according to the position of the matching region on the semantic map.

[0101] It can be understood that the embodiment of the present application can use feature objects for positioning, so that the positioning is more accurate, the positioning matching efficiency is improved, the scheme execution time is saved, the parking intelligence is improved, and the user experience is improved.

[0102] Specifically, in the embodiment of the present application, the feature objects of a certain region constructed by the vehicle are matched with the map, and at this time, multiple sets of candidate matching results can be obtained, and each target to be matched can have multiple candidates; the embodiment of the present application can take the region where the candidate frame is located as a screening condition to further select matching objects in the effective region; when the number of successful matches is greater than a threshold, the region is selected as a matching point to obtain a coarse matching coordinate, and then the ground features and the feature plane of the region on the map are extracted, the ground features and the feature plane observed at this moment are matched, and the accurate pose of the vehicle body is solved according to the matched feature points.

[0103] The vehicle positioning method of the present application will be explained below through a specific embodiment, as shown in Figure 3 The embodiment of the present application can use a vehicle positioning method, and the specific steps are as follows:

[0104] The embodiment of the present application can be used for mapping and positioning in an indoor site, and is a mapping algorithm based on pillars and walls. The embodiment of the present application adopts a specially designed surround-view camera combination scheme, including front-view cameras, rear-view cameras, left and right side-view cameras, wherein the front-view and rear-view cameras are consistent with the general scheme in the related art; the left and right side-view cameras are installed below the A-pillar of the vehicle and are the same height as the front and rear-view cameras. The embodiment of the present application also includes IMU, wheel speed meter and other automatic driving standard equipment.

[0105] Specifically, the mapping includes: using a semantic segmentation map as a template, taking out the area on the depth map corresponding to the classification as a column or a wall surface, fitting the plane therein. Then combine other features with the depth map, project or correlate to these planes. Combined with the motion state of the vehicle, monitor and track these feature planes in the front and rear frames, and update the feature planes through an optimization algorithm. Fuse multiple adjacent feature planes of the same type, column or wall surface, to generate feature objects. When all feature objects in the scene are created, the feature object map of the scene is obtained. The positioning process is similar to mapping, both of which create object maps, but positioning only needs to create a local small area map, then compare the local map with the global map to search for matching candidate areas, when the matching score is greater than the threshold, the matching area is obtained, and the pose mapping relationship is established through the matching pairs of objects in the area, and the current pose of the vehicle is solved.

[0106] The positioning includes five steps: first, collecting images of front, rear, left and right side view cameras in real time; second, performing semantic segmentation on images of the four cameras through a deep learning model; third, calculating the driving state of the vehicle using an IMU and a wheel speed meter; and fourth, estimating the depth of a spatial point in combination with the pose state and the front and rear frames of the original image.

[0107] In summary, the vehicle positioning method according to the embodiments of the present application has at least the following beneficial effects:

[0108] (1) The embodiments of the present application can perceive panoramic images by relying on the collection devices around the vehicle body, extract key semantic information from the panoramic images, construct scene stereoscopic projection planes from depth information, associate and project semantic information to predicted planes to form semantic objects, generate semantic feature objects of the global map through tracking and optimization of the semantic objects, and the semantic features are more stable, the feature elements are more complete, the feature object matching speed is faster and the success rate is higher, so as to construct a semantic map for positioning by means of spatial and ground semantic features. Since the help of spatial objects for positioning is fully considered, the accuracy of vehicle positioning can be improved, and the positioning can be used in a closed space. When applied to the scene of automatic parking, the automatic parking scheme can be more intelligent and reliable, and the user experience can be improved;

[0109] (2) The embodiments of the present application can obtain ground semantic features and spatial semantic features through semantic segmentation of panoramic images, estimate the depth of a spatial point according to the pose state of the vehicle and the obtained panoramic image, obtain depth estimation information, reduce the influence of noise on depth estimation, improve the stability of depth estimation, and further improve the stability of the fitted plane and reduce errors;

[0110] (3) The embodiments of the present application can effectively retain semantic special elements in space, improve the stability of semantic features, improve the matching speed and success rate, improve the intelligence of the scheme, and improve the user experience;

[0111] (4) The embodiment of the present application can introduce a feature object, that is, the semantic points after projection are clustered to construct an object observed at a single time, and then the objects observed at multiple times are tracked and fused to obtain a semantic feature object that can be stably represented. Thus, the embodiment of the present application can improve stability, improve feature matching speed and success rate, improve scheme intelligence, and improve user experience;

[0112] (5) The embodiment of the present application can enhance the anti-interference and stability of semantic features, reduce the complexity of semantic processing, and quickly and stably implement the calculation and extraction of semantic features;

[0113] (6) The embodiment of the present application can track and fuse features by projecting or associating features to a plane and clustering features, which makes it possible to create an omnidirectional feature object in a scene;

[0114] (7) The embodiment of the present application can use feature objects for positioning, construct a semantic map, obtain a more accurate and perfect map, improve matching speed, improve matching success rate, make the parking scheme more accurate and reliable, and the like;

[0115] (8) The embodiment of the present application can use feature objects for positioning, so that positioning is more accurate, positioning and matching efficiency is improved, scheme execution time is saved, parking intelligence is improved, and user experience is improved;

[0116] (9) The embodiment of the present application can reduce acquisition errors by uniformly controlling the installation height of multiple acquisition devices, fully utilize the acquisition range of the acquisition device, reduce costs, ensure the richness of features and the reliability of positioning, and the like;

[0117] (10) The embodiment of the present application can preferentially select a column and a wall surface in a scene as a projection surface instead of a virtual plane artificially designed, improve the consistency of plane creation at different times, and help track and fuse semantic features at multiple frames.

[0118] Secondly, a positioning device of a vehicle according to the embodiment of the present application is described with reference to the accompanying drawings.

[0119] Figure 4 is a block schematic diagram of the positioning device of the vehicle according to the embodiment of the present application.

[0120] As shown in Figure 4 , the positioning device 10 of the vehicle includes an acquisition module 100, an estimation module 200, a projection module 300, and a positioning module 400.

[0121] The acquisition module 100 is configured to acquire a panoramic image around a vehicle; the estimation module 200 is configured to perform semantic segmentation on the panoramic image to obtain ground semantic features and space semantic features, and perform depth estimation on a space point according to a pose state of the vehicle and the panoramic image to obtain depth estimation information; the projection module 300 is configured to determine a three-dimensional projection plane according to the depth estimation information, project the space semantic features to the three-dimensional projection plane to obtain space features, and project the ground semantic features to a ground projection plane to obtain ground features; and the positioning module 400 is configured to track and fuse the ground features and the space features of multiple frames of panoramic images to obtain a ground semantic feature object and a space semantic feature object, construct a semantic map based on the ground semantic feature object and the space semantic feature object, and locate a current position of the vehicle by using the semantic map and the ground semantic feature object and the space semantic feature object of a region where the vehicle is currently located.

[0122] In the embodiment of the present application, the estimation module 200 is further configured to identify a ground line of an entity around the vehicle in the panoramic image; project the ground line to the ground projection plane to obtain distribution coordinates of the ground line on the ground projection plane, construct a space straight line in a three-dimensional space according to the distribution coordinates, and obtain multiple straight lines along the ground line; fit one or more fitting planes by using randomly sampled points on the multiple straight lines, retain a fitting plane with an area greater than a preset area in the one or more fitting planes, and obtain the depth estimation information of the space point in combination with a depth estimation result of a discrete point in the space.

[0123] In the embodiment of the present application, the projection module 300 is further configured to identify a moving object feature in the space semantic features; filter the moving object feature in the space semantic features, and associate a space point in the space semantic features to a nearest fitting plane in combination with the depth estimation information; take a distance of the space point to the fitting plane and a category of the space point as the space feature, and obtain the space feature of all the space points after traversing all the space points.

[0124] In the embodiment of the present application, the positioning module 400 is further configured to obtain a fitting plane created based on the panoramic image in a motion direction of the vehicle; construct a binding box outside a ground line of the fitting plane, take the fitting plane of the first frame as a search starting point, search for a fitting plane of a neighboring frame with a similarity greater than a preset threshold to the binding box to obtain an associated plane; fuse the same space features on the associated plane, update new space features, search for a boundary range of the associated plane, and crop along the boundary range to obtain the space semantic feature object after a number of iterations is greater than a preset number.

[0125] In the embodiment of the present application, the projection module 300 is further configured to project the ground semantic features to a ground projection plane to obtain semantic information of the ground; identify ground elements of the semantic information, construct a feature area according to the ground elements, construct a feature area at a preset distance interval, and calculate the ground features by using the ground elements of the feature area.

[0126] In the embodiment of the present application, the positioning module 400 is further configured to match the sampling frame of each ground feature; merge the sampling frames at the same position and fuse the same ground elements to obtain a ground semantic feature object.

[0127] In the embodiment of the present application, the positioning module 400 is further configured to arrange the spatial semantic feature objects according to the position height, compress and encode the arranged features to obtain a spatial feature vector; splice the ground semantic feature objects according to the area, compress and encode the spliced features to obtain a ground feature vector; identify the spatial feature vector and the ground feature vector of the same road section according to the pre-constructed topological road, and arrange the feature sequences of the road sections in sequence to construct a semantic map based on the feature sequences of all the road sections.

[0128] In the embodiment of the present application, the positioning module 400 is further configured to determine one or more candidate areas from the semantic map according to the ground semantic feature objects and the spatial semantic feature objects of the current area where the vehicle is located; perform feature matching between the current area and the candidate areas, take the candidate area with the most matching points as a matching area, and determine the current position of the vehicle according to the position of the matching area on the semantic map.

[0129] In the embodiment of the present application, the acquisition module 100 is further configured to trigger multiple acquisition devices on the vehicle at the same time, and generate a panoramic image according to the images collected by the multiple acquisition devices, wherein the multiple acquisition devices have the same installation height.

[0130] In the embodiment of the present application, the three-dimensional projection plane is a solid plane around the vehicle.

[0131] It should be noted that the above explanation and description of the vehicle positioning method embodiment also applies to the vehicle positioning device of this embodiment, which will not be described here.

[0132] The positioning device of the vehicle provided by the embodiment of the present application can perceive panoramic images by means of the collection devices around the vehicle body, extract key semantic information from the panoramic images, construct a stereoscopic projection plane of a scene by means of depth information, associate and project the semantic information to the predicted plane to form semantic objects, and generate semantic feature objects of a global map by tracking and optimizing the semantic objects. The semantic features are more stable, the feature elements are more complete, the feature object matching speed is faster, and the success rate is higher, so that the semantic map is constructed by means of the semantic features of space and ground for positioning. Since the help of spatial objects for positioning is fully considered, the accuracy of vehicle positioning can be improved, and the positioning in a closed space can be used. When applied to the scenario of automatic parking, the automatic parking scheme can be more intelligent and reliable, and the user experience can be improved.

[0133] Figure 5 The vehicle provided by the embodiment of the present application is shown in the structural schematic diagram. The vehicle can include:

[0134] The memory 501, the processor 502, and the computer program stored in the memory 501 and executable on the processor 502.

[0135] The processor 502 implements the positioning method of the vehicle provided in the above embodiment when executing the program.

[0136] Further, the vehicle further includes:

[0137] The communication interface 503 is used for communication between the memory 501 and the processor 502.

[0138] The memory 501 is used for storing the computer program executable on the processor 502.

[0139] The memory 501 can include a high-speed RAM (Random Access Memory, random access memory) memory, and can also include a non-volatile memory, for example, at least one disk memory.

[0140] If the memory 501, the processor 502, and the communication interface 503 are independently implemented, the communication interface 503, the memory 501, and the processor 502 can be connected to each other through a bus and complete the communication between each other. The bus can be an ISA (Industry Standard Architecture, industrial standard architecture) bus, a PCI (Peripheral Component, peripheral component interconnect) bus, or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5Only one bus or only one type of bus can exist, however.

[0141] Optionally, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can complete the communication among each other through an internal interface.

[0142] The processor 502 can be a CPU (Central Processing Unit, central processor) or an ASIC (Application Specific Integrated Circuit, specific integrated circuit) or one or more integrated circuits configured to implement one or more embodiments of the present application.

[0143] The embodiments of the present application further provide a computer readable storage medium, which has stored a computer program, and the computer program is executed by a processor to implement the positioning method of the vehicle.

[0144] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms is not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or N embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

[0145] In addition, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0146] Any process or method described in a flowchart or otherwise described herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for implementing the specified logical functions or steps, and the various embodiments of the application include alternative implementations of the described processes or methods, in which the order of steps can be changed, including the use of simultaneous steps or reverse order of steps, where necessary and / or desirable.

[0147] It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, the steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware, and in another embodiment, any of the following technologies, or a combination thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions upon an application of data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays, field programmable gate arrays, and the like.

[0148] Those of ordinary skill in the art can understand that all or part of the steps carried out by the above-mentioned embodiment methods can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium, and when executed, include one or a combination of steps of the embodiment methods.

[0149] Although the above has shown and described the embodiments of the present application, it should be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those of ordinary skill in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application.

Claims

1. A positioning method of a vehicle, characterized by, The method comprises the following steps: acquiring a panoramic image around a vehicle; performing semantic segmentation on the panoramic image to obtain ground semantic features and space semantic features, and performing depth estimation on a space point according to a pose state of the vehicle and the panoramic image to obtain depth estimation information; determining a three-dimensional projection plane according to the depth estimation information, projecting the space semantic features to the three-dimensional projection plane to obtain space features, and projecting the ground semantic features to a ground projection plane to obtain ground features; tracking and fusing the ground features and the space features of multiple panoramic images to obtain a ground semantic feature object and a space semantic feature object, constructing a semantic map based on the ground semantic feature object and the space semantic feature object, and locating a current position of the vehicle by using the semantic map and the ground semantic feature object and the space semantic feature object of a region where the vehicle is currently located; the depth estimation on the space point according to the pose state of the vehicle and the panoramic image to obtain the depth estimation information comprises: identifying a ground line of an entity around the vehicle in the panoramic image; projecting the ground line to the ground projection plane to obtain distribution coordinates of the ground line on the ground projection plane, constructing a space straight line in a three-dimensional space according to the distribution coordinates, and obtaining multiple straight lines along the ground line; fitting one or more fitting planes by using randomly sampled points on the multiple straight lines, retaining a fitting plane with an area greater than a preset area in the one or more fitting planes, and obtaining the depth estimation information of the space point in combination with a depth estimation result of a discrete point in the space; the three-dimensional projection plane is an entity plane around the vehicle.

2. The positioning method of a vehicle according to claim 1, characterized by, the projecting the space semantic features to the three-dimensional projection plane to obtain the space features comprises: identifying a moving object feature in the space semantic features; filtering the moving object feature in the space semantic features, and associating a space point in the space semantic features to a nearest fitting plane in combination with the depth estimation information; taking a distance of the space point to the fitting plane and a category of the space point as the space features, and obtaining space features of all the space points after traversing all the space points.

3. The positioning method of a vehicle according to any one of claims 1-2, characterized in that, the tracking and fusing the ground features and the space features of multiple panoramic images to obtain the ground semantic feature object and the space semantic feature object comprises: acquiring a fitting plane created based on a panoramic image in a moving direction of the vehicle; constructing a binding box outside a ground line of the fitting plane, taking the fitting plane of a first frame as a search starting point, searching for a fitting plane of a neighboring frame with a similarity greater than a preset threshold to a binding box, and obtaining an associated plane; fusing the same space features on the associated plane, updating new space features, searching for a boundary range of the associated plane, and cropping along the boundary range, and obtaining the space semantic feature object after a number of iterations is greater than a preset number.

4. The positioning method of a vehicle according to claim 1, characterized by, the projecting the ground semantic features to the ground projection plane to obtain the ground features comprises: projecting the ground semantic features to the ground projection plane to obtain semantic information of the ground; The ground elements of the semantic information are identified, a feature area is constructed according to the ground elements, one feature area is constructed at an interval of a preset distance, and the ground features are calculated by using the ground elements of the feature area.

5. The positioning method of a vehicle according to claim 4, wherein The ground feature and the space feature of the multi-frame panoramic image are tracked and fused respectively to obtain a ground semantic feature object and a space semantic feature object, and the method comprises the following steps of: matching a sampling frame of each ground feature; merging the sampling frames at the same position and fusing the same ground elements to obtain the ground semantic feature object.

6. The positioning method of a vehicle according to claim 1, characterized by, The semantic map is constructed based on the ground semantic feature object and the space semantic feature object, and the method comprises the following steps of: arranging the space semantic feature object according to the position height, and compressively encoding the arranged feature to obtain a space feature vector; splicing the ground semantic feature object according to the area, and compressively encoding the spliced feature to obtain a ground feature vector; identifying the space feature vector and the ground feature vector of the same road section according to a pre-constructed topological road, arranging the space feature vector and the ground feature vector according to the road section sequence to obtain a feature sequence of the road section, and constructing the semantic map based on the feature sequences of all road sections.

7. The positioning method of a vehicle according to claim 1, wherein The current position of the vehicle is located by using the semantic map and the ground semantic feature object and the space semantic feature object of the current area where the vehicle is located, and the method comprises the following steps of: determining one or more candidate areas from the semantic map according to the ground semantic feature object and the space semantic feature object of the current area where the vehicle is located; performing feature matching on the current area and the candidate areas, taking the candidate area with the most matching points as a matching area, and determining the current position of the vehicle according to the position of the matching area on the semantic map.

8. The positioning method of a vehicle according to claim 1, wherein The panoramic image around the vehicle is obtained, and the method comprises the following steps of: simultaneously triggering a plurality of collection devices on the vehicle, and generating the panoramic image according to the images collected by the plurality of collection devices, wherein the installation heights of the plurality of collection devices are consistent.

9. A positioning device for a vehicle, characterized in that A positioning method for a vehicle as claimed in any one of claims 1-8, comprising: an acquisition module configured to acquire a panoramic image around the vehicle; an estimation module configured to perform semantic segmentation on the panoramic image to obtain ground semantic features and space semantic features, and to perform depth estimation on a space point to obtain depth estimation information according to a pose state of the vehicle and the panoramic image; a projection module configured to determine a three-dimensional projection plane according to the depth estimation information, project the space semantic features onto the three-dimensional projection plane to obtain space features, and project the ground semantic features onto a ground projection plane to obtain ground features; and a positioning module configured to track and fuse the ground features and the space features of a plurality of panoramic images to obtain a ground semantic feature object and a space semantic feature object, construct a semantic map based on the ground semantic feature object and the space semantic feature object, and locate a current position of the vehicle by using the semantic map and the ground semantic feature object and the space semantic feature object of a current area where the vehicle is located.

10. A vehicle characterized by comprising: comprise: - a memory, a processor, and a computer program stored on the memory and runable on the processor, the processor executing the program to implement the positioning method of the vehicle according to any one of claims 1-8.

11. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor for implementing the positioning method of the vehicle according to any one of claims 1-8.

Citation Information

Patent Citations

  • Automatic parking system and method

    CN111169468A

  • Mobile device positioning method, device and system, and mobile device

    WO2020135325A1