Parking positioning, map construction method and device, electronic equipment, and storage medium

By combining feature extraction and semantic matching of front and surround view images, and correcting pose using wheel speed and inertial data, the problem of large error and inaccurate positioning in visual SLAM localization methods is solved, achieving accurate parking positioning.

CN115760987BActive Publication Date: 2026-05-01CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING CHANGAN TECH CO LTD
Filing Date
2022-11-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing visual SLAM localization methods use monocular cameras for mapping, which results in large errors and cannot output pose in real time on the built map, leading to inaccurate autonomous parking localization.

Method used

By acquiring the target front view and surround view images of the vehicle, feature extraction and semantic matching are performed. Combined with a pre-set semantic map, the vehicle's pose is determined, and wheel speed and inertial data are used to correct pose changes, thereby achieving accurate parking positioning.

Benefits of technology

It achieves precise parking positioning at any location in a parking garage, overcoming the problem of positioning failure in the middle of the journey and improving the accuracy and robustness of parking positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760987B_ABST
    Figure CN115760987B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a parking positioning method, a map construction method and device, an electronic device and a storage medium. The parking positioning method comprises: obtaining a target front view image and a first target surround view image of a vehicle; performing feature extraction on the target front view image to obtain target front view image features; matching a target preset pose corresponding to the target front view image features according to the target front view image features; determining a first target local semantic map based on a similarity between the target preset pose and a plurality of surround view poses of a preset semantic map; performing semantic extraction on the first target surround view image to obtain first target semantic information; matching the first target local semantic map according to the first target semantic information; and determining a first time instant pose of the vehicle to position the vehicle. The parking positioning method can be initialized at any position in a parking garage scene, and can overcome the problem of positioning failure in the middle of a route, thereby ensuring the accuracy and robustness of parking positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automatic parking technology, specifically to a parking positioning and map building method, device, electronic device, and computer-readable storage medium. Background Technology

[0002] With the increasing popularity of private cars and the development of onboard sensors and processor hardware, more and more vehicles are equipped with some driver assistance functions. Among these, autonomous parking technology reduces the difficulty of parking for drivers, providing great convenience to people's lives. Currently, there are two main solutions for autonomous parking: one based on vision sensors and the other based on LiDAR. Due to the high cost of LiDAR hardware, fewer models currently install it; most vehicles use vision sensors, such as monocular cameras and fisheye cameras. These vision sensors are commonly used to perceive the vehicle's surroundings, detect and track targets, or perform simultaneous localization and mapping.

[0003] Existing visual SLAM (Simultaneous Localization and Mapping) methods primarily use monocular cameras for visual SLAM mapping, neglecting to utilize other common vehicle sensor information such as wheel speed. This results in significant mapping errors, and the presence of loops in the vehicle's trajectory also impacts mapping quality. Furthermore, while this method implements the SLAM mapping process, it does not perform localization within the constructed map. Therefore, when the vehicle re-enters the constructed map scene, it cannot output its pose in real time. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, embodiments of this application provide a parking positioning, map building method, apparatus, electronic device, and computer-readable storage medium to solve the above-mentioned technical problems.

[0005] This application provides a parking positioning method, which includes: acquiring a target front view image and a first target surround view image of a vehicle; extracting features from the target front view image to obtain target front view image features; matching a target preset pose corresponding to the target front view image features based on the target front view image features; determining a first target local semantic map based on the similarity between the target preset pose and multiple surround view poses of a preset semantic map, wherein the preset semantic map includes multiple local semantic maps, each local semantic map having a corresponding surround view pose pre-set; extracting semantics from the first target surround view image to obtain first target semantic information; matching the first target local semantic map based on the first target semantic information to determine the vehicle's first-moment pose, thereby performing parking positioning on the vehicle.

[0006] In one embodiment of this application, after determining the first moment pose of the vehicle, the parking positioning method includes: acquiring target wheel speed data, target inertial data, and a second target surround view image of the vehicle at a second acquisition moment, wherein the second acquisition moment is later than the first acquisition moment, and the first acquisition moment is the moment when the target front view image and the first target surround view image are acquired; determining the pose change amount based on the target wheel speed data and the target inertial data; determining the second moment initial pose of the vehicle based on the first moment pose and the pose change amount; determining a second target local semantic map based on the similarity between the second moment initial pose and multiple surround view poses of the preset semantic map; performing semantic extraction on the second target surround view image to obtain second target semantic information; and matching the second target local semantic map based on the second target semantic information to determine the second moment pose of the vehicle.

[0007] In one embodiment of this application, after determining the second-moment pose of the vehicle, the parking positioning method includes: acquiring new target wheel speed data, new target inertial data, and a new target surround view image of the vehicle at the next acquisition moment, wherein the next acquisition moment is later than the second acquisition moment; determining a new pose change based on the new target wheel speed data and the new target inertial data; determining the vehicle's initial pose at the next moment based on the previous pose and the new pose change; determining a new target local semantic map based on the similarity between the initial pose at the next moment and multiple surround view poses of the preset semantic map, wherein the previous pose includes the second-moment pose; performing semantic extraction on the new target surround view image to obtain new target semantic information; and matching the new target local semantic map based on the new target semantic information to determine the vehicle's next-moment pose.

[0008] In one embodiment of this application, matching a target preset pose corresponding to the target front view image features according to the target front view image features includes: acquiring multiple front view images and the front view image acquisition time of each front view image, wherein the acquisition time of each front view image is earlier than the acquisition time of the target front view image; extracting features from the front view images to obtain front view image features of the multiple front view images; determining the preset pose of the front view images based on the front view image acquisition time of the front view images, each surrounding view pose, and the surrounding view image acquisition time corresponding to each surrounding view pose, to obtain the preset poses of the multiple front view images; creating a front view image feature-preset pose database based on the front view image features of the multiple front view images and the preset poses of each front view image; and matching similar front view image features from the front view image feature-preset pose database according to the target front view image features, so as to determine the preset pose of the similar front view image features as the target preset pose corresponding to the target front view image features.

[0009] In one embodiment of this application, determining the preset pose of the front view image based on the front view image acquisition time, each surrounding view pose, and the surrounding view image acquisition time corresponding to each surrounding view pose includes at least one of the following: if the front view image acquisition time is different from the surrounding view image acquisition time, at least two surrounding view poses are determined as target surrounding view poses based on the front view image acquisition time and the surrounding view image acquisition time, and the preset pose of the front view image is determined through the target surrounding view poses; if the front view image acquisition time is the same as a surrounding view image acquisition time, the surrounding view pose corresponding to the surrounding view image acquisition time is determined as the preset pose of the front view image.

[0010] In one embodiment of this application, a map construction method is also provided. The map construction method includes: acquiring map construction reference data and image data, wherein the map construction reference data includes multiple wheel speed data and wheel speed acquisition time of each wheel speed data, multiple inertial data and inertial acquisition time of each inertial data, and the image data includes multiple surround view images and surround view image acquisition time of each surround view image; determining a reference time based on the wheel speed acquisition time of one wheel speed data or the inertial acquisition time of one inertial data, and determining a reference pose of the reference time based on the wheel speed data and the inertial data to obtain multiple reference times and reference poses of each reference time, wherein the wheel speed acquisition time is the same as the inertial acquisition time; determining the surround view initial pose of the surround view image based on the surround view image acquisition time of the surround view image, each reference time, and the reference pose of each reference time to obtain the surround view initial poses of multiple surround view images; and constructing a map based on the semantic information of the multiple surround view images and the surround view initial poses of each surround view image to obtain a preset semantic map.

[0011] In one embodiment of this application, a map is constructed based on the semantic information of multiple surround view images and the initial surround view poses of each surround view image to obtain a preset semantic map. This includes: extracting semantic information from the surround view images to obtain semantic information of multiple surround view images; matching and calculating the initial surround view pose of one surround view image with the semantic information of each surround view image to determine the surround view pose of the surround view image, thereby obtaining the surround view poses of multiple surround view images; and stitching together the semantic information of the multiple surround view images according to the surround view poses of each surround view image to obtain the preset semantic map. The semantic information includes lane line information and parking space information.

[0012] In one embodiment of this application, after obtaining the panoramic poses of multiple panoramic images, the map construction method includes: extracting features from the front view images to obtain front view image features of multiple front view images, wherein the image data further includes multiple front view images; determining a preset pose of the front view images based on the front view image acquisition time of the front view images, the panoramic poses of each panoramic image, and the acquisition time of each panoramic image, to obtain preset poses of multiple front view images, wherein the image data further includes the front view image acquisition time of each front view image; and creating a front view image feature-preset pose database based on the front view image features of multiple front view images and the preset poses of each front view image.

[0013] In one embodiment of this application, determining the preset pose of the front view image based on the front view image acquisition time, the surround view pose of each surround view image, and the surround view image acquisition time includes at least one of the following: if the front view image acquisition time is different from the surround view image acquisition time, at least two surround view poses are determined as target surround view poses based on the front view image acquisition time and the surround view image acquisition time, and the preset pose of the front view image is determined through the target surround view poses; if the front view image acquisition time is the same as the surround view image acquisition time, the surround view pose of the surround view image is determined as the preset pose of the front view image.

[0014] In one embodiment of this application, determining the initial panoramic pose of the panoramic image based on the panoramic image acquisition time, each reference time, and the reference pose at each reference time includes at least one of the following: if the panoramic image acquisition time is not the same as each reference time, at least two reference poses are determined as target reference poses based on the panoramic image acquisition time and each reference time, and the initial panoramic pose of the panoramic image is determined through the target reference poses; if the panoramic image acquisition time is the same as a reference time, the reference pose at the reference time is determined as the initial panoramic pose of the panoramic image.

[0015] In one embodiment of this application, a parking positioning device is also provided, comprising: an image acquisition module for acquiring a target front view image and a first target surround view image of a vehicle; a pose estimation module for extracting features from the target front view image to obtain target front view image features, and matching a target preset pose corresponding to the target front view image features based on the target front view image features; a local map determination module for determining a first target local semantic map based on the similarity between the target preset pose and multiple surround view poses of a preset semantic map, the preset semantic map including multiple local semantic maps, each local semantic map having a corresponding surround view pose pre-set; and a pose optimization module for extracting semantics from the first target surround view image to obtain first target semantic information, and matching the first target local semantic map based on the first target semantic information to determine the first moment pose of the vehicle for parking positioning.

[0016] In one embodiment of this application, a map building apparatus is also provided, comprising: a data acquisition module for acquiring map building reference data and image data, wherein the map building reference data includes multiple wheel speed data and the wheel speed acquisition time of each wheel speed data, multiple inertial data and the inertial acquisition time of each inertial data, and the image data includes multiple surround view images and the surround view image acquisition time of each surround view image; and a reference pose determination module for determining a reference time based on the wheel speed acquisition time of one wheel speed data or the inertial acquisition time of one inertial data, and based on the... Wheel speed data and inertial data determine the reference pose at the reference time to obtain multiple reference times and reference poses at each reference time, wherein the wheel speed acquisition time is the same as the inertial acquisition time; an initial pose determination module is used to determine the initial panoramic pose of the panoramic image based on the panoramic image acquisition time, each reference time, and the reference pose at each reference time to obtain the initial panoramic poses of multiple panoramic images; a map construction module is used to construct a map based on the semantic information of multiple panoramic images and the initial panoramic poses of each panoramic image to obtain a preset semantic map.

[0017] In one embodiment of this application, an electronic device is also provided, the electronic device comprising: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method described above.

[0018] In one embodiment of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a computer's processor, causes the computer to perform the method described above.

[0019] The beneficial effects of this application are as follows: The parking positioning method in this application matches the target's front view image with the target's preset pose, and determines the first target local semantic map from the preset semantic map based on the target's preset pose. The first target semantic information of the first target's surround view image is then matched with the first target local semantic map to obtain the vehicle's first-moment pose. This parking positioning method can be initialized at any location in the parking garage scenario, overcoming the problem of positioning failure in the middle of the journey, and ensuring the accuracy and robustness of parking positioning.

[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0022] Figure 1 This is a schematic diagram illustrating the implementation environment of the parking positioning method in an exemplary embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating a parking positioning method in an exemplary embodiment of this application;

[0024] Figure 3 This is a schematic diagram illustrating the implementation environment of the map construction method according to an exemplary embodiment of this application;

[0025] Figure 4 This is a flowchart illustrating a map construction method in an exemplary embodiment of this application;

[0026] Figure 5 This is a flowchart illustrating a mapping and positioning system according to an exemplary embodiment of this application;

[0027] Figure 6 This is a block diagram illustrating a parking positioning device in an exemplary embodiment of this application;

[0028] Figure 7 This is a block diagram illustrating a map building apparatus according to an exemplary embodiment of this application;

[0029] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0030] The embodiments of this application will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be understood that the preferred embodiments are only for illustrating this application and are not intended to limit the scope of protection of this application.

[0031] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0032] It should be noted that in this application, terms such as "first" and "second" are merely for distinguishing similar objects, and do not limit the order or sequence of similar objects. The variations of "including" and "having" indicate that the scope covered by the subject of the word is not exclusive, except for the examples shown by the word.

[0033] It is understood that the various numerical designations, step numbers, and other identifiers recorded in this application are for descriptive convenience and are not intended to limit the scope of this application. The size of the identifiers in this application does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.

[0034] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.

[0035] First, it's important to note that existing technologies offer various parking localization and mapping methods. While these methods can reduce accumulated errors in localization and mapping, they still have shortcomings. For example, some mapping methods primarily use monocular cameras for visual SLAM mapping, neglecting other common vehicle sensor information such as wheel speed. Furthermore, the trained model significantly impacts the estimated camera pose, and the presence of loops in the vehicle's trajectory also affects mapping quality. Moreover, they don't involve localization within the constructed map, meaning that when the car re-enters the map scene, it cannot output its pose in real-time. Some parking methods require complex calculations for image stitching and demand high precision in camera extrinsic parameter calibration. Additionally, real-time semantic segmentation for identifying drivable areas also involves substantial computation, significantly consuming processor resources.

[0036] To address these issues, embodiments of this application propose a parking positioning method, a map building method, a parking positioning device, a map building device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail below.

[0037] Please refer to Figure 1 , Figure 1 This is a schematic diagram illustrating the implementation environment of the parking positioning method as an exemplary embodiment of this application.

[0038] Reference Figure 1 As shown, the implementation environment may include an intelligent vehicle 101 and a computer device 102. The computer device 102 may be at least one of a microcomputer, an embedded computer, or a neural network computer. The computer device 102 is used to process the acquired front view image and the first surrounding view image of the target, and determine the pose at a first moment based on the processing results and a preset semantic map. The intelligent vehicle 101 is used to acquire the front view image and the first surrounding view image of the target through an image acquisition device and provide them to the computer device 102 for processing.

[0039] For example, after acquiring the target front view image and the first target surround view image of the vehicle, the computer device 102 performs feature extraction on the target front view image to obtain target front view image features. Based on the target front view image features, it matches the target preset pose corresponding to the target front view image features. Based on the similarity between the target preset pose and multiple surround view poses of the preset semantic map, it determines the first target local semantic map. It then performs semantic extraction on the first target surround view image to obtain first target semantic information. Based on the first target semantic information, it matches the first target local semantic map to determine the vehicle's first-moment pose for parking positioning. Therefore, the technical solution of this application embodiment can achieve initialization at any location in a parking garage scenario, overcoming the problem of positioning failure in the middle of the journey, and ensuring the accuracy and robustness of parking positioning.

[0040] It should be noted that the parking positioning method provided in this application embodiment is generally executed by computer device 102, and correspondingly, the parking positioning device is generally installed in computer device 102.

[0041] Please see Figure 2 , Figure 2 This is a flowchart illustrating a parking positioning method in an exemplary embodiment of this application. This method can be applied to... Figure 1 The implementation environment shown is specifically executed by computer device 102 within that implementation environment. It should be understood that this method can also be applied to other exemplary implementation environments and executed by devices in other implementation environments; this embodiment does not limit the implementation environment to which the method is applicable.

[0042] Reference Figure 2 As shown, in an exemplary embodiment, the parking positioning method includes at least steps S210 to S240, which are described in detail below:

[0043] Step S210: Obtain the target front view image and the first target surround view image of the vehicle.

[0044] In one embodiment of this application, when the vehicle begins localization or localization initialization, a target front view image and a first target surround view image of the vehicle are first acquired. Computer device 102 or other computer devices can acquire the target front view image and the first target surround view image of the vehicle through the image acquisition device of the connected intelligent vehicle, or the target front view image and the first target surround view image can be sent to computer device 102 or other computer devices from the cloud. The image acquisition device includes a forward-looking camera and a surround-view camera. The forward-looking camera includes at least one of a monocular camera, a binocular camera, a trinocular camera, etc., used to acquire the target front view image. The surround-view camera includes at least one of a wide-angle camera and a fisheye camera, etc., used to acquire the first target surround view image.

[0045] Step S220: Extract features from the target front view image to obtain target front view image features, and match the target preset pose corresponding to the target front view image features.

[0046] In one embodiment of this application, the target front view image can be input into a pre-trained convolutional neural network model for feature extraction to obtain the target front view image features. Then, the target preset pose corresponding to the target front view image features is matched. There are various methods for feature extraction. The target front view image features can also be obtained by calculating the HOG (Histogram of Oriented Gradient) feature descriptor of the target front view image, or by using other methods to extract features from the target front view image. No limitation is imposed here. In this embodiment, the pre-trained convolutional neural network model is trained using historical front view image data, which includes historical front view images and historical front view image features.

[0047] A new historical front view image is generated from the historical front view image through random projective transformation. One of the new historical front view image and the historical front view image is randomly selected as input information, and the HOG feature descriptor of the other image is calculated. The HOG feature descriptor is used as output information to train a convolutional neural network model for deep learning. The trained convolutional neural network model is then fed into computer device 102 or other computer devices. A 3648-dimensional vector can be used as the HOG feature descriptor, or other dimensional vectors defined by those skilled in the art; no limitation is imposed here.

[0048] In one embodiment of this application, matching a target preset pose corresponding to the target front view image features includes the following steps:

[0049] Step S221: Obtain multiple front view images and the front view image acquisition time of each front view image.

[0050] In one embodiment of this application, while the vehicle is driving in a parking lot and collecting surround view images required to construct a preset semantic map, it also collects front view images. The vehicle collects multiple front view images at different times using a front-view camera, and provides the multiple front view images and the acquisition times of each front view image to computer device 102 or other computer device. The acquisition time of each front view image is earlier than the acquisition time of the target front view image.

[0051] Step S222: Extract features from the front view image to obtain front view image features of multiple front view images.

[0052] In one embodiment of this application, the front view image can be input into a pre-trained convolutional neural network model for feature extraction to obtain front view image features of multiple front view images, or the HOG feature descriptor of the front view image can be calculated to obtain front view image features of multiple front view images, or other methods can be used to extract features from the front view image, which is not limited here.

[0053] Step S223: Based on the front view image acquisition time, each surrounding view pose, and the surrounding view image acquisition time corresponding to each surrounding view pose, determine the preset pose of the front view image to obtain the preset poses of multiple front view images.

[0054] In one embodiment of this application, the acquisition time of the surrounding view image is selected to be the same as or similar to the acquisition time of the front view image, and the preset pose of the front view image is determined according to the surrounding view pose of the surrounding view image at the acquisition time of the surrounding view image, thereby obtaining the preset poses of multiple front view images.

[0055] In one embodiment of this application, step S223 includes at least one of the following:

[0056] Step S2231: If the acquisition time of the front view image is different from the acquisition time of each surrounding view image, at least two surrounding view poses are determined as target surrounding view poses based on the acquisition time of the front view image and the acquisition time of each surrounding view image, and the preset pose of the front view image is determined through the target surrounding view poses.

[0057] In one embodiment of this application, if no surround view image acquisition time exists that is the same as the front view image acquisition time, at least two surround view image acquisition times that are close to or closest to the front view image acquisition time are selected. The surround view poses of these at least two surround view images are determined as the target surround view poses. The target surround view poses are then calculated using an interpolation algorithm to obtain the preset pose of the front view image. Alternatively, other algorithms can be used to calculate the target surround view poses to obtain the preset pose of the front view image; this is not a limitation.

[0058] Step S2232: If the acquisition time of the front view image is the same as the acquisition time of the ring view image, the ring view pose corresponding to the acquisition time of the ring view image is determined as the preset pose of the front view image.

[0059] In one embodiment of this application, if there is a surround view image acquisition time that is the same as the front view image acquisition time, then the surround view pose of the surround view image with the same acquisition time is determined as the preset pose of the front view image.

[0060] Step S224: Create a front view image feature-preset pose database based on the front view image features of multiple front view images and the preset poses of each front view image.

[0061] In one embodiment of this application, a correspondence is formed between the front view image features and the preset pose of the front view image to obtain multiple sets of corresponding front view image features-preset poses, and these multiple sets of corresponding front view image features-preset poses are stored in a database to obtain a front view image feature-preset pose database.

[0062] Step S225: Match similar front view image features from the front view image feature-preset pose database based on the target front view image features, so as to determine the preset pose of the similar front view image features as the target preset pose corresponding to the target front view image features.

[0063] In one embodiment of this application, the KD-Tree (K-Dimensional Tree) method is used to search (match) the front view image feature most similar to the target front view image feature from the front view image feature-preset pose database, indicating that the vehicle is near the location where the front view image corresponding to the most similar front view image feature was acquired. It should be understood that the KD-Tree is a data structure for fast nearest neighbor and near nearest neighbor search in high-dimensional space. Besides the KD-Tree method, other methods can be used to find similar front view image features; this is not limited here. The most similar front view image feature is taken as the similar front view image feature, and the preset pose corresponding to the similar front view image feature is extracted from the front view image feature-preset pose database. This preset pose is then determined as the target preset pose corresponding to the target front view image feature.

[0064] Step S230: Determine the first target local semantic map based on the similarity between the target preset pose and multiple look-around poses of the preset semantic map.

[0065] In one embodiment of this application, it should be understood that the target preset pose is the estimated pose of the vehicle at the first moment, which has a large error and needs further correction. The similarity between the target preset pose and multiple look-around poses of the preset semantic map is calculated, and the local semantic map corresponding to the look-around pose with the highest similarity is determined as the first target local semantic map. The preset semantic map includes multiple local semantic maps, each with a pre-set corresponding look-around pose.

[0066] Step S240: Semantic extraction is performed on the first target surround view image to obtain the first target semantic information. The first target semantic information is then matched with the local semantic map of the first target to determine the vehicle's first-moment pose for parking positioning.

[0067] In one embodiment of this application, the acquired first target surround view image is semantically segmented to obtain first target semantic information including lane line elements and parking space elements. The first target semantic information can be matched and calculated with the first target local semantic map using the ICP (Iterative Closest Point) algorithm. Specifically, the first target semantic information is processed to obtain a first point cloud, and the first target local semantic map is processed to obtain a first target local point cloud. Based on the positional correspondence between the first point cloud and the first target local point cloud, a coarse registration pose is determined. Based on the nearest point iteration algorithm, the coarse registration pose is finely registered to obtain a finely registered pose. The finely registered pose is used as the vehicle's first-moment pose, which is the first pose at the start of localization initialization or localization. When the vehicle starts moving, the first-moment pose is used as the initial value to iterate with other data to obtain the vehicle's real-time pose. If localization fails during vehicle movement, localization initialization is performed, i.e., localization can be restarted from step S210.

[0068] In one embodiment of this application, after step S240, the following steps are included:

[0069] Step S251: Obtain the target wheel speed data, target inertial data, and second target surround view image of the vehicle collected at the second acquisition time.

[0070] In one embodiment of this application, after the vehicle starts moving, target wheel speed data of the vehicle at a second acquisition time is acquired via wheel speed sensors, target inertial data of the vehicle at the second acquisition time is acquired via IMU (Inertial Measurement Unit) sensors, and a second target surround view image of the vehicle at the second acquisition time is acquired via a surround view camera. The second acquisition time is later than the first acquisition time, which is the time when the target front view image and the first target surround view image are acquired.

[0071] Step S252: Determine the pose change based on the target wheel speed data and target inertial data; determine the vehicle's initial pose at the second moment based on the pose at the first moment and the pose change; and determine the second target local semantic map based on the similarity between the initial pose at the second moment and multiple surround-view poses of the preset semantic map.

[0072] In one embodiment of this application, a translation vector is calculated from the target wheel speed data, and a rotation matrix is ​​calculated from the target inertial data. The pose change is composed of the translation vector and the rotation matrix. Based on the first-time pose and the pose change, the vehicle's second-time initial pose is calculated. It should be understood that the second-time initial pose is a predicted pose of the vehicle at the second time, which has a large error and needs further correction. The similarity between the second-time initial pose and multiple surround-view poses of a preset semantic map is calculated, and the local semantic map corresponding to the surround-view pose with the highest similarity is determined as the second target local semantic map.

[0073] Step S253: Semantic extraction is performed on the surround view image of the second target to obtain the semantic information of the second target. The local semantic map of the second target is matched according to the semantic information of the second target to determine the second time pose of the vehicle.

[0074] In one embodiment of this application, the acquired second target surround view image is semantically segmented to obtain second target semantic information including lane line elements and parking space elements. The second target semantic information is then matched with the second target local semantic map using the ICP algorithm to determine the vehicle's second-time pose. The second-time pose is the vehicle's precise pose at the second time. The matching calculation method is described in the embodiment of step S240 and will not be repeated here.

[0075] In one embodiment of this application, after step S253, the following steps are included:

[0076] Step S261: Obtain the new target wheel speed data, new target inertial data, and new target surround view image of the vehicle collected at the next acquisition time.

[0077] In one embodiment of this application, when the vehicle reaches the next time moment, the vehicle acquires new target wheel speed data for the next acquisition time via wheel speed sensors, acquires new target inertial data for the vehicle for the next acquisition time via IMU sensors, and acquires new target surround view images of the vehicle for the next acquisition time via surround view cameras. The next acquisition time is later than the aforementioned second acquisition time.

[0078] Step S262: Determine the new pose change based on the new target wheel speed data and the new target inertial data; determine the vehicle's initial pose at the next moment based on the pose at the previous moment and the new pose change; and determine the new target local semantic map based on the similarity between the initial pose at the next moment and multiple surround-view poses of the preset semantic map.

[0079] In one embodiment of this application, a new translation vector is calculated from the new target wheel speed data, and a new rotation matrix is ​​calculated from the new target inertial data. The new translation vector and the new rotation matrix constitute a new pose change. The vehicle's initial pose at the next moment is calculated based on the pose at the previous moment and the new pose change. It should be understood that the initial pose at the next moment is a predicted pose of the vehicle at the next moment, which has a large error and needs further correction. The pose at the previous moment includes the pose at the second moment mentioned above. The similarity between the initial pose at the next moment and multiple look-around poses of a preset semantic map is calculated, and the local semantic map corresponding to the look-around pose with the highest similarity is determined as the local semantic map of the new target.

[0080] Step S263: Semantic extraction is performed on the new target surround view image to obtain the new target semantic information. The local semantic map of the new target is matched based on the new target semantic information to determine the vehicle's pose at the next moment.

[0081] In one embodiment of this application, the acquired new target surround view image is semantically segmented to obtain new target semantic information including lane line elements and parking space elements. The ICP algorithm is then used to match and calculate the new target semantic information with the new target local semantic map to determine the vehicle's pose at the next moment. The pose at the next moment is the vehicle's precise pose at the next moment. The idea behind this matching calculation is described in the embodiment of step S240, and will not be repeated here.

[0082] Please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating the implementation environment of the map construction method according to an exemplary embodiment of this application.

[0083] Reference Figure 3 As shown, the implementation environment may include an intelligent vehicle 401 and a computer device 402. The computer device 402 may be at least one of a microcomputer, an embedded computer, or a neural network computer. The computer device 402 is used to process the acquired map construction reference data and image data, and to construct a map based on the processing results. The intelligent vehicle 401 is used to collect map construction reference data and image data through sensor devices and provide them to the computer device 402 for processing.

[0084] For example, after acquiring map construction reference data and image data, computer device 402 determines multiple reference times and reference poses for each reference time based on the map construction reference data. The map construction reference data includes multiple wheel speed data and the wheel speed acquisition time for each wheel speed data, multiple inertial data and the inertial acquisition time for each inertial data. Based on the image data, multiple reference times, and reference poses for the reference times, it determines the initial surround-view poses for multiple surround-view images. The image data includes multiple surround-view images and the surround-view image acquisition time for each surround-view image. Based on the semantic information of the multiple surround-view images and the initial surround-view poses for each surround-view image, map construction is performed to obtain a preset semantic map. It is evident that the technical solution of this application embodiment fully utilizes existing common vehicle-mounted sensor information, comprehensively obtaining the initial surround-view poses through wheel speed data, inertial data, and semantic information of surround-view images. This reduces mapping errors, thereby improving mapping quality. Furthermore, compared to stitching surround-view images for mapping, using the semantic information of surround-view images for mapping reduces the computational load on the processor.

[0085] It should be noted that the map building method provided in this application embodiment is generally executed by computer device 402, and correspondingly, the map building device is generally disposed in computer device 402.

[0086] Please see Figure 4 , Figure 4 This is a flowchart illustrating a map construction method in an exemplary embodiment of this application. This method can be applied to... Figure 3 The implementation environment shown is specifically executed by computer device 402 within that implementation environment. It should be understood that this method can also be applied to other exemplary implementation environments and executed by devices in other implementation environments; this embodiment does not limit the implementation environment to which the method is applicable.

[0087] Reference Figure 4 As shown, in an exemplary embodiment, the map construction method includes at least steps S510 to S540, which are described in detail below:

[0088] Step S510: Obtain map construction reference data and image data.

[0089] In one embodiment of this application, when the vehicle begins preparing for mapping, map building reference data and image data are acquired for the first time. Computer device 402 or other computer devices can acquire the map building reference data and image data through the sensor devices of the connected intelligent vehicle, or the map building reference data and image data can be sent to computer device 402 or other computer devices from the cloud. The map building reference data includes multiple wheel speed data and the wheel speed acquisition time for each wheel speed data, multiple inertial data and the inertial acquisition time for each inertial data, and the image data includes multiple surround view images and the surround view image acquisition time for each surround view image. The sensor devices include wheel speed sensors, IMU sensors, and surround view cameras, with the surround view camera including at least one of a wide-angle camera and a fisheye camera.

[0090] Step S520: Determine a reference time based on the wheel speed acquisition time of one wheel speed data or the inertial acquisition time of one inertial data, and determine the reference pose of the reference time based on the wheel speed data and the inertial data, so as to obtain multiple reference times and the reference pose of each reference time.

[0091] In one embodiment of this application, the wheel speed acquisition time or the inertial acquisition time is determined as the reference time, thereby obtaining multiple reference times, wherein the wheel speed acquisition time and the inertial acquisition time are consistent; and the reference pose change amount corresponding to the acquisition time is calculated based on the wheel speed data and inertial data with the same acquisition time, thereby obtaining the reference pose change amount of multiple reference times, wherein the reference pose change amount includes a reference translation vector and a reference rotation matrix, the reference translation vector is calculated based on the wheel speed data, and the reference rotation matrix is ​​calculated based on the inertial data. Starting with the reference pose of the initial reference time as 0, the reference pose change amount of the next reference time is superimposed with the reference pose of the initial reference time to obtain the reference pose of the next reference time; the reference pose change amount of the next reference time after that is superimposed with the reference pose of the next reference time to obtain the reference pose of the next reference time after that, and so on, to obtain the reference pose of multiple reference times.

[0092] Step S530: Based on the acquisition time of the panoramic image, each reference pose, and the reference time of each reference pose, determine the initial panoramic pose of the panoramic image to obtain the initial panoramic pose of multiple panoramic images.

[0093] In one embodiment of this application, a reference time that is the same as or similar to the acquisition time of the panoramic image is selected, and the initial panoramic pose of the panoramic image is determined based on the reference pose of the reference time, thereby obtaining the initial panoramic pose of multiple panoramic images.

[0094] In one embodiment of this application, step S530 includes at least one of the following:

[0095] Step S531: If the acquisition time of the panoramic image is different from each reference time, at least two reference poses are determined as target reference poses based on the acquisition time of the panoramic image and each reference time, and the panoramic initial pose of the panoramic image is determined through the target reference poses.

[0096] In one embodiment of this application, if no reference time exists that is the same as the acquisition time of the panoramic image, at least two reference times that are close to or closest to the acquisition time of the panoramic image are selected, and the reference poses of these at least two reference times are determined as the target reference poses. The initial panoramic pose of the panoramic image is then calculated using an interpolation algorithm. Alternatively, other algorithms can be used to calculate the target reference poses to obtain the initial panoramic pose of the panoramic image; no limitation is imposed here.

[0097] Step S532: If the acquisition time of the panoramic image is the same as a reference time, then the reference pose at the reference time is determined as the initial panoramic pose of the panoramic image.

[0098] In one embodiment of this application, if there is a reference time that is the same as the acquisition time of the panoramic image, the reference pose of the reference time is determined as the initial panoramic pose of the panoramic image.

[0099] Step S540: Based on the semantic information of multiple panoramic images and the initial panoramic pose of each panoramic image, a map is constructed to obtain a preset semantic map.

[0100] In one embodiment of this application, the surround view image is semantically segmented to obtain semantic information of multiple surround view images. The multiple semantic information are then stitched together according to the initial surround view pose of each surround view image to obtain a preset semantic map.

[0101] In one embodiment of this application, step S540 includes the following steps:

[0102] Step S541: Semantic extraction is performed on the panoramic view image to obtain the semantic information of multiple panoramic view images. The initial panoramic pose of one panoramic view image is matched and calculated with the semantic information of each panoramic view image to determine the panoramic pose of the panoramic view image, so as to obtain the panoramic pose of multiple panoramic view images.

[0103] In one embodiment of this application, semantic segmentation is performed on each surround view image to obtain semantic information of multiple surround view images. To make the constructed preset semantic map more accurate, the ICP algorithm can be used to use the initial surround view pose of one surround view image as an initial value, and match it with the semantic information of all surround view images acquired before the acquisition time of that surround view image to correct or optimize the initial surround view pose. The corrected or optimized pose is then used as the surround view pose of the surround view image, and the surround view poses of multiple surround view images are obtained in this way.

[0104] In one embodiment of this application, after step S541, the following steps are included:

[0105] Step S5411: Extract features from the front view image to obtain front view image features of multiple front view images.

[0106] In one embodiment of this application, the acquired image data also includes multiple front view images. The front view images can be input into a pre-trained convolutional neural network model for feature extraction to obtain front view image features of multiple front view images, or the HOG feature descriptor of the front view images can be calculated to obtain front view image features of multiple front view images, or other methods can be used to extract features from the front view images. No limitation is imposed here.

[0107] Step S5412: Based on the acquisition time of the front view image, the panoramic pose of each panoramic image, and the acquisition time of each panoramic image, determine the preset pose of the front view image to obtain the preset pose of multiple front view images.

[0108] In one embodiment of this application, the image data further includes the front view image acquisition time of each front view image, filtering the surrounding view image acquisition time that is the same as or similar to the front view image acquisition time, and determining the preset pose of the front view image based on the surrounding view pose of the surrounding view image at the surrounding view image acquisition time, thereby obtaining the preset poses of multiple front view images.

[0109] In one embodiment of this application, step S5412 includes at least one of the following:

[0110] Step S54121: If the acquisition time of the front view image is different from the acquisition time of each surrounding view image, at least two surrounding view poses are determined as target surrounding view poses based on the acquisition time of the front view image and the acquisition time of each surrounding view image, and the preset pose of the front view image is determined through the target surrounding view poses.

[0111] In one embodiment of this application, if no surround view image acquisition time exists that is the same as the front view image acquisition time, at least two surround view image acquisition times that are close to or closest to the front view image acquisition time are selected. The surround view poses of these at least two surround view images are determined as the target surround view poses. The target surround view poses are then calculated using an interpolation algorithm to obtain the preset pose of the front view image. Alternatively, other algorithms can be used to calculate the target surround view poses to obtain the preset pose of the front view image; this is not a limitation.

[0112] Step S54122: If the acquisition time of the front view image is the same as the acquisition time of the ring view image, the ring view pose is determined as the preset pose of the front view image.

[0113] In one embodiment of this application, if there is a surround view image acquisition time that is the same as the front view image acquisition time, then the surround view pose of the surround view image with the same acquisition time is determined as the preset pose of the front view image.

[0114] Step S5413: Create a front view image feature-preset pose database based on the front view image features of multiple front view images and the preset poses of each front view image.

[0115] In one embodiment of this application, a correspondence is formed between the front view image features and the preset pose of the front view image to obtain multiple sets of corresponding front view image features-preset poses, and these multiple sets of corresponding front view image features-preset poses are stored in a database to obtain a front view image feature-preset pose database.

[0116] Step S542: Based on the panoramic pose of each panoramic image, the semantic information of multiple panoramic images is stitched together to obtain a preset semantic map.

[0117] In one embodiment of this application, the semantic information of multiple surround view images is stitched together based on the surround view pose of each surround view image to obtain a preset semantic map. The semantic information includes lane line information and parking space information, which serve as local semantic maps within the preset semantic map. Therefore, the preset semantic map includes multiple local semantic maps and the surround view pose corresponding to each local semantic map.

[0118] Please refer to Figure 5 , Figure 5 This is a flowchart illustrating a mapping and positioning system according to an exemplary embodiment of this application.

[0119] Reference Figure 5 As shown, the mapping process is as follows: The vehicle pose (reference pose) is calculated using wheel speed and IMU input values ​​(wheel speed data and inertial data); semantic segmentation of the surround-view camera input values ​​(surround-view images) outputs semantic information including lane lines and parking spaces; the precise pose (surround-view pose) is calculated by matching the vehicle pose with the semantic information, and a semantic map (preset semantic map) is created based on the precise pose and semantic information. Furthermore, the front-view camera input values ​​(front-view images) are transformed using random projective transformation to generate corresponding new images (new front-view images); a HOG descriptor (front-view image feature) is calculated for any image, and this HOG descriptor is used as output information. Another image is used as input information and fed into a convolutional neural network model for training, enabling the convolutional neural network model to learn the geometric features in the parking scene; the precise pose is interpolated to obtain the pose at the time of front-view image acquisition (preset pose of the front-view image); the front-view image features of the front-view image (or the new front-view image) and the preset pose of the front-view image are correlated and written into a database to obtain the front-view image feature-preset pose database.

[0120] Reference Figure 5 As shown, the localization process is as follows: The input value from the forward-looking camera at the first moment of localization (target forward-looking image) is input into the learned convolutional neural network model to obtain a 3648-dimensional vector (target forward-looking image features). A KD-Tree search is used to find the most similar image (similar forward-looking image features) in the database (forward-looking image features - preset pose database), and the pose corresponding to that image (preset pose) is extracted as the initial value for localization (target preset pose). The semantic segmentation of the input value from the surround-view camera at the first moment of localization (first target surround-view image) outputs the first target semantic information, including lane lines and parking spaces. Based on the target preset pose, the first target local semantic map is determined from the preset semantic map, and the first... The target semantic information is matched with the first target local semantic map to output the pose at the first moment. The vehicle pose (pose change) is calculated by locating the wheel speed at the second moment and the IMU input values ​​(target wheel speed data and target inertial data). Semantic segmentation is then performed using the surround-view camera input values ​​at the second moment (second target surround-view image) to output the second target semantic information, including lane lines and parking spaces. The initial pose at the second moment is calculated based on the first moment pose and pose change. The second target local semantic map is determined from the preset semantic map based on the initial pose at the second moment. The second target semantic information is matched with the second target local semantic map to output the pose at the second moment. This process is repeated to obtain the pose at the next moment. If localization fails, initialization can be performed again to output the pose once more.

[0121] In one specific embodiment of this application, when creating a front view image feature-preset pose database, DeepLCD (Deep Loop Closure Detection) can be used to encode the front view image to obtain front view image features. A correspondence is established between the front view image features and the preset poses of the front view image, thereby obtaining multiple sets of corresponding front view image features-preset poses, which are then stored in the DeepLCD database. When matching a target preset pose, DeepLCD is used again to encode the acquired target front view image to obtain the target front view image features. Based on the target front view image features, the front view image features most similar to the target front view image features are matched from the DeepLCD database to extract the preset pose corresponding to the front view image features. This preset pose is then determined as the target preset pose corresponding to the target front view image features.

[0122] Please refer to Figure 6 , Figure 6 This is a block diagram illustrating a parking positioning device according to an exemplary embodiment of this application. The device can be applied to… Figure 1 The implementation environment shown is specifically configured in computer device 102. This device can also be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the device is applicable.

[0123] Reference Figure 6 As shown, the exemplary parking positioning device includes:

[0124] Image acquisition module 710 is used to acquire a target front view image and a first target surround view image of the vehicle; pose estimation module 720 is used to extract features from the target front view image to obtain target front view image features, and match the target preset pose corresponding to the target front view image features; local map determination module 730 is used to determine a first target local semantic map based on the similarity between the target preset pose and multiple surround view poses of a preset semantic map, the preset semantic map including multiple local semantic maps, each local semantic map having a corresponding surround view pose pre-set; pose optimization module 740 is used to extract semantics from the first target surround view image to obtain first target semantic information, and match the first target local semantic map based on the first target semantic information to determine the vehicle's first-moment pose for parking positioning.

[0125] It should be noted that the parking positioning device and the parking positioning method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the parking positioning device provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0126] Please refer to Figure 7 , Figure 7 This is a block diagram illustrating a map building apparatus according to an exemplary embodiment of this application. The apparatus can be applied to… Figure 3 The implementation environment shown is specifically configured in computer device 402. This device can also be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the device is applicable.

[0127] Reference Figure 7 As shown, the exemplary map building apparatus includes:

[0128] The data acquisition module 810 is used to acquire map construction reference data and image data. The map construction reference data includes multiple wheel speed data and the wheel speed acquisition time of each wheel speed data, multiple inertial data and the inertial acquisition time of each inertial data, and the image data includes multiple panoramic images and the panoramic image acquisition time of each panoramic image. The reference pose determination module 820 is used to determine a reference time based on the wheel speed acquisition time of one wheel speed data or the inertial acquisition time of one inertial data, and to determine the reference pose of the reference time based on the wheel speed data and the inertial data, so as to obtain multiple reference times and the reference pose of each reference time. The wheel speed acquisition time is the same as the inertial acquisition time. The initial pose determination module 830 is used to determine the panoramic initial pose of the panoramic image based on the panoramic image acquisition time of the panoramic image, each reference time, and the reference pose of each reference time, so as to obtain the panoramic initial pose of multiple panoramic images. The map construction module 840 is used to construct a map based on the semantic information of multiple panoramic images and the panoramic initial pose of each panoramic image to obtain a preset semantic map.

[0129] It should be noted that the map building apparatus and the map building method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit performs operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the map building apparatus provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0130] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the methods provided in the above embodiments.

[0131] Please refer to Figure 8 , Figure 8 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 8 The computer system 900 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0132] Reference Figure 8As shown, the computer system 900 includes a Central Processing Unit (CPU) 901, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 902 or programs loaded from storage portion 908 into Random Access Memory (RAM) 903, such as performing the methods described in the above embodiments. The RAM 903 also stores various programs and data required for system operation. The CPU 901, ROM 902, and RAM 903 are interconnected via a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0133] The following components are connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 910 as needed so that computer programs read from them can be installed into storage section 908 as needed.

[0134] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by central processing unit (CPU) 901, it performs various functions defined in the system of this application.

[0135] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0137] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0138] Another aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the method described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0139] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.

[0140] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A parking positioning method, characterized in that, The parking positioning method includes: Acquire the target front view image and the first target surround view image of the vehicle; Feature extraction is performed on the target front view image to obtain target front view image features. A target preset pose corresponding to the target front view image features is then matched. The matching of the target preset pose corresponding to the target front view image features includes: creating a front view image feature-preset pose database based on the front view image features of multiple front view images and the preset poses of each front view image; and matching similar front view image features from the front view image feature-preset pose database based on the target front view image features, so as to determine the preset pose of the similar front view image features as the target preset pose. A first target local semantic map is determined based on the similarity between the target preset pose and multiple look-around poses of the preset semantic map. The preset semantic map includes multiple local semantic maps, each with a corresponding look-around pose pre-set. The construction method of the preset semantic map includes: determining a reference time based on the wheel speed acquisition time of a wheel speed data or the inertial acquisition time of an inertial data, and determining a reference pose of the reference time based on the wheel speed data and the inertial data; determining the look-around pose of the look-around image based on the look-around image acquisition time of the look-around image, each reference time, and the reference pose of each reference time; and stitching together the semantic information of multiple look-around images based on the look-around poses of each look-around image to obtain the preset semantic map. Semantic extraction is performed on the first target surround view image to obtain the first target semantic information. The first target semantic information is then matched with the local semantic map of the first target to determine the first moment pose of the vehicle, so as to perform parking positioning of the vehicle.

2. The parking positioning method according to claim 1, characterized in that, After determining the vehicle's initial pose, the parking positioning method includes: The target wheel speed data, target inertial data, and second target surround view image of the vehicle are acquired at a second acquisition time. The second acquisition time is later than the first acquisition time. The first acquisition time is the time when the target front view image and the first target surround view image are acquired. The pose change is determined based on the target wheel speed data and the target inertial data. The initial pose of the vehicle at the second moment is determined based on the pose at the first moment and the pose change. The second target local semantic map is determined based on the similarity between the initial pose at the second moment and multiple surround-view poses of the preset semantic map. Semantic extraction is performed on the second target surround view image to obtain the second target semantic information. The second target semantic information is then matched with the second target local semantic map to determine the vehicle's second time pose.

3. The parking positioning method according to claim 2, characterized in that, After determining the second-moment pose of the vehicle, the parking positioning method includes: Acquire the new target wheel speed data, new target inertial data, and new target surround view image of the vehicle at the next acquisition time, wherein the next acquisition time is later than the second acquisition time; The new pose change is determined based on the new target wheel speed data and the new target inertial data. The vehicle's next initial pose is determined based on the previous pose and the new pose change. The new target local semantic map is determined based on the similarity between the next initial pose and multiple surround poses of the preset semantic map. The previous pose includes the second pose. Semantic extraction is performed on the new target surround view image to obtain new target semantic information. The local semantic map of the new target is then matched based on the new target semantic information to determine the vehicle's pose at the next moment.

4. The parking positioning method according to claim 1, characterized in that, The method further includes: Acquire multiple front view images and the acquisition time of each front view image, with the acquisition time of each front view image being earlier than the acquisition time of the target front view image; Feature extraction is performed on the front view image to obtain front view image features of multiple front view images; Based on the front view image acquisition time, each surrounding view pose, and the surrounding view image acquisition time corresponding to each surrounding view pose, the preset pose of the front view image is determined to obtain the preset poses of multiple front view images.

5. The parking positioning method according to claim 4, characterized in that, Based on the front view image acquisition time, each surrounding view pose, and the surrounding view image acquisition time corresponding to each surrounding view pose, the preset pose of the front view image is determined, including at least one of the following: If the acquisition time of the front view image is different from the acquisition time of each surrounding view image, based on the acquisition time of the front view image and the acquisition time of each surrounding view image, at least two surrounding view poses are determined as target surrounding view poses, and the preset pose of the front view image is determined through the target surrounding view poses. If the acquisition time of the front view image is the same as the acquisition time of the ring view image, the ring view pose corresponding to the acquisition time of the ring view image is determined as the preset pose of the front view image.

6. A map construction method, characterized in that, The map construction method includes: The map construction reference data and image data are obtained. The map construction reference data includes multiple wheel speed data and the wheel speed acquisition time of each wheel speed data, multiple inertial data and the inertial acquisition time of each inertial data, and the image data includes multiple panoramic images and the panoramic image acquisition time of each panoramic image. A reference time is determined based on the wheel speed acquisition time of one wheel speed data or the inertial acquisition time of one inertial data, and a reference pose of the reference time is determined based on the wheel speed data and the inertial data to obtain multiple reference times and reference poses of each reference time. The wheel speed acquisition time is the same as the inertial acquisition time. Based on the acquisition time of the panoramic image, each reference time, and the reference pose at each reference time, the panoramic initial pose of the panoramic image is determined to obtain the panoramic initial pose of multiple panoramic images. A map is constructed based on the semantic information of multiple panoramic images and the initial panoramic pose of each panoramic image to obtain a preset semantic map. The map construction includes matching and calculating the initial panoramic pose of a panoramic image with the semantic information of each panoramic image to determine the panoramic pose of the panoramic image, so as to obtain the panoramic poses of multiple panoramic images; and stitching together the semantic information of the multiple panoramic images according to the panoramic poses of each panoramic image to obtain the preset semantic map.

7. The map construction method according to claim 6, characterized in that, The method further includes: Semantic extraction is performed on the surround view images to obtain semantic information of multiple surround view images, including lane line information and parking space information.

8. The map construction method according to claim 7, characterized in that, After obtaining the panoramic poses of multiple panoramic images, the map construction method includes: Feature extraction is performed on the front view image to obtain front view image features of multiple front view images; the image data also includes multiple front view images. Based on the front view image acquisition time, the panoramic pose of each panoramic image, and the acquisition time of each panoramic image, the preset pose of the front view image is determined to obtain the preset pose of multiple front view images. The image data also includes the front view image acquisition time of each front view image. A front view image feature-preset pose database is created based on the front view image features of multiple front view images and the preset poses of each front view image.

9. The map construction method according to claim 8, characterized in that, Based on the acquisition time of the front view image, the panoramic pose of each panoramic image, and the acquisition time of each panoramic image, the preset pose of the front view image is determined, including at least one of the following: If the acquisition time of the front view image is different from the acquisition time of each surrounding view image, based on the acquisition time of the front view image and the acquisition time of each surrounding view image, at least two surrounding view poses are determined as target surrounding view poses, and the preset pose of the front view image is determined through the target surrounding view poses. If the acquisition time of the front view image is the same as the acquisition time of the ring view image, the ring view pose of the ring view image is determined as the preset pose of the front view image.

10. The map construction method according to claim 6, characterized in that, Based on the panoramic image acquisition time, reference times, and reference poses at each reference time, the initial panoramic pose of the panoramic image is determined, including at least one of the following: If the acquisition time of the panoramic image is different from each reference time, at least two reference poses are determined as target reference poses based on the acquisition time of the panoramic image and each reference time, and the panoramic initial pose of the panoramic image is determined through the target reference poses. If the acquisition time of the panoramic image is the same as a reference time, then the reference pose at the reference time is determined as the initial panoramic pose of the panoramic image.

11. A parking positioning device, characterized in that, The parking positioning device includes: The image acquisition module is used to acquire the target front view image and the first target surround view image of the vehicle; The pose prediction module is used to extract features from the target front view image to obtain target front view image features, and to match the target preset pose corresponding to the target front view image features. The matching of the target preset pose corresponding to the target front view image features includes: creating a front view image feature-preset pose database based on the front view image features of multiple front view images and the preset poses of each front view image; and matching similar front view image features from the front view image feature-preset pose database based on the target front view image features, so as to determine the preset pose of the similar front view image features as the target preset pose. A local map determination module is used to determine a first target local semantic map based on the similarity between the target preset pose and multiple look-around poses of a preset semantic map. The preset semantic map includes multiple local semantic maps, each with a corresponding look-around pose pre-set. The construction method of the preset semantic map includes: determining a reference time based on the wheel speed acquisition time of wheel speed data or the inertial acquisition time of inertial data, and determining a reference pose of the reference time based on the wheel speed data and the inertial data; determining the look-around pose of the look-around image based on the look-around image acquisition time of the look-around image, each reference time, and the reference pose of each reference time; and stitching together the semantic information of multiple look-around images based on the look-around poses of each look-around image to obtain the preset semantic map. The pose optimization module is used to extract semantic information from the first target surround view image to obtain the first target semantic information, match the first target local semantic map according to the first target semantic information, and determine the first moment pose of the vehicle to perform parking positioning of the vehicle.

12. A map building device, characterized in that, The map building device includes: The data acquisition module is used to acquire map construction reference data and image data. The map construction reference data includes multiple wheel speed data and the wheel speed acquisition time of each wheel speed data, multiple inertial data and the inertial acquisition time of each inertial data, and the image data includes multiple panoramic images and the panoramic image acquisition time of each panoramic image. The reference pose determination module is used to determine a reference time based on the wheel speed acquisition time of a wheel speed data or the inertial acquisition time of an inertial data, and to determine the reference pose of the reference time based on the wheel speed data and the inertial data, so as to obtain multiple reference times and the reference pose of each reference time, wherein the wheel speed acquisition time is the same as the inertial acquisition time. The initial pose determination module is used to determine the initial pose of the panoramic view image based on the panoramic view image acquisition time, each reference time, and the reference pose of each reference time, so as to obtain the initial pose of multiple panoramic view images. A map construction module is used to construct a map based on the semantic information of multiple surround view images and the initial surround view poses of each surround view image to obtain a preset semantic map. The map construction includes matching and calculating the initial surround view pose of one surround view image with the semantic information of each surround view image to determine the surround view pose of the surround view image, so as to obtain the surround view poses of multiple surround view images; and stitching together the semantic information of the multiple surround view images according to the surround view poses of each surround view image to obtain the preset semantic map.

13. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Parking garage mapping and positioning method and system and vehicle

    CN114693787A

  • Semantic map building and positioning method and device for indoor parking lot

    CN114863096A