Visual semantics-based mapping method, device, storage medium, and electronic device

By performing bird's-eye view stitching and texture mapping on the multi-channel surround view images of the panoramic surround view imaging system, the problem of feature matching errors caused by blind spots under the vehicle in the panoramic surround view imaging system is solved, the continuity and accuracy of semantic features are achieved, the success rate of mapping and user experience are improved, and the cost is low.

CN117351161BActive Publication Date: 2025-09-09BEIJING YINWO AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311245393.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-09-09
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

Existing panoramic surround view imaging systems have blind spots under the vehicle, which increases the error rate of feature matching, increases mapping errors, and even causes failure.

Method used

By acquiring multiple surround view images, a bird's-eye view stitching is performed to generate the first key frame image. Multiple historical frame images are used for texture mapping to eliminate the blind spots under the vehicle and generate a second key frame image without blind spots. The image is then converted into a local semantic map based on the initial pose, and finally a global map is generated.

Benefits of technology

It eliminates the effects of feature occlusion and truncation caused by visual blind spots, ensures the continuity and accuracy of semantic features, improves the success rate of mapping and user experience, and does not require additional camera modules, making it low-cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351161B_ABST
    Figure CN117351161B_ABST
Patent Text Reader

Abstract

The present application provides a mapping method, device, storage medium and electronic device based on visual semantics, which relates to the field of intelligent driving. The method includes: obtaining a multi-way surround view image collected by the target vehicle in the target environment at the current moment; performing a bird's-eye view stitching on the multi-way surround view image to generate a first key frame image; determining a plurality of historical frame images without a blind spot under the vehicle corresponding to the first key frame image; based on the plurality of historical frame images, performing texture mapping on the blind spot under the vehicle in the first key frame image to obtain a second key frame image; based on the initial position of the target vehicle, converting the second key frame image into a local semantic map; stitching the local semantic maps of the target vehicle at multiple moments in the target environment to generate a global map of the target environment. The solution in the present application can eliminate the feature occlusion and truncation effects caused by the visual blind spot, and ensure the continuity and accuracy of the semantic features of the second key frame image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and specifically to a visual semantics-based mapping method, device, storage medium, and electronic device. Background Art

[0002] With the development of intelligent driving technology, vehicles are generally equipped with panoramic surround-view imaging systems, allowing drivers to better understand the surrounding environment of the vehicle. However, the current panoramic surround-view imaging systems still have a visual blind spot under the vehicle.

[0003] In related applications, deep neural networks can be used to segment static road sign elements in panoramic surround image mosaics, using the segmentation results as features for inter-image matching. However, when panoramic surround image mosaics contain blind spots under vehicles, relevant features can be truncated or completely obscured, increasing the error rate in the feature matching process, leading to greater errors in map creation or even failure. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a visual semantics-based mapping method, device, storage medium, and electronic device.

[0005] In a first aspect, an embodiment of the present application provides a visual semantics-based mapping method, comprising: obtaining a multi-way surround view image collected by a target vehicle in a target environment at the current moment; performing bird's-eye view stitching on the multi-way surround view image to generate a first key frame image, wherein the first key frame image includes a blind spot under the vehicle; determining a plurality of historical frame images without a blind spot under the vehicle corresponding to the first key frame image; performing texture mapping on the blind spot under the vehicle in the first key frame image based on the plurality of historical frame images to obtain a second key frame image without a blind spot under the vehicle; converting the second key frame image into a local semantic map based on the initial posture of the target vehicle; and stitching the local semantic maps of the target vehicle at multiple moments in the target environment to generate a global map of the target environment.

[0006] In combination with the first aspect, in certain implementations of the first aspect, texture mapping is performed on the blind spot under the vehicle in the first key frame image based on multiple historical frame images to obtain a second key frame image without the blind spot under the vehicle, including: determining a reference frame image for mapping the blind spot under the vehicle in the first key frame image from the multiple historical frame images; determining the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value of the target vehicle at the historical moment corresponding to the reference frame image; determining the rotation angle and offset of the target vehicle in the first key frame image and the reference frame image based on the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value at the historical moment; determining a valid mapping area in the reference frame image based on the rotation angle and the offset; and texture mapping is performed on the blind spot under the vehicle in the first key frame image based on the valid mapping area to obtain a second key frame image.

[0007] In combination with the first aspect, in certain implementations of the first aspect, a reference frame image for mapping the blind spot under the vehicle in the first key frame image is determined from multiple historical frame images, including: based on the driving parameters of the target vehicle, respectively determining the transformation matrices of the multiple historical frame images and the first key frame image; based on the target sampling error, taking the feature information in the first key frame image as the standard, randomly sampling the feature information corresponding to each of the multiple historical frame images, and determining the sampling values ​​corresponding to each of the multiple historical frame images, wherein the sampling values ​​represent the degree of matching between the feature information in the historical frame images and the feature information in the first key frame image; and determining the reference frame image from the multiple historical frame images based on the transformation matrices of the multiple historical frame images and the first key frame image and the sampling values ​​corresponding to each of the multiple historical frame images.

[0008] In combination with the first aspect, in certain implementations of the first aspect, texture mapping is performed on the blind spot under the vehicle in the first key frame image based on the effective mapping area to obtain a second key frame image, including: determining a contour line of the blind spot under the vehicle; expanding the target distance outward along the contour line in a direction away from the blind spot under the vehicle to obtain an outer edge line, and using the area formed by the contour line and the outer edge line as a transition area; texture mapping is performed on the blind spot under the vehicle in the first key frame image based on the effective mapping area to obtain an image to be calibrated; determining a grayscale gradient value of the effective mapping area in the image to be calibrated; and calibrating the grayscale value of the transition area in the image to be calibrated based on the grayscale gradient value to obtain the second key frame image.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the second key frame image is converted into a local semantic map based on the initial posture of the target vehicle, including: performing semantic feature extraction on the second key frame image to obtain semantic features of the target object contained in the second key frame image, the semantic features including position features, direction features and category features; determining the posture of the target vehicle at the current moment; and converting the second key frame image into a local semantic map based on the semantic features of the target object contained in the second key frame image, the posture of the target vehicle at the current moment and the initial posture.

[0010] In combination with the first aspect, in certain implementations of the first aspect, local semantic maps of the target vehicle in the target environment at multiple times are spliced ​​to generate a global map of the target environment, including: determining at least two local semantic maps to be matched in the local semantic maps at multiple times; determining a target point cloud and a source point cloud in the at least two local semantic maps to be matched respectively; determining the nearest neighbor of each point in the target point cloud to the source point cloud based on the geometric features of the target point cloud; and splicing the local semantic maps of the target vehicle in the target environment at multiple times based on the nearest neighbor of each point in the target point cloud to the source point cloud to generate a global map of the target environment.

[0011] In combination with the first aspect, in certain implementations of the first aspect, the visual semantics-based mapping method further includes: matching features in the global map based on the initial pose of the target vehicle to determine the current pose of the target vehicle.

[0012] In the second aspect, an embodiment of the present application provides a mapping device based on visual semantics, including: an acquisition module for acquiring a multi-way surround view image collected by a target vehicle in a target environment at the current moment; a first splicing module for performing a bird's-eye view splicing of the multi-way surround view images to generate a first key frame image, wherein the first key frame image includes a blind spot under the vehicle; a determination module for determining a plurality of historical frame images without a blind spot under the vehicle corresponding to the first key frame image; a mapping module for performing texture mapping on the blind spot under the vehicle in the first key frame image based on a plurality of historical frame images to obtain a second key frame image without a blind spot under the vehicle; a conversion module for converting the second key frame image into a local semantic map based on the initial posture of the target vehicle; a second splicing module for splicing the local semantic maps of the target vehicle at multiple moments in the target environment to generate a global map of the target environment.

[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program for executing the method described in the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; and the processor for executing the method described in the first aspect.

[0015] In this embodiment, motion compensation is used to map the historical frame image to the blind spot under the vehicle of the first key frame image according to specific transformation conditions. This eliminates the effects of feature occlusion and truncation caused by the visual blind spot, ensures the continuity and accuracy of the semantic features of the second key frame image, and avoids feature truncation or loss during the mapping process, providing driving assistance and improving the user experience. In addition, this application does not require the installation of an additional camera module, which is low-cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other purposes, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 Shown is a flowchart of a map construction process provided by an exemplary embodiment of the present application.

[0018] Figure 2 FIG2 is a flow chart of obtaining a second key frame image provided by an exemplary embodiment of the present application.

[0019] Figure 3 Shown is a flowchart of converting to a local semantic map provided by an exemplary embodiment of the present application.

[0020] Figure 4 Shown is a schematic diagram of a process for generating a global map provided by an exemplary embodiment of the present application.

[0021] Figure 5 The figure shows a schematic diagram of the structure of a mapping device based on a semantic map provided in one embodiment of the present application.

[0022] Figure 6 Shown is a structural schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] Traditional SLAM (Simultaneous Localization and Mapping) technology only contains low-level information and is unable to meet the demands of modern computer vision. With the rise of artificial intelligence, neural network technology has surpassed traditional image processing in image classification, detection, and segmentation, and has initially demonstrated significant advantages in industries such as autonomous driving, robotics, drones, and healthcare. Compared to traditional vSLAM (visual Simultaneous Localization and Mapping), semantic vSLAM not only captures the geometric structure of the environment but also extracts semantic information about individual objects. In mapping, semantic information provides rich object information for constructing different types of semantic maps, such as pixel-level and object-level maps. In localization, semantic vSLAM leverages semantic constraints to improve positioning accuracy and robustness. Therefore, semantic vSLAM can help robots improve their ability to accurately perceive and adapt to unknown and complex environments, enabling them to perform more complex tasks. However, when blind spots exist under the vehicle in the local semantic map, relevant features can be truncated or completely obscured, increasing the error rate in feature matching and leading to increased mapping errors or even failure.

[0025] In view of this, this application provides a visual semantics-based mapping method. Motion compensation is performed based on the acquired first keyframe image and vehicle body motion information, thereby filling semantic features into the blind spot under the vehicle in the first keyframe image. Ultimately, the semantic information is more comprehensive and robust, and the feature matching success rate is higher.

[0026] Figure 1 The figure shows a flowchart of a map-building process provided by an exemplary embodiment of the present application. Figure 1 As shown, in an embodiment of the present application, the visual semantics-based mapping method includes the following steps.

[0027] Step S110 , obtaining a multi-way surround view image collected by the target vehicle in the target environment at the current moment.

[0028] Exemplarily, a panoramic surround view system of the target vehicle is used to generate a panoramic surround view overhead view. For example, video streams captured by multiple cameras installed around the target vehicle are input into the panoramic surround view system for preprocessing to generate a multi-channel surround view image. Alternatively, multiple images captured by cameras a, b, c, and d installed around the target vehicle are used as the multi-channel surround view image. The multiple images can be two-dimensional or three-dimensional, and this embodiment of the present application is not limited to this.

[0029] Step S120 : performing bird's-eye view stitching on the multiple surround view images to generate a first key frame image.

[0030] The first key frame image includes the blind spot under the vehicle. For example, the cameras in different directions and positions of the target vehicle are first subjected to bird's-eye view changes, and then spliced ​​together to obtain the first key frame image of the target vehicle and the surrounding area from a bird's-eye view. More specifically, the camera in the target vehicle is generally installed at an oblique downward angle. Therefore, the original surround view image output by the camera is not a top-down view. In order to achieve the effect of a bird's-eye view, the surround view images output by different cameras need to be projected to a new top-down perspective. For example, a matrix can be used to represent the change of a point in the original surround view image output by the camera. Specifically, its expression is as follows. Afterwards, the surround view images of different roads are spliced ​​together through calibration to obtain the first key frame image.

[0031]

[0032] Step S130 , determining a plurality of historical frame images without a vehicle bottom blind spot corresponding to the first key frame image.

[0033] The first key frame image and the multiple historical frame images can be sequentially or discontinuously timed. If the first key frame image and the multiple historical frame images are discontinuous, the time interval between each frame must be within a preset range to prevent excessive differences between the point cloud data of the two frames and improve the timeliness of the target object information subsequently inherited from the multiple historical frame images. It should be noted that the first key frame image is the image captured at the current moment, while the multiple historical frame images are images captured before the current moment.

[0034] Step S140 : Based on the multiple historical frame images, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain a second key frame image without the blind spot under the vehicle.

[0035] Exemplarily, a historical frame image with the highest correlation with the first key frame image is selected from multiple historical frame images, and a region for texture mapping is determined within the historical frame image. A three-dimensional projection model corresponding to the panoramic surround view imaging system is obtained, and the world coordinates of the model points of the three-dimensional projection model in the world coordinate system are obtained. The world coordinates corresponding to the aforementioned region are calculated. Using texture mapping, the first key frame image and the aforementioned region are attached to the three-dimensional projection model, and a three-dimensional panoramic image is obtained after splicing. Furthermore, the three-dimensional panoramic image is converted into a two-dimensional second key frame image based on the internal and external parameters of the panoramic surround view system camera. The second key frame image is identical to the first key frame image, both being images from a bird's-eye view.

[0036] Step S150: Convert the second key frame image into a local semantic map based on the initial posture of the target vehicle.

[0037] For example, the initial position and posture of the target vehicle when it starts to start in the target environment are obtained using sensors installed on the target vehicle. The initial position and posture include initial position data and initial posture data. The initial position data can be absolute position data (for example, the longitude and latitude directly obtained by GPS (Global Positioning System)), or relative position data (for example, the cumulative motion parameters obtained by the number of rotations and speed of the target vehicle wheels, such as distance). The initial posture data can be the absolute heading angle obtained by differential GPS, or the relative heading angle (for example, the cumulative motion parameters obtained by the rotation angle of the target vehicle wheels, such as direction).

[0038] For example, the local pose of the target vehicle can be determined based on the initial pose of the target vehicle when it starts and the accumulated movement distance and direction parameters obtained by the IMU (Inertial Measurement Unit). Since this pose is obtained only by movement changes, it is a relative local pose.

[0039] Furthermore, detection, tracking, and recognition are performed on the second keyframe image, and semantic entities (equivalent to the target object) in the target environment are determined based on the detection, tracking, and recognition results. The second keyframe image is converted into a local semantic map based on the initial position of the target vehicle, the semantic entities in the target environment, and the local pose of the target vehicle.

[0040] Step S160 , stitching the local semantic maps of the target vehicle in the target environment at multiple moments to generate a global map of the target environment.

[0041] The semantic entities in the local semantic maps at different times are matched to achieve the splicing of the same semantic entities at the same location, thereby obtaining a global map of the target environment.

[0042] In this embodiment, motion compensation is used to map the historical frame image to the blind spot under the vehicle of the first key frame image according to specific transformation conditions. This eliminates the occlusion and truncation effects caused by the visual blind spot, ensures the continuity and accuracy of the semantic features of the second key frame image, and avoids feature truncation or loss during the mapping process, providing driving assistance and improving the user experience. Furthermore, this application does not require the installation of an additional camera module, which is cost-effective.

[0043] exist Figure 1 On the basis of the illustrated embodiment, the current posture of the target vehicle can also be determined by matching features in the global map based on the initial posture of the target vehicle.

[0044] Specifically, in this embodiment, the initial position and posture of the target vehicle refers to the position and posture of the target vehicle at the current moment collected by the odometer during the driving process of the target vehicle.

[0045] Furthermore, a feature truth database is obtained based on the global map. During the driving process of the target vehicle, the semantic features of the target object contained in the image are determined based on the multi-way surround view images collected by the target vehicle. Based on the semantic features of the target object, the positioning feature position is found in the feature truth database. For example, the semantic features of the target object are converted to the global coordinate system corresponding to the global map, and then the semantic features of the target object and the feature truth values ​​in the global map are matched for feature type. Afterwards, the positioning feature position is solved to obtain a solved pose. For example, by using the static feature point information in the positioning feature position, triangulation positioning is used to calculate the current global position of the static feature point information and the relative position to the laser radar, thereby solving the current global position of the laser radar, and the current position of the vehicle is obtained by the positioning feature position, and then the solved pose is obtained.

[0046] In addition, the global position of the target vehicle collected by the satellite positioning system, the calculated pose of the target vehicle and the initial pose of the target vehicle are fused, and the fusion result is used as the current pose of the target vehicle.

[0047] In this embodiment, features of the target environment are derived from multiple surround view images and then registered within the global map. Because global information is known, search speed is fast. Once the static global position is determined, the target vehicle's global position information is deduced based on the geometric relationship between the static global position and the target vehicle, improving positioning accuracy.

[0048] Figure 2 FIG. 1 is a flow chart of obtaining a second key frame image provided by an exemplary embodiment of the present application. Figure 1 Based on the embodiment shown, Figure 2The embodiment shown is described below in detail. Figure 2 The embodiment shown is Figure 1 The differences and similarities between the illustrated embodiments are not described in detail.

[0049] like Figure 2 As shown, in an embodiment of the present application, based on multiple historical frame images, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain a second key frame image without the blind spot under the vehicle, which includes the following steps.

[0050] Step S210 : determining a reference frame image for mapping the vehicle bottom blind spot in the first key frame image from a plurality of historical frame images.

[0051] In one implementation, for the blind spot under the vehicle in the first key frame image, multiple historical frame images are sorted in chronological order and each provides a valid mapping area corresponding to part or all of the blind spot under the vehicle. The historical frame image that provides the largest valid mapping area is determined and used as the reference frame image to avoid poor display effects caused by combining and mapping the blind spot under the vehicle from too many historical frame images.

[0052] If multiple historical frames don't provide the entire valid mapping area, the one with the most pixels corresponding to the blind spot under the vehicle in the first key frame can be selected as a reference frame to obtain the valid mapping area. This method of selecting a larger historical frame image to fill the blind spot under the vehicle minimizes the number of stitching operations and effectively avoids image misalignment and brightness inconsistencies caused by stitching multiple images, improving the display of the blind spot under the vehicle.

[0053] In another implementation method, the transformation matrices of multiple historical frame images and the first key frame image can be determined based on the driving parameters of the target vehicle; based on the target sampling error, the feature information corresponding to each of the multiple historical frame images is randomly sampled with the feature information in the first key frame image as the standard to determine the sampling values ​​corresponding to each of the multiple historical frame images; based on the transformation matrices of the multiple historical frame images and the first key frame image and the sampling values ​​corresponding to each of the multiple historical frame images, a reference frame image is determined from the multiple historical frame images.

[0054] The driving parameters include steering wheel angle information and vehicle speed information, and the sampling value represents the degree of matching between the feature information in the historical frame image and the feature information in the first key frame image. Furthermore, for each of the multiple historical frame images, the computing resources (expressed as numerical values) consumed by the historical frame image using its corresponding transformation matrix under the sampling value are determined. The weights corresponding to the computing resources and sampling values ​​are determined, and based on the weights, the score of the historical frame image is obtained, and the historical frame image with the highest score is determined as the reference frame image. The higher the score, the higher the feature similarity between the corresponding historical frame image and the first key frame image. When it is subsequently used for texture mapping, the computing resources consumed are less, which ensures the feature similarity between the reference frame image and the first key frame image while reducing the resources consumed in the subsequent texture mapping.

[0055] Step S220 , determining the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value of the target vehicle at the historical moment corresponding to the reference frame image.

[0056] The wheel pulse count value refers to the number of pulses in one wheel rotation. For example, 1080 pulses can be obtained in one wheel rotation.

[0057] Step S230 : determining the rotation angle and offset of the target vehicle in the first key frame image and the reference frame image based on the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count values ​​at the historical moments.

[0058] For example, the relative offset and rotation angle of the target vehicle between the reference frame image and the second key frame image are calculated using the principles of inertial navigation. An inertial navigation system (INS) is an autonomous navigation system that does not rely on external information or radiate energy. Its basic operating principle is based on Newtonian mechanics. By measuring the acceleration of a vehicle in an inertial reference frame, integrating it over time, and transforming it into a navigation coordinate system, information such as velocity, yaw angle, and position in the navigation coordinate system can be obtained.

[0059] Step S240: determining a valid mapping area in the reference frame image based on the rotation angle and the offset.

[0060] The reference frame image is overlaid on the first key frame image, and the target vehicle in the reference frame image is rotated by the aforementioned rotation angle and offset by the aforementioned offset relative to the target vehicle in the first key frame image, thereby obtaining a valid mapping area in the reference frame image that maps the blind spot under the vehicle onto the first key frame image. For example, image B is the reference frame image, and image A is the first key frame image. Based on the relative offset and rotation angle, image B is offset and rotated with the target vehicle in image B as the base point, thereby mapping image B onto image A, thereby obtaining the valid mapping area.

[0061] Step S250 : Based on the effective mapping area, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain a second key frame image.

[0062] In one implementation, a contour line of the blind spot under the vehicle in a first key frame image is determined; an outer edge line is obtained by expanding a target distance outward in a direction away from the blind spot under the vehicle along the contour line, and an area formed by the contour line and the outer edge line is used as a transition area; based on a valid mapping area, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain an image to be calibrated; a grayscale gradient value of the valid mapping area in the image to be calibrated is determined; and based on the grayscale gradient value, the grayscale value of the transition area in the image to be calibrated is calibrated to obtain a second key frame image.

[0063] It is understandable that the shape of the transition area is not limited here, as long as the inner frame line of the transition area can coincide with the outline of the blind spot under the vehicle.

[0064] By setting the transition area and adjusting the grayscale gradient of the pixels in the transition area, the visual effect of the fusion boundary between the effective mapping area and other areas in the second key frame image can be eliminated.

[0065] In this embodiment, a reference frame image is determined from multiple historical frames. An image with some degree of similarity to the first key frame image can be selected to ensure feature consistency during subsequent texture mapping of the blind spot under the vehicle. The effective mapping area within the reference frame image is determined based on the rotation angle and offset, further ensuring the continuity of semantic features during the texture mapping process and improving the visual quality of the second key frame image.

[0066] Figure 3 The figure shows a flow chart of converting into a local semantic map provided by an exemplary embodiment of the present application. Figure 1 Based on the embodiment shown, Figure 3 The embodiment shown is described below in detail. Figure 3 The embodiment shown is Figure 1 The differences and similarities between the illustrated embodiments are not described in detail.

[0067] like Figure 3 As shown, in an embodiment of the present application, based on the initial posture of the target vehicle, the second key frame image is converted into a local semantic map, including the following steps.

[0068] Step S310 : performing semantic feature extraction on the second key frame image to obtain semantic features of the target object contained in the second key frame image.

[0069] Semantic features include position features, direction features and category features. Exemplarily, semantic feature extraction is performed by a deep neural network, including semantic entities (i.e., target objects) such as parking lines, parking corners, speed bumps, lane lines, ground signs, buildings, and wheel blocks. Further, the attribute information of the semantic entity is determined. For example, the attribute information can indicate the physical characteristics of the semantic entity, or the semantic entity may affect the properties of the target vehicle's own movement. For example, the attribute information can be spatial attribute information such as the position, shape, size, and orientation of each semantic entity, or it can be category attribute information of each semantic entity (such as, whether each semantic entity is a feasible road, curb, lane and lane line, traffic sign, road surface sign, traffic light, stop line, crosswalk, roadside tree or pillar, etc.). The relative position relationship between the semantic entity and the target vehicle is determined based on the second key frame image, and the spatial attribute information of the semantic entity is determined based on the local posture information and the relative position relationship. For example, the spatial attribute can include various attributes related to spatial characteristics such as the size, shape, orientation, height, and occupancy of the semantic mark. In addition to the spatial attribute information, for example, the category of each semantic entity may be further determined based on the second key frame image.

[0070] Step S320: Determine the position and posture of the target vehicle at the current moment.

[0071] In this embodiment, the target vehicle's current position is equivalent to Figure 1 Local pose information in the illustrated embodiment.

[0072] Step S330 : converting the second key frame image into a local semantic map based on the semantic features of the target object contained in the second key frame image, the current position and initial position of the target vehicle.

[0073] After determining the semantic entities and their attribute information included in the second keyframe image, this information can be integrated to construct a local semantic map based on the second keyframe image. In other words, the semantic landmark results of the second keyframe image are reconstructed and attributes such as position and size are added to obtain a semantic landmark map with absolute attributes.

[0074] In this embodiment, the local semantic map is obtained by using the second key frame image without the blind spot under the vehicle, which ensures the integrity, continuity and accuracy of the semantic features in the local semantic map, making the subsequent mapping and positioning more accurate.

[0075] Figure 4 The figure shows a flow chart of generating a global map provided by an exemplary embodiment of the present application. Figure 1 Based on the embodiment shown, Figure 4 The embodiment shown is described below in detail. Figure 4 The embodiment shown is Figure 1 The differences and similarities between the illustrated embodiments are not described in detail.

[0076] like Figure 4 As shown, in an embodiment of the present application, local semantic maps of a target vehicle in a target environment at multiple moments are spliced ​​to generate a global map of the target environment, including the following steps.

[0077] Step S410 : determining at least two local semantic maps to be matched from the local semantic maps at multiple moments.

[0078] At least two local semantic maps to be matched are continuous in time sequence.

[0079] Step S420 : determining a target point cloud and a source point cloud in at least two local semantic maps to be matched.

[0080] If there are two local semantic maps to be matched, the target point cloud and the source point cloud are the point clouds in these two local semantic maps respectively. For example, the local semantic map containing the source point cloud can be used as the reference, and the local semantic map containing the target point cloud can be used for stitching according to the semantic features of the local semantic map containing the target point cloud. If there are more than two local semantic maps, one of the local semantic maps can be selected as the source point cloud, and the remaining local semantic maps to be matched can be stitched with it.

[0081] Step S430 : determining the nearest neighbor point of each point in the target point cloud to the source point cloud based on the geometric features of the target point cloud.

[0082] Step S440 , based on the nearest neighbor of each point in the target point cloud to the source point cloud, local semantic maps of the target vehicle in the target environment at multiple moments are spliced ​​to generate a global map of the target environment.

[0083] The source point cloud is denoted as P, and the target point cloud is denoted as Q. For each point cloud in Q, find the corresponding nearest point cloud in P to form a matching point pair. The sum of the Euclidean distances of all matching point pairs is used as the objective function to be solved. The rotation matrix R and the translation matrix t are obtained by singular value decomposition to minimize the objective function. According to R and t, Q is transformed (including rotation and translation) to obtain the new Q ′ , and find the corresponding point pairs again, and iterate until the error is minimized. Using the rotation matrix R and translation matrix t obtained when the error is minimized, the local semantic maps are spliced ​​to generate a global map of the target environment.

[0084] In this embodiment, matching point pairs are obtained between the source and target point clouds. A rotation matrix and a translation matrix are constructed based on these matching point pairs. The source point cloud is then transformed into the coordinate system of the target point cloud using these rotation and translation matrices. The error function between the transformed source and target point clouds is estimated. If the error function exceeds a threshold, the iteration process continues until the error value after concatenation based on the rotation and translation matrices meets a given error requirement. This method is simple and offers good accuracy.

[0085] Combined with the above Figures 1 to 4 , describes in detail the embodiment of the visual semantics-based mapping method of this application, and Figure 5 , describes in detail the embodiment of the visual semantics-based mapping device of the present application. It should be understood that the description of the embodiment of the visual semantics-based mapping method corresponds to the description of the embodiment of the visual semantics-based mapping device. Therefore, for parts not described in detail, reference can be made to the above method embodiments.

[0086] Figure 5 FIG. 1 is a schematic diagram of a structure of a visual semantics-based mapping device provided by an exemplary embodiment of the present application. Figure 5 As shown, the visual semantics-based mapping device 50 provided in the embodiment of the present application includes:

[0087] An acquisition module 510 is configured to acquire a multi-way surround view image collected by the target vehicle in the target environment at the current moment;

[0088] A first stitching module 520 is configured to stitch the multiple surround view images together to generate a first key frame image, where the first key frame image includes a blind spot under the vehicle;

[0089] A determination module 530 is configured to determine a plurality of historical frame images without a vehicle bottom blind spot corresponding to the first key frame image;

[0090] A mapping module 540 is configured to perform texture mapping on the blind spot under the vehicle in the first key frame image based on multiple historical frame images to obtain a second key frame image without the blind spot under the vehicle;

[0091] a conversion module 550 for converting the second key frame image into a local semantic map based on the initial pose of the target vehicle;

[0092] The second splicing module 560 is used to splice the local semantic maps of the target vehicle in the target environment at multiple moments to generate a global map of the target environment.

[0093] In one embodiment of the present application, the mapping module 540 is also used to determine a reference frame image for mapping the blind spot under the vehicle in the first key frame image from multiple historical frame images; determine the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value of the target vehicle at the historical moment corresponding to the reference frame image; determine the rotation angle and offset of the target vehicle in the first key frame image and the reference frame image based on the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value at the historical moment; determine a valid mapping area in the reference frame image based on the rotation angle and the offset; and perform texture mapping on the blind spot under the vehicle in the first key frame image based on the valid mapping area to obtain a second key frame image.

[0094] In one embodiment of the present application, the mapping module 540 is further used to determine, based on the driving parameters of the target vehicle, the transformation matrices of the multiple historical frame images and the first key frame image respectively; based on the target sampling error, with the feature information in the first key frame image as the standard, randomly sample the feature information corresponding to each of the multiple historical frame images, and determine the sampling values ​​corresponding to each of the multiple historical frame images, where the sampling values ​​represent the degree of matching between the feature information in the historical frame images and the feature information in the first key frame image; based on the transformation matrices of the multiple historical frame images and the first key frame image and the sampling values ​​corresponding to each of the multiple historical frame images, determine the reference frame image from the multiple historical frame images.

[0095] In one embodiment of the present application, the mapping module 540 is further used to determine the contour line of the blind spot under the vehicle; expand the target distance outward along the contour line in the direction away from the blind spot under the vehicle to obtain an outer edge line, and use the area composed of the contour line and the outer edge line as the transition area; based on the effective mapping area, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain an image to be calibrated; determine the grayscale gradient value of the effective mapping area in the image to be calibrated; based on the grayscale gradient value, calibrate the grayscale value of the transition area in the image to be calibrated to obtain a second key frame image.

[0096] In one embodiment of the present application, the conversion module 550 is also used to extract semantic features from the second key frame image to obtain semantic features of the target object contained in the second key frame image, the semantic features including position features, direction features and category features; determine the position and posture of the target vehicle at the current moment; and convert the second key frame image into a local semantic map based on the semantic features of the target object contained in the second key frame image, the position and posture of the target vehicle at the current moment and the initial posture.

[0097] In one embodiment of the present application, the second stitching module 560 is further used to determine at least two local semantic maps to be matched in the local semantic maps at multiple times; determine the target point cloud and the source point cloud respectively in the at least two local semantic maps to be matched; determine the nearest neighbor point of each point in the target point cloud to the source point cloud based on the geometric features of the target point cloud; stitch the local semantic maps of the target vehicle at multiple times in the target environment based on the nearest neighbor point of each point in the target point cloud to the source point cloud to generate a global map of the target environment.

[0098] In one embodiment of the present application, a positioning module is further included, which is used to match features in the global map based on the initial position of the target vehicle to determine the current position of the target vehicle.

[0099] Below, reference Figure 6 To describe the electronic device according to the embodiment of the present application. Figure 6 Shown is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.

[0100] like Figure 6 As shown, the electronic device 60 includes one or more processors 601 and a memory 602 .

[0101] The processor 601 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 60 to perform desired functions.

[0102] The memory 602 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 601 may execute the program instructions to implement the methods of the various embodiments of the present application described above and / or other desired functions. Various contents such as multiple surround view images, a first key frame image, a second key frame image, multiple historical frame images, a local semantic map, a global map, etc. may also be stored in the computer-readable storage medium.

[0103] In one example, the electronic device 60 may further include an input device 603 and an output device 604 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0104] The input device 603 may include, for example, a keyboard, a mouse, and the like.

[0105] The output device 604 can output various information to the outside, including multi-way surround view images, first key frame images, second key frame images, multiple historical frame images, local semantic maps, global maps, etc. The output device 604 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0106] Of course, to simplify, Figure 6 Only some of the components related to the present application in the electronic device 60 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 60 may further include any other appropriate components according to specific application scenarios.

[0107] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present application described above in this specification.

[0108] The computer program product may be written in any combination of one or more programming languages ​​to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0109] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the method according to various embodiments of the present application described above in this specification.

[0110] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0111] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.

[0112] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0113] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0114] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0115] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A visual semantics-based mapping method, characterized in that: include: Obtaining the current multi-way surround view images collected by the target vehicle in the target environment; Performing bird's-eye view stitching on the multiple surround view images to generate a first key frame image, wherein the first key frame image includes a blind spot under the vehicle; Determine a plurality of historical frame images without a vehicle bottom blind spot corresponding to the first key frame image; Based on the multiple historical frame images, texture mapping is performed on the blind spot under the vehicle in the first key frame image to obtain a second key frame image without the blind spot under the vehicle; Based on the initial pose of the target vehicle, converting the second key frame image into a local semantic map; splicing local semantic maps of the target vehicle in the target environment at multiple moments to generate a global map of the target environment; The step of performing texture mapping on the blind spot under the vehicle in the first key frame image based on the multiple historical frame images to obtain a second key frame image without the blind spot under the vehicle includes: Determining a reference frame image for mapping the vehicle bottom blind spot in the first key frame image from the plurality of historical frame images; Determining a wheel speed pulse count value of the target vehicle at a current moment and a wheel speed pulse count value of the target vehicle at a historical moment corresponding to the reference frame image; Determining a rotation angle and an offset of the target vehicle in the first key frame image and the reference frame image based on the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value at the historical moment; Determining a valid mapping area in the reference frame image based on the rotation angle and the offset; Based on the effective mapping area, texture mapping is performed on the blind area under the vehicle in the first key frame image to obtain the second key frame image.

2. The method according to claim 1, characterized in that The determining of a reference frame image for mapping the vehicle bottom blind spot in the first key frame image from the plurality of historical frame images includes: Determining transformation matrices of the plurality of historical frame images and the first key frame image based on the driving parameters of the target vehicle; Based on a target sampling error, taking the feature information in the first key frame image as a criterion, randomly sampling the feature information corresponding to each of the plurality of historical frame images, and determining a sampling value corresponding to each of the plurality of historical frame images, wherein the sampling value represents a degree of matching between the feature information in the historical frame image and the feature information in the first key frame image; The reference frame image is determined from the multiple historical frame images based on the transformation matrix between the multiple historical frame images and the first key frame image and the sampling values ​​corresponding to each of the multiple historical frame images.

3. The method according to claim 1, characterized in that The step of performing texture mapping on the vehicle bottom blind area in the first key frame image based on the effective mapping area to obtain the second key frame image includes: Determining the contour line of the blind spot under the vehicle; Along the contour line, the target distance is expanded outward in a direction away from the blind spot under the vehicle to obtain an outer edge line, and an area formed by the contour line and the outer edge line is used as a transition area; Based on the effective mapping area, texture mapping is performed on the blind area under the vehicle in the first key frame image to obtain an image to be calibrated; Determining the grayscale gradient value of the effective mapping area in the image to be calibrated; Based on the grayscale gradient value, the grayscale value of the transition area in the image to be calibrated is calibrated to obtain the second key frame image.

4. The method according to any one of claims 1 to 3, characterized in that The converting the second key frame image into a local semantic map based on the initial pose of the target vehicle includes: Performing semantic feature extraction on the second key frame image to obtain semantic features of the target object contained in the second key frame image, wherein the semantic features include position features, direction features, and category features; Determining the position and posture of the target vehicle at the current moment; Based on the semantic features of the target object contained in the second key frame image, the current position and initial position of the target vehicle, the second key frame image is converted into the local semantic map.

5. The method according to any one of claims 1 to 3, characterized in that The step of splicing the local semantic maps of the target vehicle in the target environment at multiple moments to generate a global map of the target environment includes: Determining at least two local semantic maps to be matched among the local semantic maps at the multiple moments; Determining a target point cloud and a source point cloud in the at least two local semantic maps to be matched respectively; Determining the nearest neighbor point of each point in the target point cloud to the source point cloud based on the geometric features of the target point cloud; Based on the nearest neighbor of each point in the target point cloud to the source point cloud, local semantic maps of the target vehicle in the target environment at multiple moments are spliced ​​to generate a global map of the target environment.

6. The method according to any one of claims 1 to 3, characterized in that Also includes: Based on the initial position and posture of the target vehicle, features in the global map are matched to determine the current position and posture of the target vehicle.

7. A mapping device based on visual semantics, characterized in that: include: An acquisition module is used to acquire the multi-way surround view images collected by the target vehicle in the target environment at the current moment; A first stitching module is configured to stitch the multi-channel surround view images together to generate a first key frame image, wherein the first key frame image includes a blind spot under the vehicle; A determination module, configured to determine a plurality of historical frame images without a vehicle bottom blind spot corresponding to the first key frame image; a mapping module, configured to perform texture mapping on the blind spot under the vehicle in the first key frame image based on the plurality of historical frame images, to obtain a second key frame image without the blind spot under the vehicle; a conversion module, configured to convert the second key frame image into a local semantic map based on the initial pose of the target vehicle; A second splicing module is used to splice the local semantic maps of the target vehicle in the target environment at multiple moments to generate a global map of the target environment; The mapping module is further configured to determine a reference frame image for mapping the blind spot under the vehicle in the first key frame image from among the plurality of historical frame images; determine a wheel speed pulse count value of the target vehicle at a current moment and a wheel speed pulse count value of the target vehicle at a historical moment corresponding to the reference frame image; Determining a rotation angle and an offset of the target vehicle in the first key frame image and the reference frame image based on the wheel speed pulse count value of the target vehicle at the current moment and the wheel speed pulse count value at the historical moment; Determining a valid mapping area in the reference frame image based on the rotation angle and the offset; Based on the effective mapping area, texture mapping is performed on the blind area under the vehicle in the first key frame image to obtain the second key frame image.

8. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mapping method and system based on visual semantic point cloud

    CN112348921A

  • Blind area image acquisition method and related terminal device

    CN113228135A