Map information processing method, map information acquisition method, and device
By combining two-dimensional and three-dimensional information to generate map information including the elevation of location points, the problem of insufficient accuracy of existing two-dimensional maps in complex traffic environments is solved, and a more accurate reflection of the traffic environment is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-04-02
AI Technical Summary
Existing two-dimensional map information cannot accurately reflect the real situation in the traffic environment, especially in complex scenarios such as uphill and downhill, where the lack of height information of location points leads to inaccurate map information.
By acquiring three-dimensional information about the traffic environment and combining it with two-dimensional map information, the elevation of each location point is determined, thereby generating updated map information containing two-dimensional coordinates and elevation.
The generated updated map information can more accurately reflect the traffic environment and improve the accuracy of map information, especially in scenarios with complex terrain changes.
Smart Images

Figure CN2025106587_02042026_PF_FP_ABST
Abstract
Description
A map information processing method, a map information acquisition method, and a device
[0001] The present application claims priority from the Chinese patent application No. 202411355609.0 filed on September 26, 2024, and entitled "A map information processing method, a map information acquisition method, and a device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to intelligent driving technology, and in particular, to a map information processing method, a map information acquisition method, and a device. BACKGROUND
[0003] With the rapid development of intelligent driving technology, more and more scenarios need to use map information. Currently, the map information used is usually two-dimensional (2D) map information of a traffic environment, which includes two-dimensional coordinates of position points in road elements in the traffic environment. However, in some traffic scenarios, such as uphill and downhill traffic scenarios, the 2D map information cannot accurately reflect the real traffic environment because it does not include the height of each position point. Therefore, a generation scheme of map information that can contain both two-dimensional coordinates and height of position points is urgently needed. SUMMARY
[0004] The present application provides a map information processing method, a map information acquisition method, and a device, and provides a generation scheme of map information that can contain both two-dimensional coordinates and height of position points. Thus, the updated first map information can more accurately reflect the real traffic environment.
[0005] The present application provides the following technical solutions:
[0006] In a first aspect, the present application provides a method for processing map information, which can be used in the field of intelligent driving. In the method, a first device not only obtains first map information of a first traffic environment, but also obtains three-dimensional (3D) information of the first traffic environment. The first map information of the first traffic environment includes a two-dimensional coordinate of each first position point in at least one first position point, and the at least one first position point is a position point in a road element included in the first traffic environment. The 3D information of the first traffic environment includes a height of each first position region in at least one position region, and the at least one position region is a position region in an object included in the first traffic environment. Then, the first device can obtain the height of each first position point in the at least one first position point based on the first map information of the first traffic environment and the 3D information of the first traffic environment, that is, complete the high value adsorption of each first position point, so as to obtain updated first map information, and the updated first map information includes the two-dimensional coordinate and the height of each first position point in the at least one first position point.
[0007] Exemplarily, the first traffic environment in the present application can also be referred to as a first road environment, and the first map information of the first traffic environment can be understood as two-dimensional (2D) map information of the first traffic environment. Optionally, the first map information can be 2D map information in a bird eye view (BEV) (hereinafter referred to as "in a BEV view"). The BEV view in the present application can also be referred to as a top view.
[0008] Optionally, the road element can be a road element on the ground, for example, the road element can include a lane boundary line, and the lane boundary line in the present application can be a line existing in a real road to divide different lanes. Optionally, the road element can also include at least one of the following: a lane center line, a lane, a boundary line of a road, an intersection, a sidewalk, a turning indication line or other road elements, etc. The lane center line in the present application can be understood as a virtual central axis of the lane, and the lane center line can be a line that does not exist in a real road. The lane center line can be used to position the lane.
[0009] Exemplarily, each position region can be represented as a cube, and the at least one position region can be understood as dividing the object in the first traffic environment into the at least one position region. Exemplarily, the object in the first traffic environment can include a road element in the first traffic environment, and the object in the first traffic environment can also include other objects other than the road element in the first traffic environment, for example, the other objects other than the road element in the first traffic environment can include a traffic signal, a street lamp, a tree, a building, a fire hydrant, a triangular board, a cone barrel, a warning column, a crash barrel, a construction board, a guide board or other types of static obstacles, and the like. Optionally, the other objects other than the road element in the first traffic environment can also include dynamic obstacles in the first traffic environment, for example, the dynamic obstacles can include a vehicle, an electric vehicle, a bicycle, a pedestrian, an animal or other dynamic obstacles, and the like.
[0010] In the implementation, not only the first map information of the traffic environment is obtained, but also 3D information of the traffic environment is obtained, the 3D information includes a height of at least one position region included by an object in the traffic environment, and then the height of the at least one first position point is obtained based on the first map information and the 3D information, and the updated first map information can include the two-dimensional coordinates and the height of the at least one position point, thereby providing a generation scheme of the map information including the two-dimensional coordinates and the height of the position point, and the updated first map information can more accurately reflect the real traffic environment.
[0011] In a possible implementation, any one of the at least one first position point is referred to as a second position point, and the first device obtains the updated first map information based on the first map information of the first traffic environment and the 3D information of the first traffic environment, which can include: the first device determines the height of the second position point based on a height of a target position region in the at least one position region, wherein the target position region is a position region corresponding to the second position point in the at least one position region; and the first device can repeatedly execute the foregoing step at least once to obtain the height of each first position point in the at least one first position point.
[0012] In the implementation, for any one of the at least one first position point (i.e., the second position point), the height of a target position region corresponding to the first position point in the at least one position region is obtained, and the height of the second position point is determined based on the height of the target position region, that is, each first position point in the at least one first position point is traversed to obtain the height of each first position point one by one, that is, the process of obtaining the height of the at least one first position point is more finely managed, and the accuracy of the obtained height of each first position point is improved.
[0013] In an implementation, the two-dimensional coordinate of the first position point may, for example, include first coordinate information in an X-axis direction and a Y-axis direction, and the X-axis and the Y-axis are perpendicular to each other. For example, the two-dimensional coordinate of the first position point may be first coordinate information in a first coordinate system, and the first coordinate system may be a vehicle body coordinate system. The 3D information of the first traffic environment further includes two-dimensional coordinates of each of the at least one position region, in other words, the 3D information of the first traffic environment may include second coordinate information of each of the at least one position region in a second coordinate system, and the second coordinate information of each of the at least one position region may include two-dimensional coordinates and a height of each of the at least one position region, and the second coordinate system may be a vehicle body coordinate system.
[0014] The first device obtains updated first map information based on the first map information and the 3D information, which may include: the first device determines at least one target position region corresponding to a second position point (i.e., any one of the at least one first position point) from the at least one position region based on the two-dimensional coordinate of the second position point and the two-dimensional coordinates of each of the at least one position region, and then determines the height of the second position point based on the height of the at least one target position region.
[0015] For example, if the at least one target position region includes only one target position region, the first device may directly determine the height of the one target position region as the height of the second position point. If the at least one target position region includes at least two target position regions, in an implementation, the first device may determine the smallest one of at least two heights corresponding to the at least two target position regions as the height of the second position point, since the second position point may be a position point on the road surface, it is feasible to directly determine the smallest one of the at least two heights as the height of the second position point; or in another implementation, the first device may determine the average height of the at least two heights corresponding to the at least two target position regions as the height of the second position point, etc.
[0016] In the implementation, it is determined that the target position region corresponding to each first position point is found based on the two-dimensional coordinate of the first position point and the two-dimensional coordinate of the position region in the 3D information, and a simple and effective method for determining the target position region is provided.
[0017] In a possible implementation, the first device determines the target location region corresponding to the second location point from the at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, which can include that the first device determines the target location region corresponding to the second location point from the at least one location region based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region, and the semantic type of each location region. For example, the semantic type of a location region can be the semantic type of a lane boundary line, a boundary line of a road, a sidewalk, an intersection, a turning indication line, a traffic signal, a street lamp, a tree, a building, a fire hydrant, a triangular board, a cone barrel, a warning column, a crash barrel, a construction board, a guide board, or other objects that can appear in a traffic environment, and the like; optionally, if the objects in the first traffic environment also include dynamic obstacles in the first traffic environment, the semantic type of a location region can also be the semantic type of a vehicle, an electric vehicle, a bicycle, a pedestrian, an animal, or other dynamic obstacles, and the like.
[0018] Optionally, the first device can determine the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, select at least one candidate location region closest to the two-dimensional coordinates of the second location point from the at least one location region, and the distance between the two-dimensional coordinates of each candidate location region in the at least one candidate location region and the two-dimensional coordinates of the second location point can be consistent, for example, if the at least one candidate location region includes at least two candidate location regions, the two-dimensional coordinates of different candidate location regions in the at least two candidate location regions can be consistent, and the difference is that the heights are different. The first device can determine one target location region corresponding to the semantic type of the second location point from the at least one candidate location region based on the semantic type of each candidate location region.
[0019] In the implementation, in the process of determining the target location region corresponding to any one first location point, not only the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location region are used, but also the semantic type of each location region is used. Since the matching is performed only based on the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location region, one first location point can correspond to multiple location regions, the two-dimensional coordinates of the multiple location regions are consistent, but the heights are different, and then the target location region corresponding to the first location point can be further determined based on the semantic type of each location region, which is beneficial to accurately find the target location region corresponding to each first location point, and further beneficial to obtain the accurate height of each first location point, so as to improve the accuracy of the updated first map information.
[0020] In a possible implementation, the first device determines the target location region corresponding to the second location point from the at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, which can include that the first device determines the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, and selects at least one target location region closest to the two-dimensional coordinates of the second location point from the at least one location region, and the distance between the two-dimensional coordinates of each target location region in the at least one target location region and the two-dimensional coordinates of the second location point can be uniform, for example, if the at least one target location region includes at least two target location regions, the two-dimensional coordinates of different target location regions in the at least two target location regions can be uniform, and the difference is that the heights are different. Alternatively, the first device can round up the position parameters in the X-axis direction and the position parameters in the Y-axis direction included in the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region. Alternatively, the first device can round down the position parameters in the X-axis direction and the position parameters in the Y-axis direction included in the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region.
[0021] In a possible implementation, the first map information is map information in a BEV view, and the method can further include that the first device determines evaluation information corresponding to the updated first map information based on a first image in a perspective view (PV) of the first traffic environment, and the evaluation information indicates the accuracy of the updated first map information.
[0022] In the implementation, since the image in the PV view of the first traffic environment can accurately reflect the situation of the first traffic environment, and the updated first map information is used to more accurately reflect the first traffic environment, it is feasible to determine the accuracy of the updated first map information by means of the image in the PV view of the first traffic environment, and the image in the PV view of the first traffic environment is easy to obtain, thereby providing an easily implemented detection scheme for the accuracy of the updated first map information.
[0023] In a possible implementation, the at least one first position point includes a position point on a lane boundary line in the first traffic environment, and the first device determines the evaluation information corresponding to the updated first map information based on the first image in the perspective view of the first traffic environment, including: the first device projects the first position point into the first image in the perspective view based on the two-dimensional coordinates of the first position point and the height of the first position point to obtain first position information, the first position information indicating a position of a projection point corresponding to the first position point in the first image in the perspective view, and the projection point can be understood as a projection point obtained by projecting a position point on a lane boundary line in a BEV view into a first image in a PV view; and then the evaluation information can be determined based on the first position information and the first image in the perspective view.
[0024] For example, in a case, the at least one first position point can all be position points on a lane boundary line included in the first traffic environment, and then the first device projects each of the at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the at least one first position point to obtain first position information, and the first position information can include a position of each of at least one projection point corresponding to the at least one first position point in the first image in the perspective view.
[0025] In another case, if the at least one first position point included in the updated first map information includes not only position points on a lane boundary line included in the first traffic environment but also position points in other road elements included in the first traffic environment, the first device can screen the at least one first position point included in the updated first map information to obtain screened at least one first position point, and the screened at least one first position point includes only position points on the lane boundary line; and the first device projects each of the screened at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the screened at least one first position point to obtain first position information, and the first position information can include a position of each of at least one projection point corresponding to the screened at least one first position point in the first image in the perspective view.
[0026] In the present implementation, a scheme is proposed for projecting the first position point in the BEV perspective into the first image in the PV perspective based on the two-dimensional coordinates and the height of the first position point in the updated first map information, obtaining the position information of the projection point, and further evaluating the accuracy of the updated first map information based on the position information of the projection point. Since the first image in the PV perspective of the first traffic environment can accurately reflect the first traffic environment, the manner of projecting the first position point in the BEV perspective into the first image in the PV perspective can maximize the use of the first image in the PV perspective, which is conducive to obtaining more accurate evaluation information.
[0027] In a possible implementation, the first device determines the evaluation information based on the first position information and the first image in the perspective view, which can include: the first device performs a semantic recognition operation based on the first image in the perspective view to obtain second position information, the second position information including position information of pixel points in the first image in the perspective view with a semantic type of lane boundary line; and exemplarily, the second position information can include coordinate information of the pixel points in the first image in the perspective view with the semantic type of lane boundary line. Then, the first device can determine the evaluation information according to the first position information and the second position information.
[0028] Optionally, the first position information can indicate position information of at least one first lane boundary line composed of projection points obtained by projecting the first position point on the lane boundary line in the BEV perspective into the first image in the PV perspective, and the second position information can indicate position information of at least one second lane boundary line in the first image in the PV perspective. Exemplarily, the first device can determine a first distance between the at least one first lane boundary line and the at least one second lane boundary line according to the first position information and the second position information, and further determine the evaluation information based on the first distance between the at least one first lane boundary line and the at least one second lane boundary line.
[0029] In this implementation, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line composed of the projection points corresponding to the first position points in the first image under the PV perspective and the position of the lane boundary line in the first image under the PV perspective, that is, a specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, and the realizability of the scheme is improved. In addition, the accuracy of the lane boundary line in the updated first map information can be obtained based on the position of the lane boundary line composed of the projection points and the position of the original lane boundary line in the first image under the PV perspective. Since the lane boundary line in the traffic environment is often long, and it is often difficult to accurately restore the lane boundary line in the actual environment in the map information of traffic scenes such as uphill and downhill, the accuracy of the updated first map information is evaluated from the dimension of the lane boundary line, more accurate evaluation information can be obtained, and the accuracy of the position information of the lane boundary line in the obtained updated first map information is improved.
[0030] In a possible implementation, the first device determines the evaluation information based on the first position information and the first image under the perspective view, which can include that the first device performs an image segmentation operation to obtain third position information based on the first image under the perspective view, the third position information indicating the position of at least one image region where the lane boundary line is located in the first image under the perspective view. For example, each first image region where the lane boundary line is located can be represented as a polygon, and the third position information can indicate the position of each first image region in the at least one first image region in the form of a polygon where the lane boundary line is located. The first device can determine the evaluation information according to the first position information and the third position information.
[0031] Optionally, the first device can also determine at least one second image region of the at least one first lane boundary line in the first image under the PV perspective according to the first position information, each second image region can be obtained by increasing the width of each first lane boundary line, for example, the width of the first lane boundary line can be increased by a preset width to obtain the second image region corresponding to each first lane boundary line. The first device can determine the intersection over union (IoU) between the at least one first image region and the at least one second image region, that is, determine the ratio between the intersection and the union of all first image regions and all second image regions; and then determine the evaluation information based on the intersection over union between the at least one first image region and the at least one second image region.
[0032] In this implementation, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line composed of the projection points corresponding to the first position points in the first image in the PV perspective and the position of the image region where the lane boundary line is located in the first image in the PV perspective, that is, another specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, and the implementation flexibility of the scheme is improved.
[0033] In a possible implementation, the first map information of the first traffic environment can be obtained based on second map information and third map information, where the second map information and the third map information are both map information in a bird's eye view (BEV) of the first traffic environment, the two-dimensional coordinates included in the second map information are in an absolute world coordinate system, the two-dimensional coordinates included in the third map information are in a relative world coordinate system, the second map information and the third map information are both obtained based on the first image and / or the point cloud data of the first traffic environment collected by the vehicle, and the origin of the relative world coordinate system is obtained based on the position of the vehicle. In this implementation, the first map information of the first traffic environment is obtained based on map information in two different coordinate systems, which is beneficial to obtaining more accurate map information.
[0034] In a possible implementation, the 3D information of the first traffic environment is obtained based on the second image in the perspective view of the first traffic environment and / or the point cloud data of the first traffic environment. In this implementation, it is specified that the 3D information of the first traffic environment is obtained based on which information, which is beneficial to providing the implementability of the scheme.
[0035] In a second aspect, the present application provides a method for obtaining map information, which can be used in the field of intelligent driving. In the method, a second device can obtain an image in a perspective view (PV) of a second traffic environment; based on the image in the perspective view, obtain map information of the second traffic environment, the map information of the second traffic environment including two-dimensional coordinates and a height of at least one third position point in a road element included in the second traffic environment, and the map information being generated by a machine learning model; and the second device can input at least one image in the perspective view of the second traffic environment (optionally, also including point cloud data of the second traffic environment) into the machine learning model that has performed a training operation, to obtain the map information of the second traffic environment generated by the machine learning model.
[0036] The training data of the machine learning model includes updated first map information of the first traffic environment, the updated first map information of the first traffic environment includes two-dimensional coordinates and a height of at least one first position point in road elements included in the first traffic environment, the height of the at least one first position point is obtained based on first map information of the first traffic environment and 3D information of the first traffic environment, the first map information includes two-dimensional coordinates of the at least one first position point, and the 3D information of the first traffic environment includes a height of at least one position region of an object in the first traffic environment.
[0037] In a possible implementation, the second position point is any one of the at least one first position point, and the height of the second position point is obtained based on a height of a target position region in the at least one position region corresponding to the second position point.
[0038] The updated first map information in the second aspect can be obtained by the first aspect and various possible implementation manners of the first aspect, the meanings of the terms in the second aspect and various possible implementation manners of the second aspect, and the beneficial effects brought by each possible implementation manner can be referred to the description of the various possible implementation manners in the first aspect, which will not be repeated here.
[0039] In a third aspect, the present application provides a map information processing apparatus, which can be used in the field of intelligent driving, and includes an acquisition module, a processing module, and the like. The acquisition module is configured to acquire first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in road elements included in the first traffic environment. The acquisition module is further configured to acquire three-dimensional (3D) information of the first traffic environment, the 3D information including a height of at least one position region, and an object in the first traffic environment including the at least one position region. The processing module is configured to obtain a height of the at least one first position point based on the first map information and the 3D information, and to update the first map information to include the two-dimensional coordinates of the at least one first position point and the height of the at least one first position point.
[0040] In a possible implementation, the second position point is any one of the at least one first position point, and the processing module is specifically configured to determine the height of the second position point based on a height of a target position region in the at least one position region, where the target position region is a position region in the at least one position region corresponding to the second position point.
[0041] In a possible implementation, the 3D information further includes two-dimensional coordinates of each position region, and the processing module is specifically further configured to determine the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each position region.
[0042] In a possible implementation, the processing module is specifically configured to determine, from the at least one location region, a target location region corresponding to the second location point based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region, and the semantic type of each location region.
[0043] In a possible implementation, the first map information is map information in a bird's eye view (BEV), and the processing apparatus of the map information further includes a determination module configured to determine, based on the first image in a perspective view (PV) of the first traffic environment, evaluation information corresponding to the updated first map information, the evaluation information indicating accuracy of the updated first map information.
[0044] In a possible implementation, the at least one first location point includes a location point on a lane boundary line in the first traffic environment, and the determination module is specifically configured to project the first location point into the first image in the perspective view based on the two-dimensional coordinates of the first location point and a height of the first location point to obtain first location information, the first location information indicating a position of a projection point corresponding to the first location point in the first image in the perspective view; and determine the evaluation information based on the first location information and the first image in the perspective view.
[0045] In a possible implementation, the determination module is specifically configured to perform a semantic recognition operation based on the first image in the perspective view to obtain second location information, the second location information including position information of a pixel point with a semantic type of a lane boundary line in the first image in the perspective view; and determine the evaluation information according to the first location information and the second location information.
[0046] In a possible implementation, the determination module is specifically configured to perform an image segmentation operation based on the first image in the perspective view to obtain third location information, the third location information indicating a position of an image region in which the lane boundary line is located in the first image in the perspective view; and determine the evaluation information according to the first location information and the third location information.
[0047] In a possible implementation, the first map information is obtained based on second map information and third map information, the second map information and the third map information are both map information in a bird's eye view (BEV) of the first traffic environment, the two-dimensional coordinates included in the second map information are in an absolute world coordinate system, the two-dimensional coordinates included in the third map information are in a relative world coordinate system, the second map information and the third map information are both obtained based on first image and / or point cloud data of the first traffic environment collected by the vehicle, and an origin of the relative world coordinate system is obtained based on a position of the vehicle.
[0048] In a possible implementation, the 3D information is obtained based on a second image in a perspective view of the first traffic environment and / or point cloud data of the first traffic environment.
[0049] The meanings of the terms in the third aspect of the present application and the various possible implementation manners of the third aspect, and the beneficial effects brought by each possible implementation manner, can refer to the description in the various possible implementation manners of the first aspect, which will not be repeated here.
[0050] In a fourth aspect, the present application provides a map information acquisition device, which can be used in the field of intelligent driving. The map information acquisition device comprises: an acquisition module configured to acquire an image under a perspective view (PV) of a second traffic environment; and a processing module configured to obtain map information of the second traffic environment based on the image under the perspective view, the map information comprising a two-dimensional coordinate and a height of at least one third position point in a road element included in the second traffic environment, the map information being generated by a machine learning model; wherein the training data of the machine learning model comprises updated first map information of a first traffic environment, the updated first map information of the first traffic environment comprising a two-dimensional coordinate and a height of at least one first position point in a road element included in the first traffic environment, the height of the at least one first position point being obtained based on first map information of the first traffic environment and three-dimensional (3D) information of the first traffic environment, the first map information comprising a two-dimensional coordinate of the at least one first position point, and the 3D information of the first traffic environment comprising a height of at least one position region of an object in the first traffic environment.
[0051] In a possible implementation manner, the second position point is any one of the at least one first position point, and the height of the second position point is obtained based on a height of a target position region corresponding to the second position point in the at least one position region.
[0052] The updated first map information in the fourth aspect of the present application can be obtained by the first aspect and the various possible implementation manners of the first aspect. The meanings of the terms in the fourth aspect and the various possible implementation manners of the fourth aspect, and the beneficial effects brought by each possible implementation manner, can refer to the description in the various possible implementation manners of the first aspect, which will not be repeated here.
[0053] In a fifth aspect, the embodiments of the present application provide a device, which comprises a processor and a memory, the processor being coupled to the memory, the memory being configured to store a program, and the processor being configured to execute the program in the memory, so that the device executes the method of the first aspect or the second aspect.
[0054] In a sixth aspect, the embodiments of the present application provide a vehicle, which comprises a processor and a memory, the processor being coupled to the memory, the memory being configured to store a program, and the processor being configured to execute the program in the memory, so that the device executes the method of the second aspect.
[0055] In a seventh aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. When the computer program is run on a computer, the computer program enables the computer to perform the method of the first aspect or the second aspect.
[0056] In an eighth aspect, an embodiment of the present application provides a computer program product, which includes a program. When the program is run on a computer, the program enables the computer to perform the method of the first aspect or the second aspect.
[0057] In a ninth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above methods. In a possible design, the chip system further includes a memory for storing the necessary program instructions and data of the terminal device or the communication device. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0058] The fifth aspect to the ninth aspect of the present application correspond to the first aspect or the various possible manners of the first aspect, and have corresponding beneficial effects. BRIEF DESCRIPTION OF DRAWINGS
[0059] FIG. 1 is a flow diagram of a method for processing map information according to an embodiment of the present application;
[0060] FIG. 2 is a schematic diagram of first map information according to an embodiment of the present application;
[0061] FIG. 3 is a schematic diagram of obtaining first map information of a first traffic environment according to an embodiment of the present application;
[0062] FIG. 4 is another flow diagram of a method for processing map information according to an embodiment of the present application;
[0063] FIG. 5 is another flow diagram of a method for processing map information according to an embodiment of the present application;
[0064] FIG. 6 is a schematic diagram of a first image region in which a lane boundary line is located in a first image of a perspective view according to an embodiment of the present application;
[0065] FIG. 7 is a schematic diagram of generating evaluation information according to an embodiment of the present application;
[0066] FIG. 8 is a schematic diagram of a beneficial effect according to an embodiment of the present application;
[0067] FIG. 9 is a schematic diagram of a first score according to an embodiment of the present application;
[0068] FIG. 10 is a flow diagram of a method for training a model according to an embodiment of the present application;
[0069] FIG. 11 is a flow diagram of a method for acquiring map information according to an embodiment of the present application;
[0070] FIG. 12 is a flow diagram of another method for acquiring map information according to an embodiment of the present application;
[0071] FIG. 13 is a block diagram of a map information processing device according to an embodiment of the present application;
[0072] FIG. 14 is a block diagram of a map information acquisition device according to an embodiment of the present application;
[0073] FIG. 15 is a block diagram of an apparatus according to an embodiment of the present application;
[0074] FIG. 16 is a block diagram of a vehicle according to an embodiment of the present application. DETAILED DESCRIPTION
[0075] Embodiments of the present application will be described below with reference to the accompanying drawings. It is obvious that the described embodiments are only a part of the embodiments of the present application, but not all of the embodiments of the present application. It is obvious for those skilled in the art that the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems as new application scenarios appear.
[0076] The terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, products or equipment.
[0077] In the embodiments of the present application, the indication can include direct indication and indirect indication, and can also include explicit indication and implicit indication. The information indicated by certain information (indication information described below) is referred to as to-be-indicated information. In the implementation process, there are many ways to indicate the to-be-indicated information, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or the index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, where the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be indicated only by a part of the to-be-indicated information, and the other part of the to-be-indicated information is known or agreed in advance. For example, the arrangement order of each information agreed in advance (for example, protocol predefined) can be used to indicate a specific information, thereby reducing the indication overhead to a certain extent. The present application does not limit the specific manner of indication. It can be understood that the indication information can be used to indicate the to-be-indicated information for the sender of the indication information, and the indication information can be used to determine the to-be-indicated information for the receiver of the indication information.
[0078] The method provided by the present application can be applied in the field of intelligent driving. For example, the method can be applied in a scenario requiring map information in the field of intelligent driving. For example, when the intelligent driving system in the vehicle performs trajectory planning, the map information corresponding to the surrounding traffic environment needs to be obtained. For another example, when the intelligent driving system in the vehicle provides a navigation function, the map information corresponding to the surrounding traffic environment needs to be obtained. The present application does not exhaustively list the scenarios requiring map information. For example, the vehicle can be a car, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, an entertainment vehicle, an amusement park vehicle, a trolley, a golf cart or a train, and the like, which are not particularly limited in the present application.
[0079] In the related art, the map information used is often two-dimensional (2D) map information of a traffic environment, which includes two-dimensional coordinates of position points in road elements in the traffic environment. However, in some traffic scenarios, such as uphill and downhill traffic scenarios, the 2D map information cannot accurately reflect the real traffic environment because the 2D map information does not include the height of each position point. To solve the foregoing problem, the present application discloses that a first device not only obtains first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment, but also obtains three-dimensional (3D) information of the first traffic environment, the 3D information including the height of at least one position region, an object in the first traffic environment including the at least one position region. Then, the first device can obtain the height of the at least one first position point based on the first map information and the 3D information, and update the first map information to include the two-dimensional coordinates of the at least one first position point and the height of the at least one first position point. That is, a generation scheme of map information including both the two-dimensional coordinates of a position point and the height of the position point is provided, which is beneficial to making the updated first map information more accurately reflect the real traffic environment.
[0080] Optionally, the updated first map information obtained by the method provided in the present application can be used as training data of a machine learning model (hereinafter referred to as a "first machine learning model" for convenience of distinction). The input of the first machine learning model can include an image in a perspective view (PV) of a traffic environment (hereinafter referred to as an "image in a PV view"), and the output of the first machine learning model can include map information including two-dimensional coordinates of position points in road elements included in the traffic environment and the height of the position points. That is, the updated first map information generated by the method provided in the present application is used to train the first machine learning model, and the trained machine learning model can directly obtain map information carrying the height based on the image in the PV view.
[0081] In combination with the foregoing description, the specific implementation process of the map information processing method provided in the present application will be described below. Please refer to FIG. 1, which is a flowchart of the map information processing method provided in an embodiment of the present application. The map information processing method provided in the embodiment of the present application can include the following steps.
[0082] 101. Obtain first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment.
[0083] The first traffic environment in the present application can also be referred to as a first road environment. The first map information of the first traffic environment can be understood as two-dimensional (2D) map information of the first traffic environment.
[0084] Alternatively, the first map information can be 2D map information in a bird eye view (BEV) (hereinafter referred to as "BEV view"). The BEV view in the present application can also be referred to as a top view. Alternatively, the first map information can also be map information in a perspective view (PV), and the like. The specific determination can be combined with the actual application scenario.
[0085] The first map information can indicate the position of a road element in the first traffic environment. The road element can be a road element on the ground. For example, the road element can include a lane boundary line. The lane boundary line in the present application can be a line that actually exists in a road and is used to divide different lanes. Alternatively, the road element can also include at least one of the following: a lane center line, a lane, a boundary line of a road, an intersection, a sidewalk, a turning indication line, or other road elements. The lane center line in the present application can be understood as a virtual central axis of a lane. The lane center line can be a line that does not actually exist in a road. The lane center line can be used to position a lane. For example, the first map information can include 2D coordinates of at least one first position point. Each first position point is a position point in a road element included in the first traffic environment. The position of the road element in the first traffic environment can be indicated by the 2D coordinates of the at least one first position point. Alternatively, the first map information further includes a connection relationship between different first position points in the at least one first position point. The at least one first position point can be converted into a line segment based on the connection relationship, so as to more accurately indicate the position of the road element in the first traffic environment.
[0086] The 2D coordinates of the first position point can include first coordinate information in the X-axis direction and the Y-axis direction. The X-axis and the Y-axis are perpendicular to each other. For example, the 2D coordinates of the first position point can be coordinate information in a first coordinate system. If the first coordinate system is a vehicle body coordinate system, the X-axis direction and the Y-axis direction in the vehicle body coordinate system can be the X-axis direction and the Y-axis direction. Alternatively, if the first coordinate system is a world coordinate system, which can also be referred to as an earth coordinate system, the X-axis direction can be consistent with the longitude direction of the earth coordinate system, and the Y-axis direction can be consistent with the latitude direction of the earth coordinate system. Alternatively, the first coordinate system can be coordinate information in other coordinate systems (such as a camera coordinate system or a radar sensor, etc.). The specific determination can be combined with the actual application scenario.
[0087] For a more intuitive understanding of the scheme, please refer to FIG. 2, which is a schematic diagram of the first map information provided by the embodiment of the application. In FIG. 2, the first map information is taken as an example of the map information in the BEV perspective, as shown in FIG. 2, the first map information includes the two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment, and the first map information can indicate the position of the road elements in the first traffic environment. In FIG. 2, the road elements in the first traffic environment include lane boundary lines, lanes, intersections, and road boundary lines, for example. It should be understood that the examples in FIG. 2 are only for the convenience of understanding the scheme and do not limit the scheme.
[0088] For example, the first map information of the first traffic environment can be obtained based on at least one image of the first traffic environment and / or point cloud data of the first traffic environment. The at least one image of the first traffic environment can be at least one second image in the PV perspective of the first traffic environment, for example, the second image in the PV perspective of the first traffic environment can be obtained by an optical sensor, such as a camera or an event camera, and the like. The aforementioned point cloud data can be obtained by an ultrasonic sensor, a laser radar sensor, a millimeter wave radar sensor, or other sensors capable of measuring point cloud data. It should be noted that the concept of "first image" will be used in the subsequent description, and the concept of "first image" will not be introduced here.
[0089] In one case, the first map information of the first traffic environment is obtained based on second map information and third map information, wherein the second map information and the third map information are both map information in the BEV perspective of the first traffic environment, that is, the second map information and the third map information both include the two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment, and the two-dimensional coordinates of each position point included in the second map information are in the absolute world coordinate system, and the two-dimensional coordinates of each position point included in the third map information are in the relative world coordinate system. For example, the X-axis direction of the absolute world coordinate system and the relative world coordinate system are consistent with the longitude direction in the earth coordinate system, and the Y-axis direction of the absolute world coordinate system and the relative world coordinate system are consistent with the latitude direction in the earth coordinate system, but the origins of the absolute world coordinate system and the relative world coordinate system are different.
[0090] For example, the second map information and the third map information are both based on at least one second image in a PV perspective of the first traffic environment and / or point cloud data of the first traffic environment collected by a sensor deployed on a vehicle (hereinafter referred to as "first vehicle" for convenience of description), and the origin of the relative world coordinate system can be obtained based on the position of the first vehicle. Optionally, the origin of the relative world coordinate system can be the starting position of the first vehicle. In other words, the origin of the relative world coordinate system can be the position at which the first vehicle needs to run in order to collect at least one second image in a PV perspective of the first traffic environment and / or point cloud data of the first traffic environment. The starting position of the first vehicle can be the position at which the first vehicle starts to run.
[0091] For example, the second map information is obtained by inputting at least one second image in a PV perspective of the first traffic environment and / or point cloud data of the first traffic environment into a second machine learning model, and the second map information generated by the second machine learning model includes two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment, and the two-dimensional coordinates include coordinate information in the X-axis direction and the Y-axis direction in the absolute world coordinate system. For example, the second machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, a neural network based on an attention mechanism, a multi-layer perceptron or other types of machine learning models, etc. For example, the second machine learning model can adopt a U-shaped coding and decoding network. It should be understood that the examples herein are only used to prove the feasibility of the scheme, and the specific model of the second machine learning model can be determined in combination with the actual application scenario.
[0092] For example, the third map information is obtained by inputting at least one second image in a PV perspective of the first traffic environment and / or point cloud data of the first traffic environment into a third machine learning model, and the third map information generated by the third machine learning model includes two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment, and the two-dimensional coordinates include coordinate information in the X-axis direction and the Y-axis direction in the relative world coordinate system. The specific form of the third machine learning model can refer to the description of the second machine learning model above, which will not be repeated here.
[0093] Optionally, if the two-dimensional coordinates included in the first map information are coordinate information in a vehicle body coordinate system, step 101 can include: the first device can convert the two-dimensional coordinates included in the second map information from the absolute world coordinate system to the vehicle body coordinate system to obtain updated second map information, the updated second map information including two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment in the vehicle body coordinate system; and the vehicle body coordinate system can be a vehicle body coordinate system of the first vehicle. The first device can also convert the two-dimensional coordinates included in the third map information from the relative world coordinate system to the vehicle body coordinate system to obtain updated third map information, the updated third map information including two-dimensional coordinates of at least one position point in the road elements included in the first traffic environment in the vehicle body coordinate system. The first device can determine the first map information of the first traffic environment according to the updated second map information and the updated third map information; and the first device can perform a fusion operation according to the updated second map information and the updated third map information to obtain the first map information of the first traffic environment.
[0094] For a more intuitive understanding of the present scheme, please refer to FIG. 3, which is a schematic diagram provided by an embodiment of the present application for obtaining the first map information of the first traffic environment. As shown in FIG. 3, the second map information in the absolute world coordinate system can be obtained, and the second map information can be converted from the absolute world coordinate system to the vehicle body coordinate system to obtain updated second map information. The third map information in the relative world coordinate system can be obtained, and the third map information can be converted from the relative world coordinate system to the vehicle body coordinate system to obtain updated third map information. The updated second map information and the updated third map information can be fused to obtain the first map information in the vehicle body coordinate system. It should be understood that the examples in FIG. 3 are only for the convenience of understanding the present scheme and are not used to limit the present scheme.
[0095] Alternatively, if the two-dimensional coordinates included in the first map information are coordinate information in the absolute world coordinate system, step 101 can include: the first device can convert the two-dimensional coordinates included in the third map information from the relative world coordinate system to the absolute world coordinate system to obtain updated third map information. The first device can determine the first map information of the first traffic environment based on the second map information and the updated third map information; and the first device can perform a fusion operation according to the second map information and the updated third map information to obtain the first map information of the first traffic environment.
[0096] In the embodiments of the present application, the first map information of the first traffic environment is obtained based on map information in two different coordinate systems, which is beneficial to obtain more accurate map information.
[0097] In another case, the first map information of the first traffic environment can be directly obtained based on the at least one second image under the PV perspective of the first traffic environment and / or the point cloud data of the first traffic environment; for example, the first map information of the first traffic environment can be obtained by inputting the at least one second image under the PV perspective of the first traffic environment and / or the point cloud data of the first traffic environment into a machine learning model, and obtaining the first map information generated by the machine learning model; the specific form of the machine learning model can refer to the description of the second machine learning model above, which will not be repeated here.
[0098] 102. Obtain 3D information of the first traffic environment, the three-dimensional information comprising a height of at least one location region, the object in the first traffic environment comprising the at least one location region.
[0099] Exemplarily, the 3D information of the first traffic environment can comprise second coordinate information of each location region in the at least one location region in the first traffic environment under a second coordinate system, and the second coordinate information of each location region can comprise two-dimensional coordinates and a height of each location region. The height in the present application can also be referred to as an elevation. The two-dimensional coordinates of each location region can comprise position parameters in the X-axis direction and the Y-axis direction of the second coordinate system. The height of each location region can be understood as a position parameter in the Z-axis direction of the second coordinate system. The Z-axis direction can be perpendicular to the plane formed by the X-axis direction and the Y-axis direction. In other words, the second coordinate information of each location region can comprise position parameters of each location region in the X-axis direction, the Y-axis direction and the Z-axis direction of the second coordinate system.
[0100] Exemplarily, each location region can be represented as a cube, and the at least one location region can be understood as a division of the object in the first traffic environment into the at least one location region. Exemplarily, the 3D information of the first traffic environment can be obtained based on the at least one second image under the PV perspective of the first traffic environment and / or the point cloud data of the first traffic environment. Alternatively, the 3D information of the first traffic environment can be obtained based on a plurality of second images under a perspective view (PV) of the first traffic environment and / or the point cloud data of the first traffic environment; for example, the plurality of second images under the PV perspective are obtained by a plurality of first sensors of the first vehicle, and the plurality of first sensors can comprise at least two of the following: a camera in the left front, a camera in the front, a camera in the right front, a camera in the left rear, a camera in the rear, and a camera in the right rear, etc. The specific images under the PV perspective can be determined in combination with the actual application scenario.
[0101] Exemplarily, the object in the first traffic environment can include a road element in the first traffic environment, and the object in the first traffic environment can also include other objects other than the road element, for example, the other objects other than the road element in the first traffic environment can include a traffic signal, a street lamp, a tree, a building, a fire hydrant, a triangular board, a cone barrel, a warning column, a crash barrel, a construction board, a guide board or other types of static obstacles, and the like. Optionally, the other objects other than the road element in the first traffic environment can also include dynamic obstacles in the first traffic environment, for example, the aforementioned dynamic obstacles can include a vehicle, an electric vehicle, a bicycle, a pedestrian, an animal or other dynamic obstacles, and the like. It should be understood that the examples herein are only for the convenience of understanding the present solution, and the specific information contained in the object in the first traffic environment can be determined in combination with the actual application scenario.
[0102] Optionally, the 3D information of the first traffic environment can be understood as being obtained by performing three-dimensional modeling on the object in the first traffic environment based on at least one second image under the PV perspective of the first traffic environment and / or point cloud data of the first traffic environment, and each position region can also be understood as a voxel. Exemplarily, the at least one second image under the PV perspective of the first traffic environment and / or the point cloud data of the first traffic environment can be input into a machine learning model to obtain the 3D information of the first traffic environment generated by the machine learning model. The specific form of the aforementioned machine learning model can be referred to the description of the second machine learning model described above, which is not repeated here.
[0103] For example, the 3D information of the first traffic environment can be general obstacle detection (GOD) data of the first traffic environment, and the GOD data of the first traffic environment is obtained based on at least one second image under the PV perspective of the first traffic environment and / or point cloud data of the first traffic environment. In the embodiments of the present application, it is clear that the 3D information of the first traffic environment is obtained based on which information, which is beneficial to provide the realizability of the present solution.
[0104] Optionally, the second coordinate system can be a vehicle body coordinate system, for example, the at least one second image under the PV perspective of the first traffic environment and / or the point cloud data of the first traffic environment is obtained by a sensor deployed in the first vehicle, and the vehicle body coordinate system can be a vehicle body coordinate system of the first vehicle. Alternatively, the second coordinate system can also be a world coordinate system or other types of coordinate systems (such as a camera coordinate system or a radar sensor, etc.), which can be determined in combination with the actual application scenario.
[0105] Optionally, the position parameters of each first position point in the X-axis direction and the Y-axis direction of the first coordinate system can be real numbers, in other words, the position parameters of each first position point in the X-axis direction and the Y-axis direction of the first coordinate system can not be integers; and the position parameters of each position region in the X-axis direction, the Y-axis direction and the Z-axis direction of the second coordinate system are all integers.
[0106] For example, the first coordinate information of a certain first position point can be (5.6, 7.8), representing that the position parameter of the first position point in the X-axis direction of the first coordinate system is 5.6, and the position parameter of the first position point in the Y-axis direction of the first coordinate system is 7.8; and the second coordinate information of a certain position region can be (6, 8, 2), representing that the position parameter of the position region in the X-axis direction of the second coordinate system is 3, the position parameter of the position region in the Y-axis direction of the second coordinate system is 4, and the position parameter of the position region in the Z-axis direction is 2. It should be understood that the examples herein are only for the convenience of understanding the difference between the first coordinate information of the first position point and the second coordinate information of the position region.
[0107] 103. Based on the first map information and the 3D information, the height of at least one first position point is obtained, and the updated first map information includes the two-dimensional coordinates of at least one first position point and the height of at least one first position point.
[0108] For example, after obtaining the first map information of the first traffic environment and the 3D information of the first traffic environment, the first device can obtain the height of each first position point in the at least one first position point based on the first map information of the first traffic environment and the 3D information of the first traffic environment, that is, obtain the updated first map information, and the updated first map information includes the two-dimensional coordinates and the height of each first position point in the at least one first position point.
[0109] Exemplarily, in an implementation, the present application refers to any one of the at least one first position point as a "second position point", and step 103 can include that the first device can determine the height of the second position point based on the height of a target position area in the at least one position area, the target position area being a position area in the at least one position area corresponding to the second position point, in other words, the foregoing step can be understood as that the height of the second position point is absorbed; the first device can repeatedly perform the foregoing step at least once to obtain the height of each of the at least one first position point. In the embodiment of the present application, for any one of the at least one first position point (i.e., the second position point), the height of the target position area in the at least one position area corresponding to the first position point is obtained, and then the height of the second position point is determined based on the height of the target position area, that is, each of the at least one first position point is traversed, and the height of each of the at least one first position point is obtained one by one, that is, the process of obtaining the height of the at least one first position point is more finely managed, which is beneficial to improve the accuracy of the height of each of the at least one first position point obtained.
[0110] Exemplarily, the first device can determine at least one target position area corresponding to the second position point from the at least one position area based on the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each of the at least one position area; and then the first device can determine the height of the second position point based on the height of the at least one target position area. In the embodiment of the present application, the two-dimensional coordinates of the first position point and the two-dimensional coordinates of the position area in the 3D information are used to find the target position area corresponding to each of the at least one first position point, and a simple and effective method for determining the target position area is provided.
[0111] Optionally, the first device can determine the distance between the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each of the at least one position area, and select at least one target position area closest to the two-dimensional coordinates of the second position point from the at least one position area, the distance between the two-dimensional coordinates of each of the at least one target position area and the two-dimensional coordinates of the second position point can be consistent, for example, if the at least one target position area includes at least two target position areas, the two-dimensional coordinates of different target position areas of the at least two target position areas can be consistent, and the difference lies in the different heights. Alternatively, the first device can round up the position parameters in the X-axis direction and the position parameters in the Y-axis direction included in the two-dimensional coordinates of the second position point to obtain the two-dimensional coordinates of each of the at least one target position area. Alternatively, the first device can round down the position parameters in the X-axis direction and the position parameters in the Y-axis direction included in the two-dimensional coordinates of the second position point to obtain the two-dimensional coordinates of each of the at least one target position area, and the specific implementation mode can be determined in combination with the actual application scenario.
[0112] For example, if the at least one target location area includes only one target location area, the first device can directly determine the height of the one target location area as the height of the second location point. If the at least one target location area includes at least two target location areas, in one implementation, the first device can determine the minimum height of the at least two heights corresponding to the at least two target location areas as the height of the second location point, since the second location point can be a location point on the road surface, it is feasible to directly determine the minimum height as the height of the second location point; or in another implementation, the first device can determine the average height of the at least two heights corresponding to the at least two target location areas as the height of the second location point, etc.
[0113] Optionally, the first device determines the target location area corresponding to the second location point from the at least one location area based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location area, can include: the first device determines the target location area corresponding to the second location point from the at least one location area based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location area in the at least one location area, and the semantic type of each location area in the at least one location area.
[0114] For example, the semantic type of a location area can be the semantic type of a lane boundary line, a boundary line of a road, a sidewalk, an intersection, a turning indication line, a traffic signal, a street lamp, a tree, a building, a fire hydrant, a triangular board, a cone barrel, a warning column, an anti-collision barrel, a construction board, a guide board, or other objects that can appear in a traffic environment, etc. Optionally, if the objects in the first traffic environment also include dynamic obstacles in the first traffic environment, the semantic type of a location area can also be the semantic type of a vehicle, an electric vehicle, a bicycle, a pedestrian, an animal, or other dynamic obstacles, etc. The examples herein are only for easy understanding of the scheme.
[0115] Exemplarily, the first device can determine distances between the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each position region, select at least one candidate position region closest to the two-dimensional coordinates of the second position point from the at least one position region, and the distance between the two-dimensional coordinates of each candidate position region in the at least one candidate position region and the two-dimensional coordinates of the second position point can be uniform, for example, if the at least one candidate position region includes at least two candidate position regions, the two-dimensional coordinates of different candidate position regions in the at least two candidate position regions can be uniform, and the difference is that the heights are different. In one case, the first device can determine one target position region consistent with the semantic type of the second position point from the at least one candidate position region based on the semantic type of each candidate position region. In another case, optionally, if each of the at least one first position point is a position point on the lane boundary line, the first device can also determine one target position region with the semantic type of the lane boundary line from the at least one candidate position region based on the semantic type of each candidate position region. Further, the first device can determine the height of the one target position region as the height of the second position point.
[0116] For a more intuitive understanding of the present scheme, please refer to FIG. 4, which is another flowchart of the method for processing map information provided by the embodiment of the present application. The contents in FIG. 4 can be understood in combination with the above description of FIG. 3, and the repeated parts will not be described here. The first map information of the first traffic environment in the vehicle coordinate system and the 3D information of the first traffic environment in the vehicle coordinate system are obtained. The first map information of the first traffic environment includes the two-dimensional coordinates of at least one first position point, and the 3D information of the first traffic environment includes the height of at least one position region. The height of the target position region corresponding to each first position point can be obtained from the height of the at least one position region. Further, based on the height of the target position region corresponding to each first position point, the height of each first position point is determined, that is, the high value adsorption of each first position point is completed, and the updated first map information of the first traffic environment is obtained. It should be understood that the example in FIG. 4 is only for the convenience of understanding the present scheme, and is not used to limit the present scheme.
[0117] In the embodiment of the present application, in the process of determining the target location area corresponding to any one first location point, not only the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area are used, but also the semantic type of each location area is used. Since matching is only based on the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area, one first location point can correspond to multiple location areas, the two-dimensional coordinates of the multiple location areas are consistent, but the heights are inconsistent, and then the target location area corresponding to the first location point can be further determined based on the semantic type of each location area, which is beneficial to accurately find the target location area corresponding to each first location point, and then is beneficial to obtain the accurate height of each first location point, so as to improve the accuracy of the updated first map information.
[0118] In another implementation manner, the step 103 can include that the first device can also input the first map information of the first traffic environment and the 3D information of the first traffic environment into the machine learning model to obtain the updated first map information generated by the machine model.
[0119] Optionally, the first vehicle can collect continuous multiple sets of images in the PV perspective and / or continuous multiple point cloud data of the first traffic environment during the travel, each set of images in the PV perspective can include at least one second image in the PV perspective, and then the first device can obtain continuous multiple first map information and continuous multiple 3D information, wherein each first map information in the continuous multiple first map information is obtained based on one set of images in the PV perspective of the first traffic environment and / or one point cloud data of the first traffic environment, and each 3D information in the continuous multiple 3D information is obtained based on one set of images in the PV perspective of the first traffic environment and / or one point cloud data of the first traffic environment. Correspondingly, the first device can obtain one updated first map information based on one first map information and one 3D information, and then the first device can repeatedly execute the steps 101 to 103 based on the continuous multiple first map information and the continuous multiple 3D information, to obtain continuous multiple updated first map information. Since there are overlapping location areas in two adjacent updated first map information in the continuous multiple updated first map information, the continuous multiple updated first map information can include repeated first location points, in other words, the continuous multiple updated first map information can include at least two two-dimensional coordinates and heights of the repeated first location points.
[0120] The first device can perform a smoothing operation based on at least two heights of the repeated first position point in the plurality of continuous updated first map information, to obtain a final height of the repeated first position point; and can perform a smoothing operation based on at least two two-dimensional coordinates of the repeated first position point, to obtain a final two-dimensional coordinate of the repeated first position point.
[0121] For example, the first device can average the at least two heights of the first position point to obtain a final height of the first position point; and can average at least two position parameters of the first position point in the X-axis direction of the first coordinate system to obtain a final position parameter of the first position point in the X-axis direction of the first coordinate system; and average at least two position parameters of the first position point in the Y-axis direction of the first coordinate system to obtain a final position parameter of the first position point in the Y-axis direction of the first coordinate system. For another example, the first device can determine a median value of the at least two heights of the first position point as a final height of the first position point; and can determine a median value of at least two position parameters of the first position point in the X-axis direction of the first coordinate system as a final position parameter of the first position point in the X-axis direction of the first coordinate system; and determine a median value of at least two position parameters of the first position point in the Y-axis direction of the first coordinate system as a final position parameter of the first position point in the Y-axis direction of the first coordinate system. The specific smoothing method can be determined according to actual application scenarios, and is not limited herein.
[0122] Optionally, based on the description of the embodiment corresponding to FIG. 1, refer to FIG. 5, which is another flowchart of the method for processing map information provided by the embodiment. The method for processing map information provided by the embodiment can include:
[0123] 501. Obtain first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment.
[0124] 502. Obtain three-dimensional information of the first traffic environment, the three-dimensional information including a height of at least one position region, and objects in the first traffic environment including the at least one position region.
[0125] 503. Obtain a height of the at least one first position point based on the first map information and the three-dimensional information, and update the first map information to include the two-dimensional coordinates of the at least one first position point and the height of the at least one first position point.
[0126] The specific implementation of the first device performing steps 501 to 503 and the meaning of the terms in steps 501 to 503 can be referred to the description in the embodiment corresponding to FIG. 1, which will not be repeated here.
[0127] 504. Determine, based on the first image in the perspective view of the first traffic environment, evaluation information corresponding to the updated first map information, the evaluation information indicating accuracy of the updated first map information.
[0128] Step 504 is an optional step. In one case, the at least one first position point comprises a position point in a lane boundary line included in the first traffic environment, in other words, the at least one first position point can all be position points in a lane boundary line included in the first traffic environment, and step 504 can comprise: projecting, by the first device, each of the at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the at least one first position point, to obtain first position information, the first position information can comprise a position of each of at least one projection point corresponding to each of the at least one first position point in the first image in the perspective view, the at least one projection point can be understood as a projection point obtained by projecting a position point on a lane boundary line in a BEV view into a first image in a PV view; and further determining, by the first device, the evaluation information based on the first position information and the first image in the perspective view.
[0129] In the embodiments of the present application, a scheme is proposed for projecting a first position point in a BEV view into a first image in a PV view based on the two-dimensional coordinates and the height of the first position point in the updated first map information, to obtain position information of a projection point, and further evaluating the accuracy of the updated first map information based on the position information of the projection point. Since the first image in the PV view of the first traffic environment can accurately reflect the first traffic environment, the manner of projecting the first position point in the BEV view into the first image in the PV view can maximize the use of the first image in the PV view, and is conducive to obtaining more accurate evaluation information.
[0130] Exemplarily, the first device can project each of the at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the at least one first position point, the first image in the perspective view, and the intrinsic parameters of a camera for capturing the first image in the perspective view, to obtain first position information.
[0131] Next, the specific implementation manner of the first device obtaining the evaluation information based on the first position information is introduced. In one implementation manner, the first device performs a semantic recognition operation based on the first image in the perspective view to obtain second position information, and the second position information includes position information of pixel points of a semantic type of lane boundary line in the first image in the perspective view. Optionally, the first device can input the first image in the perspective view into a machine learning model, and perform target detection on the lane boundary line in the first image in the perspective view through the machine learning model. In other words, the machine learning model is used to identify the pixel points of the semantic category of the lane boundary line in the first image in the perspective view, and the second position information generated by the machine learning model is obtained. For example, the machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, a neural network based on an attention mechanism, or other types of neural networks, etc. The specific implementation can be determined according to actual application scenarios.
[0132] For example, the second position information can include coordinate information of the pixel points of the semantic type of the lane boundary line in the first image in the perspective view.
[0133] Optionally, the first position information can indicate position information of at least one first lane boundary line composed of projection points obtained by projecting first position points on the lane boundary line in the BEV view to the first image in the PV view, and the second position information can indicate position information of at least one second lane boundary line in the first image in the PV view. For example, the first device can determine a first distance between the at least one first lane boundary line and the at least one second lane boundary line according to the first position information and the second position information, and then determine the evaluation information based on the first distance between the at least one first lane boundary line and the at least one second lane boundary line.
[0134] For example, the first device can determine a second distance between each first lane boundary line and the nearest one of the at least one second lane boundary line, that is, at least one second distance corresponding to the at least one first lane boundary line is obtained. For example, the distance can be obtained by calculating a cosine distance, an Euclidean distance, a Mahalanobis distance, an L1 distance, or other distance calculation algorithms between lines. Then, the first device can obtain the first distance based on the at least one second distance. For example, the first distance can be an average value of the at least one second distance, the first distance can be a median value of the at least one second distance, the first distance can be a maximum value of the at least one second distance, etc. The specific implementation can be determined according to actual application scenarios.
[0135] In one case, the first distance can be directly determined as the evaluation information, wherein the smaller the first distance is, the higher the accuracy of the updated first map information is, and the larger the first distance is, the lower the accuracy of the updated first map information is. In another case, the evaluation information can include a first score indicating the accuracy of the updated first map information, wherein the higher the first score is, the higher the accuracy of the updated first map information is, and the lower the first score is, the lower the accuracy of the updated first map information is. The first score can be obtained based on the first distance, wherein the smaller the first distance is, the higher the first score is, representing the higher the accuracy of the updated first map information is, and the larger the first distance is, the lower the first score is, representing the lower the accuracy of the updated first map information is.
[0136] Further, for example, after the first device determines the at least one first lane boundary based on the first position information and determines the at least one second lane boundary based on the second position information, the first device can first pair the at least one first lane boundary and the at least one second lane boundary to obtain a second distance between each first lane boundary and the nearest second lane boundary. For another example, in this application, any one of the at least one first lane boundary is referred to as a target lane boundary. After the first device determines the at least one first lane boundary based on the first position information and determines the at least one second lane boundary based on the second position information, the first device can calculate a distance between the target lane boundary and each second lane boundary of the at least one second lane boundary to obtain at least one distance corresponding to the at least one second lane boundary. The smallest distance of the at least one distance is determined as a second distance between the target lane boundary and the nearest second lane boundary of the at least one second lane boundary. The first device can repeat the above steps at least once to obtain the second distance between each first lane boundary and the nearest second lane boundary, and the like. It should be understood that the above examples are only used to prove the feasibility of the present scheme, and the specific way of calculating the second distance can be determined in combination with the actual application scenario.
[0137] In the embodiments of the present application, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line composed of the projection points corresponding to the first position points in the first image under the PV perspective and the position of the lane boundary line in the first image under the PV perspective, that is, a specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, and the realizability of the scheme is improved. In addition, the accuracy of the lane boundary line in the updated first map information can be obtained based on the position of the lane boundary line composed of the projection points and the position of the original lane boundary line in the first image under the PV perspective. Since the lane boundary line in the traffic environment is often long, and it is often difficult to accurately restore the lane boundary line in the actual environment in the map information of traffic scenes such as uphill and downhill, the accuracy of the updated first map information is evaluated from the dimension of the lane boundary line, which can obtain more accurate evaluation information and is beneficial to improving the accuracy of the position information of the lane boundary line in the obtained updated first map information.
[0138] In another implementation manner, the first device performs an image segmentation operation based on the first image under the perspective view to obtain third position information, the third position information indicating the position of the lane boundary line in the first image under the perspective view. Illustratively, the third position information indicates the position of at least one first image region where the lane boundary line is located in the first image under the perspective view. Each first image region where the lane boundary line is located can be represented as a polygon, and the third position information can indicate the position of each first image region in at least one first image region in the form of a polygon where the lane boundary line is located. Optionally, the first device can input the first image under the perspective view into a machine learning model, perform image segmentation on the first image under the perspective view through the machine learning model, and obtain third position information generated by the machine learning model. For example, the machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, a neural network based on an attention mechanism, or other types of neural networks, which can be determined in combination with actual application scenarios. Then, the first device can determine the evaluation information according to the first position information and the third position information.
[0139] For a more intuitive understanding of the scheme, please refer to FIG. 6, which is a schematic diagram of a first image region where a lane boundary line is located in a first image in a perspective view according to an embodiment of the present application. FIG. 6 is a schematic diagram obtained by performing image segmentation on the first image in the perspective view of the first traffic environment. As shown in FIG. 6, the first image in the perspective view is divided into four regions. The image region in pure black represents a background region in the first image. The background region can include other objects in the first traffic environment other than road elements, such as vehicles, blue sky, or buildings, etc. The image region in white represents a fence at the road boundary. The image region in light gray represents a road. The image region in dark gray represents a lane boundary line. That is, the dark gray region in FIG. 6 represents at least one first image region where the lane boundary line is located in the first image in the perspective view. FIG. 6 includes multiple first image regions. Since the original FIG. 6 has been subjected to grayscale processing, the color between the pure black region and the dark gray region is somewhat similar. In order to distinguish between the two different colors, the background region is marked in the pure black region in FIG. 6. An arrow indicates a certain first image region among the multiple first image regions. It should be understood that the example in FIG. 6 is only for the convenience of understanding the scheme and does not limit the scheme.
[0140] Optionally, the first device can further determine, according to the first position information, at least one second image region of the at least one first lane boundary line in the first image in the PV view. Each second image region can be obtained by increasing the width of each first lane boundary line. For example, the width of the first lane boundary line can be increased by a predetermined width to obtain a second image region corresponding to each first lane boundary line. The first device can determine the intersection over union (IoU) between the at least one first image region and the at least one second image region, that is, determine the ratio between the intersection and the union of all first image regions and all second image regions. Then, based on the intersection over union between the at least one first image region and the at least one second image region, the evaluation information is determined.
[0141] In one case, the aforementioned intersection-over-union ratio can be directly determined as the evaluation information, wherein the greater the intersection-over-union ratio between the at least one first image region and the at least one second image region, the higher the accuracy of the updated first map information, and the smaller the intersection-over-union ratio between the at least one first image region and the at least one second image region, the lower the accuracy of the updated first map information. In another case, the evaluation information can include a second score indicating the accuracy of the updated first map information, and the higher the second score, the higher the accuracy of the updated first map information, and the lower the second score, the lower the accuracy of the updated first map information. The aforementioned second score can be obtained based on the intersection-over-union ratio between the at least one first image region and the at least one second image region, and the greater the aforementioned intersection-over-union ratio, the higher the second score, representing the higher the accuracy of the updated first map information, and the smaller the aforementioned intersection-over-union ratio, the lower the second score, representing the lower the accuracy of the updated first map information.
[0142] In the embodiments of the present application, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line composed of the projection points corresponding to the first position points in the first image under the PV perspective and the position of the image region where the lane boundary line is located in the first image under the PV perspective, that is, another specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, and the implementation flexibility of the present scheme is improved.
[0143] In another implementation manner, the first device determines the second distance between each first lane boundary line and the nearest second lane boundary line according to the first position information and the second position information, and then obtains the first distance based on the at least one second distance corresponding to the at least one first lane boundary line; and determines the intersection-over-union ratio between the at least one first image region and the at least one second image region according to the first position information and the second position information. The evaluation information includes the first distance and the intersection-over-union ratio between the at least one first image region and the at least one second image region; or the first device further obtains the first score based on the first distance and the second score based on the intersection-over-union ratio between the at least one first image region and the at least one second image region, and the evaluation information includes the first score and the second score.
[0144] For a more intuitive understanding of the scheme, please refer to FIG. 7, which is a schematic diagram for generating evaluation information provided by an embodiment of the present application. As shown in FIG. 7, the first device projects the first position point on the lane boundary line in the first map information to the first image in the perspective view based on the two-dimensional coordinates and the height of the first position point on the lane boundary line in the updated first map information in the BEV view, to obtain the first position information of the projection point. The first position information can indicate the position of at least one first lane boundary line composed of the projection point obtained by projecting the first position point on the lane boundary line in the BEV view to the first image in the perspective view. The first device performs target detection on the first image in the perspective view to obtain second position information, which indicates the position of at least one second lane boundary line in the first image in the perspective view. The first device obtains the first distance between the at least one first lane boundary line and the at least one second lane boundary line based on the first position information and the second position information. The first device can also perform image segmentation on the first image in the perspective view to obtain third position information, which indicates the position of each first image region in at least one first image region where the lane boundary line is located in the first image in the perspective view. The first device can also determine at least one second image region in the first image in the PV view according to the first position information. Then, the intersection over union between the at least one first image region and the at least one second image region is determined. The evaluation information can include the first distance and the intersection over union. It should be understood that the example in FIG. 7 is only for the convenience of understanding the scheme and is not used to limit the scheme.
[0145] Optionally, the first device can also filter the updated first map information based on the evaluation information, and then delete the updated first map information that does not meet the filtering conditions. For example, the filtering conditions can include: the first distance is less than or equal to a distance threshold, and / or, the first score is greater than or equal to a first score threshold, and / or, the IoU between the at least one first image region and the at least one second image region is greater than or equal to a preset ratio, and / or, the second score is greater than or equal to a second score threshold.
[0146] In another implementation, after obtaining the first position information, the first device can determine, based on the first position information, a position of at least one first lane boundary line composed of the projection points in the first image in the PV perspective, and then can display the first image in the PV perspective containing the display of the at least one first lane boundary line, in other words, the first image in the PV perspective is added with the at least one first lane boundary line, and the evaluation information can be the first image in the PV perspective containing the display of the at least one first lane boundary line; then the technician can determine the accuracy of the updated first map information by visually checking the fitting degree between the at least one first lane boundary line and the lane boundary line in the first image in the PV perspective.
[0147] In another case, if the updated first map information includes at least one first position point including not only the position point in the lane boundary line included in the first traffic environment but also the position point in other road elements included in the first traffic environment, the step 504 can further include: the first device screening the at least one first position point included in the updated first map information to obtain screened at least one first position point, the screened at least one first position point including only the position point of the lane boundary line; the first device projecting each of the screened at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the screened at least one first position point to obtain the first position information, the first position information can include the position of each of the at least one projection point corresponding to the screened at least one first position point in the first image in the perspective view, the at least one projection point can be understood as the projection point obtained by projecting the position point on the lane boundary line in the BEV perspective into the first image in the PV perspective; and then the first device can determine the evaluation information based on the first position information and the first image in the perspective view.
[0148] Exemplarily, the first device can project each of the screened at least one first position point into the first image in the perspective view based on the two-dimensional coordinates and the height of each of the screened at least one first position point, the first image in the perspective view, and the intrinsic parameters of the camera for shooting the first image in the perspective view to obtain the first position information.
[0149] The specific implementation of the first device determining the evaluation information based on the first position information and the first image in the perspective view can refer to the description above, which will not be repeated here.
[0150] In the embodiments of the present application, since the image under the PV perspective of the first traffic environment can accurately reflect the situation of the first traffic environment, and the updated first map information is used to more accurately reflect the first traffic environment, it is feasible to determine the accuracy of the updated first map information by means of the image under the PV perspective of the first traffic environment, and the image under the PV perspective of the first traffic environment is easy to obtain, thereby providing an easy-to-implement detection scheme for the accuracy of the updated first map information.
[0151] In order to more intuitively understand the beneficial effects brought by the method provided by the present application, please refer to FIG. 8, which is a schematic diagram of the beneficial effects provided by the embodiments of the present application, and FIG. 8 includes two sub-diagrams, the upper sub-diagram of FIG. 8 shows the two-dimensional coordinates of the position points on the lane boundary line in the map information under the BEV perspective of the traffic environment, and the aforementioned map information does not include the height of the position points on the lane boundary line, and the position of the lane boundary line composed of the projection points of the aforementioned position points on the lane boundary line projected into the image under the PV perspective; 100 meters in the upper sub-diagram of FIG. 8 represent the position of the projection point obtained by projecting the position point on the lane boundary line 100 meters away from the current position in the aforementioned map information into the image under the PV perspective, and 200 meters represent the position of the projection point obtained by projecting the position point on the lane boundary line 200 meters away from the current position in the aforementioned map information into the image under the PV perspective. The lower sub-diagram of FIG. 8 shows the two-dimensional coordinates and height of the first position points on the lane boundary line in the updated first map information under the BEV perspective of the traffic environment, and the position of the first lane boundary line composed of the projection points of the aforementioned position points on the lane boundary line projected into the image under the PV perspective; 100 meters in the upper sub-diagram of FIG. 8 represent the position of the projection point obtained by projecting the first position point on the lane boundary line 100 meters away from the current position in the updated first map information into the image under the PV perspective, and 200 meters represent the position of the projection point obtained by projecting the first position point on the lane boundary line 200 meters away from the current position in the updated first map information into the image under the PV perspective. By comparing the upper sub-diagram and the lower sub-diagram of FIG. 8, it can be seen that in the uphill scene of FIG. 8, since the updated first map information obtained based on the method provided by the present application carries the height of each first position point, the first lane boundary line obtained based on the updated first map information is more consistent with the position of the lane boundary line in the first image under the perspective view, and it should be understood that the examples in FIG. 8 are only for the convenience of understanding the present scheme, and are not used to limit the present scheme.
[0152] For a more intuitive understanding of the scheme, please refer to FIG. 9, which is a schematic diagram of the first score provided by an embodiment of the present application. After obtaining a plurality of updated first map information by the method provided by the present application, the first score of each updated first map information in the plurality of updated first map information is obtained. The vertical coordinate of FIG. 9 represents the quantity of the updated first map information corresponding to each first score, and the horizontal coordinate of FIG. 9 represents the first score. The higher the first score corresponding to a certain updated first map information is, the more consistent the at least one first lane boundary line obtained based on the updated first map information is with the at least one second lane boundary line in the image of the PV perspective. When the first score corresponding to a certain updated first map information is 1, it means that the at least one first lane boundary line obtained based on the updated first map information is completely consistent with the at least one second lane boundary line in the image of the PV perspective. When the first score corresponding to a certain updated first map information is greater than or equal to 0, it means that the first distance between the at least one first lane boundary line obtained based on the updated first map information and the at least one second lane boundary line in the image of the PV perspective is less than or equal to the distance threshold, that is, the accuracy of the updated first map information is qualified. As shown in FIG. 9, the accuracy of most of the updated first map information in the plurality of updated first map information obtained by the method provided by the present application is qualified. It should be understood that the example in FIG. 9 is only for the convenience of understanding the scheme and does not limit the scheme.
[0153] In an embodiment of the present application, not only the first map information of the traffic environment including the two-dimensional coordinates of at least one first position point in the road element included in the traffic environment is obtained, but also the 3D information of the traffic environment is obtained, the 3D information including the height of at least one position region included in the object in the traffic environment. Then, based on the foregoing first map information and 3D information, the height of at least one first position point is obtained, and the updated first map information can include the two-dimensional coordinates and the height of at least one position point. A generation scheme of map information containing the two-dimensional coordinates and the height of the position point is provided, which is conducive to making the updated first map information more accurately reflect the real traffic environment.
[0154] Optionally, if the updated first map information obtained by the method provided by the present application is also used as the training data of the first machine learning model, the present application further provides a specific implementation process of the training stage and the application stage of the first machine learning model. The training stage and the application stage of the first machine learning model are described as follows. Please refer to FIG. 10, which is a flowchart of a model training method provided by an embodiment of the present application. The model training method provided by the present application can include:
[0155] 1001、inputting the image under the perspective view of the first traffic environment into the first machine learning model to obtain predicted map information generated by the first machine learning model, the predicted map information comprising predicted two-dimensional coordinates of at least one first position point and a predicted height of the at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment.
[0156] Illustratively, the training device can input the image under the perspective view of the first traffic environment into the first machine learning model to obtain predicted map information generated by the first machine learning model; optionally, the training device can input the image under the perspective view of the first traffic environment and point cloud data of the first traffic environment into the first machine learning model to obtain predicted map information generated by the first machine learning model.
[0157] 1002、training the first machine learning model based on the predicted map information and updated first map information of the first traffic environment, wherein the updated first map information of the first traffic environment comprises two-dimensional coordinates of at least one first position point in a road element included in the first traffic environment and an expected height, the height of the at least one first position point being obtained based on first map information of the first traffic environment and three-dimensional (3D) information of the first traffic environment, the first map information comprising two-dimensional coordinates of the at least one first position point, and the 3D information of the first traffic environment comprising a height of at least one position region of an object in the first traffic environment.
[0158] Illustratively, the updated first map information of the first traffic environment can be understood as training data of the first machine learning model, and the updated first map information of the first traffic environment can be obtained through the above-mentioned corresponding embodiments of FIG. 1. Optionally, the second position point is any one of the at least one first position point, the height of the second position point in the updated first map information is obtained based on a height of a target position region corresponding to the second position point in the at least one position region, and the updated first map information of the first traffic environment can be implemented in the above-mentioned manner, which will not be repeated here.
[0159] The training device can calculate a similarity between the predicted map information and the updated first map information of the first traffic environment to obtain a function value of a loss function, the loss function indicating the similarity between the predicted map information and the updated first map information of the first traffic environment, and a target of training the first machine learning model using the loss function comprises improving the similarity between the predicted map information and the updated first map information of the first traffic environment. The training device can update weight parameters of the first machine learning model based on the function value of the loss function using a back propagation algorithm to complete one training of the first machine learning model.
[0160] The training device can repeatedly perform steps 1001 and 1002 multiple times to realize iterative training of the first machine learning model until a convergence condition is met, obtaining the trained first machine learning model; the aforementioned convergence condition can include: meeting the convergence condition of the loss function, and / or the number of times of iterative training of the first machine learning model reaching a preset number of times.
[0161] Please continue to refer to FIG. 11, which is a flowchart of a method for acquiring map information provided by an embodiment of the present application. The method for acquiring map information provided by the embodiment of the present application can include:
[0162] 1101. Acquire an image under a perspective view of a second traffic environment.
[0163] Illustratively, the second device can acquire at least one image under a perspective view of the second traffic environment through the first sensor; optionally, the second device can also acquire point cloud data of the second traffic environment through the second sensor.
[0164] Illustratively, the second device can be a vehicle, and steps 1101 and 1102 can be performed by an intelligent driving system in the second device; or the second device can also be a cloud server in communication connection with the vehicle.
[0165] 1102. Obtain map information of the second traffic environment based on the image under the perspective view of the second traffic environment, the map information of the second traffic environment including a two-dimensional coordinate and a height of at least one third position point in a road element included in the second traffic environment, the map information of the second traffic environment being generated by a first machine learning model, wherein the training data of the first machine learning model includes updated first map information of a first traffic environment, the updated first map information of the first traffic environment including a two-dimensional coordinate and a height of at least one first position point in a road element included in the first traffic environment, the height of the at least one first position point being obtained based on first map information of the first traffic environment and three-dimensional (3D) information of the first traffic environment, the first map information including a two-dimensional coordinate of the at least one first position point, and the 3D information of the first traffic environment including a height of at least one position region of an object in the first traffic environment.
[0166] Exemplarily, the second device can input at least one image under the perspective view of the second traffic environment (optionally, also including point cloud data of the second traffic environment) into the first machine learning model that has performed the training operation, to obtain map information of the second traffic environment generated by the first machine learning model, the map information of the second traffic environment including two-dimensional coordinates and height of at least one third position point in road elements included in the second traffic environment. The training data of the first machine learning model includes images under the perspective view of the first traffic environment and the updated first map information of the first traffic environment. The updated first map information of the first traffic environment can be obtained in the manner described in the corresponding embodiment of FIG. 1, which is not repeated here.
[0167] For a more intuitive understanding of the present scheme, please refer to FIG. 12, which is another flowchart of the method for obtaining map information provided by the embodiments of the present application. As shown in FIG. 12, multiple images under the perspective view of the second traffic environment and point cloud data of the second traffic environment are input into the first machine learning model, to obtain map information of the second traffic environment generated by the first machine learning model, the map information of the second traffic environment including two-dimensional coordinates and height of at least one third position point in road elements included in the second traffic environment. It should be understood that the example in FIG. 12 is only for the convenience of understanding the present scheme, and is not used to limit the present scheme.
[0168] In the embodiments of the present application, the updated first map information is used as the training data of the first machine learning model, the input of the first machine learning model includes images under the perspective view of the traffic environment, and the output of the first machine learning model includes two-dimensional coordinates and height of position points in the traffic environment, that is, the first machine learning model is trained to directly generate map information carrying two-dimensional coordinates and height of position points, which is beneficial to realize real-time generation of map information carrying two-dimensional coordinates and height of position points based on images under the perspective view of the traffic environment, so as to obtain accurate map information more conveniently and timely.
[0169] Based on the embodiments corresponding to FIG. 1 to FIG. 12, in order to better implement the above scheme of the embodiments of the present application, the related equipment for implementing the above scheme is further provided below. Referring to FIG. 13, FIG. 13 is a structural schematic diagram of a map information processing apparatus provided by the embodiments of the present application. The map information processing apparatus 1300 includes: an acquisition module 1301 configured to acquire first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment; the acquisition module 1301 is further configured to acquire three-dimensional (3D) information of the first traffic environment, the 3D information including a height of at least one position region, an object in the first traffic environment including the at least one position region; and a processing module 1302 configured to obtain the height of the at least one first position point based on the first map information and the 3D information, and update the first map information to include the two-dimensional coordinates of the at least one first position point and the height of the at least one first position point.
[0170] Optionally, the second position point is any one of the at least one first position point, and the processing module is specifically configured to determine the height of the second position point based on the height of a target position region in the at least one position region, where the target position region is a position region in the at least one position region corresponding to the second position point.
[0171] Optionally, the 3D information further includes two-dimensional coordinates of each position region, and the processing module is specifically further configured to determine the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each position region.
[0172] Optionally, the processing module 1302 is specifically configured to determine the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point, the two-dimensional coordinates of each position region, and a semantic type of each position region.
[0173] Optionally, the first map information is map information in a bird's eye view (BEV), and the map information processing apparatus 1300 further includes a determination module 1303 configured to determine evaluation information corresponding to the updated first map information based on a first image in a perspective view (PV) of the first traffic environment, the evaluation information indicating an accuracy of the updated first map information.
[0174] Optionally, the at least one first position point comprises a position point on a lane boundary line in the first traffic environment, the determining module 1303 is specifically configured to project the first position point into the first image under the perspective view to obtain first position information based on the two-dimensional coordinates of the first position point and the height of the first position point, the first position information indicating a position of a projection point corresponding to the first position point in the first image under the perspective view; and determine the evaluation information based on the first position information and the first image under the perspective view.
[0175] Optionally, the determining module 1303 is specifically configured to perform a semantic recognition operation based on the first image under the perspective view to obtain second position information, the second position information comprising position information of a pixel point of a semantic type of a lane boundary line in the first image under the perspective view; and determine the evaluation information according to the first position information and the second position information.
[0176] Optionally, the determining module 1303 is specifically configured to perform an image segmentation operation based on the first image under the perspective view to obtain third position information, the third position information indicating a position of an image region in which the lane boundary line is located in the first image under the perspective view; and determine the evaluation information according to the first position information and the third position information.
[0177] Optionally, the first map information is obtained based on second map information and third map information, the second map information and the third map information are both map information under a bird's eye view BEV of the first traffic environment, the two-dimensional coordinates included in the second map information are under an absolute world coordinate system, the two-dimensional coordinates included in the third map information are under a relative world coordinate system, the second map information and the third map information are both obtained based on the first image and / or the point cloud data of the first traffic environment collected by the vehicle, and the origin of the relative world coordinate system is obtained based on the position of the vehicle.
[0178] Optionally, the 3D information is obtained based on the second image under the perspective view of the first traffic environment and / or the point cloud data of the first traffic environment.
[0179] It should be noted that the information interaction and execution process between the modules / units in the map information processing apparatus 1300 are based on the same concept as the various method embodiments corresponding to FIGS. 1 to 12 of the present application, and the specific content can be referred to the description of the method embodiments in the foregoing description of the present application, which will not be described here.
[0180] Please continue to refer to FIG. 14, which is a structural schematic diagram of the apparatus for acquiring map information provided in the embodiment of the present application. The apparatus for acquiring map information 1400 includes: an acquisition module 1401, configured to acquire an image under a perspective view PV of a second traffic environment; and a processing module 1402, configured to obtain map information of the second traffic environment based on the image under the perspective view, the map information including a two-dimensional coordinate and a height of at least one third position point in a road element included in the second traffic environment, the map information being generated by a machine learning model; wherein training data of the machine learning model includes updated first map information of a first traffic environment, the updated first map information of the first traffic environment including a two-dimensional coordinate and a height of at least one first position point in a road element included in the first traffic environment, the height of the at least one first position point being obtained based on first map information of the first traffic environment and three-dimensional (3D) information of the first traffic environment, the first map information including a two-dimensional coordinate of the at least one first position point, and the 3D information of the first traffic environment including a height of at least one position region of an object in the first traffic environment.
[0181] Optionally, the second position point is any one of the at least one first position point, and the height of the second position point is obtained based on a height of a target position region corresponding to the second position point in the at least one position region.
[0182] It should be noted that the information interaction and execution process between the modules / units in the apparatus for acquiring map information 1400 are based on the same concept as the respective method embodiments corresponding to FIGS. 1 to 12 of the present application, and the specific content can be referred to the description in the foregoing method embodiments of the present application, which will not be described here again.
[0183] The embodiment of the present application also provides a device, as shown in FIG. 15, which is a structural schematic diagram of the device provided in the embodiment of the present application. Optionally, the device 1500 performs the functions of the first device, the second device, or the training device in the respective method embodiments corresponding to FIGS. 1 to 12.
[0184] The device 1500 includes a memory 1502 and at least one processor 1501. Optionally, the processor 1501 implements the method in the foregoing embodiments by reading instructions saved in the memory 1502, or the processor 1501 can also implement the method in the foregoing embodiments by internally stored instructions. In the case where the processor 1501 implements the method in the foregoing embodiments by reading instructions saved in the memory 1502, the memory 1502 saves instructions for implementing the method provided in the foregoing embodiments of the present application.
[0185] Optionally, the at least one processor 1501 is one or more CPUs, or is a single core CPU, or is a multi-core CPU. The memory 1502 includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, or an optical memory, etc. The memory 1502 stores instructions of an operating system. After the program instructions stored in the memory 1502 are read by the at least one processor 1501, the device 1500 performs the corresponding operations in the foregoing embodiments.
[0186] Optionally, the device 1500 further includes a network interface 1503, which can be a wired interface or a wireless interface. The network interface 1503 is configured to perform the receiving and sending of data in the execution of the various method embodiments of FIGS. 1-12.
[0187] It should be understood that the network interface 1503 has the functions of receiving data and sending data. The function of receiving data and the function of sending data can be integrated in the same transceiver interface, or the function of receiving data and the function of sending data can be implemented in different interfaces respectively, which is not limited here. In other words, the network interface 1503 can include one or more interfaces for implementing the function of receiving data and the function of sending data.
[0188] After the processor 1501 reads the program instructions in the memory 1502, the device 1500 can perform other functions, which can refer to the descriptions in the foregoing various method embodiments.
[0189] Optionally, the device 1500 further includes a bus 1504. The processor 1501 and the memory 1502 are usually connected to each other through the bus 1504, or can be connected to each other in other manners.
[0190] The device 1500 provided by the embodiments of the present application is configured to perform the method performed by the first device, the second device, or the training device in the foregoing various method embodiments, and achieve the corresponding beneficial effects. The specific implementation modes of the device 1500 shown in FIG. 15 can refer to the descriptions in the foregoing various method embodiments, which will not be repeated here.
[0191] The embodiments of the present application also provide a vehicle. Referring to FIG. 16, FIG. 16 is a schematic structural diagram of a vehicle provided by the embodiments of the present application. The vehicle 100 is configured to be in a full or partial intelligent driving mode. For example, the vehicle 100 can control itself while being in the intelligent driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the possibility of the other vehicle performing the possible behavior, and control the vehicle 100 based on the determined information. When the vehicle 100 is in the intelligent driving mode, the vehicle 100 can also be configured to operate without human interaction.
[0192] The vehicle 100 can include various subsystems, such as a travel system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power source 110, a computer system 112, and a user interface 116. Alternatively, the vehicle 100 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and component of the vehicle 100 can be interconnected by wire or wirelessly.
[0193] The travel system 102 can include components that provide powered movement for the vehicle 100. In one embodiment, the travel system 102 can include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.
[0194] The engine 118 can be an internal combustion engine, an electric motor, an air compression engine, or other types of engine combinations, such as a hybrid engine composed of a gasoline engine and an electric motor, a hybrid engine composed of an internal combustion engine and an air compression engine. The engine 118 converts the energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source 119 can also provide energy for other systems of the vehicle 100. The transmission 120 can transmit mechanical power from the engine 118 to the wheels 121. The transmission 120 can include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 can also include other devices, such as a clutch. The drive shaft can include one or more shafts that can be coupled to one or more wheels 121.
[0195] The sensor system 104 can include several sensors that sense information about the environment surrounding the vehicle 100. For example, the sensor system 104 can include a positioning system 122 (which can be a global positioning GPS system, a Beidou system, or other positioning system), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. The sensor system 104 can also include sensors that monitor internal systems of the vehicle 100 (e.g., in-vehicle air quality monitors, fuel gauges, oil temperature gauges, etc.). Sensory data from one or more of these sensors can be used to detect objects and their respective characteristics (location, shape, direction, velocity, etc.). Such detection and identification is a critical function for the safe operation of the autonomous vehicle 100.
[0196] The positioning system 122 can be used to estimate the geographic location of the vehicle 100. The IMU 124 is used to perceive changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. The radar 126 can utilize radio signals to perceive objects within the surrounding environment of the vehicle 100, which can be manifested as a millimeter wave radar or a laser radar. In some embodiments, in addition to perceiving objects, the radar 126 can also be used to perceive the velocity and / or heading of the objects. The laser rangefinder 128 can utilize laser light to perceive objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 can include one or more laser sources, a laser scanner, and one or more detectors, among other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.
[0197] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 can include various components, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a line control system 142, and an obstacle avoidance system 144.
[0198] The steering system 132 is operable to adjust the heading of the vehicle 100. In one embodiment, the steering system 132 can be a steering wheel system. The throttle 134 is used to control the speed of the engine 118 and, in turn, the speed of the vehicle 100. The braking unit 136 is used to control the deceleration of the vehicle 100. The braking unit 136 can use friction to slow the wheels 121. In other embodiments, the braking unit 136 can convert the kinetic energy of the wheels 121 into electrical current. The braking unit 136 can also take other forms to slow the wheels 121 and, in turn, control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the cameras 130 to identify objects and / or features in the environment surrounding the vehicle 100. The objects and / or features can include traffic signals, road boundaries, and obstacles. The computer vision system 140 can use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate the speed of objects, and the like. The route control system 142 is used to determine the route and speed of travel of the vehicle 100. In some embodiments, the route control system 142 can include a lateral planning module 1421 and a longitudinal planning module 1422 that are used to determine the route and speed of travel of the vehicle 100 in conjunction with data from the obstacle avoidance system 144, the GPS 122, and one or more predetermined maps, respectively. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise navigate around obstacles in the environment of the vehicle 100, which can be actual obstacles and virtual moving obstacles that can collide with the vehicle 100. In one instance, the control system 106 can include additional components in addition to those shown and described, or some of the components shown can be reduced.
[0199] The vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users through the peripherals 108. The peripherals 108 can include a wireless communication system 146, an on-board computer 148, a microphone 150, and / or a speaker 152. In some embodiments, the peripherals 108 provide a means for a user of the vehicle 100 to interact with the user interface 116. For example, the on-board computer 148 can provide information to a user of the vehicle 100. The user interface 116 can also operate the on-board computer 148 to receive input from the user. The on-board computer 148 can be operated through a touch screen. In other cases, the peripherals 108 can provide a means for the vehicle 100 to communicate with other devices located within the vehicle. For example, the microphone 150 can receive audio (e.g., voice commands or other audio input) from a user of the vehicle 100. Similarly, the speaker 152 can output audio to a user of the vehicle 100. The wireless communication system 146 can wirelessly communicate with one or more devices, either directly or via a communication network. For example, the wireless communication system 146 can use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE. Or 5G cellular communication. The wireless communication system 146 can utilize wireless local area network (WLAN) communication. In some embodiments, the wireless communication system 146 can utilize an infrared link, Bluetooth, or ZigBee to communicate directly with devices. Other wireless protocols, such as various vehicle communication systems, for example, the wireless communication system 146 can include one or more dedicated short range communications (DSRC) devices, which can include public and / or private data communication between vehicles and / or roadside stations.
[0200] The power supply 110 can provide power to various components of the vehicle 100. In one embodiment, the power supply 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such a battery can be configured to power the various components of the vehicle 100. In some embodiments, the power supply 110 and the energy source 119 can be implemented together, such as in some all-electric vehicles.
[0201] Some or all of the functionality of vehicle 100 is controlled by computer system 112. Computer system 112 can include at least one processor 113 that executes instructions 115 stored in a non-transitory computer readable medium such as memory 114. Computer system 112 can also be a plurality of computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 can be any conventional processor such as a commercially available central processing unit (CPU). Alternatively, processor 113 can be a dedicated device such as an application specific integrated circuit (ASIC) or other hardware-based processor. Although FIG. 16 functionally illustrates the processor, memory, and other components of computer system 112 within the same block, it will be understood by those skilled in the art that the processor, or memory, can actually comprise a plurality of processors, or memories, that can or can not be stored within the same physical housing. For example, memory 114 can be a hard drive or other storage media located in a housing different from that of computer system 112. Accordingly, references to processor 113 or memory 114 will be understood to include references to a collection of processors or memories that can or can not operate in parallel. Rather than using a single processor to perform the steps described herein, some components such as the steering assembly and the deceleration assembly can each have their own processor that only performs calculations related to the functionality specific to that component.
[0202] In various aspects described herein, processor 113 can be located remotely from vehicle 100 and in wireless communication with vehicle 100. In other aspects, some of the processes described herein are performed on processor 113 disposed within vehicle 100 while others are performed by a remote processor 113, including taking the necessary steps to perform a single maneuver.
[0203] In some embodiments, the memory 114 can include instructions 115 (e.g., program logic) that can be executed by the processor 113 to perform various functions of the vehicle 100, including those described above. The memory 114 can also include additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the travel system 102, the sensor system 104, the control system 106, and the peripherals 108. In addition to the instructions 115, the memory 114 can also store data, such as road maps, route information, the vehicle's location, direction, speed, and other such vehicle data, and other information. Such information can be used by the vehicle 100 and the computer system 112 during operation of the vehicle 100 in autonomous, semi-autonomous, and / or manual modes. The user interface 116 is used to provide information to or receive information from a user of the vehicle 100. Optionally, the user interface 116 can include one or more input / output devices within the set of peripherals 108, such as the wireless communication system 146, the on-board computer 148, the microphone 150, and the speaker 152.
[0204] The computer system 112 can control the functions of the vehicle 100 based on inputs received from various subsystems (e.g., the travel system 102, the sensor system 104, and the control system 106), as well as from the user interface 116. For example, the computer system 112 can utilize inputs from the control system 106 in order to control the steering system 132 to avoid obstacles detected by the sensor system 104 and the obstacle avoidance system 144. In some embodiments, the computer system 112 can be operable to provide control over many aspects of the vehicle 100 and its subsystems.
[0205] Optionally, one or more of the above-described components can be installed separately from or associated with the vehicle 100. For example, the memory 114 can exist partially or entirely separately from the vehicle 100. The above-described components can be communicatively coupled together in a wired and / or wireless manner.
[0206] Optionally, the above-described components are just one example, and in actual applications, components in each of the above-described modules can be added or deleted according to actual needs, and FIG. 16 should not be understood as a limitation on the embodiments of the present application. A vehicle traveling on a road, such as the vehicle 100 above, can identify objects within its surrounding environment to determine an adjustment to a current speed. The objects can be other vehicles, traffic control devices, or other types of objects. In some examples, each identified object can be considered independently, and based on respective characteristics of the object, such as its current speed, acceleration, spacing from the vehicle, etc., can be used to determine a speed at which the vehicle is to adjust.
[0207] Optionally, the vehicle 100 or a computing device associated with the vehicle 100, such as the computer system 112, the computer vision system 140, the memory 114 of FIG. 16, can predict the behavior of the identified object based on the characteristics of the identified object and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object is dependent on the behavior of the other identified objects, so all of the identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified object. In other words, the vehicle 100 can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the object. Other factors can also be considered in determining the speed of the vehicle 100 during this process, such as the lateral position of the vehicle 100 in the road it is traveling on, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 100 to cause the vehicle 100 to follow a given trajectory and / or maintain a safe lateral and longitudinal distance from objects in the vicinity of the vehicle 100 (e.g., a car in the adjacent lane on the road).
[0208] In the embodiments of the present application, the processor 113 in the vehicle 100 is configured to perform the method performed by the second device in the embodiments corresponding to FIG. 1 to FIG. 12. It should be noted that the specific manners in which the processor 113 performs the foregoing steps are based on the same concept as the method embodiments corresponding to FIG. 1 to FIG. 12 in the present application, and bring the same technical effects as the method embodiments corresponding to FIG. 1 to FIG. 12 in the present application. For details, refer to the descriptions of the method embodiments in the foregoing present application, which will not be described here again.
[0209] In the embodiments of the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a program. When the program runs on a computer, the computer is caused to perform the steps performed by the first device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12, or the computer is caused to perform the steps performed by the second device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12, or the computer is caused to perform the steps performed by the training device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12.
[0210] In the embodiments of the present application, a computer program product is also provided, and the computer program product includes a program. When the program runs on a computer, the computer is caused to perform the steps performed by the first device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12, or the computer is caused to perform the steps performed by the second device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12, or the computer is caused to perform the steps performed by the training device in the method described in the foregoing embodiments of FIG. 1 to FIG. 12.
[0211] The embodiments of the present application further provide a circuit system, which comprises a processing circuit configured to perform the steps performed by the first device in the method described in the foregoing embodiments of the present application shown in FIG. 1 to FIG. 12, or the processing circuit is configured to perform the steps performed by the second device in the method described in the foregoing embodiments of the present application shown in FIG. 1 to FIG. 12, or the processing circuit is configured to perform the steps performed by the training device in the method described in the foregoing embodiments of the present application shown in FIG. 1 to FIG. 12.
[0212] The first device, the second device or the training device provided by the embodiments of the present application can be a chip, which comprises a processing unit, for example, a processor, and optionally comprises a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit, so as to enable the chip to perform the method described in the foregoing embodiments of the present application shown in FIG. 1 to FIG. 12. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0213] In addition, it should be noted that the apparatus embodiments described above are merely schematic, wherein the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application. In addition, in the apparatus embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0214] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CLU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, the software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, server or network device, etc.) execute the method described in various embodiments of the application.
[0215] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be in the form of computer program product entirely or partially.
[0216] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the flow or function described in the embodiments of the application is entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as server, data center, etc. integrated with one or more available media. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.
Claims
1. A method of processing map information, characterized by, The method comprises: obtaining first map information of a first traffic environment, the first map information comprising two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in a road element included in the first traffic environment; obtaining three-dimensional (3D) information of the first traffic environment, the 3D information comprising a height of at least one position region, an object in the first traffic environment comprising the at least one position region; based on the first map information and the 3D information, obtaining a height of the at least one first position point, the updated first map information comprising two-dimensional coordinates of the at least one first position point and the height of the at least one first position point.
2. The method of claim 1, wherein, The second position point is any one of the at least one first position point, and the obtaining of the updated first map information based on the first map information and the 3D information comprises: determining the height of the second position point based on a height of a target position region in the at least one position region, wherein the target position region is a position region in the at least one position region corresponding to the second position point.
3. The method of claim 2, wherein, The 3D information further comprises two-dimensional coordinates of each position region, and the obtaining of the updated first map information based on the first map information and the 3D information further comprises: determining the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each position region.
4. The method of claim 3, wherein, The determining of the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point and the two-dimensional coordinates of each position region comprises: determining the target position region corresponding to the second position point from the at least one position region based on the two-dimensional coordinates of the second position point, the two-dimensional coordinates of each position region, and a semantic type of each position region.
5. The method according to any one of claims 1 to 4, characterized in that, The first map information is map information in a bird's eye view (BEV), and the method further comprises: determining evaluation information corresponding to the updated first map information based on a first image in a perspective view (PV) of the first traffic environment, the evaluation information indicating accuracy of the updated first map information.
6. The method of claim 5, wherein, The at least one first position point comprises a position point on a lane boundary line in the first traffic environment, and the determining of the evaluation information corresponding to the updated first map information based on the first image in the perspective view of the first traffic environment comprises: projecting the first position point into the first image in the perspective view to obtain first position information based on the two-dimensional coordinates of the first position point and the height of the first position point, the first position information indicating a position of a projection point corresponding to the first position point in the first image in the perspective view; determining the evaluation information based on the first position information and the first image in the perspective view.
7. The method of claim 6, wherein, The determining of the evaluation information based on the first position information and the first image in the perspective view comprises: perform a semantic recognition operation based on the first image under the perspective view to obtain second position information, the second position information including position information of pixel points of a semantic type of lane boundary line in the first image under the perspective view; determine the evaluation information according to the first position information and the second position information.
8. The method of claim 6, wherein, The determination of the evaluation information based on the first position information and the first image under the perspective view includes: perform an image segmentation operation based on the first image under the perspective view to obtain third position information, the third position information indicating a position of an image region in which the lane boundary line is located in the first image under the perspective view; determine the evaluation information according to the first position information and the third position information.
9. The method according to any one of claims 1 to 4, characterized in that, The first map information is obtained based on second map information and third map information, the second map information and the third map information are both map information under a bird's eye view (BEV) of the first traffic environment, the two-dimensional coordinates included in the second map information are under an absolute world coordinate system, the two-dimensional coordinates included in the third map information are under a relative world coordinate system, the second map information and the third map information are both obtained based on first images and / or point cloud data of the first traffic environment collected by a vehicle, and the origin of the relative world coordinate system is obtained based on a position of the vehicle.
10. The method according to any one of claims 1 to 4, characterized in that, The 3D information is obtained based on second images under a perspective view of the first traffic environment and / or point cloud data of the first traffic environment.
11. A map information acquisition method characterized by comprising: The method includes: obtaining an image under a perspective view (PV) of a second traffic environment; obtaining map information of the second traffic environment based on the image under the perspective view, the map information including two-dimensional coordinates and a height of at least one third position point in road elements included in the second traffic environment, and the map information being generated by a machine learning model; wherein training data of the machine learning model includes updated first map information of a first traffic environment, the updated first map information of the first traffic environment including two-dimensional coordinates and a height of at least one first position point in road elements included in the first traffic environment, the height of the at least one first position point being obtained based on first map information of the first traffic environment and three-dimensional (3D) information of the first traffic environment, the first map information including two-dimensional coordinates of the at least one first position point, and the 3D information of the first traffic environment including a height of at least one position region of an object in the first traffic environment.
12. The method of claim 11, wherein, The second position point is any one of the at least one first position point, and the height of the second position point is obtained based on a height of a target position region corresponding to the second position point in the at least one position region.
13. A map information processing device, characterized by comprising: The device includes: an obtaining module configured to obtain first map information of a first traffic environment, the first map information including two-dimensional coordinates of at least one first position point, the at least one first position point being a position point in road elements included in the first traffic environment; The acquisition module is further configured to acquire three-dimensional (3D) information of the first traffic environment, the 3D information comprising a height of at least one location region, and the object in the first traffic environment comprising the at least one location region. The processing module is configured to obtain the height of the at least one first location point based on the first map information and the 3D information, and update the first map information to comprise the 2D coordinates of the at least one first location point and the height of the at least one first location point.
14. A map information acquisition apparatus characterized by comprising: The apparatus comprises: An acquisition module configured to acquire an image under a perspective view (PV) of a second traffic environment. A processing module configured to obtain map information of the second traffic environment based on the image under the PV, the map information comprising 2D coordinates and a height of at least one third location point in a road element comprised by the second traffic environment, and the map information being generated by a machine learning model. The training data of the machine learning model comprises updated first map information of a first traffic environment, the updated first map information of the first traffic environment comprising 2D coordinates and a height of at least one first location point in a road element comprised by the first traffic environment, the height of the at least one first location point being obtained based on first map information of the first traffic environment and 3D information of the first traffic environment, the first map information comprising 2D coordinates of the at least one first location point, and the 3D information of the first traffic environment comprising a height of at least one location region of an object in the first traffic environment.
15. An apparatus, comprising: A processor coupled to a memory, the memory storing program instructions that, when executed by the processor, implement the method of any one of claims 1-12.
16. A vehicle characterized by comprising: A processor coupled to a memory, the memory storing program instructions that, when executed by the processor, implement the method of claim 11 or 12.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and when the program is executed on a computer, the computer is caused to perform the method of any one of claims 1-12.
18. A computer program product, characterised in that, The computer program product comprises a program, and when the program is executed on a computer, the computer is caused to perform the method of any one of claims 1-12.
19. A chip, characterized by The chip comprises a processor configured to perform the steps of the method of any one of claims 1-12.
Citation Information
Patent Citations
Map generation method and device, driving control method and device, electronic equipment and system
CN112069856A
Three-dimensional map generation method and device thereof
CN112308969A
Map generation method and electronic equipment
CN114677486A
Acquisition method of lane marking data, computer equipment and storage medium
CN115451977A
Device, program, and method for map information creation
JP2018106017A