Map information processing method, map information acquisition method and equipment
By combining two-dimensional and three-dimensional information to generate map information including the elevation of location points, the problem of insufficient accuracy of two-dimensional maps in complex traffic environments is solved, and a more accurate reflection of the traffic environment is achieved.
Patent Information
- Application Number
- CN202411355609.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-11-28
AI Technical Summary
Existing two-dimensional map information cannot accurately reflect the real situation in the traffic environment, especially in complex scenarios such as uphill and downhill sections, where the lack of height information of location points leads to inaccurate map information.
By acquiring three-dimensional information about the traffic environment and combining it with two-dimensional map information, the elevation of each location point is determined, thereby generating updated map information containing two-dimensional coordinates and elevation.
The generated updated map information can more accurately reflect the traffic environment and improve the accuracy of map information, especially in scenarios with complex terrain changes.
Smart Images

Figure CN121026164A_ABST
Abstract
Description
Technical Field
[0001] This application relates to intelligent driving technology, and more particularly to a method for processing map information, a method for acquiring map information, and a device. Background Technology
[0002] With the rapid development of intelligent driving technology, map information is needed in more and more scenarios. Currently, the map information used is often two-dimensional (2D) map information of the traffic environment. This 2D map information includes the two-dimensional coordinates of the location points in the road elements of the traffic environment. However, in some traffic scenarios, such as uphill and downhill traffic scenarios, the 2D map information does not include the height of each location point, which makes it impossible for the 2D map information to accurately reflect the real traffic environment. Therefore, a map generation scheme that can simultaneously include the two-dimensional coordinates and height of the location points is urgently needed. Summary of the Invention
[0003] This application provides a method for processing map information, a method for acquiring map information, and a device for doing so. It provides a scheme for generating map information that simultaneously includes the two-dimensional coordinates and elevation of a location point, thereby enabling the updated first map information to more accurately reflect the real traffic environment.
[0004] This application provides the following technical solution:
[0005] Firstly, this application provides a method for processing map information, which can be used in the field of intelligent driving. In this method, a first device not only acquires first map information of a first traffic environment, but also acquires three-dimensional (3D) information of the first traffic environment. The first map information of the first traffic environment includes the two-dimensional coordinates of each of at least one first location points, where each first location point is a location point among road elements included in the first traffic environment. The 3D information of the first traffic environment includes the height of each of at least one location regions, where each of the at least one location regions is a location region among objects included in the first traffic environment. Furthermore, the first device can obtain the height of each of the at least one first location points based on the first map information and the 3D information of the first traffic environment, thus completing the high-value snapping of each first location point, thereby obtaining updated first map information. The updated first map information includes the two-dimensional coordinates and height of each of the at least one first location points.
[0006] For example, the first traffic environment in this application can also be referred to as the first road environment, and the first map information of the first traffic environment can be understood as two-dimensional (2D) map information of the first traffic environment. Optionally, the first map information can be 2D map information under a bird's-eye view (BEV) (hereinafter referred to as "BEV view"), and the BEV view in this application can also be referred to as a top-down view.
[0007] Optionally, the road element can be a road element on the ground. For example, the road element can include lane boundary lines. In this application, the lane boundary lines can be lines that exist in real roads to divide different lanes. Optionally, the road element can also include at least one of the following: lane center line, lane, road boundary line, intersection, sidewalk, turn indicator line or other road elements. In this application, the lane center line can be understood as the virtual centerline of the lane. The lane center line can be a line that does not exist in real roads. The lane center line can be used to locate the lane.
[0008] For example, each location region can be represented as a cube, and the aforementioned at least one location region can be understood as dividing the objects in the first traffic environment into the aforementioned at least one location region. For example, the objects in the first traffic environment can include road elements in the first traffic environment, and the objects in the first traffic environment can also include other objects besides road elements. For example, other objects besides road elements in the first traffic environment can include traffic lights, streetlights, trees, buildings, fire hydrants, warning triangles, cones, warning posts, crash barriers, construction signs, traffic guidance signs, or other types of static obstacles, etc. Optionally, other objects besides road elements in the first traffic environment can also include dynamic obstacles in the first traffic environment. For example, the aforementioned dynamic obstacles can include: vehicles, electric vehicles, bicycles, pedestrians, animals, or other dynamic obstacles, etc.
[0009] In this implementation, not only is the first map information of the traffic environment obtained, which includes the two-dimensional coordinates of at least one first location point among the road elements included in the traffic environment, but also the 3D information of the traffic environment is obtained, which includes the height of at least one location area included by objects in the traffic environment. Based on the aforementioned first map information and 3D information, the height of at least one first location point is obtained. The updated first map information may include the two-dimensional coordinates and height of at least one location point, providing a map information generation scheme that simultaneously includes the two-dimensional coordinates and height of location points, thereby enabling the updated first map information to more accurately reflect the real traffic environment.
[0010] In one possible implementation, any one of the at least one first location points is referred to as a second location point. The first device obtains updated first map information based on the first map information and 3D information of the first traffic environment. This may include: the first device determining the height of the second location point based on the height of a target location area in at least one location area, wherein the aforementioned target location area is the location area corresponding to the second location point in at least one location area; the first device may repeat the aforementioned steps at least once to obtain the height of each of the at least one first location points.
[0011] In this implementation, for any one of the at least one first position point (i.e., the second position point), the height of the target position area corresponding to the first position point in at least one position area is obtained. Based on the height of the target position area, the height of the second position point is determined. That is, each of the at least one first position points is traversed, and the height of each first position point is obtained one by one. This more refined management of the process of obtaining the height of at least one first position point is beneficial to improving the accuracy of the height of each first position point.
[0012] In one possible implementation, exemplarily, the two-dimensional coordinates of the first location point may include first coordinate information in the X-axis direction and the Y-axis direction, wherein the X-axis and Y-axis are perpendicular to each other. For example, the two-dimensional coordinates of the first location point may be first coordinate information in a first coordinate system, optionally, the first coordinate system may be a vehicle coordinate system. The 3D information of the first traffic environment also includes the two-dimensional coordinates of each location region in at least one location region. In other words, the 3D information of the first traffic environment may include the second coordinate information of each location region in at least one location region in the first traffic environment in a second coordinate system. The second coordinate information of each location region may include the two-dimensional coordinates and height of each location region. Optionally, the second coordinate system may be a vehicle coordinate system.
[0013] The first device obtains updated first map information based on first map information and 3D information, which may include: the first device determining at least one target location area corresponding to the second location point from at least one location area based on the two-dimensional coordinates of the second location point (i.e., any one of at least one first location point) and the two-dimensional coordinates of each location area in at least one location area; and then determining the height of the second location point based on the height of the at least one target location area.
[0014] For example, if the at least one target location area includes only one target location area, the first device can directly determine the height of the single target location area as the height of the second location point. If the at least one target location area includes at least two target location areas, in one implementation, the first device can determine the minimum height among the at least two heights corresponding to the at least two target location areas as the height of the second location point. Since the second location point can be a location point on the road surface, directly determining the minimum height as the height of the second location point is feasible. Alternatively, in another implementation, the first device can determine the average height of the at least two heights corresponding to the at least two target location areas as the height of the second location point, etc.
[0015] This implementation method clarifies the two-dimensional coordinates of the first position point and the two-dimensional coordinates of the position region in the 3D information to find the target position region corresponding to each first position point, providing a simple and effective method for determining the target position region.
[0016] In one possible implementation, the first device determines a target location area corresponding to the second location point from at least one location area based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location area. This can include: the first device determining the target location area corresponding to the second location point from at least one location area based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location area, and the semantic type of each location area. For example, the semantic type of a location area can be the semantic type of lane boundary lines, road boundary lines, sidewalks, intersections, turn indicators, traffic lights, streetlights, trees, buildings, fire hydrants, warning triangles, cones, warning posts, crash barriers, construction signs, traffic guidance signs, or other objects that may appear in the traffic environment. Optionally, if the objects in the first traffic environment also include dynamic obstacles in the first traffic environment, the semantic type of a location area can also be the semantic type of vehicles, electric vehicles, bicycles, pedestrians, animals, or other dynamic obstacles.
[0017] Optionally, the first device can determine the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, and select at least one candidate location region from at least one location region that is closest to the two-dimensional coordinates of the second location point. The distance between the two-dimensional coordinates of each of the aforementioned at least one candidate location region and the two-dimensional coordinates of the second location point can be consistent. For example, if the aforementioned at least one candidate location region includes at least two candidate location regions, then the two-dimensional coordinates of the different candidate location regions of the aforementioned at least two candidate location regions can be consistent, differing only in height. The first device can determine a target location region from the at least one candidate location region that has the same semantic type as the second location point, based on the semantic type of each candidate location region.
[0018] In this implementation, in determining the target location area corresponding to any first location point, not only the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area are used, but also the semantic type of each location area is used. Since the matching is based solely on the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area, a certain first location point may correspond to multiple location areas. The two-dimensional coordinates of the aforementioned multiple location areas are consistent, but their heights are inconsistent. Therefore, the target location area corresponding to the first location point can be further determined based on the semantic type of each location area. This is beneficial for accurately finding the target location area corresponding to each first location point, and thus for obtaining the accurate height of each first location point, thereby improving the accuracy of the updated first map information.
[0019] In one possible implementation, the first device determines a target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region. This may include: the first device determining the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region; and selecting at least one target location region from the at least one location region that is closest to the two-dimensional coordinates of the second location point. The distance between the two-dimensional coordinates of each target location region and the two-dimensional coordinates of the second location point can all be the same. For example, if the at least one target location region includes at least two target location regions, the two-dimensional coordinates of the different target location regions in the at least two target location regions can all be the same, differing only in height. Alternatively, the first device can round up both the X-axis and Y-axis position parameters of the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region. Alternatively, the first device can round down both the X-axis and Y-axis position parameters of the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region.
[0020] In one possible implementation, the first map information is map information from the BEV perspective, and the method may further include: the first device determining evaluation information corresponding to the updated first map information based on the first image from the perspective view (PV) of the first traffic environment, the evaluation information indicating the accuracy of the updated first map information.
[0021] In this implementation, since the image from the PV perspective of the first traffic environment can accurately reflect the situation of the first traffic environment, and the updated first map information is to more accurately reflect the first traffic environment, it is feasible to determine the accuracy of the updated first map information by using the image from the PV perspective of the first traffic environment. The image from the PV perspective of the first traffic environment is easy to obtain, thus providing an easy-to-implement detection scheme for the accuracy of the updated first map information.
[0022] In one possible implementation, the aforementioned at least one first location point includes a location point on a lane boundary line in a first traffic environment. The first device determines evaluation information corresponding to the updated first map information based on a first image from a perspective view of the first traffic environment. This includes: the first device projects the first location point onto the first image from a perspective view based on the two-dimensional coordinates and height of the first location point to obtain first location information. The first location information indicates the position of the projection point corresponding to the first location point in the first image from a perspective view. The aforementioned projection point can be understood as the projection point obtained by projecting the location point on the lane boundary line from the BEV perspective onto the first image from the PV perspective. Then, the evaluation information can be determined based on the first location information and the first image from a perspective view.
[0023] For example, in one case, the above-mentioned at least one first location point can all be a location point in the lane boundary line included in the first traffic environment. Then, the first device projects each of the at least one first location point into a first image under the perspective view based on the two-dimensional coordinates and height of each of the at least one first location point to obtain first location information. The first location information can include the position of each of the at least one projection point corresponding to the at least one first location point in the first image under the perspective view.
[0024] In another scenario, if the updated first map information includes at least one first location point not only in the lane boundary lines of the first traffic environment but also in other road elements within the first traffic environment, then the first device can filter the updated first map information to obtain at least one filtered first location point, which only includes lane boundary line locations. Based on the two-dimensional coordinates and height of each of the filtered first location points, the first device projects each of the filtered first location points onto a first image under a perspective view to obtain first location information. The first location information may include the position of each of the at least one projection points corresponding one-to-one with the filtered first location points in the first image under a perspective view.
[0025] In this implementation, a scheme is proposed to back-project the first location point from the BEV perspective onto the first image from the PV perspective based on the two-dimensional coordinates and height of the first location point in the updated first map information, thereby obtaining the location information of the projected point. Then, the accuracy of the updated first map information is evaluated based on the location information of the projected point. Since the first image from the PV perspective of the first traffic environment can accurately reflect the first traffic environment, the method of back-projecting the first location point from the BEV perspective onto the first image from the PV perspective can make the most of the first image from the PV perspective, which is conducive to obtaining more accurate evaluation information.
[0026] In one possible implementation, the first device determines evaluation information based on first location information and a first image from a perspective view. This may include: the first device performing a semantic recognition operation on the first image from the perspective view to obtain second location information, which includes the location information of pixels in the first image from the perspective view whose semantic type is lane boundary line; for example, the second location information may include the coordinate information of pixels in the first image from the perspective view whose semantic type is lane boundary line. The first device can then determine the evaluation information based on the first location information and the second location information.
[0027] Optionally, the first position information may indicate the position information of at least one first lane boundary line formed by projecting a first position point on the lane boundary line from the BEV perspective onto a first image from the PV perspective, and the second position information may indicate the position information of at least one second lane boundary line in the first image from the PV perspective. For example, the first device may determine a first distance between at least one first lane boundary line and at least one second lane boundary line based on the first and second position information, and then determine the evaluation information based on the first distance between the at least one first lane boundary line and at least one second lane boundary line.
[0028] In this implementation, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line formed by the projection points corresponding to the first location point in the first image under the PV view, and the position of the lane boundary line in the first image under the PV view. This provides a specific implementation scheme for obtaining the evaluation information of the updated first map information, improving the feasibility of this scheme. In addition, based on the position of the lane boundary line formed by the projection points and the original position of the lane boundary line in the first image under the PV view, the accuracy of the lane boundary line in the updated first map information can be obtained. Since lane boundary lines in traffic environments are often long, and it is often difficult to accurately reproduce the lane boundary line situation in the actual environment in the map information of traffic scenarios such as uphill and downhill, evaluating the accuracy of the updated first map information from the dimension of lane boundary lines can obtain more accurate evaluation information and is conducive to improving the accuracy of the position information of lane boundary lines in the updated first map information.
[0029] In one possible implementation, the first device determines evaluation information based on first location information and a first image from a perspective view. This may include: the first device performing an image segmentation operation on the first image from the perspective view to obtain third location information. The third location information indicates the position of at least one image region in the first image from the perspective view where the lane boundary line is located. For example, the lane boundary line in each first image region may be represented as a polygon, and the third location information may indicate the position of each first image region in at least one polygonal first image region where the lane boundary line is located. The first device can determine the evaluation information based on the first location information and the third location information.
[0030] Optionally, the first device can also determine at least one second image region in the first image under the PV view based on the first location information. Each second image region can be obtained by increasing the width of each first lane boundary line. For example, the width of the first lane boundary line can be increased by a preset width to obtain the second image region corresponding to each first lane boundary line. The first device can determine the intersection-over-union (IoU) ratio between at least one first image region and at least one second image region, that is, determine the ratio between the intersection and union of all first image regions and all second image regions; and then determine the evaluation information based on the IoU ratio between at least one first image region and at least one second image region.
[0031] In this implementation, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line formed by the projection points corresponding to the first position point in the first image under the PV view, and the position of the image area where the lane boundary line is located in the first image under the PV view. This provides another specific implementation scheme for obtaining the evaluation information of the updated first map information, which improves the implementation flexibility of this scheme.
[0032] In one possible implementation, the first map information of the first traffic environment can be obtained based on second and third map information. Both the second and third map information are bird's-eye view (BEV) map information of the first traffic environment. The second map information includes two-dimensional coordinates in an absolute world coordinate system, while the third map information includes two-dimensional coordinates in a relative world coordinate system. Both the second and third map information are obtained based on first images and / or point cloud data of the first traffic environment collected by the vehicle, and the origin of the relative world coordinate system is obtained based on the vehicle's position. In this implementation, obtaining the first map information of the first traffic environment based on map information in two different coordinate systems is beneficial for obtaining more accurate map information.
[0033] In one possible implementation, the 3D information of the first traffic environment is obtained based on a second image from the perspective of the first traffic environment and / or point cloud data of the first traffic environment. This implementation clarifies the information sources upon which the 3D information of the first traffic environment is obtained, which helps to improve the feasibility of this solution.
[0034] Secondly, this application provides a method for acquiring map information, which can be used in the field of intelligent driving. In this method, a second device can acquire an image from a perspective view (PV) of a second traffic environment; based on the image from the perspective view, map information of the second traffic environment is obtained. The map information of the second traffic environment includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information is generated by a machine learning model. For example, the second device can input at least one image from the perspective view of the second traffic environment (optionally, also including point cloud data of the second traffic environment) into a machine learning model that has undergone training operations to obtain the map information of the second traffic environment generated by the machine learning model.
[0035] The training data for the machine learning model includes updated first map information of the first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of at least one first location point is obtained based on the first map information and 3D information of the first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
[0036] In one possible implementation, the second location point is any one of at least one first location point, and the height of the second location point is obtained based on the height of the target location region corresponding to the second location point in at least one location region.
[0037] The updated first map information in the second aspect of this application can be obtained through the first aspect and its various possible implementations. The meanings of the terms in the second aspect and its various possible implementations, as well as the beneficial effects of each possible implementation, can be found in the descriptions of the various possible implementations in the first aspect, and will not be repeated here.
[0038] Thirdly, this application provides a map information processing apparatus that can be used in the field of intelligent driving. The map information processing apparatus includes: an acquisition module for acquiring first map information of a first traffic environment, the first map information including the two-dimensional coordinates of at least one first location point, the at least one first location point being a location point among road elements included in the first traffic environment; the acquisition module is further configured to acquire three-dimensional (3D) information of the first traffic environment, the 3D information including the height of at least one location region, objects in the first traffic environment including at least one location region; and a processing module for obtaining the height of at least one first location point based on the first map information and the 3D information, the updated first map information including the two-dimensional coordinates of at least one first location point and the height of at least one first location point.
[0039] In one possible implementation, the second location point is any one of at least one first location point, and the processing module is specifically used to determine the height of the second location point based on the height of a target location region in at least one location region, wherein the target location region is the location region corresponding to the second location point in at least one location region.
[0040] In one possible implementation, the 3D information also includes two-dimensional coordinates of each location region. The processing module is specifically used to determine the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region.
[0041] In one possible implementation, the processing module is specifically used to determine the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region, and the semantic type of each location region.
[0042] In one possible implementation, the first map information is map information under the bird's-eye view (BEV). The map information processing device further includes: a determination module, used to determine evaluation information corresponding to the updated first map information based on the first image under the perspective view (PV) of the first traffic environment, and the evaluation information indicates the accuracy of the updated first map information.
[0043] In one possible implementation, at least one first location point includes a location point on a lane boundary line in a first traffic environment. The determining module is specifically used to project the first location point onto a first image under a perspective view based on the two-dimensional coordinates and height of the first location point to obtain first location information. The first location information indicates the position of the projection point corresponding to the first location point in the first image under the perspective view. Evaluation information is determined based on the first location information and the first image under the perspective view.
[0044] In one possible implementation, a determining module is specifically used to perform semantic recognition operations based on a first image from a perspective view to obtain second position information, the second position information including the position information of pixels in the first image from the perspective view whose semantic type is lane boundary line; and to determine evaluation information based on the first position information and the second position information.
[0045] In one possible implementation, a determining module is specifically used to perform image segmentation operations based on a first image from a perspective view to obtain third position information, the third position information indicating the position of the lane boundary line in the image region of the first image from a perspective view; and to determine evaluation information based on the first position information and the third position information.
[0046] In one possible implementation, the first map information is obtained based on the second and third map information. Both the second and third map information are map information from the bird's-eye view of the first traffic environment under the BEV. The two-dimensional coordinates included in the second map information are in the absolute world coordinate system, and the two-dimensional coordinates included in the third map information are in the relative world coordinate system. Both the second and third map information are obtained based on the first image and / or point cloud data of the first traffic environment collected by the vehicle. The origin of the relative world coordinate system is obtained based on the position of the vehicle.
[0047] In one possible implementation, the 3D information is obtained based on a second image and / or point cloud data of the first traffic environment from a perspective view of the first traffic environment.
[0048] The meanings of the terms in the third aspect of this application and the various possible implementations of the third aspect, as well as the beneficial effects of each possible implementation, can be found in the descriptions of the various possible implementations in the first aspect, and will not be repeated here.
[0049] Fourthly, this application provides a map information acquisition device that can be used in the field of intelligent driving. The map information acquisition device includes: an acquisition module for acquiring an image under a perspective view (PV) of a second traffic environment; and a processing module for obtaining map information of the second traffic environment based on the image under the perspective view. The map information includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information is generated by a machine learning model. The training data of the machine learning model includes updated first map information of a first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional (3D) information of the first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
[0050] In one possible implementation, the second location point is any one of at least one first location point, and the height of the second location point is obtained based on the height of the target location region corresponding to the second location point in at least one location region.
[0051] The updated first map information in the fourth aspect of this application can be obtained through the first aspect and its various possible implementations. The meanings of the terms in the fourth aspect and its various possible implementations, as well as the beneficial effects of each possible implementation, can be found in the descriptions of the various possible implementations in the first aspect, and will not be repeated here.
[0052] Fifthly, embodiments of this application provide an apparatus including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; the processor being used to execute the program in the memory, causing the apparatus to perform the methods described in the first or second aspect.
[0053] In a sixth aspect, embodiments of this application provide a vehicle including a processor and a memory, the processor being coupled to the memory for storing a program; the processor being used to execute the program in the memory, causing the device to perform the method described in the second aspect above.
[0054] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect.
[0055] Eighthly, embodiments of this application provide a computer program product, which includes a program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect.
[0056] Ninthly, this application provides a chip system including a processor for supporting the implementation of the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the terminal device or communication device. This chip system may be composed of chips or may include chips and other discrete devices.
[0057] The fifth to ninth aspects of this application correspond to the first aspect or multiple possible ways of the first aspect, and have corresponding beneficial effects. Attached Figure Description
[0058] Figure 1 A flowchart illustrating a map information processing method provided in an embodiment of this application;
[0059] Figure 2 A schematic diagram of the first map information provided in an embodiment of this application;
[0060] Figure 3 A schematic diagram illustrating the acquisition of first map information of a first traffic environment provided in an embodiment of this application;
[0061] Figure 4 Another flowchart illustrating the map information processing method provided in this application embodiment;
[0062] Figure 5 Another flowchart illustrating the map information processing method provided in this application embodiment;
[0063] Figure 6 A schematic diagram of the area in the first image where the lane boundary line is located in the first image from a perspective provided in an embodiment of this application;
[0064] Figure 7 A schematic diagram illustrating the generation of evaluation information provided in an embodiment of this application;
[0065] Figure 8 A schematic diagram illustrating the beneficial effects provided by the embodiments of this application;
[0066] Figure 9 A schematic diagram of the first scoring provided in an embodiment of this application;
[0067] Figure 10 A schematic flowchart illustrating a training method for a model provided in an embodiment of this application;
[0068] Figure 11 A flowchart illustrating a method for obtaining map information provided in an embodiment of this application;
[0069] Figure 12 Another flowchart illustrating the method for obtaining map information provided in this application embodiment;
[0070] Figure 13 A schematic diagram of a map information processing apparatus provided in an embodiment of this application;
[0071] Figure 14 A schematic diagram of a map information acquisition device provided in an embodiment of this application;
[0072] Figure 15 A schematic diagram of the device provided in an embodiment of this application;
[0073] Figure 16 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation
[0074] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0075] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0076] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.
[0077] The method provided in this application can be applied to the field of intelligent driving. For example, it can be applied to scenarios in intelligent driving that require map information, such as when the intelligent driving system in a vehicle performs trajectory planning, it needs to obtain map information corresponding to the surrounding traffic environment; or, for example, when the intelligent driving system in a vehicle provides navigation functions, it needs to obtain map information corresponding to the surrounding traffic environment, etc. This application does not exhaustively list all scenarios requiring map information. For example, the vehicle can be a car, truck, motorcycle, bus, ship, airplane, helicopter, recreational vehicle, amusement park vehicle, tram, golf cart, or train, etc., and this application does not impose any particular limitation.
[0078] In related technologies, the map information used is often two-dimensional (2D) map information of the traffic environment. This 2D map information includes the two-dimensional coordinates of the location points in the road elements of the traffic environment. However, in some traffic scenarios, such as uphill and downhill traffic scenarios, the 2D map information does not include the height of each location point, which makes it impossible for the 2D map information to accurately reflect the real traffic environment. In order to solve the aforementioned problem, this application discloses that: the first device not only acquires the first map information of the first traffic environment, which includes the two-dimensional coordinates of at least one first location point, and the at least one first location point is a location point in the road elements included in the first traffic environment; it also acquires the three-dimensional (3D) information of the first traffic environment, which includes the height of at least one location area, and the objects in the first traffic environment include at least one location area; then the first device can obtain the height of at least one first location point based on the first map information and the 3D information, and the updated first map information includes the two-dimensional coordinates of at least one first location point and the height of at least one first location point; that is, it provides a map information generation scheme that simultaneously includes the two-dimensional coordinates and height of location points, which helps to make the updated first map information more accurately reflect the real traffic environment.
[0079] Optionally, the updated first map information obtained by the method provided in this application can be used as training data for a machine learning model (hereinafter referred to as the "first machine learning model" for ease of distinction). The input of the first machine learning model can include an image under the perspective view (PV) of the traffic environment (hereinafter referred to as "PV view"). The output of the first machine learning model can include map information, which includes the two-dimensional coordinates and height of the location points in the road elements included in the traffic environment. That is, the updated first map information generated by the method provided in this application is used to train the first machine learning model. The trained machine learning model can directly obtain map information carrying height based on the image under the PV view.
[0080] Based on the above description, the specific implementation flow of the map information processing method provided in this application will now be described. Please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic flowchart of a map information processing method provided in an embodiment of this application. The map information processing method provided in an embodiment of this application may include:
[0081] 101. Obtain first map information of the first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point, and the at least one first location point is a location point in the road elements included in the first traffic environment.
[0082] In this application, the first traffic environment can also be referred to as the first road environment, and the first map information of the first traffic environment can be understood as the two-dimensional (2D) map information of the first traffic environment.
[0083] Optionally, the first map information can be 2D map information from a bird's-eye view (BEV) (hereinafter referred to as "BEV view"). The BEV view in this application can also be called a top-down view. Alternatively, the first map information can also be map information from a perspective view (PV), etc. The specific information can be determined based on the actual application scenario.
[0084] The first map information can indicate the location of road elements in the first traffic environment. Optionally, the road element can be a road element on the ground. For example, the road element can include lane boundary lines. In this application, the lane boundary lines can be lines existing in real roads used to divide different lanes. Optionally, the road element can also include at least one of the following: lane center lines, lanes, road boundary lines, intersections, sidewalks, turn indicators, or other road elements, etc., which are not exhaustively listed here. In this application, the lane center line can be understood as the virtual centerline of the lane. The lane center line can be a line that does not exist in real roads. The lane center line can be used to locate the lane. For example, the first map information can include the two-dimensional coordinates of at least one first location point. Each first location point is a location point among the road elements included in the first traffic environment. The location of the road element in the first traffic environment can be indicated by the two-dimensional coordinates of the aforementioned at least one first location point. Optionally, the first map information also includes the connection relationship between different first location points among the aforementioned at least one first location point, so that at least one first location point can be converted into a line segment based on the aforementioned connection relationship, thereby more accurately indicating the location of the road element in the first traffic environment.
[0085] The two-dimensional coordinates of the first position point can include coordinate information along the X-axis and Y-axis, with the X and Y axes perpendicular to each other. For example, the two-dimensional coordinates of the first position point can be coordinate information in a first coordinate system. If the first coordinate system is a vehicle coordinate system, then the aforementioned X-axis and Y-axis directions can be the X-axis and Y-axis directions in the vehicle coordinate system. Alternatively, if the first coordinate system is a world coordinate system (which can also be called an Earth coordinate system), then the X-axis direction can be consistent with the longitude direction in the Earth coordinate system, and the Y-axis direction can be consistent with the latitude direction in the Earth coordinate system. Alternatively, the first coordinate system can also be coordinate information in other coordinate systems (such as a camera coordinate system or a radar sensor coordinate system), depending on the specific application scenario.
[0086] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of the first map information provided in an embodiment of this application. Figure 2 Taking the first map information as the map information from the BEV's perspective as an example, such as Figure 2 As shown, the first map information includes the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment. The first map information can indicate the location of the road elements in the first traffic environment. Figure 2 Taking the road elements in the first traffic environment, including lane boundary lines, lanes, intersections, and road boundary lines, as an example, it should be understood that... Figure 2 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0087] For example, the first map information of the first traffic environment can be obtained based on at least one image of the first traffic environment and / or point cloud data of the first traffic environment. The aforementioned at least one image of the first traffic environment can be at least one second image from the PV perspective of the first traffic environment. For example, the second image from the PV perspective of the first traffic environment can be acquired by a photoelectric sensor, such as a camera or event camera. The aforementioned point cloud data can be obtained by an ultrasonic sensor, a lidar sensor, a millimeter-wave radar sensor, or other sensors capable of measuring and acquiring point cloud data. It should be noted that since the concept of "first image" will be used in the following description, the concept of "first image" will not be introduced here.
[0088] In one scenario, the first map information of the first traffic environment is obtained based on the second and third map information. Both the second and third map information are from the BEV (Browser-Electric Vehicle) perspective of the first traffic environment; that is, both include the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment. The two-dimensional coordinates of each location point included in the second map information are in an absolute world coordinate system, while the two-dimensional coordinates of each location point included in the third map information are in a relative world coordinate system. For example, the X-axis direction of both the absolute and relative world coordinate systems is consistent with the longitude direction in the Earth coordinate system, and the Y-axis direction of both the absolute and relative world coordinate systems is consistent with the dimensional direction in the Earth coordinate system, but the origins of the absolute and relative world coordinate systems are different.
[0089] For example, both the second and third map information are obtained based on at least one second image and / or point cloud data of the first traffic environment from a PV perspective, collected by sensors deployed on the vehicle (hereinafter referred to as the "first vehicle" for convenience). The origin of the relative world coordinate system can be obtained based on the position of the first vehicle. Optionally, the origin of the relative world coordinate system can be the starting position of the first vehicle. In other words, the origin of the relative world coordinate system can be the position that the first vehicle needs to operate in order to collect at least one second image and / or point cloud data of the first traffic environment from a PV perspective. The starting position of the first vehicle can be the position where the first vehicle starts running.
[0090] For example, the second map information is obtained by inputting at least one second image and / or point cloud data of the first traffic environment from the PV perspective of the first traffic environment into the second machine learning model, thereby generating the second map information of the first traffic environment. The second map information includes the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment. The aforementioned two-dimensional coordinates include coordinate information in the X-axis and Y-axis directions in the absolute world coordinate system. For example, the second machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, an attention-based neural network, a multilayer perceptron, or other types of machine learning models. For example, the second machine learning model can adopt a U-shaped encoder-decoder network. It should be understood that the example here is only to demonstrate the feasibility of this solution, and the specific model used for the second machine learning model can be determined in combination with the actual application scenario.
[0091] For example, the third map information is obtained by inputting at least one second image and / or point cloud data of the first traffic environment from the PV perspective of the first traffic environment into the third machine learning model, thereby generating the third map information of the first traffic environment. The third map information includes the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment. The aforementioned two-dimensional coordinates include coordinate information in the X-axis direction and Y-axis direction relative to the world coordinate system. The specific manifestation of the third machine learning model can be referred to the above description of the second machine learning model, which will not be repeated here.
[0092] Optionally, if the two-dimensional coordinates included in the first map information are coordinates in a vehicle coordinate system, step 101 may, for example, include: the first device can convert the two-dimensional coordinates included in the second map information from an absolute world coordinate system to a vehicle coordinate system to obtain updated second map information. The updated second map information includes the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment in the vehicle coordinate system. For example, the aforementioned vehicle coordinate system may be the vehicle coordinate system of the first vehicle. The first device can also convert the two-dimensional coordinates included in the third map information from a relative world coordinate system to a vehicle coordinate system to obtain updated third map information. The updated third map information includes the two-dimensional coordinates of at least one location point among the road elements included in the first traffic environment in the vehicle coordinate system. The first device can determine the first map information of the first traffic environment based on the updated second map information and the updated third map information. For example, the first device can perform a fusion operation based on the updated second map information and the updated third map information to obtain the first map information of the first traffic environment.
[0093] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram illustrating the acquisition of first map information of a first traffic environment provided in an embodiment of this application, such as... Figure 3 As shown, second map information in the absolute world coordinate system can be obtained, and then transformed from the absolute world coordinate system to the vehicle coordinate system to obtain updated second map information; third map information in the relative world coordinate system can be obtained, and then transformed from the relative world coordinate system to the vehicle coordinate system to obtain updated third map information; the updated second map information and the updated third map information can be merged to obtain first map information in the vehicle coordinate system. It should be understood that... Figure 3 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0094] Alternatively, if the two-dimensional coordinates included in the first map information are coordinates in an absolute world coordinate system, step 101 may, for example, include: the first device can convert the two-dimensional coordinates included in the third map information from a relative world coordinate system to an absolute world coordinate system to obtain updated third map information. The first device determines first map information for the first traffic environment based on the second map information and the aforementioned updated third map information; for example, the first device can perform a fusion operation based on the second map information and the aforementioned updated third map information to obtain first map information for the first traffic environment.
[0095] In this embodiment, the first map information of the first traffic environment is obtained based on map information in two different coordinate systems, which is beneficial to obtaining more accurate map information.
[0096] In another scenario, the first map information of the first traffic environment can be obtained directly based on at least one second image from the PV perspective of the first traffic environment and / or point cloud data of the first traffic environment. For example, the first map information of the first traffic environment can be obtained by inputting at least one second image from the PV perspective of the first traffic environment and / or point cloud data of the first traffic environment into a machine learning model, thereby generating the first map information. The specific manifestation of the aforementioned machine learning model can be found in the above description of the second machine learning model, which will not be repeated here.
[0097] 102. Obtain 3D information of the first traffic environment, the 3D information including the height of at least one location region, and objects in the first traffic environment including at least one location region.
[0098] For example, the 3D information of the first traffic environment may include the second coordinate information of each location region in the first traffic environment in a second coordinate system. The second coordinate information of each location region may include the two-dimensional coordinates and height of each location region. In this application, the height may also be referred to as elevation. The two-dimensional coordinates of each location region may include position parameters in the X-axis direction and Y-axis direction of the second coordinate system. The height of each location region can be understood as the position parameter in the Z-axis direction of the second coordinate system. The Z-axis direction may be perpendicular to the plane formed by the X-axis direction and the Y-axis direction. In other words, the second coordinate information of each location region may include the position parameters of each location region in the X-axis direction, Y-axis direction and Z-axis direction of the second coordinate system.
[0099] For example, each location region can be represented as a cube, and the aforementioned at least one location region can be understood as dividing an object in the first traffic environment into the aforementioned at least one location region. For example, the 3D information of the first traffic environment can be obtained based on at least one second image from the perspective (PV) view of the first traffic environment and / or point cloud data of the first traffic environment. Optionally, the 3D information of the first traffic environment can be obtained based on multiple second images from the perspective (PV) view of the first traffic environment and / or point cloud data of the first traffic environment; for example, the multiple second images from the PV view are acquired by multiple first sensors of the first vehicle, and the multiple first sensors may include at least two of the following: a camera at the left front, a camera at the front, a camera at the right front, a camera at the left rear, a camera at the rear, and a camera at the right rear, etc. The specific images from the PV view used can be determined in conjunction with the actual application scenario.
[0100] For example, objects in the first traffic environment may include road elements in the first traffic environment. Objects in the first traffic environment may also include other objects besides road elements. For example, other objects besides road elements in the first traffic environment may include traffic lights, streetlights, trees, buildings, fire hydrants, warning triangles, cones, warning posts, crash barriers, construction signs, traffic guidance signs, or other types of static obstacles. Optionally, other objects besides road elements in the first traffic environment may also include dynamic obstacles in the first traffic environment. For example, the aforementioned dynamic obstacles may include vehicles, electric vehicles, bicycles, pedestrians, animals, or other dynamic obstacles. It should be understood that the examples here are only for the convenience of understanding this solution, and the specific information included in the objects in the first traffic environment can be determined in combination with the actual application scenario.
[0101] Optionally, the 3D information of the first traffic environment can be understood as a 3D model of objects in the first traffic environment based on at least one second image and / or point cloud data of the first traffic environment from a PV perspective. Each location region can also be understood as a volume pixel. For example, at least one second image and / or point cloud data of the first traffic environment from a PV perspective can be input into a machine learning model to obtain the 3D information of the first traffic environment generated by the machine learning model. The specific manifestation of the aforementioned machine learning model can be referred to the above description of the second machine learning model, and will not be repeated here.
[0102] For example, the 3D information of the first traffic environment can be general obstacle detection (GOD) data of the first traffic environment; the GOD data of the first traffic environment is obtained based on at least one second image from the PV viewpoint of the first traffic environment and / or point cloud data of the first traffic environment. In the embodiments of this application, the information on which the 3D information of the first traffic environment is obtained is clearly defined, which is beneficial to improving the feasibility of this solution.
[0103] Optionally, the second coordinate system can be a vehicle coordinate system. For example, at least one second image and / or point cloud data of the first traffic environment from the PV perspective of the first traffic environment are acquired by sensors deployed in the first vehicle, and the vehicle coordinate system can be the vehicle coordinate system of the first vehicle; or, the second coordinate system can also be a world coordinate system or other types of coordinate systems (such as camera coordinate system or radar sensor, etc.), which can be determined in combination with the actual application scenario.
[0104] Optionally, the position parameters of each first position point in the X-axis and Y-axis directions of the first coordinate system can be real numbers; in other words, the position parameters of each first position point in the X-axis and Y-axis directions of the first coordinate system can be non-integers. The position parameters of each position region in the X-axis, Y-axis and Z-axis directions of the second coordinate system are all integers.
[0105] For example, the first coordinate information of a first position point can be (5.6, 7.8), which means that the position parameter of the first position point in the X-axis direction of the first coordinate system is 5.6, and the position parameter of the first position point in the Y-axis direction of the first coordinate system is 7.8; the second coordinate information of a position region can be (6,8,2), which means that the position parameter of the position region in the X-axis direction of the second coordinate system is 3, the position parameter of the position region in the Y-axis direction of the second coordinate system is 4, and the position parameter of the position region in the Z-axis direction is 2. It should be understood that the examples here are only for the convenience of understanding the difference between the first coordinate information of the first position point and the second coordinate information of the position region.
[0106] 103. Based on the first map information and 3D information, obtain the height of at least one first location point. The updated first map information includes the two-dimensional coordinates of at least one first location point and the height of at least one first location point.
[0107] For example, after acquiring the first map information and the 3D information of the first traffic environment, the first device can obtain the height of each of the at least one first location points based on the first map information and the 3D information of the first traffic environment, that is, obtain the updated first map information. The updated first map information includes the two-dimensional coordinates and height of each of the at least one first location points.
[0108] For example, in one implementation, any one of the at least one first location points is referred to as a "second location point". Step 103 may include: the first device can determine the height of the second location point based on the height of a target location region in at least one location region, where the target location region is the location region corresponding to the second location point in at least one location region. In other words, the aforementioned step can be understood as performing high-value adsorption on the second location point; the first device can repeat the aforementioned step at least once to obtain the height of each of the at least one first location points. In the embodiments of this application, for any one of the at least one first location points (i.e., the second location point), the height of the target location region corresponding to the first location point in at least one location region will be obtained, and then the height of the second location point will be determined based on the height of the target location region. That is, each of the at least one first location points will be traversed, and the height of each first location point will be obtained one by one. This more refined management of the process of obtaining the height of at least one first location point is beneficial to improving the accuracy of the height of each first location point.
[0109] For example, the first device can determine at least one target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region in at least one location region; furthermore, the first device can determine the height of the second location point based on the height of the aforementioned at least one target location region. In this embodiment, a simple and effective method for determining target location regions is provided, based on the two-dimensional coordinates of the first location point and the two-dimensional coordinates of the location regions in 3D information to find the target location region corresponding to each first location point.
[0110] Optionally, the first device can determine the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, and select at least one target location region from at least one location region that is closest to the two-dimensional coordinates of the second location point. The distance between the two-dimensional coordinates of each target location region in the at least one target location region and the two-dimensional coordinates of the second location point can all be the same. For example, if the at least one target location region includes at least two target location regions, the two-dimensional coordinates of the different target location regions in the at least two target location regions can all be the same, the difference being that they are different in height. Alternatively, the first device can round up both the position parameters in the X-axis direction and the position parameters in the Y-axis direction of the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region. Alternatively, the first device can round down both the position parameters in the X-axis direction and the position parameters in the Y-axis direction of the two-dimensional coordinates of the second location point to obtain the two-dimensional coordinates of each target location region in the at least one target location region. The specific implementation method can be determined according to the actual application scenario.
[0111] For example, if the at least one target location area includes only one target location area, the first device can directly determine the height of the single target location area as the height of the second location point. If the at least one target location area includes at least two target location areas, in one implementation, the first device can determine the minimum height among the at least two heights corresponding to the at least two target location areas as the height of the second location point. Since the second location point can be a location point on the road surface, directly determining the minimum height as the height of the second location point is feasible. Alternatively, in another implementation, the first device can determine the average height of the at least two heights corresponding to the at least two target location areas as the height of the second location point, etc.
[0112] Optionally, the first device determines the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region. This may include: the first device determining the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region in at least one location region, and the semantic type of each location region in at least one location region.
[0113] For example, the semantic type of a location area can be the semantic type of lane boundary lines, road boundary lines, sidewalks, intersections, turn indicators, traffic lights, streetlights, trees, buildings, fire hydrants, warning triangles, cones, warning posts, crash barriers, construction signs, traffic guidance signs, or other objects that may appear in the traffic environment. Optionally, if the objects in the first traffic environment also include dynamic obstacles in the first traffic environment, the semantic type of a location area can also be the semantic type of vehicles, electric vehicles, bicycles, pedestrians, animals, or other dynamic obstacles, etc. The example here is only for the convenience of understanding this scheme.
[0114] For example, the first device can determine the distance between the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, and select at least one candidate location region from at least one location region that is closest to the two-dimensional coordinates of the second location point. The distance between the two-dimensional coordinates of each of the at least one candidate location regions and the two-dimensional coordinates of the second location point can all be the same. For example, if the at least one candidate location region includes at least two candidate location regions, the two-dimensional coordinates of the different candidate location regions of the at least two candidate location regions can all be the same, differing only in height. In one case, the first device can determine a target location region from the at least one candidate location region that has the same semantic type as the second location point, based on the semantic type of each candidate location region. In another case, optionally, if each of the at least one first location points is a location point on a lane boundary line, the first device can also determine a target location region with the semantic type of a lane boundary line from the at least one candidate location region based on the semantic type of each candidate location region. Furthermore, the first device can determine the height of the aforementioned target location region as the height of the second location point.
[0115] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 4 , Figure 4 This is another flowchart illustrating the map information processing method provided in an embodiment of this application. Figure 4 The content can be combined with the above. Figure 3 The description is provided below, and repeated parts will not be elaborated here. The process involves obtaining the first map information and 3D information of the first traffic environment in the vehicle coordinate system. The first map information includes the two-dimensional coordinates of at least one first location point, and the 3D information includes the height of at least one location region. The height of the target location region corresponding to each first location point can be obtained from the height of the at least one location region. Based on the height of the target location region corresponding to each first location point, the height of each first location point is determined. This means that height snapping is performed on each first location point, resulting in the updated first map information of the first traffic environment. This should be understood as... Figure 4 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0116] In this embodiment of the application, in the process of determining the target location area corresponding to any first location point, not only the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area are used, but also the semantic type of each location area is used. Since the matching is based solely on the two-dimensional coordinates of the first location point and the two-dimensional coordinates of each location area, a certain first location point may correspond to multiple location areas. The two-dimensional coordinates of the aforementioned multiple location areas are consistent, but the heights are inconsistent. Therefore, the target location area corresponding to the first location point can be further determined based on the semantic type of each location area. This is beneficial for accurately finding the target location area corresponding to each first location point, and thus for obtaining the accurate height of each first location point, thereby improving the accuracy of the updated first map information.
[0117] In another implementation, step 103 may include: the first device may also input the first map information and 3D information of the first traffic environment into the machine learning model to obtain the updated first map information generated by the machine model.
[0118] Optionally, during its journey, the first vehicle can acquire multiple sets of images from a PV perspective and / or multiple sets of point cloud data of the first traffic environment. Each set of images from a PV perspective may include at least one second image from the PV perspective. In this case, the first device can acquire multiple sets of first map information and multiple sets of 3D information. Each set of first map information is obtained based on a set of images from a PV perspective of the first traffic environment and / or a point cloud data of the first traffic environment. Similarly, each set of 3D information is obtained based on a set of images from a PV perspective of the first traffic environment and / or a point cloud data of the first traffic environment. Correspondingly, the first device can obtain an updated first map information based on a first map information and a 3D information. Then, the first device can obtain multiple updated first map information by repeatedly executing steps 101 to 103 based on multiple consecutive first map information and multiple consecutive 3D information. Since there are overlapping position areas in two adjacent updated first map information in multiple consecutive updated first map information, multiple consecutive updated first map information may include duplicate first position points. In other words, multiple consecutive updated first map information may include at least two two-dimensional coordinates and height of duplicate first position points.
[0119] The first device can perform a smoothing operation based on at least two heights of repeated first location points in multiple consecutive updated first map information to obtain the final height of the repeated first location points; optionally, the first device can also perform a smoothing operation based on at least two two-dimensional coordinates of the repeated first location points to obtain the final two-dimensional coordinates of the repeated first location points.
[0120] For example, the first device can average at least two heights of the first position point to obtain the final height of the first position point; the first device can also average at least two position parameters of the first position point in the X-axis direction of the first coordinate system to obtain the final position parameter of the first position point in the X-axis direction of the first coordinate system; and average at least two position parameters of the first position point in the Y-axis direction of the first coordinate system to obtain the final position parameter of the first position point in the Y-axis direction of the first coordinate system. As another example, the first device can determine the median value of at least two heights of the first position point as the final height of the first position point; the first device can also determine the median value of at least two position parameters of the first position point in the X-axis direction of the first coordinate system as the final position parameter of the first position point in the X-axis direction of the first coordinate system; and so on. The specific smoothing method used can be determined based on the actual application scenario and is not limited here.
[0121] Optionally, in the above Figure 1 Based on the description of the corresponding embodiments, please refer to Figure 5 , Figure 5 This is another schematic flowchart illustrating the map information processing method provided in this application embodiment. The map information processing method provided in this application embodiment may include:
[0122] 501. Obtain first map information of the first traffic environment, the first map information including the two-dimensional coordinates of at least one first location point, the at least one first location point being a location point in the road elements included in the first traffic environment.
[0123] 502. Obtain three-dimensional information of the first traffic environment, the three-dimensional information including the height of at least one location region, and objects in the first traffic environment including at least one location region.
[0124] 503. Based on the first map information and 3D information, obtain the height of at least one first location point. The updated first map information includes the two-dimensional coordinates of at least one first location point and the height of at least one first location point.
[0125] The specific implementation of steps 501 to 503 performed by the first device, as well as the meanings of the terms used in steps 501 to 503, can be found in the above-mentioned text. Figure 1 The descriptions in the corresponding embodiments will not be repeated here.
[0126] 504. Based on the first image under the perspective of the first traffic environment, determine the evaluation information corresponding to the updated first map information. The evaluation information indicates the accuracy of the updated first map information.
[0127] Step 504 is an optional step. In one case, at least one first location point includes a location point within the lane boundary line of the first traffic environment. In other words, all of the above-mentioned at least one first location point can be a location point within the lane boundary line of the first traffic environment. Step 504 may include: the first device projects each of the at least one first location point onto a first image under a perspective view based on the two-dimensional coordinates and height of each of the at least one first location point to obtain first location information. The first location information may include the position of each of the at least one projection point corresponding to the at least one first location point in the first image under the perspective view. The aforementioned at least one projection point can be understood as a projection point obtained by projecting a location point on the lane boundary line under the BEV perspective onto the first image under the PV perspective. Then, the first device can determine evaluation information based on the first location information and the first image under the perspective view.
[0128] In this embodiment of the application, a scheme is proposed to back-project the first location point from the BEV perspective onto the first image from the PV perspective based on the two-dimensional coordinates and height of the first location point in the updated first map information, thereby obtaining the location information of the projected point. Then, the accuracy of the updated first map information is evaluated based on the location information of the projected point. Since the first image from the PV perspective of the first traffic environment can accurately reflect the first traffic environment, the method of back-projecting the first location point from the BEV perspective onto the first image from the PV perspective can make the most of the first image from the PV perspective, which is conducive to obtaining more accurate evaluation information.
[0129] For example, the first device can project each of the at least one first position points into the first image under the perspective view based on the two-dimensional coordinates and height of each of the at least one first position points, the first image under the perspective view, and the intrinsic parameters of the camera that captured the first image under the perspective view.
[0130] The following describes a specific implementation method for the first device to obtain evaluation information based on the first location information. In one implementation, the first device performs semantic recognition operations on a first image from a perspective view to obtain second location information. The second location information includes the location information of pixels in the first image from the perspective view whose semantic type is lane boundary line. Optionally, the first device can input the first image from the perspective view into a machine learning model, and use the aforementioned machine learning model to perform target detection on the lane boundary line in the first image from the perspective view. In other words, the machine learning model identifies pixels in the first image from the perspective view whose semantic category is lane boundary line, and obtains the second location information generated by the machine learning model. For example, the aforementioned machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, an attention-based neural network, or other types of neural networks, which can be determined based on the actual application scenario. Then, the first device can determine the evaluation information based on the first location information and the second location information.
[0131] For example, the second location information may include the coordinate information of pixels in the first image with the semantic type of lane boundary line in the perspective view.
[0132] Optionally, the first position information may indicate the position information of at least one first lane boundary line formed by projecting a first position point on the lane boundary line from the BEV perspective onto a first image from the PV perspective, and the second position information may indicate the position information of at least one second lane boundary line in the first image from the PV perspective. For example, the first device may determine a first distance between at least one first lane boundary line and at least one second lane boundary line based on the first and second position information, and then determine the evaluation information based on the first distance between the at least one first lane boundary line and at least one second lane boundary line.
[0133] For example, the first device can determine a second distance between each first lane boundary line and the nearest second lane boundary line among the aforementioned at least one second lane boundary lines, that is, obtain at least one second distance corresponding one-to-one with at least one first lane boundary line; for example, the aforementioned distance can be obtained by calculating cosine distance, Euclidean distance, Mahalanobis distance, L1 distance, or other algorithms used to calculate the distance between lines. Furthermore, the first device can obtain a first distance based on the aforementioned at least one second distance, for example, the first distance can be the average of at least one second distance, the first distance can be the median of at least one second distance, the first distance can be the maximum value among at least one second distance, etc., which can be determined according to the actual application scenario.
[0134] In one scenario, the aforementioned first distance can be directly determined as the evaluation information, where a smaller first distance indicates higher accuracy of the updated first map information, and a larger first distance indicates lower accuracy of the updated first map information. In another scenario, the evaluation information may include a first score indicating the accuracy of the updated first map information. A higher first score indicates higher accuracy of the updated first map information, and a lower first score indicates lower accuracy of the updated first map information. Therefore, the aforementioned first score can be obtained based on the first distance; a smaller first distance results in a higher first score, representing higher accuracy of the updated first map information, and a larger first distance results in a lower first score, representing lower accuracy of the updated first map information.
[0135] Furthermore, for example, after the first device determines at least one first lane boundary line based on the first location information and at least one second lane boundary line based on the second location information, it can first pair the aforementioned at least one first lane boundary line and at least one second lane boundary line to obtain the nearest second lane boundary line for each first lane boundary line, and then calculate the second distance between each first lane boundary line and the nearest second lane boundary line. For example, in this application, any one of the at least one first lane boundary lines is referred to as the target lane boundary line. After the first device determines at least one first lane boundary line based on the first location information and at least one second lane boundary line based on the second location information, it can calculate the distance between the target lane boundary line and each of the at least one second lane boundary lines to obtain at least one distance corresponding to the at least one second lane boundary line. The smallest of the aforementioned at least one distance is determined as the second distance between the target lane boundary line and the nearest second lane boundary line among the at least one second lane boundary lines. The first device can repeat the aforementioned steps at least once to obtain the second distance between each first lane boundary line and the nearest second lane boundary line, etc. It should be understood that the example here is only to demonstrate the feasibility of this solution, and the specific method of calculating the second distance can be determined in combination with the actual application scenario.
[0136] In this embodiment, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line formed by the projection points corresponding to the first location point in the first image under the PV view, and the position of the lane boundary line in the first image under the PV view. That is, a specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, which improves the feasibility of this scheme. In addition, based on the position of the lane boundary line formed by the projection points and the original position of the lane boundary line in the first image under the PV view, the accuracy of the lane boundary line in the updated first map information can be obtained. Since the lane boundary lines in the traffic environment are often long, and it is often difficult to accurately reproduce the lane boundary line situation in the actual environment in the map information of traffic scenarios such as uphill and downhill, evaluating the accuracy of the updated first map information from the dimension of the lane boundary line can obtain more accurate evaluation information, and it is beneficial to improve the accuracy of the position information of the lane boundary line in the updated first map information.
[0137] In another implementation, the first device performs image segmentation based on a first image from a perspective view to obtain third location information. This third location information indicates the position of the lane boundary line within the first image from the perspective view. For example, the third location information indicates the position of at least one first image region containing the lane boundary line in the first image from the perspective view. Each first image region containing the lane boundary line can be represented as a polygon, thus the third location information can indicate the position of each first image region within the at least one polygonal first image region containing the lane boundary line. Optionally, the first device can input the first image from the perspective view into a machine learning model, and perform image segmentation on the first image from the perspective view using the machine learning model to obtain the third location information generated by the machine learning model. For example, the machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, an attention-based neural network, or other types of neural networks, which can be determined based on the actual application scenario. The first device can then determine evaluation information based on the first location information and the third location information.
[0138] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 6 , Figure 6 This is a schematic diagram of the area in the first image where the lane boundary line is located, provided as an embodiment of this application, from a perspective view. Figure 6 This is a schematic diagram obtained after image segmentation of the first image from the perspective of the first traffic environment. Figure 6The image shows the first image from a perspective view divided into four regions: a pure black region representing the background area, which can include objects other than road elements in the first traffic environment, such as vehicles, blue sky, or buildings; a white region representing the fence at the road boundary; a light gray region representing the road; and a dark gray region representing the lane boundary lines. Figure 6 The medium-dark gray area represents at least one first image region containing the lane boundary line in the first image from a perspective view. Figure 6 It includes multiple first image regions, due to the original Figure 6 After grayscale processing, the colors of the pure black areas and the dark gray areas are somewhat similar. To distinguish these two different colors, Figure 6 The background area is marked within a pure black area, and an arrow indicates a specific first image area among multiple first image areas. It should be understood that... Figure 6 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0139] Optionally, the first device can also determine at least one second image region in the first image under the PV view based on the first location information. Each second image region can be obtained by increasing the width of each first lane boundary line. For example, the width of the first lane boundary line can be increased by a preset width to obtain the second image region corresponding to each first lane boundary line. The first device can determine the intersection-over-union (IoU) ratio between at least one first image region and at least one second image region, that is, determine the ratio between the intersection and union of all first image regions and all second image regions; and then determine the evaluation information based on the IoU ratio between at least one first image region and at least one second image region.
[0140] In one scenario, the aforementioned intersection-over-union ratio (IoU) can be directly determined as the evaluation information. A higher IoU between at least one first image region and at least one second image region indicates higher accuracy of the updated first map information, while a lower IoU between the two regions indicates lower accuracy. In another scenario, the evaluation information may include a second score indicating the accuracy of the updated first map information. A higher second score indicates higher accuracy, and a lower second score indicates lower accuracy. The second score can be obtained based on the IoU between at least one first image region and at least one second image region. A higher IoU results in a higher second score, indicating higher accuracy of the updated first map information, while a lower IoU results in a lower second score, indicating lower accuracy.
[0141] In this embodiment of the application, the evaluation information of the updated first map information is obtained based on the position of the lane boundary line formed by the projection points corresponding to the first position point in the first image under the PV view, and the position of the image area where the lane boundary line is located in the first image under the PV view. That is, another specific implementation scheme for obtaining the evaluation information of the updated first map information is provided, which improves the implementation flexibility of this scheme.
[0142] In another implementation, the first device determines a second distance between each first lane boundary line and its nearest second lane boundary line based on first and second location information, and then obtains a first distance based on at least one second distance corresponding to at least one first lane boundary line; it also determines the intersection-over-union ratio (IoU) between at least one first image region and at least one second image region based on the first and second location information; the evaluation information includes the first distance and the IoU between at least one first image region and at least one second image region; or, the first device also obtains a first score based on the first distance and a second score based on the IoU between at least one first image region and at least one second image region, the evaluation information including the first score and the second score.
[0143] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 7 , Figure 7 A schematic diagram illustrating the generation of evaluation information provided in an embodiment of this application, such as... Figure 7As shown, the first device projects the first position point on the lane boundary line in the updated first map information from the BEV perspective onto the first image from the perspective view, based on the two-dimensional coordinates and height of the first position point on the lane boundary line in the updated first map information, to obtain the first position information of the projected point. This first position information can indicate the position of at least one first lane boundary line formed by the projected points obtained by projecting the first position point on the lane boundary line from the BEV perspective onto the first image from the perspective view. The first device performs target detection on the first image from the perspective view to obtain second position information, which indicates the position of at least one second lane boundary line in the first image from the perspective view. Based on the first position information and the second position information, the first device obtains a first distance between at least one first lane boundary line and at least one second lane boundary line. The first device can also perform image segmentation on the first image under the perspective view to obtain third position information. The third position information indicates the position of each first image region in at least one first image region where the lane boundary line is located in the first image under the perspective view. The first device can also determine at least one second image region in the first image under the PV view based on the first position information. Furthermore, it determines the intersection-over-union ratio (IoU) between at least one first image region and at least one second image region. This evaluation information may include the aforementioned first distance and the aforementioned IoU. It should be understood that... Figure 7 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0144] Optionally, the first device may also filter the updated first map information based on evaluation information, thereby deleting the updated first map information that does not meet the filtering conditions. For example, the filtering conditions may include: a first distance less than or equal to a distance threshold, and / or, a first score greater than or equal to a first score threshold, and / or, and / or, an IoU between at least one first image region and at least one second image region greater than or equal to a preset ratio, and / or, a second score greater than or equal to a second score threshold.
[0145] In another implementation, after obtaining the first location information, the first device can determine the position of at least one first lane boundary line composed of projection points in the first image under the PV view based on the first location information. Then, it can display the first image under the PV view containing the aforementioned at least one first lane boundary line. In other words, the first image under the PV view has the aforementioned at least one first lane boundary line added to it. The evaluation information can be the first image under the PV view containing the aforementioned at least one first lane boundary line. Then, the technician can visually check the degree of fit between the at least one first lane boundary line and the lane boundary line in the first image under the PV view to determine the accuracy of the updated first map information.
[0146] In another scenario, if the updated first map information includes at least one first location point, which includes not only location points within the lane boundary lines of the first traffic environment but also location points within other road elements within the first traffic environment, step 504 may further include: the first device filtering the updated first map information to obtain at least one filtered first location point, wherein the filtered at least one first location point only includes location points within the lane boundary lines; the first device, based on the two-dimensional coordinates and height of each of the filtered at least one first location points, projects each of the filtered at least one first location points onto a first image under a perspective view to obtain first location information, wherein the first location information may include the position of each of at least one projection point corresponding one-to-one with the filtered at least one first location point in the first image under a perspective view, wherein the aforementioned at least one projection point can be understood as a projection point obtained by projecting the location point on the lane boundary line under the BEV perspective onto the first image under the PV perspective; and then the first device may determine evaluation information based on the first location information and the first image under a perspective view.
[0147] For example, the first device can project each of the at least one selected first location points into the first image under the perspective view based on the two-dimensional coordinates and height of each first location point in the at least one selected first location point, the first image under the perspective view, and the intrinsic parameters of the camera that captured the first image under the perspective view.
[0148] The specific implementation method of the first device determining the evaluation information based on the first position information and the first image under the perspective view can be found in the above description, and will not be repeated here.
[0149] In this embodiment, since the image from the PV perspective of the first traffic environment can accurately reflect the situation of the first traffic environment, and the updated first map information is to more accurately reflect the first traffic environment, it is feasible to determine the accuracy of the updated first map information by using the image from the PV perspective of the first traffic environment. The image from the PV perspective of the first traffic environment is easy to obtain, thus providing an easy-to-implement detection scheme for the accuracy of the updated first map information.
[0150] To more intuitively understand the beneficial effects of the method provided in this application, please refer to [link / reference needed]. Figure 8 , Figure 8 A schematic diagram illustrating the beneficial effects provided by the embodiments of this application. Figure 8 Includes two sub-diagrams, upper and lower. Figure 8The above diagram shows the two-dimensional coordinates of the position points on the lane boundary line in the map information based on the BEV perspective of the traffic environment. The aforementioned map information does not include the height of the position points on the lane boundary line. The position of the lane boundary line is formed by projecting the position points on the aforementioned lane boundary line onto the image under the PV perspective. Figure 8 In the above diagram, 100 meters represents the position of the projection point obtained by projecting the position point on the lane boundary line 100 meters away from the current position in the aforementioned map information onto the image under the PV view, and 200 meters represents the position of the projection point obtained by projecting the position point on the lane boundary line 200 meters away from the current position in the aforementioned map information onto the image under the PV view. Figure 8 The following diagram illustrates the two-dimensional coordinates and height of the first location point on the lane boundary line in the updated first map information based on the traffic environment from the BEV perspective. The position of the first lane boundary line is formed by projecting the aforementioned location point on the lane boundary line onto the image under the PV perspective. Figure 8 In the above diagram, 100 meters represents the position of the first location point on the lane boundary line 100 meters from the current position in the updated first map information, projected onto the image under the PV view. 200 meters represents the position of the first location point on the lane boundary line 200 meters from the current position in the updated first map information, projected onto the image under the PV view. (Comparison) Figure 8 As can be seen from the upper and lower part diagrams, in Figure 8 In this uphill scenario, since the updated first map information obtained based on the method provided in this application carries the height of each first location point, the position of the first lane boundary line obtained based on the updated first map information is closer to the position of the lane boundary line in the first image from the perspective view. It should be understood that... Figure 8 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0151] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 9 , Figure 9 This diagram illustrates a first score provided in an embodiment of this application. It describes how, by using the method provided in this application to obtain multiple updated first map information, a first score is obtained for each of the multiple updated first map information. Figure 9 The vertical axis represents the number of updated first map information corresponding to each first score. Figure 9The horizontal axis represents the first score. The higher the first score corresponding to a certain updated first map information, the closer the at least one first lane boundary line obtained based on the updated first map information is to at least one second lane boundary line in the PV view image. Specifically, when the first score corresponding to a certain updated first map information is 1, it means that the at least one first lane boundary line obtained based on the updated first map information is completely aligned with at least one second lane boundary line in the PV view image. When the first score corresponding to a certain updated first map information is greater than or equal to 0, it means that the first distance between the at least one first lane boundary line obtained based on the updated first map information and the at least one second lane boundary line in the PV view image is less than or equal to a distance threshold, that is, the accuracy of the updated first map information is qualified. Figure 9 As shown, among the multiple updated first map information obtained by the method provided in this application, the accuracy of most of the updated first map information is satisfactory. It should be understood that... Figure 9 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0152] In this embodiment, not only is first map information of the traffic environment obtained, which includes the two-dimensional coordinates of at least one first location point among the road elements included in the traffic environment, but also 3D information of the traffic environment is obtained, which includes the height of at least one location area of objects in the traffic environment. Based on the aforementioned first map information and 3D information, the height of at least one first location point is obtained. The updated first map information may include the two-dimensional coordinates and height of at least one location point, providing a map information generation scheme that simultaneously includes the two-dimensional coordinates and height of location points, thereby enabling the updated first map information to more accurately reflect the real traffic environment.
[0153] Optionally, if the updated first map information obtained by the method provided in this application is also used as training data for the first machine learning model, then this application also provides specific implementation processes for the training and application stages of the first machine learning model. The training and application stages of the first machine learning model are described below. Please refer to [link / reference]. Figure 10 , Figure 10 This is a flowchart illustrating a training method for a model provided in an embodiment of this application. The training method for a model provided in an embodiment of this application may include:
[0154] 1001. Input the image from the perspective of the first traffic environment into the first machine learning model to obtain the predicted map information generated by the first machine learning model. The predicted map information includes the predicted two-dimensional coordinates of at least one first location point and the predicted height of at least one first location point. The at least one first location point is a location point among the road elements included in the first traffic environment.
[0155] For example, the training device can input an image from the perspective of the first traffic environment into the first machine learning model to obtain the predicted map information generated by the first machine learning model; alternatively, the training device can input an image from the perspective of the first traffic environment and point cloud data of the first traffic environment into the first machine learning model to obtain the predicted map information generated by the first machine learning model.
[0156] 1002. Based on the predicted map information and the updated first map information of the first traffic environment, a first machine learning model is trained, wherein the updated first map information of the first traffic environment includes the two-dimensional coordinates and expected height of at least one first location point among the road elements included in the first traffic environment, and the height of at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional 3D information of the first traffic environment, wherein the first map information includes the two-dimensional coordinates of at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of objects in the first traffic environment.
[0157] For example, the updated first map information of the first traffic environment can be understood as the training data of the first machine learning model, and the updated first map information of the first traffic environment can be obtained through the above... Figure 1 According to the corresponding embodiment, optionally, the second location point is any one of at least one first location point. The height of the second location point in the updated first map information is obtained based on the height of the target location area corresponding to the second location point in at least one location area. The specific implementation of the updated first map information of the first traffic environment can be referred to the above description, and will not be repeated here.
[0158] The training device can calculate the similarity between the predicted map information and the updated first map information of the first traffic environment, obtaining the value of the loss function. This loss function indicates the similarity between the predicted map information and the updated first map information of the first traffic environment. The goal of training the first machine learning model using the loss function is to improve the similarity between the predicted map information and the updated first map information of the first traffic environment. Based on the value of the loss function, the training device can update the weight parameters of the first machine learning model using the backpropagation algorithm to complete one training iteration of the first machine learning model.
[0159] The training device can repeat steps 1001 and 1002 multiple times to iteratively train the first machine learning model until the convergence condition is met, thus obtaining the trained first machine learning model. The aforementioned convergence condition may include: meeting the convergence condition of the loss function, and / or, the number of times the first machine learning model is iteratively trained reaches a preset number.
[0160] Please continue reading. Figure 11 , Figure 11 This is a flowchart illustrating a method for obtaining map information provided in an embodiment of this application. The method for obtaining map information provided in an embodiment of this application may include:
[0161] 1101. Obtain images from the perspective of the second traffic environment.
[0162] For example, the second device can acquire at least one image of the second traffic environment from a perspective view using the first sensor; alternatively, the second device can also acquire point cloud data of the second traffic environment using the second sensor.
[0163] For example, the second device can be a vehicle, in which case steps 1101 and 1102 can be performed by the intelligent driving system in the second device; or, the second device can also be a cloud server that is communicatively connected to the vehicle.
[0164] 1102. Based on the image from the perspective of the second traffic environment, map information of the second traffic environment is obtained. The map information of the second traffic environment includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information of the second traffic environment is generated by a first machine learning model. The training data of the first machine learning model includes the updated first map information of the first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional 3D information of the first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point. The 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
[0165] For example, the second device can input at least one image (optionally including point cloud data of the second traffic environment) from a perspective view of the second traffic environment into a first machine learning model that has undergone training operations, to obtain map information of the second traffic environment generated by the first machine learning model. The map information of the second traffic environment includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The training data of the first machine learning model includes the image from the perspective view of the first traffic environment and the updated first map information of the first traffic environment. The method for obtaining the updated first map information of the first traffic environment can be referred to the above. Figure 1 The descriptions in the corresponding embodiments will not be repeated here.
[0166] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 12 , Figure 12 Another flowchart illustrating the map information acquisition method provided in this application embodiment is shown below. Figure 12 As shown, multiple images from the perspective of the second traffic environment and point cloud data of the second traffic environment are input into the first machine learning model to obtain map information of the second traffic environment generated by the first machine learning model. The map information of the second traffic environment includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. It should be understood that... Figure 12 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0167] In this embodiment, the updated first map information is used as the training data for the first machine learning model. The input of the first machine learning model includes an image from the perspective of the traffic environment, and the output map information of the first machine learning model includes the two-dimensional coordinates and height of the location points in the traffic environment. That is, the first machine learning model trained can directly generate map information carrying the two-dimensional coordinates and height of the location points. This is beneficial for realizing the real-time generation of map information carrying the two-dimensional coordinates and height of the location points based on the image from the perspective of the traffic environment, so as to obtain accurate map information more conveniently and timely.
[0168] exist Figures 1 to 12 Based on the corresponding embodiments, in order to better implement the above-described solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 13 , Figure 13This is a schematic diagram of a map information processing device provided in an embodiment of this application. The map information processing device 1300 includes: an acquisition module 1301, used to acquire first map information of a first traffic environment, the first map information including the two-dimensional coordinates of at least one first location point, the at least one first location point being a location point among road elements included in the first traffic environment; the acquisition module 1301 is also used to acquire three-dimensional 3D information of the first traffic environment, the 3D information including the height of at least one location region, objects in the first traffic environment including at least one location region; and a processing module 1302, used to obtain the height of at least one first location point based on the first map information and the 3D information, the updated first map information including the two-dimensional coordinates of at least one first location point and the height of at least one first location point.
[0169] Optionally, the second location point is any one of at least one first location point. The processing module is specifically used to determine the height of the second location point based on the height of a target location region in at least one location region, wherein the target location region is the location region corresponding to the second location point in at least one location region.
[0170] Optionally, the 3D information also includes two-dimensional coordinates of each location region. The processing module is specifically used to determine the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region.
[0171] Optionally, the processing module 1302 is specifically used to determine the target location region corresponding to the second location point from at least one location region based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region, and the semantic type of each location region.
[0172] Optionally, the first map information is map information under the bird's-eye view BEV. The map information processing device 1300 further includes: a determination module 1303, used to determine evaluation information corresponding to the updated first map information based on the first image under the perspective view PV of the first traffic environment. The evaluation information indicates the accuracy of the updated first map information.
[0173] Optionally, at least one first location point includes a location point on the lane boundary line in the first traffic environment. The determining module 1303 is specifically used to project the first location point onto a first image under a perspective view based on the two-dimensional coordinates and height of the first location point to obtain first location information. The first location information indicates the position of the projection point corresponding to the first location point in the first image under the perspective view. Evaluation information is determined based on the first location information and the first image under the perspective view.
[0174] Optionally, the determining module 1303 is specifically used to perform semantic recognition operations based on the first image under the perspective view to obtain second position information, the second position information including the position information of pixels with the semantic type of lane boundary line in the first image under the perspective view; and to determine evaluation information based on the first position information and the second position information.
[0175] Optionally, the determining module 1303 is specifically used to perform image segmentation operation based on the first image under the perspective view to obtain third position information, the third position information indicating the position of the lane boundary line in the image region of the first image under the perspective view; and to determine evaluation information based on the first position information and the third position information.
[0176] Optionally, the first map information is obtained based on the second map information and the third map information. Both the second map information and the third map information are map information under the bird's-eye view of the first traffic environment in the BEV. The two-dimensional coordinates included in the second map information are in the absolute world coordinate system, and the two-dimensional coordinates included in the third map information are in the relative world coordinate system. Both the second map information and the third map information are obtained based on the first image and / or point cloud data of the first traffic environment collected by the vehicle. The origin of the relative world coordinate system is obtained based on the position of the vehicle.
[0177] Optionally, the 3D information is obtained based on a second image from the perspective of the first traffic environment and / or point cloud data of the first traffic environment.
[0178] It should be noted that the information interaction and execution process between the modules / units in the map information processing device 1300 are different from those in this application. Figures 1 to 12 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0179] Please continue reading. Figure 14 , Figure 14This is a schematic diagram of a map information acquisition device provided in an embodiment of this application. The map information acquisition device 1400 includes: an acquisition module 1401, used to acquire an image under a perspective view PV of a second traffic environment; and a processing module 1402, used to obtain map information of the second traffic environment based on the image under the perspective view. The map information includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information is generated by a machine learning model. The training data of the machine learning model includes updated first map information of a first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional 3D information of the first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
[0180] Optionally, the second position point is any one of at least one first position point, and the height of the second position point is obtained based on the height of the target position area corresponding to the second position point in at least one position area.
[0181] It should be noted that the information interaction and execution process between the modules / units in the map information acquisition device 1400 are different from those in this application. Figures 1 to 12 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0182] This application also provides a device, see [link to relevant documentation] Figure 15 As shown, Figure 15 This is a schematic diagram of the structure of a device provided in an embodiment of this application. Optionally, device 1500 performs... Figures 1 to 12 The functions of the first device, the second device, or the training device in the corresponding method embodiments.
[0183] The device 1500 includes a memory 1502 and at least one processor 1501. Optionally, the processor 1501 implements the methods in the above embodiments by reading instructions stored in the memory 1502, or the processor 1501 may also implement the methods in the above embodiments by internally stored instructions. When the processor 1501 implements the methods in the above embodiments by reading instructions stored in the memory 1502, the memory 1502 stores instructions for implementing the methods provided in the above embodiments of this application.
[0184] Optionally, at least one processor 1501 is one or more CPUs, either a single-core CPU or a multi-core CPU. The memory 1502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, or optical memory. The memory 1502 stores the instructions of the operating system. After the program instructions stored in the memory 1502 are read by the at least one processor 1501, the device 1500 executes the corresponding operations in the foregoing embodiments.
[0185] Optionally, device 1500 also includes a network interface 1503, which can be a wired interface or a wireless interface, and is used for... Figures 1 to 12 The corresponding method implementations execute the sending and receiving of data.
[0186] It should be understood that network interface 1503 has the functions of receiving and sending data. The functions of "receiving data" and "sending data" can be integrated into the same transceiver interface, or the functions of "receiving data" and "sending data" can be implemented in different interfaces, which is not limited here. In other words, network interface 1503 may include one or more interfaces for implementing the functions of "receiving data" and "sending data".
[0187] After the processor 1501 reads the program instructions from the memory 1502, other functions that the device 1500 can perform are described in the preceding method embodiments.
[0188] Optionally, the device 1500 also includes a bus 1504, through which the processor 1501 and memory 1502 are typically interconnected, or in other ways.
[0189] The device 1500 provided in this application embodiment is used to execute the methods executed by the first device, the second device, or the training device in the above-described method embodiments, and to achieve the corresponding beneficial effects. Figure 15 The specific implementation of the device 1500 shown can be referred to the descriptions in the aforementioned method embodiments, and will not be repeated here.
[0190] This application also provides a vehicle; please refer to [link / reference]. Figure 16 , Figure 16This is a schematic diagram of a vehicle structure provided in an embodiment of this application. The vehicle 100 is configured in a fully or partially intelligent driving mode. For example, the vehicle 100 can control itself while in intelligent driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behaviors of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of the other vehicle performing the possible behavior, and control the vehicle 100 based on the determined information. When the vehicle 100 is in intelligent driving mode, the vehicle 100 can also be set to operate without human interaction.
[0191] Vehicle 100 may include various subsystems, such as a mobility system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power supply 110, a computer system 112, and a user interface 116. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 100 may be interconnected via wired or wireless means.
[0192] The mobility system 102 may include components that provide powered motion to the vehicle 100. In one embodiment, the mobility system 102 may include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.
[0193] Engine 118 can be an internal combustion engine, an electric motor, an air-compressed engine, or other combinations of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. Engine 118 converts energy source 119 into mechanical energy. Examples of energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy source 119 can also provide energy to other systems of vehicle 100. Transmission 120 transmits mechanical power from engine 118 to wheels 121. Transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, transmission 120 may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels 121.
[0194] Sensor system 104 may include several sensors for sensing information about the environment surrounding vehicle 100. For example, sensor system 104 may include a positioning system 122 (which may be a GPS system, a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. Sensor system 104 may also include sensors for the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensing data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the autonomous vehicle 100.
[0195] The positioning system 122 can be used to estimate the geographical location of the vehicle 100. An IMU 124 is used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. A radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, specifically millimeter-wave radar or lidar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of objects. A laser rangefinder 128 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. A camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.
[0196] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 may include various components, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a trajectory control system 142, and an obstacle avoidance system 144.
[0197] The steering system 132 is operable to adjust the forward direction of the vehicle 100. For example, in one embodiment, it may be a steering wheel system. The throttle 134 controls the operating speed of the engine 118 and thus the speed of the vehicle 100. The braking unit 136 controls the deceleration of the vehicle 100. The braking unit 136 may use friction to slow down the wheels 121. In other embodiments, the braking unit 136 may convert the kinetic energy of the wheels 121 into electrical current. The braking unit 136 may also take other forms to slow down the rotational speed of the wheels 121 to control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 140 may use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 may be used to map the environment, track objects, estimate the speed of objects, etc. The route control system 142 is used to determine the driving route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to combine data from the obstacle avoidance system 144, GPS 122, and one or more predetermined maps to determine the driving route and speed for the vehicle 100. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise traverse obstacles in the environment of the vehicle 100, which may specifically be physical obstacles and virtual moving bodies that may collide with the vehicle 100. In one example, the control system 106 may add or alternatively include components other than those shown and described. Alternatively, some of the components shown above may be reduced.
[0198] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users via peripheral device 108. Peripheral device 108 may include wireless communication system 146, on-board computer 148, microphone 150, and / or speaker 152. In some embodiments, peripheral device 108 provides a means for a user of vehicle 100 to interact with user interface 116. For example, on-board computer 148 may provide information to a user of vehicle 100. User interface 116 may also operate on-board computer 148 to receive user input. On-board computer 148 may be operated via a touchscreen. In other cases, peripheral device 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio (e.g., voice commands or other audio input) from a user of vehicle 100. Similarly, speaker 152 may output audio to a user of vehicle 100. Wireless communication system 146 may communicate wirelessly with one or more devices, either directly or via a communication network. For example, the wireless communication system 146 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. The wireless communication system 146 may utilize a wireless local area network (WLAN) for communication. In some embodiments, the wireless communication system 146 may utilize an infrared link, Bluetooth, or ZigBee to communicate directly with the device. Other wireless protocols, such as various vehicle communication systems, may also be used. For example, the wireless communication system 146 may include one or more dedicated short-range communications (DSRC) devices that may enable public and / or private data communication between the vehicle and / or a roadside station.
[0199] Power source 110 can provide power to various components of vehicle 100. In one embodiment, power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more such battery packs can be configured to provide power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 can be implemented together, as is the case in some fully electric vehicles.
[0200] Some or all of the functions of vehicle 100 are controlled by computer system 112. Computer system 112 may include at least one processor 113, which executes instructions 115 stored in a non-transitory computer-readable medium such as memory 114. Computer system 112 may also be multiple computing devices controlling individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, processor 113 may be a dedicated device such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although... Figure 16 The processor, memory, and other components of computer system 112 within the same block are functionally illustrated; however, those skilled in the art will understand that the processor or memory may actually include multiple processors or memories not stored in the same physical housing. For example, memory 114 may be a hard disk drive or other storage media located in a housing different from that of computer system 112. Therefore, references to processor 113 or memory 114 will be understood to include a collection of processors or memories that may or may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as steering and deceleration components, may each have their own processor that performs calculations only related to the component's specific function.
[0201] In all the aspects described herein, processor 113 may be located remotely from vehicle 100 and may communicate wirelessly with vehicle 100. In other aspects, some of the processes described herein are executed on processor 113 located within vehicle 100, while others are executed by remote processor 113, including taking the necessary steps to perform a single operation.
[0202] In some embodiments, memory 114 may contain instructions 115 (e.g., program logic) that can be executed by processor 113 to perform various functions of vehicle 100, including those described above. Memory 114 may also contain additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the mobility system 102, sensor system 104, control system 106, and peripheral devices 108. In addition to instructions 115, memory 114 may also store data such as road maps, route information, vehicle position, direction, speed, and other such vehicle data, as well as other information. This information may be used by vehicle 100 and computer system 112 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is provided to or receives information from a user of vehicle 100. Optionally, user interface 116 may include one or more input / output devices within the set of peripheral devices 108, such as wireless communication system 146, on-board computer 148, microphone 150, and speaker 152.
[0203] Computer system 112 can control the functions of vehicle 100 based on input received from various subsystems (e.g., driving system 102, sensor system 104, and control system 106) and from user interface 116. For example, computer system 112 can utilize input from control system 106 to control steering system 132 to avoid obstacles detected by sensor system 104 and obstacle avoidance system 144. In some embodiments, computer system 112 is operable to provide control over many aspects of vehicle 100 and its subsystems.
[0204] Alternatively, one or more of these components may be installed separately from or associated with vehicle 100. For example, memory 114 may exist partially or completely separately from vehicle 100. The components may be communicatively coupled together in a wired and / or wireless manner.
[0205] Optionally, the components described above are merely examples. In actual applications, components in each of the above modules may be added or removed as needed. Figure 16 This should not be construed as a limitation on the embodiments of this application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine an adjustment to its current speed. These objects can be other vehicles, traffic control equipment, or other types of objects. In some examples, each identified object can be considered independently, and based on the object's individual characteristics, such as its current speed, acceleration, and distance from the vehicle, the speed adjustment to be made by the vehicle can be determined.
[0206] Optionally, the vehicle 100 or a computing device associated with the vehicle 100, such as Figure 16 The computer system 112, computer vision system 140, and memory 114 can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object depends on the behavior of each other, so all identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, the vehicle 100 can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. In this process, other factors can also be considered in determining the speed of the vehicle 100, such as the lateral position of the vehicle 100 in the road, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 100 so that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).
[0207] In this embodiment of the application, the processor 113 in the vehicle 100 is used to execute... Figures 1 to 12 The method executed by the second device in the corresponding embodiment. It should be noted that the specific manner in which the processor 113 executes the aforementioned steps differs from that in this application. Figures 1 to 12 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 12 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0208] This application also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figures 1 to 12 The steps performed by the first device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figures 1 to 12 The steps performed by the second device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figures 1 to 12 The steps performed by the training device in the method described in the illustrated embodiment.
[0209] This application also provides a computer program product, which includes a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figures 1 to 12 The steps performed by the first device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figures 1 to 12The steps performed by the second device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figures 1 to 12 The steps performed by the training device in the method described in the illustrated embodiment.
[0210] This application embodiment also provides a circuit system, the circuit system including a processing circuit, the processing circuit being configured to perform the aforementioned... Figures 1 to 12 The steps performed by the first device in the method described in the illustrated embodiment, or the processing circuit configured to perform the steps as described above. Figures 1 to 12 The steps performed by the second device in the method described in the illustrated embodiment, or the processing circuit configured to perform the steps as described above. Figures 1 to 12 The steps performed by the training device in the method described in the illustrated embodiment.
[0211] The first device, second device, or training device provided in this application embodiment can specifically be a chip. The chip includes a processing unit, such as a processor. Optionally, the chip also includes a communication unit, such as an input / output interface, pins, or circuitry. This processing unit can execute computer execution instructions stored in a storage unit to cause the chip to perform the aforementioned operations. Figures 1 to 12 The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM), etc.
[0212] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0213] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CLUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0214] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0215] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. A method for processing map information, characterized in that, The method includes: Obtain first map information of a first traffic environment, the first map information including the two-dimensional coordinates of at least one first location point, the at least one first location point being a location point among the road elements included in the first traffic environment; Acquire three-dimensional 3D information of the first traffic environment, the 3D information including the height of at least one location region, and objects in the first traffic environment including the at least one location region; Based on the first map information and the 3D information, the height of the at least one first location point is obtained, and the updated first map information includes the two-dimensional coordinates of the at least one first location point and the height of the at least one first location point.
2. The method according to claim 1, characterized in that, The second location point is any one of the at least one first location point, and the step of obtaining the updated first map information based on the first map information and the 3D information includes: The height of the second location point is determined based on the height of the target location region in the at least one location region, wherein the target location region is the location region corresponding to the second location point in the at least one location region.
3. The method according to claim 2, characterized in that, The 3D information also includes the two-dimensional coordinates of each location region. The step of obtaining the updated first map information based on the first map information and the 3D information further includes: Based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region, the target location region corresponding to the second location point is determined from the at least one location region.
4. The method according to claim 3, characterized in that, Determining the target location region corresponding to the second location point from the at least one location region based on the two-dimensional coordinates of the second location point and the two-dimensional coordinates of each location region includes: Based on the two-dimensional coordinates of the second location point, the two-dimensional coordinates of each location region, and the semantic type of each location region, the target location region corresponding to the second location point is determined from the at least one location region.
5. The method according to any one of claims 1 to 4, characterized in that, The first map information is map information from a bird's-eye view BEV, and the method further includes: Based on the first image under the perspective view PV of the first traffic environment, evaluation information corresponding to the updated first map information is determined, and the evaluation information indicates the accuracy of the updated first map information.
6. The method according to claim 5, characterized in that, The at least one first location point includes a location point on the lane boundary line in the first traffic environment. The determination of evaluation information corresponding to the updated first map information based on the first image from the perspective view of the first traffic environment includes: Based on the two-dimensional coordinates and the height of the first position point, the first position point is projected onto the first image under the perspective view to obtain the first position information. The first position information indicates the position of the projection point corresponding to the first position point in the first image under the perspective view. The evaluation information is determined based on the first location information and the first image under the perspective view.
7. The method according to claim 6, characterized in that, Determining the evaluation information based on the first location information and the first image under the perspective view includes: Based on the first image under the perspective view, a semantic recognition operation is performed to obtain second location information, which includes the location information of pixels in the first image under the perspective view whose semantic type is lane boundary line. The evaluation information is determined based on the first location information and the second location information.
8. The method according to claim 6, characterized in that, Determining the evaluation information based on the first location information and the first image under the perspective view includes: Based on the first image under the perspective view, an image segmentation operation is performed to obtain third position information, which indicates the position of the lane boundary line in the image region of the first image under the perspective view. The evaluation information is determined based on the first location information and the third location information.
9. The method according to any one of claims 1 to 4, characterized in that, The first map information is obtained based on the second map information and the third map information. The second map information and the third map information are both map information under the bird's-eye view of the first traffic environment in BEV. The two-dimensional coordinates included in the second map information are in the absolute world coordinate system, and the two-dimensional coordinates included in the third map information are in the relative world coordinate system. The second map information and the third map information are both obtained based on the first image and / or point cloud data of the first traffic environment collected by the vehicle. The origin of the relative world coordinate system is obtained based on the position of the vehicle.
10. The method according to any one of claims 1 to 4, characterized in that, The 3D information is obtained based on a second image from the perspective of the first traffic environment and / or point cloud data of the first traffic environment.
11. A method for acquiring map information, characterized in that, The method includes: Obtain images from the perspective (PV) of the second traffic environment; Based on the image from the perspective view, map information of the second traffic environment is obtained. The map information includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information is generated by a machine learning model. The training data for the machine learning model includes updated first map information of a first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of the at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional 3D information of the first traffic environment. The first map information includes the two-dimensional coordinates of the at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
12. The method according to claim 11, characterized in that, The second location point is any one of the at least one first location point, and the height of the second location point is obtained based on the height of the target location region corresponding to the second location point in the at least one location region.
13. A map information processing device, characterized in that, The device includes: The acquisition module is used to acquire first map information of a first traffic environment. The first map information includes the two-dimensional coordinates of at least one first location point, and the at least one first location point is a location point among the road elements included in the first traffic environment. The acquisition module is further configured to acquire three-dimensional 3D information of the first traffic environment, the 3D information including the height of at least one location region, and objects in the first traffic environment including the at least one location region; The processing module is configured to obtain the height of the at least one first location point based on the first map information and the 3D information, wherein the updated first map information includes the two-dimensional coordinates of the at least one first location point and the height of the at least one first location point.
14. A device for acquiring map information, characterized in that, The device includes: The acquisition module is used to acquire images from the perspective view (PV) of the second traffic environment; The processing module is used to obtain map information of the second traffic environment based on the image under the perspective view. The map information includes the two-dimensional coordinates and height of at least one third location point among the road elements included in the second traffic environment. The map information is generated by a machine learning model. The training data for the machine learning model includes updated first map information of a first traffic environment. The updated first map information of the first traffic environment includes the two-dimensional coordinates and height of at least one first location point among the road elements included in the first traffic environment. The height of the at least one first location point is obtained based on the first map information of the first traffic environment and the three-dimensional 3D information of the first traffic environment. The first map information includes the two-dimensional coordinates of the at least one first location point, and the 3D information of the first traffic environment includes the height of at least one location area of an object in the first traffic environment.
15. A device, characterized in that, The method includes a processor coupled to a memory storing program instructions, which, when executed by the processor, implement the method of any one of claims 1 to 12.
16. A vehicle, characterized in that, The method includes a processor coupled to a memory storing program instructions that, when executed by the processor, implement the method of claim 11 or 12.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 12.
18. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 12.
19. A chip, characterized in that, The chip includes a processor for performing the steps of the method according to any one of claims 1 to 12.