Map generation method and device based on roadside camera, equipment, medium and program

By generating local maps using roadside cameras and stitching them together to form a global map, the problem of high-precision map generation costs has been solved, enabling low-cost and high-timeliness map updates.

CN120953473APending Publication Date: 2025-11-14LINKTECH NAVI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410592492.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing high-precision map generation methods require expensive map acquisition vehicles and a large amount of manual operation, resulting in high generation costs and low efficiency.

Method used

Location information and camera parameters are obtained using roadside cameras to generate local map data. Multiple local maps are then stitched together in the same coordinate system through point cloud registration and coordinate system transformation to generate a global map.

Benefits of technology

It reduces map generation costs, improves map timeliness and coverage, reduces manual operations, and enables real-time updates of high-precision maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953473A_ABST
    Figure CN120953473A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a map generation method and device based on roadside cameras, equipment, a medium and a program, and the method can be applied to the field of maps, and comprises the steps: obtaining the position information of a plurality of roadside cameras and the internal and external parameters of the cameras, generating local map data corresponding to each roadside camera according to the environment image sequence of the road acquired by each roadside camera and the camera internal and external parameters; according to the position information and the coordinate system of the plurality of roadside cameras, converting the local map data corresponding to the plurality of roadside cameras to the same coordinate system; and under the same coordinate system, splicing the converted local map data corresponding to the plurality of roadside cameras to obtain a global map. According to the invention, the road map is generated based on the environment image sequence acquired by the roadside cameras arranged beside the road, the cost of acquiring the environment information by adopting a map acquisition vehicle is saved, and the roadside cameras can acquire the environment images in real time for map generation, so that the timeliness of the generated map can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of map generation technology, and in particular to a map generation method, apparatus, device, medium and program based on roadside cameras. Background Technology

[0002] 3D high-precision maps refer to electronic maps that provide higher accuracy and more data dimensions. The accuracy can reach the centimeter level, and the data dimensions include not only road information but also a large amount of semantic information and driving assistance information. They are widely used in autonomous driving, intelligent transportation (such as traffic flow analysis and route planning), and urban planning (such as urban modeling and landscape modeling).

[0003] Currently, the generation of high-precision maps mainly utilizes map collection vehicles to gather environmental data, including images captured by the vehicle's onboard cameras and point cloud data collected by its onboard LiDAR. Then, map generation equipment performs real-time or offline fusion processing on the point cloud data and image data to generate a global map. Finally, map elements and topological connections in the point cloud are labeled manually or automatically.

[0004] Traditional map generation methods require expensive map acquisition vehicles and consume considerable time and manpower to complete the point cloud acquisition of the entire environment. Matching and stitching the point clouds and labeling map elements also require significant manual labor. Therefore, existing map generation methods are costly. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program for map generation based on roadside cameras, which can provide users with clustering mining that simultaneously meets both user needs and content relevance.

[0006] In a first aspect, embodiments of this application provide a map generation method based on roadside cameras. The method includes: acquiring location information and intrinsic and extrinsic parameters of multiple roadside cameras, wherein the fields of view of adjacent roadside cameras overlap; acquiring local map data corresponding to each roadside camera, wherein the local map data is generated based on the environmental image sequence of the road and the intrinsic and extrinsic parameters of the camera collected by each roadside camera within a first time period; transforming the local map data corresponding to the multiple roadside cameras to the same coordinate system according to the location information and coordinate system of the multiple roadside cameras; and stitching the transformed local map data corresponding to the multiple roadside cameras together in the same coordinate system to obtain a global map.

[0007] In an exemplary embodiment, the above-mentioned acquisition of local map data corresponding to each roadside camera includes: receiving local map data corresponding to each roadside camera sent by each roadside camera, wherein the local map data corresponding to each roadside camera is generated by each roadside camera based on the collected environmental image sequence and camera intrinsic and extrinsic parameters.

[0008] In an exemplary embodiment, the above-mentioned acquisition of local map data corresponding to each roadside camera includes: for each roadside camera, receiving an environmental image sequence sent by the roadside camera, the environmental image sequence including N frames of environmental images, where N is greater than or equal to 2; generating initial local map data corresponding to the roadside camera using a map generation algorithm based on the environmental image sequence of the roadside camera and the camera's intrinsic and extrinsic parameters, the initial local map data including map data of N frames of initial local maps, wherein each frame of environmental image corresponds to the generation of one frame of initial local map; and optimizing the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera.

[0009] In an exemplary embodiment, the map data of the initial local map includes the three-dimensional coordinates of map elements and the attribute information of map elements; correspondingly, the above-mentioned optimization of the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera includes: traversing the map elements in the initial local map data corresponding to the roadside camera, filtering the three-dimensional coordinates of the current map element in the N frames of the initial local map of the roadside camera to obtain the three-dimensional coordinates of the current map element in the local map corresponding to the roadside camera; and determining the attribute information of the current map element in the local map corresponding to the roadside camera based on the attribute information of the current map element in the N frames of the initial local map of the roadside camera.

[0010] In an exemplary embodiment, the filtering process is mean filtering or median filtering.

[0011] In an exemplary embodiment, determining the attribute information of the current map element in the local map corresponding to the roadside camera based on the attribute information of the current map element in the initial N-frame local map of the roadside camera includes: when the attribute information of the current map element in the initial N-frame local map of the roadside camera is the same, the attribute information of the current map element in the initial N-frame local map of the roadside camera is used as the attribute information of the current map element in the local map corresponding to the roadside camera; when the attribute information of the current map element in the initial N-frame local map of the roadside camera is different, a vote is performed on the attribute information of the current map element in the initial N-frame local map of the roadside camera to obtain the unique attribute information of the current map element, and the unique attribute information is determined as the attribute information of the current map element in the local map corresponding to the roadside camera.

[0012] In an exemplary embodiment, after traversing the map elements in the initial local map data corresponding to the roadside cameras, the method further includes: detecting the vehicle trajectory based on the environmental image sequence of the plurality of roadside cameras; and correcting the map elements in the local map corresponding to the plurality of roadside cameras based on the vehicle trajectory.

[0013] In an exemplary embodiment, the above-mentioned detection of vehicle trajectory based on the environmental image sequence of the plurality of roadside cameras includes: for each roadside camera, inputting the environmental image sequence of the roadside camera into a trajectory detection model to obtain the vehicle trajectory in the image coordinate system; converting the vehicle trajectory in the image coordinate system into the vehicle trajectory in a bird's-eye view; and superimposing the vehicle trajectories in the bird's-eye view corresponding to multiple frames of images to obtain the vehicle trajectory.

[0014] Accordingly, the above-mentioned correction of map elements in the local map corresponding to the plurality of roadside cameras based on the vehicle's running trajectory includes: correcting map elements in the local map corresponding to the plurality of roadside cameras based on the vehicle's running trajectory in the bird's-eye view.

[0015] In an exemplary embodiment, the above-mentioned transformation of the local map data corresponding to the plurality of roadside cameras to the same coordinate system based on the location information and coordinate system of the plurality of roadside cameras includes: determining a first coordinate system transformation relationship of each roadside camera relative to a reference coordinate system based on the location information and coordinate system of the plurality of roadside cameras; and transforming the local map data corresponding to each roadside camera to the reference coordinate system based on the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system.

[0016] In an exemplary embodiment, the above-described method of transforming the local map data corresponding to the plurality of roadside cameras to the same coordinate system based on the location information and coordinate system of the plurality of roadside cameras includes: determining a first coordinate system transformation relationship of each roadside camera relative to a reference coordinate system based on the location information and coordinate system of the plurality of roadside cameras; starting from the roadside camera at the starting position, sequentially accessing two adjacent roadside cameras, performing point cloud registration on the two adjacent roadside cameras based on the features of the point sets of the two adjacent roadside cameras within their overlapping fields of view, optimizing the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the registration result, and obtaining a second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system; and transforming the local map data corresponding to each roadside camera to the reference coordinate system based on the second coordinate system transformation relationship of each roadside camera relative to the reference coordinate system.

[0017] In an exemplary embodiment, the point cloud registration employs the Iterative Nearest Neighbor (ICP) algorithm; the features of the point sets of the two adjacent roadside cameras within their overlapping fields of view include the point's location, geometric features, and bird's-eye view features.

[0018] In an exemplary embodiment, the step of performing point cloud registration on the two adjacent roadside cameras based on the features of their respective point sets within the overlapping field of view, and optimizing the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the registration result to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system, includes: determining the initial coordinate system transformation relationship between the two adjacent roadside cameras based on the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system; using the initial coordinate system transformation relationship between the two adjacent roadside cameras as the initial value, performing multiple iterations to obtain the optimized coordinate system transformation relationship of the two adjacent roadside cameras; and correcting the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the optimized coordinate system transformation relationship to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system.

[0019] In an exemplary embodiment, each iteration process includes: performing feature matching on a first point set of the first roadside camera and a second point set of the second roadside camera in the overlapping area based on the distance between points, the similarity of geometric features, and the similarity of bird's-eye view features, to obtain multiple matching point pairs; calculating a new coordinate system transformation relationship between the two adjacent roadside cameras based on the matching point pairs; and transforming the second point set of the second roadside camera in the overlapping area based on the new coordinate transformation relationship to obtain the point set used in the next iteration.

[0020] In an exemplary embodiment, the above-described method of stitching together the converted local map data corresponding to the plurality of roadside cameras in the same coordinate system to obtain a global map includes: stitching together points within non-overlapping areas of the converted local maps corresponding to the plurality of roadside cameras; and fusing the points within overlapping areas of the converted local maps corresponding to two adjacent roadside cameras based on the map data of the points within the overlapping areas of the two adjacent roadside cameras to obtain the map data of the points within the overlapping areas in the global map.

[0021] In an exemplary embodiment, the above-mentioned fusion of map data of points in the overlapping area of ​​the two adjacent roadside cameras to obtain map data of points in the overlapping area of ​​the field of view in the global map includes: traversing the map elements in the overlapping area of ​​the field of view, calculating the mean of the three-dimensional coordinates of the points of each map element in the transformed local map corresponding to the two adjacent roadside cameras, and obtaining the three-dimensional coordinates of the points in each map element in the global map; and determining the attribute information of the points in each map element in the global map based on the attribute information of the points of each map element in the transformed local map corresponding to the two adjacent roadside cameras.

[0022] Secondly, embodiments of this application provide a map generation device based on roadside cameras. The device includes: a first acquisition module, a second acquisition module, a coordinate transformation module, and a stitching module. The first acquisition module is used to acquire the location information and intrinsic / extrinsic parameters of multiple roadside cameras, wherein the fields of view of adjacent roadside cameras overlap. The second acquisition module is used to acquire local map data corresponding to each roadside camera, wherein the local map data is generated based on the environmental image sequence of the road and the intrinsic / extrinsic parameters of the camera collected by each roadside camera within a first time period. The coordinate transformation module is used to transform the local map data corresponding to the multiple roadside cameras to the same coordinate system based on the location information and coordinate system of the multiple roadside cameras. The stitching module is used to stitch the transformed local map data corresponding to the multiple roadside cameras in the same coordinate system to obtain a global map.

[0023] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method as described in the first aspect or various implementations above.

[0024] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the methods described in the first aspect or implementations above.

[0025] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in the first aspect or various implementations above.

[0026] The technical solution provided in this application generates a local map corresponding to the viewpoint of each roadside camera based on the environmental image sequence collected by each roadside camera set up on the roadside. Then, the local maps corresponding to multiple roadside cameras are stitched and optimized to obtain a global map. This saves the cost of using a map collection vehicle to collect environmental information. Moreover, the roadside cameras can collect environmental images in real time for map generation, thereby ensuring the timeliness of the generated map. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of the architecture of a map generation system applicable to the embodiments of this application;

[0029] Figure 2 A flowchart of the map generation method based on roadside cameras provided in Embodiment 1 of this application;

[0030] Figure 3 A flowchart of the map generation method based on roadside cameras provided in Embodiment 2 of this application;

[0031] Figure 4 A flowchart illustrating the process of building a local map;

[0032] Figure 5 A flowchart of the map generation method based on roadside cameras provided in Embodiment 3 of this application;

[0033] Figure 6 A flowchart illustrating the process of building a global map;

[0034] Figure 7 This is a schematic diagram of the structure of the map generation device based on a roadside camera provided in Embodiment 4 of this application;

[0035] Figure 8This is a schematic diagram of the structure of the map generation device provided in the embodiments of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0038] To facilitate understanding of the embodiments of this application, before describing the various embodiments of this application, some concepts involved in all embodiments of this application will be appropriately explained.

[0039] This application provides a map generation method based on roadside cameras, used to generate high-definition or standard-definition maps. The method described in this application can be applied to the fields of mapping or transportation.

[0040] Intelligent Traffic Systems (ITS), also known as Intelligent Transportation Systems, effectively integrate advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. This strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, enhances the environment, and saves energy.

[0041] Figure 1 This is a schematic diagram of the architecture of a map generation system applicable to the embodiments of this application, such as... Figure 1As shown, the map generation system includes: multiple roadside cameras 10 and map generation devices 20 installed on the side of the road, and the roadside cameras 10 and map generation devices 20 can communicate with each other via wired or wireless means.

[0042] Roadside cameras 10 refer to cameras installed on poles (or camera poles) along roads (such as highways, urban roads, etc.) to capture images of roads, vehicles, and pedestrians. Adjacent roadside cameras have overlapping fields of view; this overlap reduces blind spots and ensures the entire road area is covered.

[0043] A roadside camera, also known as a roadside monocular camera, is a device based on image acquisition and processing technology. It typically consists of an image sensor, lens, image processing chip, and data transmission equipment. It is used to monitor traffic conditions, provide real-time traffic data and image information, and support traffic management and safety. For example, roadside cameras use image recognition algorithms to identify targets such as traffic signs, vehicles, and pedestrians on the road. The identified data can be used to calculate information such as traffic flow and vehicle speed, and then transmitted to the traffic management system via data transmission equipment.

[0044] In this embodiment, environmental images collected by roadside cameras are used to generate local maps corresponding to the viewpoints of each roadside camera. Then, the local maps corresponding to multiple roadside cameras are stitched together to obtain a global map.

[0045] In one scenario, each roadside camera 10 sends the environmental images it captures to a map generation device 20. The map generation device 20 determines the local map corresponding to each roadside camera 10 based on the environmental images captured by each roadside camera 10, and generates a global map based on the local maps corresponding to all roadside cameras 10.

[0046] A local map refers to the map within the field of view of a single roadside camera, generated from environmental images captured by that single camera. A global map is a large-scale map stitched together from multiple consecutive local maps.

[0047] In another scenario, each roadside camera 10 has computing power and can generate a corresponding local map based on the environmental images it collects. The generated local map data is then sent to the map generation device 20, which generates a global map based on the local map data corresponding to each roadside camera 10.

[0048] The map generation device 20 can be an electronic device such as a laptop, desktop computer, or server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0049] Roadside cameras are inexpensive and easy to install. New roadside cameras can be installed for map generation, and existing roadside cameras can be reused. Therefore, in this embodiment, road maps are generated based on environmental images collected by multiple roadside cameras. Compared to using a dedicated map collection vehicle to collect environmental information, the cost of road map generation is significantly reduced. Furthermore, roadside cameras can collect environmental images in real time, and the map generation device can update the map instantly based on these images, ensuring the map's timeliness.

[0050] The road map generated by the method in this application embodiment can be applied in the following scenarios: (1) As input for high-precision maps, combined with sensor inputs such as laser point clouds, it provides richer road information during generation and can provide real-time maps, improving the timeliness of high-precision maps and reducing the generation cost of high-precision maps. (2) The road map can also be applied in intelligent transportation systems or roadside sensing products, removing the dependence of intelligent transportation systems or roadside sensing products on high-precision maps, making the system more lightweight. (3) In twin or simulation scenarios, it provides road elements and maintains high-frequency updates. This is just an example; the road map can have many more applications, which will not be listed here.

[0051] The technical solutions of this application will be described in detail below through some embodiments. The embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0052] Figure 2 This is a flowchart of a map generation method based on a roadside camera provided in Embodiment 1 of this application. The method in this embodiment is executed by a map generation device. Figure 2 As shown, the method provided in this embodiment includes the following steps:

[0053] S101. Obtain the location information and internal and external parameters of multiple roadside cameras, wherein the fields of view of adjacent roadside cameras overlap.

[0054] This step is used to collect the location information and internal and external parameters of all roadside cameras within the global map area. The location of the roadside camera can be the latitude and longitude of the pole where the roadside camera is located.

[0055] For example, map generation devices can obtain the latitude and longitude information of the poles where roadside cameras are located in the following ways: selecting points using high-precision maps, selecting points using satellite maps, or marking points using handheld Global Positioning System (GPS) devices.

[0056] High-precision map point selection refers to using the latitude and longitude coordinates of camera pole positions displayed on the high-precision map as the location of roadside cameras. Satellite map point selection refers to using the latitude and longitude coordinates of camera pole positions displayed on the satellite map as the location of roadside cameras; the camera pole position of a particular roadside camera refers to the pole position where that roadside camera is installed. For handheld GPS device point recording, the location of the roadside camera must first be determined, then the GPS device is placed at the location of the roadside camera and its latitude and longitude coordinates are recorded.

[0057] Camera intrinsic and extrinsic parameters include intrinsic and extrinsic parameters. Intrinsic parameters refer to the camera's own parameters such as focal length and optical center. Extrinsic parameters refer to the transformation relationship between the camera coordinate system and the camera pole coordinate system. This transformation relationship includes the rotation and translation matrices of the camera coordinate system relative to the camera pole coordinate system. The camera pole coordinate system can be a local coordinate system established based on the latitude and longitude coordinates of the camera pole position.

[0058] Camera internal and external parameters can be obtained through camera calibration methods, including but not limited to: marker calibration method, semi-automatic calibration method and automatic calibration method.

[0059] S102. Obtain the local map data corresponding to each roadside camera. The local map data is determined based on the environmental image sequence of the road and the camera's intrinsic and extrinsic parameters collected by each roadside camera in the first time period.

[0060] In one implementation, the map generation device receives local map data corresponding to each roadside camera sent by each roadside camera. The local map data corresponding to each roadside camera is generated by each roadside camera based on the collected environmental image sequence and camera intrinsic and extrinsic parameters.

[0061] In another implementation, the map generation device receives the environmental image sequence sent by each roadside camera, and generates local map data corresponding to each roadside camera based on the environmental image sequence and the camera's intrinsic and extrinsic parameters.

[0062] Roadside cameras and map generation devices can use the same method to generate local maps. The initial time interval can be 1 minute, 2 minutes, 5 minutes, or even longer or shorter intervals. Each roadside camera captures an environmental image sequence within its initial time interval, consisting of N frames of environmental images, where N is an integer greater than or equal to 2. The environmental images include images of objects such as vehicles and roads.

[0063] This local map is generated based on N initial local maps from roadside cameras. These N initial local maps are generated from N environmental images and camera intrinsic and extrinsic parameters captured by the roadside cameras. Each environmental image corresponds to one initial local map. For each roadside camera, N initial local maps are generated from the N environmental images. These N initial local maps are then fused to obtain a single local map corresponding to that roadside camera.

[0064] The local map data corresponding to the roadside camera includes the map data of that local map. The map data of the local map includes information on each map element, including but not limited to roads, lane lines, sidewalks, traffic lights, vegetation, etc.

[0065] Map element information includes its identifier, location information, and attribute information. The map element identifier uniquely identifies a map element on the map. The location information of a map element can be understood as the three-dimensional coordinates of each point on the map element. The attribute information of a map element, also known as its semantic information, includes its type and characteristics. Map element types include, but are not limited to, roads, lane lines, sidewalks, traffic lights, and vegetation on the map. Map element characteristics include lane type, road width, road height, speed limits, and driving direction.

[0066] Information about map elements in a local map can be stored in the form of point sets or in a planar form.

[0067] Taking the generation of local maps by a map generation device as an example, each roadside camera can send the collected images to the map generation device in real time. For each roadside camera, the map generation device can generate local map data in the following way: receive the environmental image sequence sent by the roadside camera, which includes N frames of environmental images; based on the environmental image sequence of the roadside camera and the camera's internal and external parameters, use a map generation algorithm to generate initial local map data corresponding to the roadside camera, which includes map data of N frames of initial local maps, wherein each frame of environmental image corresponds to one frame of initial local map; optimize the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera.

[0068] Optionally, the map generation device can employ an end-to-end map generation algorithm to generate local maps. An end-to-end algorithm is a method that directly maps the original input to the desired output, eliminating the need for manually designed intermediate feature extraction or processing steps. In this embodiment, the input to the end-to-end map generation algorithm is each frame of the environmental image and camera intrinsic and extrinsic parameters, and the output is the map data of the initial local map corresponding to each frame of the environmental image. End-to-end algorithms can make the local map generation process more automated and efficient.

[0069] Taking machine learning as an example, some models in machine learning take image features as input, thus requiring a separate feature extraction model to extract features from the image. End-to-end algorithms, however, do not require feature extraction or other processing of the original image; the original image can be directly input into the end-to-end algorithm to obtain the desired output.

[0070] The map generation algorithm includes, but is not limited to: a view frustum-based bird's-eye view method, a transformer-based map detection method, or an image segmentation-based method.

[0071] Optionally, the initial local map data corresponding to each roadside camera also includes a bird's-eye view feature matrix corresponding to each frame of the initial local map. For each roadside camera, N frames of initial local maps correspond to N frames of bird's-eye view feature matrices. Optimizing or fusing these N frames of bird's-eye view feature matrices yields a single frame of bird's-eye view feature matrix for the local map corresponding to the roadside camera. The bird's-eye view feature matrix of the local map can be used for optimization during subsequent generation of the global map.

[0072] Bird's Eye View (BEV) is a perspective from above when viewing an object or scene, much like a bird looking down at the ground from the air.

[0073] In this embodiment, the bird's-eye view feature matrix corresponding to the initial local map of each frame can be understood as an intermediate processing result of the map generation algorithm. The map generation algorithm simplifies the environmental image into a BEV representation (i.e., the bird's-eye view feature matrix) through the BEV algorithm, which facilitates tasks such as object detection and trajectory detection.

[0074] BEV has the following advantages: (1) It simplifies the perspective by converting three-dimensional space into two-dimensional space, which can save a lot of resources in terms of computing and storage. (2) BEV provides a unique visual effect, making the objects and spatial relationships in the scene clearer. (3) It facilitates data processing. Using BEV features to process tasks such as object detection and trajectory detection is simpler than processing directly in 3D data.

[0075] S103. Based on the location information and coordinate system of multiple roadside cameras, transform the local map data corresponding to multiple roadside cameras to the same coordinate system.

[0076] In this embodiment, each roadside camera has an independent coordinate system, and the local maps corresponding to each roadside camera are generated in their own independent coordinate system. Therefore, it is not possible to directly stitch together the local maps corresponding to each roadside camera. Thus, point cloud registration is required to unify the point sets on the local maps of each roadside camera into the same coordinate system, forming a complete data point cloud, which can then be stitched together.

[0077] The coordinate system transformed from each roadside camera is called the reference coordinate system. This reference coordinate system can be the coordinate system of any one of the multiple roadside cameras, an industry-standard coordinate system, or a coordinate system specified according to the usage scenario.

[0078] Typically, when establishing a coordinate system, each roadside camera uses the lane line direction as the Y-axis, the direction perpendicular to the lane line as the X-axis, and the direction perpendicular to the XY plane as the Z-axis. However, because roads are curved, the lane line direction also changes; therefore, the coordinate systems of each roadside camera have different orientations.

[0079] Taking the coordinate system of the first roadside camera, with the road's starting position as the reference coordinate system, as an example, assuming there are 10 roadside cameras in the road, the coordinate system transformation relationships of the remaining 9 roadside cameras relative to the first roadside camera are calculated. This coordinate system transformation relationship includes displacement and rotation information; the rotation information can be understood as the rotation angle between the two coordinate systems. This coordinate system transformation relationship is calculated based on the position information of the roadside cameras and the coordinate system. Based on the coordinate system transformation relationship, the local map data corresponding to multiple roadside cameras can be transformed to the same coordinate system.

[0080] S104. Under the same coordinate system, the transformed local map data corresponding to multiple roadside cameras are stitched together to obtain a global map.

[0081] In one implementation, points within the non-overlapping areas of the converted local maps corresponding to multiple roadside cameras are stitched together; points within the overlapping areas of the converted local maps corresponding to two adjacent roadside cameras are fused based on the map data of the points within the overlapping areas of the two adjacent roadside cameras to obtain the map data of the points within the overlapping areas in the global map.

[0082] For example, points within the overlapping area of ​​two adjacent roadside cameras are fused as follows: traverse the map elements within the overlapping area of ​​view, calculate the average of the 3D coordinates of each point in the transformed local map corresponding to the two adjacent roadside cameras, and obtain the 3D coordinates of the points within each map element in the global map; determine the attribute information of the points within each map element in the global map based on the attribute information of the points in the transformed local map corresponding to the two adjacent roadside cameras.

[0083] Points within an overlapping region may have the same or different 3D coordinates and attribute information in two adjacent local maps. For a point within an overlapping region, the average of its 3D coordinates in the two adjacent local maps is calculated as its 3D coordinates in the global map. If its attribute information is the same in the two adjacent local maps, this same attribute information is used as its attribute information in the global map. If its attribute information is different in the two adjacent local maps, an attribute information is selected from the attribute information in the two adjacent local maps according to a preset rule as its attribute information in the global map.

[0084] The method in this embodiment acquires the location information and intrinsic and extrinsic parameters of multiple roadside cameras, as well as the local map data corresponding to each roadside camera. The local map data is generated based on the environmental image sequence of the road captured by each roadside camera within the first time period and the camera's intrinsic and extrinsic parameters. Based on the location information and coordinate system of the multiple roadside cameras, the local map data corresponding to the multiple roadside cameras are transformed to the same coordinate system. Under the same coordinate system, the transformed local map data corresponding to the multiple roadside cameras are stitched together to obtain a global map. This method generates local maps corresponding to the viewpoints of each roadside camera based on the environmental image sequence captured by each roadside camera set up along the roadside, and then stitches and optimizes the local maps corresponding to multiple roadside cameras to obtain a global map. This saves the cost of using map acquisition vehicles to collect environmental information, and the roadside cameras can collect environmental images in real time for map generation, thus ensuring the timeliness of the generated map.

[0085] Figure 3 This flowchart illustrates a roadside camera-based map generation method provided in Embodiment 2 of this application, serving to provide a detailed description of a specific method for generating a local map in Embodiment 1. Figure 4 A flowchart illustrating the process of building a local map, refer to... Figure 3 and Figure 4 The method provided in this embodiment includes the following steps.

[0086] S201. Obtain the location information and internal and external parameters of multiple roadside cameras, where the fields of view of adjacent roadside cameras overlap.

[0087] Reference Figure 4 As shown, the location information of the roadside camera is obtained by collecting the camera pole coordinates, and the camera's intrinsic and extrinsic parameters are calculated by the camera's intrinsic and extrinsic parameters. The specific acquisition method is described in the relevant description of Embodiment 1, and will not be repeated here.

[0088] S202, Receive an environmental image sequence sent by multiple roadside cameras, the environmental image sequence including N frames of environmental images.

[0089] Each roadside camera collects environmental image sequences and sends the collected environmental image sequences to the map generation device.

[0090] S203. For each roadside camera, based on the environmental image sequence and internal and external parameters of the roadside camera, a map generation algorithm is used to generate the initial local map data corresponding to the roadside camera. The initial local map data includes map data of N frames of initial local maps.

[0091] Each frame of the environment image corresponds to the generation of an initial local map. The map generation device can use an end-to-end map generation algorithm to generate the local map. This map generation algorithm includes, but is not limited to, the following: a view frustum point cloud-based bird's-eye view method, a transformer-based map detection method, or an image segmentation-based method.

[0092] In this embodiment, after constructing N initial local maps for each roadside camera, the initial local map data corresponding to each roadside camera is optimized to obtain the local map data corresponding to each roadside camera.

[0093] Optimizing the initial local map data corresponding to the roadside cameras includes fusing the map data of N frames of the initial local map corresponding to the roadside cameras to obtain the local map corresponding to the roadside cameras, as described in steps S204-S205 below. Optionally, vehicle trajectory detection can also be performed based on the environmental image sequences of multiple roadside cameras, and the map elements in the local maps corresponding to the multiple roadside cameras can be corrected based on the vehicle trajectory, as described in steps S206-S207 below.

[0094] S204. Traverse the map elements in the initial local map data corresponding to the roadside camera, and filter the three-dimensional coordinates of the current map element in the N frames of the initial local map of the roadside camera to obtain the three-dimensional coordinates of the current map element in the local map corresponding to the roadside camera.

[0095] The initial local map data corresponding to the roadside camera includes many map elements. All map elements are traversed sequentially, and N frames of map data for each map element are fused. The initial local map data includes the 3D coordinates and attribute information of the map elements. Fusing the N frames of map data for each map element includes fusing the 3D coordinates and attribute information of points on the map elements in the N frames of the initial local map.

[0096] For the current map element, which comprises many points, this can be called the point set of the current map element. The 3D coordinates of each point in the initial N-frame local map will not be exactly the same; there will be some deviations. Taking a road as an example, when generating the initial local map, one point can be taken every 1 meter. When the coverage area of ​​the roadside camera is 200 meters, the road's point set will include 200 points. The 3D coordinates of these points in the initial N-frame local map will have some deviations.

[0097] In this embodiment, the three-dimensional coordinates of the current map element in the initial local map of the roadside camera in N frames are filtered. After filtering, the N three-dimensional coordinates of the same point on the current map element are transformed into one three-dimensional coordinate.

[0098] The filtering process can be either mean filtering or median filtering. Mean filtering takes the average of N 3D coordinates of the same point in a map element, while median filtering sorts the N 3D coordinates of the same point in a map element and then takes the median value as the final 3D coordinate of that point.

[0099] S205. Based on the attribute information of the current map element in the initial local map of the N frames of the roadside camera, determine the attribute information of the current map element in the local map corresponding to the roadside camera.

[0100] The attribute information of the same point in the current map element may be the same or different in the initial local map of N frames. When fusing the N attribute information of the same point, a voting method can be used. The voting method selects the attribute information with the most votes from the N attribute information of the same point as the final attribute information of that point. For example, if N is 10, for the same point, if the map element type of that point is lane line in 3 initial local maps, and the map element type of that point is pedestrian in the remaining 7 initial local maps, then the voting result is that the map element type of that point is pedestrian.

[0101] Alternatively, if the attribute information of the current map element is the same in the initial N frames of the local map from the roadside camera, the attribute information of the current map element in the initial N frames of the local map from the roadside camera is used as the attribute information of the current map element in the corresponding local map from the roadside camera. If the attribute information of the current map element is different in the initial N frames of the local map from the roadside camera, a vote is taken to obtain the unique attribute information of the current map element, and this unique attribute information is determined as the attribute information of the current map element in the corresponding local map from the roadside camera.

[0102] S206. Detect vehicle trajectory based on environmental image sequences from multiple roadside cameras.

[0103] In the first moment, multiple vehicles may pass through the road. The map generation device can detect the trajectories of all or some of the vehicles. For example, when the road has four lanes, at least one vehicle can be selected in each lane to detect its trajectory, so that the map elements in that lane can be corrected based on the vehicle trajectory in each lane.

[0104] The vehicles are in motion, and they appear sequentially in the field of view of multiple consecutive roadside cameras in chronological order. The vehicle's trajectory can be detected based on the continuous multi-frame images.

[0105] In one exemplary approach, for each roadside camera, the environmental image sequence of the roadside camera is input into the trajectory detection model to obtain the vehicle's trajectory in the image coordinate system; the vehicle's trajectory in the image coordinate system is converted into the vehicle's trajectory in the bird's-eye view; and the vehicle's trajectory in the bird's-eye view corresponding to multiple frames of images is superimposed to obtain the vehicle's running trajectory.

[0106] The trajectory detection model can be a deep learning model, or other neural network models. Of course, other algorithms can also be used to detect the vehicle's trajectory in the image coordinate system.

[0107] A sequence of environmental images from a roadside camera may include multiple vehicles. The trajectory of a vehicle in the image coordinate system detected by this trajectory detection model is the trajectory of the vehicle in a single frame image. The trajectory of the vehicle in the image coordinate system can be converted into the trajectory of the vehicle in the bird's-eye view through homography transformation, solid geometry methods, etc. The vehicle's running trajectory is obtained by superimposing the trajectories of the vehicle in the bird's-eye view corresponding to multiple frames of images.

[0108] S207. Based on the vehicle's trajectory, correct the map elements in the local map corresponding to multiple roadside cameras.

[0109] When making corrections, the coordinate system of the vehicle's trajectory needs to be aligned with the coordinate system of the local map corresponding to the roadside camera. The view of the local map can be taken as a bird's-eye view. By converting the vehicle's trajectory in the image coordinate system into the trajectory in the bird's-eye view, the coordinate system of the vehicle's trajectory is aligned with the coordinate system of the local map. Only after the coordinate system of the vehicle's trajectory is aligned with the coordinate system of the local map can the map elements of the local map be corrected using the vehicle's trajectory.

[0110] Vehicles typically travel along lane lines during operation. In the local map generated by the above method, some lane lines may be missing or inaccurate in position due to factors such as shooting light, occlusion, and algorithms. Based on the vehicle's trajectory, the missing lane lines can be filled in or the position of the lane lines can be corrected.

[0111] S208. Based on the location information and coordinate system of multiple roadside cameras, transform the local map data corresponding to multiple roadside cameras to the same coordinate system.

[0112] S209. Under the same coordinate system, the transformed local map data corresponding to multiple roadside cameras are stitched together to obtain a global map.

[0113] The specific implementation of steps S208-S209 is described in the relevant description of Embodiment 1, and will not be repeated here.

[0114] In this embodiment, for each roadside camera, the map generation device receives an environmental image sequence sent by the roadside camera. This environmental image sequence includes N frames of environmental images. Based on the environmental image sequence of the roadside camera and the camera's intrinsic and extrinsic parameters, a map generation algorithm is used to generate map data for N frames of initial local maps corresponding to the roadside camera. The map data of these N frames of initial local maps is then optimized to obtain the local map data corresponding to the roadside camera. By using multiple frames of images and optimization methods, the generated local map is made more accurate.

[0115] Based on Embodiments 1 and 2, Embodiment 3 of this application provides a map generation method based on roadside cameras, which is used to provide a detailed description of the construction of a global map. Figure 5 This is a flowchart of the map generation method based on roadside cameras provided in Embodiment 3 of this application. Figure 6 A flowchart illustrating the global map construction process, see reference. Figure 5 and Figure 6 The method provided in this embodiment includes the following steps.

[0116] S301. Obtain the location information and internal and external parameters of multiple roadside cameras, where the fields of view of adjacent roadside cameras overlap.

[0117] S302. Obtain the local map data corresponding to each roadside camera. The local map data is generated based on the environmental image sequence of the road and the camera's intrinsic and extrinsic parameters collected by each roadside camera in the first time period.

[0118] The specific implementation methods of steps S301-S302 are described in the relevant descriptions of Embodiment 1 and Embodiment 2, and will not be repeated here.

[0119] S303. Based on the position information and coordinate system of multiple roadside cameras, determine the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system.

[0120] This step is an initial estimation of the coordinate system transformation relationship of each roadside camera. The initial estimation yields the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system. The reference coordinate system can be the coordinate system of any one of the multiple roadside cameras, an industry-standard coordinate system, or a coordinate system specified according to the usage scenario.

[0121] For example, taking the coordinate system of the first roadside camera at the starting position of the road as the reference coordinate system, the transformation relationship between the remaining roadside cameras and the first roadside camera in the first coordinate system is calculated respectively. The coordinate system transformation relationship of each roadside camera relative to the reference coordinate system includes displacement information and rotation information. The rotation information can be understood as the rotation angle between the two cameras.

[0122] The displacement information of the roadside camera's coordinate system and the reference coordinate system can be calculated based on the latitude and longitude information of the camera pole position and the reference coordinate system. It can be understood that the latitude and longitude information of the camera pole position and the reference coordinate system is in degrees, while the displacement information in the coordinate system transformation relationship is in length units (e.g., meters).

[0123] The rotation information between the roadside camera's coordinate system and the reference coordinate system can be represented by the rotation angle between the two coordinate systems, which can be expressed by a rotation matrix. The rotation angle between the two coordinate systems is calculated based on the orientation of the reference coordinate system and the orientation of the roadside camera's coordinate system, where the roadside camera's coordinate system is the coordinate system used by the roadside camera to construct the local map.

[0124] By using the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system, the coordinate system transformation relationship between any two roadside cameras can be obtained, including the coordinate system transformation relationship between any two adjacent roadside cameras.

[0125] In some cases, the first coordinate system transformation relationship estimated through this step is an initial assumption and has a certain degree of error. For example, inaccurate measurement of the camera pole position may lead to a certain error in the final calculated coordinate system transformation relationship.

[0126] In this embodiment, the transformation relationship of each roadside camera relative to the reference coordinate system can be optimized through the following step S304. (Refer to...) Figure 6 As shown, the transformation relationship of the first coordinate system can be optimized based on the geometric and visual features of the local map.

[0127] Optionally, in some other embodiments, the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system may not be optimized. Instead, the local map data corresponding to each roadside camera is transformed to the reference coordinate system based on the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system. Under this reference coordinate system, the transformed local map data corresponding to multiple roadside cameras are stitched together to obtain a global map.

[0128] S304. Starting from the roadside camera at the initial position, visit two adjacent roadside cameras in sequence. Based on the characteristics of the point sets of the two adjacent roadside cameras in the overlapping field of view, perform point cloud registration on the two adjacent roadside cameras. Based on the registration result, optimize the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system.

[0129] There is overlap in the field of view between any two adjacent roadside cameras. Point cloud registration can be performed based on the characteristics of the point sets in the overlapping area of ​​the two adjacent roadside cameras. The purpose of point cloud registration is to determine a suitable coordinate system transformation relationship to transform point clouds or point sets captured from multiple different perspectives into a unified coordinate system.

[0130] The Iterative Closest Point (ICP) algorithm is a commonly used point cloud registration method. Of course, other point cloud registration methods can also be used, and this application does not limit them.

[0131] When performing point cloud registration, the features of the point sets within the overlapping fields of view of two adjacent roadside cameras can include point location, point geometry, and point bird's-eye view features. Optionally, point cloud registration can also be performed using only the point geometry and point bird's-eye view features.

[0132] The bird's-eye view features of the points are obtained from the bird's-eye view feature matrix of the local map data corresponding to the roadside cameras. At the same time as obtaining the local map data corresponding to each roadside camera, the local map data corresponding to each roadside camera can also be obtained.

[0133] In one implementation, when generating the initial local map corresponding to each frame of the environmental image, the map generation algorithm generates a bird's-eye view feature matrix corresponding to each frame of the initial local map. When optimizing the initial local map data corresponding to each roadside camera, the bird's-eye view features are also optimized. For example, the average value of the bird's-eye view feature matrices of N frames of the roadside camera is taken to obtain the bird's-eye view feature matrix of the local map data corresponding to that roadside camera.

[0134] The geometric features of a point can be the curvature of the curve of the map element where the point is located, or the curvature of the curve formed by multiple points near the point, or the vector or distance between the point and multiple nearby points.

[0135] The ICP algorithm works as follows: Given a reference point set P and a data point set Q, where Q is the set of points obtained from a given initial coordinate system transformation (including rotation matrix R and translation vector t), the algorithm finds the nearest point in P for each point in Q, forming a matching point pair. Then, the sum of the Euclidean distances of all matching point pairs is used as the objective function to be solved. Singular value decomposition is used to find R and t to minimize the objective function. A new Q′ is obtained based on R and t, and the point pairs corresponding to P and Q′ are found again, and this process is iterated.

[0136] In this embodiment, the process of point cloud registration using the ICP algorithm includes the following steps:

[0137] (1) Determine the initial coordinate system transformation relationship between two adjacent roadside cameras based on the first coordinate system transformation relationship between the two adjacent roadside cameras and the reference coordinate system. The initial coordinate system transformation relationship includes the initial rotation matrix R and the initial translation vector t.

[0138] (2) Using the initial coordinate system transformation relationship between two adjacent roadside cameras as the initial value, perform multiple iterations to obtain the optimized coordinate system transformation relationship between the two adjacent roadside cameras.

[0139] Each iteration process includes:

[0140] Based on the distance between points, the similarity of geometric features, and the similarity of bird's-eye view features, feature matching is performed on the first point set of the first roadside camera in the overlapping area and the second point set of the second roadside camera in the overlapping area to obtain multiple matching point pairs.

[0141] Calculate the new coordinate system transformation relationship between two adjacent roadside cameras based on the matching point pairs, and transform the second point set of the second roadside camera in the overlapping area according to the new coordinate transformation relationship to obtain the point set used in the next iteration;

[0142] (3) Based on the optimized coordinate system transformation relationship, the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system is corrected to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system.

[0143] Assuming two adjacent roadside cameras are the first roadside camera and the second roadside camera, in each iteration, feature matching is performed on the first point set of the first roadside camera and the second point set of the second roadside camera to find all matching point pairs or a preset number of point pairs. A matching point pair includes two points, one of which belongs to the first point set and the other of which belongs to the second point set.

[0144] When performing feature matching on any two points, the distance between the two points is calculated based on their positions, the similarity of their geometric features is calculated based on their geometric features, and the similarity of their bird's-eye view features is calculated based on their bird's-eye view features. Then, the distance between the two points, the similarity of their geometric features, and the similarity of their bird's-eye view features are added together or weighted and summed to obtain the similarity between the two points. The similarity between the two points is then used to determine whether the two points are a matching pair.

[0145] The iteration can end when the sum of the distances between all matching point pairs in the first and second point sets is less than a preset distance.

[0146] S305. Based on the second coordinate system transformation relationship of each roadside camera relative to the reference coordinate system, transform the local map data corresponding to each roadside camera to the reference coordinate system.

[0147] The second coordinate system transformation relationship is obtained by optimizing the first coordinate system transformation relationship.

[0148] S306. Under this reference coordinate system, the transformed local map data corresponding to multiple roadside cameras are stitched together to obtain a global map.

[0149] In this embodiment, based on the location information and coordinate system of multiple roadside cameras, a first coordinate system transformation relationship of each roadside camera relative to a reference coordinate system is determined. Starting from the roadside camera at the initial position, two adjacent roadside cameras are accessed sequentially. Point cloud registration is performed based on the location features, geometric features, and visual features (i.e., bird's-eye view features) of the point sets of the two adjacent roadside cameras within their overlapping fields of view. The first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system is optimized based on the registration results. Map stitching is then performed based on the optimized coordinate system transformation relationship, resulting in a more accurate global map.

[0150] To facilitate better implementation of the roadside camera-based map generation method of this application, this application also provides a roadside camera-based map generation device. Figure 7 This is a schematic diagram of the map generation device based on a roadside camera provided in Embodiment 4 of this application, as shown below. Figure 7 As shown, the roadside camera-based map generation device 100 may include: a first acquisition module 11, a second acquisition module 12, a coordinate transformation module 13, and a stitching module 14.

[0151] The system comprises the following modules: a first acquisition module 11, used to acquire the location information and intrinsic / extrinsic parameters of multiple roadside cameras, wherein the fields of view of adjacent roadside cameras overlap. A second acquisition module 12, used to acquire local map data corresponding to each roadside camera, wherein the local map data is generated based on the environmental image sequence of the road and the intrinsic / extrinsic parameters of each roadside camera collected within a first time period. A coordinate transformation module 13, used to transform the local map data corresponding to the multiple roadside cameras to the same coordinate system based on the location information and coordinate system of the multiple roadside cameras. A stitching module 14, used to stitch the transformed local map data corresponding to the multiple roadside cameras in the same coordinate system to obtain a global map.

[0152] In some exemplary embodiments, the second acquisition module 12 is specifically used to: receive local map data corresponding to each roadside camera sent by each roadside camera, wherein the local map data corresponding to each roadside camera is generated by each roadside camera based on the collected environmental image sequence and camera intrinsic and extrinsic parameters.

[0153] In some other exemplary embodiments, the second acquisition module 12 is specifically configured to: for each roadside camera, receive an environmental image sequence sent by the roadside camera, the environmental image sequence including N frames of environmental images, where N is greater than or equal to 2; generate initial local map data corresponding to the roadside camera using a map generation algorithm based on the environmental image sequence of the roadside camera and the camera's intrinsic and extrinsic parameters, the initial local map data including map data of N frames of initial local maps, wherein each frame of environmental image corresponds to one frame of initial local map; optimize the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera.

[0154] In an exemplary embodiment, the map data of the initial local map includes the three-dimensional coordinates of map elements and the attribute information of map elements; correspondingly, the second acquisition module 12 optimizes the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera, including: traversing the map elements in the initial local map data corresponding to the roadside camera, filtering the three-dimensional coordinates of the current map element in the N frames of the initial local map of the roadside camera to obtain the three-dimensional coordinates of the current map element in the local map corresponding to the roadside camera; and determining the attribute information of the current map element in the local map corresponding to the roadside camera based on the attribute information of the current map element in the N frames of the initial local map of the roadside camera.

[0155] In an exemplary embodiment, the filtering process is mean filtering or median filtering.

[0156] In an exemplary embodiment, the second acquisition module 12 determines the attribute information of the current map element in the local map corresponding to the roadside camera based on the attribute information of the current map element in the initial N-frame local map of the roadside camera. This includes: when the attribute information of the current map element in the initial N-frame local map of the roadside camera is the same, the attribute information of the current map element in the initial N-frame local map of the roadside camera is used as the attribute information of the current map element in the local map corresponding to the roadside camera; when the attribute information of the current map element in the initial N-frame local map of the roadside camera is different, the attribute information of the current map element in the initial N-frame local map of the roadside camera is voted on to obtain the unique attribute information of the current map element, and the unique attribute information is determined as the attribute information of the current map element in the local map corresponding to the roadside camera.

[0157] In an exemplary embodiment, after traversing the map elements in the initial local map data corresponding to the roadside cameras, the second acquisition module 12 is further configured to: detect the vehicle's trajectory based on the environmental image sequence of the plurality of roadside cameras; and correct the map elements in the local map corresponding to the plurality of roadside cameras based on the vehicle's trajectory.

[0158] In an exemplary embodiment, the second acquisition module 12 detects the vehicle's trajectory based on the environmental image sequence from the plurality of roadside cameras, including: for each roadside camera, inputting the environmental image sequence from the roadside camera into a trajectory detection model to obtain the vehicle's trajectory in the image coordinate system; converting the vehicle's trajectory in the image coordinate system into the vehicle's trajectory from a bird's-eye view; superimposing the vehicle's trajectories from multiple frames of images corresponding to the bird's-eye view to obtain the vehicle's trajectory. Based on the vehicle's trajectory from the bird's-eye view, map elements in the local maps corresponding to the plurality of roadside cameras are corrected.

[0159] In some exemplary embodiments, the coordinate transformation module 13 is specifically used to: determine a first coordinate system transformation relationship of each roadside camera relative to a reference coordinate system based on the location information and coordinate system of the plurality of roadside cameras; and transform the local map data corresponding to each roadside camera to the reference coordinate system based on the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system.

[0160] In some other exemplary embodiments, the coordinate transformation module 13 is specifically used to: determine a first coordinate system transformation relationship of each roadside camera relative to a reference coordinate system based on the position information and coordinate system of the plurality of roadside cameras; starting from the roadside camera at the starting position, sequentially access two adjacent roadside cameras, perform point cloud registration on the two adjacent roadside cameras based on the characteristics of the point sets of the two adjacent roadside cameras in the overlapping field of view, optimize the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the registration result, and obtain a second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system; and transform the local map data corresponding to each roadside camera to the reference coordinate system based on the second coordinate system transformation relationship of each roadside camera relative to the reference coordinate system.

[0161] In an exemplary embodiment, the point cloud registration employs the Iterative Nearest Neighbor (ICP) algorithm; the features of the point sets of the two adjacent roadside cameras within their overlapping fields of view include the point's location, geometric features, and bird's-eye view features.

[0162] In an exemplary embodiment, the coordinate transformation module 13 performs point cloud registration on the two adjacent roadside cameras based on the characteristics of the point sets of each of the two adjacent roadside cameras within the overlapping field of view. Based on the registration result, it optimizes the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system. Specifically, it determines the initial coordinate system transformation relationship between the two adjacent roadside cameras based on the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system; using the initial coordinate system transformation relationship as the initial value, it iterates multiple times to obtain the optimized coordinate system transformation relationship of the two adjacent roadside cameras; and corrects the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the optimized coordinate system transformation relationship to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system.

[0163] In an exemplary embodiment, each iteration process includes: performing feature matching on a first point set of the first roadside camera and a second point set of the second roadside camera in the overlapping area based on the distance between points, the similarity of geometric features, and the similarity of bird's-eye view features, to obtain multiple matching point pairs; calculating a new coordinate system transformation relationship between the two adjacent roadside cameras based on the matching point pairs; and transforming the second point set of the second roadside camera in the overlapping area based on the new coordinate transformation relationship to obtain the point set used in the next iteration.

[0164] In an exemplary embodiment, the stitching module 14 is specifically used to: stitch together points in the non-overlapping areas of the converted local maps corresponding to the plurality of roadside cameras; and fuse the points in the overlapping areas of the converted local maps corresponding to two adjacent roadside cameras based on the map data of the points in the overlapping areas of the two adjacent roadside cameras to obtain the map data of the points in the overlapping areas of the global map.

[0165] In an exemplary embodiment, the stitching module 14 fuses the map data of points in the overlapping area from the two adjacent roadside cameras to obtain map data of points in the overlapping area in the global map. Specifically, it traverses the map elements in the overlapping area, calculates the average of the three-dimensional coordinates of points of each map element in the converted local maps corresponding to the two adjacent roadside cameras, and obtains the three-dimensional coordinates of points in each map element in the global map. Based on the attribute information of points of each map element in the converted local maps corresponding to the two adjacent roadside cameras, it determines the attribute information of points in each map element in the global map.

[0166] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here.

[0167] The apparatus 100 of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly manifested as execution by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0168] This application also provides a map generation device. Figure 8 This is a schematic diagram of the structure of the map generation device provided in the embodiments of this application, such as... Figure 8 As shown, the map generation device 400 includes a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, and a computer program stored in the memory 402 and executable on the processor. The processor 401 is electrically connected to the memory 402. Those skilled in the art will understand that the map generation device structure shown in the figures does not constitute a limitation on the map generation device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0169] The processor 401 is the control center of the map generation device 400. It connects various parts of the map generation device 400 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, it executes various functions of the map generation device 400 and processes data, thereby performing overall processing of the map generation device 400.

[0170] In this embodiment, the processor 401 in the map generation device 400 loads the instructions corresponding to the processes of one or more applications into the memory 402 according to the following steps, and the processor 401 runs the applications stored in the memory 402 to achieve the following functions:

[0171] The system acquires the location information and intrinsic and extrinsic parameters of multiple roadside cameras, where the fields of view of adjacent roadside cameras overlap. It then acquires local map data corresponding to each roadside camera, generated based on the environmental image sequence of the road and the camera's intrinsic and extrinsic parameters collected by each roadside camera within a first time period. Based on the location information and coordinate system of the multiple roadside cameras, the local map data corresponding to the multiple roadside cameras are transformed to the same coordinate system. Finally, the transformed local map data corresponding to the multiple roadside cameras are stitched together in the same coordinate system to obtain a global map.

[0172] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0173] Optional, such as Figure 8 As shown, the map generation device 400 also includes: a display screen 403, a radio frequency circuit 404, an audio circuit 405, an input unit 406, and a power supply 407. The processor 401 is electrically connected to the display screen 403, the radio frequency circuit 404, the audio circuit 405, the input unit 406, and the power supply 407. Those skilled in the art will understand that... Figure 8 The map generation device structure shown does not constitute a limitation on the map generation device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0174] Display screen 403 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. Display screen 403 can be a touch screen, which may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the map generation device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 401. It can also receive and execute commands from the processor 401. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 401 to determine the type of touch event. Subsequently, the processor 401 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and display panel can be integrated into a touch display screen to achieve input and output functions. However, in some embodiments, the touch panel and display panel can be implemented as two independent components to achieve input and output functions. That is, the touch display screen can also be used as part of the input unit 406 to achieve input functions.

[0175] The radio frequency circuit 404 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other map generation devices, and to transmit and receive signals with network devices or other map generation devices.

[0176] Audio circuitry 405 can be used to provide an audio interface between the user and the map generation device via a speaker and a microphone. Audio circuitry 405 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 405, converted back into audio data, and then processed by processor 401 before being transmitted via radio frequency circuitry 404 to, for example, another map generation device, or output to memory 402 for further processing. Audio circuitry 405 may also include an earphone jack to provide communication between peripheral headphones and the map generation device.

[0177] The input unit 406 can be used to receive input numbers, character information or object feature information (such as fingerprints, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0178] Power supply 407 is used to power the various components of map generation device 400. Optionally, power supply 407 can be logically connected to processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 407 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0179] although Figure 8 As not shown in the diagram, the map generation device 400 may also include a camera, sensors, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0180] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.

[0181] This application also provides a computer program product comprising a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the corresponding processes in the above method embodiments; for brevity, these will not be elaborated further here.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0183] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0184] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A map generation method based on roadside cameras, characterized in that, The method includes: The location information and internal and external parameters of multiple roadside cameras are obtained, among which the fields of view of adjacent roadside cameras overlap. Acquire local map data corresponding to each roadside camera, wherein the local map data is generated based on the road environment image sequence and camera intrinsic and extrinsic parameters collected by each roadside camera in the first time period; Based on the location information and coordinate system of the multiple roadside cameras, the local map data corresponding to the multiple roadside cameras are all transformed to the same coordinate system; Under the same coordinate system, the transformed local map data corresponding to the multiple roadside cameras are stitched together to obtain a global map.

2. The method according to claim 1, characterized in that, The acquisition of local map data corresponding to each roadside camera includes: The system receives local map data corresponding to each roadside camera sent by each roadside camera, wherein the local map data corresponding to each roadside camera is generated by each roadside camera based on the collected environmental image sequence and camera intrinsic and extrinsic parameters.

3. The method according to claim 1, characterized in that, The acquisition of local map data corresponding to each roadside camera includes: For each roadside camera, receive an environmental image sequence sent by the roadside camera, the environmental image sequence including N frames of environmental images, where N is greater than or equal to 2; Based on the environmental image sequence and camera intrinsic and extrinsic parameters of the roadside camera, an initial local map data corresponding to the roadside camera is generated using a map generation algorithm. The initial local map data includes map data of N frames of initial local maps, wherein each frame of environmental image corresponds to the generation of one frame of initial local map. The initial local map data corresponding to the roadside camera is optimized to obtain the local map data corresponding to the roadside camera.

4. The method according to claim 3, characterized in that, The map data of the initial local map includes the three-dimensional coordinates of the map elements and the attribute information of the map elements; The optimization of the initial local map data corresponding to the roadside camera to obtain the local map data corresponding to the roadside camera includes: Traverse the map elements in the initial local map data corresponding to the roadside camera, and filter the three-dimensional coordinates of the current map element in the N frames of the initial local map of the roadside camera to obtain the three-dimensional coordinates of the current map element in the local map corresponding to the roadside camera. Based on the attribute information of the current map element in the initial local map of the roadside camera in N frames, the attribute information of the current map element in the local map corresponding to the roadside camera is determined.

5. The method according to claim 4, characterized in that, The filtering process is either mean filtering or median filtering.

6. The method according to claim 4, characterized in that, The step of determining the attribute information of the current map element in the local map corresponding to the roadside camera based on the attribute information of the current map element in the initial local map of the N frames of the roadside camera includes: When the attribute information of the current map element is the same as that of the N-frame initial local map of the roadside camera, the attribute information of the current map element in the N-frame initial local map of the roadside camera shall be used as the attribute information of the current map element in the local map corresponding to the roadside camera. When the attribute information of the current map element is different in the N-frame initial local map of the roadside camera, the attribute information of the current map element in the N-frame initial local map of the roadside camera is voted on to obtain the unique attribute information of the current map element, and the unique attribute information is determined as the attribute information of the current map element in the local map corresponding to the roadside camera.

7. The method according to claim 4, characterized in that, After traversing the map elements in the initial local map data corresponding to the roadside camera, the method further includes: The vehicle trajectory is detected based on the environmental image sequence from the multiple roadside cameras; Based on the vehicle's trajectory, map elements in the local maps corresponding to the multiple roadside cameras are corrected.

8. The method according to claim 7, characterized in that, The step of detecting the vehicle trajectory based on the environmental image sequence from the multiple roadside cameras includes: For each roadside camera, the environmental image sequence of the roadside camera is input into the trajectory detection model to obtain the vehicle's trajectory in the image coordinate system; The trajectory of the vehicle in the image coordinate system is converted into the trajectory of the vehicle from the bird's-eye view. The vehicle's trajectory is obtained by superimposing the trajectories of the vehicle from the bird's-eye view corresponding to multiple frames of images. The step of correcting map elements in the local map corresponding to the plurality of roadside cameras based on the vehicle's trajectory includes: Based on the vehicle's trajectory from a bird's-eye view, map elements in the local maps corresponding to the multiple roadside cameras are corrected.

9. The method according to any one of claims 1-8, characterized in that, The step of transforming the local map data corresponding to the multiple roadside cameras to the same coordinate system based on the location information and coordinate system of the multiple roadside cameras includes: Based on the position information and coordinate system of the plurality of roadside cameras, determine the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system; Based on the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system, the local map data corresponding to each roadside camera is transformed to the reference coordinate system.

10. The method according to any one of claims 1-8, characterized in that, The step of transforming the local map data corresponding to the multiple roadside cameras to the same coordinate system based on the location information and coordinate system of the multiple roadside cameras includes: Based on the position information and coordinate system of the plurality of roadside cameras, determine the first coordinate system transformation relationship of each roadside camera relative to the reference coordinate system; Starting from the roadside camera at the initial position, two adjacent roadside cameras are visited sequentially. Based on the characteristics of the point sets of the two adjacent roadside cameras in the overlapping field of view, point cloud registration is performed on the two adjacent roadside cameras. Based on the registration result, the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system is optimized to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system. Based on the second coordinate system transformation relationship of each roadside camera relative to the reference coordinate system, the local map data corresponding to each roadside camera is transformed to the reference coordinate system.

11. The method according to claim 10, characterized in that, The point cloud registration uses the Iterative Nearest Neighbor (ICP) algorithm. The features of the point sets of the two adjacent roadside cameras within their overlapping field of view include the point's location, geometric features, and bird's-eye view features.

12. The method according to claim 11, characterized in that, The step involves registering point clouds of the two adjacent roadside cameras based on the features of their respective point sets within the overlapping field of view, and optimizing the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system based on the registration result to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system, including: Based on the first coordinate system transformation relationship between the two adjacent roadside cameras relative to the reference coordinate system, the initial coordinate system transformation relationship between the two adjacent roadside cameras is determined; Using the initial coordinate system transformation relationship between the two adjacent roadside cameras as the initial value, multiple iterations are performed to obtain the optimized coordinate system transformation relationship between the two adjacent roadside cameras; Each iteration process includes: Based on the distance between points, the similarity of geometric features, and the similarity of bird's-eye view features, feature matching is performed on the first point set of the first roadside camera and the second point set of the second roadside camera in the overlapping area to obtain multiple matching point pairs. Calculate the new coordinate system transformation relationship between the two adjacent roadside cameras based on the matching point pairs, and transform the second point set of the second roadside camera in the overlapping area according to the new coordinate transformation relationship to obtain the point set used in the next iteration. Based on the optimized coordinate system transformation relationship, the first coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system is corrected to obtain the second coordinate system transformation relationship of the two adjacent roadside cameras relative to the reference coordinate system.

13. The method according to any one of claims 1-8, characterized in that, In the same coordinate system, the transformed local map data corresponding to the multiple roadside cameras are stitched together to obtain a global map, including: The points within the non-overlapping areas of the converted local maps corresponding to the multiple roadside cameras are stitched together. For points within the overlapping area of ​​the transformed local map corresponding to two adjacent roadside cameras, the map data of the points within the overlapping area of ​​the two adjacent roadside cameras are fused to obtain the map data of the points within the overlapping area in the global map.

14. The method according to claim 13, characterized in that, The step of fusing map data of points within the overlapping area from two adjacent roadside cameras to obtain map data of points within the overlapping field of view in the global map includes: Traverse the map elements within the overlapping field of view, calculate the average of the three-dimensional coordinates of each map element's points in the transformed local maps corresponding to the two adjacent roadside cameras, and obtain the three-dimensional coordinates of each map element's points in the global map. Based on the attribute information of the points of each map element in the transformed local map corresponding to the two adjacent roadside cameras, the attribute information of the points within each map element in the global map is determined.

15. A map generation device based on a roadside camera, characterized in that, include: The first acquisition module is used to acquire the location information and internal and external parameters of multiple roadside cameras, wherein the fields of view of adjacent roadside cameras overlap. The second acquisition module is used to acquire local map data corresponding to each roadside camera, wherein the local map data is generated based on the environmental image sequence of the road and the camera's internal and external parameters collected by each roadside camera in the first time period; The coordinate transformation module is used to transform the local map data corresponding to the multiple roadside cameras to the same coordinate system based on the location information and coordinate system of the multiple roadside cameras. The stitching module is used to stitch together the transformed local map data corresponding to the multiple roadside cameras in the same coordinate system to obtain a global map.

16. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 14.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 14.