COLLABORATIVE PERCEPTION SYSTEM FOR CREATING A COOPERATIVE PERCEPTION MAP
The collaborative perception system addresses misaligned data sharing by centralizing perception data analysis to rank roadside objects for improved object detection and reduced computational load on vehicle controllers.
Patent Information
- Application Number
- DE102024100566
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2044-01-10
AI Technical Summary
Existing perception systems face challenges in seamlessly sharing perception data between vehicles due to misaligned data arising from localization errors and temporal asynchrony, especially in limited bandwidth networks, which complicates data registration and object detection.
A collaborative perception system that generates a cooperative perception map by central computers analyzing perception data from multiple vehicles, ranking roadside static objects based on stability using a utility importance function, and transmitting this map to vehicle controllers for enhanced object detection.
The system effectively reduces computational load on vehicle controllers by prioritizing stable roadside objects, improving data alignment and object detection accuracy through centralized data processing and ranking.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
INITIATIONThe present invention relates to a collaborative perception system for creating a cooperative perception map based on perception data collected from a plurality of vehicles.For background information, reference may be made in advance to the publications U.S. Pat. No. 2020 / 0 278 217 A1 and U.S. Pat. No. 2020 / 0 219 386 A1 and to the article "Robust real-time multi-vehicle collaboration on asynchronous sensors" by ZHANG, Qingzhao [et al.] (published in: Proceedings of the 29th annual international conference on mobile computing and networking, Oct. 2-6, 2023, Art. No. 57, pages 858-872 - ISBN 978-1-4503-990-6) and the article "SwarmMap: scaling up real-time collaborative visual SLAM at the edge" by XU, Jingao [et al.] (published in: Proceedings of the 19th USENIX symposium on networked systems design and implementation, April 4-6, 2022, pages 977-993-ISBN 978-1-939133-27-4).An autonomous vehicle performs various tasks such as perception, localization, mapping, path planning, decision making, and motion control. For example, an autonomous vehicle may include perception sensors to acquire perception data about the environment of the vehicle.Sometimes, objects in the environment may not be seen or detected by the autonomous vehicle perception sensors for various reasons. For example, an object may not be in the line of sight of the perception sensors or may be outside of the respective range of the perception sensor. One approach to solving this problem involves the partial sharing of perception data by multiple vehicles in a limited bandwidth wireless network. However, it may be difficult to seamlessly share the perception data collected from multiple vehicles without artifacts arising from misaligned data. This is because the perception data shared by the vehicles may have not inconsiderable deviations due to localization errors and temporal asynchrony. Further, since the network bandwidth is limited, the data registration cannot be used to solve the problem of data misalignment issue (data misalignment issue), because two whole data frames are required for the data registration.While current perception systems achieve their purpose, there is a need for an improved approach to exchanging perception data between vehicles.SUMMARYIn several aspects, a collaborative perception system is disclosed that generates a cooperative perception map based on perception data collected from a plurality of vehicles. The collaborative perception system includes one or more central computers in wireless communication with one or more controllers for each of the plurality of vehicles located in an environment that includes a plurality of road-side static objects. The one or more central computers execute instructions to receive an individual perception map from each of the plurality of vehicles and determine an object set including a plurality of object identifiers, a size set including a plurality of size identifiers, and a duration set including a plurality of duration identifiers based on the individual perception maps from each of the plurality of vehicles. The one or more central computers determine a corresponding stability for each of the static objects on the roadside that are located in the environment based on a utility importance function calculated based on each of the object identifiers that are part of the object set, the largest singular size identifier that is part of the size set, and the largest singular time duration identifier that is part of the time duration set. The one or more central computers rank each surrounding road edge static object based on a respective utility value and generate the cooperative perception map by annotation surrounding map data based on a respective ranking and geographic location of each surrounding road edge static object.In another aspect, the utility importance function is calculated based on a majority selection function that takes into account each of the object identifiers that are part of the object set, a norm function that determines the largest singular time duration identifier that is part of the time duration set, and a norm function that determines the largest singular time duration identifier that is part of the time duration set.In another aspect, the utility importance function is calculated based on:R(o, s, t) represents the utility importance function, o represents one of the object identifiers, s represents the largest singular size identifier that is part of the size set, t represents the largest singular time duration identifier that is part of the time duration set, and a, b, and c represent weights, respectively.In one aspect, the individual perception map includes map data provided with semantic data corresponding to each of the static objects on the roadside at a corresponding location in the environment.In another aspect, the static objects on the roadside each represent an object that has a fixed geographical location in the environment.In another aspect, each object identifier represents a respective road-edge static object located in the environment, each size identifier represents a size of one of the respective road-edge static objects that are part of the object set, and each time duration identifier represents a time duration in which a respective road-edge static object that is part of the object set is observed by a plurality of perception sensors of a respective vehicle.In one aspect, the stability of a respective road-side static object refers to the likelihood of detection by a plurality of perception sensors of each of the plurality of vehicles and the likelihood that the road-side static object changes its geographical location.In another aspect, the one or more central computers execute instructions to transmit the cooperative perception map to the one or more controllers of each of the plurality of vehicles.In another aspect, a collaborative perception system generates a cooperative perception map. The collaborative perception system includes a plurality of vehicles each including a plurality of perception sensors in electronic communication with one or more controllers, the plurality of perception sensors corresponding to each vehicle collecting perception data representing an environment including a plurality of road-edge static objects and one or more central computers in wireless communication with the one or more controllers of each of the plurality of vehicles located in an environment including a plurality of road-edge static objects. The one or more central computers execute instructions to receive an individual perception map from each of the plurality of vehicles and determine an object set including a plurality of object identifiers, a size set including a plurality of size identifiers, and a duration set including a plurality of duration identifiers based on the individual perception maps from each of the plurality of vehicles. The one or more central computers determine a corresponding stability for each of the static objects on the roadside that are located in the environment based on a utility importance function calculated based on each of the object identifiers that are part of the object set, the largest singular size identifier that is part of the size set, and the largest singular time duration identifier that is part of the time duration set. The one or more central computers rank each static object at the roadside in the environment based on a corresponding payload. The one or more central computers generate the cooperative perception map by annotation map data of the environment based on a respective rank and geographic location for each of the static roadside objects located in the environment, and transmitting the cooperative perception map to the one or more controllers of each of the plurality of vehicles.In another aspect, the one or more controllers of an ego vehicle that is part of the plurality of vehicles execute instructions to determine a subset of the road-edge static objects of the cooperative perception map based on the respective rank of each road-edge static object included as part of the cooperative perception map, the subset of the road-edge static objects having a corresponding minimum rank.In another aspect, the one or more controllers of the ego vehicle receive three-dimensional perception data collected from the plurality of perception sensors corresponding to the ego vehicle and determine a set of three-dimensional perception points that are in a predetermined proximity to the subset of objects on the roadside.In one aspect, the one or more controllers of the ego vehicle execute instructions to estimate an ego-based relative pose corresponding to the ego vehicle for each static object on the roadside that is part of the subset of objects on the roadside based on the set of three-dimensional perception points and the subset of objects on the roadside by executing one or more point cloud adaptation algorithms.In another aspect, the one or more controllers of the ego vehicle execute instructions to receive a set of adjacent three-dimensional perception points and adjacent relative poses corresponding to an adjacent vehicle, each adjacent relative pose corresponding to one of the static objects on the roadside that are part of the subset of objects on the roadside, and the set of adjacent three-dimensional perception points are collected by corresponding perception sensors of the adjacent vehicle, execute a transformation function to convert the set of three-dimensional perception points from a local coordinate system of the ego vehicle to a world coordinate system, and execute a transformation function to convert the set of adjacent three-dimensional perception points from a local coordinate system of the adjacent vehicle to the world coordinate system.In another aspect, the one or more controllers of the ego vehicle execute instructions to merge the set of adjacent three-dimensional perception points with the set of three-dimensional perception points based on a matrix stack to generate a merged matrix.In one aspect, the one or more controllers of the ego vehicle execute instructions to analyze the set of three-dimensional perception points and the set of adjacent three-dimensional perception points of the merged matrix based on a three-dimensional object detection model to predict one or more bounding boxes located in an immediate environment of the ego vehicle, each bounding box representative of a corresponding dynamic object in the immediate environment.In another aspect, the utility importance function is calculated based on a majority selection function that takes into account each of the object identifiers that are part of the object set, a norm function that determines the largest singular time duration identifier that is part of the time duration set, and a norm function that determines the largest singular time duration identifier that is part of the time duration set.In another aspect, the utility importance function is calculated based on:R(o, s, t) represents the utility importance function, o represents one of the object identifiers, s represents the largest singular size identifier that is part of the size set, t represents the largest singular time duration identifier that is part of the time duration set, and a, b, and c represent weights, respectively.In one aspect, the individual perception map includes map data provided with semantic data corresponding to each of the static objects on the roadside at a corresponding location in the environment.In another aspect, the static objects on the roadside each represent an object that has a fixed geographical location in the environment.In another aspect, each object identifier represents a respective road-edge static object located in the environment, each size identifier represents a size of one of the respective road-edge static objects that are part of the object set, and each time duration identifier represents a time duration in which a respective road-edge static object that is part of the object set is observed by the plurality of perception sensors of a respective vehicle.In one aspect, a collaborative perception system is disclosed that generates a cooperative perception map. The collaborative perception system includes a plurality of vehicles each including a plurality of perception sensors in electronic communication with one or more controllers, the plurality of perception sensors corresponding to each vehicle collecting perception data representing an environment including a plurality of road-edge static objects, and one or more central computers in wireless communication with the one or more controllers of each of the plurality of vehicles located in an environment including a plurality of road-edge static objects, the road-edge static objects each representing an object having a fixed geographic location within the environment. The one or more central computers execute instructions to receive an individual perception map from each of the plurality of vehicles, the individual perception map including map data provided with semantic data corresponding to each of the static roadside objects at a respective location within the environment. The one or more central computers determine an object set having a plurality of object identifiers, a size set having a plurality of size identifiers, and a duration set having a plurality of duration identifiers based on the individual perception maps of each of the plurality of vehicles. The one or more central computers determine a corresponding stability for each of the static objects on the roadside that are located in the environment based on a utility importance function calculated based on each of the object identifiers that are part of the object set, the largest singular size identifier that is part of the size set, and the largest singular time duration identifier that is part of the time duration set. The one or more central computers rank each static object at the roadside in the environment based on a corresponding payload. The one or more central computers generate the cooperative perception map by annotation map data of the environment based on a respective rank and geographic location for each of the static roadside objects located in the environment, and transmitting the cooperative perception map to the one or more controllers of each of the plurality of vehicles.BRIEF DESCRIPTION OF THE DRAWINGSThe drawings described herein are for illustrative purposes only. FIG. 1 shows a schematic diagram of the disclosed collaborative perception system having one or more central computers in wireless communication with a plurality of vehicles, according to an example embodiment; FIG. 2 is a diagram of one of the vehicles shown in FIG. 1 traveling in an environment with a static object on the roadside, according to an example embodiment; FIG. 3 illustrates the software architecture of the one or more controllers shown in FIG. 1, according to an example embodiment; and FIG. 4 illustrates the software architecture of one or more controllers that are part of one of the vehicles shown in FIG. 1, according to an example embodiment.DETAILED DESCRIPTIONThe following description is merely illustrative in nature.FIG. 1 illustrates an exemplary collaborative perception system 10 for creating a cooperative perception map 12. Collaborative perception system 10 includes one or more central computers 20 located in a back end office 22. The one or more central computers 20 are in wireless communication with a plurality of vehicles 24 located in an environment 26 via a communication network 28. It will be appreciated that the vehicle 24 may be any type of vehicle, such as a sedan, truck, sport utility vehicle (SUV), van, or recreational vehicle. In the embodiment shown in FIG. 1, each vehicle 24 includes one or more controllers 30 in electronic communication with a plurality of perception sensors 32 that collect perception data about the environment 26. The communication network 28 wirelessly connects each of the one or more controllers 30 of each vehicle 24 to the one or more central computers 20 and the one or more controllers 30 corresponding to the one or more remaining vehicles 24. The perception sensors 32 corresponding to each vehicle 24 capture perception data representing the environment 26 that includes a plurality of static objects on the roadside 40. As discussed further below, the one or more central computers 20 generate the cooperative perception map 12 by crowdsourced the perception data collected by multiple vehicles 24 and ranking the road-side static objects 40 based on their corresponding stability.FIG. 2 shows one of the vehicles 24 moving in the environment 26. In the embodiment shown in FIG. 2, the plurality of perception sensors 32 include one or more cameras 42 for capturing image data, an inertial measurement unit (IMU) 44, a global positioning system (GPS) 46, radar 48, and LiDAR 51, but other or additional perception sensors may be used. The plurality of perception sensors 32 collect perception data representative of the environment 26 to which the plurality of road-side static objects 40 also belong. The road-side static objects 40 represent objects that have a fixed geographic location within the environment 26. This means that the static objects 40 are stationary on the road edge and do not change their geographical position within the environment 26. Some examples of road-side static objects 40 include a traffic sign, a building, a tree, and a light or utility line mast.As shown in FIGS. 1 and 2, the plurality of perception sensors 32 collect, for each of the vehicles 24, perception data representative of the environment 26 to which the road-side static objects 40 also belong. The one or more controllers 30 of each vehicle 24 combine the perception data collected by the plurality of perception sensors 32 with map data representing the environment 26 to generate an individual perception map. In one embodiment, the map data may be high resolution map data, but other types of map data may be used. The individual perception map includes the map data representative of the environment 26 provided with semantic data corresponding to each of the road-edge static objects 40 at their respective location, the individual perception map being based on a world coordinate system W.The semantic data indicates an object type, a geographic location, and the perception data corresponding to a respective object at the roadside 40. The object type indicates the size of the road-edge static object 40 and the amount of time that the road-edge static object 40 was detected by the perception sensors 32. The size of the road-edge static object 40 indicates the number of perception data points detected by the perception sensors 32 of the respective vehicle 24, or alternatively, the size of a bounding box corresponding to the road-edge static object 40. For example, if the environment 26 includes a traffic sign at a particular location, the individual perception map is provided with the semantic data representative of the traffic sign at the particular location, the individual perception map being based on the world coordinate system W.As shown in FIGS. 1 and 2, each of the controllers 30 of the plurality of vehicles 24 may geohash the respective individual perception map and transmit the respective individual perception map to the one or more central computers 20 via the communication network 28. The one or more central computers 20 receive an individual perception map from each of the plurality of vehicles 24 via the communication network 28, each individual perception map including map data provided with semantic data corresponding to each of the road-side static objects 40 at their respective locations within the environment 26.FIG. 3 is a block diagram illustrating the software architecture of the one or more central computers 20 shown in FIG. 1. The one or more central computers 20 include an object set block 50, a size set block 52, a duration block 54, an evaluation block 56, and a ranking block 58. the object set block 50 of the one or more central computers 20 determines an object set {o 1, o 2,... o n} based on the individual perception maps of each of the plurality of vehicles 24. the object set includes a plurality of object identifiers o 1, o 2,... o n, each representing a corresponding road-edge static object 40 located in the environment 26 (FIG. 2 ), where n represents the number of road-edge static objects 40 located in the environment 26. The size set block 52 of the one or more central computers 20 determines a size set {s 1, s 2,... s n} based on the individual perception maps of each of the plurality of vehicles 24. the size set includes a plurality of size identifiers s 1, s 2,... s n, each representing the size of one of the respective road-edge static objects 40 that are part of the object set.The duration block 54 of the one or more central computers 20 determines a duration set {t 1, t 2,... t n} based on the individual perception maps of each of the plurality of vehicles 24. The duration set includes a plurality of duration identifiers t 1, t 2,... t n, each representing the duration during which the respective static object 40 is observed at the roadside by the plurality of perception sensors 32 of a respective vehicle 24. The duration block 54 compares the duration identifiers to a threshold duration. The threshold duration represents the minimum duration of time that the perception sensors 32 need to observe the respective static object 40 at the roadside. If the respective object on the roadside 40 is not observed for the minimum duration, the perception data may not be stable.The evaluation block 56 of the one or more central computers 20 receives as input the object set from the object set block 50, the size set from the size set block 52, and the time duration from the time duration block 54, and determines a corresponding stability for each of the road-side static objects 40 located in the environment 26. The stability of a respective roadside static object 40 refers to the likelihood of detection by the perception sensors 32 of each of the plurality of vehicles 24 (FIG. 1 ) and the likelihood that the roadside static object 40 changes its corresponding geographic location.The detection probability is based on the visibility of the road-edge static object 40 by the perception sensors 32 of each of the plurality of vehicles 24. the detection probability is determined based on factors such as the overall size of the road-edge static object 40, the amount of time the road-edge static object 40 was detected by the perception sensors 32, and the number of times the road-edge static object 40 was detected by two or more vehicles 24. For example, a large building for the perception sensors 32 is easier to detect than an object such as a traffic sign. The likelihood that the roadside static object 40 will change the corresponding geographical location is based on the difficulty level in moving the geographical location of the roadside static object 40. For example, a building would be more difficult to move the corresponding geographical location than a traffic sign or bush that is part of the environment 26, as it is much less difficult to move a traffic sign or bush compared to a building.The evaluation block 56 determines the respective stability for each of the road-edge static objects 40 located in the environment 26 based on a utility importance function. The utility importance function determines a utility function value for each of the road-edge static objects 40 represented by an object identifier o 1, o 2,... o n that is part of the object set, with a higher value indicating higher stability. In one embodiment, the utility importance function is calculated based on a majority vote function (majority vote function) that takes into account each of the object identifiers o 1, o 2,... o n that are part of the object set, a norm function that determines the largest singular size identifier s 1, s 2,... s n that is part of the size set, and a norm function that determines the largest singular time duration identifier t 1, t 2,... t n that is part of the time duration set. The size identifiers s 1, s 2,... s n, which are part of the size set, and the time duration identifier t 1, t 2,... t n, which is part of the time duration set, have a one-to-one correspondence with one of the object identifiers o 1, o 2,... o n, which are part of the object set. In one embodiment, the utility importance function in Equation 1 is expressed as: where R(o, s, t) represents the utility importance function, o represents one of the object identifiers, s represents the largest singular size identifier that is part of the size set, t represents the largest singular time duration identifier that is part of the time duration set, and a, b, and c each represent weights having a value between 0 and 1, where the sum of the weights is 1, or a+b+c=1. In one embodiment, the respective values for the weights a, b, and c are determined empirically.The ranking block 58 of the one or more central computers 20 receives the payload function value for each of the road-edge static objects 40 that are part of the object set, and arranges each road-edge static object 40 in an order based on the respective payload function value. It can be seen that a higher utility function value indicates a higher stability of the respective static roadside object 40 (e.g., more data observations or a larger overall physical size of the static roadside object 40). The ranking block 58 of the one or more central computers 20 then provides the map data with a corresponding ranking and geographic location for each of the road-edge static objects 40 located in the environment 26 to generate the cooperative perception map 12, where the cooperative perception map 12 is expressed in the world coordinate system W. The one or more central computers 20 then transmit the cooperative perception map 12 to the one or more controllers 30 of each of the plurality of vehicles 24 via the communication network 28.FIG. 4 is a diagram illustrating the software architecture of the one or more controllers 30 of an ego vehicle A that is one of the plurality of vehicles 24. In the present example, the ego vehicle 24 is labeled A and an adjacent vehicle 24 that is part of the plurality of vehicles 24 is labeled B. The one or more controllers 30 include a subset block 70, a point cloud block 72, a relative pose block 74, a transformation block 76, and a merging block 78 As will be discussed below, the one or more controllers 30 of the ego vehicle A merges the three-dimensional perception data captured by the perception sensors 32 (FIG. 2 ) of the ego vehicle 24 with the three-dimensional perception data captured by the perception sensors 32 of the adjacent vehicle B that is part of the plurality of vehicles 24. The three-dimensional perception data includes, among other things, radar or LiDAR point clouds and image data captured by stereo or depth cameras.While the present example describes merging perception data between the ego vehicle A and the neighbor vehicle B, the ego vehicle A may also merge perception data from more than one vehicle 24. In another embodiment, the ego vehicle A may also merge data from another adjacent vehicle C, for example.The subset block 70 of the one or more controllers 30 of the ego vehicle A receives the cooperative perception map 12 as input and determines a subset of the road-edge static objects 40 that are part of the cooperative perception map 12 based on the respective ranking of each road-edge static object 40 that is included as part of the cooperative perception map 12. The subset of the roadside static objects 40 is referred to as O', and the entire set of the roadside static objects 40 included in the cooperative perception map 12 is referred to as O. The subset O' of the road-edge static objects 40 includes road-edge static objects 40 that are part of the cooperative perception map 12 and have a corresponding minimum ranking. The subset block 70 of the one or more controllers 30 selects the corresponding minimum rank based on a computational capacity of the one or more controllers 30, where a higher computational capacity of the one or more controllers 30 results in a larger subset O'. It should be noted that the subset O' of the roadside static objects 40 is selected to reduce the computational load on the one or more controllers 30 of the ego vehicle A, as numerous roadside static objects 40 may be included as part of the cooperative perception map 12.The subset block 70 of the one or more controllers 30 of the ego vehicle A sends the subset O' of the road-edge static objects 40 to the point cloud block 72 of the one or more controllers 30. the point cloud block 72 also receives the three-dimensional perception data 60 captured by the perception sensors 32 (FIG. 2 ) of the ego vehicle A. The point cloud block 72 then determines a set of three-dimensional perception points L A', which are in a predetermined proximity to the subset of the objects on the roadside O'. The set of three-dimensional perception points L A' is expressed in matrix form and is based on a local coordinate system of the ego vehicle A. In one embodiment, the set of three-dimensional perception points L A' is expressed as, for example, an nx4 matrix. The predetermined proximity is chosen to include three-dimensional perception data points that potentially represent one of the road-edge static objects 40 that are part of the road-edge subset O'. It should be noted that the three-dimensional perception points may not match the static objects at the roadside 40 due to localization errors and deviations from GPS and IMU sensor deviations.The relative pose block 74 of the one or more controllers 30 of the ego vehicle A receives the set of three-dimensional perception points L A' and the subset O' of the roadside objects as input. The relative pose block 74 of the one or more controllers 30 estimates an ego-based relative pose T A corresponding to the ego vehicle A for each static roadside object 40 that is part of the subset O' of roadside objects based on the set of three-dimensional perception points L A' and the subset O' of roadside objects by executing one or more point cloud adaptation algorithms. The point cloud adaptation algorithms determine the ego-based relative pose T A for each road-edge static object 40 that is part of the road-edge subset O' by determining a minimum distance between a transformation of the set of three-dimensional perception points L A' T and a corresponding location of each road-edge static object 40 that is part of the road-edge subset O' of objects. An example of a point cloud adaptation algorithm that may be used is the normal distribution transform (NDT) algorithm. Note that the ego-based relative pose T A is expressed in matrix form. In one embodiment, the ego-based relative pose T A is in the form of a 4x4 matrix.The transformation block 76 of the one or more controllers 30 of the ego vehicle A receives the ego-based relative pose T A and the set of three-dimensional perception points L A' for each road-edge static object 40 that is part of the subset O' of the road-edge objects. The transformation block 76 of the one or more controllers 30 of the ego vehicle A also receives, via the communication network 28, a set of adjacent three-dimensional perception points L B' and a set of adjacent relative poses T B, corresponding to the adjacent vehicle B. The adjacent relative poses T B each correspond to a road-edge static object 40 that is part of the subset O' of the road-edge objects, and the set of the adjacent three-dimensional perception points L B' is detected by the respective perception sensors 32 (FIG. 2 ) of the adjacent vehicle B. Note that the adjacent relative pose T B and the set of the adjacent three-dimensional perception points L B' are also expressed in matrix form. In addition, the adjacent relative pose T B and the set of the adjacent three-dimensional perception points L B' are based on a local coordinate system of the adjacent vehicle B.The transformation block 76 of the one or more controllers 30 performs a transformation function to convert the set of three-dimensional perception points L A' from the local coordinate system of the ego vehicle A to the world coordinate system W. The transformation function is a matrix multiplication function that multiplies the ego-based relative position T A and the set of three-dimensional perception points L A' corresponding to each road-edge static object 40 that is part of the subset O' of the road-edge objects. Similarly, the transformation block 76 of the one or more controllers 30 performs a transformation function to convert the set of the neighboring three-dimensional perception points L B' from the local coordinate system of the neighboring vehicle B to the world coordinate system W.The merging block 78 of the one or more controllers 30 receives the set of three-dimensional perception points L A' and the set of adjacent three-dimensional perception points L B', both expressed in the world coordinate system W, as input from the transformation block 76. The merging block 78 of the one or more controllers 30 then merges the adjacent three-dimensional perception points L B' with the set of three-dimensional perception points L A' based on matrix stacking to create a merged matrix (L A'+ L B'). Specifically, matrix stacking includes concatenating a matrix representing the set of three-dimensional perception points L A' with a matrix representing the adjacent three-dimensional perception points L B', to produce the merged matrix (L A'+ L B').The one or more controllers 30 store a three-dimensional object detection model in memory. The three-dimensional object detection model predicts one or more road-side static objects 40 located in the immediate vicinity of the ego vehicle A based on the merged matrix (L A'+ L B'). An example of a three-dimensional object detection model is the PointP point cloud coder and the PointP network, but other three-dimensional object detection models may be used. The merging block 78 of the one or more controllers 30 analyzes the set of three-dimensional perception points L A' and the adjacent three-dimensional perception points L B, of the merged matrix (L A'+ L B') based on the three-dimensional object detection model to predict one or more bounding boxes located in the immediate environment of the ego vehicle A, each bounding box representative of a corresponding dynamic object (e.g., a vehicle, a pedestrian, or a cyclist) in the immediate environment. The ego vehicle A may then perform one or more perception-related tasks based on the corresponding dynamic objects predicted in the immediate environment.With general reference to the figures, the disclosed collaborative perception system provides various technical effects and advantages.In particular, the cooperative perception map created by the one or more central computers provides an approach to overcoming the problems encountered in attempting to share perception data collected from multiple vehicles, such as the occurrence of artifacts arising from misaligned data. The cooperative perception map primarily utilizes the data collected from multiple vehicles. Because the stability of the roadside static objects is evaluated in the cloud (i.e., in one or more central computers), the vehicle controllers may take into account a portion or subset of the roadside static objects based on their respective rankings, thereby reducing the computational load on the vehicle controllers.The controllers may refer to or be part of an electronic circuit, a combinational logic circuit, a field programmable gate array (FPGA), a (shared, dedicated, or grouped) processor executing code, or a combination of some or all of the above elements, such as in a system-on-a-chip. In addition, the controllers may be microprocessor controlled, such as a computer having at least one processor, a memory (RAM and / or ROM), and associated input and output buses. The processor may operate under the control of an operating system residing in memory. The operating system may manage the computing resources such that the computer program code embodied as one or more computer software applications, e.g., an application residing in memory, may execute instructions from the processor. In an alternative embodiment, the processor may execute the application directly; in this case, the operating system may be omitted.
Claims
A collaborative perception system (10) that generates a cooperative perception map (12) based on perception data collected from a plurality of vehicles (24), the collaborative perception system (10) comprising: one or more central computers (20) in wireless communication with one or more controllers (30) of each of the plurality of vehicles (24) located in an environment that includes a plurality of roadside static objects (40), the one or more central computers (20) executing instructions to: receive an individual perception map from each of the plurality of vehicles (24); determining an object set including a plurality of object identifiers, a size set including a plurality of size identifiers, and a duration set including a plurality of duration identifiers based on the individual perception maps of each of the plurality of vehicles (24); determining a respective stability for each of the road-edge static objects (40) located in the environment based on a utility importance function calculated based on each of the object identifiers that are part of the object set, the largest singular size identifier that is part of the size set, and the largest singular duration identifier that is part of the duration set; ranking each road-edge static object (40) in the environment based on a respective utility function value; and creating the cooperative perception map (12) by annotation map data of the environment based on a respective rank and geographic location of each of the road-side static objects (40) located in the environment.The collaborative perception system (10) of claim 1, wherein the utility importance function is calculated based on a majority tuning function that considers each of the object identifiers that are part of the object set, a norm function that determines the largest singular duration identifier that is part of the duration set, and a norm function that determines the largest singular duration identifier that is part of the duration set.The collaborative perception system (10) of claim 2, wherein the utility importance function is calculated based on: R ( o, s, t ) = a * M a j o r i t y V o t e ( o ) + b * N o r m ( s ) + c * N o r m ( t ) where R ( o, s, t) represents the utility importance function, 0 represents one of the object identifiers, s represents the largest singular size identifier that is part of the size set, t represents the largest singular time duration identifier that is part of the time duration set, and a, b, and c represent weights, respectively.The collaborative perception system (10) of claim 1, wherein the individual perception map includes map data annotated with semantic data corresponding to each of the road-edge static objects (40) at a corresponding location within the environment.The collaborative perception system (10) of claim 1, wherein the static objects (40) on the roadside each represent an object (40) having a fixed geographic location in the environment.The collaborative perception system (10) of claim 1, wherein each object identifier represents a respective roadside static object (40) located in the environment, each size identifier represents a size of one of the respective roadside static objects (40) that are part of the object set, and each time duration identifier represents a time duration in which a respective roadside static object (40) that is part of the object set is observed by a plurality of perception sensors (32) of a respective vehicle (24).The collaborative perception system (10) of claim 1, wherein the stability of a respective roadside static object (40) relates to a likelihood of detection by a plurality of perception sensors (32) of each of the plurality of vehicles (24) and a likelihood that the roadside static object (40) changes its geographical location.The collaborative perception system (10) of claim 1, wherein the one or more central computers (20) execute instructions to: transmit the cooperative perception map (12) to the one or more controllers (30) of each of the plurality of vehicles (24).A collaborative perception system (10) that generates a cooperative perception map (12), the collaborative perception system (10) comprising: a plurality of vehicles (24) each comprising a plurality of perception sensors (32) in electronic communication with one or more controllers (30), wherein the plurality of perception sensors (32) corresponding to the respective vehicles (24) collect perception data representing an environment including a plurality of road-edge static objects (40); and one or more central computers (20) in wireless communication with the one or more controllers (30) of each of the plurality of vehicles (24) located in an environment including a plurality of road-side static objects (40), the one or more central computers (20) executing instructions to: receive an individual perception map from each of the plurality of vehicles (24); determine an object set including a plurality of object identifiers, a size set including a plurality of size identifiers, and a duration set including a plurality of duration identifiers based on the individual perception maps of each of the plurality of vehicles (24); determining a respective stability for each of the roadside static objects (40) located in the environment based on a utility importance function calculated based on each of the object identifiers that are part of the object set, the largest singular size identifier that is part of the size set, and the largest singular time duration identifier that is part of the time duration set; ranking each of the roadside static objects (40) in the environment based on a respective utility function value; creating the cooperative perception map (12) by annotation map data of the environment based on a respective rank and geographic location for each of the roadside static objects (40) located in the environment; and communicating the cooperative perception map (12) to the one or more controllers (30) of each of the plurality of vehicles (24).The collaborative perception system (10) of claim 9, wherein the one or more controllers (30) of an ego vehicle (A) that is part of the plurality of vehicles execute instructions to: determine a subset of the road-edge static objects (40) of the cooperative perception map (12) based on the respective rank of each road-edge static object (40) included as part of the cooperative perception map (12), the subset of the road-edge static objects (40) having a corresponding minimum rank, receive three-dimensional perception data collected from the plurality of perception sensors (32) corresponding to the ego vehicle (A); and determining a set of three-dimensional perception points that are in a predetermined proximity to the subset of objects (40) on the roadside.
Citation Information
Patent Citations
Method, apparatus, and computer program for determining a plurality of traffic situations
US20200219386A1
Method and apparatus for a context-aware crowd-sourced sparse high definition map
US20200278217A1