Collaborative awareness system for creating collaborative awareness maps
Through the collaborative perception system, a central computer is used to process the perceived data of multiple vehicles and create a collaborative perception map, which solves the misalignment problem during the sharing of perceived data between autonomous vehicles, improves the accuracy and consistency of data, and reduces the computing load of the vehicle controller.
Patent Information
- Application Number
- CN202311826920.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2023-12-27
- Publication Date
- 2025-05-30
AI Technical Summary
When sharing perceptual data between autonomous driving vehicles, due to positioning error and time asynchronous, there may be an unnegligible amount of misalignment of perceptual data shared between vehicles, resulting in artifacts. Due to limited network bandwidth, data registration cannot effectively solve the data misalignment problem.
A collaborative perception system is adopted to communicate wirelessly with multiple vehicles through one or more central computers, receive individual perception maps of each vehicle, determine the object set, size set and duration set, and calculate the stability of each static roadside object based on the utility importance function, sort and map data annotation to create a collaborative perception map.
It effectively solves the misalignment problem during perceived data sharing, reduces the generation of artifacts, improves the accuracy and consistency of perceived data, and evaluates the stability of static roadside objects through the cloud, reducing the computing load of the vehicle controller.
Smart Images

Figure CN120071664A_ABST
Abstract
Description
Technical Field The present disclosure relates to a collaborative perception system for creating a collaborative perception map based on perception data collected by multiple vehicles. Background Art Autonomous vehicles perform various tasks, such as but not limited to: perception, localization, mapping, path planning, decision making, and motion control. As an example, an autonomous vehicle may include perception sensors for collecting perception data about the environment around the vehicle. Sometimes, due to various reasons, the perception sensors corresponding to autonomous vehicles may not see or detect objects located in the surrounding environment. For example, an object may not be within the line of sight of the perception sensor or may be outside the corresponding range of the perception sensor. One way to mitigate this problem involves partially sharing perception data between multiple vehicles over a wireless network with limited bandwidth. However, it is challenging to seamlessly share the perception data collected from multiple vehicles without encountering artifacts due to data misalignment. This is because due to positioning errors and time asynchrony, there may be a non-negligible amount of misalignment in the perception data shared between vehicles. Additionally, due to limited network bandwidth, data registration cannot be used to solve the data misalignment problem because data registration requires two complete data frames. Therefore, while current perception systems achieve their intended purposes, there is still a need in the art for an improved method for sharing perception data between vehicles. Summary of the Invention According to several aspects, a collaborative perception system for creating a collaborative perception map based on perception data collected by multiple vehicles is disclosed. The collaborative perception system includes one or more central computers that wirelessly communicate with one or more controllers of each of multiple vehicles located in an environment that includes a plurality of static roadside objects. The one or more central computers execute instructions to receive individual perception maps from each of the multiple vehicles, and determine an object set including a plurality of object identifiers, a dimension set including a plurality of dimension identifiers, and a duration set including duration identifiers based on the individual perception maps from each of the multiple vehicles. The one or more central computers determine a respective stability of each static roadside object located in the environment based on a utility importance function, the calculation of the utility importance function being based on each object identifier that is part of the object set, the maximum singular dimension identifier that is part of the dimension set, and the maximum singular duration identifier that is part of the duration set. The one or more central computers sort each static roadside object in the environment based on their respective utility function values, and annotate the map data of the environment according to the respective sorting and geographical location of each static roadside object in the environment to create a collaborative perception map. On the other hand, the calculation of the utility importance function is based on: a majority vote function considering each object identifier that is part of the set of objects, a norm function determining the largest singular duration identifier that is part of the set of durations, and a norm function determining the largest singular duration identifier that is part of the set of durations. In yet another aspect, the utility importance function is calculated based on the following equation: R(o,s,t) = a * MajorityVote(o) + b * Norm(s) + c * Norm(t) where R(o,s,t) represents the utility importance function, o represents one of the object identifiers, s represents the largest singular size identifier that is part of the set of sizes, t represents the largest singular duration identifier that is part of the set of durations, and a, b, and c all represent weights. In one aspect, the individual perception map includes map data annotated with semantic data corresponding to each static roadside object at a corresponding location in the environment. On the other hand, each static roadside object represents an object having a fixed geographical location within the environment. In yet another aspect, each object identifier represents a corresponding static roadside object located in the environment, each size identifier represents the size of one of the corresponding static roadside objects that is part of the set of objects, and each duration identifier represents the observation duration of the corresponding static roadside object observed by multiple sensing sensors of a corresponding vehicle, where the static roadside object is part of the set of objects. In one aspect, the stability of a corresponding static roadside object refers to the detection probability of the static roadside object by multiple sensing sensors of each vehicle among multiple vehicles, and the possibility of the static roadside object changing its geographical location. On the other hand, one or more central computers execute instructions to transmit the collaborative perception map to one or more controllers of each vehicle among multiple vehicles. In yet another aspect, a cooperative perception system creates a cooperative perception map. The cooperative perception system includes: multiple vehicles, each vehicle including: multiple perception sensors in electronic communication with one or more controllers, where the multiple perception sensors corresponding to each vehicle collect perception data representing an environment including multiple static roadside objects; and one or more central computers in wireless communication with one or more controllers of each of the multiple vehicles located in an environment including multiple static roadside objects. The one or more central computers execute instructions to receive individual perception maps from each of the multiple vehicles, and based on the individual perception maps from each of the multiple vehicles, determine an object set including multiple object identifiers, a dimension set including multiple dimension identifiers, and a duration set including multiple duration identifiers. The one or more central computers determine the respective stability of each static roadside object located in the environment based on a utility importance function, the calculation of which is based on each object identifier that is part of the object set, the maximum singular dimension identifier that is part of the dimension set, and the maximum singular duration identifier that is part of the duration set. The one or more central computers rank each static roadside object in the environment based on the respective utility function values. The one or more central computers annotate map data of the environment to create a cooperative perception map based on the respective rankings and geographical locations of each static roadside object located in the environment, and transmit the cooperative perception map to one or more controllers of each of the multiple vehicles. In another aspect, one or more controllers of a host vehicle that is part of the multiple vehicles execute instructions to determine a subset of the static roadside objects of the cooperative perception map based on the respective rankings of each static roadside object, where each of the static roadside objects is included in the map as part of the cooperative perception map, and the subset of the static roadside objects has the smallest respective rankings. In yet another aspect, one or more controllers of the host vehicle receive three-dimensional perception data collected by the multiple perception sensors corresponding to the host vehicle, and determine a set of three-dimensional perception points, where the distances between the perception points in the set of three-dimensional perception points and the subset of the roadside objects are within a predetermined proximity range. In one aspect, one or more controllers of the host vehicle execute instructions to evaluate a relative pose of the host vehicle corresponding to each static roadside object that is part of the subset of the roadside objects based on the set of three-dimensional perception points and the subset of the roadside objects by performing one or more point cloud matching algorithms. In another aspect, one or more controllers of the host vehicle execute instructions to: receive a set of neighboring three-dimensional perception points and a neighboring relative pose corresponding to a neighboring vehicle, where each neighboring relative pose corresponds to one of the static roadside objects that is part of a subset of roadside objects, and the set of neighboring three-dimensional perception points is collected by a corresponding perception sensor of the neighboring vehicle; execute a transformation function to transform the set of three-dimensional perception points in the local coordinate system of the host vehicle into the world coordinate system; and execute a transformation function to transform the set of neighboring three-dimensional perception points from the local coordinate system of the neighboring vehicle into the world coordinate system. In yet another aspect, one or more controllers of the host vehicle execute instructions to merge the set of neighboring three-dimensional perception points with the set of three-dimensional perception points based on matrix stacking to create a fusion matrix. In one aspect, one or more controllers of the host vehicle execute instructions to analyze the set of three-dimensional perception points and the set of neighboring three-dimensional perception points of the fusion matrix based on a three-dimensional object detection model to predict one or more bounding boxes in the immediate environment around the host vehicle, where each bounding box represents a corresponding dynamic object in the immediate environment. In another aspect, the calculation of the utility importance function is based on: a majority vote function considering each object identifier that is part of a set of objects, a norm function determining the maximum singular duration identifier that is part of a set of durations, and a norm function determining the maximum singular duration identifier that is part of a set of durations. In yet another aspect, the utility importance function is calculated based on the following equation: R(o,s,t) = a * MajorityVote(o) + b * Norm(s) + c * Norm(t) where R(o,s,t) represents the utility importance function, o represents one of the object identifiers, s represents the maximum singular size identifier that is part of a set of sizes, t represents the maximum singular duration identifier that is part of a set of durations, and a, b, and c all represent weights. In one aspect, the individual perception map includes map data annotated with semantic data corresponding to each of the static roadside objects at corresponding locations in the environment. In another aspect, each static roadside object represents an object having a fixed geographical location within the environment. In yet another aspect, each object identifier represents a corresponding static roadside object located in the environment, each size identifier represents the size of one of the corresponding static roadside objects that is part of a set of objects, and each duration identifier represents the observation duration of the corresponding static roadside object that is part of a set of objects as observed by multiple perception sensors of the corresponding vehicle. In one aspect, a collaborative awareness system for creating a collaborative awareness map is disclosed. The collaborative awareness system includes: multiple vehicles, each vehicle including multiple sensing sensors that communicate electronically with one or more controllers, wherein the multiple sensing sensors corresponding to each vehicle collect sensing data representing an environment containing multiple static roadside objects; and one or more central computers that communicate wirelessly with one or more controllers of each of the multiple vehicles located in an environment containing multiple static roadside objects, wherein each static roadside object represents an object having a fixed geographical location within the environment. The one or more central computers execute instructions to receive individual awareness maps from each of the multiple vehicles, wherein the individual awareness maps include map data annotated with semantic data corresponding to each static roadside object at a corresponding location within the environment. The one or more central computers determine an object set including multiple object identifiers, a dimension set including multiple dimension identifiers, and a duration set including multiple duration identifiers based on the individual awareness maps from each of the multiple vehicles. The one or more central computers determine the respective stability of each static roadside object located in the environment based on a utility importance function, the calculation of which is based on: each object identifier that is part of the object set, the largest singular dimension identifier that is part of the dimension set, and the largest singular duration identifier that is part of the duration set. The one or more central computers rank each static roadside object in the environment based on the respective utility function values. The one or more central computers create a collaborative awareness map by annotating the map data of the environment based on the respective rankings and geographical locations of each static roadside object located in the environment and transmit the collaborative awareness map to one or more controllers of each of the multiple vehicles. Further application areas will become apparent from the description provided herein. It should be understood that these descriptions and specific examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Figure 1 A schematic diagram showing the disclosed collaborative awareness system according to an exemplary embodiment, the collaborative awareness system including one or more central computers that communicate wirelessly with multiple vehicles; Figure 2 is according to an exemplary embodiment of Figure 1 A schematic diagram of one of the vehicles shown traveling in an environment including static roadside objects; Figure 3 Shows according to an exemplary embodiment of Figure 1 The software architecture of the one or more central computers shown; Figure 4illustrates, according to an exemplary embodiment, a software architecture of one or more controllers that are part of one of the vehicles shown as Figure 1 a part of one of the vehicles shown. DETAILED DESCRIPTION The following description is merely exemplary in nature and is not intended to limit the present disclosure, application, or uses. Referring to Figure 1 , an exemplary cooperative perception system 10 for creating a cooperative perception map 12 is shown. The cooperative perception system 10 includes one or more central computers 20 located in a back office 22. The one or more central computers 20 wirelessly communicate with a plurality of vehicles 24 located in an environment 26 via a communication network 28. It should be understood that each of the plurality of vehicles 24 can be any type of vehicle, such as but not limited to a sedan, a truck, a sport utility vehicle, a van, or a recreational vehicle. In a non-limiting embodiment as shown in Figure 1 , each vehicle 24 includes one or more controllers 30 that are in electronic communication with a plurality of perception sensors 32 that collect perception data about the environment 26. The communication network 28 wirelessly connects each of the one or more controllers 30 of each vehicle 24 to the one or more central computers 20 and to one or more controllers 30 corresponding to one or more of the remaining vehicles 24. The perception sensors 32 corresponding to each vehicle 24 collect perception data representing the environment 26 that includes a plurality of static roadside objects 40. As explained below, the one or more central computers 20 create the cooperative perception map 12 by crowdsourcing the perception data collected by the plurality of vehicles 24 and ranking the static roadside objects 40 based on their respective stability. Figure 2 is an illustration of one of the plurality of vehicles 24 traveling in the environment 26. In a non-limiting embodiment as shown in Figure 2 , the plurality of perception sensors 32 includes one or more cameras 42 for collecting image data, an inertial measurement unit (IMU) 44, a global positioning system (GPS) 46, a radar 48, and a LiDAR 51. However, it should be understood that different or additional perception sensors can also be used. The plurality of perception sensors 32 collect perception data representing the environment 26 that includes a plurality of static roadside objects 40. The static roadside objects 40 represent objects that have a fixed geographical location within the environment 26. That is, the static roadside objects 40 are stationary and do not change their geographical location within the environment 26. Some examples of the static roadside objects 40 include but are not limited to traffic signs, buildings, trees, and lamp posts or utility poles. Referring to Figure 1 and Figure 2, Multiple sensing sensors 32 of each vehicle 24 among multiple vehicles 24 collect sensing data representing an environment 26 including multiple static roadside objects 40. One or more controllers 30 of each vehicle 24 respectively combine the sensing data collected by the multiple sensing sensors 32 with map data representing the environment 26 to create an individual perception map. In one embodiment, the map data may be high-definition map data. However, it should be understood that other types of map data may also be used. The individual perception map includes map data annotated with semantic data, and the semantic data corresponds to each of the static roadside objects 40 at corresponding positions in the environment 26, where the individual perception map is based on the world coordinate system W. The semantic data indicates the object type, geographical location, and sensing data corresponding to the corresponding roadside object 40. The object type represents the size of the static roadside object 40 and the duration of the static roadside object 40 captured by the sensing sensor 32. The size of the static roadside object 40 represents the number of sensing data points collected by the sensing sensor 32 of the corresponding vehicle 24, or, in an alternative, the size of the bounding box corresponding to the static roadside object 40. For example, if the environment 26 includes a traffic sign at a specified location, the individual perception map is annotated with semantic data representing the traffic sign at the corresponding location, where the individual perception map is based on the world coordinate system W. Continuing to refer to Figure 1 and Figure 2 , Each controller 30 of the multiple vehicles 24 may perform geographical hashing on the corresponding individual perception map and transmit the corresponding individual perception map to one or more central computers 20 via the communication network 28. One or more central computers 20 receive individual perception maps from each of the multiple vehicles 24 via the communication network 28, where each individual perception map includes map data annotated with semantic data, and the semantic data corresponds to each static roadside object 40 at corresponding positions in the environment 26. Figure 3 is a diagram showing Figure 1 the software architecture of the one or more central computers 20 shown. The one or more central computers 20 include an object set block 50, a size set block 52, a duration block 54, a scoring block 56, and a sorting block 58. The object set block 50 of the one or more central computers 20 determines an object set {o 1 , o 2 ,... o n} based on the individual perception maps from each of the multiple vehicles 24. The object set includes multiple object identifiers o 1 , o 2 ,... o n , and each object identifier represents a location in the environment 26 ( Figure 2) corresponding static roadside object 40 therein, where n represents the number of static roadside objects 40 located in the environment 26. The size set block 52 of one or more central computers 20 determines a size set {s 1 , s 2 ,... s n} based on the individual perception maps from each of the multiple vehicles 24. The size set includes a plurality of size identifiers s 1 , s 2 ,... s n , where each size identifier represents the size of one of the respective static roadside objects 40 that is part of the object set. The duration block 54 of one or more central computers 20 determines a duration set {t 1 , t 2 ,... t n} based on the individual perception maps from each of the multiple vehicles 24. The duration set includes a plurality of duration identifiers t 1 , t 2 ,... t n , where each duration identifier represents the duration for which the respective static roadside object 40 is observed by the plurality of sensing sensors 32 of the respective vehicle 24. The duration block 54 compares the duration identifier with a threshold duration. The threshold duration represents the minimum length of time required for the sensing sensors 32 to observe the respective static roadside object 40. If the respective roadside object 40 is not observed for the required minimum length of time, the sensed data may be unstable. The scoring module 56 of one or more central computers 20 receives as input the object set from the object set block 50, the size set from the size set block 52, and the duration set from the duration block 54, and determines the respective stability of each static roadside object 40 located in the environment 26. The stability of the respective static roadside object 40 refers to the detection probability of the detected object by the sensing sensors 32 of each of the multiple vehicles 24( Figure 1 ) and the likelihood that the static roadside object 40 changes its respective geographical location. The detection probability is based on the visibility of the static roadside object 40 by the sensing sensors 32 of each vehicle among multiple vehicles 24. The detection probability is based on, such as but not limited to, the overall physical size of the static roadside object 40, the detection duration during which the static roadside object 40 is detected by the sensing sensor 32, and the frequency at which the static roadside object 40 is detected by two or more vehicles 24. As an example, a large building is more likely to be detected by the sensing sensor 32 compared to an object such as a traffic sign. The likelihood that the static roadside object 40 changes its corresponding geographical location is based on the level of difficulty in moving the geographical location of the static roadside object 40. For example, when compared to a traffic sign or a shrub that is part of the environment 26, a building has a higher level of difficulty in moving its corresponding geographical location because it is much easier to move a traffic sign or a shrub compared to a building. The scoring block 56 determines the corresponding stability of each static roadside object 40 located in the environment 26 based on a utility importance function. The utility importance function determines the utility function value of each static roadside object 40 represented by the object identifier o 1 ,o 2 ,...o n , where a higher value indicates higher stability. In one embodiment, the calculation of the utility importance function is based on a majority vote function that considers each object identifier o 1 ,o 2 ,...o n that is part of the object set, a norm function that determines the maximum singular size identifier s 1 ,s 2 ,...s n that is part of the size set, and a norm function that determines the maximum singular duration identifier t 1 ,t 2 ,...t n that is part of the duration set. It should be understood that the size identifier s 1 ,s 2 ,...s n that is part of the size set and the duration identifier t 1 ,t 2 ,...t n have a one-to-one correspondence with one of the object identifiers o 1 ,o 2 ,...o n that is part of the object set. In a non-limiting embodiment, the utility importance function is expressed in Equation 1 as: R(o, s, t) = a * MajorityVote(o) + b * Norm(s) + c * Norm(t) Equation 1 Where R(o, s, t) represents the utility importance function, o represents one of the object identifiers, s represents the largest singular size identifier as part of a set of sizes, t represents the largest singular duration identifier as part of a set of durations, and a, b, and c all represent weights with a value range from 0 to 1, where the sum of the weights equals 1, or a + b + c = 1. In one embodiment, the respective values of each of the weights a, b, and c are determined empirically. The sorting block 58 of one or more central computers 20 receives the utility function values of each static roadside object 40 that is part of a set of objects, and sorts each static roadside object 40 in order based on the respective utility function values of each static roadside object 40. It should be understood that the higher the utility function value, the higher the stability of the corresponding static roadside object 40 (e.g., the more data observations, or the larger the overall physical size of the static roadside object 40). Then, the sorting block 58 of one or more central computers 20 annotates the map data with the respective sorting and geographical locations of each static roadside object 40 located in the environment 26 to create a cooperative perception map 12, where the cooperative perception map 12 is represented in the world coordinate system W. Then, one or more central computers 20 transmit the cooperative perception map 12 to one or more controllers 30 of each vehicle among the plurality of vehicles 24 via the communication network 28. Figure 4 is a diagram showing the software architecture of one or more controllers 30 of the host vehicle A, which is one of the plurality of vehicles 24. In this example, the host vehicle 24 is represented as A, and the neighboring vehicle 24 as part of the plurality of vehicles 24 is represented as B. One or more controllers 30 include a subset block 70, a point cloud block 72, a relative pose block 74, a transformation block 76, and a fusion block 78. As described below, one or more controllers 30 of the host vehicle A fuse the three-dimensional perception data collected by the perception sensors 32 of the host vehicle 24 ( Figure 2 ) with the three-dimensional perception data collected by the perception sensors 32 of the neighboring vehicle B as part of the plurality of vehicles 24. The three-dimensional perception data includes, but is not limited to, radar or LiDAR point clouds and image data collected by stereo or depth cameras. It should be understood that although this example describes fusing the perception data between the host vehicle A and the neighboring vehicle B, the host vehicle A can also fuse the perception data from more than one vehicle 24. For example, in another embodiment, the host vehicle A can also fuse the data from another neighboring vehicle C. A subset block 70 of one or more controllers 30 of the host vehicle A receives the cooperative perception map 12 as an input and determines a subset of the static roadside objects 40 that are part of the cooperative perception map 12 according to the respective rankings of each static roadside object 40 included as part of the cooperative perception map 12. It should be understood that the subset of the static roadside objects 40 is denoted as O', and the entire set of the static roadside objects 40 included in the cooperative perception map 12 is denoted as O. The subset O' of the static roadside objects 40 includes the static roadside objects 40 with the smallest respective rankings, as part of the cooperative perception map 12. The subset block 70 of one or more controllers 30 selects the smallest respective ranking based on the computing power of the one or more controllers 30, where the higher the computing power of the one or more controllers 30, the larger the subset O'. It should be understood that the subset O' of the static roadside objects 40 is selected to reduce the computational complexity on one or more controllers 30 of the host vehicle A, because there may be a large number of static roadside objects 40 included in its cooperative perception map as part of the cooperative perception map 12.
[0046] A subset block 70 of one or more controllers 30 of the host vehicle A sends the subset O' of the static roadside objects 40 to a point cloud block 72 of the one or more controllers 30. The point cloud block 72 also receives the three-dimensional perception data 60 collected by the perception sensors 32 ( Figure 2 ) of the host vehicle A. Then, the point cloud block 72 determines a three-dimensional perception point set L A' , where the distance between the three-dimensional perception point set L A' and the subset O' of the roadside objects is within a predetermined proximity range. It should be understood that the three-dimensional perception point set L A’ is represented in matrix form and is based on the local coordinate system of the host vehicle A. For example, in one embodiment, the three-dimensional perception point set L A' is represented as an n×4 matrix. The predetermined proximity range is selected to cover the three-dimensional perception data points that may represent one of the static roadside objects 40, which are part of the subset O' of the roadside objects. It should be understood that due to the mismatches caused by the positioning errors and biases caused by GPS and IMU sensor deviations, the three-dimensional perception points may not be aligned with the static roadside objects 40.
[0047] A relative pose block 74 of one or more controllers 30 of the host vehicle A receives the three-dimensional perception point set L A' and the subset O' of the roadside objects as inputs. The relative pose block 74 of the one or more controllers 30 estimates the relative pose T A' of each static roadside object 40 that is part of the subset O' of the roadside objects with respect to the host vehicle A based on the three-dimensional perception point set L A. The point cloud matching algorithm determines the relative pose T of each static roadside object 40 based on its own vehicle by determining the minimum distance between the transformation L of the three-dimensional perception point set and the corresponding positions of each static roadside object 40 that is part of the subset O' of the roadside object. A' The static roadside object is part of the subset O' of the roadside object. An example of a point cloud matching algorithm that can be used is the Normal Distribution Transform (NDT) algorithm. It should be understood that the relative pose T based on its own vehicle is expressed in matrix form. In one embodiment, the relative pose T based on its own vehicle takes the form of a 4x4 matrix. A , which is part of the subset O' of the roadside object. An example of a point cloud matching algorithm that can be used is the Normal Distribution Transform (NDT) algorithm. It should be understood that the relative pose T based on its own vehicle is expressed in matrix form. A In one embodiment, the relative pose T based on its own vehicle takes the form of a 4x4 matrix. A The transformation block 76 of one or more controllers 30 of the vehicle A itself receives the relative pose T based on its own vehicle for each static roadside object 40 that is part of the subset O' of the roadside object. and the three-dimensional perception point set L. The transformation block 76 of one or more controllers 30 of the vehicle A itself also receives the adjacent three-dimensional perception point set L corresponding to the adjacent vehicle B and the adjacent relative pose set T through the communication network 28. A and the three-dimensional perception point set L. A' . The transformation block 76 of one or more controllers 30 of the vehicle A itself also receives the adjacent three-dimensional perception point set L corresponding to the adjacent vehicle B and the adjacent relative pose set T through the communication network 28. B' and the adjacent relative pose set T. B . Each adjacent relative pose T corresponds to a static roadside object 40 that is part of the subset O' of the roadside object, and the adjacent three-dimensional perception point set L is collected by the respective perception sensors 32 of the adjacent vehicle B. B It should be noted that the adjacent relative pose T and the adjacent three-dimensional perception point set L are also represented in matrix form. In addition, the adjacent relative pose T and the adjacent three-dimensional perception point set L are based on the local coordinate system of the adjacent vehicle B. B' The transformation block 76 of one or more controllers 30 performs a transformation function to transform the three-dimensional perception point set L from the local coordinate system of the vehicle A itself to the world coordinate system W. The transformation function is a matrix multiplication function that multiplies the relative pose T based on its own vehicle and the three-dimensional perception point set L corresponding to each static roadside object 40 that is part of the subset O' of the road object with each other. Similarly, the transformation block 76 of one or more controllers 30 performs a transformation function to transform the adjacent three-dimensional perception point set L from the local coordinate system of the adjacent vehicle B to the world coordinate system W. Figure 2 ) B and the adjacent three-dimensional perception point set L. B’ It should be noted that the adjacent relative pose T and the adjacent three-dimensional perception point set L are also represented in matrix form. In addition, the adjacent relative pose T and the adjacent three-dimensional perception point set L are based on the local coordinate system of the adjacent vehicle B. B and the adjacent three-dimensional perception point set L. B' Based on the local coordinate system of the adjacent vehicle B. The transformation block 76 of one or more controllers 30 performs a transformation function to transform the three-dimensional perception point set L from the local coordinate system of the vehicle A itself to the world coordinate system W. The transformation function is a matrix multiplication function that multiplies the relative pose T based on its own vehicle and the three-dimensional perception point set L corresponding to each static roadside object 40 that is part of the subset O' of the road object with each other. Similarly, the transformation block 76 of one or more controllers 30 performs a transformation function to transform the adjacent three-dimensional perception point set L from the local coordinate system of the adjacent vehicle B to the world coordinate system W. A’ from the local coordinate system of the vehicle A itself to the world coordinate system W. The transformation function is a matrix multiplication function that multiplies the relative pose T based on its own vehicle and the three-dimensional perception point set L corresponding to each static roadside object 40 that is part of the subset O' of the road object with each other. Similarly, the transformation block 76 of one or more controllers 30 performs a transformation function to transform the adjacent three-dimensional perception point set L from the local coordinate system of the adjacent vehicle B to the world coordinate system W. A and the three-dimensional perception point set L corresponding to each static roadside object 40 that is part of the subset O' of the road object. A' Multiply each other. Similarly, the transformation block 76 of one or more controllers 30 performs a transformation function to transform the adjacent three-dimensional perception point set L from the local coordinate system of the adjacent vehicle B to the world coordinate system W. B’ from the local coordinate system of the adjacent vehicle B to the world coordinate system W. The fusion block 78 of one or more controllers 30 receives the three-dimensional perception point set L and the adjacent three-dimensional perception point set L. A' and the adjacent three-dimensional perception point set L. B', which are all represented in the world coordinate system W and serve as the input to the transformation block 76. Then, the fusion block 78 of one or more controllers 30 creates a fused matrix (L B’ by combining the neighboring three-dimensional perception points L A’ with this set of three-dimensional perception points L A’ +L B’ ). Specifically, matrix stacking involves concatenating the matrix representing the set of three-dimensional perception points L A’ with the matrix representing the neighboring three-dimensional perception points L B’ to determine the fused matrix (L A’ +L B’ ). One or more controllers 30 store a three-dimensional object detection model in the memory. The three-dimensional object detection model predicts one or more static roadside objects 40 in the immediate environment around the host vehicle A based on the fused matrix (L A’ +L B’ ). An example of a three-dimensional object detection model is the PointPillars point cloud encoder and network. However, it should be understood that other three-dimensional object detection models can also be used. The fusion block 78 of one or more controllers 30 analyzes the set of three-dimensional perception points L A' and the neighboring three-dimensional perception points L B' in the fused matrix (L A' +L B' based on the three-dimensional object detection model to predict one or more bounding boxes in the immediate environment around the host vehicle A, where each bounding box represents a corresponding dynamic object (e.g., a vehicle, a pedestrian, or a cyclist) in the immediate environment. Then, the host vehicle A can perform one or more perception-related tasks based on the corresponding dynamic objects predicted within the immediate environment. Generally referring to the accompanying drawings, the disclosed cooperative perception system provides various technical effects and benefits. Specifically, the cooperative perception map created by one or more central computers provides a way to overcome the challenges faced when attempting to share the perception data collected from multiple vehicles, such as experiencing artifacts caused by data misalignment. In particular, the cooperative perception map utilizes the crowdsourced data collected from multiple vehicles. Additionally, since the stability of the static roadside objects is evaluated in the cloud (i.e., one or more central computers), this enables the vehicle controller to consider a part or subset of the static roadside objects based on their respective rankings, thereby reducing the computational load on the vehicle controller. The controller may refer to an electronic circuit, combinational logic circuit, field programmable gate array (FPGA), processor (shared, dedicated, or group) executing code, or some or all of the foregoing, such as a combination in a system-on-chip. Additionally, the controller may be based on a microprocessor, such as a computer having at least one processor, memory (RAM and / or ROM), and associated input and output buses. The processor may operate under the control of an operating system resident in the memory. The operating system may manage the computer resources so that computer process code embodied as one or more computer software application processes, such as application processes resident in the memory, may have instructions executed by the processor. In an alternative embodiment, the processor may directly execute the application process, in which case the operating system may be omitted. The description of the present disclosure is merely exemplary in nature and variations that do not depart from the gist of the present disclosure are intended to fall within the scope of the present disclosure. These variations should not be regarded as departing from the spirit and scope of the present disclosure.
Claims
1. A collaborative perception system that creates a collaborative perception map based on perception data collected by multiple vehicles, the collaborative perception system comprises: one or more central computers that communicate wirelessly with one or more controllers of each of the multiple vehicles, the multiple vehicles being located in an environment containing multiple static roadside objects, the one or more central computers executing instructions to: receive individual perception maps from each of the multiple vehicles; based on the individual perception maps from each of the multiple vehicles, determine an object set including multiple object identifiers, a size set including multiple size identifiers, and a duration set including multiple duration identifiers; based on a utility importance function, determine the respective stability of each of the static roadside objects located in the environment, the utility importance function being calculated based on: each object identifier that is part of the object set, the largest singular size identifier that is part of the size set, and the largest singular duration identifier that is part of the duration set; sort each of the static roadside objects in the environment based on their respective utility function values; and annotate the map data of the environment according to the respective sorting and geographical locations of each of the static roadside objects in the environment to create the collaborative perception map.
2. The collaborative perception system according to claim 1, wherein, the utility importance function is calculated based on: a majority vote function that considers each object identifier that is part of the object set, a norm function that determines the largest singular duration identifier that is part of the duration set, and a norm function that determines the largest singular duration identifier that is part of the duration set.
3. The collaborative perception system according to claim 2, wherein, the utility importance function is calculated based on the following equation: R(o,s,t) = a * MajorityVote(o) + b * Norm(s) + c * Norm(t) where R(o,s,t) represents the utility importance function, o represents one of the object identifiers, s represents the identifier of the largest singular size that is part of the size set, t represents the identifier of the largest singular duration that is part of the duration set, and a, b, and c all represent weights.
4. The collaborative perception system according to claim 1, wherein, the individual perception map includes map data annotated with semantic data corresponding to each of the static roadside objects at corresponding locations in the environment.
5. The collaborative perception system according to claim 1, wherein, each of the static roadside objects represents an object having a fixed geographical location within the environment.
6. The collaborative perception system according to claim 1, wherein, Each object identifier represents a corresponding static roadside object located in the environment, each size identifier represents the size of a corresponding static roadside object that is part of the object set, and each duration identifier represents the observation duration of a corresponding static roadside object that is part of the object set as observed by multiple sensing sensors of a corresponding vehicle.
7. The cooperative perception system according to claim 1, wherein, the stability of the corresponding static roadside object refers to the detection probability of the corresponding static roadside object detected by multiple sensing sensors of each vehicle among the multiple vehicles, and the possibility of the static roadside object changing its geographical location.
8. The cooperative perception system according to claim 1, wherein, the one or more central computers execute instructions to: transmit the cooperative perception map to the one or more controllers of each vehicle among the multiple vehicles.
9. A cooperative perception system for creating a cooperative perception map, the cooperative perception system comprising: multiple vehicles, each vehicle including multiple sensing sensors that communicate electronically with one or more controllers, wherein the multiple sensing sensors corresponding to each vehicle collect sensing data, and the sensing data represents an environment containing multiple static roadside objects; and one or more central computers that communicate wirelessly with the one or more controllers of each vehicle among the multiple vehicles, the multiple vehicles being located in an environment containing multiple static roadside objects, the one or more central computers execute instructions to: receive individual perception maps from each vehicle among the multiple vehicles; based on the individual perception maps from each vehicle among the multiple vehicles, determine an object set including multiple object identifiers, a size set including multiple size identifiers, and a duration set including multiple duration identifiers; based on a utility importance function, determine the corresponding stability of each static roadside object located in the environment, the utility importance function being calculated based on: each of the object identifiers that is part of the object set, the largest singular size identifier that is part of the size set, and the largest singular duration identifier that is part of the duration set; sort each of the static roadside objects in the environment based on the corresponding utility function values; annotate the map data of the environment based on the corresponding sorting and geographical location of each of the static roadside objects located in the environment to create the cooperative perception map; and transmit the cooperative perception map to the one or more controllers of each vehicle among the multiple vehicles.
10. The cooperative perception system according to claim 9, wherein, one or more controllers of the host vehicle that is part of the multiple vehicles execute instructions to: determine a subset of the static roadside objects of the cooperative perception map based on the corresponding sorting of each static roadside object, wherein each of the static roadside objects is included in the cooperative perception map as part of the cooperative perception map, and the subset of the static roadside objects has the smallest respective sorting; Receive three-dimensional perception data collected by the plurality of perception sensors corresponding to the host vehicle; And Determine a three-dimensional perception point set, wherein the distance between the three-dimensional perception points in the three-dimensional perception point set and a subset of the roadside objects is within a predetermined proximity range.