Method and system for identifying an object
By using a general statistical model to integrate multiple sensor data in autonomous vehicles and fusion with historical data in real time, the problem of difficulty in supporting multiple sensor data at the same time in the prior art is solved, and positioning accuracy and data accuracy are improved.
Patent Information
- Application Number
- CN201910525944.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-06-18
AI Technical Summary
The prior art is difficult to support multiple different types of sensor data in one model at the same time, resulting in low accuracy in positioning the vehicle in the specific location of the road model, and the sensor data at a single moment is unreliable.
A common statistical model is used to integrate multiple types of sensor data into a data model, and the model is updated to identify objects by fusing real-time data with historical data.
Improve the accuracy of positioning of vehicles in road model locations, and consider historical data, significantly improve the accuracy of sensor data.
Smart Images

Figure CN112101392B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to identifying objects, and more particularly, to using real-time sensor data to identify objects. Background Art
[0002] An autonomous vehicle (also known as a driverless car, self-driving car, robotic car) is a vehicle capable of sensing its surrounding environment and navigating without human input. Autonomous vehicles (hereinafter referred to as "ADVs") use various technologies to detect their surrounding environment, such as radar, lidar, GPS, rangefinding, and computer vision. Advanced control systems interpret the sensed information to identify appropriate navigation paths, as well as obstacles and relevant signs.
[0003] More specifically, an ADV collects sensor data from various in-vehicle sensors, such as vision sensors (e.g., cameras), radar-like ranging sensors (such as lidar, millimeter-wave radar, ultrasonic radar), etc.
[0004] As is known in the art, different types of sensors produce different forms or formats of data. When processing sensor data from different sensors, each type of sensor data must be processed separately. Therefore, for each type of sensor data, one or more models for storing that type of sensor data must be established for object identification. Currently, there does not exist a model that can support multiple different types of sensor data simultaneously.
[0005] In addition, a single set of sensor data obtained for a single moment is unstable and unreliable. It is desirable to fuse the sensor data for that moment with multiple sets of sensor data for previous multiple moments to fit the data and thus identify the object.
[0006] Therefore, it is desirable to provide a solution for real-time object identification using a model that can support multiple types of sensor data simultaneously to overcome the above-mentioned deficiencies. Summary of the Invention
[0007] The Summary of the Invention is provided to introduce in a simplified form some concepts that will be further described in the following Detailed Description. The Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.
[0008] According to an embodiment of the present invention, a method for constructing a real-time road model is provided, including: at a first moment, acquiring a first set of sensor data output by a plurality of different types of sensors loaded on a vehicle; feeding the first set of sensor data into a general statistical model to form general statistical model format data for the first moment, the general statistical model format including a plurality of sensor data sets, wherein each of the plurality of sensor data sets is output by one of the different types of sensors; fusing the general statistical model format data for the first moment with historical general statistical model format data to update the general statistical model format data for the first moment; comparing the updated general statistical model format data for the first moment with an implicit model to identify an object, the implicit model including a plurality of sensor data sample sets for describing predefined objects, wherein each of the plurality of sensor data sample sets includes pre-acquired sensor data sample sets for describing the predefined objects output by one of the different types of sensors.
[0009] According to an embodiment of the present invention, a device for constructing a real-time road model is provided, including: a sensor data acquisition module configured to acquire a first set of sensor data output by different types of sensors loaded on a vehicle at a first moment; a data feeding module configured to feed the first set of sensor data into a general statistical model to form general statistical model format data for the first moment, the general statistical model format including a plurality of sensor data sets, wherein each of the plurality of sensor data sets is output by one of the different types of sensors; a data fusion module configured to fuse the general statistical model format data for the first moment with historical general statistical model format data to update the general statistical model format data for the first moment; an object identification module configured to compare the updated general statistical model format data for the first moment with an implicit model to identify an object, the implicit model including a plurality of sensor data sample sets for describing predefined objects, wherein each of the plurality of sensor data sample sets includes pre-acquired sensor data sample sets for describing the predefined objects output by one of the different types of sensors.
[0010] According to an embodiment of the present invention, a vehicle is provided, including: a plurality of different types of sensors; and the above-mentioned device for constructing a real-time road model. The plurality of different types of sensors include vision sensors and radar ranging sensors, the vision sensors include cameras, and the radar includes one or more of lidar, ultrasonic radar, and millimeter-wave radar.
[0011] By adopting the methods, devices, and vehicles disclosed in the present invention, it is possible to support various different types of sensor data in one model, improving the accuracy of positioning the vehicle at a specific location in the road model. In addition, by taking historical data into account, the accuracy of the obtained sensor data can be greatly improved.
[0012] These and other features and advantages will become apparent by reading the following detailed description and referring to the associated drawings. It should be understood that the foregoing general description and the following detailed description are illustrative only and do not limit the various aspects claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to understand in detail the manner in which the above-described features of the present invention are used, the above briefly summarized content may be described more specifically with reference to the various embodiments, some of which are shown in the drawings. However, it should be noted that the drawings only show some typical aspects of the present invention and should not be considered as limiting its scope, as the description may allow other equally effective aspects.
[0014] Figure 1 A schematic diagram showing an autonomous vehicle 100 with different types of sensors traveling on a road according to an embodiment of the present invention is shown.
[0015] Figure 2 A flowchart of a method 200 for identifying an object according to an embodiment of the present invention is shown.
[0016] Figure 3 Schematic diagrams 301 and 302 for general statistical model format data fusion according to an embodiment of the present invention are shown.
[0017] Figure 4 Shown according to Figure 3 An embodiment of a flowchart of a method 400 for fusing general statistical model format data for a first moment with historical general statistical model format data is shown.
[0018] Figure 5 A block diagram of a device 500 for identifying an object according to an embodiment of the present invention is shown.
[0019] Figure 6A block diagram of an exemplary computing device 600 in accordance with one embodiment of the present invention is shown. DETAILED DESCRIPTION
[0020] The present invention will be described in detail below in conjunction with the accompanying drawings, and the features of the present invention will be further manifested in the following detailed description.
[0021] The following detailed description refers to the accompanying drawings that illustrate exemplary embodiments of the present invention. However, the scope of the present invention is not limited to these embodiments, but is defined by the appended claims. Thus, embodiments other than those shown in the drawings, such as modified versions of the illustrated embodiments, are still encompassed by the present invention.
[0022] References in this specification to "one embodiment", "an embodiment", "example embodiment", etc., mean that the described embodiment may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is understood that the particular feature, structure, or characteristic can be implemented in combination with other embodiments within the knowledge of those skilled in the relevant art, whether or not explicitly described.
[0023] For ease of explanation, only embodiments in which the technical solution of the present invention is applied to a "vehicle" or an "autonomous vehicle" (the two terms may be used interchangeably hereinafter) are described in detail herein. However, those skilled in the art can fully understand that the technical solution of the present invention can be applied to any vehicle capable of realizing unmanned autonomous driving, such as an airplane, a helicopter, a train, a subway, a ship, etc. Unless otherwise specified, the term "A or B" used in this specification refers to "A and B" and "A or B", rather than meaning that A and B are exclusive.
[0024] General statistical model
[0025] When an autonomous vehicle is driving on the road, it needs to know the real-time road conditions. For this purpose, various types of sensors are installed on the vehicle to act as the vehicle's "eyes". Currently, widely used sensors include vision sensors (e.g., cameras) and radar ranging sensors (such as lidar, millimeter-wave radar, ultrasonic radar). Each sensor has its own advantages and disadvantages. For example, cameras are low-cost, can identify different objects, and have advantages in aspects such as the measurement accuracy of object height and width, lane line recognition, and pedestrian recognition accuracy. They are indispensable sensors for functions such as lane departure warning and traffic sign recognition. However, their operating range and ranging accuracy are not as good as those of radar, and they are easily affected by factors such as light and weather. Millimeter-wave radar is suitable for all-weather environments, not affected by adverse weather such as light, haze, and sandstorms, and can identify moving and static obstacles. However, in the driving environment, the coexistence of multiple frequency bands has a greater impact on millimeter-wave radar, resulting in a lower data accuracy. Lidar has a wide detection range and high detection accuracy, but its performance is poor in extreme weather such as rain, snow, and fog, and its cost is relatively high. Therefore, it is ideal to use two or more sensors for mutual verification during driving. Currently, autonomous vehicles widely adopt the following integrated solutions to achieve the purpose of safety redundancy: (1) combining a camera and millimeter-wave radar; (2) combining a camera and lidar; (3) combining a camera, millimeter-wave radar, and lidar.
[0026] Reference Figure 1 , which shows a schematic diagram of an autonomous vehicle 100 with different types of sensors driving on the road. Although, for the purpose of illustration, Figure 1 the vehicle 100 in [reference] uses a camera 101, a millimeter-wave radar 102, and a lidar 103 to identify objects, those skilled in the art can fully understand that the technical solution of the present invention can use more or fewer sensors. For example, using only the camera 101 and the millimeter-wave radar 102 or only the camera 101 and the lidar 103 is also within the scope of the present invention.
[0027] When the vehicle has multiple types of sensors, each sensor records its own sensor data and provides it to the vehicle's central processing unit. The formats of sensor data provided by various types or various sensor manufacturers are usually different. Generally speaking, according to the settings of the vehicle manufacturer or the sensor manufacturer, the sensor outputs the original sensor data or the data after preprocessing the original sensor data (e.g., feature extraction, target extraction, etc.). For example, for Figure 1For the same object such as the lane line 104 in [ ], the camera 101 outputs camera data representing the lane line 104, such as raw image data or image features extracted from the raw image data. The millimeter-wave radar 102 outputs millimeter-wave radar data representing the lane line 104, such as raw millimeter-wave radar data or a sequence of polygon data constructed from the raw millimeter-wave radar data. The lidar 103 outputs lidar data representing the lane line 104, such as raw lidar data or three-dimensional point cloud data constructed from the raw lidar data. Of course, the data formats of the sensor outputs listed above are merely illustrative, and those skilled in the art can fully understand that any data format of the sensor outputs is within the scope of the present invention.
[0028] Generally, for each type of sensor data, one or more models are used to record the sensor data of that type. However, this solution will generate multiple sensor data records. For example, at a certain time t, for the same object, multiple data records will be generated that respectively record the camera data output by the camera 101, the millimeter-wave radar data output by the millimeter-wave radar 102, and the lidar data output by the lidar 103. This solution not only consumes too much memory space but also potentially causes a delay in data processing speed.
[0029] The present invention defines a general statistical model that can support multiple types of sensor data simultaneously. The format of the general statistical model is as follows: {t, d s1 , d s2 ... d sn}, which represents a set of sensor data sets output by sensors 1 to n at time t. Among them, s1 represents sensor 1 installed on the vehicle, s2 represents sensor 2 installed on the vehicle, and sn represents sensor n installed on the vehicle (n is any integer greater than 1). The n sensors are different types of sensors used to identify objects, such as cameras, millimeter-wave radars, lidars, etc. Those skilled in the art can fully understand that depending on the specific configuration and requirements of the autonomous vehicle manufacturer, other types of sensors (such as ultrasonic sensors) are also included in the scope of the present invention.
[0030] d s1 represents the data set output by s1 at time t, d s2 represents the data set output by s2 at time t, d sn represents the data set output by sn (n is any integer greater than 1) at time t. Generally, d s1 ……d snEach data set in it has a different data format, such as camera data, millimeter-wave radar data, lidar data, etc. Those skilled in the art can fully understand that according to the sensors adopted, the above data sets can completely include sensor data output by other types of sensors.
[0031] By adopting this general statistical model, data sets output by multiple different types of sensors at a single moment can be integrated into a data model, thereby reducing the storage pressure and being more efficient when performing data processing.
[0032] Hidden model
[0033] In the present invention, multiple implicit models are predefined, and each implicit model includes various different types of sensor data describing an object. Just to give a few examples for illustration, the object can be: road signs, lane lines, buildings, pedestrians, road edges, bridges, utility poles, elevated structures, traffic lights, or traffic signs, etc.
[0034] The format of the implicit model is: {pd s1 , pd s2 ... pd sn , object}, which represents a set of different types of sensor data sample sets used to describe an object. Among them, s1……sn respectively correspond to the sensors s1……sn in the above general statistical model. pd s1 represents a sample set output by s1 obtained in advance to describe the object, pd s2 represents a sample set output by s2 obtained in advance to describe the object, and pd sn represents a sample set output by sn obtained in advance to describe the object. For example, if the object is a building, and s1 is a camera, s2 is a millimeter-wave radar, and s3 is a lidar, then the implicit model is instantiated as: {pd 摄像头 , pd 毫米波雷达 , pd 激光雷达 , building}, from which pd 摄像头 can include a camera data sample set describing the building, pd 毫米波雷达 can include a millimeter-wave radar data sample set describing the building, and pd 激光雷达 can include a lidar data sample set describing the building.
[0035] The sensor data sample set can be pre-obtained by the autonomous vehicle manufacturer, the sensor manufacturer, or the user. For example, as is known to those skilled in the art, vehicle manufacturers and sensor manufacturers can collect a large number of sensor data samples during the training of the target model or the road model, and label the collected sensor data samples through various algorithms or feature extraction methods. Thus, the sample data output by a certain sensor and labeled as describing the same object can be grouped together to form the sensor data sample set of the sensor for that object.
[0036] For further illustration, by way of example only, {pd 摄像头 , pd 毫米波雷达 , pd 激光雷达 , building} can be constructed in the following way. For example, during training, the vehicle manufacturer can drive the vehicle from point A to point B, during which sensors such as cameras, millimeter-wave radars, and lidars obtain a large number of sensor data samples. After processing each of these sensor data samples, each sensor data sample can be labeled to identify a certain object (for example, road signs, lane lines, buildings, pedestrians, road edges, bridges, utility poles, elevated structures, or traffic signs, etc.). Then, multiple camera data samples identifying the same object (for example, a building) are grouped into the pd 摄像头 for that object, multiple millimeter-wave radar data samples identifying the same object (for example, a building) are grouped into the pd 毫米波雷达 for that object, and multiple lidar data samples identifying the same object are grouped into the pd 激光雷达 for that object. Finally, the above three sensor data sample sets are fed into the implicit model to be instantiated as the implicit model for the building, that is, {pd 摄像头 , pd 毫米波雷达 , pd 激光雷达 , building}.
[0037] In addition, the number of samples included in the sensor data sample set can vary according to specific hardware limitations, usage scenarios, and user requirements. According to an embodiment of the present invention, the sensor data sample set can be pre-stored in the storage device of the autonomous vehicle or can be obtained in real time through the network from the server of the autonomous vehicle manufacturer, the server of the sensor manufacturer, or various cloud services.
[0038] According to an embodiment of the present invention, after obtaining the real-time sensor, the object can be identified by comparing the sensor data with the sensor data sample set. The specific method will be described in detail below.
[0039] Implementation method
[0040] Figure 2FIG. 200 is a flowchart of a method for identifying an object according to an embodiment of the present invention. For example, method 200 may be implemented within at least one processor (e.g., Figure 6 processor 604), which may be located in an in-vehicle computer system, a remote server, or a combination thereof. Of course, in various aspects of the present invention, method 200 may also be implemented by any suitable device capable of performing related operations.
[0041] Method 200 begins at step 210. At step 210, at a first moment (in the following description, the "first moment" is understood as "in real time"), a first set of sensor data output by a plurality of different types of sensors loaded on the vehicle at the first moment is acquired. The vehicle may employ two or more different types of sensors. According to an embodiment of the present invention, the plurality of different types of sensors may include a camera and one or more of a millimeter-wave sensor and a lidar sensor. For example, the vehicle may employ a camera and a millimeter-wave radar, a camera and a lidar, or a camera, a millimeter-wave radar, and a lidar. The first set of sensor data includes a plurality of sensor data sets corresponding to the real-time sensor data output by the camera, the millimeter-wave sensor, and / or the lidar sensor, respectively. Of course, as those skilled in the art can understand, other quantities and other types of sensors are also within the scope of the present invention.
[0042] At step 220, the acquired first set of sensor data is fed into a general statistical model to form general statistical model format data for the first moment. That is, by feeding time information (such as a timestamp) and the plurality of sensor data sets included in the first set of sensor data into the general statistical model {t, d s1 , d s2 ... d sn} to instantiate the model. For example, in the case where the vehicle employs a camera, a millimeter-wave radar, and a lidar, the general statistical model format data for the first moment is {t 第一时刻 , d 摄像头 , d 毫米波雷达 , d 激光雷达}. As described above, the data format of d 摄像头 , d 毫米波雷达 , d 激光雷达 may be set by the vehicle manufacturer or the sensor manufacturer. Generally speaking, d 摄像头 , d 毫米波雷达 , d 激光雷达 have different data formats.
[0043] In step 230, the general statistical model format data for the first moment is fused with the historical general statistical model format data to update the general statistical model format data for the first moment. In practice, the real-time sensor data obtained at a single moment alone cannot accurately depict the object. In particular, for continuous objects such as lane lines, the fusion of sensor data at multiple consecutive moments is required to describe the lane line.
[0044] According to one embodiment of the present invention, it is assumed that various types of sensors loaded on the vehicle are synchronized in time and output sensor data at the same time interval. According to different actual requirements, the time interval can be different, such as 0.1 second, 0.2 second, 0.5 second, 1 second, etc. Of course, time intervals of other lengths are also within the scope of the present invention. According to one or more embodiments of the present invention, the historical general statistical model format data can be formed at several consecutive moments before the first moment and stored in the vehicle's memory or cached for quick reading. As should be understood, the historical general statistical model format data has the same data format as the general statistical model format data for the first moment and is formed in the same manner as the general statistical model format data for the first moment at one or more moments before the first moment.
[0045] Figure 3 Sketches 301 and 302 for general statistical model format data fusion according to one or more embodiments of the present invention are shown. Briefly, what is shown in sketch 301 is single fusion, while what is shown in sketch 302 is iterative multiple fusions.
[0046] Sketch 301 is a sketch showing the fusion of the general statistical model format data for the first moment with the historical general statistical model format data including multiple general statistical model format data for multiple previous moments within a threshold time period before the first moment. That is, {t 第一时刻 ,d s1 ,d s2 ...d sn} is fused with {t 第一时刻-1 ,d s1 ,d s2 ...d sn}, {t 第一时刻-2 ,d s1 ,d s2 ...d sn}……{t 第一时刻-tn ,d s1 ,d s2 ...d sn} to update {t 第一时刻 ,d s1 ,ds2 ...d sn}. Among them, between two adjacent moments, for example, t 第一时刻 and t 第一时刻-1 between, t 第一时刻-1 and t 第一时刻-2 between, the interval is the predetermined time interval as described above, and the threshold time period elapsed between t 第一时刻-tn and t 第一时刻 can also be selected according to actual needs. For example, when the predetermined time interval is 0.1 second, the threshold time period can be selected as 1 second, and thus 10 (i.e., in this case, tn is 10) historical general statistical model format data within 1 second before the first moment are selected for fusion. For example, {t 第一时刻 ,d s1 ,d s2 ...d sn} is fused with the 10 historical general statistical model format data within the previous 1 second to update {t 第一时刻 ,d s1 ,d s2 ...d sn} to obtain {t 第一时刻 ,d s1’ ,d s2’ ...d sn’}, where each of the 10 historical general statistical model format data corresponds to the general statistical model format data obtained at intervals of 0.1 meter within 1 second before the first moment. Thus, it can be seen that the method of FIG. 301 is to fuse {t 第一时刻 ,d s1 ,d s2 ...d sn} with the historical general statistical model format data once to update {t 第一时刻 ,d s1 ,d s2 ...d sn}.
[0047] FIG. 302 is a schematic diagram showing iterative fusion of multiple general statistical model format data for multiple previous moments within the threshold time period before the first moment. Continuing with the above example, assume that the threshold time period is 1 second and the predetermined time interval between two adjacent moments is 0.1 second. Iteratively fuse the general statistical model format data for the previous moment with the general statistical model format data for the next moment to update the general statistical model format data for the next moment until the general statistical model format data for the first moment is updated, thereby obtaining {t 第一时刻 ,d s1’ ,d s2’ ...dsn’}.
[0048] For example, first, fuse {t 第一时刻-tn ,d s1 ,d s2 ...d sn} with {t 第一时刻-tn+1 ,d s1 ,d s2 ...d sn} to update {t 第一时刻-tn+1 ,d s1 ,d s2 ...d sn}, obtaining the updated {t 第一时刻-tn+1 ,d s1’ ,d s2’ ...d sn’}. Then, fuse {t 第一时刻-tn+1 ,d s1’ ,d s2’ ...d sn’} with {t 第一时刻-tn+2 ,d s1 ,d s2 ...d sn} to update {t 第一时刻-tn+2 ,d s1 ,d s2 ...d sn}, obtaining the updated {t 第一时刻-tn+2 ,d s1’ ,d s2’ ...d sn’}. Then, fuse {t 第一时刻-tn+2 ,d s1’ ,d s2’ ...d sn’} with {t 第一时刻-tn+3 ,d s1 ,d s2 ...d sn} to update {t 第一时刻-tn+3 ,d s1 ,d s2 ...d sn}, obtaining the updated {t 第一时刻-tn+3 ,d s1’ ,d s2’ ...d sn’}. And so on, until fuse {t 第一时刻-1 ,d s1’ ,d s2’ ...d sn’} with {t 第一时刻 ,d s1 ,d s2 ...d sn} to update {t第一时刻 , d s1 , d s2 ... d sn} to obtain the updated {t 第一时刻 , d s1’ , d s2’ ... d sn’}.
[0049] In order to make the fused data more accurate, the following mathematical methods can be adopted during the fusion process. Figure 4 shows a flowchart of method 400 for fusing general statistical model format data for the first moment and historical general statistical model format data according to an embodiment of Figure 3 . At step 410, historical general statistical model format data is obtained, and the historical general statistical model format data includes multiple general statistical model format data for multiple previous moments within a threshold time period before the first moment.
[0050] At step 420, the general statistical model format data for the first moment and the historical general statistical model format data are converted to the same coordinate system. For example, assume that the vehicle is the origin of the local coordinate system, the traveling direction of the vehicle is the x-axis of the local coordinate system, and the direction perpendicular to the traveling direction of the vehicle is the y-axis of the local coordinate system. Then, as the vehicle travels from time t 第一时刻-1 to t 第一时刻 , and the vehicle travels a distance L in the traveling direction, it can be understood that the origin of the local coordinate system at t 第一时刻 has moved (L 第一时刻-1 , L x , L y ) compared to the origin of the local coordinate system at t 第一时刻 . Through coordinate transformation, the sensor data sets in the collected historical general statistical model format data are converted to the local coordinate system at t 第一时刻 , thereby making all the sensor data sets for fusion in the same coordinate system. According to another embodiment of the present invention, the various coordinates adopted for the collected historical general statistical model format data and the general statistical model format data for the first moment can be uniformly converted to coordinates in the world coordinate system, thereby making all the sensor data sets for fusion in the same coordinate system. Various coordinate transformation methods include, but are not limited to, translation and rotation of coordinates in two-dimensional space, translation and rotation of coordinates in three-dimensional space, etc.
[0051] At step 430, according to Figure 3Any one of the two fusion methods 301 and 302 shown fuses the general statistical model format data for the first moment and the historical general statistical model format data with the sensor data sets in the same coordinate system, so that the general statistical model format data for the first moment is updated to include the fused sensor data. As is known to those skilled in the art, in order to obtain smooth and coherent data, the fusion process includes the aggregation and denoising of the data sets. For example, in the fusion method using 301, it is respectively assumed that the threshold time period is 1 second and the predetermined time interval between two adjacent moments is 0.1 second. The d 第一时刻 ,d s1 ,d s2 ...d sn} included in the sensor data set and the d s1 ,d s2 ...d sn included in the first 10 historical general statistical model format data are aggregated, and then the duplicate data or abnormal data in the aggregated sensor data set are removed or filtered to obtain the updated {t s1 ,d s2 ...d sn} after fusion. Another example is that in the fusion method using 302, the sensor data sets included in the general statistical model format data of two consecutive moments are also similarly aggregated and denoised, thereby updating the general statistical model format data for the latter moment until the general statistical model format data for the first moment is updated. 第一时刻 ,d s1’ ,d s2’ ...d sn’} after fusion. Another example is that in the fusion method using 302, the sensor data sets included in the general statistical model format data of two consecutive moments are also similarly aggregated and denoised, thereby updating the general statistical model format data for the latter moment until the general statistical model format data for the first moment is updated.
[0052] In one embodiment, a weighted average algorithm can also be used for fusion. For example, during aggregation, the historical general statistical model format data recorded closer to the first moment in time is given a higher weight, while the general statistical model format data recorded farther from the first moment in time is given a lower weight. Of course, other weighted methods can also be conceived.
[0053] Return Figure 2 , at step 240, the updated general statistical model format data for the first moment is compared with the implicit model to identify the object. As described above, the implicit model format is {pd s1 ,pd s2 ...pd sn , object}, which represents a set of sensor data sample sets of different types used to describe an object. By comparing the {t 第一时刻 ,d s1’ ,d s2’ ...dsn’}, with one or more {pd s1 , pd s2 ... pd sn , object} for comparison, it can be concluded that {t 第一时刻 , d s1’ , d s2’ ... d sn’} describes a specific object. According to an embodiment of the present invention, assume that the vehicle uses three sensors, and {t 第一时刻 , d s1’ , d s2’ , d s3’} is compared with {pd s1 , pd s2 , pd s3 , object 1}. In this example, d s1’ is compared with pd s1 , d s2’ is compared with pd s2 , d s3’ is compared with pd s3 respectively to determine whether {t 第一时刻 , d s1’ , d s2’ , d s3’} describes the object 1 . In specific practice, it is very likely that not all three comparison results are true among the three comparison results for the three sensors. In this regard, the vehicle manufacturer can pre-define a determination criterion. For example, if two of the three comparison results of the three sensor data sets are true, it is regarded as the overall comparison result being true, or all three comparison results of the three sensor data sets must be true to be regarded as the overall comparison result being true.
[0054] Alternatively, according to pre - settings, the vehicle can automatically select a determination criterion based on the current climate environment, road environment, or the object to be identified. For example, as described above, different types of sensors have different advantages and disadvantages. Depending on the different adaptabilities of the sensors to environmental conditions, in an environment with poor visibility such as haze or heavy fog, the confidence level of the sensor data set output by the millimeter - wave radar can be specified as relatively high, while in a better environmental condition, the confidence levels of the sensor data sets output by lidar and cameras can be specified as relatively high. In addition, depending on the ways in which different types of sensors obtain data, different confidence levels of sensor data sets can be set for different types of objects. For example, for some objects with three - dimensional characteristics such as buildings, the confidence level of the sensor data set output by the camera can be set lower than that of the sensor data sets output by the laser sensor and the millimeter - wave radar sensor. However, for some planar objects such as lane lines and ground traffic signs, the confidence level of the sensor data set output by the camera can also be set lower than that of the sensor data sets output by the laser sensor and the millimeter - wave radar sensor. In this way, generally speaking, the overall determination result can be calculated by the following equation:
[0055] Overall = Confidence s1 *S1 + Confidence s2 *S2 + …… + Confidence sn *Sn.
[0056] Wherein, Confidence s1 + Confidence s2 + …… Confidence sn = 1, S1 is the comparison result between d s1’ and pd s1 , S2 is the comparison result between d s2’ and pd s2 , Sn is the comparison result between d sn’ and pd sn . Among them, S1, S2 …… Sn are 1 or 0. For example, S1 = 1 means that after comparing d s1’ with pd s1 , it is obtained that d s1’ identifies the object described by pd s1 , while S1 = 0 means that after comparing d s1’ with pd s1 , it is obtained that d s1’ does not identify the object described by pd s1 . The same is true for S2…Sn. Thus, the vehicle manufacturer or user can set that if the overall > a predetermined value (for example, 50%), then it can be determined that {t 第一时刻 , d s1’ , d s2’ ... d sn’} identifies {pds1 , pd s2 ... pd sn , the object described by {the object}. Of course, various other different determination criteria can also be conceived.
[0057] After comparing the updated general statistical model format data for the first moment with the implicit model, it is also possible to obtain a situation where the updated general statistical model format data for the first moment does not successfully identify the object. In this case, the map data within the threshold range around the vehicle can be extracted in various ways as the real-time road model for the vehicle to travel. For example, the map can be pre-installed in the memory of the autonomous vehicle or obtained from a map provider through the network. The map can include offline maps such as OSM (Open Street Map) offline maps. Or, in order to obtain further more accurate centimeter-level positioning, a high-precision map (such as the high-precision maps provided by map providers such as Google and HERE) can be used. Those skilled in the art can understand that other types of maps can also be used.
[0058] Thus, by using the method of the present invention, compared with obtaining sensor data separately from various different types of sensors and processing them separately, by putting different types of sensor data sets into a unified data model, the sensor data sets can be processed more quickly.
[0059] Figure 5 is a block diagram of a device 500 for identifying an object according to an embodiment of the present invention. All functional blocks of the device 500 (including each unit in the device 500) can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art should understand that Figure 5 the functional blocks described in can be combined into a single functional block or divided into multiple sub-functional blocks.
[0060] The apparatus 500 may include a sensor data acquisition module 510 configured to acquire a first set of sensor data output by different types of sensors mounted on a vehicle at a first moment. The apparatus 500 may further include a data feeding module 520 configured to feed the first set of sensor data into a general statistical model to form general statistical model format data for the first moment. The apparatus 500 may further include a data fusion module 530 configured to fuse the general statistical model format data for the first moment with historical general statistical model format data to update the general statistical model format data for the first moment. The apparatus 500 may further include an object identification module 540 configured to compare the updated general statistical model format data for the first moment with an implicit model to identify an object.
[0061] Figure 6 FIG. shows a block diagram of an exemplary computing device according to an embodiment of the present invention, which is an example of a hardware device applicable to various aspects of the present invention.
[0062] Referring Figure 6 , a computing device 600 will now be described, which is an example of a hardware device applicable to various aspects of the present invention. The computing device 600 may be any machine configured to implement processing and / or computing, and may be, but is not limited to, a workstation, a server, a desktop computer, a laptop computer, a tablet computer, a personal digital assistant, a smartphone, an in-vehicle computer, or any combination thereof. The foregoing various methods / apparatuses / servers / client devices may be implemented in whole or at least in part by the computing device 600 or a similar device or system.
[0063] The computing device 600 may include components that can be connected or communicate via one or more interfaces and a bus 602. For example, the computing device 600 may include a bus 602, one or more processors 604, one or more input devices 606, and one or more output devices 608. The one or more processors 604 can be any type of processor and may include, but are not limited to, one or more general-purpose processors and / or one or more dedicated processors (e.g., specialized processing chips). The input device 606 can be any type of device capable of inputting information into the computing device and may include, but are not limited to, a mouse, a keyboard, a touch screen, a microphone, and / or a remote controller. The output device 608 can be any type of device capable of presenting information and may include, but are not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The computing device 600 may also include a non-transitory storage device 610 or be connected to the non-transitory storage device, which can be any storage device that is non-transitory and capable of implementing data storage, and the non-transitory storage device may include, but are not limited to, a disk drive, an optical storage device, a solid-state memory, a floppy disk, a flexible disk, a hard disk, a magnetic tape, or any other magnetic medium, an optical disk, or any other optical medium, a ROM (read-only memory), a RAM (random access memory), a cache memory, and / or any storage chip or cartridge, and / or any other medium from which a computer can read data, instructions, and / or code. The non-transitory storage device 610 can be separated from the interface. The non-transitory storage device 610 may have data / instructions / code for implementing the above methods and steps. The computing device 600 may also include a communication device 612. The communication device 612 can be any type of device or system capable of implementing communication with internal devices and / or communication with a network and may include, but are not limited to, a modem, a network card, an infrared communication device, a wireless communication device, and / or a chipset, such as a Bluetooth device, an IEEE 1302.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or a similar device.
[0064] When the computing device 600 is used as an in-vehicle device, it can also be connected to external devices (e.g., a GPS receiver, sensors for sensing different environmental data (such as an acceleration sensor, a wheel speed sensor, a gyroscope, etc.)). In this way, the computing device 600 can receive, for example, positioning data and sensor data indicating the driving condition of the vehicle. When the computing device 600 is used as an in-vehicle device, it can also be connected to other devices for controlling the driving and operation of the vehicle (e.g., an engine system, a windshield wiper, an anti-lock braking system, etc.).
[0065] In addition, the non-transitory storage device 610 may have map information and software components, so that the processor 604 can implement route guidance processing. In addition, the output device 606 may include a display for displaying maps, displaying positioning markers of the vehicle, and displaying images indicating the driving status of the vehicle. The output device 606 may also include a speaker or a headphone jack for audio guidance.
[0066] The bus 602 may include, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus. In particular, for in-vehicle devices, the bus 602 may also include a Controller Area Network (CAN) bus or other architectures designed for automotive applications.
[0067] The computing device 600 may also include a working memory 614, which can be any type of working memory capable of storing instructions and / or data beneficial to the operation of the processor 604 and may include, but is not limited to, random access memory and / or read-only storage devices.
[0068] The software components may be located in the working memory 614, and these software components include, but are not limited to, an operating system 616, one or more application programs 618, drivers, and / or other data and code. The instructions for implementing the above methods and steps may be included in the one or more application programs 618, and the modules / units / components of the foregoing various devices / servers / client devices may be implemented by the processor 604 reading and executing the instructions of the one or more application programs 618.
[0069] It should also be recognized that changes can be made according to specific requirements. For example, custom hardware may also be used, and / or specific components may be implemented in hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. In addition, connections with other computing devices, such as network input / output devices, etc., may be adopted. For example, part or all of the disclosed methods and devices may be implemented by programming hardware (e.g., programmable logic circuits including Field Programmable Gate Arrays (FPGAs) and / or Programmable Logic Arrays (PLAs)) with an assembly language or a hardware programming language (e.g., VERILOG, VHDL, C++) using the logics and algorithms according to the present invention.
[0070] Although aspects of the present invention have been described so far with reference to the accompanying drawings, the above methods, systems, and devices are merely examples, and the scope of the present invention is not limited to these aspects, but is defined only by the appended claims and their equivalents. Various components may be omitted or may be replaced by equivalent components. Additionally, the steps may be implemented in an order different from the order described in the present invention. Further, the various components may be combined in various ways. Also, importantly, as technology develops, many of the components described may be replaced by equivalent components that emerge later.
Claims
1. A method for identifying an object, comprising: At a first moment, obtaining a first set of sensor data output by a plurality of different types of sensors loaded via a vehicle; Feeding the first set of sensor data into a general statistical model to form general statistical model format data for the first moment, the general statistical model format data including a plurality of sensor data sets, wherein each of the plurality of sensor data sets is output by one of the different types of sensors; Fusing the general statistical model format data for the first moment with historical general statistical model format data to update the general statistical model format data for the first moment; and Comparing the updated general statistical model format data for the first moment with an implicit model to identify the object, the implicit model including a plurality of sensor data sample sets for describing a predefined object, wherein each of the plurality of sensor data sample sets includes a pre-acquired sensor data sample set for describing the predefined object output by one of the different types of sensors.
2. The method according to claim 1, wherein fusing the general statistical model format data for the first moment with historical general statistical model format data includes: fusing the general statistical model format data for the first moment with historical general statistical model format data including a plurality of general statistical model format data for a plurality of previous moments within a threshold time period before the first moment.
3. The method according to claim 1, wherein fusing the general statistical model format data for the first moment with historical general statistical model format data includes: For a plurality of general statistical model format data for a plurality of previous moments within a threshold time period before the first moment, iteratively performing: fusing the general statistical model format data for each moment with the general statistical model format data for the next moment after a predetermined time interval to update the general statistical model format data for the next moment, until the general statistical model format data for the first moment is updated.
4. The method according to claim 1, wherein the fusing further includes: making the historical general statistical model format data and the general statistical model format data for the first moment both represented in the same coordinate system by any one of the following methods: taking the position of the vehicle at the first moment as the origin of the local coordinate system to transform the historical general statistical model format data into the local coordinate system, or, uniformly transforming various coordinates used in the historical general statistical model format data and the general statistical model format data for the first moment into coordinates in the world coordinate system.
5. The method according to claim 4, wherein The fusion further includes: correspondingly aggregating the historical general statistical model format data and the sensor data sets output by the same sensors in the general statistical model format data for the first moment, and removing duplicate data in each aggregated sensor data set.
6. The method according to claim 1, wherein, the object includes at least one of the following: road signs, lane lines, buildings, pedestrians, another vehicle, road edges, bridges, utility poles, elevated structures, or traffic signs.
7. The method according to claim 1, wherein, the multiple different types of sensors include vision sensors and radar ranging sensors, the vision sensors include cameras, and the radar includes one or more of lidar, ultrasonic radar, and millimeter wave radar.
8. The method according to claim 1, wherein, comparing the updated general statistical model format data for the first moment with the implicit model to identify an object further includes: depending on the current climate environment or the object to be identified, different confidence levels are assigned to the sensor data sets output by each sensor.
9. An apparatus for identifying an object, comprising: a sensor data acquisition module configured to acquire a first set of sensor data output by different types of sensors mounted on a vehicle at a first moment; a data feeding module configured to feed the first set of sensor data into a general statistical model to form general statistical model format data for the first moment, the general statistical model format data including multiple sensor data sets, wherein each sensor data set in the multiple sensor data sets is output by one of the different types of sensors; a data fusion module configured to fuse the general statistical model format data for the first moment with historical general statistical model format data to update the general statistical model format data for the first moment; and an object identification module configured to compare the updated general statistical model format data for the first moment with an implicit model to identify an object, the implicit model including multiple sensor data sample sets for describing predefined objects, wherein each sensor data sample set in the multiple sensor data sample sets includes a pre-acquired sensor data sample set output by one of the different types of sensors for describing the predefined object.
10. The apparatus according to claim 9, wherein, the data fusion module is further configured to: fuse the general statistical model format data for the first moment with historical general statistical model format data including multiple general statistical model format data for multiple previous moments within a threshold time period before the first moment.
11. The apparatus according to claim 9, wherein, the data fusion module is further configured to: Iteratively perform, for a plurality of general statistical model format data for a plurality of previous moments within a threshold time period before the first moment: fuse the general statistical model format data for each moment with the general statistical model format data for the subsequent moment after a predetermined time interval to update the general statistical model format data for the subsequent moment, until the general statistical model format data for the first moment is updated.
12. The apparatus according to claim 9, wherein, the object identification module is further configured to: depending on the current weather condition or the object to be identified, different confidence levels are assigned to the sensor data sets output by each sensor.
13. The apparatus according to claim 9, wherein, the plurality of different types of sensors include vision sensors and radar ranging sensors, the vision sensors include cameras, and the radar includes one or more of lidar, ultrasonic radar, and millimeter wave radar.
14. A vehicle, comprising: a plurality of different types of sensors; and the apparatus according to any one of claims 9-13.
15. The vehicle according to claim 14, wherein, the plurality of different types of sensors include vision sensors and radar ranging sensors, the vision sensors include cameras, and the radar includes one or more of lidar, ultrasonic radar, and millimeter wave radar.
Citation Information
Patent Citations
Distribution transform-based multi-sensor image fusion method
CN102013095A
Multi-sensor information fusion-based collision and departure pre-warning device and method
CN102303605A