Automatic driving SOTIF via signal representation
By combining camera and radar modules to generate multiple detection representations and using machine learning models to fuse signals from different sensors, the safety issues of autonomous vehicles in the event of sensor failure or misuse are solved, achieving robust object detection and environmental modeling that meets the SOTIF standard.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-09-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing autonomous and semi-autonomous vehicles struggle to meet safety (SOTIF) requirements when detecting environmental information, especially in cases of sensor failure or misuse, where they are unable to effectively identify potential hazards and ensure safe navigation paths.
By combining the signal paths of the camera and radar modules, multiple detection representations are generated. The inputs from different sensors are fused using a machine learning model to achieve redundant detection, thereby improving the sensitivity and robustness of object detection and meeting the SOTIF standard.
The effectiveness of the perception module of the vehicle has been improved, the ability to detect small objects and objects outside the image processing model has been enhanced, a robust environment model has been generated, and safety has been ensured in the event of sensor failure or misuse.
Smart Images

Figure CN122003620A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Patent Application No. 18 / 825,645, filed September 5, 2024, entitled “AUTOMATED DRIVING SOTIF VIASIGNAL REPRESENTATION,” and U.S. Provisional Application No. 63 / 590,899, filed October 17, 2023, entitled “AUTOMATED DRIVING SOTIF VIA SIGNAL REPRESENTATION,” both of which have been assigned to the assignee of this application, and the entire contents of both applications are incorporated herein by reference for all purposes. Background Technology
[0003] As industry increasingly deploys more sophisticated self-driving technologies, vehicles are becoming more intelligent, capable of operating with little or no human input, and thus semi-autonomous or autonomous. Autonomous and semi-autonomous vehicles may be able to detect information about their location and surrounding environment (e.g., using ultrasonic, radar, lidar, SPS (Satellite Positioning System), and / or odometers, and / or one or more sensors such as accelerometers, cameras, etc.). Autonomous and semi-autonomous vehicles typically include a control system to interpret information about the vehicle's environment, thereby identifying hazards and determining a navigation path to follow. The design of autonomous vehicles can leverage industry standards to guide the verification and validation measures required to achieve the intended functionality safety (SOTIF). SOTIF is generally defined as the absence of unreasonable risk due to hazard caused by a malfunction of the intended functionality or by reasonably foreseeable human misuse. Industry standards such as the International Organization for Standardization (ISO) 21448 can provide additional requirements for achieving SOTIF in autonomous and semi-autonomous vehicles. Summary of the Invention
[0004] An example method for generating an object representation having multiple signal paths according to the present disclosure includes: obtaining image information from at least one camera module disposed on a vehicle; obtaining target information from at least one radar module disposed on the vehicle; generating a first detection representation having a first signal path based on the image information and the target information; generating a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and outputting the first detection representation and the second detection representation.
[0005] An example apparatus according to this disclosure includes at least one memory, at least one camera module, at least one radar module, and at least one processor, the at least one processor being communicatively coupled to at least one memory, at least one camera, and at least one radar module, and configured to: acquire image information from at least one camera module disposed on a vehicle; acquire target information from at least one radar module disposed on a vehicle; generate a first detection representation having a first signal path based on the image information and the target information; generate a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and output the first detection representation and the second detection representation.
[0006] The projects and / or technologies described herein can provide one or more of the following capabilities, as well as others not mentioned: Multiple sensors, such as cameras, radar, and lidar, can acquire target information about objects in proximity to autonomous or semi-autonomous vehicles. Sensor inputs can be evaluated via different signal paths. Machine learning models can be implemented along different signal paths. Parametric and non-parametric representations of object data can be generated. Fusion of signals from different sensors can improve the sensitivity of object detection. Multiple signal paths can improve the robustness of object detection and the corresponding environment model. The SOTIF standard can be implemented. Other capabilities can be provided, and not every specific embodiment of this disclosure is required to provide any, let alone all, of the capabilities discussed. Attached Figure Description
[0007] Figure 1 This is a top view of an example of a vehicle.
[0008] Figure 2 yes Figure 1 The self-driving vehicle shown can be a block diagram of the components of an example device.
[0009] Figure 3 This is a block diagram of the components of an example send / receive point.
[0010] Figure 4 This is a block diagram of the server components.
[0011] Figure 5 This is a block diagram of the example device.
[0012] Figure 6 This is a map illustrating an example geographical environment.
[0013] Figure 7 It is divided into grids Figure 6 The map shows the geographical environment.
[0014] Figure 8 Is with Figure 7An example of the occupancy map corresponding to the grid shown.
[0015] Figure 9 This is a block diagram of an example system for robust object detection using different signal representations.
[0016] Figure 10 This is a block diagram of a first example functional architecture for implementing redundancy in object detection signals.
[0017] Figure 11 This is a block diagram of a second example functional architecture for implementing redundancy in object detection signals.
[0018] Figure 12A This is an example deep learning architecture for object detection.
[0019] Figure 12B This is another example of a deep learning architecture used for object detection.
[0020] Figure 13 This is a flowchart of an example method for generating an object representation with multiple signal paths.
[0021] Figure 14 This is a flowchart of a sample method for generating a list of one or more objects based on multiple signal paths. Detailed Implementation
[0022] This paper discusses techniques for detecting objects near vehicles with multiple signal paths. Building robust environment models is an important aspect of autonomous driving systems. Industry standards may require some degree of redundancy to meet SOTIF requirements. In one example, sensor-based redundancy can be implemented to mitigate the impact of sensor failure. Redundancy can also be achieved through specific implementations of different signal paths. Signals received from various sensors, such as cameras, radar modules, and lidar modules, can be jointly processed and fused in separate signal paths. In one example, a first signal path can be configured to generate a parametric representation of the object based on the fusion of camera and radar inputs, and a second signal path can be configured to generate a non-parametric representation of the object based on both camera and radar inputs. Other sensor inputs can also be used to generate both parametric and non-parametric representations of the object. For example, various combinations of image, radar, and lidar signals can be fused to generate representations. Machine learning models can be implemented to generate representations of the detected objects. Different signal paths can be configured to use different backbone networks in the machine learning model. In one example, a common backbone network can be utilized, and separate training for each head can be enforced to handle redundancy. However, other techniques can be used.
[0023] Specific aspects of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages: Redundancy requirements of the SOTIF standard can be met. The performance of object detection based on multi-sensor fusion can be maintained compared to single-sensor object detection techniques, and the effectiveness of perception modules on vehicles can be improved. The fusion of object detection results from different types of sensors enables the detection of smaller objects or objects beyond the training of image processing models. Robust environment models can be generated based on improved object detection and redundant signal paths. Other advantages can also be achieved.
[0024] refer to Figure 1 The self-driving vehicle 100 includes a self-driving vehicle driver assistance system 110. The driver assistance system 110 may include multiple sensors of different types mounted at appropriate locations on the self-driving vehicle 100. For example, the system 110 may include: a pair of diverging and outward-pointing radar sensors 121 mounted at respective front corners of the vehicle 100; a similar pair of diverging and outward-pointing radar sensors 122 mounted at respective rear corners of the vehicle 100; a forward-pointing LRR sensor 123 (long-range radar) centrally mounted at the front of the vehicle 100; and a pair of generally forward-pointing optical sensors 124 (cameras) forming part of an SVS 126 (stereo vision system), which may be mounted, for example, in the area of the upper edge of the windshield 128 of the vehicle 100. Each of the sensors 121, 122 may include LRR and / or SRR (short-range radar). Various sensors 121 to 124 are operatively connected to a central electronic control system, which is typically provided in the form of an ECU 140 (Electronic Control Unit) installed in a convenient location within the vehicle 100. In the particular arrangement illustrated, the front sensor 121 and the rear sensor 122 are connected to the ECU 140 via one or more conventional Controller Area Network (CAN) buses 150, and the sensors of the LRR sensor 123 and the SVS 126 are connected to the ECU 140 via a serial bus 160 (e.g., a faster FlexRay serial bus).
[0025] Together, and under the control of ECU 140, various sensors 121 to 124 can be used to provide a variety of different types of driver assistance functions. For example, sensors 121 to 124 and ECU 140 can provide blind spot monitoring, adaptive cruise control, collision prevention assist, lane departure protection, and / or rear collision mitigation.
[0026] The CAN bus 150 can be regarded by the ECU 140 as a sensor that provides self-vehicle parameters to the ECU 140. For example, a GPS module can also be connected to the ECU 140 as a sensor to provide geographic location parameters to the ECU 140.
[0027] Also refer to Figure 2Device 200 (which may be a mobile device such as User Equipment (UE) or a Vehicle Equipment (VUE)) includes a computing platform containing processor 210, a memory 211 containing software (SW) 212, one or more sensors 213, a transceiver interface 214 for transceivers 215 (which includes a wireless transceiver 240 and a wired transceiver 250), a user interface 216, a satellite positioning system (SPS) receiver 217, a camera 218, and a positioning device (PD) 219. The terms “User Equipment” or “UE” (or variations thereof) are not specific to or otherwise limited to any particular radio access technology (RAT) unless otherwise indicated. Processor 210, memory 211, sensors 213, transceiver interface 214, user interface 216, SPS receiver 217, camera 218, and positioning device 219 may be communicatively coupled to each other via bus 220 (which may be configured for, for example, optical and / or electrical communication). One or more of the devices shown in the apparatus (e.g., camera 218, positioning device 219, and / or one or more sensors in sensor 213, etc.) may be omitted from device 200. Processor 210 may include one or more hardware devices, such as a central processing unit (CPU), microcontroller, application-specific integrated circuit (ASIC), etc. Processor 210 may include multiple processors, including a general-purpose / application processor 230, a digital signal processor (DSP) 231, a modem processor 232, a video processor 233, and / or a sensor processor 234. One or more of processors 230 to 234 may include multiple devices (e.g., multiple processors). For example, sensor processor 234 may include processors for, for example, RF (radio frequency) sensing (where one or more transmitted (cellular) wireless signals and reflections are used to identify, map, and / or track objects) and / or ultrasound, etc. Modem processor 232 may support dual SIM / dual connectivity (or even more SIMs). For example, a SIM (subscriber identity module or subscriber identification module) may be used by an original equipment manufacturer (OEM), and another SIM may be used by the end user of device 200 to obtain connectivity. Memory 211 may be a non-transitory storage medium including random access memory (RAM), flash memory, disk storage, and / or read-only memory (ROM). Memory 211 may store software 212, which may be processor-readable, processor-executable software code containing instructions that can be configured to cause processor 210 to perform the various functions described herein when executed. Alternatively, software 212 may not be directly executable by processor 210, but may be configured to cause processor 210 to perform these functions, for example, when compiled and executed. The description herein may refer to processor 210 performing functions, but this includes other specific implementations, such as specific implementations of instructions by processor 210 executing software and / or firmware.The description herein may refer to the functions performed by processor 210 as a shorthand for one or more processor functions performed by processors 230 to 234. The description herein may also refer to the functions performed by device 200 as a shorthand for one or more suitable components of device 200 performing that function. In addition to and / or instead of memory 211, processor 210 may include memory containing stored instructions. The functionality of processor 210 is discussed more fully below.
[0028] Figure 2 The configuration of device 200 shown is exemplary and not intended to limit this disclosure (including the claims), and other configurations may be used. For example, an example configuration of the UE may include one or more of processors 230 to 234 in processor 210, memory 211, and wireless transceiver 240. Other example configurations may include one or more of processors 230 to 234 in processor 210, memory 211, wireless transceiver, and one or more of the following devices: sensor 213, user interface 216, SPS receiver 217, camera 218, PD 219, and / or wired transceiver.
[0029] Device 200 may include a modem processor 232 capable of performing baseband processing on signals received and downconverted by transceiver 215 and / or SPS receiver 217. Modem processor 232 may also perform baseband processing on signals to be upconverted for transmission by transceiver 215. Alternatively or additionally, baseband processing may be performed by general-purpose / application processor 230 and / or DSP 231. However, other configurations may be used to perform baseband processing.
[0030] Device 200 may include sensor 213, which may include one or more sensors of various types, such as one or more inertial sensors, one or more magnetometers, one or more environmental sensors, one or more optical sensors, one or more weight sensors, and / or one or more radio frequency (RF) sensors. An inertial measurement unit (IMU) may include, for example, one or more accelerometers (e.g., collectively responding to acceleration of device 200 in three dimensions) and / or one or more gyroscopes (e.g., three-dimensional gyroscopes). Sensor 213 may include one or more magnetometers (e.g., three-dimensional magnetometers) to determine orientation (e.g., relative to magnetic north and / or true north), which can be used for any of a variety of purposes (e.g., to support one or more compass applications). Environmental sensors may include, for example, one or more temperature sensors, one or more barometric pressure sensors, one or more ambient light sensors, one or more camera imagers, and / or one or more microphones. Sensor 213 may generate analog and / or digital signals, indications of which may be stored in memory 211 and processed by DSP 231 and / or general-purpose / application processor 230 to support one or more applications (e.g., applications involving positioning and / or navigation operations).
[0031] Sensor 213 can be used for relative position measurement, relative position determination, motion determination, etc. Information detected by sensor 213 can be used for motion detection, relative displacement, dead reckoning, sensor-based position determination, and / or sensor-assisted position determination. Sensor 213 can be used to determine whether device 200 is stationary or mobile and / or whether certain useful information related to the mobility of device 200 needs to be reported to the LMF (Location Management Function), for example. For example, based on information obtained / measured by sensor 213, device 200 can notify / report to the LMF that device 200 has detected movement or that device 200 has moved, and report relative displacement / distance (e.g., via dead reckoning implemented by sensor 213, or sensor-based position determination, or sensor-assisted position determination). In another example, for relative positioning information, the sensor / IMU can be used to determine the angle and / or orientation of another object (another device) relative to device 200, etc.
[0032] The IMU can be configured to provide measurements of the direction and / or velocity of motion of device 200, which can be used for relative position determination. For example, one or more accelerometers and / or one or more gyroscopes of the IMU can detect the linear acceleration and rotational velocity of device 200, respectively. The linear acceleration and rotational velocity measurements of device 200 can be integrated over time to determine the instantaneous direction of motion and displacement of device 200. The instantaneous direction of motion and displacement can be integrated to track the position of device 200. For example, a reference position of device 200 at a certain moment can be determined, for example, using SPS receiver 217 (and / or by some other means), and measurements acquired from the accelerometers and gyroscopes after that moment can be used for dead reckoning to determine the current position of device 200 based on the movement (direction and distance) of device 200 relative to that reference position.
[0033] A magnetometer can determine the strength of a magnetic field in different directions, which can be used to determine the orientation of device 200. For example, this orientation can be used to provide a digital compass for device 200. The magnetometer may include a two-dimensional magnetometer configured to detect and provide an indication of the magnetic field strength in two orthogonal dimensions. The magnetometer may also include a three-dimensional magnetometer configured to detect and provide an indication of the magnetic field strength in three orthogonal dimensions. The magnetometer may provide components for sensing the magnetic field and, for example, providing an indication of the magnetic field to processor 210.
[0034] Transceiver 215 may include a wireless transceiver 240 and a wired transceiver 250 configured to communicate with other devices via wireless and wired connections, respectively. For example, wireless transceiver 240 may include a wireless transmitter 242 and a wireless receiver 244 coupled to antenna 246 for transmitting (e.g., on one or more uplink channels and / or one or more sidelink channels) and / or receiving (e.g., on one or more downlink channels and / or one or more sidelink channels) wireless signals 248 and converting signals from wireless signals 248 to wired (e.g., electrical and / or optical) signals and from wired (e.g., electrical and / or optical) signals to wireless signals 248. Wireless transmitter 242 includes suitable components (e.g., power amplifiers and digital-to-analog converters). Wireless receiver 244 includes suitable components (e.g., one or more amplifiers, one or more frequency filters, and analog-to-digital converters). Wireless transmitter 242 may include multiple transmitters that may be discrete components or combined / integrated components, and / or wireless receiver 244 may include multiple receivers that may be discrete components or combined / integrated components. The wireless transceiver 240 can be configured to transmit signals according to a variety of radio access technologies (RATs) (e.g., with TRP and / or one or more other devices), such as 5G New Radio (NR), GSM (Global System for Mobile Communications), UMTS (Universal Mobile Telecommunications System), AMPS (Advanced Mobile Telephone Systems), CDMA (Code Division Multiple Access), WCDMA (Wideband CDMA), LTE (Long Term Evolution), LTE Direct (LTE-D), 3GPP LTE-V2X (PC5), IEEE 802.11 (including IEEE 802.11p), and WiFi. ® Short-range wireless communication technology, WiFi ® Direct connection (WiFi-D), Bluetooth ® Short-range wireless communication technology, Zigbee ®Short-range wireless communication technologies, etc. The new radio can use millimeter-wave frequencies and / or frequencies below 6 GHz. Wired transceiver 250 may include a wired transmitter 252 and a wired receiver 254 configured for wired communication, for example, a network interface that can be used to communicate with NG-RAN (Next Generation Radio Access Network) to transmit communications to and receive communications from NG-RAN. Wired transmitter 252 may include multiple transmitters that can be discrete components or combined / integrated components, and / or wired receiver 254 may include multiple receivers that can be discrete components or combined / integrated components. Wired transceiver 250 may be configured, for example, for optical and / or electrical communication. Transceiver 215 may be communicatively coupled to transceiver interface 214, for example, via optical and / or electrical connections. Transceiver interface 214 may be at least partially integrated with transceiver 215. The wireless transmitter 242, the wireless receiver 244, and / or the antenna 246 may each include multiple transmitters, multiple receivers, and / or multiple antennas for transmitting and / or receiving appropriate signals, respectively.
[0035] User interface 216 may include one or more of a number of devices, such as speakers, microphones, display devices, vibration devices, keyboards, touchscreens, etc. User interface 216 may include more than one of these devices. User interface 216 may be configured to enable a user to interact with one or more applications hosted by device 200. For example, user interface 216 may store indications of analog and / or digital signals in memory 211 in response to actions from the user, for processing by DSP 231 and / or general-purpose / application processor 230. Similarly, applications hosted on device 200 may store indications of analog and / or digital signals in memory 211 to present output signals to the user. User interface 216 may include audio input / output (I / O) devices, including, for example, speakers, microphones, digital-to-analog circuitry, analog-to-digital circuitry, amplifiers, and / or gain control circuitry (including more than one of these devices). Other configurations of the audio I / O devices may be used. Additionally or alternatively, the user interface 216 may include one or more touch sensors that respond to touch and / or pressure on, for example, the keyboard and / or touchscreen of the user interface 216.
[0036] SPS receiver 217 (e.g., a Global Positioning System (GPS) receiver) may be able to receive and acquire SPS signal 260 via SPS antenna 262. SPS antenna 262 is configured to convert SPS signal 260 from a wireless signal to a wired signal (e.g., an electrical or optical signal) and may be integrated with antenna 246. SPS receiver 217 may be configured to process the acquired SPS signal 260 fully or partially for estimating the location of device 200. For example, SPS receiver 217 may be configured to determine the location of device 200 by performing trilateration using SPS signal 260. SPS receiver 217 may be used in conjunction with general-purpose / application processor 230, memory 211, DSP 231, and / or one or more dedicated processors (not shown) to process the acquired SPS signal fully or partially and / or calculate the estimated location of device 200. Memory 211 may store indications (e.g., measurements) of SPS signal 260 and / or other signals (e.g., signals acquired from wireless transceiver 240) for use in performing positioning operations. General-purpose / application processor 230, DSP 231, and / or one or more dedicated processors, and / or memory 211 may provide or support a position engine for use in processing measurements to estimate the position of device 200.
[0037] Device 200 may include a camera 218 for capturing still or moving images. Camera 218 may include, for example, an imaging sensor (e.g., a charge-coupled device or CMOS (complementary metal-oxide-semiconductor) imager), lenses, analog-to-digital circuitry, frame buffers, etc. Additional processing, conditioning, encoding, and / or compression of signals representing the captured images may be performed by general-purpose / application processor 230 and / or DSP 231. Additionally or alternatively, video processor 233 may perform conditioning, encoding, compression, and / or manipulation of signals representing the captured images. Video processor 233 may decode / decompress stored image data for presentation on a display device (not shown), for example, user interface 216.
[0038] Location device (PD) 219 may be configured to determine the location of device 200, the movement of device 200, and / or the relative location of device 200, and / or time. For example, PD 219 may communicate with SPS receiver 217 and / or include part or all of the SPS receiver. PD 219 may, where appropriate, operate in conjunction with processor 210 and memory 211 to perform at least a portion of one or more location methods, although the description herein may refer to PD 219 being configured to perform according to a location method or the PD performing according to a location method. PD 219 may additionally or alternatively be configured to use terrestrial signals (e.g., at least some of the radio signals in radio signals 248) for trilateration, assisted acquisition, and use of SPS signal 260, or both, to determine the location of device 200. PD 219 may be configured to determine the location of device 200 based on the coverage area of a serving base station and / or another technology (such as E-CID). PD 219 may be configured to determine the location of device 200 using one or more images from camera 218 and image recognition combined with the known location of landmarks (e.g., natural landmarks such as mountains and / or man-made landmarks such as buildings, bridges, streets, etc.). PD 219 may be configured to determine the location of device 200 using one or more other technologies (e.g., relying on the UE's self-reported location (e.g., part of the UE's positioning beacon)), and may use a combination of these technologies (e.g., SPS and terrestrial positioning signals) to determine the location of device 200. PD 219 may include one or more sensors among sensors 213 (e.g., gyroscopes, accelerometers, magnetometers, etc.) that can sense the orientation and / or motion of device 200 and provide an indication of such orientation and / or motion, and processor 210 (e.g., general-purpose / application processor 230 and / or DSP 231) may be configured to use this indication to determine the motion of device 200 (e.g., velocity vector and / or acceleration vector). PD 219 can be configured to provide an indication of uncertainty and / or error in the determined positioning and / or motion. The functionality of PD 219 can be provided in a variety of ways and / or configurations, such as by a general-purpose / application processor 230, transceiver 215, SPS receiver 217 and / or another component of device 200, and can be provided by hardware, software, firmware or various combinations thereof.
[0039] Also refer to Figure 3Examples of TRP 300 (such as gNB (General Node B) and / or ng-eNB (Next Generation Evolved Node B) base stations) may include a computing platform containing processor 310, memory 311 containing software (SW) 312, and transceiver 315. Even when cited in the singular, processor 310 may include one or more processors, transceiver 315 may include one or more transceivers (e.g., one or more transmitters and / or one or more receivers), and / or memory 311 may include one or more memories. Processor 310, memory 311, and transceiver 315 may be communicatively coupled to each other via bus 320 (which may be configured for, for example, optical and / or electrical communications). One or more devices in the illustrated apparatus (e.g., wireless transceivers) may be omitted from TRP 300. Processor 310 may include one or more hardware devices, such as a central processing unit (CPU), microcontroller, application-specific integrated circuit (ASIC), etc. Processor 310 may include multiple processors (e.g., including general-purpose / application processors, DSPs, modem processors, video processors, and / or sensor processors, such as...) Figure 2 (As shown). Memory 311 may be a non-transitory storage medium including random access memory (RAM), flash memory, disk storage, and / or read-only memory (ROM). Memory 311 may store software 312, which may be processor-readable, processor-executable software code containing instructions configured to cause processor 310 to perform the various functions described herein when executed. Alternatively, software 312 may not be directly executable by processor 310, but may be configured to cause processor 310 to perform these functions, for example, when compiled and executed.
[0040] The description herein may refer to the functionality performed by processor 310, but this includes other specific implementations, such as specific implementations of software and / or firmware performed by processor 310. The description herein may refer to the functionality performed by processor 310 as an abbreviation for the functionality performed by one or more processors included in processor 310. The description herein may refer to the functionality performed by TRP 300 as an abbreviation for the functionality performed by one or more suitable components of TRP 300 (e.g., processor 310 and memory 311). In addition to and / or instead of memory 311, processor 310 may include memory with stored instructions. The functionality of processor 310 is discussed more fully below.
[0041] Transceiver 315 may include a wireless transceiver 340 and / or a wired transceiver 350 configured to communicate with other devices via wireless and wired connections, respectively. For example, wireless transceiver 340 may include a wireless transmitter 342 and a wireless receiver 344 coupled to one or more antennas 346 for transmitting (e.g., on one or more uplink channels and / or one or more downlink channels) and / or receiving (e.g., on one or more downlink channels and / or one or more uplink channels) wireless signals 348 and converting signals from wireless signals 348 into guided (e.g., electromagnetic, electrical, and / or optical) signals and from guided (e.g., electromagnetic, electrical, and / or optical) signals into wireless signals 348. Therefore, wireless transmitter 342 may include multiple transmitters that may be discrete components or combined / integrated components, and / or wireless receiver 344 may include multiple receivers that may be discrete components or combined / integrated components. The wireless transceiver 340 can be configured to transmit signals according to a variety of radio access technologies (RATs) (e.g., with device 200, one or more other UEs, and / or one or more other devices), such as 5G New Radio (NR), GSM (Global System for Mobile Communications), UMTS (Universal Mobile Telecommunications System), AMPS (Advanced Mobile Telephone Systems), CDMA (Code Division Multiple Access), WCDMA (Wideband CDMA), LTE (Long Term Evolution), LTE Direct (LTE-D), 3GPP LTE-V2X (PC5), IEEE 802.11 (including IEEE 802.11p), and WiFi. ® Short-range wireless communication technology, WiFi ® Direct connection (WiFi) ® -D), Bluetooth ® Short-range wireless communication technology, Zigbee ® Short-range wireless communication technologies, etc. The wired transceiver 350 may include a wired transmitter 352 and a wired receiver 354 configured for wired communication, for example, a network interface used to communicate with NG-RAN to transmit and receive communications to, for example, LMF and / or one or more other network entities. The wired transmitter 352 may include multiple transmitters that can be discrete components or combined / integrated components, and / or the wired receiver 354 may include multiple receivers that can be discrete components or combined / integrated components. The wired transceiver 350 may be configured, for example, for optical communication and / or electrical communication.
[0042] Figure 3The configuration of TRP 300 shown is illustrative and not intended to limit this disclosure (including the claims), and other configurations may be used. For example, the description herein discusses that TRP 300 may be configured to perform several functions or that the TRP performs several functions, but one or more of these functions may be performed by LMF and / or device 200 (i.e., LMF and / or device 200 may be configured to perform one or more of these functions).
[0043] Also refer to Figure 4 Server 400 (LMF is an example thereof) may include a computing platform containing processor 410, memory 411 containing software (SW) 412, and transceiver 415. Even when cited in the singular, processor 410 may include one or more processors, transceiver 415 may include one or more transceivers (e.g., one or more transmitters and / or one or more receivers), and / or memory 411 may include one or more memories. Processor 410, memory 411, and transceiver 415 may be communicatively coupled to each other via bus 420 (which may be configured for, for example, optical communication and / or electrical communication). One or more devices in the illustrated apparatus (e.g., wireless transceivers) may be omitted from server 400. Processor 410 may include one or more hardware devices, such as a central processing unit (CPU), microcontroller, application-specific integrated circuit (ASIC), etc. Processor 410 may include multiple processors (e.g., including general-purpose / application processors, DSPs, modem processors, video processors, and / or sensor processors, such as... Figure 2 (As shown). Memory 411 may be a non-transitory storage medium including random access memory (RAM), flash memory, disk storage, and / or read-only memory (ROM). Memory 411 may store software 412, which may be processor-readable, processor-executable software code containing instructions configured to cause processor 410 to perform the various functions described herein when executed. Alternatively, software 412 may not be directly executable by processor 410, but may be configured to cause processor 410 to perform these functions, for example, when compiled and executed. The description herein may refer to processor 410 performing functions, but this includes other specific implementations, such as specific implementations of processor 410 performing software and / or firmware. The description herein may refer to the function performed by processor 410 as an abbreviation for one or more processors included in processor 410 performing functions. The description herein may refer to the function performed by server 400 as an abbreviation for one or more suitable components of server 400 performing functions. In addition to and / or instead of memory 411, processor 410 may include memory with stored instructions. The functionality of the processor 410 will be discussed more comprehensively below.
[0044] Transceiver 415 may include a wireless transceiver 440 and / or a wired transceiver 450 configured to communicate with other devices via wireless and wired connections, respectively. For example, wireless transceiver 440 may include a wireless transmitter 442 and a wireless receiver 444 coupled to one or more antennas 446 for transmitting (e.g., on one or more downlink channels) and / or receiving (e.g., on one or more uplink channels) wireless signals 448 and converting signals from wireless signals 448 into guided (e.g., electromagnetic, electrical, and / or optical) signals and from guided (e.g., electromagnetic, electrical, and / or optical) signals into wireless signals 448. Therefore, wireless transmitter 442 may include multiple transmitters that may be discrete components or combined / integrated components, and / or wireless receiver 444 may include multiple receivers that may be discrete components or combined / integrated components. The wireless transceiver 440 can be configured to transmit signals according to a variety of radio access technologies (RATs) (e.g., with device 200, one or more other UEs, and / or one or more other devices), such as 5G New Radio (NR), GSM (Global System for Mobile Communications), UMTS (Universal Mobile Telecommunications System), AMPS (Advanced Mobile Telephone Systems), CDMA (Code Division Multiple Access), WCDMA (Wideband CDMA), LTE (Long Term Evolution), LTE Direct (LTE-D), 3GPP LTE-V2X (PC5), IEEE 802.11 (including IEEE 802.11p), and WiFi. ® Short-range wireless communication technology, WiFi ® Direct connection (WiFi) ® -D), Bluetooth ® Short-range wireless communication technology, Zigbee ® Short-range wireless communication technologies, etc. Wired transceiver 450 may include a wired transmitter 452 and a wired receiver 454 configured for wired communication, for example, a network interface used to communicate with NG-RAN to transmit and receive communications to, for example, LMF 300 and / or one or more other network entities. Wired transmitter 452 may include multiple transmitters that can be discrete components or combined / integrated components, and / or wired receiver 454 may include multiple receivers that can be discrete components or combined / integrated components. Wired transceiver 450 may be configured, for example, for optical communication and / or electrical communication.
[0045] The description herein may refer to the functionality performed by processor 410, but this includes other specific implementations, such as specific implementations of software and / or firmware (stored in memory 411) performed by processor 410. The description herein may refer to the functionality performed by server 400 as an abbreviation for the functionality performed by one or more appropriate components of server 400 (e.g., processor 410 and memory 411).
[0046] Figure 4 The configuration of server 400 shown is exemplary and not intended to limit this disclosure (including the claims), and other configurations may be used. For example, wireless transceiver 440 may be omitted. Additionally or alternatively, this specification discusses server 400 being configured to perform certain functions or the server performing certain functions, but one or more of these functions may be performed by TRP 300 and / or device 200 (i.e., TRP 300 and / or device 200 may be configured to perform one or more of these functions).
[0047] refer to Figure 5 Device 500 includes a processor 510, a transceiver 520, a memory 530, and a sensor 540 communicatively coupled to each other via a bus 550. Even when cited in the singular, processor 510 may include one or more processors, transceiver 520 may include one or more transceivers (e.g., one or more transmitters and / or one or more receivers), and memory 530 may include one or more memories. Device 500 may take any of a variety of forms, such as a mobile device, such as a vehicle UE (VUE). Device 500 may include... Figure 5 The components shown may include one or more other components, such as Figure 2 Any of the components shown makes device 200 an example of device 500. For example, processor 510 may include one or more components of processor 210. Transceiver 520 may include one or more components of transceiver 215, such as wireless transmitter 242 and antenna 246, or wireless receiver 244 and antenna 246, or wireless transmitter 242, wireless receiver 244 and antenna 246. Additionally or alternatively, transceiver 520 may include wired transmitter 252 and / or wired receiver 254. Memory 530 may be configured similarly to memory 211, for example including software having processor-readable instructions configured to cause processor 510 to perform functions. Sensor 540 includes one or more radar sensors 542 and one or more cameras 544. Sensor 540 may include one or more other sensors, such as lidar, Hall effect sensors, ultrasonic sensors, and / or one or more other sensors configured to assist in vehicle operation.
[0048] The description herein may refer to processor 510 performing a function, but this includes other specific implementations, such as specific implementations of software and / or firmware (stored in memory 530) performed by processor 510. The description herein may refer to device 500 performing a function as a shorthand for one or more suitable components of device 500 (e.g., processor 510 and memory 530) performing that function. Processor 510 (possibly in conjunction with memory 530 and, where appropriate, transceiver 520) may include an occupying grid unit 560 (which may include ADAS (Advanced Driver Assistance Systems) for VUE). Occupying grid unit 560 is discussed further below, and the description herein may refer to occupying grid unit 560 performing one or more functions, and / or may generally refer to processor 510 or device 500 as performing any function of occupying grid unit 560, wherein device 500 is configured to perform those functions.
[0049] One or more functions performed by device 500 (e.g., occupying grid cell 560) may be performed by another entity. For example, sensor measurements (e.g., radar measurements, camera measurements (e.g., pixels, images)) and / or processed sensor measurements (e.g., camera images converted into bird's-eye view images) may be provided to another entity, such as server 400, and the other entity may perform one or more functions discussed herein with respect to occupying grid cell 560 (e.g., using machine learning to determine and / or apply observation models, analyzing measurements from different sensors to determine the current occupied grid, etc.).
[0050] Also refer to Figure 6The geographic environment 600 (in this example, a driving environment) includes multiple mobile wireless communication devices (here, vehicles 601, 602, 603, 604, 605, 606, 607, 608, 609), buildings 610, RSUs 612 (roadside units), and street signs 620 (e.g., stop signs). RSUs 612 may be configured similarly to TRPs 300, but may have less functionality and / or shorter range than TRPs 300 (e.g., base station-based TRPs). One or more vehicles among vehicles 601 to 609 may be configured to perform autonomous driving. A vehicle considering its perspective (e.g., for environmental assessment, autonomous driving, etc.) may be referred to as an observer vehicle or a self-vehicle. Self-vehicles such as vehicle 601 may assess the area around the self-vehicle for one or more desired purposes (e.g., to facilitate autonomous driving). Vehicle 601 may be an example of device 500. The vehicle 601 can divide the area around the vehicle into multiple sub-areas and assess whether an object occupies each sub-area, and if so, determine one or more characteristics of the object (e.g., size, shape (e.g., dimensions (possibly including height)), speed (rate and direction), object type or category (bicycle, car, truck, etc.) etc.).
[0051] Also refer to Figure 7 and Figure 8In this example, a region 700 spanning a portion of environment 600 can be evaluated to determine an occupancy grid 800 (also called an occupancy map), which indicates multiple probabilities for each cell of grid 800 whether the cell is occupied or vacant and whether the occupancy object is static or dynamic. For example, region 700 may be divided into grids with sub-regions 710 (which may be called occupancy grids), which may be similar (e.g., identical) in size and shape, or may have two or more sizes and / or shapes (e.g., where sub-regions are smaller near the vehicle (e.g., vehicle 601) and larger further away from the vehicle, and / or where sub-regions near the vehicle have a different shape than sub-regions further away from the vehicle). Region 700 and grid 800 can be of regular shapes (e.g., rectangles, triangles, hexagons, octagons, etc.) and / or, for convenience (e.g., to simplify calculations), can be divided into sub-regions of the same or regular shape, but other shapes of regions / grids (e.g., irregular shapes) and / or sub-regions (e.g., irregular shapes, multiple different regular shapes, or a combination of one or more irregular shapes and one or more regular shapes) can be used. For example, sub-region 710 can have a rectangular (e.g., square) shape. Region 700 can be any of a variety of sizes and can have any of a variety of granularity of sub-regions. For example, region 700 can be a rectangle (e.g., a square) with approximately 100m on each side. As another example, although region 700 is shown as having a sub-region 710 with a square of approximately 1m on each side, other sizes of sub-regions can be used, including much smaller sub-regions. For example, a square sub-region with approximately 25cm on each side can be used. In this example, region 700 is divided into M rows (here, parallel to...). Figure 8 The indicated x-axis has 24 rows, each with N columns (here, parallel to...). Figure 8 (The y-axis has 23 columns shown). As another example, the grid can include sub-regions of a 512 x 512 array. Other specific implementations that occupy the grid are also possible.
[0052] Each subregion in subregion 710 may correspond to a corresponding cell 810 in the occupancy map, and information about what (if any) occupies each subregion in subregion 710 and whether the occupant is static or dynamic can be obtained so that cells 810 of the occupancy grid 800 are filled with the probability that the cell is occupied (O) or free (F) (i.e., unoccupied) and the probability that the object occupying the cell is at least partially static (S) or dynamic (D). Each probability may be a floating-point value. Information about what (if any) occupies each subregion in subregion 710 may be obtained from a variety of sources. For example, occupancy information may be obtained from sensor measurements from sensor 540 of device 500. As another example, occupancy information may be obtained by one or more other devices and communicated to device 500. For example, one or more vehicles among vehicles 602 to 609 may communicate occupancy information to vehicle 601, for example, via C-V2X communication. As another example, RSU 612 may (e.g., from one or more sensors of RSU 612 and / or from communication with one or more vehicles of vehicles 602 to 609 and / or one or more other devices) collect occupancy information and communicate the collected information to vehicle 601, for example, directly and / or through one or more network entities (e.g., TRP).
[0053] like Figure 8 As shown, each cell in cell 810 may include a set of occupancy information 820, which indicates a dynamic probability 821 (P). D ), static probability 822 (P) S ), idle probability 823 (P) F ), the probability of being occupied is 824 (P) P The probability 821 indicates the probability that an object (if any) in the corresponding sub-region 710 is dynamic. The probability 822 indicates the probability that an object (if any) in the corresponding sub-region 710 is static. The probability 823 indicates the probability that no object exists in the corresponding sub-region 710. The probability 824 indicates the probability that an object exists in (any part of) the corresponding sub-region 710. Each cell in cell 810 may include the corresponding probabilities 821 to 824 of whether the object in cell 810 is static, dynamic, non-existent, or present, where the sum of the probabilities is 1. Figure 8 In the example shown, for the sake of diagram simplicity and readability of occupied grid 800, cells that are more likely to be free (empty) than occupied are not marked in occupied grid 800. Furthermore, as... Figure 8As shown, cells that are more likely to be occupied than free and occupied by objects that are more likely to be dynamic than static are marked with "D", and cells that are more likely to be occupied than free and occupied by objects that are more likely to be static than dynamic are marked with "S". A self-driving vehicle may not be able to determine whether a cell is occupied (e.g., behind the visible surface of an object and not based on observations of the object (e.g., if the size and shape of the detected object are unknown)), and such cells may be marked as unoccupied.
[0054] Constructing a dynamic occupancy grid (an occupancy grid with dynamic occupancy types) can be helpful, or even necessary, for understanding the environment of the device (e.g., environment 600) to facilitate or even enable further processing. For example, a dynamic occupancy grid can be useful for predicting occupancy, motion planning, etc. A dynamic occupancy grid may include one or more cells of static occupancy type and / or one or more cells of dynamic occupancy type at any given time. Dynamic objects can be represented as a collection of one or more velocity vectors. For example, an occupancy grid cell may make some or all of the occupancy probabilities dynamic, and within the dynamic occupancy probabilities, there may be multiple (e.g., four) velocity vectors, each with a corresponding probability, which together sum to the dynamic occupancy probability of that cell 810. The dynamic occupancy grid can be obtained, for example, by occupancy grid cell 560 by processing information from multiple sensors (such as those from a radar system, e.g., in sensor 540). Adding data from one or more cameras to determine the dynamic occupancy grid can provide significant improvements to the grid, such as the accuracy of probabilities and / or velocities in the grid cells.
[0055] refer to Figure 9A block diagram of an example system 900 for robust object detection via different signal representations is shown. System 900 includes multiple sensors, such as a radar module 902 and a camera 904. In one example, system 900 may be implemented in device 500, where radar module 902 may be one or more radar sensors 542, and camera 904 may be one or more cameras 544. System 900 also includes functional blocks including LLP functional block 906 (Low-Level Perception Functional Block), Dynamic Occupied Grid (DoG) functional block 908, Robust Fusion Functional Block 910, and Environment Model 912. LLP functional block 906 is a first signal path configured to receive input data from radar module 902 and camera 904, and generate parametric representations of objects in the received radar and image data. In one example, the parametric representation may be a structured list of elements associated with objects detected in the radar and camera data. The list of elements may include coordinate information of the detected object (e.g., grid coordinates x, y, z), object size information (e.g., length, width, height), object pose information, and object classification information. DoG function block 908 is the second signal path, configured to receive input data from radar module 902 and camera 904, and to generate nonparametric representations of objects in the received radar and image data. The nonparametric representations can be occupancy maps, such as those relating to... Figure 8 As described. For example, a nonparametric representation could be a floating-point number for each cell in a 512x512 grid, indicating the probability that the cell is occupied or free, and the probability that the object occupying the cell is static or dynamic. Other information, such as the motion of dynamic cells (e.g., velocities in the x, y, and z directions) and categorical information, could be included in the nonparametric representation.
[0056] Robust fusion function block 910 is configured to fuse parametric and nonparametric representations into one or more object lists for environment model 912. In one example, the fusion process in robust fusion function block 910 may utilize object coordinate information from the parametric representation received from LLP function block 906 and cell positions from the nonparametric representation received from DoG function block 908. Robust fusion function block 910 may be configured to identify clusters (e.g., clusters of dynamic grid cells) within the nonparametric representation that have similar properties (e.g., similar object classifications and / or similar velocities), and to identify (e.g., from LLP function block 906) indications of the identified objects in the parametric representation to track the objects, for example, using a Kalman filter (and / or one or more other algorithms). Robust fusion function block 910 may be configured to output a list of object trajectories indicating the tracked objects to environment model 912. The object trajectory list may include the position, velocity, length, and width (and possibly other information) of each object in the object trajectory list. The object trajectory list may include a shape for representing each object, such as a closed polygon or other shape (e.g., an ellipse (e.g., indicated by the values of the major and minor axes)). The robust fusion function block 910 can be configured to determine static objects (e.g., road boundaries, traffic signs, etc.) based on parametric and / or nonparametric representations and provide static object information to the environment model 912.
[0057] refer to Figure 10This paper illustrates a first example functional architecture 1000 for implementing signal redundancy in object detection. System 900 can be configured to utilize one or more features in architecture 1000. In one example, architecture 1000 can be implemented in one or more software modules 1006 configured to receive signals from radar module 902 and camera 904. Software module 1006 may include a first signal path 1008 configured to generate parametric representations and a second signal path 1010 configured to generate nonparametric representations. The first signal path 1008 and the second signal path 1010 may each include one or more modules comprising machine learning models configured to output parametric and nonparametric representations, respectively. The machine learning model may be based on deep learning techniques. For example, the first signal path 1008 may include a camera deep learning (DL) detection module 1012, a low-level (LL) fusion object module 1014, and a radar detection module 1016. One or more of modules 1012, 1014, and 1016 may be trained to output parametric representations based on inputs received from radar module 902 and / or camera 904. The second signal path 1010 may include a camera drivable space module 1018, a camera-based semantic segmentation (Camer SemSeg) module 1020, a radar point cloud module 1022, and a low-level bird's-eye view (BEV) segmentation and occupancy flow module 1024. One or more of modules 1018, 1020, 1022, and 1024 may be trained to output nonparametric representations based on inputs received from radar module 902 and / or camera 904. The number of modules and the type of machine learning model shown are illustrative and not limiting, as other modules and machine learning techniques may also be used in the signal path to generate corresponding parametric and nonparametric representations.
[0058] Architecture 1000 may include a fusion functional block 1026, which includes an object tracking module 1030 and an occupancy grid module 1032. The object tracking module 1030 may be configured to track objects using parametric representations received via a first signal path 1008 and nonparametric representations received from a second signal path 1010, for example, using a Kalman filter (and / or one or more other algorithms), and to output a list of object trajectories indicating the tracked objects to an environment model. The object trajectory list may include the position, velocity, length, and width (and possibly other information) of each object in the object trajectory list. The object trajectory list may include a shape for representing each object, such as a closed polygon or other shape (e.g., an ellipse (e.g., indicated by the values of the major and minor axes)). The occupancy grid module 1032 may be configured to determine static objects (e.g., road boundaries, traffic signs, etc.) in the parametric and / or nonparametric representations provided by the respective first signal path 1008 and second signal path 1010. Static object information may be provided to the environment model.
[0059] refer to Figure 11 And further reference Figure 10 A second example functional architecture 1100 for implementing signal redundancy in object detection is illustrated. The second example architecture 1100 includes features of the first example architecture 1000 and adds additional sensors and signal paths. In one example, a lidar module 1102 may be configured to provide signals (e.g., target information) to a secondary path 1104 separate from the first signal path 1008 and the second signal path 1010. The secondary path 1104 may include one or more additional modules with a machine learning model configured to provide object detection information to an environment model based on target information received from the lidar module 1102. In one example, the secondary path 1104 may be configured to receive signals from a radar module 902 and / or a camera 904 to generate object information for the environment model. Architecture 1100 further improves the robustness of the object detection functionality by adding additional sensors and signal paths.
[0060] refer to Figure 12A and Figure 12B This illustrates an example deep learning (DL) architecture for object detection. In general, Figure 12A and Figure 12B The deep learning architecture includes one or more backbone network models, neck models, and head models. The backbone network model can be configured to extract and encode features from input signals (e.g., from radar module 902 and camera 904). The neck model can be configured to transform and refine the features extracted by the backbone network model. The head model can be configured for a task-specific layer designed to produce a final prediction or inference based on information extracted by the backbone network model and the neck model. The head model may correspond to one or more modules 1012, 1014, 1016, 1018, 1020, 1022, and 1024 in the corresponding first signal path 1008 and / or second signal path 1010.
[0061] In the first DL architecture 1200, a common backbone network 1202 can be configured to receive signals from radar module 902 and camera 904. A first set 1204 of neck models and a first set 1206 of head models can be configured to generate parametric and / or non-parametric representations associated with one or more modules 1012, 1014, 1016, 1018, 1020, 1022, and 1024. In one example, separate training for each head model can be enforced to handle redundancy. In the second DL architecture 1250, the backbone network model, neck model, and head model can be separated based on a first signal path and a second signal path. For example, the first backbone network 1252 can be configured to receive signal input from radar module 902 and camera 904, and the second set 1254 of neck models and the second set 1256 of head models can be trained to generate parametric representations. The second backbone network 1258 can also be configured to receive signal input from the radar module 902 and the camera 904, and the third set 1260 of the neck model and the third set 1262 of the head model can be trained to generate nonparametric representations. Other deep learning architectures can also be used to generate redundancy in the signal stream and improve the robustness of object detection capabilities.
[0062] refer to Figure 13 And further reference Figures 1 to 12B Method 1300 for generating an object representation with multiple signal paths includes the stages shown. However, method 1300 is illustrative and not limiting. Method 1300 can be modified, for example, by adding, removing, rearranging, combining, performing one or more stages concurrently, and / or splitting one or more stages into multiple stages.
[0063] At stage 1302, method 1300 includes acquiring image information using at least one camera module mounted on the vehicle. Device 500, including processor 510 and sensor 540, is the component for acquiring the image information. In one example, camera 544 may acquire images of the environment adjacent to the vehicle. The images may include static objects such as road signs, trees, barriers, and other non-moving objects, as well as dynamic objects such as other vehicles, bicycles, pedestrians, and other moving objects. In one example, image information may be acquired at a frame rate of approximately 40 ms. Other frame rates may be used.
[0064] At stage 1304, method 1300 includes obtaining target information from at least one radar module mounted on a vehicle. Device 500, including processor 510 and sensor 540, is the component for obtaining the target information. In one example, one or more radar sensors 542 may be configured to provide range, azimuth, and velocity information of an object that generates a returned radar signal (e.g., a radar echo). In one example, the target information may be a range map based on the echo signal. Radar target information can be obtained at a frame rate of approximately 40 ms. Other frame rates may be used.
[0065] At stage 1306, method 1300 includes generating a first detection representation having a first signal path based on image information and target information. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for generating the first detection representation. In one example, the first detection representation includes one or more parameter representations generated via the first signal path 1008. The first signal path 1008 may include one or more machine learning models, such as those described in modules 1012, 1014, and 1016, to generate the parameter representations. These modules and corresponding parameter representations are examples and not limitations, as the first signal path 1008 may utilize other modules to generate other detection representations.
[0066] At stage 1308, method 1300 includes generating a second detection representation with a second signal path based on image information and target information, wherein the second signal path is different from the first signal path. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for generating the second detection representation. In one example, the second detection representation includes one or more nonparametric representations generated via the second signal path 1010. The second signal path 1010 may include one or more machine learning models, such as those described in modules 1018, 1020, 1022, and 1024, to generate the nonparametric representations. These modules and corresponding nonparametric representations are examples and not limitations, as the second signal path 1010 may utilize other modules to generate other detection representations.
[0067] At stage 1310, method 1300 includes outputting a first detection representation and a second detection representation. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for outputting the detection representations. In one example, the first and second detection representations may be output to a fusion module configured to generate an object list based on the first and second representations. Other modules in the autonomous vehicle perception architecture may be configured to receive the first and second detection representations (e.g., before fusion).
[0068] refer to Figure 14And further reference Figures 1 to 13 Method 1400 for generating one or more lists of objects based on multiple signal paths includes the stages shown. However, method 1400 is illustrative and not restrictive. Method 1400 can be modified, for example, by adding, removing, rearranging, combining, performing one or more stages concurrently, and / or splitting one or more stages into multiple stages.
[0069] At stage 1402, method 1400 includes receiving a first detection representation via a first signal path and receiving a second detection representation via a second signal path. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for receiving the first and second detection representations. In one example, fusion function block 1026, including object tracking module 1030 and occupancy grid module 1032, is configured to receive parametric and nonparametric representations as corresponding first and second detection representations.
[0070] At stage 1404, method 1400 includes generating one or more object lists based at least in part on a first detection representation and a second detection representation. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for generating one or more object lists. Object tracking module 1030 is configured to track objects using parametric representations received via a first signal path 1008 and nonparametric representations received from a second signal path 1010, for example, using a Kalman filter (and / or one or more other algorithms), and to generate a list of object trajectories indicating the tracked objects. The object trajectory list may include the position, velocity, length, and width (and possibly other information) of each object in the object trajectory list. The object trajectory list may include a shape for representing each object, such as a closed polygon or other shape (e.g., an ellipse (e.g., indicated by the values of the major and minor axes)). Occupied grid module 1032 is configured to determine static objects (e.g., road boundaries, traffic signs, etc.) in the parametric and / or nonparametric representations provided by the respective first signal path 1008 and second signal path 1010, and to generate static object information.
[0071] At stage 1406, method 1400 includes outputting one or more lists of objects. Device 500, including processor 510, sensor 540, and architecture 1000, is a component for outputting one or more lists of objects. In one example, fusion function block 1026 may be configured to output a list of object trajectories and static object information to an environment model. Other modules in the autonomous vehicle perception architecture may be configured to receive one or more lists of objects (e.g., fusion of parametric and nonparametric representations based on sensor information).
[0072] Other examples and specific implementations are within the scope of this disclosure and the appended claims. For example, due to the nature of software and computers, the functions described above can be implemented using software, hardware, firmware, hardwiring, or any combination thereof executed by a processor. Features implementing the functions can also be physically located in various locations, including portions distributed such that the functions are implemented in different physical locations.
[0073] As used herein, the singular forms “a,” “an,” and “the” also include the plural forms, unless the context clearly indicates otherwise. Thus, references to a device in the singular form included in the claims (e.g., “device,” “the / said device”) include at least one of such devices (i.e., one or more) (e.g., “processor” includes at least one processor (e.g., one processor, two processors, etc.), “the / said processor” includes at least one processor, “memory” includes at least one memory, “the / said memory” includes at least one memory, etc.). The phrases “at least one” and “one or more” are used interchangeably, and such that the object referred to by “at least one” and the object referred to by “one or more” include embodiments having one referred object and embodiments having multiple referred objects. For example, “at least one processor” and “one or more processors” each include embodiments having one processor and embodiments having multiple processors. Furthermore, as used herein, “set” includes one or more members, and a “subset” contains all members of a set less than the set referred to by the subset.
[0074] As used herein, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0075] Furthermore, as used herein, an item enumeration followed by "at least one of" or "one or more of" indicates a disjunctive enumeration, such that an enumeration of, for example, "at least one of A, B, or C," or "at least one of A, B, and C," or "one or more of A, B, or C," or "one or more of A, B, and C," or "A or B or C" represents A or B or C, or AB (A and B), or AC (A and C), or BC (B and C), or ABC (i.e., A and B and C), or a combination having more than one feature (e.g., AA, AAB, ABBC, etc.). Therefore, a statement that an item (e.g., a processor) is configured to perform a function relating to at least one of A or B, or a statement that an item is configured to perform function A or function B, indicates that the item can be configured to perform a function relating to A, or can be configured to perform a function relating to B, or can be configured to perform a function relating to A and B. For example, the phrase "a processor configured to measure at least one of A or B" or "a processor configured to measure A or measure B" means that the processor can be configured to measure A (and may or may not be configured to measure B), or can be configured to measure B (and may or may not be configured to measure A), or can be configured to measure both A and B (and can be configured to select which of A and B or measure both). Similarly, a description of a component for measuring at least one of A or B includes: a component for measuring A (which may or may not be able to measure B), or a component for measuring B (which may or may not be configured to measure A), or a component for measuring A and B (which may be able to select which of A and B or measure both). As another example, a description of an item (e.g., a processor) being configured to perform at least one of function X or function Y means that the item can be configured to perform function X, or can be configured to perform function Y, or can be configured to perform both functions X and Y. For example, the phrase "processor configured to measure at least one of X or Y" means that the processor can be configured to measure X (and may or may not be configured to measure Y), or can be configured to measure Y (and may or may not be configured to measure X), or can be configured to measure both X and Y (and can be configured to select which of X and Y or measure both).
[0076] As used herein, unless otherwise stated, a description of a function or operation as “based on” an item or condition means that the function or operation is based on the described item or condition and may be based on one or more items and / or conditions other than the described item or condition.
[0077] Substantial changes can be made depending on specific requirements. For example, custom hardware may be used, and / or specific elements may be implemented in the hardware, in software executed by the processor (including portable software such as applets), or both. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Unless otherwise specified, components shown in the figures and / or discussed herein that are connected or communicate with each other (functionally or otherwise) are communicatively coupled. That is, these components may be connected directly or indirectly to enable communication between them.
[0078] The systems and devices discussed above are examples. Various configurations may appropriately omit, substitute, or add various processes or components. For example, features described with respect to certain configurations may be combined in various other configurations. Different aspects and elements of configurations may be combined in a similar manner. Furthermore, technology is constantly evolving, and therefore many elements are examples and do not limit the scope of this disclosure or the claims.
[0079] Specific details are provided in this description to offer a thorough understanding of the example configurations, including specific implementations. However, the configurations can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail to avoid obscuring these configurations. The description herein provides example configurations and does not limit the scope, applicability, or configuration of the claims. Rather, the preceding description of the configurations provides a description for implementing the described techniques. Various changes can be made to the function and arrangement of the elements.
[0080] As used herein, the terms “processor-readable medium,” “machine-readable medium,” and “computer-readable medium” refer to any medium that participates in providing data that enables a machine to operate in a particular manner. Using a computing platform, various processor-readable media may involve providing instructions / code to a processor for execution, and / or may be used to store and / or carry such instructions / code (e.g., as signals). In many specific implementations, processor-readable media are physical and / or tangible storage media. Such media can take many forms, including but not limited to non-volatile and volatile media. Non-volatile media include, for example, optical discs and / or magnetic disks. Volatile media include, but are not limited to, dynamic memory.
[0081] Having described several example configurations, various modifications, alternative constructions, and equivalents can be used. For example, the above elements can be components of a larger system, where other rules may take precedence over or otherwise modify the application of this disclosure. Furthermore, several operations may be performed before, during, or after considering the above elements. Accordingly, the above description does not limit the scope of the claims.
[0082] Unless otherwise indicated, the terms "about" and / or "approximately" as used herein when referring to measurable values (such as quantities, durations of time, etc.) cover variations of ±20%, ±10%, ±5%, or ±0.1% from the specified value, as appropriate in the context of the systems, devices, circuits, methods, and other specific embodiments described herein. Similarly, unless otherwise indicated, the term "substantially" as used herein when referring to measurable values (such as quantities, durations of time, physical properties (such as frequencies), etc.) also covers variations of ±20%, ±10%, ±5%, or ±0.1% from the specified value, as appropriate in the context of the systems, devices, circuits, methods, and other specific embodiments described herein.
[0083] A statement that a value exceeds (or is greater than or higher than) a first threshold is equivalent to a statement that a value meets or exceeds a second threshold slightly greater than the first threshold. For example, in the resolution of the computing system, the second threshold is one value higher than the first threshold. A statement that a value is less than the first threshold (or within or below the first threshold) is equivalent to a statement that a value is less than or equal to a second threshold slightly lower than the first threshold. For example, in the resolution of the computing system, the second threshold is one value lower than the first threshold.
[0084] Specific implementation examples are described in the following numbered clauses:
[0085] Clause 1. A method for generating an object representation having multiple signal paths, the method comprising: obtaining image information from at least one camera module disposed on a vehicle; obtaining target information from at least one radar module disposed on the vehicle; generating a first detection representation having a first signal path based on the image information and the target information; generating a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and outputting the first detection representation and the second detection representation.
[0086] Clause 2. The method according to Clause 1, wherein the first detection representation includes a parametric representation of the target object, and the second detection representation includes a nonparametric representation of the target object.
[0087] Clause 3. The method according to Clause 2, wherein the parameters of the target object represent the coordinate information of the target object and the size information of the target object.
[0088] Clause 4. The method according to Clause 2, wherein the nonparametric representation of the target object is an occupancy graph.
[0089] Clause 5. The method according to Clause 2, wherein the first signal path includes at least a first machine learning model configured to generate the parametric representation based at least in part on the image information and the target information, and the second signal path includes at least a second machine learning model configured to generate the nonparametric representation based at least in part on the image information and the target information.
[0090] Clause 6. The method according to Clause 5, wherein the first machine learning model and the second machine learning model utilize a public backbone network.
[0091] Clause 7. The method according to Clause 5, wherein the first machine learning model utilizes at least a first backbone network, and the second machine learning model utilizes at least a second backbone network.
[0092] Clause 8. The method according to Clause 1, the method further comprising: receiving the first detection representation via the first signal path and receiving the second detection representation via the second signal path; generating one or more object lists based at least in part on the first detection representation and the second detection representation; and outputting the one or more object lists.
[0093] Clause 9. The method described in Clause 8, wherein the list of one or more objects includes a list of object trajectories indicating the position and velocity of the objects.
[0094] Clause 10. The method according to Clause 9, wherein the list of object trajectories indicates the shape of the object.
[0095] Clause 11. The method described in Clause 9, wherein the list of one or more objects includes static object information.
[0096] Clause 12. The method according to Clause 8, wherein outputting the list of one or more objects includes providing the list of one or more objects to the environment model.
[0097] Clause 13. The method according to Clause 8, the method further comprising: receiving target information from a lidar module disposed on the vehicle via a secondary path separate from the first signal path and the second signal path; generating object detection information based on the target information; and outputting the object detection information.
[0098] Clause 14. The method according to Clause 13, the method further comprising: receiving image information from the at least one camera module disposed on the vehicle; generating the object detection information based on the target information and the image information; and outputting the object detection information.
[0099] Clause 15. An apparatus comprising: at least one memory; at least one camera module; at least one radar module; at least one processor communicatively coupled to the at least one memory, the at least one camera, and the at least one radar module, and configured to: acquire image information from the at least one camera module disposed on a vehicle; acquire target information from the at least one radar module disposed on the vehicle; generate a first detection representation having a first signal path based on the image information and the target information; generate a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and output the first detection representation and the second detection representation.
[0100] Clause 16. The apparatus according to Clause 15, wherein the first detection representation includes a parametric representation of the target object, and the second detection representation includes a nonparametric representation of the target object.
[0101] Clause 17. The apparatus according to Clause 16, wherein the parameters of the target object represent the coordinate information of the target object and the size information of the target object.
[0102] Clause 18. The apparatus according to Clause 16, wherein the nonparametric representation of the target object is an occupancy map.
[0103] Clause 19. The apparatus of Clause 16, wherein the first signal path includes at least a first machine learning model and the at least one processor is further configured to generate the parametric representation based at least in part on the image information and the target information, and the second signal path includes at least a second machine learning model and the at least one processor is further configured to generate the nonparametric representation based at least in part on the image information and the target information.
[0104] Clause 20. The apparatus of Clause 19, wherein the first machine learning model and the second machine learning model utilize a public backbone network.
[0105] Clause 21. The apparatus of Clause 19, wherein the first machine learning model utilizes at least a first backbone network and the second machine learning model utilizes at least a second backbone network.
[0106] Clause 22. The apparatus according to Clause 15, wherein the at least one processor is further configured to: receive the first detection representation via the first signal path and receive the second detection representation via the second signal path; generate one or more object lists based at least in part on the first detection representation and the second detection representation; and output the one or more object lists.
[0107] Clause 23. The apparatus according to Clause 22, wherein the list of one or more objects includes a list of object trajectories indicating the position and velocity of the objects.
[0108] Clause 24. The apparatus according to Clause 23, wherein the object trajectory list indicates the shape of the object.
[0109] Clause 25. The apparatus according to Clause 23, wherein the list of one or more objects includes static object information.
[0110] Clause 26. The apparatus according to Clause 22, wherein the at least one processor is further configured to output the list of one or more objects to an environment model.
[0111] Clause 27. The apparatus according to Clause 22, the apparatus further comprising at least one lidar module disposed on the vehicle, wherein the at least one processor is further configured to: receive additional target information from the at least one lidar module via a secondary path separate from the first signal path and the second signal path; generate object detection information based on the additional target information; and output the object detection information.
[0112] Clause 28. The apparatus according to Clause 27, wherein the at least one processor is further configured to: generate the object detection information based on the target information and the image information; and output the object detection information.
[0113] Clause 29. An apparatus for generating an object representation having multiple signal paths, the apparatus comprising: components for acquiring image information from at least one camera module disposed on a vehicle; components for acquiring target information from at least one radar module disposed on the vehicle; components for generating a first detection representation having a first signal path based on the image information and the target information; components for generating a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and components for outputting the first detection representation and the second detection representation.
[0114] Clause 30. The apparatus according to Clause 29, further comprising: means for receiving the first detection representation via the first signal path and receiving the second detection representation via the second signal path; means for generating one or more object lists based at least in part on the first detection representation and the second detection representation; and means for outputting the one or more object lists.
[0115] Clause 31. The apparatus according to Clause 30, further comprising: means for receiving target information from a lidar module disposed on the vehicle via a secondary path separate from the first signal path and the second signal path; means for generating object detection information based on the target information; and means for outputting the object detection information.
[0116] Clause 32. A non-transitory processor-readable storage medium comprising processor-readable instructions configured to cause one or more processors to generate object representations having multiple signal paths, the processor-readable instructions comprising code for: acquiring image information from at least one camera module disposed on a vehicle; acquiring target information from at least one radar module disposed on the vehicle; generating a first detection representation having a first signal path based on the image information and the target information; generating a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and outputting the first detection representation and the second detection representation.
[0117] Clause 33. The non-transitory processor-readable storage medium according to Clause 32, the non-transitory processor-readable storage medium further comprising: code for receiving the first detection representation via the first signal path and receiving the second detection representation via the second signal path; code for generating one or more object lists based at least in part on the first detection representation and the second detection representation; and code for outputting the one or more object lists.
[0118] Clause 34. The non-transitory processor-readable storage medium according to Clause 33, the non-transitory processor-readable storage medium further comprising: code for receiving target information from a lidar module disposed on the vehicle via a secondary path separate from the first signal path and the second signal path; code for generating object detection information based on the target information; and code for outputting the object detection information.
[0119] Clause 35. A method for generating an object representation having multiple signal paths, the method comprising: receiving a first detection representation via a first signal path and receiving a second detection representation via a second signal path; generating one or more object lists based at least in part on the first detection representation and the second detection representation; and outputting the one or more object lists.
[0120] Clause 36. An apparatus comprising: at least one memory; at least one camera module; at least one radar module; at least one processor communicatively coupled to the at least one memory, the at least one camera, and the at least one radar module, and configured to: receive a first detection representation via a first signal path and receive a second detection representation via a second signal path; generate one or more object lists based at least in part on the first detection representation and the second detection representation; and output the one or more object lists.
[0121] Clause 37. An apparatus comprising: means for receiving a first detection representation via a first signal path and a second detection representation via a second signal path; means for generating one or more object lists based at least in part on the first detection representation and the second detection representation; and means for outputting the one or more object lists.
[0122] Clause 38. A non-transitory processor-readable storage medium comprising processor-readable instructions configured to cause one or more processors to generate object representations having a plurality of signal paths, the processor-readable instructions comprising code for: receiving a first detection representation via a first signal path and receiving a second detection representation via a second signal path; generating one or more object lists based at least in part on the first detection representation and the second detection representation; and outputting the one or more object lists.
Claims
1. An apparatus, the apparatus comprising: At least one memory; At least one camera module; At least one radar module; At least one processor, communicatively coupled to the at least one memory, the at least one camera module, and the at least one radar module, and configured to: Image information is obtained from the at least one camera module installed on the vehicle; Target information is obtained from the at least one radar module installed on the vehicle; A first detection representation with a first signal path is generated based on the image information and the target information; A second detection representation with a second signal path is generated based on the image information and the target information, wherein the second signal path is different from the first signal path; as well as Output the first detection representation and the second detection representation.
2. The apparatus of claim 1, wherein the first detection representation includes a parametric representation of the target object, and the second detection representation includes a non-parametric representation of the target object.
3. The apparatus according to claim 2, wherein the parameter representation of the target object includes the coordinate information of the target object and the size information of the target object.
4. The apparatus of claim 2, wherein the nonparametric representation of the target object is an occupancy map.
5. The apparatus of claim 2, wherein the first signal path includes at least a first machine learning model and the at least one processor is further configured to generate the parametric representation based at least in part on the image information and the target information, and the second signal path includes at least a second machine learning model and the at least one processor is further configured to generate the nonparametric representation based at least in part on the image information and the target information.
6. The apparatus of claim 5, wherein the first machine learning model and the second machine learning model utilize a common backbone network.
7. The apparatus of claim 5, wherein the first machine learning model utilizes at least a first backbone network, and the second machine learning model utilizes at least a second backbone network.
8. The apparatus of claim 1, wherein the at least one processor is further configured to: The first detection representation is received via the first signal path and the second detection representation is received via the second signal path; Generate one or more object lists based at least in part on the first detection representation and the second detection representation; and Output a list of the one or more objects.
9. The apparatus of claim 8, wherein the list of one or more objects includes a list of object trajectories indicating the position and velocity of the objects.
10. The apparatus of claim 9, wherein the object trajectory list indicates the shape of the object.
11. The apparatus of claim 9, wherein the list of one or more objects includes static object information.
12. The apparatus of claim 8, wherein the at least one processor is further configured to output the one or more object lists to an environment model.
13. The apparatus of claim 8, further comprising at least one lidar module disposed on the vehicle, wherein the at least one processor is further configured to: Additional target information is received from the at least one lidar module via a secondary path separate from the first signal path and the second signal path; Based on the additional target information, generate object detection information; and Output the object detection information.
14. The apparatus of claim 13, wherein the at least one processor is further configured to: The object detection information is generated based on the target information and the image information; and Output the object detection information.
15. A method for generating an object representation having multiple signal paths, the method comprising: Image information is obtained from at least one camera module installed on the vehicle; Target information is obtained from at least one radar module installed on the vehicle. A first detection representation with a first signal path is generated based on the image information and the target information; A second detection representation with a second signal path is generated based on the image information and the target information, wherein the second signal path is different from the first signal path; as well as Output the first detection representation and the second detection representation.
16. The method of claim 15, wherein the first detection representation includes a parametric representation of the target object, and the second detection representation includes a nonparametric representation of the target object.
17. The method of claim 16, wherein the parameter representation of the target object includes the coordinate information of the target object and the size information of the target object.
18. The method of claim 16, wherein the nonparametric representation of the target object is an occupancy map.
19. The method of claim 16, wherein the first signal path includes at least a first machine learning model configured to generate the parametric representation based at least in part on the image information and the target information, and the second signal path includes at least a second machine learning model configured to generate the nonparametric representation based at least in part on the image information and the target information.
20. An apparatus for generating an object representation having multiple signal paths, the apparatus comprising: A component for obtaining image information from at least one camera module installed on a vehicle; Components for obtaining target information from at least one radar module mounted on the vehicle; A component for generating a first detection representation having a first signal path based on the image information and the target information; A component for generating a second detection representation having a second signal path based on the image information and the target information, wherein the second signal path is different from the first signal path; and A component for outputting the first detection representation and the second detection representation.