Method for low-level fusion of camera and ultrasound data
By aligning and fusing camera and ultrasound data into a unified format, the method addresses the inefficiencies of separate processing, resulting in improved accuracy and comprehensive environmental perception for semi-autonomous vehicles.
Patent Information
- Application Number
- PCT/EP2024/086489
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-24
AI Technical Summary
Existing sensor fusion methods for environmental detection in vehicles are computationally intense and inaccurate due to the need for separate processing of differently formatted data from various sensor types, such as cameras and ultrasound, lacking a unified data processing approach.
Transform data from different sensor types, like cameras and ultrasound, into a unified format early in the processing pipeline, allowing for efficient low-level fusion by aligning and combining pixel-based two-dimensional images, thereby facilitating seamless integration and processing.
Enhances the accuracy and comprehensiveness of environmental perception by integrating data from multiple sensors, improving the safety and efficiency of semi-autonomous vehicle operations, particularly in parking scenarios.
Smart Images

Figure EP2024086489_24072025_PF_FP_ABST
Abstract
Description
[0001] Method for low-level fusion of camera and ultrasound data
[0002] The invention is directed at a method, in particular at a computer-implemented method, for fusing camera data and ultrasound data to generate input data for an electronic vehicle guidance system of a motor vehicle, wherein the camera data and the ultrasound data describe a surrounding environment of the motor vehicle. Further aspects of the invention are directed at a data processing device, which is designed to perform the inventive method, at a computer program comprising instructions which, when the program is executed by such a data processing device, cause it to carry out the steps of the inventive method and at a computer-readable storage medium, having stored thereon such a computer program. Further aspects of the invention are directed at an electronic vehicle guidance system comprising such a data processing device, at a motor vehicle with such an electronic vehicle guidance system and at a method for guiding a motor vehicle at least semi-automatically based on input data, which is generated according to the inventive method.
[0003] Environmental detection plays a key role in assisted driving or (semi-)automated driving. A large number of driver assistance systems for motor vehicles are operated based on the detection or recognition of a surrounding environment of the motor vehicle. The environmental detection can comprise object detection, object classification and / or segmentation, i.e. the assignment of objects to predefined object classes.
[0004] To generate input data for the environmental detection, typically different sensor systems of the motor vehicle are used to collect data, which describe the surrounding environment and / or objects in the surrounding environment. For example in parking assistance systems, ultrasound sensor systems can be used in order to determine a distance between an ultrasound sensor of the motor vehicle, which is typically located in or at a bumper of the motor vehicle, and an object in the surrounding environment. Based on the distance, different assistance measures can be initiated, i.e. the triggering of an acoustic warning signal to a driver of the motor vehicle when the distance to the object decreases below a critical distance.
[0005] In order to enhance the quality of the input data, different sensor types of the vehicle can be used alternately or simultaneously to capture characteristics or objects in the surrounding environment of the motor vehicle. For example, CN 110488319 A describes a collision distance calculation method based on the combined use of ultrasound and camera data. Here, based on ultrasound and camera data, a barrier is detected. In conjunction with further sensor data, such as vehicle speed and steering wheel angle, as well as together with an ultrasonic and camera delay time information, the ultrasound and camera data are synchronized in order to locate the barrier based on the time- synchronized data. In other words, the results of different measurements are combined in order to enhance environmental detection.
[0006] Another approach for enhancing environmental detection during a driving maneuver of a motor vehicle using several sensors of different sensor types of the vehicle may comprise the detection of an object in the surrounding environment by a first sensor type, for example a camera sensor, and the detection of the object or a further object in the surrounding environment by a second sensor type, for example an ultrasound sensor. The object information from these two sensor types, for example a class of the object, which can be derived from the camera data, and a distance to the object, which can be derived from the ultrasound data, can be combined in the analysis of the current driving maneuver.
[0007] The known combination of object information coming from different sensor types can be referred to as high-level fusion. In the context of this disclosure, high-level fusion means the fusion of already pre-processed object information. In other words, sensor data from different types of sensors can be pre-processed in order to provide a first analysis of the objects that are described by the sensor data. For example, a pre-processing step can comprise running an object detection algorithm, a segmentation algorithm and / or an object classification algorithm and the like on raw data, like for example camera data. Comparable pre-processing steps to enhance object perception can also be run on ultrasound data, lidar data and / or radar data. Thus pre-processed sensor data can be combined or high-level fused in order to obtain environmental perception of the surrounding environment of a motor vehicle.
[0008] Adversely, according to the known approaches, the data obtained by the different sensor types cannot be processed together on a lower data processing level, as they are of different formats. Like this, additional data processing steps for each type of data are necessary, which render the known methods computationally intense, elaborate and inaccurate. It is therefore an object of the present invention, to provide a more accurate and preferably also more comprehensive representation of the surrounding environment of a motor vehicle.
[0009] The object is solved by the subject-matter of the independent claims. Further embodiments of the invention are described in the dependent claims, the description and the figures.
[0010] The invention is based on the realization that data of different format coming from different sensor types should be transformed to the same format early on during the overall process of environment perception. Like this, the transformed data can be treated as if they virtually come from the same sensor type. Advantageously, data in the same format can be easily processed together, which renders separate processing steps for each type of data obsolete or unnecessary.
[0011] A first aspect of the invention is directed at a method for fusing camera data and ultrasound data to generate input data for an electronic vehicle guidance system of a motor vehicle. Therein, the camera data and the ultrasound data describe a surrounding environment of the motor vehicle, for example a parking lot, where the electronic vehicle guidance system intends to park the vehicle.
[0012] To collect the data, the motor vehicle itself may comprise a camera sensor unit and an ultrasound sensor unit. In other words, the motor vehicle is preferably equipped with an environmental sensor system. In the context of this disclosure, an environmental sensor system can be understood as a sensor system, which is able to generate sensor data or sensor signals, which depict, represent or image an environment of the environmental sensor system. In particular, the ability to capture or detect electromagnetic or other signals from the environment, cannot be considered a sufficient condition for qualifying a sensor system as an environmental sensor system. For example, cameras, lidar systems, radar systems or ultrasonic sensor systems may be considered as environmental sensor systems. As such, an environmental sensor system can comprise a camera sensor unit and an ultrasonic or ultrasound sensor unit.
[0013] The environmental sensor system can be integrated into a (semi-)autonomous driving system, such as an electronic vehicle guidance system of the motor vehicle.
[0014] The inventive method comprises the steps of - receiving the camera data from the camera sensor unit of the motor vehicle,
[0015] - receiving the ultrasound data from the ultrasound sensor unit of the motor vehicle,
[0016] - generating a first pixel-based two-dimensional image with a first pixel resolution of the surrounding environment based on the camera data,
[0017] - generating a second pixel-based two-dimensional image with a second pixel resolution of the surrounding environment based on the ultrasound data,
[0018] - aligning the first and second pixel resolutions to generate aligned two-dimensional images of the surrounding environment,
[0019] - fusing the aligned two-dimensional images of the surrounding environment into one augmented image of the surrounding environment, and
[0020] - providing the augmented image of the surrounding environment as input data for the electronic vehicle guidance system.
[0021] The received camera data are preferably transmitted to an image processing unit, where the first pixel-based two-dimensional image can be generated. In the image processing unit, one or more pre-processing steps can be carried out, for example in order to gain a first perception of the surrounding environment of the motor vehicle from the camera data.
[0022] The received ultrasound data are preferably transmitted to an ultrasound processing unit, which can comprise a beamforming control unit, where the second pixel-based two- dimensional image can be generated. Preferably, generating the second pixel-based two- dimensional image comprises running a beamforming algorithm on the ultrasound data. For example, the beamforming algorithm can process the received ultrasound signals using a delay-and-sum technique with apodization and dynamic focusing to create the two-dimensional ultrasound image with a high resolution and quality. Optionally, feature extraction or any other type of image perception can be applied to the second pixel-based two-dimensional image before further processing.
[0023] In a next step, the pixel resolutions of the first and second image are aligned. This can be done by applying some sort of spatial interpolation. For example, the ultrasound measurement after beamforming may result in a two-dimensional image having a pixel resolution of 100x100 pixels. On the other hand, the processing of the camera data may result in a two-dimensional image having a pixel resolution of 1000x1000 pixels.
[0024] Alignment can be done to downscale or downsample from 1000x1000 to 100x100 (each 10 pixels will be corresponding to 1). This alignment process results in two two- dimensional images with the same pixel resolution. These images can then be mapped to each other 1 :1 . The aligned images will then be fused to augment each other. In the context of the present disclosure, this fusion is referred to as low-level fusion or physical fusion of the two-dimensional images, in contrast to the high-level fusion as outlined above.
[0025] The augmented image combines information from the camera sensor unit and from the ultrasound sensor unit. Ultrasound sensors have advantages over optical sensors such as cameras or radars, such as being able to penetrate through air and reflect off solid surfaces, being immune to interference from other sources, and being able to be steered and focused using beamforming techniques. Beamforming techniques can improve the resolution and quality of the ultrasound image by delaying and summing the signals from multiple transducers to form a directional beam, applying apodization and dynamic focusing to reduce sidelobes and enhance lateral resolution, and using frequency and bandwidth modulation to increase contrast and depth resolution.
[0026] Especially in parking situations, the cameras can have a blind spot, while the area covered by the blind spot may be accessible by the ultrasound sensors. Likewise, the ultrasound sensors may not be able to gather information on objects, which are located far from the motor vehicle, for example one hundred or more meters from the motor vehicle, while it may be possible to gather these information by the camera sensors. The described fusion will make use of information from both sensor systems: Closer, but less information from the ultrasound sensors, far and with rich information from the camera sensors.
[0027] The augmented image of the surrounding environment is provided as input data for the electronic vehicle guidance system. In the electronic vehicle guidance system, based on the input data, at least one control signal may be generated for (semi-)automatically or (semi-)autonomously driving the motor vehicle. The at least one control signal may for example be provided to one or more actuators of the motor vehicle, including for example one or more braking actuators and / or one or more steering actuators and / or one or more propulsion motors of the motor vehicle. The one or more actuators may affect a longitudinal and / or lateral control of the motor vehicle in order to guide the motor vehicle at least in part automatically. Like this, the motor vehicle can for example be parked in the parking lot (semi-)automatically or (semi-)autonomously by the electronic vehicle guidance system. In the electronic vehicle guidance system, the input data may also be used to generate an assistance information for a driver of the motor vehicle. The assistance information may be output by means of an output device of the motor vehicle, for example a display and / or an audio output system and / or a haptic output system.
[0028] The invention can provide more effective and user-friendly feedback or guidance to the driver based on the augmented image. The feedback or guidance, i.e. the assistance information, can be visual, auditory and / or haptic, depending on the driver’s preference and situation. The feedback or guidance can also be adaptive and personalized, depending on the driver’s behavior and profile. Overall, the invention can improve the safety and efficiency of parking, especially of (semi-)autonomous parking, by providing a more accurate and preferably also more comprehensive representation of the surrounding environment of a motor vehicle.
[0029] The invention includes further embodiments by which additional advantages are obtained.
[0030] According to an embodiment, generating the second pixel-based two-dimensional image of the surrounding environment based on the ultrasound data comprises running a beamforming algorithm on the ultrasound data. As described, the received ultrasound data may be transferred to a beamforming control unit, where the beamforming algorithm may be run on the ultrasound data. Preferably, the beamforming control unit is situated within the ultrasound sensor unit of the motor vehicle. The beamforming control unit may be designed to control one or more ultrasound transducers of the ultrasound sensor unit, such that it may not only be able to run the beamforming algorithm on the received ultrasound data or signals, but such that it may also be able to control the transmitted ultrasound signals. The control of the transmitted ultrasound signals may for example comprise wave focusing and / or wave steering techniques. By using wave focusing and / or wave steering, an ultrasound beam can be constructed and steered in order to scan the surrounding environment, for example the parking lot, without moving the motor vehicle.
[0031] Therefore, according to a further embodiment, the ultrasound data comprise (i) ultrasound signals that are transmitted by an array of ultrasound transducers of the ultrasound sensor unit and (ii) ultrasound signals that are received by the array of ultrasound transducers of the ultrasound sensor unit. According to this embodiment, the beamforming algorithm is run on the transmitted ultrasound signals to achieve wave focusing and / or steering. Alternatively or additionally, the beamforming algorithm may also be run on the received ultrasound signals to achieve focusing and beamforming of the received signals. Like this, a high-resolution ultrasound image of the surrounding environment may be generated.
[0032] The ultrasound sensor unit may also comprise one or more of the usual hard- and / or firmware components, such as memory, buffers and / or amplifier stages, switches and / or filters, such as bandwidth filters, and the needed circuitry for transmitting and receiving control signals to / from the beamforming control unit. In order to facilitate data management and processing, at least parts of the described method may be outsourced to an external computing facility to use the computing capacity there.
[0033] Preferably, the beamforming algorithm comprises delaying and summing of the ultrasound signals to form a directional ultrasound beam (delay and sum - DAS). Like this, the resolution and quality of the ultrasound image can be improved. Delay and / or summing may be performed by the usual firmware components, such as a delay generator, a delay memory and / or a beam summer, some or all of which can be controlled by the beamforming control unit.
[0034] Preferably, the beamforming algorithm may be adjusted depending on a predetermined area of interest within the surrounding environment of the motor vehicle. For example, frequency and bandwidth modulation may be adjusted in order to increase contrast and depth resolution of objects in the predetermined area of interest. In the context of the present disclosure, this may be referred to as dynamic focusing.
[0035] As outlined above, the pixel resolutions of the pixel-based two-dimensional images are aligned before the low-level fusion. This alignment or cross-alignment may be accomplished by spatially interpolating a predetermined number of pixels of at least one of the pixel resolutions. This is an easy and efficient way of aligning the images, i.e. for obtaining an equal number of pixels in each image.
[0036] A motor vehicle typically carries sensors of additional types, such as lidar sensors and / or radar sensors. According to a further embodiment, the augmented image is fused with data from at least one further sensor unit of the motor vehicle, in particular with data from a lidar sensor unit and / or with data from a radar sensor unit of the motor vehicle. The fusion of the augmented image with the additional sensor data may be referred to as high- level fusion as described in the context of the present disclosure. By the method described herein, a pixel-based two dimensional image is generated from the ultrasound data. In other words, the ultrasound data are brought into the same format as the camera data. This is advantageous in that known computer vision algorithms, such as deep learning algorithms for object detection, segmentation and / or classification from the analysis of camera data can be used for analyzing the ultrasound data.
[0037] Computer vision algorithms, which may also be denoted as machine vision algorithms or algorithms for automatic visual perception, may be considered as computer algorithms for performing a visual perception task automatically. A visual perception task, also denoted as computer vision task, may for example be understood as a task for extracting visual information from image data. In particular, the visual perception task may in several cases be performed by a human in principle, who is able to visually perceive an image corresponding to the image data. In the present context, however, visual perception tasks are performed automatically without requiring the support by a human.
[0038] For example, a computer vision algorithm may be understood as an image processing algorithm or an algorithm for image analysis, which is trained using machine learning and may for example be based on an artificial neural network, in particular a convolutional neural network.
[0039] For example, the computer vision algorithm may include an object detection algorithm, an obstacle detection algorithm, an object tracking algorithm, a classification algorithm, a segmentation algorithm, and / or a depth estimation algorithm.
[0040] Corresponding algorithms may analogously be performed based on input data other than images visually perceivable by a human. For example, point clouds or images of infrared cameras et cetera may also be analyzed by means of correspondingly adapted computer algorithms. Strictly speaking, however, the corresponding algorithms are not visual perception algorithms since the corresponding sensors may operate in domains, which are not perceivable by the human eye, such as the infrared range. Therefore, here and in the following, such algorithms are denoted as perception or automatic perception algorithms. Perception algorithms therefore include visual perception algorithms, but are not restricted to them with respect to a human perception. Consequently, a perception algorithm according to this understanding may be considered as computer algorithm for performing a perception task automatically, for example, by using an algorithm for sensor data analysis or processing sensor data, which may, for example, be trained using machine learning and may, for example, be based on an artificial neural network. Also, the generalized perception algorithms may include object detection algorithms, object tracking algorithms, classification algorithms and / or segmentation algorithms, such as semantic segmentation algorithms.
[0041] In case an artificial neural network is used to implement a visual perception algorithm, a commonly used architecture is a convolutional neural network, CNN. In particular, a 2D- CNN may be applied to respective 2D-camera images. Also for other perception algorithms, CNNs may be used. For example, 3D-CNNs, 2D-CNNs or 1 D-CNNs may be applied to point clouds depending on the spatial dimensions of the point cloud and the details of processing.
[0042] The output of a perception algorithm depends on the specific underlying perception task. For example, an output of an object detection algorithm may include one or more bounding boxes defining a spatial location and, optionally, orientation of one or more respective objects in the environment and / or corresponding object classes for the one or more objects. An output of a semantic segmentation algorithm applied to a camera image may include a pixel level class for each pixel of the camera image. Analogously, an output of a semantic segmentation algorithm applied to a point cloud may include a corresponding point level class for each of the points. The pixel level classes or point level classes may, for example, define a type of object the respective pixel or point belongs to.
[0043] Accordingly, according to a further embodiment, in order to create a semantic map of the surrounding environment of the motor vehicle, at least one deep learning algorithm in this sense is run on the first pixel-based two-dimensional image and / or on the second pixelbased two-dimensional image and / or on the augmented image of the surrounding environment of the motor vehicle. In other words, each of the individual images can be further processed using the deep learning algorithm. Alternatively or additionally, the augmented image, i.e. the image that results from the low-level fusion as described herein, can be processed and evaluated using the deep learning algorithm. This facilitates environmental detection with regard to the known approaches and results in robust object detection.
[0044] According to a further embodiment, the deep learning algorithm is at least one of an object detection algorithm, a computer vision algorithm, a segmentation algorithm, and / or a classification algorithm. In the context of the present disclosure, an object detection algorithm may be understood as a computer algorithm, which is able to identify and localize one or more objects within a provided input dataset, for example input image, by specifying respective bounding boxes or regions of interest, ROI, and, in particular, assigning a respective object class to each of the bounding boxes, wherein the object classes may be selected from a predefined set of object classes. Therein, assigning an object class to a bounding box may be understood such that a corresponding confidence value or probability for the object identified within the bounding box being of the corresponding object class is provided. For example, the algorithm may provide such a confidence value or probability for each of the object classes for a given bounding box. Assigning the object class may for example include selecting or providing the object class with the largest confidence value or probability. Alternatively, the algorithm can specify only the bounding boxes without assigning a corresponding object class.
[0045] The deep learning algorithms can also be used to fuse the ultrasound and camera images using a convolutional neural network (CNN) with cross-alignment to create a fused image that highlights the objects and obstacles in front of or behind the motor vehicle in its surrounding environment. The deep learning algorithms can also segment, classify, and / or detect the objects and obstacles in the fused image using another CNN that can output a semantic map and bounding boxes with confidence scores. The CNN may follow the usual auto-encoder architecture with fused features and multiple decoders for object detection and semantic segmentation.
[0046] Unless stated otherwise, all steps of the computer-implemented method may be performed by a data processing device or apparatus, in particular by a data processing device of the vehicle, which preferably comprises at least one computing unit. In particular, the at least one computing unit is configured or adapted to perform the steps of the computer-implemented method. For this purpose, the at least one computing unit may for example store a computer program comprising instructions which, when executed by the at least one computing unit, cause the at least one computing unit to execute the computer-implemented method.
[0047] All computing units of the at least one computing unit may be comprised by the vehicle. However, it is also possible that all computing units of the at least one computing unit are part of an external computing system external to the vehicle, for example a backend server or a cloud computing system. It is also possible that the at least one computing unit comprises at least one vehicle computing unit of the vehicle as well as at least one external computing unit comprised by the external computing system. The at least one vehicle computing unit may for example be comprised by one or more electronic control units, ECUs, and / or one or more zone control units, ZCUs, and / or one or more domain control units, DCUs, of the vehicle and / or the camera.
[0048] In the present disclosure, a computing unit may for example be understood as a data processing device with processing circuitry. A computing unit can therefore perform computing operations in order to process data. The computing operations may also include indexed accesses to a data structure, for example a look-up table, LUT.
[0049] In particular, a computing unit may include one or more computers, one or more microcontrollers, and / or one or more integrated circuits, for example, one or more application-specific integrated circuits, ASIC, one or more field-programmable gate arrays, FPGA, and / or one or more systems on a chip, SoC. The computing unit may also include one or more processors, for example one or more microprocessors, one or more central processing units, CPU, one or more graphics processing units, GPU, and / or one or more signal processors, in particular one or more digital signal processors, DSP. The computing unit may also include a physical or a virtual cluster of computers or other of said units.
[0050] A computing unit may also comprise one or more hardware and / or software interfaces and / or one or more memory units. Therein, a memory unit may be implemented as a volatile data memory, for example a dynamic random access memory, DRAM, or a static random access memory, SRAM, or as a non-volatile data memory, for example a readonly memory, ROM, a programmable read-only memory, PROM, an erasable programmable read-only memory, EPROM, an electrically erasable programmable readonly memory, EEPROM, a flash memory or flash EEPROM, a ferroelectric random access memory, FRAM, a magnetoresistive random access memory, MRAM, or a phase-change random access memory, PCRAM.
[0051] A further aspect of the invention is directed at a data processing device, comprising an image processing unit, a beamforming control unit and a data processing unit. Therein, the image control unit is designed to generate a first pixel-based two-dimensional image with a first pixel resolution of the surrounding environment based on the camera data. The beamforming control unit is designed to generate a second pixel-based two-dimensional image with a second pixel resolution of the surrounding environment based on the ultrasound data, preferably by running a beamforming algorithm on the ultrasound data. The data processing unit is designed to align the first and second pixel resolutions to generate aligned two-dimensional images of the surrounding environment, to fuse the aligned two-dimensional images of the surrounding environment into one augmented image, and to provide the augmented image as input data for the electronic vehicle guidance system.
[0052] A further aspect of the invention is directed at an electronic vehicle guidance system, for a motor vehicle. The electronic vehicle guidance system comprises a camera sensor unit, comprising at least one camera sensor for receiving camera data describing a surrounding environment of the motor vehicle, an ultrasound sensor unit, comprising an array of ultrasound transducers for transmitting ultrasound signals to the surrounding environment of the motor vehicle and / or for receiving ultrasound signals from the surrounding environment of the motor vehicle, and a data processing device of the abovedescribed type.
[0053] An electronic vehicle guidance system may be understood as an electronic system, configured to guide a vehicle in a fully automated or a fully autonomous manner and, in particular, without a manual intervention or control by a driver or user of the vehicle being necessary. The vehicle carries out all required functions, such as steering maneuvers, deceleration maneuvers and / or acceleration maneuvers as well as monitoring and recording the road traffic and corresponding reactions automatically. In particular, the electronic vehicle guidance system may implement a fully automatic or fully autonomous driving mode according to level 5 of the SAE J3016 classification.
[0054] An electronic vehicle guidance system may also be implemented as an advanced driver assistance system, ADAS, assisting a driver for partially automatic or partially autonomous driving. In particular, the electronic vehicle guidance system may implement a partly automatic or partly autonomous driving mode according to levels 1 to 4 of the SAE J3016 classification. Here and in the following, SAE J3016 refers to the respective standard dated April 2021 .
[0055] Guiding the vehicle at least in part automatically may therefore comprise guiding the vehicle according to a fully automatic or fully autonomous driving mode according to level 5 of the SAE J3016 classification. Guiding the vehicle at least in part automatically may also comprise guiding the vehicle according to a partly automatic or partly autonomous driving mode according to levels 1 to 4 of the SAE J3016 classification. According to a further aspect of the invention, a computer program comprising instructions is provided. When the instructions are executed by the data processing device or by at least one computing unit thereof, the instructions cause the at least one computing unit to carry out a method according to the invention.
[0056] The instructions may be provided as program code, for example. The program code can for example be provided as binary code or assembler and / or as source code of a programming language, for example C, and / or as program script, for example Python.
[0057] According to a further aspect of the invention, a computer-readable storage medium storing a computer program according to the invention is provided.
[0058] The computer program and the computer-readable storage medium are respective computer program products comprising the instructions.
[0059] Further aspects of the invention are directed at a motor vehicle with an electronic vehicle guidance system and at a method for guiding the motor vehicle at least semi-automatically based on input data comprising an augmented image of a surrounding environment of the motor vehicle, wherein the augmented image is generated by a method according to the invention.
[0060] Further implementations of the various aspects according to the invention follow directly from the various embodiments of the method according to the invention and vice versa. In particular, individual features and corresponding explanations as well as advantages relating to the various implementations of the method according to the invention can be transferred analogously to corresponding implementations of the various other aspects according to the invention. In particular, the data processing device according to the invention is designed or programmed to carry out the method according to the invention. In particular, the data processing device according to the invention carries out the method according to the invention.
[0061] Further features of the invention are apparent from the claims, the figures and the figure description. The features and combinations of features mentioned above in the description as well as the features and combinations of features mentioned below in the description of figures and / or shown in the figures may be comprised by the invention not only in the respective combination stated, but also in other combinations. In particular, embodiments and combinations of features, which do not have all the features of an originally formulated claim, may also be comprised by the invention. Moreover, embodiments and combinations of features, which go beyond or deviate from the combinations of features set forth in the recitations of the claims may be comprised by the invention.
[0062] In the figures:
[0063] Fig. 1 shows a schematic view of a motor vehicle with an electronic vehicle guidance system according to an embodiment of the invention;
[0064] Fig. 2 shows a schematic view of an electronic vehicle guidance system according to an embodiment of the invention;
[0065] Fig. 3 shows a schematic view of a deep learning network using an autoencoder architecture to fuse camera and ultrasound data; and
[0066] Fig. 4 shows a schematic view of a data fusion method according to an embodiment of the invention.
[0067] Fig. 1 shows a schematic view of a motor vehicle 10 with an electronic vehicle guidance system 12 according to an embodiment of the invention. The electronic vehicle guidance system 12 comprises an environmental sensor system, comprising a plurality of sensors, which are designed to capture sensor data describing objects in a surrounding environment of the motor vehicle 10 and one or more computing units.
[0068] For the sake of clarity, Fig. 1 only shows a camera sensor unit 14, comprising at least one camera sensor 16, an ultrasound sensor unit 18 comprising an array of ultrasound transducers 20, and a data processing device 22, which is designed to receive sensor data from the camera sensor unit 14 and from the ultrasound sensor unit 18. Fig. 1 shows an exemplary data connection 24, which connects the camera sensor unit 14 and the ultrasound sensor unit 18 with the data processing device 22. Of course, the camera sensor unit 14 and the ultrasound sensor unit 18 can also be connected individually and independently of each other to the data processing device 22. A respective connection can be realized as a wireless connection or as a wired connection. The camera sensor 16 and / or the array of ultrasound transducers 20 may be arranged on a rear and / or on a front side of the motor vehicle 10, for example within one or more bumpers of the motor vehicle 10. The camera sensor 16 and / or the ultrasound transducers 20 may be arranged in order to observe a predetermined area of interest in front of or behind the motor vehicle 10. In other words, their respective fields of view may be directed at the predetermined area of interest. By using the described beamforming techniques, a directed ultrasound beam may be generated by the ultrasound transducers 20, which can for example be directed at different areas of interest depending on a current preference of a driver of the motor vehicle 10.
[0069] Overall, the hardware components of the electronic vehicle guidance system 12 may include an array of ultrasound transducers 20 that can transmit and receive ultrasound signals, a camera sensor 16 that can capture optical images, and a processing unit 30 that can fuse and interpret the ultrasound and camera data. Fig. 2 shows a schematic detailed view of the electronic vehicle guidance system 12 according to an embodiment of the invention. The electronic vehicle guidance system 12 as shown in Fig. 2 comprises at least a camera sensor unit 14, an ultrasound sensor unit 18 and a data processing device 22.
[0070] The data processing device 22 can comprise an image processing unit 26, a beamforming control unit 28 with a transmit beamformer unit 28.1 and a receive beamformer unit 28.2 and a data processing unit 30. The data processing device 22 may further comprise a preprocessing unit 32 and a computer vision unit 34. The data processing device 22 may be connected to an alert and guidance system 36 of the electronic vehicle guidance system 12.
[0071] Apart from the ultrasound transducers 20, the ultrasound sensor unit 18 may comprise a number of hard- and / or firmware components, like a buffer 38, switches 40 and an amplifier and filter stage 42.
[0072] The shown components of the electronic vehicle guidance system 12 can be embedded in a larger sensor system and computing environment of the electronic vehicle guidance system 12, comprising further sensors for localizing the motor vehicle 10, memory space for storing digital maps of the surrounding environment, computing units for path planning and / or decision making and / or motion control and the like. Fig. 3 shows a schematic view of a deep learning network 44 using an autoencoder architecture to fuse camera and ultrasound data. In a feature extraction or encoder stage 46, the deep learning network 44 may extract image features from the pixel-based two- dimensional ultrasound and camera images. A camera feature extractor 46.1 of the feature extraction stage 46 of the deep learning network 44 may extract features from the first pixel-based two-dimensional image, which was generated based on the camera data and having a first pixel resolution. Further, an ultrasound feature extractor 46.2 of the feature extraction stage 46 of the deep learning network 44 may extract features from the second pixel-based two-dimensional image, which was generated based on the ultrasound data and having a second pixel resolution. In an alignment and fusion stage 48, the deep learning network 44 may align the first and second pixel resolutions, for example by performing a method of spatial interpolation. The alignment may be performed in an alignment unit 48.1 of the alignment and fusion stage 48 and the fusion may be performed in a fusion unit 48.2 of the alignment and fusion stage 48. In a decoder stage 50, the deep learning network 44 may comprise a decoder 50.1 for object detection and a decoder 50.2 for semantic segmentation and / or any other type of computer vision application.
[0073] With reference to Figs. 1 to 3, Fig. 4 shows a schematic view of a data fusion method according to an embodiment of the invention, wherein camera data and ultrasound data are fused to generate input data for an electronic vehicle guidance system 12 of a motor vehicle 10, wherein the camera data and the ultrasound data describe a surrounding environment of the motor vehicle 10.
[0074] In a step S1 , camera data from the camera sensor unit 14 of the motor vehicle 10 are received and transmitted to an image processing unit 26, where one or more preprocessing steps can be carried out, for example in order to gain a first perception of the surrounding environment of the motor vehicle 10 from the camera data.
[0075] In a step S2, which may be carried out before, after or simultaneously with step 1 , ultrasound data from the ultrasound sensor unit 18 of the motor vehicle 10 are received and transmitted to an ultrasound processing unit, which can comprise a beamforming control unit 28.
[0076] In a step S4, a first pixel-based two-dimensional image with a first pixel resolution of the surrounding environment based on the camera data is generated by the image processing unit 26. In a parallel or subsequent step S5, a second pixel-based two-dimensional image with a second pixel resolution of the surrounding environment based on the ultrasound data is generated by the beamforming control unit 28.
[0077] In a step S6, the first and second pixel resolutions are aligned to generate aligned two- dimensional images of the surrounding environment. This can be done by applying some sort of spatial interpolation. For example, the ultrasound measurement after beamforming may result in a two-dimensional image having a pixel resolution of 100x100 pixels. On the other hand, the processing of the camera data may result in a two-dimensional image having a pixel resolution of 1000x1000 pixels. Alignment can be done to downscale or downsample from 1000x1000 to 100x100 (each 10 pixels will be corresponding to 1). This alignment process results in two two-dimensional images with the same pixel resolution. These images can then be mapped to each other 1 :1.
[0078] In a step S7, the aligned images may be fused to augment each other. This fusion, which may be referred to as “low-level” fusion, will lead to an augmented image of the surrounding environment of the motor vehicle 10, comprising information from the camera data and from the ultrasound data. The augmented image may be understood as an image that was generated from data from the same sensor type, i.e. from a virtual sensor type.
[0079] Finally, in a step S8, the augmented image of the surrounding environment is provided as input data for the electronic vehicle guidance system 12. The electronic vehicle guidance system 12 may use the augmented image as such, or it may also be designed to run further object detection and / or object classification algorithms on the augmented image. The electronic vehicle guidance system 12 may also use the augmented image and fuse it with a semantic map of the surrounding environment, which may be stored in a memory unit of the electronic vehicle guidance system 12 or which may be stored in an internetbased cloud service, which may be accessed by the electronic vehicle guidance system 12 on demand.
[0080] The electronic vehicle guidance system 12 may also fuse the augmented image in the course of “high-level” fusion with data from other sensors of the motor vehicle 10, such as lidar data, radar data or the like. The electronic vehicle guidance system 12 may comprise a feedback or guidance component. The feedback or guidance component can provide visual, auditory, or haptic feedback or guidance to the driver based on the interpretation results from the deep learning algorithms. For example, the electronic vehicle guidance system 12 can transmit the interpretation result to a display unit of the motor vehicle 10, where the fused augmented image may be displayed with semantic map information and bounding boxes on a screen in front of or behind the driver, along with some text or icons indicating the distance and / or direction to each object or obstacle in the surrounding environment of the motor vehicle 10. The electronic vehicle guidance system 12 can also play some sounds or voice messages for warning or instructing the driver about potential collisions or maneuvers. The electronic vehicle guidance system 12 can also induce vibrations or activate some actuators on the steering wheel or pedals to alert or assist the driver. The guidance component can also be used to guide or steer the motor vehicle 10 at least in parts automatically based on the interpretation results, for example to execute an automatic parking maneuver of the motor vehicle 10 in a parking lot.
[0081] Overall, the examples show, how a more accurate and preferably more comprehensive representation of the surrounding environment of a motor vehicle can be provided by fusing camera and ultrasound data.
Claims
Claims1 . Method for fusing camera data and ultrasound data to generate input data for an electronic vehicle guidance system (12) of a motor vehicle (10), wherein the camera data and the ultrasound data describe a surrounding environment of the motor vehicle (10), the method comprising the steps of- receiving (S1) the camera data from a camera sensor unit (14) of the motor vehicle (10),- receiving (S2) the ultrasound data from an ultrasound sensor unit (18) of the motor vehicle (10),- generating (S3) a first pixel-based two-dimensional image with a first pixel resolution of the surrounding environment based on the camera data,- generating (S4) a second pixel-based two-dimensional image with a second pixel resolution of the surrounding environment based on the ultrasound data,- aligning (S5) the first and second pixel resolutions to generate aligned two- dimensional images of the surrounding environment,- fusing (S6) the aligned two-dimensional images of the surrounding environment into one augmented image of the surrounding environment, and- providing (S7) the augmented image of the surrounding environment as input data for the electronic vehicle guidance system (12).
2. Method according to claim 1 , wherein generating the second pixel-based two- dimensional image of the surrounding environment based on the ultrasound data comprises running a beamforming algorithm on the ultrasound data.
3. Method according to claim 2, wherein the ultrasound data comprise (i) ultrasound signals that are transmitted by an array of ultrasound transducers (20) of the ultrasound sensor unit (18) and (ii) ultrasound signals that are received by the array of ultrasound transducers (20) of the ultrasound sensor unit (18), wherein the beamforming algorithm is run on the transmitted ultrasound signals and / or on the received ultrasound signals.
4. Method according to claim 3, wherein the beamforming algorithm comprises delaying and summing of the ultrasound signals to form a directional ultrasound beam.
5. Method according to any of claims 3 or 4, wherein the beamforming algorithm is adjusted depending on a predetermined area of interest within the surrounding environment of the motor vehicle (10).
6. Method according to any of the preceding claims, wherein the pixel resolutions are aligned by spatially interpolating a predetermined number of pixels of at least one of the pixel resolutions.
7. Method according to any of the preceding claims, wherein the augmented image is fused with data from at least one further sensor unit of the motor vehicle (10), in particular with data from a lidar sensor unit and / or with data from a radar sensor unit of the motor vehicle (10).
8. Method according to any of the preceding claims, wherein in order to create a semantic map of the surrounding environment of the motor vehicle (10), at least one deep learning algorithm is run on the first pixel-based two-dimensional image and / or on the second pixel-based two-dimensional image and / or on the augmented image of the surrounding environment of the motor vehicle (10).
9. Method according to claim 8, wherein the at least one deep learning algorithm is- an object detection algorithm,- a computer vision algorithm,- a segmentation algorithm, and / or- a classification algorithm.
10. Data processing device (22), comprising- an image processing unit (26),- a beamforming control unit (28), and- a data processing unit (30),wherein the data processing device (22) is designed to perform a method according to any of the preceding claims.11 . Computer program comprising instructions which, when the program is executed by a data processing device (22), in particular by a data processing device (22) according to claim 10, cause the data processing device (22) to carry out the steps of the method according to any of claims 1 to 9.
12. Computer-readable storage medium having stored thereon a computer program according to claim 11 .
13. Electronic vehicle guidance system (12) for a motor vehicle (10), comprising- a camera sensor unit (14), comprising at least one camera sensor (16) for receiving camera data describing a surrounding environment of the motor vehicle (10),- an ultrasound sensor unit (18), comprising an array of ultrasound transducers (20) for transmitting ultrasound signals to the surrounding environment of the motor vehicle (10) and / or for receiving ultrasound signals from the surrounding environment of the motor vehicle (10), and- a data processing device (22) according to claim 10.
14. Motor vehicle (10) with an electronic vehicle guidance system (12) according to claim 13.
15. Method for guiding a motor vehicle (10) at least semi-automatically based on input data comprising an augmented image of a surrounding environment of the motor vehicle (10), wherein the augmented image is generated by a method according to any of claims 1 to 9.
Citation Information
Patent Citations
Collision distance calculation method and system based on fusion of ultrasound wave and camera
CN110488319A
Detecting a trailer hitch in the vicinity of a vehicle
DE102022121778A1
Network training process for hardware definition
US11151447B1
Stereo depth estimation using deep neural networks
US20190295282A1
Visual perception using a vehicle on the basis of a camera image and an ultrasonic map
WO2024041833A1