Driver assistance device and method
A sound detection and localization system using strategically placed microphones and neural networks addresses the limitations of existing ADAS systems by enhancing auditory awareness, improving road safety through real-time sound analysis and localization.
Patent Information
- Application Number
- PCT/ES2025/070446
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-07-21
- Publication Date
- 2026-01-29
AI Technical Summary
Existing Advanced Driver Assistance Systems (ADAS) in vehicles are ineffective in detecting and localizing auditory signals, such as emergency vehicle sirens, due to improved sound insulation and background noise, leading to driver stress and potential accidents.
A comprehensive sound detection and localization system using strategically placed microphones with aerodynamic structures and deep neural networks for real-time sound analysis, preprocessing, and localization, integrated with vehicle's existing ADAS systems.
Enhances driver awareness of critical auditory signals, reducing stress and improving road safety by accurately detecting and localizing relevant sounds in real-time, complementing existing ADAS systems.
Smart Images

Figure ES2025070446_29012026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] Driving assistance device and procedure
[0003] TECHNICAL SECTOR
[0004] The present invention relates to a device that defines a sound detection and localization system for driving assistance. The information thus received is transmitted to the driver to enable them to make the right decisions.
[0005] STATE OF THE ART
[0006] Safe driving depends on the driver's visual perception, but also on their ability to detect and respond appropriately to auditory signals, such as emergency vehicle sirens, vehicle horns, etc.
[0007] In many situations where an emergency vehicle is trying to make its way through traffic, drivers don't clearly hear the sound and are almost never able to discern the direction it's coming from. This is due to the increasingly better sound insulation in vehicles, or also to the playing of music or radio inside.
[0008] These situations can cause stress and lead to incorrect reactions from the driver. Some newer vehicles use camera-based visual systems, but these are only effective at detecting very close vehicles and are not useful for the use cases mentioned. ADAS systems with precise audible information could reduce the risk of accidents and increase road safety.
[0009] The use of microphone systems to recognize sounds and their origin, associated with driving, is known in the prior art. Examples can be found in US2021358300, US2019294169, and US2022284919. These inventions have not been incorporated into ADAS systems.
[0010] The present invention therefore addresses the development of a comprehensive solution that complements the other ADAS systems of a vehicle, providing the driver with relevant auditory information for safer and more efficient driving.
[0011] The applicant is unaware of any device that could be considered similar to the invention. BRIEF EXPLANATION OF THE INVENTION
[0012] The invention relates to a driving assistance device and procedure according to the independent claims and whose features solve the problems of the state of the art.
[0013] According to a first aspect, the present invention discloses a driver assistance device. The device is a comprehensive hardware and software system designed to enhance Advanced Driver Assistance Systems (ADAS) capabilities in vehicles. The system complements the vehicle's existing visual systems by adding real-time sound detection and analysis capabilities through the installation of strategically placed microphones around the vehicle, which capture sounds from the surrounding environment. This sound data is processed by specific hardware and dedicated software, which analyzes the direction, distance, movement, and proximity of sound sources, particularly those of interest such as emergency vehicle sirens or other events relevant to driving safety.
[0014] Specifically, the driver assistance device in a vehicle incorporates a series of microphones and a control unit that runs sound recognition software to detect the presence of another vehicle. The device comprises microphones distributed around the vehicle in aerodynamic structures. Preferably, the device comprises at least three microphones. According to a preferred embodiment, the microphones are located at the outer corners of the vehicle's passenger compartment. In this document, the term "outer corners of the passenger compartment" is understood to mean the following four areas: the front and rear of the bodywork area enclosing the passenger compartment on both sides of the vehicle. In the most preferred embodiment, it comprises a single microphone in each outer corner of the passenger compartment, for a total of four microphones.
[0015] Preferably, aerodynamic structures comprise a filling of absorbent material.
[0016] Preferably, the aerodynamic structures include two inlet openings. Preferably, the inlet openings to the aerodynamic structures are equipped with wind filters.
[0017] Preferably, the aerodynamic structures have a shape selected from the group comprising conical, hemispherical, or tubular. According to a second aspect, the present invention discloses a driving assistance method using the device of the first aspect. The method comprises the following steps:
[0018] - Preprocessing of signals captured by the microphones, to reduce wind noise and rolling noise.
[0019] - Calculation of the MEL spectrogram of the microphone signals.
[0020] - Analysis using a first deep neural network of transformers for sound detection inference, including the calculation of the GCC-PHAT of the different microphone pairs.
[0021] - Processing using a second deep convolutional neural network of the result of the GCC-PHAT calculation for the inference of sound localization.
[0022] - Review using an algorithm to estimate the movement of the sound source.
[0023] - Transmission of the result to the user.
[0024] Other variations can be seen in the rest of the memory.
[0025] In this document, the word "comprises" and its variants are to be interpreted as open-ended expressions that do not preclude the possibility of other technical features or components beyond those explicitly mentioned. Furthermore, the word "comprises" includes the case "consists of," which is interpreted as a closed-ended expression limited solely to the technical features or components explicitly mentioned. For those skilled in the art, other objects, advantages, and features of the invention will become apparent partly from the description and partly from the practice of the invention. Moreover, the present invention covers all possible combinations of embodiments described herein.
[0026] DESCRIPTION OF THE FIGURES
[0027] For a better understanding of the invention, the following figures are included, showing exemplary embodiments.
[0028] Figure 1: View of a vehicle where an example of an embodiment of the invention has been installed, in a conical shape.
[0029] Figure 2: Examples of aerodynamic structures: (A) tubular, (B) hemispherical.
[0030] Figure 3: Schematic cross-section of an example of an aerodynamic structure. Figure 4: Schematic of the procedure.
[0031] Figure 5: Spectrograms of the sound of a siren, without wind noise (top of the figure) and with wind noise (bottom of the figure), obtained in an experimental study.
[0032] MODES OF REALIZING THE INVENTION
[0033] Next, a brief description is given of one way of carrying out the invention, as an illustrative and non-limiting example thereof.
[0034] The device described in Figure 1 consists of a series of hardware components and software. The hardware includes four microphones (1) positioned approximately at the outer corners of the vehicle's passenger compartment (2). The microphones (1) are mounted on aerodynamic structures (3) to reduce both wind noise and noise from vibrations of the moving vehicle. The microphones (1) are connected to converters (4) that digitize the signals and send them to a control unit (5) that executes the software in real time. A user-accessible screen (6) displays the conclusions drawn by the software from the acoustic signals. In other words, it is configured to show the vehicle information detected by the microphones (1).
[0035] The procedure of the present invention, schematically illustrated in Figure 4, is explained below. The signal processing program begins with a signal preprocessing stage (101) to further reduce wind noise and eliminate rolling noise. Next, the MEL spectrogram of the microphone signals (1) is calculated (102). This spectrogram is then processed by a first deep transformer neural network (103) for sound detection inference, determining whether any sounds are relevant. This first neural network (103) selects the noises considered relevant, discarding the rest. The selected noises then proceed through the stages of the procedure. If relevant, the type of sound is identified and displayed on the screen (6).In parallel, an initial study (108) of the sound origin is performed (distance, direction from which it originates, and, if the source is mobile, its movement). This study is carried out using an Encoder-Decoder network (107) that performs a noise removal process on the signal of interest. Subsequently, the GCC-PHAT calculation (104) is processed between each possible pair of microphones, ensuring that it is performed for all possible pairs of microphones (six combinations in the case of four microphones).
[0036] The result of the GCC-PHAT calculation (104) is processed by a second deep convolutional neural network (105) for sound localization inference. This inference is checked by a sound source motion estimation algorithm (106) and displayed on the screen (6) with the appropriate interface.
[0037] In light of the above, the procedure includes the use of artificial intelligence in the form of neural networks. In a preferred embodiment, the artificial intelligence comprises:
[0038] 1. Detector: Audio Spectrogram Transformer (AST), adapting the last layers to have two dense layers. The first is a dense layer of 768 neurons activated with a PRelu function. The second has 3 neurons, activated with the SoftMax function. Training was performed over 100 epochs with an early stop of 10 epochs.
[0039] 2. Cleaner: This is an Encoder-Decoder Network. Encoder-Decoder Networks are a type of deep neural network architecture widely used in tasks where it is necessary to transform an input into an output in a structured way. They are especially effective for problems requiring complex mapping between input and output, such as machine translation, text generation, and signal processing, including cleaning audio spectrograms, which is our case. The Encoder-Decoder architecture consists of two main parts:
[0040] • Encoder: a. Takes the input (in this case, a noisy audio spectrogram) and transforms it into a lower-dimensional latent representation. b. This process involves several layers (which may be convolutional, recurrent, or transformer layers) that compress the relevant information into a compact representation.
[0041] • Decoder: a. Takes the latent representation generated by the encoder and transforms it back into a comprehensible form (in this case, a clean audio spectrogram). b. Similarly, it involves several layers that expand the latent representation to reconstruct the desired output.
[0042] An example of a U-NET type encoder-decoder network can be seen in O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation", in Lecture Notes in Computer Science (including subsehes Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2015. doi: 10.1007 / 978-3- 319-24574-4_28.
[0043] The network has 2D matrices as input and output, creating a bottleneck where the extracted information is concentrated, and subsequently this information is expanded again as a result of the bottleneck.
[0044] 3. Locator: This is a ResNet18 network with adapted final layers. A Dense Layers network of 512 neurons activated with the PRelu function is connected to the ResNet output. This layer is then connected to two more parallel layers. Each of these parallel networks is another Dense Layers network of 512 neurons with PRelu activation. Finally, a single neuron is connected to the output of each parallel network, applying the arctangent of the values of the parallel networks.
[0045] 4. Movement estimation: Using the location values over time, the movement of the sound is calculated.
[0046] Next, the level of detail in the explanation will be increased.
[0047] The microphones (1) are strategically distributed, ideally placing one in approximately each of the four outer corners of the vehicle's passenger compartment. This arrangement allows the system to use only four microphones (1), resulting in a significant reduction in the amount of data to be processed and an increase in response speed. This optimization is crucial in real-time applications, improving the efficiency and speed of the various systems involved.
[0048] The microphones (1) must be able to withstand the adverse temperature and humidity conditions typical of outdoor environments. They are powered by the vehicle itself, receiving a power cable and transmitting a sound signal, usually also via wiring.
[0049] The microphones (1) are housed within aerodynamic structures (3), specially designed to minimize wind resistance and reduce vibrations. This aerodynamic structure (3), preferably made of any waterproof material, can take various geometric shapes, including spheres or cones. Figure 2 shows two preferred examples. A notch for cable routing is visible.
[0050] The aerodynamic structures (3) improve the sound capture quality. To achieve this, the aerodynamic structures (3) are filled with an absorbent material (31). This absorbent material (31) enhances the acoustic and mechanical insulation function of the structure. Preferably, the absorbent material (31) is a thermoacoustic insulator with the following characteristics:
[0051] - Acoustic absorption coefficient greater than 0.9, indicating an excellent ability to absorb sound waves across a wide range of frequencies, rather than reflecting or amplifying them within the enclosure.
[0052] - Acoustic insulation against airborne noise exceeding 60 dB, guaranteeing effective protection against unwanted sounds from the outside environment.
[0053] The preferred density is between 70 and 90 kg / m³ to effectively support the microphone without the need for an additional structure, reducing rigid contact points that could transmit vibrations. This feature not only simplifies the design but also improves the microphone's mechanical independence from the housing, enhancing the quality of the captured signal.
[0054] For safety, it is desirable that its fire reaction be B-s1,d0, meaning it does not propagate flame, generates little smoke, and does not produce flammable droplets, making it suitable for automotive applications. Furthermore, it is desirable that its thermal resistance be at least 1.00 m. 2 K / W, which gives it good stability against temperature changes, preventing deformations or degradations that could affect its long-term performance.
[0055] An example of suitable absorbent material (31) is agglomerated polyurethane foam. This absorbent material (31) helps to further reduce the effects of vehicle vibrations and wind noise.
[0056] The aerodynamic structures comprise inlet openings (32). The inlet openings (32) of the aerodynamic structures (3) are equipped with windscreens (33), such as the one shown in Figure 3, which allow sound to pass through but block wind and airborne particles. The windscreens (33) consist of one or more layers of acoustically transparent material, such as nylon fabric stretched over a support assembly. This arrangement almost completely eliminates the effect of residual wind on the system, especially under high-speed driving conditions, thus ensuring accurate detection of sounds of interest on the road.
[0057] The signals, pre-amplified if necessary, are digitized by analog-to-digital converters (ADCs) (4) at a specific sampling frequency (fs). An fs of 16000 Hz is used as the baseline, but the invention supports any sampling frequency, higher or lower than this. The converters (4) can be located within the aerodynamic structure (3), as shown in Figure 3, or associated with the control unit (5).
[0058] The control unit (5) will carry out the following steps: noise reduction algorithms, analysis of audio samples by frames for the extraction of sound features, execution of inference algorithms for the detection and localization of sound sources using artificial intelligence models
[0059] It is crucial that this system functions optimally in real time, therefore requiring very high processing capacity in terms of floating-point operations (FLOPs). Since it will be the core of the entire system, its performance will be fundamental to the overall operation of the vehicle. In general, the control unit (5) will be capable of deep parallel processing, so architectures based wholly or partially on GPUs are recommended. Using only four microphones (1) also limits processing requirements.
[0060] The control unit (5) may be a separate component from the vehicle's onboard central computer, or it may be integrated within it, depending on the architecture and specifications of the particular system. In vehicles with a powerful computer system with parallel processing, based wholly or partially on GPUs, and with available processor time, the control unit (5) could be integrated into it.
[0061] Once the information is processed, the system displays the results visually or audibly within the vehicle to alert the driver. These alerts include graphical information on instrument panel displays (6) and, optionally, audible alerts for the driver using the vehicle's sound system speakers. Additionally, automatic vehicle behavior adjustments can be implemented in response to the detected sounds. In this way, the central computer plays a crucial role in improving the vehicle's perception and safety within its auditory environment.
[0062] The following are some observations on different aspects of the present invention, according to a preferred embodiment:
[0063] Importance of the aerodynamic structure in motion
[0064] It is important to emphasize that the aerodynamic structure (3) of the present invention is not simply a protective housing for the microphones (1), but an essential element for ensuring their proper functioning under moving vehicle conditions. The main function of this structure is to mitigate the impact of aerodynamic wind noise, which is generated at high speeds and directly masks the fundamental frequencies of sounds relevant to road safety, such as emergency vehicle sirens. This phenomenon has been studied experimentally and can be observed, for example, in Figure 5, which shows a spectrogram representing the variation of frequencies (vertical axis) over time (horizontal axis) and illustrating how wind noise overlaps critical spectral bands of a siren's sound, drastically reducing the efficiency of the procedure in subsequent stages.
[0065] This masking has the following negative repercussions:
[0066] - Sound detection: may not be recognized or may generate false positives because the fundamental frequencies that Artificial Intelligence has learned to recognize from the sounds of interest do not appear due to this masking.
[0067] - Sound localization: The GCC-PHAT algorithm used to infer direction is degraded when the signal is hidden by background noise. The signal-to-noise ratio (SNR) is drastically reduced, greatly increasing the error rate.
[0068] Thus, the present invention includes an acoustic pickup device (microphone in an aerodynamic structure) based on an external architecture whose shape and components are optimized not only for the physical protection of the microphone (1), but also to significantly improve the quality of the captured signal before digital processing. This improvement has a direct and measurable impact on the performance of subsequent detection and localization algorithms.
[0069] Role of the aerodynamic structure alongside the wiper (Red Encoder-Decoder)
[0070] In addition to the physical structure, the system incorporates a cleaning block based on an Encoder-Decoder neural network (e.g., U-NET). This network acts digitally on the spectrograms to eliminate residual noise that has not been fully attenuated by the structure, including not only lingering wind noise but also tire noise produced by the contact of the tires with the asphalt—another significant component that interferes with the quality of the captured audio.
[0071] Both elements — physical structure and digital cleaner — are complementary and important to ensure the reliability of the system under real driving conditions.
[0072] Mitigation of mechanical vibrations of the vehicle
[0073] A critical point that must be highlighted is the transmission of vehicle vibrations through the body to the microphone. This can compromise both the integrity of the sensor and the quality of the captured sound.
[0074] - At a mechanical level, constant vibrations can deteriorate the microphone and generate non-linear responses or unwanted artifacts.
[0075] - At an acoustic level, these vibrations generate structural noises that, once again, mask frequencies of interest, hindering correct detection and localization.
[0076] According to the proposed solution, the microphone is suspended by means of an acoustic insulation medium (absorbent material (31)) within the aerodynamic structure, which prevents structural vibrations of the vehicle from being transmitted to the sensor.
[0077] Structural safety and aerodynamic design
[0078] When the vehicle is in motion, especially at high speeds, it is crucial that the structure remains firmly attached to prevent displacement or detachment. For this reason, the structure has been designed with an aerodynamic shape that takes advantage of the air pressure generated by the vehicle's movement. This pressure helps the structure remain more securely attached to the body.
[0079] Thanks to this design, wind is not only no longer a problem, but it is also used to improve the structure's stability while driving. This allows the device to function correctly without risk of coming loose.
[0080] Aerodynamic conical structure (implementation example)
[0081] Conical shape: flow guidance and dynamic pressure reduction on the central axis
[0082] The conical shape allows the incident airflow to be naturally deflected towards the outer edges of the cone, rather than directly impacting the microphone diaphragm. This creates a zone of zero or very low turbulence at the center of the cone's inlet—that is, on the microphone's pickup axis—protecting the signal from aerodynamic interference. Unlike flat or open shapes, no eddies or direct air impacts are generated in the most sensitive area, significantly reducing unwanted acoustic disturbances.
[0083] It should be noted that the structure is not a perfect cone, but rather a cone truncated laterally, creating a flat longitudinal surface on one side. This flat area allows the structure to fit securely to the vehicle, facilitating its attachment without compromising the overall aerodynamic performance. By maintaining the conical shape at the front, airflow guidance and aerodynamic stability are preserved, while the flat section ensures a safe and practical installation.
[0084] Furthermore, thanks to its progressive and continuous geometry, the air flows laminarly and stably around the entire surface of the cone, without encountering abrupt changes, sharp angles, or irregularities that could cause flow separation, turbulence, or vortices. This not only protects the microphone from the wind but also prevents the structure from vibrating due to air pressure, which is especially important when the vehicle is traveling at high speed. By reducing flow-induced vibrations, the microphone is mechanically protected, and the generation of low-frequency noise or internal resonances is prevented.
[0085] The conical shape, therefore, acts as a passive aerodynamic profile, channeling and stabilizing the air surrounding the microphone, improving sound capture from a physical point of view even before applying any type of digital processing.
[0086] This effect cannot be achieved with other types of structures, as these shapes tend to generate unstable pressures, internal resonances, and stagnant zones that distort the signal. In many cases, these shapes also fail to prevent vibrations transmitted through the air or the vehicle's body itself, further compromising the quality of the captured audio.
[0087] Peripheral rim: turbulence containment
[0088] Ideally, the cone's inlet should have a peripheral lip. This lip acts as a physical barrier, slowing the spread of turbulence generated at the edge of the airflow. Tests have shown that turbulence—inevitable at any open-geometry edge—occurs precisely at the margins. This lip allows any potential flow disturbances to dissipate before reaching the front of the microphone.
[0089] Without this rim, or in solutions without defined geometry, these turbulences propagate inwards and contaminate the signal with low-frequency components and pulsating noise, which hinder the extraction of relevant features in spectrograms, especially in the early stages of digital processing.
[0090] Thanks to the aerodynamics of the conical shape and the peripheral rim, the flow pressure is virtually zero Pascals across the entire structure, completely eliminating turbulence and vibrations on the cone's surface. This not only improves the stability of the captured acoustic signal but also contributes to a better physical attachment of the structure to the vehicle by preventing the generation of unwanted forces.
[0091] Furthermore, pressure differences are concentrated at the edge of the cone's inlet, precisely where the peripheral safety flange is located. This configuration ensures that the entire central area of the cone remains turbulence-free, guaranteeing a stable and protected front opening for sound capture.
[0092] Wind filter: passive suppression of fluctuating pressure
[0093] A windscreen filter (33) is placed at the inlet to the cone. This filter acts as an acoustically transparent mesh that stabilizes the flow, dampening any gusts or residual flows that may have overcome the aerodynamic barrier of the cone and the rim. This component is important for eliminating high-frequency microturbulence, which in the digital signal translates into white noise or unstructured spectra that are difficult to clean up even with neural networks trained for this purpose.
[0094] Although the present invention has been described with reference to particular options and embodiments thereof, those skilled in the art may make modifications and variations to the foregoing teachings without departing from the scope and spirit of the present invention.
Claims
CLAIMS 1- Driving assistance device, in a vehicle (2), incorporating a series of microphones (1), a control unit (5), characterized in that: ■ the device comprises microphones (1) distributed around the vehicle (2) in separate aerodynamic structures (3); ■ The control unit (5) is configured to run a sound recognition program using neural networks to check for the presence of another vehicle. 2- Driving assistance device, according to claim 1, characterized in that the distribution of the microphones consists of a single microphone (1) in each outer corner of the vehicle cabin (2). 3- Driving assistance device, according to any of the preceding claims, characterized in that the aerodynamic structures (3) comprise a filling of absorbent material (31). 4- Driving assistance device, according to any of the preceding claims, characterized in that the aerodynamic structures (3) comprise respective inlet holes (32). 5- Driving assistance device, according to claim 4, characterized in that the inlet holes (32) to the aerodynamic structures (3) have wind filters (33). 6- Driving assistance device, according to any of the preceding claims, characterized in that the aerodynamic structures (3) have a shape selected from the group comprising conical, hemispherical or tubular. 7- Driving assistance device, according to claims 4 and 6, characterized in that the aerodynamic structures (3) are conical in shape, and comprise a peripheral flange in the inlet hole (32). 8- Driving assistance device, according to any of claims 6 or 7, characterized in that the aerodynamic structures (3) have a defined conical shape like a laterally truncated cone, generating a longitudinal flat surface on one of its sides, where it attaches to the vehicle. 9- Driving assistance device, according to any of the preceding claims, characterized in that each microphone (1) comprises an analog / digital converter (4) in the aerodynamic structure (3) itself. 10- Driving assistance device, according to any of the preceding claims, characterized in that the control unit (5) is integrated into the central computer on board the vehicle (2) and connected to a user-accessible screen (6) configured to display information about the detected vehicle. 11- Driving assistance procedure, with the device of any one of the preceding claims, characterized in that it comprises the steps of: ■ preprocessing (101) of signals captured by the microphones (1) to reduce wind noise and rolling noise; ■ calculation (102) of the MEL spectrogram of the microphone signals (1); ■ analysis using a first deep transformer neural network (103) for sound detection inference, including the calculation of the GCC-PHAT (104) of different microphone pairs (1); ■ treatment by means of a second deep convolutional neural network (105) of the result of the calculation of the GCC-PHAT (104) for the inference of the sound localization; ■ review using an algorithm for estimating the movement of the sound source (106); ■ transmission of the result to the user.
Citation Information
Patent Citations
Method and apparatus for detecting a proximate emergency vehicle
US20190294169A1
Method for identifying sirens of priority vehicles and warning a hearing-impaired driver of the presence of a priority vehicle
US20210358300A1
Detection and classification of siren signals and localization of siren signal sources
US20220284919A1
Closed sound sensor with sound-permeable interface
DE102022205148B3
Acoustic equipment for vehicle
JP1986048210A