Vehicle-mounted voice identification and positioning intelligent early warning system and method

Through the intelligent early warning system for on-board sound recognition and positioning, combined with the fusion of multi-sensor data and deep learning model, the shortcomings of existing systems in identifying and positioning obstacles in complex environments are solved, and the early warning effect with higher accuracy and rapid response is achieved, reducing the risk of traffic accidents.

CN120048287AActive Publication Date: 2025-05-27RIVOTEK TECH (JIANGSU) CO LTD

Patent Information

Application Number
CN202510238418.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-27
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing vehicle-mounted intelligent early warning system cannot accurately and timely identify and locate obstacles in complex environments, such as occlusion, poor visual distance or environmental noise interference, resulting in insufficient system response speed and accuracy, increasing the risk of traffic accidents.

Method used

The intelligent early warning system for on-board sound identification and positioning is adopted to collect audio signals in real time through the sound detector module, combine deep learning models to identify the sound source type, and perform spatiotemporal calibration with on-board camera and radar data. Through the fusion and precision positioning of multi-sensor data, collision risks are evaluated and warning signals are generated.

Benefits of technology

It significantly improves the response speed and accuracy of the on-board intelligent early warning system in complex environments, enhances the system's environmental adaptability, and reduces the risk of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048287A_ABST
    Figure CN120048287A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted voice identification and positioning intelligent early warning system and method, and relates to the technical field of vehicle-mounted intelligent early warning. The system comprises a sound detector module which collects an audio signal of a surrounding environment in real time, and a sound recognition module which analyzes the audio signal through a deep learning model and recognizes a sound source type. And the calibration module performs space-time calibration to ensure data synchronization and alignment of the sensors. The positioning module combines time difference of arrival (TDOA) and a sound intensity attenuation model, and optimizes sound source positioning through particle swarm optimization (PSO) and extended Kalman filter (EKF) technologies. The risk assessment module assesses the collision risk based on the sound source position, the shielding state and the object type, gives an alarm to a driver in time through the early warning module, and automatically triggers a deceleration instruction. And the self-learning module dynamically adjusts the early warning sensitivity and the noise compensation coefficient through a federated learning framework updating model. The environmental adaptability of the vehicle-mounted early warning system is improved, and the vehicle-mounted early warning system has remarkable safety improvement and wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of in-vehicle intelligent warning, and particularly to an in-vehicle sound recognition and positioning intelligent warning system and method. Background Art

[0002] With the rapid development of intelligent driving technology, autonomous driving systems and assisted driving systems have become an indispensable part of the modern transportation field. Sensors such as in-vehicle cameras, radars, and lidar play important roles in these systems, capable of providing visual information about the surrounding environment and object detection capabilities for the vehicle. However, existing systems still have significant limitations in certain specific environments and complex scenarios. Especially in situations where the line of sight is limited, the sensors are blocked, or there is environmental noise interference, the detection capabilities of these traditional sensors are often severely affected, resulting in the system being unable to detect potential obstacles or hazard sources in a timely and accurate manner.

[0003] For example, in narrow urban roads, intersections, or situations blocked by other vehicles, the fields of view of in-vehicle cameras and radars may be partially blocked, making it impossible to effectively detect pedestrians, bicycles, tricycles, or other obstacles ahead. This "sudden appearance" phenomenon is particularly likely to occur, resulting in the vehicle being unable to make timely warning or avoidance actions during driving, thus increasing the risk of traffic accidents. Traditional in-vehicle warning systems usually rely on visual sensors and radars to evaluate the collision risk by calculating the relative position between the object and the vehicle. However, these systems have significant deficiencies in complex or blocked environments. Especially when there are potential high-risk obstacles in front of the vehicle, existing systems are difficult to provide accurate warning information. In addition, existing systems also face the problem of environmental noise interference. Especially in noisy environments such as urban streets and busy intersections, the signals collected by the sensors may be interfered with, resulting in the system being unable to correctly identify and locate obstacles, thereby affecting the accuracy and timeliness of the warning.

[0004] In addition, the warning accuracy of the system is often limited by the performance of a single sensor. Especially when the detection area of an in-vehicle camera or radar is limited, "blind spots" are likely to occur, causing potential dangers to not be recognized in a timely manner. Therefore, how to achieve a warning system with higher accuracy and faster response in complex environments has become a key challenge in the current technological development. Summary of the Invention

[0005] The object of the present invention is to provide an in-vehicle sound recognition and positioning intelligent warning system and method, which solves the problem that in the prior art, in complex environments such as occlusion, poor line of sight, or environmental noise interference, in-vehicle cameras and radars cannot accurately and timely identify and position obstacles, thereby significantly improving the response speed and accuracy of the in-vehicle intelligent warning system in variable and complex driving environments, enhancing the environmental adaptability of the system, and effectively reducing the risk of traffic accidents.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] An in-vehicle sound recognition and positioning intelligent warning system, comprising:

[0008] A sound detector module, configured to collect audio signals of the surrounding environment in real time;

[0009] A sound recognition module, configured to analyze the collected audio signals through a deep learning model, identify the sound source type, and determine whether the sound source is occluded by combining in-vehicle camera and radar data;

[0010] A calibration module, configured to perform precise spatio-temporal calibration on the audio signals and in-vehicle camera and radar detection data, optimize signal alignment through time synchronization and spatial mapping, and determine whether the sound source object is occluded;

[0011] A positioning module, configured to calculate the azimuth of the sound source object through the time difference of arrival TDOA of signals from multiple sound detectors, calculate the distance of the sound source object by combining the sound intensity attenuation model and the particle swarm optimization PSO algorithm, and correct the positioning result by fusing camera and radar data through the extended Kalman filter EKF, and dynamically adjust the noise compensation coefficient according to the environmental noise intensity;

[0012] A risk assessment module, configured to assess the collision risk based on the spatial position, occlusion state, object type, and distance of the sound source, and generate warning signals of different levels;

[0013] A warning module, configured to display the sound source position on the in-vehicle display screen and issue a voice prompt, and trigger an automatic deceleration command when the collision risk meets a preset condition;

[0014] A self-learning module, configured to update the model through a federated learning framework, and dynamically adjust the warning sensitivity and noise compensation coefficient according to the regional type.

[0015] As a preferred solution of the present invention, the sound recognition module includes:

[0016] The timbre feature extraction unit extracts the feature vector of the audio signal through the fusion algorithm of Mel Frequency Cepstral Coefficients (MFCC) and Linear Predictive Coding (LPC). The fusion algorithm combines the feature vectors of MFCC and LPC in a weighted splicing manner. The specific steps are as follows:

[0017] a. Preprocess the audio signal, including denoising and framing operations;

[0018] b. Use the MFCC algorithm to extract the spectral features of the audio signal to obtain the feature vector:

[0019] MFCC = [mfcc 1 , mfcc 2 ,..., mfcc N ;

[0020] Where: mfcc N represents the Nth eigenvalue of the Mel Frequency Cepstral Coefficients; N is the dimension of the MFCC feature;

[0021] c. Use the LPC algorithm to extract the prediction coefficient features of the audio signal to obtain the feature vector:

[0022] LPC = [lpc 1 , lpc 2 ,..., lpc M ;

[0023] Where: lpc M represents the Mth eigenvalue of the LPC feature; M is the dimension of the LPC feature;

[0024] d. Combine the MFCC and LPC feature vectors through weighted splicing to obtain the combined feature vector:

[0025] F fusion = [mfcc 1 , mfcc 2 ,..., mfcc N , lpc 1 , lpc 2 ,..., lpc M ;

[0026] Where: F fusion is the final combined feature vector, N and M are the dimensions of the MFCC and LPC features respectively. In weighted splicing, the weight coefficients of MFCC and LPC are adjusted according to the application environment or training data;

[0027] A classification unit is used to classify the fused audio feature vectors based on the Deep Residual Network (ResNet) model and output the sound source type labels, including: pedestrians, tricycles, motorcycles, and other motor vehicles. The classification method uses a deep learning model and is adapted to variable and complex noise environments;

[0028] An occlusion judgment unit is used to determine that the sound source object is occluded when the in-vehicle camera or radar fails to detect a target in the predicted sound source area. The occlusion judgment method improves the traditional occlusion detection based on visual sensors by combining the fusion of sound signals and visual signals.

[0029] As a preferred solution of the present invention, the calibration module includes:

[0030] A time synchronization unit: uses the Precision Time Protocol (PTP) to calibrate the time of each sensor, and the synchronization error Δt ≤ 1 ms;

[0031] A space mapping unit: maps the angle of the sound source position to the angle of the in-vehicle camera's field of view. The mapping formula is:

[0032]

[0033] where: θ is the horizontal angle of the sound source relative to the in-vehicle camera, φ is the vertical angle of the sound source relative to the in-vehicle camera, X source and Y source are the lateral and longitudinal coordinates of the sound source; D is the distance between the sound source and the in-vehicle camera;

[0034] An occlusion detection unit, if the in-vehicle camera or radar fails to detect a target in the predicted sound source area and the sound source category is a person or other motor vehicle, marks the source as an occluded target.

[0035] As a preferred solution of the present invention, the positioning module includes:

[0036] An azimuth calculation unit: uses the Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm to calculate the time delay difference between the sound detectors;

[0037] A distance estimation unit, based on the sound intensity attenuation model I(d) = I 0 -20log 10 (d) - αd + ò, uses the Particle Swarm Optimization (PSO) algorithm to fit and calculate the sound source distance; where: I(d) represents the signal intensity at distance d; I 0 represents the signal intensity at the reference distance; d represents the distance between the sound source and the detector; α represents the sound intensity attenuation coefficient, which is usually related to the propagation environment and frequency; ò represents the noise term, which is a random error;

[0038] Position correction unit, when the camera / radar detects that the target is not blocked, the multi-source data is fused through the Extended Kalman Filter (EKF) to optimize the position calculation;

[0039] Environmental noise compensation unit, which monitors the environmental noise intensity in real time, dynamically adjusts the noise source term, and the noise model is σ~N(0,σ 2 ), where: σ represents the noise standard deviation; N(0,σ 2 ) represents a Gaussian distribution with a mean of 0 and a variance of σ 2 , representing the environmental noise.

[0040] As a preferred solution of the present invention, the risk assessment module includes:

[0041] Risk classification unit: Set three-level warning thresholds, classify according to the distance d: close d≤5m; medium 5<d≤15m; far d>15m, and calculate the corresponding alarm intensity according to the distance d, using the following weight formula:

[0042]

[0043] where: d 0 =10m, k=0.5m -1 , representing the relationship between the distance and the alarm intensity;

[0044] Trajectory prediction unit, based on the Long Short-Term Memory Network (LSTM) to predict the sound source position in the future T = 2s, and predict the future movement trajectory of the sound source;

[0045] Collision probability calculation unit, using Monte Carlo simulation to generate N = 1000 sets of trajectories, and the collision probability P calculation formula is:

[0046]

[0047] where: x car (t) is the vehicle position, x src (t) is the sound source position, ||·|| represents the distance between two points, and the collision probability reflects the proximity of the sound source to the vehicle;

[0048] Dynamic decision-making unit, if P>70% and the remaining time is less than 2s, trigger an emergency warning.

[0049] As a preferred solution of the present invention, the warning module includes:

[0050] Multi-state prompt unit, which displays the sound source position in a 30° fan-shaped area on the in-vehicle screen and issues a voice prompt at the same time;

[0051] Emergency braking interface: If the collision probability P>90% and the driver's response time is greater than 1s, issue a deceleration command;

[0052] The logging unit saves data with a cycle of 30 days.

[0053] As a preferred solution of the present invention, the self-learning module includes:

[0054] The federated learning unit updates the model through the federated learning framework and optimizes the objective function:

[0055]

[0056] where: θ is the parameter of the model, which needs to be optimized through training; represents the total size of all data sets; L represents the loss function; f(x (i) ; θ) is the output of the i-th client under the global model parameter θ; y (i) is the corresponding label; D i is the sample set of the i-th client;

[0057] The scene adaptation unit adjusts the warning sensitivity according to the area type;

[0058] The model calibration unit optimizes the noise compensation coefficient every 24 hours.

[0059] A vehicle-mounted sound identification and positioning intelligent warning method includes the following steps:

[0060] Step 1: Real-time collect the audio signals of the surrounding environment through multiple sound detectors, and perform spatio-temporal calibration on the signals of multiple detectors to synchronize the data of different devices;

[0061] Step 2: Analyze the collected audio signals, identify the sound source type, and combine the vehicle-mounted camera and radar data to determine whether the sound source is blocked by the front. The sound source types include pedestrians, tricycles, motorcycles, and other motor vehicles;

[0062] Step 3: Calculate the azimuth of the sound source according to the time difference of arrival TDOA of the signals of multiple sound detectors, and calculate the distance of the sound source in combination with the sound intensity attenuation model. The particle swarm optimization PSO algorithm is used to optimize the distance estimation result;

[0063] Step 4: Use the extended Kalman filter EKF to fuse the signal data of the vehicle-mounted camera and radar, correct the position of the sound source and optimize the positioning result, and dynamically adjust the noise compensation coefficient according to the environmental noise intensity;

[0064] Step 5: Evaluate the collision risk based on the spatial position, occlusion situation, object type, and distance of the sound source, and generate warning signals of different levels. When the collision risk exceeds the set threshold and the remaining time is less than 2 seconds, an emergency warning is triggered;

[0065] Step 6: Send a warning message to the driver through the in-vehicle display or voice prompt. If the collision risk meets the preset conditions, trigger an automatic deceleration command;

[0066] Step 7: Update the model through the federated learning framework and dynamically adjust the warning sensitivity and noise compensation coefficient according to the regional type.

[0067] As a preferred solution of the present invention, the sound source recognition step includes:

[0068] Denoise and preprocess the audio signals collected by multiple sound detectors, and extract the Mel Frequency Cepstral Coefficients MFCC and Linear Prediction Coding LPC features of each signal;

[0069] Fuse the extracted audio features, and combine the MFCC and LPC feature vectors by weighted splicing to improve the expression ability of the audio features;

[0070] Use the Deep Residual Network ResNet to classify the fused feature vectors, identify the sound source type, and output the sound source type label, including pedestrians, tricycles, motorcycles, and other motor vehicles;

[0071] Combine the visual signals of the in-vehicle camera and radar to determine whether the sound source object is blocked by the front. If the target cannot be detected in the prediction area, it is determined to be blocked.

[0072] As a preferred solution of the present invention, the positioning step includes:

[0073] Calculate the Time Difference of Arrival TDOA of the signals of multiple sound detectors, and use the Generalized Cross-Correlation Phase Transform GCC-PHAT algorithm to obtain the azimuth information of the sound source;

[0074] According to the sound intensity attenuation model, calculate the distance of the sound source, and optimize the distance estimation result through the Particle Swarm Optimization PSO algorithm;

[0075] Through the Extended Kalman Filter EKF algorithm, fuse the signal data of the in-vehicle camera and radar, correct the sound source position, and optimize the positioning result;

[0076] Adjust the noise compensation coefficient according to the real-time environmental noise intensity to improve the positioning accuracy.

[0077] Compared with the prior art, the beneficial effects of the present invention are as follows: By integrating data from multiple sensors, especially combining sound detection technology with in-vehicle camera and radar data, the performance of the in-vehicle intelligent warning system is significantly improved. The system can collect and analyze audio signals of the surrounding environment in real time, and combine the visual and radar data of other sensors, effectively overcoming the collision detection blind spots caused by limited sensor line of sight or occlusion in traditional technologies. By analyzing the audio signals through a deep learning model, the system can accurately identify and judge the sound source type in complex environments, improving the recognition accuracy and reliability of the system. At the same time, through the precise spatio-temporal calibration of the calibration module, the system can accurately align the audio signals with the in-vehicle camera and radar detection data, ensuring the accurate fusion of multi-sensor data, solving the spatio-temporal synchronization problem between sensor data, and improving the accuracy of sound source position calculation. The positioning module calculates the azimuth of the sound source through the time difference of arrival TDOA of signals from multiple sound detectors, and combines the sound intensity attenuation model and the particle swarm optimization algorithm for distance estimation. The extended Kalman filter further corrects the positioning result to ensure the accuracy of sound source positioning. The system dynamically adjusts the noise compensation coefficient according to the environmental noise intensity, optimizes the noise compensation in real time, reduces the interference of environmental noise on the system recognition and positioning results, and enhances the stability of the system in noisy environments. The risk assessment module generates warning signals of different levels by evaluating the spatial position, occlusion state, object type and distance of the sound source, and timely triggers the automatic deceleration instruction to reduce the risk of collision. Through the federated learning framework, the self-learning module can continuously update the model and dynamically adjust the warning sensitivity and noise compensation coefficient to ensure that the system can adapt to different driving environments and continuously improve its performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0079] Among them:

[0080] Figure 1 is a schematic diagram of the system modular structure of the present invention;

[0081] Figure 2 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.

[0083] As Figure 1 shown, this is an embodiment of the present invention, which provides an in-vehicle sound recognition and positioning intelligent warning system, including:

[0084] (1) Sound detector module

[0085] For real-time collection of audio signals in the surrounding environment;

[0086] In this embodiment, this module includes multiple sound detectors installed at different positions of the in-vehicle system for real-time collection of audio signals in the surrounding environment. The sound detectors can sense various sound signals from outside the vehicle, including the engine sound of the vehicle during driving, and the sounds of traffic participants such as pedestrians, tricycles, and motorcycles. Through a high-sensitivity microphone array, the sound detectors can still effectively capture the surrounding sound signals and transmit them to the subsequent sound recognition module under the conditions of high vehicle speed or noisy external environment.

[0087] (2) Sound recognition module

[0088] For analyzing the collected audio signals through a deep learning model, identifying the sound source type, and combining the in-vehicle camera and radar data to determine whether the sound source is blocked;

[0089] The sound recognition module includes:

[0090] Timbre feature extraction unit, which extracts the feature vectors of the audio signals through the Mel frequency cepstral coefficient MFCC and linear prediction coding LPC fusion algorithm. The fusion algorithm combines the feature vectors of MFCC and LPC in a weighted splicing manner. The specific steps are as follows:

[0091] a. Preprocess the audio signals, including denoising and framing operations;

[0092] b. Use the Mel frequency cepstral coefficient MFCC algorithm to extract the spectral features of the audio signals to obtain the feature vectors:

[0093] MFCC = [mfcc 1 , mfcc 2 ,..., mfcc N ;

[0094] Where: mfccN represents the Nth eigenvalue of the Mel Frequency Cepstral Coefficients; N is the dimension of the MFCC features;

[0095] c. Use the Linear Predictive Coding (LPC) algorithm to extract the predictive coefficient features of the audio signal, obtaining a feature vector:

[0096] LPC = [lpc 1 , lpc 2 ,..., lpc M ;

[0097] where: lpc M represents the Mth eigenvalue of the LPC features; M is the dimension of the LPC features;

[0098] d. Fuse the MFCC and LPC feature vectors through weighted concatenation to obtain a fused feature vector:

[0099] F fusion = [mfcc 1 , mfcc 2 ,..., mfcc N , lpc 1 , lpc 2 ,..., lpc M ;

[0100] where: F fusion is the final fused feature vector, N and M are the dimensions of the MFCC and LPC features respectively, and in the weighted concatenation, the weight coefficients of MFCC and LPC are adjusted according to the application environment or training data;

[0101] A classification unit, which is used to classify the fused audio feature vector based on the Deep Residual Network (ResNet) model and output the sound source type label, including: pedestrians, tricycles, motorcycles, and other motor vehicles. The classification method is through a deep learning model and is adapted to a variable and complex noise environment;

[0102] An occlusion judgment unit, which is used to judge that when the in-vehicle camera or radar fails to detect a target in the predicted sound source area, it is determined that the sound source object is occluded. The occlusion judgment method improves the traditional occlusion detection based on visual sensors and combines the fusion of sound signals and visual signals.

[0103] In this embodiment, the voice recognition module is used to analyze the collected audio signal and identify the type of sound source. This module uses deep learning models (such as convolutional neural network CNN, recurrent neural network RNN, etc.) to extract features and classify the audio signal. First, the audio signal extracts audio features through Mel Frequency Cepstral Coefficients (MFCC) and Linear Predictive Coding (LPC), and then inputs the features into the deep learning model for sound source classification. The types of sound sources include pedestrians, tricycles, motorcycles, and other motor vehicles. The voice recognition module also combines the data of in-vehicle cameras and radars, and through the multi-modal fusion of image and sound data, determines whether the sound source is blocked. For example, in the case where the in-vehicle camera cannot clearly identify, the system can rely on the sound source positioning information provided by the voice recognition module to supplement the deficiency of visual data.

[0104] The core idea of the weighted splicing method is that when splicing the MFCC and LPC feature vectors, considering that the importance of these two features may be different in different driving environments, we introduce a dynamic weighting coefficient to optimize the combination of these two features. The MFCC feature has a good performance in the spectral characteristics of sound, while the LPC feature can capture more detailed changes in speech signals. According to the noise intensity, vehicle speed, and external interference conditions in different environments, dynamically adjust the weight coefficient of each feature, so as to improve the accuracy and precision of sound source recognition.

[0105] MFCC: Mel Frequency Cepstral Coefficients, which simulate the auditory characteristics of the human ear and extract spectral features through steps such as frame division, windowing, FFT, Mel filter bank, and discrete cosine transform.

[0106] LPC: Linear Predictive Coding, which predicts the current value through linear combination of past samples and extracts the channel characteristic parameters.

[0107] (3) Calibration module

[0108] Used to perform precise spatio-temporal calibration on the audio signal and the detection data of in-vehicle cameras and radars, optimize signal alignment through time synchronization and space mapping, and determine whether the sound source object is blocked;

[0109] The calibration module includes:

[0110] Time synchronization unit: Use the Precision Time Protocol PTP (PTP is a high-precision time synchronization protocol used to calibrate the sensor clock) to perform time calibration on each sensor, and the synchronization error Δt ≤ 1ms;

[0111] Space mapping unit: Map the angle of the sound source position to the angle of the in-vehicle camera's field of view, and the mapping formula is:

[0112]

[0113] Where: θ is the horizontal angle of the sound source relative to the vehicle-mounted camera, φ is the vertical angle of the sound source relative to the vehicle-mounted camera, X source and Y source are the lateral and longitudinal coordinates of the sound source; D is the distance between the sound source and the vehicle-mounted camera;

[0114] Occlusion detection unit: If the vehicle-mounted camera or radar fails to detect the target within the predicted area of the sound source, and the sound source category is a person or another motor vehicle, then mark the source as an occluded target.

[0115] In this embodiment, the calibration module is used to perform precise spatio-temporal calibration on the collected audio signal and the detection data of the vehicle-mounted camera and radar. This module uses time synchronization and space mapping technologies to align the data from different sensors, ensuring that the data of each sensor can be processed within the same spatio-temporal framework. Spatio-temporal calibration realizes time synchronization through the Precision Time Protocol (PTP), and at the same time uses a space mapping algorithm to correspond the angle of the sound source position to the field of view angle of the vehicle-mounted camera. The calibration module can accurately determine whether the sound source object is occluded and provide precise spatial coordinate data for subsequent sound source positioning.

[0116] (4) Positioning module

[0117] It is used to calculate the azimuth of the sound source object through the time difference of arrival TDOA of multiple sound detector signals, combine the sound intensity attenuation model and the Particle Swarm Optimization PSO algorithm to calculate the distance of the sound source object, and use the Extended Kalman Filter EKF to fuse the camera and radar data to correct the positioning result. At the same time, the noise compensation coefficient is dynamically adjusted according to the environmental noise intensity;

[0118] The positioning module includes:

[0119] Azimuth calculation unit: It uses the Generalized Cross-Correlation Phase Transform GCC-PHAT algorithm to calculate the time delay difference between sound detectors;

[0120] Distance estimation unit: Based on the sound intensity attenuation model I(d) = I 0 -20log 10 (d)-αd+ò, use the Particle Swarm Optimization PSO algorithm to fit and calculate the sound source distance; where: I(d) represents the signal intensity at distance d; I 0 represents the signal intensity at the reference distance; d represents the distance between the sound source and the detector; α represents the sound intensity attenuation coefficient, which is usually related to the propagation environment and frequency; ò represents the noise term, which is a random error;

[0121] Position correction unit: When the camera / radar detects that the target is not occluded, it uses the Extended Kalman Filter EKF to fuse multi-source data to optimize the position calculation;

[0122] Ambient noise compensation unit, which monitors the ambient noise intensity in real time, dynamically adjusts the noise source term, and the noise model is σ~N(0,σ 2 ), where: σ represents the noise standard deviation; N(0,σ 2 ) represents a Gaussian distribution with a mean of 0 and a variance of σ 2 , representing the ambient noise.

[0123] In this embodiment, the function of the positioning module is to calculate the azimuth of the sound source object through the time difference of arrival (TDOA) of signals from multiple sound detectors. This module uses an acoustic intensity attenuation model combined with a particle swarm optimization (PSO) algorithm to calculate the distance of the sound source. The acoustic intensity attenuation model is used to describe the law of sound attenuation with increasing distance, and the PSO algorithm optimizes the distance estimation result by simulating the search of a particle swarm, thereby achieving more accurate distance measurement. In addition, the positioning module fuses the data of the vehicle-mounted camera and radar through an extended Kalman filter (EKF) algorithm to further correct the position of the sound source and improve the positioning accuracy. According to the ambient noise intensity, the system can dynamically adjust the noise compensation coefficient, optimize the processing effect of the sound signal, and reduce the influence of ambient noise on the positioning accuracy.

[0124] (5) Risk assessment module

[0125] Used to evaluate the collision risk based on the spatial position, occlusion state, object type, and distance of the sound source, and generate warning signals of different levels;

[0126] The risk assessment module includes:

[0127] Risk classification unit: Set three-level warning thresholds, classify according to the distance d: near d≤5m; medium 5<d≤15m; far d>15m, and calculate the corresponding alarm intensity according to the distance d, using the following weight formula:

[0128]

[0129] where: d 0 =10m, k=0.5m -1 , representing the relationship between distance and alarm intensity;

[0130] Trajectory prediction unit, based on the long short-term memory network LSTM, predicts the position of the sound source in the future T = 2s, and predicts the future movement trajectory of the sound source;

[0131] Collision probability calculation unit, using Monte Carlo simulation to generate N = 1000 sets of trajectories, and the collision probability P calculation formula is:

[0132]

[0133] where: x car (t) is the vehicle position, xsrc (t) is the sound source position, ||·|| represents the distance between two points, and the collision probability reflects the proximity of the sound source to the vehicle;

[0134] The dynamic decision-making unit triggers an emergency warning if P > 70% and the remaining time is less than 2 s.

[0135] In this embodiment, the risk assessment module is used to evaluate potential collision risks and generate warning signals of different levels according to the risk degree. This module comprehensively considers factors such as the spatial position of the sound source, occlusion state, object type, and relative distance from the vehicle. The risk assessment module determines the severity of the collision risk by setting different warning thresholds and triggers corresponding warning responses according to the evaluation results. When the collision risk exceeds the set threshold, the system issues a high-risk warning signal. The role of this module is to ensure that the driver can timely understand potential dangers and take corresponding avoidance measures.

[0136] (6) Warning module

[0137] It is used to display the sound source position on the in-vehicle display screen and issue a voice prompt, and trigger an automatic deceleration command when the collision risk meets the preset conditions;

[0138] The warning module includes:

[0139] The multi-state prompt unit displays the sound source position in a 30° fan-shaped area on the in-vehicle screen and issues a voice prompt at the same time;

[0140] Emergency braking interface: If the collision probability P > 90% and the driver response time is greater than 1 s, issue a deceleration command;

[0141] The log recording unit saves data with a cycle of 30 days.

[0142] In this embodiment, the warning module is responsible for displaying the position of the sound source on the in-vehicle display screen and prompting the driver through voice. When the collision risk meets the set conditions, the system triggers an automatic deceleration command to help the vehicle decelerate and avoid danger. The warning module helps the driver understand potential obstacles in the surrounding environment by real-time displaying the position information of the sound source, and provides voice prompts or deceleration reminders at critical moments, thereby reducing the probability of accidents.

[0143] (7) Self-learning module

[0144] It is used to update the model through the federated learning framework and dynamically adjust the warning sensitivity and noise compensation coefficient according to the region type.

[0145] The self-learning module includes:

[0146] The federated learning unit updates the model through the federated learning framework and optimizes the objective function:

[0147]

[0148] Where: θ is a parameter of the model and needs to be optimized through training; represents the total size of all datasets; L represents the loss function; f(x (i) ; θ) is the output of the i-th client under the global model parameter θ; y (i) is the corresponding label; D i is the sample set of the i-th client;

[0149] The scene adaptive unit adjusts the warning sensitivity according to the area type;

[0150] The model calibration unit optimizes the noise compensation coefficient every 24 hours.

[0151] In this embodiment, the self-learning module continuously updates and optimizes the system model through the federated learning framework. This module enables the system to dynamically adjust the warning sensitivity and noise compensation coefficient according to the area type, thereby achieving adaptive learning. The federated learning framework enables different vehicles to share experiences while ensuring data privacy, thus continuously improving the system performance. The addition of the self-learning module enables the system to continuously optimize according to different environments and driving scenarios, improving its performance under various complex driving conditions.

[0152] Federated learning allows multiple clients (vehicles) to collaboratively train and share the model without sharing the original data. In the self-learning module, each vehicle locally updates the model parameters and uploads them to the server for aggregation to improve the global model performance.

[0153] Such as Figure 2 shown, is another embodiment of the present invention. This embodiment provides an in-vehicle sound identification and localization intelligent warning method, including the following steps:

[0154] Step 1: Real-time collect the audio signals of the surrounding environment through multiple sound detectors, and perform spatio-temporal calibration on the signals of multiple detectors to synchronize the data of different devices;

[0155] Step 2: Analyze the collected audio signals, identify the sound source type, and combine the in-vehicle camera and radar data to determine whether the sound source is blocked by the front. The sound source types include pedestrians, tricycles, motorcycles, and other motor vehicles;

[0156] Step 3: Calculate the azimuth of the sound source according to the time difference of arrival TDOA of the signals of multiple sound detectors, and calculate the distance of the sound source in combination with the sound intensity attenuation model. The particle swarm optimization PSO algorithm is used to optimize the distance estimation result;

[0157] Step 4: Fusion of the signal data from the in-vehicle camera and radar through the Extended Kalman Filter (EKF) to correct the sound source position and optimize the positioning result, and dynamically adjust the noise compensation coefficient according to the environmental noise intensity;

[0158] Step 5: Based on the spatial position, occlusion situation, object type, and distance of the sound source, evaluate the collision risk and generate warning signals of different levels. Trigger an emergency warning when the collision risk exceeds the set threshold and the remaining time is less than 2 seconds;

[0159] Step 6: Send warning information to the driver through the in-vehicle display screen or voice prompt. If the collision risk meets the preset conditions, trigger an automatic deceleration command;

[0160] Step 7: Update the model through the federated learning framework and dynamically adjust the warning sensitivity and noise compensation coefficient according to the area type.

[0161] Specifically, the sound source recognition steps include:

[0162] Denoise and preprocess the audio signals collected by multiple sound detectors, and extract the Mel Frequency Cepstral Coefficients (MFCC) and Linear Prediction Coding (LPC) features of each signal;

[0163] Fuse the extracted audio features, and combine the MFCC and LPC feature vectors through weighted splicing to improve the expression ability of the audio features;

[0164] Use the Deep Residual Network (ResNet) to classify the fused feature vectors, identify the sound source type, and output the sound source type label, including pedestrians, tricycles, motorcycles, and other motor vehicles;

[0165] Combine the visual signals of the in-vehicle camera and radar to determine whether the sound source object is blocked by the front. If the target cannot be detected in the prediction area, it is determined to be blocked.

[0166] Among them, the positioning steps include:

[0167] Calculate the Time Difference of Arrival (TDOA) of the signals from multiple sound detectors, and use the Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm to obtain the azimuth information of the sound source;

[0168] According to the sound intensity attenuation model, calculate the distance of the sound source, and optimize the distance estimation result through the Particle Swarm Optimization (PSO) algorithm;

[0169] Through the Extended Kalman Filter (EKF) algorithm, fuse the signal data of the in-vehicle camera and radar, correct the sound source position, and optimize the positioning result;

[0170] Adjust the noise compensation coefficient according to the real-time environmental noise intensity to improve the positioning accuracy.

[0171] In summary, by innovatively integrating sound detection technology with in-vehicle camera and radar data, the present invention proposes a novel intelligent early warning system for in-vehicle sound identification and localization. This system can effectively solve the problem of collision detection blind spots in the prior art due to limited sensor line of sight, occlusion, or environmental noise interference, and significantly improve the response speed and accuracy of the in-vehicle intelligent early warning system in complex environments. Through the fusion of multi-sensor data and the introduction of deep learning models, the system can analyze and locate potential obstacles in real time, evaluate collision risks, and provide effective early warnings in a timely manner. In particular, through the self-learning module and dynamic adjustment mechanism, the system can continuously optimize the early warning sensitivity and noise compensation coefficient to ensure efficient operation in various complex scenarios. This technical solution not only has strong environmental adaptability but also can significantly improve the safety guarantee of drivers, with broad application prospects and commercial value.

[0172] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0173] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed.

[0174] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A vehicle-mounted sound recognition and positioning intelligent warning system, characterized in that: include: Sound detector module, used to collect audio signals from the surrounding environment in real time; The sound recognition module is used to analyze the collected audio signals through a deep learning model, identify the type of sound source, and determine whether the sound source is blocked by combining the on-board camera and radar data; A calibration module, used to perform precise spatiotemporal calibration of the audio signal with the vehicle-mounted camera and radar detection data, optimize signal alignment through time synchronization and spatial mapping, and determine whether the sound source object is blocked; The positioning module is used to calculate the direction of the sound source object through the arrival time difference TDOA of multiple sound detector signals, calculate the distance of the sound source object by combining the sound intensity attenuation model and particle swarm optimization PSO algorithm, and correct the positioning result by fusing the camera and radar data through the extended Kalman filter EKF, and dynamically adjust the noise compensation coefficient according to the intensity of the ambient noise; The risk assessment module is used to assess the collision risk based on the spatial location of the sound source, occlusion status, object type and distance, and generate warning signals of different levels; The warning module is used to display the location of the sound source on the vehicle display screen and issue a voice prompt, triggering an automatic deceleration command when the collision risk meets the preset conditions; The self-learning module is used to update the model through the federated learning framework and dynamically adjust the warning sensitivity and noise compensation coefficient according to the area type.

2. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The sound recognition module comprises: The timbre feature extraction unit extracts the feature vector of the audio signal through the fusion algorithm of Mel frequency cepstral coefficient MFCC and linear predictive coding LPC. The fusion algorithm combines the feature vectors of MFCC and LPC through weighted splicing. The specific steps are as follows: a. Preprocess the audio signal, including denoising and framing operations; b. Use the Mel-frequency cepstral coefficient MFCC algorithm to extract the spectral features of the audio signal and obtain the feature vector: MFCC=[mfcc1,mfcc2,...,mfcc N ]; Among them: mfcc N Represents the Nth eigenvalue of the Mel-frequency cepstral coefficient; N is the dimension of the MFCC feature; c. Use the linear predictive coding (LPC) algorithm to extract the prediction coefficient features of the audio signal and obtain the feature vector: LPC=[lpc1,lpc2,...,lpc M ]; Among them: lpc M Represents the Mth eigenvalue of the LPC feature; M is the dimension of the LPC feature; d. Fuse the MFCC and LPC feature vectors by weighted concatenation to obtain the fused feature vector: F fusion =[mfcc1,mfcc2,...,mfcc N ,lpc1,lpc2,...,lpc M ]; Among them: F fusion is the final fusion feature vector, N and M are the dimensions of MFCC and LPC features respectively. In weighted concatenation, the weight coefficients of MFCC and LPC are adjusted according to the application environment or training data; The classification unit is used to classify the fused audio feature vector based on the deep residual network ResNet model and output the sound source type label, including: pedestrians, tricycles, motorcycles and other motor vehicles. The classification method uses a deep learning model to adapt to variable and complex noise environments; The occlusion judgment unit is used to determine that the sound source object is occluded when the vehicle-mounted camera or radar fails to detect the target in the predicted sound source area. The occlusion judgment method improves the traditional occlusion detection based on visual sensors and combines the fusion of sound signals and visual signals.

3. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The calibration module comprises: Time synchronization unit: Use the precision time protocol PTP to calibrate the time of each sensor, with a synchronization error of Δt≤1ms; Spatial mapping unit: maps the angle of the sound source position to the angle of the vehicle camera's field of view. The mapping formula is: Where: θ is the horizontal angle of the sound source relative to the vehicle camera, φ is the vertical angle of the sound source relative to the vehicle camera, X source and Y source are the lateral and longitudinal coordinates of the sound source; D is the distance between the sound source and the vehicle-mounted camera; The occlusion detection unit marks the source as an occluded target if the vehicle-mounted camera or radar fails to detect a target within the sound source prediction area and the sound source category is a person or other motor vehicle.

4. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The positioning module comprises: The azimuth calculation unit uses the generalized cross-correlation phase transform GCC-PHAT algorithm to calculate the time delay difference between sound detectors; Distance estimation unit, based on the sound intensity attenuation model I(d) = I0-20log 10 (d)-αd+ò, the particle swarm optimization PSO algorithm is used to fit and calculate the sound source distance; where: I(d) represents the signal strength at distance d; I0 represents the signal strength at the reference distance; d represents the distance between the sound source and the detector; α represents the sound intensity attenuation coefficient, which is usually related to the propagation environment and frequency; ò represents the noise term, which is a random error; Position correction unit, when the camera / radar detects that the target is not blocked, it optimizes the position calculation by fusing multi-source data through the extended Kalman filter EKF; The environmental noise compensation unit monitors the environmental noise intensity in real time and dynamically adjusts the noise source term. The noise model is σ~N(0,σ 2 ), where: σ represents the noise standard deviation; N(0,σ 2 ) means the mean is 0 and the variance is σ 2 The Gaussian distribution represents the environmental noise.

5. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The risk assessment module includes: Risk grading unit: set three levels of warning thresholds and classify them according to the distance d: near d≤5m; medium 5<d≤15m; far d>15m, and calculate the corresponding alarm intensity according to the distance d, using the following weight formula: Where: d0 = 10m, k = 0.5m -1 , which represents the relationship between distance and alarm intensity; The trajectory prediction unit predicts the location of the sound source in the future T = 2s based on the long short-term memory network LSTM, and predicts the future movement trajectory of the sound source; The collision probability calculation unit uses Monte Carlo simulation to generate N = 1000 sets of trajectories. The collision probability P calculation formula is: Where: x car (t) is the vehicle position, x src (t) is the location of the sound source, ||·|| represents the distance between the two points, and the collision probability reflects the proximity between the sound source and the vehicle; Dynamic decision-making unit, if P>70% and the remaining time is less than 2s, triggers an emergency warning.

6. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The early warning module comprises: The multi-state prompt unit displays the location of the sound source in a 30° fan-shaped area on the vehicle screen and issues voice prompts at the same time; Emergency braking interface: If the collision probability P>90% and the driver's response time is greater than 1s, a deceleration command is issued; Log recording unit, data storage period is 30 days.

7. The vehicle-mounted sound recognition and positioning intelligent warning system according to claim 1, characterized in that: The self-learning module comprises: The federated learning unit updates the model through the federated learning framework and optimizes the objective function: Among them: θ is the parameter of the model, which needs to be optimized through training; represents the total size of all data sets; L represents the loss function; f(x (i) ; θ) is the output of the i-th client under the global model parameter θ; y (i) is the corresponding label; D i is the sample set of the i-th client; The scene adaptation unit adjusts the warning sensitivity according to the area type; Model calibration unit, optimizes noise compensation coefficients every 24 hours.

8. A vehicle-mounted sound recognition and positioning intelligent warning method according to any one of claims 1-7, characterized in that: The following steps are involved: Step 1: Use multiple sound detectors to collect audio signals from the surrounding environment in real time, and perform spatiotemporal calibration on the signals of multiple detectors to synchronize data from different devices; Step 2: Analyze the collected audio signals, identify the type of sound source, and determine whether the sound source is blocked by the front in combination with the vehicle-mounted camera and radar data. The sound source types include pedestrians, tricycles, motorcycles and other motor vehicles; Step 3: Calculate the direction of the sound source based on the arrival time difference TDOA of multiple sound detector signals, and calculate the distance of the sound source in combination with the sound intensity attenuation model, and use the particle swarm optimization PSO algorithm to optimize the distance estimation result; Step 4: The extended Kalman filter (EKF) is used to fuse the signal data of the vehicle camera and radar, correct the location of the sound source and optimize the positioning result, and dynamically adjust the noise compensation coefficient according to the intensity of the ambient noise; Step 5: Based on the spatial location, occlusion, object type and distance of the sound source, the collision risk is assessed and warning signals of different levels are generated. When the collision risk exceeds the set threshold and the remaining time is less than 2 seconds, an emergency warning is triggered; Step 6: Send a warning message to the driver through the vehicle display or voice prompt. If the collision risk meets the preset conditions, the automatic deceleration command is triggered; Step 7: Update the model through the federated learning framework and dynamically adjust the warning sensitivity and noise compensation coefficient according to the area type.

9. The vehicle-mounted sound recognition and positioning intelligent warning method according to claim 8, characterized in that: The sound source identification step comprises: De-noise and pre-process the audio signals collected by multiple sound detectors, and extract the Mel-frequency cepstral coefficients MFCC and linear predictive coding LPC features of each signal; The extracted audio features are fused, and the MFCC and LPC feature vectors are combined by weighted concatenation to improve the expressiveness of the audio features; Use the deep residual network ResNet to classify the fused feature vectors, identify the sound source type, and output the sound source type label, including pedestrians, tricycles, motorcycles and other motor vehicles; The visual signals from the on-board camera and radar are combined to determine whether the sound source object is blocked in front. If the target cannot be detected in the predicted area, it is judged to be blocked.

10. The vehicle-mounted sound recognition and positioning intelligent warning method according to claim 8, characterized in that: The positioning step comprises: Calculate the time difference of arrival (TDOA) of multiple sound detector signals and use the generalized cross-correlation phase transform (GCC-PHAT) algorithm to obtain the direction information of the sound source; According to the sound intensity attenuation model, the distance of the sound source is calculated, and the distance estimation result is optimized by the particle swarm optimization (PSO) algorithm. The extended Kalman filter (EKF) algorithm is used to fuse the signal data of the vehicle camera and radar, correct the location of the sound source, and optimize the positioning result. According to the real-time environmental noise intensity, the noise compensation coefficient is adjusted to improve positioning accuracy.

Citation Information

Patent Citations

  • Vehicle perception and danger early warning method and system based on environmental sound analysis

    CN114067612A

  • 3D space sound source positioning method and device combined with audio and video signals

    CN118708918A

  • Early warning method and system for obstacles in dead zone in front of crane, crane and medium

    CN119389946A

  • Method And Apparatus for Video Coding Using Improved Cross-Component Linear Model Prediction

    KR1020240043043A

  • File management system with automated folder creation and file classification

    KR1020250112126A

Cited By

  • Vehicle scratch early warning system based on dynamic risk assessment and damage positioning method

    CN120823723A

  • In-vehicle intelligent mosquito repelling method

    CN120910790A

  • 10kV switch cabinet partial discharge remote detection system based on sensor technology

    CN121385564A