Vehicle control method and device, medium, electronic equipment and vehicle
By collecting and processing environmental audio, and using microphone arrays and deep learning models to identify the type, movement trend, and location of proprietary vehicles, the problem of accurately locating and avoiding proprietary vehicles in existing technologies has been solved, enabling intelligent obstacle avoidance and decision support in complex road conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAOMI EV TECH CO LTD
- Filing Date
- 2026-03-04
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are unable to accurately locate and effectively avoid specialized vehicles in complex road conditions, resulting in a lack of critical decision-making information for drivers or intelligent driving systems.
By collecting ambient audio, using microphone arrays and deep learning models to identify and locate the type, movement trend, and orientation of proprietary vehicles, and combining existing audio acquisition devices, the system complexity is reduced and robustness is improved, enabling autonomous control actuators to avoid obstacles.
It enables precise identification and avoidance of private vehicles in complex road conditions, providing effective decision-making information and improving the safety and practicality of drivers and assisted driving modes.
Smart Images

Figure CN122009235A_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of vehicle control technology, and particularly relates to a vehicle control method, a vehicle control device, a computer-readable storage medium, an electronic device, and a vehicle. Background Technology
[0002] The relevant audio processing solutions can only provide extremely crude distance estimates, cannot achieve accurate positioning, and are difficult to provide effective information for drivers or intelligent driving systems to make decisions. Summary of the Invention
[0003] To overcome the problems existing in the related technologies, this disclosure provides a vehicle control method, a vehicle control device, a computer-readable storage medium, an electronic device, and a vehicle.
[0004] According to a first aspect of the present disclosure, a vehicle control method is provided, comprising: The system collects ambient audio during the vehicle's operation and performs identification and localization processing on the ambient audio to obtain the current identification result. The ambient audio includes the alarm prompt sound of a dedicated vehicle, which is a dedicated vehicle that emits the alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle. The actuators of the current vehicle are controlled based on the current identification result.
[0005] Optionally, the step of performing identification and localization processing on the environmental audio to obtain the current identification result includes: The environmental audio is processed to obtain a discrimination result; Based on the discrimination result, the ambient audio is subjected to noise reduction processing to obtain the noise-reduced ambient audio; The noise-reduced environmental audio is then processed for identification and localization to obtain the current identification result.
[0006] Optionally, the step of performing discrimination processing on the environmental audio to obtain a discrimination result includes: Measure the sound pressure level of the ambient audio and determine the distance between the current vehicle and the dedicated vehicle based on the sound pressure level; Based on the distance meeting the preset conditions, the current recognition result of the environmental audio is determined to be valid.
[0007] Optionally, the step of performing noise reduction processing on the ambient audio to obtain the noise-reduced ambient audio includes: The ambient audio is input into a preset deep learning model to obtain the noise-reduced ambient audio.
[0008] Optionally, before inputting the ambient audio into a preset deep learning model, the method further includes: A first audio sample is acquired, and a type recognition process is performed on the first audio sample to obtain a classification result. The first audio sample includes alarm samples and noise samples. The classification result indicates that the first audio sample is the alarm sample or the noise sample. A second audio sample is obtained, and the alarm sample in the second audio sample is determined according to the classification result. The second audio sample is a sample that mixes the alarm sample and the noise sample. A first loss function is determined based on the alarm sample and the second audio sample, and the preset deep learning model is trained using the first loss function.
[0009] Optionally, the step of performing identification and localization processing on the denoised environmental audio to obtain the current identification result includes: The noise-reduced ambient audio is subjected to feature extraction processing to obtain audio features; The audio features are encoded to obtain the current recognition result.
[0010] Optionally, the step of encoding the audio features to obtain the current recognition result includes: The audio features are input to the encoder to output the vehicle type using the encoder's first activation function, and to output the motion trend and the orientation information using the encoder's second activation function.
[0011] Optionally, the step of controlling the actuators of the current vehicle based on the current identification result includes: Obtain other recognition results corresponding to the current recognition result, wherein the other recognition results are recognition results determined at an adjacent time to the current recognition result; The target identification result is determined based on the current identification result and the other identification results, and the actuators of the current vehicle are controlled based on the target identification result.
[0012] Optionally, controlling the actuators of the current vehicle based on the current identification result includes any of the following methods: In the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the first driving mode of the current vehicle, the steering wheel of the current vehicle is controlled to vibrate according to the current recognition result to alert the driver to avoid the private vehicle; In the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0013] Optionally, controlling the actuators of the current vehicle based on the current identification result includes any of the following methods: In the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes based on the current recognition result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current identification result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0014] According to a second aspect of the present disclosure, a vehicle control device is provided, comprising: The audio acquisition module is configured to acquire ambient audio during the current vehicle's driving process, and to perform identification and positioning processing on the ambient audio to obtain a current identification result. The ambient audio includes the alarm prompt sound of a dedicated vehicle, which is a dedicated vehicle that emits the alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle. The component control module is configured to control the actuators of the current vehicle based on the current identification result.
[0015] Optionally, the audio acquisition module is configured as follows: The environmental audio is processed to obtain a discrimination result; Based on the discrimination result, the ambient audio is subjected to noise reduction processing to obtain the noise-reduced ambient audio; The noise-reduced environmental audio is then processed for identification and localization to obtain the current identification result.
[0016] Optionally, the component control module is configured to: In the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the first driving mode of the current vehicle, the steering wheel of the current vehicle is controlled to vibrate according to the current recognition result to alert the driver to avoid the private vehicle; In the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0017] Optionally, the component control module is configured to: In the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes based on the current recognition result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0018] According to a third aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the vehicle control method provided in any of the first aspects of the present disclosure.
[0019] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to execute the executable instructions to implement the vehicle control method provided in any of the first aspects of this disclosure.
[0020] According to a fifth aspect of the present disclosure, a vehicle is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to execute executable instructions stored in the memory to implement the steps of the vehicle control method provided in any of the first aspects of this disclosure.
[0021] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In the exemplary embodiments of this disclosure, the methods and apparatus can acquire ambient audio by utilizing existing or equipped audio acquisition devices in the vehicle (such as vehicle-mounted microphone arrays), effectively reducing system implementation complexity and additional hardware costs, and improving robustness and practicality under complex road conditions. Furthermore, by accurately determining the location and movement trends of the vehicle using ambient audio, a data foundation and theoretical support are provided for avoiding the vehicle. This is an automated and intelligent control method in non-standard scenarios, enabling the self-controlled actuators to avoid the vehicle, and providing effective decision-making information for the driver and assisted driving modes.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0024] Figure 1 The schematic diagram illustrates a flow chart of a vehicle control method according to an exemplary embodiment of the present disclosure; Figure 2 The schematic diagram illustrates a flowchart of a method for identifying and locating environmental audio in an exemplary embodiment of this disclosure; Figure 3 The schematic diagram illustrates a flowchart of a method for discriminating and processing ambient audio in an exemplary embodiment of the present disclosure; Figure 4 The illustration shows a flowchart of a method for training a preset deep learning model in an exemplary embodiment of the present disclosure. Figure 5 The illustration schematically shows a flowchart of a method for identifying and locating noise-reduced ambient audio in an exemplary embodiment of the present disclosure; Figure 6 The schematic diagram illustrates a flowchart of a method for further determining target identification results in an exemplary embodiment of this disclosure; Figure 7 The illustration schematically shows a flow diagram of a method for controlling an actuator in a first driving mode in an exemplary embodiment of the present disclosure; Figure 8 The schematic diagram illustrates a flow chart of a method for controlling an actuator in a second driving mode in an exemplary embodiment of the present disclosure; Figure 9 This schematic diagram illustrates the application of a vehicle control system in an exemplary embodiment of the present disclosure. Figure 10The schematic diagram illustrates the interface of a vehicle control system in an application scenario of an exemplary embodiment of this disclosure; Figure 11 The schematic diagram illustrates the flow chart of the identification and positioning module in an application scenario of an exemplary embodiment of this disclosure; Figure 12 This schematic diagram illustrates the structure of a vehicle control device according to an exemplary embodiment of the present disclosure; Figure 13 This schematic diagram illustrates the structure of another vehicle control device according to an exemplary embodiment of the present disclosure; Figure 14 The schematic diagram illustrates the structure of another vehicle control device in an exemplary embodiment of the present disclosure. Detailed Implementation
[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0026] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0027] Specialized vehicles (such as ambulances) often engage in unusual driving maneuvers, such as driving against traffic, running red lights, or changing lanes, when performing emergency missions. These non-standard scenarios pose significant challenges to existing intelligent driving systems. Since mainstream intelligent driving solutions are mostly optimized for conventional road conditions, their coverage of such high-risk, dynamically changing abnormal behaviors is limited, necessitating specialized identification and processing technologies.
[0028] Against this backdrop, multimodal perception technology exhibits unique advantages: sound modality possesses non-direct detection capabilities, enabling it to penetrate visual obstructions, identify and locate horn sources in advance, and provide effective warnings even when the vehicle is obstructed by other vehicles or obstacles. Therefore, sound modality perception capabilities are becoming an important research direction for improving the safety of intelligent driving in complex scenarios.
[0029] With the development of artificial intelligence (AI) technology, more and more researchers are using audio signal processing techniques to detect police sirens. Traditional methods mainly rely on spectral analysis (such as Mel-Frequency Cepstral Coefficients) combined with machine learning classifiers (such as Support Vector Machines (SVMs)).
[0030] However, traditional audio processing solutions suffer from severe performance degradation in complex urban road noise environments. For example, they have difficulty distinguishing between noises with similar frequencies and sirens, have poor robustness to sudden short-term interference (such as horn sounds), and rely on a large number of rules, resulting in weak generalization ability.
[0031] Furthermore, simple sound pressure level assessment can only provide a very rough distance estimate, and cannot achieve precise location positioning, making it difficult to provide effective information for drivers or intelligent driving systems to make decisions.
[0032] Drawing inspiration from the physiological mechanisms by which human drivers judge the direction and distance of sound sources through hearing, using microphone arrays for sound source localization (SSL) is a more intuitive solution. With the development of smart cockpit technology, microphone arrays are gradually becoming standard equipment in vehicles due to their widespread application in functions such as active noise cancellation (ANC) and voice assistants, providing a hardware foundation for audio modality-based active safety applications.
[0033] Thanks to the rapid development of deep learning technology, AI's ability to recognize and decompose complex audio scenes has been greatly improved, and the use of end-to-end deep learning models to identify and locate sound events has gradually become a research hotspot. However, directly deploying massive deep learning models on vehicle-grade computing platforms still faces multiple constraints, including computing resources, real-time performance, and cost.
[0034] The relevant technologies can usually only provide extremely crude distance estimates, cannot achieve accurate positioning, and cannot provide drivers with key decision-making information such as direction and distance, thus limiting their application value.
[0035] To address the problems existing in related technologies, this disclosure provides a vehicle control method. Figure 1 This is a flowchart illustrating a vehicle control method according to an exemplary embodiment, such as... Figure 1 As shown, the method may include at least the following steps: Step S110. Collect ambient audio during the current vehicle's driving process, and perform identification and localization processing on the ambient audio to obtain the current identification result. The ambient audio includes the alarm prompt sound of the dedicated vehicle. The dedicated vehicle is a special vehicle that issues an alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the location information of the dedicated vehicle relative to the current vehicle.
[0036] Step S120. Control the actuators of the current vehicle based on the current identification result.
[0037] In the exemplary embodiments of this disclosure, ambient audio can be acquired by utilizing existing or equipped audio acquisition devices in the vehicle (such as an onboard microphone array), effectively reducing system implementation complexity and additional hardware costs, and improving robustness and practicality in complex road conditions. Furthermore, by accurately determining the location and movement trends of the vehicle using ambient audio, a data foundation and theoretical support are provided for avoiding the vehicle. This is an automated and intelligent control method in non-standard scenarios, enabling the self-controlled actuators to avoid the vehicle and providing effective decision-making information for the driver and assisted driving modes.
[0038] The following is a detailed explanation of each step in the vehicle control method.
[0039] In step S110, the ambient audio of the current vehicle during its driving process is collected, and the ambient audio is processed for identification and positioning to obtain the current identification result. The ambient audio includes the alarm prompt sound of the dedicated vehicle. The dedicated vehicle is a special vehicle that issues an alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle.
[0040] In an exemplary embodiment of this disclosure, the ambient audio used can be acquired by a microphone system deployed in the vehicle. This system can simultaneously acquire audio signals from multiple channels of the vehicle; for example, it can simultaneously acquire audio signals from eight channels: the front, rear, left, and right sides of the vehicle. It is worth noting that the multi-channel audio acquisition capability utilized in this solution is commonly found in or optionally integrated into the vehicle's existing perception system. By reusing such existing audio acquisition devices (if the vehicle is already equipped with them), this solution effectively reduces reliance on new dedicated sensors, optimizing implementation costs and system integration complexity.
[0041] To train a model capable of recognizing specific vehicle types, determining their movement trends, and performing directional positioning, vehicles recording audio and specific vehicles (including ambulances) with sirens blaring can be arranged to travel together on regular urban roads at speeds ranging from 0 to 60 kilometers per hour in various motion states. Typical motion states include common road scenarios such as both vehicles being stationary, traveling in the same direction in a straight line, traveling in opposite directions, and traveling perpendicularly.
[0042] To enhance the diversity of audio signals, data acquisition can cover various road sections with and without interfering vehicles nearby, and can include diverse weather conditions such as sunny days, strong winds, and rain. During the data recording process used for model building, the recording vehicle can be kept free of other noise interference to ensure the accuracy of the modeling data.
[0043] In an optional embodiment, Figure 2 A flowchart illustrating a method for identifying and locating environmental audio is shown, such as... Figure 2 As shown, the method may include at least the following steps: In step S210, the ambient audio is processed to obtain a discrimination result.
[0044] In an optional embodiment, Figure 3 A flowchart illustrating a method for discriminative processing of ambient audio is shown, such as... Figure 3 As shown, the method may include at least the following steps: in step S310, the sound pressure level of the ambient audio is measured, and the distance between the current vehicle and the dedicated vehicle is determined based on the sound pressure level.
[0045] It is worth noting that since the alarm sounds of dedicated vehicles are mainly signals distributed in the 600–1200Hz frequency band, the corresponding sound pressure level can be measured.
[0046] Furthermore, the distance between the current vehicle and the dedicated vehicle can be calculated using formula (1): (1) in, Sound pressure level (unit: decibel / dB) This indicates the distance between the recording vehicle and the sound source (unit: meters / m). Parameter and The model is determined by regression fitting with a certain confidence level (e.g., 0.8) based on actual collected data (accurately aligning the real-time location, speed, and audio time series of the recording vehicle and the siren playback vehicle through the GPS (Global Positioning System) timestamp synchronization system), ensuring the credibility and accuracy of the model in real-world application scenarios.
[0047] In step S320, based on the distance meeting the preset conditions, the current recognition result of the environmental audio can be determined to be valid.
[0048] The preset condition can be set by a distance threshold, which is used to determine whether the sound source of the collected environmental audio is emitted from within the effective distance range.
[0049] When the distance between the current vehicle and the dedicated vehicle Less than a pre-given distance threshold When the distance between the dedicated vehicle and the current vehicle is such that a situation needs to be considered for avoidance, it can be determined that the collected environmental audio is valid and can be included in the subsequent processing flow.
[0050] When the distance is greater than or equal to the corresponding distance threshold, the current recognition result of the ambient audio can be determined to be invalid.
[0051] When the distance between the current vehicle and the dedicated vehicle Greater than or equal to a pre-defined distance threshold If the vehicle exceeds the preset effective range, the processing flow can be interrupted to save system computing power.
[0052] In this exemplary embodiment, the distance between the two vehicles is determined by the sound pressure level of the ambient audio, and the valid ambient audio is identified based on the distance between the two vehicles. This is equivalent to providing a preprocessing layer for filtering valid ambient audio, achieving the effects of subsequent noise reduction and identification and localization processing of valid ambient audio, as well as terminating the processing of invalid ambient audio. This reduces the occurrence of incorrect control of the actuator due to the processing of invalid ambient audio, ensures the reliability of the actuator control, and also helps to effectively allocate various resources.
[0053] In step S220, based on the discrimination result, the ambient audio is denoised to obtain the denoised ambient audio.
[0054] Once the current recognition result of the environmental audio is determined to be valid, further noise reduction processing can be performed on the valid environmental audio.
[0055] In an optional embodiment, the ambient audio is input into a preset deep learning model to obtain the noise-reduced ambient audio.
[0056] The preset deep learning model can be a pre-trained Transformer-Unet (a deep learning model that combines the architecture of Transformer and U-Net) denoising model, or other deep learning models. This exemplary embodiment does not impose any special limitations on it.
[0057] In this exemplary embodiment, a deep learning model is used to denoise the ambient audio, leveraging the advantages of deep learning models in suppressing complex non-stationary noise, preserving audio details, and improving overall audio quality. This results in better noise reduction and provides high signal-to-noise ratio audio data for subsequent recognition tasks.
[0058] In an optional embodiment, Figure 4 The flowchart illustrates the method for training a pre-defined deep learning model, as shown below. Figure 4 As shown, the method may include at least the following steps: In step S410, a first audio sample is obtained, and the first audio sample is subjected to type recognition processing to obtain a classification result. The first audio sample includes alarm samples and noise samples, and the classification result indicates that the first audio sample is an alarm sample or a noise sample.
[0059] The core task of the deep learning model is to identify and reproduce the siren sound from multi-channel noisy frequencies. Therefore, the model training is divided into two stages. In the first stage, a classification model is used to determine the type of the first audio sample. The input audio sample contains the siren sound and noise from various scenes. The characteristics of the siren and environmental noise in the frequency domain can be learned through the cross-entropy loss function, providing prior knowledge for the second stage.
[0060] In step S420, a second audio sample is obtained, and an alarm sample in the second audio sample is determined according to the classification result. The second audio sample is a sample that combines alarm samples and noise samples.
[0061] The second stage involves inputting the complex spectrum of the mixed signal to reconstruct a clean siren signal from the noisy signal, thereby enhancing the siren signal.
[0062] In step S430, a first loss function is determined based on the alarm sample and the second audio sample, and a preset deep learning model is trained using the first loss function.
[0063] In this stage, the complex MSE Loss (Mean Squared Error Loss) frequency domain loss function can be used to optimize the noise reduction effect. The calculation formula is shown in formula (2): (2) in, The complex spectrum of the pure siren audio collected at the test site. The complex spectrum output by the model. For time frames, For frequency channels, Represents the square of the modulus of the difference between complex numbers. For the current time frame, For the current channel, For time ,aisle The module receives the complex spectrum of the audio. After denoising the multi-channel audio, it provides high signal-to-noise ratio audio data for subsequent recognition tasks.
[0064] In this exemplary embodiment, prior knowledge of alarms and noise is learned using alarm samples and noise samples from the first audio sample. Then, a clean alarm sample corresponding to the second audio sample is reconstructed based on the classification result. This two-stage model training method enhances the alarm signal and makes it easier to reproduce the alarm sound in the presence of environmental noise. Furthermore, the first loss function can further optimize the noise reduction effect.
[0065] Therefore, after training the preset deep learning model, effective environmental audio can be input into the deep learning model to obtain denoised environmental audio.
[0066] In step S230, the noise-reduced ambient audio is identified and located to obtain the current identification result.
[0067] In an optional embodiment, Figure 5 A flowchart illustrating a method for identifying and locating noise-reduced environmental audio is shown, such as... Figure 5 As shown, the method may include at least the following steps: In step S510, feature extraction processing is performed on the noise-reduced environmental audio to obtain audio features.
[0068] Specifically, convolutional neural networks can be used to extract audio features from the denoised environmental audio. The extracted audio features include multi-dimensional acoustic features such as cross-correlation features, spectral shift features, and frequency domain phase difference features.
[0069] In step S520, the audio features are encoded to obtain the current recognition result.
[0070] In an optional embodiment, audio features are input to an encoder to output the vehicle type using a first activation function of the encoder, and to output motion trend and orientation information using a second activation function of the encoder. The vehicle type may include one or more specialized vehicles such as ambulances and roadside assistance vehicles; the motion trend may include stationary, moving away, or approaching; and the orientation information may include forward or backward.
[0071] The encoder can be a shared encoder based on the CED (Consistent Ensemble Distillation for AudioTagging) architecture, followed by parallel inference through three independent task-specific output heads.
[0072] For example, the vehicle type classification head can use the Sigmoid (S-shaped growth curve) activation function for multi-label classification; the motion trend estimation head and the orientation information discrimination head can both use the Softmax (normalized exponential function) activation function for mutually exclusive classification.
[0073] The total loss function during the training phase can be constructed as a weighted sum of the losses of the three tasks, as shown in formula (3): (3) in, α, β, γ Each can be 1 / 3. Furthermore, the loss function for each task can be precisely selected based on the functional characteristics of its output head to ensure the overall optimal performance of the multi-head mechanism, as follows: Type classification loss ( For the multi-label classification characteristic that "a sound event may contain multiple proprietary vehicle alarm sounds", binary cross-entropy loss (BCE Loss) is adopted. The specific calculation method is shown in formula (4):
[0074] in, For the first The category to which each sample belongs; For the first The predicted probability of a sample.
[0075] Location estimation loss ( ) and proximity discriminative loss ( For single-label multi-class classification characteristics of "unique orientation" and "unique proximity state", cross-entropy loss (CE Loss) is used, and its general formula is: ,in, For the cross-entropy loss of any task, For the first The category to which each sample belongs. For the first The loss function measures the predicted probability of each sample. By measuring the difference between the predicted probability distribution and the true label distribution, this loss function effectively outputs the predicted probability of the true class.
[0076] In this exemplary embodiment, a customized activation function can output results in three dimensions—vehicle information, motion trend, and orientation information—in parallel, which not only enables efficient and accurate completion of composite perception tasks but also improves the accuracy of each output result.
[0077] Therefore, the process of filtering effective environmental audio, denoising the effective environmental audio, and identifying and locating the denoised environmental audio to obtain the current identification result provides a solution for multi-dimensional and accurate identification of dedicated vehicles. This provides support and basis for controlling the actuators of the current vehicle, and improves the accuracy and robustness of actuator control.
[0078] In step S120, the actuators of the current vehicle are controlled based on the current identification result.
[0079] In an exemplary embodiment of this disclosure, after determining the current recognition result based on the ambient audio, the actuators of the current vehicle can be controlled based on the current recognition result.
[0080] In an optional embodiment, Figure 6 A flowchart illustrating the method for further determining the target recognition results is shown, such as... Figure 6 As shown, the method may include at least the following steps: In step S610, other recognition results corresponding to the current recognition result are obtained, wherein the other recognition results are recognition results determined at an adjacent time to the current recognition result.
[0081] For example, when the ambient audio is continuously collected every 4 seconds or at other intervals, the process of identification and localization based on the collected ambient audio shown in step S110 can be continuously performed. Therefore, after determining the current identification result, other identification results at adjacent times can also be obtained.
[0082] In step S620, the target identification result is determined based on the current identification result and other identification results, and the actuators of the current vehicle are controlled based on the target identification result.
[0083] Furthermore, a simple moving average (SMA) can be used for post-processing. This method smooths the current recognition result by calculating the arithmetic mean of data points within a fixed-size sliding window, effectively filtering out noise and instantaneous outliers. The specific calculation method is shown in formula (5): (5) in, This is the current predicted value, and N is the window size. N can take any value such as 1, 2, 3, or 4, and this value is related to the distance threshold. For example, the farther the distance, the larger the value of N; the closer the distance, the smaller the value of N.
[0084] In this exemplary embodiment, after determining the current identification result, the execution element may not be controlled. Instead, other acquired identification results are used to smooth the current identification result, thereby controlling the execution element based on the smoothed target identification result. This effectively filters out noise and abnormal current identification results, improving the stability and reliability of the control of the execution element.
[0085] In an optional embodiment, Figure 7 A flowchart illustrating the method for controlling the actuators in the first driving mode is shown, such as... Figure 7 As shown, the method may include at least the following steps: In step S710, in the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention to avoid the private vehicle.
[0086] The first driving mode can be a manual driving mode. In this case, the driver can be alerted to avoid a private vehicle by controlling the actuators of the current vehicle. Therefore, this can be achieved by displaying a warning message on the screen. For example, the warning message could be "An ambulance is approaching from behind, please give way," etc. This exemplary embodiment does not impose any special limitations on this.
[0087] In step S720, in the first driving mode of the current vehicle, the steering wheel vibration of the current vehicle is controlled according to the current recognition result to prompt attention to avoid the private vehicle.
[0088] The first driving mode can be a manual driving mode. In this mode, the driver can be prompted to avoid private vehicles by controlling the actuators of the current vehicle. Therefore, this can be achieved by vibrating the steering wheel of the current vehicle, and the specific vibration frequency and form are not limited.
[0089] In step S730, in the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound according to the current recognition result. The warning sound is used to remind the driver to pay attention to avoid the private vehicle.
[0090] The first driving mode can be a manual driving mode. In this case, the driver can be alerted to avoid a private vehicle by controlling the actuators of the current vehicle. Therefore, this can be achieved by playing a warning sound through the vehicle's speakers. For example, the content of the warning sound could be "An ambulance is approaching from behind, please give way," etc. This exemplary embodiment does not impose any special limitations on this.
[0091] It is worth noting that, based on Figure 6 In the case where the target recognition result is determined in the manner shown, it can also be achieved through the target recognition result. Figure 7The control method of the actuator shown is not limited to the control of the current identification result.
[0092] besides, Figure 7 The control method of the actuator shown can be a single selection or a combination of selections, and this exemplary embodiment does not impose any special limitations on this.
[0093] In this exemplary embodiment, three control methods for the actuators are provided in the first driving mode. The display screen, steering wheel and speaker are used to promptly remind the driver to avoid the private vehicle. The driver can be reminded to avoid the private vehicle in a timely manner through one or more sensory channels such as vision, touch and hearing, which enhances and optimizes the warning effect and effectively provides decision information for the driver and the assisted driving mode.
[0094] Figure 8 A flowchart illustrating the method for controlling the actuators in the second driving mode is shown, such as... Figure 8 As shown, the method may include at least the following steps: In step S810, in the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle.
[0095] The second driving mode can be an intelligent assisted driving mode. In this mode, the vehicle's actuators can be controlled to avoid obstacles from other vehicles; therefore, deceleration can be achieved by controlling the accelerator pedal. For example, deceleration can be achieved by controlling the accelerator or the power switch.
[0096] In step S820, in the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes to avoid the private vehicle based on the current recognition result.
[0097] The second driving mode can be an intelligent assisted driving mode. In this mode, the vehicle's actuators can be controlled to avoid a special vehicle. Therefore, lane changing can be achieved by controlling the vehicle's steering wheel and headlights to make way for a special vehicle in an emergency.
[0098] In step S830, in the second driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle.
[0099] The second driving mode can be an intelligent assisted driving mode. In this case, the driver can be prompted to intervene and avoid a vehicle by controlling the actuators of the current vehicle. Therefore, this can be achieved by displaying a warning message on the screen. For example, the warning message could be "An ambulance is approaching from behind, please give way," etc. This exemplary embodiment does not impose any special limitations on this.
[0100] In step S840, in the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound according to the current identification result. The warning sound is used to remind the driver to pay attention to avoid the private vehicle.
[0101] The second driving mode can be an intelligent assisted driving mode. In this case, the driver can be prompted to intervene and avoid a vehicle by controlling the actuators of the current vehicle. Therefore, this can be achieved by playing a warning sound through the vehicle's speakers. For example, the content of the warning sound could be "An ambulance is approaching from behind, please give way," etc. This exemplary embodiment does not impose any special limitations on this.
[0102] It is worth noting that, based on Figure 6 In the case where the target recognition result is determined in the manner shown, it can also be achieved through the target recognition result. Figure 8 The control method of the actuator shown is not limited to the control of the current identification result.
[0103] besides, Figure 8 The control method of the actuator shown can be a single selection or a combination of selections, and this exemplary embodiment does not impose any special limitations on this.
[0104] In this exemplary embodiment, four control methods for the actuators are provided in the second driving mode: controlling the accelerator pedal to decelerate, controlling the steering wheel and headlights to change lanes, controlling the display screen to show warning information, and controlling the speaker to play warning sounds. This not only supports the automatic actuators to avoid private vehicles, but also provides prompts through both visual and auditory dimensions, facilitating manual control or intervention by the driver. It supports avoidance of private vehicles in various road conditions or scenarios, optimizing the avoidance effect and user experience.
[0105] The vehicle control method in this embodiment will be described in detail below with reference to an application scenario.
[0106] Figure 9 The diagram illustrates the application of a vehicle control system in a given scenario, such as... Figure 9As shown, a lightweight artificial intelligence model (AI model) built using deep learning technology is trained based on multi-channel audio data. This multi-channel audio data can be collected by an onboard external microphone array. The main application scenarios include identifying and locating approaching private vehicles (such as ambulances) in advance in complex road conditions such as intersections and traffic congestion, thereby providing early warnings to drivers or providing decision-making basis for intelligent driving systems to improve road safety and traffic efficiency.
[0107] This system is deployed locally in the vehicle's infotainment system. It can use the audio acquisition device integrated into the vehicle's existing perception system to analyze ambient sound in real time. The recognition results can be used to prompt the type and location of the vehicle through the cockpit interface or voice, or in assisted driving mode, the vehicle control domain can trigger active safety strategies (such as deceleration and avoidance) to give way to the vehicle.
[0108] Figure 10 The diagram illustrates the interface of a vehicle control system in an application scenario, such as... Figure 10 As shown, the vehicle control system includes a signal acquisition and storage module, a pre-threshold judgment module, an audio noise reduction module, an identification and positioning module, a post-processing module, and an execution module. The four core components involved are the pre-threshold judgment model, the audio noise reduction model, the audio feature recognition model, and the judgment post-processing module.
[0109] Among them, the pre-threshold judgment model is a lightweight decision unit based on the mapping relationship between sound pressure level and distance; the audio noise reduction model is a signal preprocessing module based on the Transformer-UNet architecture; the audio feature recognition model is a deep neural network (DNN) model trained on data, and the network structure is based on common structures such as Transformer and Convolutional Neural Networks (CNN); the post-processing module uses a moving average to output the results.
[0110] Specifically, the pre-threshold judgment model can perform real-time calculations and threshold comparisons on specific frequency bands of the input audio signal based on a pre-fitted sound pressure level-distance function. If the sound source is determined to be within a preset effective range and is emitted by a real ambulance, the subsequent deep processing flow is activated. The audio noise reduction model can learn the mapping from noisy frequencies to clean audio, effectively suppressing non-stationary interferences such as wind noise, tire noise and environmental background noise, significantly improving the signal-to-noise ratio of the input signal, and providing high-fidelity audio data for subsequent feature recognition; The audio feature recognition model can receive multi-channel audio signals after noise reduction and use AI models to extract high-order acoustic features to accurately identify the type, movement direction and location of a specific vehicle. The post-processing module uses a moving average method to smooth the recognition results over time, thereby improving the stability and continuity of the output.
[0111] Of the six modules in the system, the signal acquisition and storage module can collect signals from existing external audio acquisition devices in the vehicle and temporarily store them in the vehicle's infotainment signal storage module. This module's function can also be achieved through the equipped audio acquisition device; this exemplary embodiment does not specifically limit its implementation.
[0112] Specifically, the audio data used is collected by a microphone system deployed outside the vehicle, which can simultaneously acquire audio signals from eight channels: front, rear, left, and right. To train a model capable of recognizing specific vehicle types, determining movement trends, and directional positioning, the recording vehicle and specific vehicles (including ambulances) with sirens blaring can be arranged to travel together on regular urban roads at speeds ranging from 0 to 60 km / h in various motion states. Typical motion states include common road scenarios such as both vehicles being stationary, traveling in the same direction in a straight line, traveling in opposite directions, and traveling perpendicularly. To enhance the diversity of audio signals, data collection can cover multiple road sections with and without interfering vehicles nearby, and include various weather conditions such as sunny days, strong winds, and rain. During the data recording process used for model building, the recording vehicle should be kept free of other noise interference to ensure the accuracy of the modeling data.
[0113] The pre-threshold judgment module is a lightweight real-time decision unit based on a physical acoustic model and a negative logarithmic mapping relationship between sound pressure level and distance. The theoretical relationship between sound pressure level and distance is expressed as shown in formula (1). During system operation, this module receives audio input and can measure the sound pressure level of signals in the 600–1200Hz frequency band, where standard dedicated vehicle siren is mainly distributed. When the estimated distance... Less than a pre-given distance At this point, the audio segment enters the subsequent processing flow. The core function of this module is to perform distance threshold judgment: if the sound source exceeds the preset effective range, the processing flow is interrupted to save system computing power; only when the sound source is within the set range are subsequent noise reduction, feature recognition and other deep processing units activated, thereby achieving high-efficiency allocation of system resources and low-power operation.
[0114] The audio denoising module employs a frequency-domain Transformer-UNet denoising model based on Transformer understanding. Its core task is to identify and reproduce siren sounds from multi-channel noisy audio. Model training is divided into two stages: the first stage uses a classification model to determine the type of audio. The input audio samples include siren sounds and noise from multiple scenes. The cross-entropy loss function is used to learn the characteristics of siren and environmental noise in the frequency domain, providing prior knowledge for the second stage. The second stage inputs the complex spectrum of the mixed signal and reconstructs the clean siren signal from the noisy signal to enhance the siren signal. This stage uses the complex MSE Loss frequency domain loss function to optimize the denoising effect. The calculation method is shown in formula (2). After denoising the multi-channel audio, this module provides high signal-to-noise ratio audio data for subsequent recognition tasks.
[0115] The identification and positioning module can achieve synchronous, end-to-end output of three key perception information types: vehicle type, location, and proximity status, through a multi-channel input and multi-head output mechanism.
[0116] Figure 11 The flowchart of the identification and positioning module in the application scenario is shown, such as... Figure 11 As shown, this module is implemented by using convolution to extract multi-dimensional acoustic features, including cross-correlation features, spectral shift features, and frequency domain phase difference features, after inputting multi-channel audio. These features are then input to a shared encoder based on a CED structure, and subsequently, parallel inference is performed through three independent task-specific output heads. Type classification head: Multi-label classification is performed using the Sigmoid activation function.
[0117] Both the orientation estimation head and the proximity discriminator head use the Softmax activation function for mutually exclusive classification.
[0118] Crucially, the total loss function during the training phase is constructed as a weighted sum of the losses of the three tasks, as shown in Equation (3). Furthermore, the loss function of each task is precisely selected based on the functional characteristics of its output head to ensure the overall optimal performance of the multi-head mechanism.
[0119] The post-processing module is implemented to improve the stability of the model output by using a simple moving average (SMA) post-processing method. This method smooths the original prediction sequence by calculating the arithmetic mean of the data points within a fixed-size sliding window, effectively filtering out noise and instantaneous outliers. The specific calculation method is shown in formula (5).
[0120] The execution module receives real-time analysis results from the identification and positioning system. If it determines that a nearby private vehicle is present, it autonomously activates differentiated response strategies based on the current vehicle's driving mode: In manual driving mode, it can promptly remind the driver to avoid the obstacle through a combination of pop-up warnings on the central control screen, speaker voice prompts, and high-frequency micro-vibration tactile alarms on the steering wheel; in intelligent assisted driving mode, it can automatically initiate coordinated control commands to decelerate and smoothly change lanes to a safe lane, making way for emergency vehicles behind. In addition, it can also prompt the driver to intervene through pop-up warnings on the central control screen and speaker voice prompts.
[0121] This solution employs an external microphone array optimized for wideband audio acquisition, enabling stable and reliable sound source identification and localization in all weather conditions. Furthermore, it can not only identify vehicle types but also determine their movement trends and approximate directions, enhancing the perception of specific vehicle behavior and providing drivers or intelligent driving systems with more comprehensive and forward-looking decision-making support.
[0122] Therefore, this solution mainly addresses the limitations of perception modes and high hardware costs of existing proprietary vehicle recognition solutions; it achieves low-cost and highly compatible audio perception by using existing microphone hardware in vehicles or by using equipped audio acquisition devices; and it proposes a pure audio deep learning recognition and positioning system to achieve all-weather, highly reliable early warning.
[0123] In the exemplary embodiments disclosed herein, the technical aspects include sound signal acquisition and storage, pre-threshold judgment, sound event detection, and AI model training and deployment. It provides a comprehensive and accurate identification scheme for common proprietary vehicle types, movement trends, and directions, effectively solving the problems of limited scenarios and high costs caused by reliance on vision or incompatibility with sensors in related technologies. By reusing existing microphone hardware and deploying lightweight models, the robustness and practicality of the system under adverse weather and complex road conditions are greatly improved.
[0124] Furthermore, in an exemplary embodiment of this disclosure, a vehicle control device is also provided. Figure 12 A schematic diagram of the vehicle control device is shown, such as... Figure 12 As shown, the vehicle control device 1200 may include: an audio acquisition module 1210 and a component control module 1220. Wherein: The audio acquisition module 1210 is configured to acquire ambient audio during the current vehicle's driving process, and to perform identification and positioning processing on the ambient audio to obtain a current identification result. The ambient audio includes an alarm prompt sound from a dedicated vehicle, which is a dedicated vehicle that emits the alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle. The component control module 1220 is configured to control the actuators of the current vehicle based on the current identification result.
[0125] In some embodiments of this disclosure, the audio acquisition module 1210 is configured to: The environmental audio is processed to obtain a discrimination result; Based on the discrimination result, the ambient audio is subjected to noise reduction processing to obtain the noise-reduced ambient audio; The noise-reduced environmental audio is then processed for identification and localization to obtain the current identification result.
[0126] In some embodiments of this disclosure, the audio acquisition module 1210 is configured to: Measure the sound pressure level of the ambient audio and determine the distance between the current vehicle and the dedicated vehicle based on the sound pressure level; Based on the distance meeting the preset conditions, the current recognition result of the environmental audio is determined to be valid.
[0127] In some embodiments of this disclosure, the audio acquisition module 1210 is configured to: The ambient audio is input into a preset deep learning model to obtain the noise-reduced ambient audio.
[0128] In some embodiments of this disclosure, the vehicle control device 1200 is further configured to: A first audio sample is acquired, and a type recognition process is performed on the first audio sample to obtain a classification result. The first audio sample includes alarm samples and noise samples. The classification result indicates that the first audio sample is the alarm sample or the noise sample. A second audio sample is obtained, and the alarm sample in the second audio sample is determined according to the classification result. The second audio sample is a sample that mixes the alarm sample and the noise sample. A first loss function is determined based on the alarm sample and the second audio sample, and the preset deep learning model is trained using the first loss function.
[0129] In some embodiments of this disclosure, the audio acquisition module 1210 is configured to: The noise-reduced ambient audio is subjected to feature extraction processing to obtain audio features; The audio features are encoded to obtain the current recognition result.
[0130] In some embodiments of this disclosure, the audio acquisition module 1210 is configured to: The audio features are input to the encoder to output the vehicle type using the encoder's first activation function, and to output the motion trend and the orientation information using the encoder's second activation function.
[0131] In some embodiments of this disclosure, the component control module 1220 is configured to: Obtain other recognition results corresponding to the current recognition result, wherein the other recognition results are recognition results determined at an adjacent time to the current recognition result; The target identification result is determined based on the current identification result and the other identification results, and the actuators of the current vehicle are controlled based on the target identification result.
[0132] In some embodiments of this disclosure, the component control module 1220 is configured to: In the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the first driving mode of the current vehicle, the steering wheel of the current vehicle is controlled to vibrate according to the current recognition result to alert the driver to avoid the private vehicle; In the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0133] In some embodiments of this disclosure, the component control module 1220 is configured to: In the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes based on the current recognition result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current identification result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
[0134] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0135] This disclosure also provides a vehicle, including: processor; Memory used to store processor-executable instructions; The processor is configured to execute executable instructions stored in the memory to implement the steps of any of the vehicle control methods provided in this disclosure.
[0136] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the vehicle control method provided in this disclosure.
[0137] Figure 13 This is a block diagram illustrating another vehicle control device 1300 according to an exemplary embodiment. For example, device 1300 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, in-vehicle infotainment system, etc.
[0138] Reference Figure 13 The device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power supply component 1306, a multimedia component 1308, an audio component 1310, an input / output interface 1312, a sensor component 1314, and a communication component 1316.
[0139] Processing component 1302 typically controls the overall operation of device 1300, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 1302 may include one or more processors 1320 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 1302 may include one or more modules to facilitate interaction between processing component 1302 and other components. For example, processing component 1302 may include a multimedia module to facilitate interaction between multimedia component 1308 and processing component 1302.
[0140] Memory 1304 is configured to store various types of data to support the operation of device 1300. Examples of such data include instructions for any application or method operating on device 1300, contact data, phonebook data, messages, pictures, videos, etc. Memory 1304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0141] Power supply component 1306 provides power to various components of device 1300. Power supply component 1306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 1300.
[0142] Multimedia component 1308 includes a screen that provides an output interface between the device 1300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 1308 includes a front-facing camera and / or a rear-facing camera. When the device 1300 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0143] Audio component 1310 is configured to output and / or input audio signals. For example, audio component 1310 includes a microphone (MIC) configured to receive external audio signals when device 1300 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 1304 or transmitted via communication component 1316. In some embodiments, audio component 1310 also includes a speaker for outputting audio signals.
[0144] Input / output interface 1312 provides an interface between processing component 1302 and peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0145] Sensor assembly 1314 includes one or more sensors for providing status assessments of various aspects of device 1300. For example, sensor assembly 1314 may detect the on / off state of device 1300, the relative positioning of components such as the display and keypad of device 1300, changes in the position of device 1300 or a component of device 1300, the presence or absence of user contact with device 1300, the orientation or acceleration / deceleration of device 1300, and temperature changes of device 1300. Sensor assembly 1314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1314 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1314 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0146] Communication component 1316 is configured to facilitate wired or wireless communication between device 1300 and other devices. Device 1300 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 1316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 1316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0147] In an exemplary embodiment, the apparatus 1300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0148] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1304 including instructions, which can be executed by a processor 1320 of the device 1300 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0149] The aforementioned device can be a standalone electronic device or a part of a standalone electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip, wherein the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and SoC (System on Chip). The aforementioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the aforementioned vehicle control method. The executable instructions can be stored in the integrated circuit or chip or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-described vehicle control method can be implemented; alternatively, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-described method.
[0150] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the vehicle control method described above when executed by the programmable device.
[0151] Figure 14 This is a block diagram illustrating another vehicle control device 1400 according to an exemplary embodiment. For example, device 1400 may be provided as a server. (Refer to...) Figure 14 The device 1400 includes a processing component 1422, which further includes one or more processors, and memory resources represented by memory 1432 for storing instructions, such as application programs, that can be executed by the processing component 1422. The application programs stored in memory 1432 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1422 is configured to execute instructions to perform the vehicle control method described above.
[0152] Device 1400 may also include a power supply component 1426 configured to perform power management of device 1400, a wired or wireless network interface 1450 configured to connect device 1400 to a network, and an input / output interface 1458. Device 1400 can operate on an operating system stored in memory 1432.
[0153] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0154] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A vehicle control method, characterized in that, include: The system collects ambient audio during the vehicle's operation and performs identification and localization processing on the ambient audio to obtain the current identification result. The ambient audio includes the alarm prompt sound of a dedicated vehicle, which is a dedicated vehicle that emits the alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle. The actuators of the current vehicle are controlled based on the current identification result.
2. The vehicle control method according to claim 1, characterized in that, The step of identifying and locating the environmental audio to obtain the current identification result includes: The environmental audio is processed to obtain a discrimination result; Based on the discrimination result, the environmental audio is subjected to noise reduction processing to obtain the noise-reduced environmental audio; The noise-reduced environmental audio is then processed for identification and localization to obtain the current identification result.
3. The vehicle control method according to claim 2, characterized in that, The discrimination process of the environmental audio to obtain the discrimination result includes: Measure the sound pressure level of the ambient audio and determine the distance between the current vehicle and the dedicated vehicle based on the sound pressure level; Based on the distance meeting the preset conditions, the current recognition result of the environmental audio is determined to be valid.
4. The vehicle control method according to claim 2, characterized in that, The step of performing noise reduction processing on the ambient audio to obtain the noise-reduced ambient audio includes: The ambient audio is input into a preset deep learning model to obtain the noise-reduced ambient audio.
5. The vehicle control method according to claim 4, characterized in that, Before inputting the ambient audio into a preset deep learning model, the method further includes: A first audio sample is acquired, and a type recognition process is performed on the first audio sample to obtain a classification result. The first audio sample includes alarm samples and noise samples. The classification result indicates that the first audio sample is the alarm sample or the noise sample. A second audio sample is obtained, and the alarm sample in the second audio sample is determined according to the classification result. The second audio sample is a sample that mixes the alarm sample and the noise sample. A first loss function is determined based on the alarm sample and the second audio sample, and the preset deep learning model is trained using the first loss function.
6. The vehicle control method according to claim 2, characterized in that, The step of identifying and locating the noise-reduced environmental audio to obtain the current identification result includes: The noise-reduced ambient audio is subjected to feature extraction processing to obtain audio features; The audio features are encoded to obtain the current recognition result.
7. The vehicle control method according to claim 6, characterized in that, The process of encoding the audio features to obtain the current recognition result includes: The audio features are input to the encoder to output the vehicle type using the encoder's first activation function, and to output the motion trend and the orientation information using the encoder's second activation function.
8. The vehicle control method according to claim 1, characterized in that, The actuation element that controls the current vehicle based on the current identification result includes: Obtain other recognition results corresponding to the current recognition result, wherein the other recognition results are recognition results determined at an adjacent time to the current recognition result; The target identification result is determined based on the current identification result and the other identification results, and the actuators of the current vehicle are controlled based on the target identification result.
9. The vehicle control method according to claim 1, characterized in that, The method of controlling the actuators of the current vehicle based on the current identification result includes any of the following: In the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the first driving mode of the current vehicle, the steering wheel of the current vehicle is controlled to vibrate according to the current recognition result to alert the driver to avoid the private vehicle; In the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
10. The vehicle control method according to claim 1, characterized in that, The method of controlling the actuators of the current vehicle based on the current identification result includes any of the following: In the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes based on the current recognition result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current identification result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
11. A vehicle control device, characterized in that, include: The audio acquisition module is configured to acquire ambient audio during the current vehicle's driving process, and to perform identification and positioning processing on the ambient audio to obtain a current identification result. The ambient audio includes the alarm prompt sound of a dedicated vehicle, which is a dedicated vehicle that emits the alarm prompt sound when performing a specific task. The current identification result includes one or more of the following: the vehicle type of the dedicated vehicle, the movement trend of the dedicated vehicle, and the orientation information of the dedicated vehicle relative to the current vehicle. The component control module is configured to control the actuators of the current vehicle based on the current identification result.
12. The vehicle control device according to claim 11, characterized in that, The audio acquisition module is configured as follows: The environmental audio is processed to obtain a discrimination result; Based on the discrimination result, the environmental audio is subjected to noise reduction processing to obtain the noise-reduced environmental audio; The noise-reduced environmental audio is then processed for identification and localization to obtain the current identification result.
13. The vehicle control device according to claim 11, characterized in that, The component control module is configured as follows: In the first driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current recognition result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the first driving mode of the current vehicle, the steering wheel of the current vehicle is controlled to vibrate according to the current recognition result to alert the driver to avoid the private vehicle; In the first driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
14. The vehicle control device according to claim 11, characterized in that, The component control module is configured as follows: In the second driving mode of the current vehicle, the accelerator pedal of the current vehicle is controlled to decelerate according to the current identification result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the steering wheel and headlights of the current vehicle are controlled to change lanes based on the current recognition result in order to avoid the private vehicle; In the second driving mode of the current vehicle, the display screen of the current vehicle is controlled to display a warning message based on the current identification result. The warning message is used to remind the driver to pay attention and avoid the private vehicle. In the second driving mode of the current vehicle, the speaker of the current vehicle is controlled to play a warning sound based on the current identification result. The warning sound is used to remind the driver to pay attention and avoid the private vehicle.
15. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1 to 10.
16. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1 to 10.
17. A vehicle, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute executable instructions stored in the memory to implement the steps of the vehicle control method according to any one of claims 1 to 10.