In-vehicle sound effect equalization method and device, storage medium and program product
By acquiring the gain difference spectrum and performing track-by-track processing on the media audio, and making fine gain adjustments for each independent track, the problem of sound quality loss in in-vehicle noise environments is solved, achieving clear, natural, and balanced sound effects in complex noise environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies struggle to achieve precise gain adjustment for media audio in noisy in-vehicle environments, resulting in sound quality degradation or insufficient handling of masking effects in certain frequency bands. Consequently, they are unable to achieve clear, natural, and content-appropriate listening optimization in complex noisy environments.
By acquiring the gain difference spectrum between the background noise signal in the vehicle and the reference noise signal in a quiet environment, the media audio signal is processed into tracks. For each independent track, the gain adjustment amount is determined from the gain difference spectrum according to its frequency range, and the corresponding gain adjustment is applied under the condition that it is satisfied. Finally, the output signals are merged.
It achieves differentiated enhancement of different types of sound elements in complex noise environments, improves speech clarity and the naturalness of music listening, and significantly improves the accuracy and naturalness of sound effect equalization.
Smart Images

Figure CN122294046A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to one or more embodiments in the field of vehicle technology, and more particularly to an in-vehicle sound equalization method, device, storage medium, and program product. Background Technology
[0002] With the increasing popularity of automobiles and the upgrading of driving experience, in-vehicle noise management has become a key challenge for improving comfort and safety. The background sound pressure levels generated during vehicle operation, such as engine noise, wind noise, road noise, and tire noise, are constantly changing, often leading to a mismatch between multimedia audio volume and the environment. Relying on manual adjustments by the driver not only distracts them but also increases driving risks. Therefore, there is an urgent need for an intelligent sound equalization solution that can automatically adapt to the noise environment without human intervention.
[0003] In related technologies, gain adjustment is typically performed by estimating the overall amplitude or spectrum of in-vehicle noise. One type of method maps the overall gain based on a lookup table of noise amplitude, while another detects whether each frequency point in the noise spectrum exceeds a threshold and performs gain compensation for the corresponding frequency band. However, these methods often fail to fully consider the differences in the frequency domain structure of noise and the acoustic characteristics of the audio content, easily leading to coarse adjustments, degraded sound quality, or insufficient handling of masking effects in certain frequency bands. They still struggle to achieve clear, natural, and content-appropriate listening optimization in complex noise environments. Summary of the Invention
[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, an in-vehicle sound equalization method is proposed, the method comprising: Acquire the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment; The currently playing media audio signal is processed into multiple independent audio tracks corresponding to different frequency ranges. For each independent audio track, the corresponding gain adjustment amount is determined from the gain difference spectrum according to its frequency range, and the corresponding gain adjustment is applied to the independent audio track when the gain adjustment amount meets the gain condition. The individual audio tracks are merged to output the processed media audio signal.
[0005] According to a second aspect of one or more embodiments of this specification, an in-vehicle sound equalization device is provided, the device comprising: The gain difference spectrum determination unit is used to acquire the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment. The track splitting unit is used to split the currently playing media audio signal into multiple independent audio tracks corresponding to different frequency ranges. The gain adjustment unit is used to determine the corresponding gain adjustment amount from the gain difference spectrum according to the frequency range of each independent audio track, and to apply the corresponding gain adjustment to the independent audio track when the gain adjustment amount meets the gain condition. The audio track merging unit is used to merge the signals of each independent audio track so as to output the processed media audio signal.
[0006] According to a third aspect of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0007] According to a fourth aspect of this specification, a computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0008] As can be seen from the above embodiments, this specification achieves fine gain adjustment based on the noise spectrum structure and matching the acoustic characteristics of the audio content by obtaining the gain difference spectrum characterizing the noise frequency domain characteristics and performing track-by-track processing on the media audio. For each independent audio track, the adjustment amount is extracted and determined according to the gain difference spectrum within its corresponding frequency range. This avoids the problems of coarse adjustment and sound quality loss caused by relying solely on the overall noise amplitude or simple frequency threshold judgment in the prior art. It can perform differentiated and adaptive enhancement of different types of sound elements in complex noise environments, thereby significantly improving the accuracy of sound effect equalization and the naturalness of the listening experience while improving speech clarity and music listening experience. Attached Figure Description
[0009] Figure 1 This is a diagram illustrating the architecture of an in-vehicle sound equalization system according to the embodiments disclosed in this specification; Figure 2 This is a flowchart illustrating an in-vehicle sound equalization method according to an embodiment disclosed in this specification; Figure 3 This is a flowchart illustrating another in-vehicle sound equalization method according to the embodiments disclosed in this specification; Figure 4 This is a schematic structural diagram of an electronic device shown in the embodiments of this specification; Figure 5 This is a block diagram illustrating an in-vehicle sound equalization device as shown in an embodiment of this specification. Detailed Implementation
[0010] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification.
[0011] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0012] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0013] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0014] Figure 1 This is a schematic diagram of the architecture of an in-vehicle sound equalization system provided in an exemplary embodiment. For example... Figure 1 As shown, the system may include at least an audio receiving device 12, a sound equalizer 14, and an audio playback device 16 within the vehicle 10.
[0015] The audio receiving device 12 may be one or more microphone arrays installed within the vehicle 10, used to acquire raw audio signals from the in-vehicle environment in real time. This device is located in a suitable location within the passenger compartment, such as the center of the ceiling or the dashboard, to effectively capture the mixed signal containing background noise and speaker playback, and transmit it to a subsequent processing unit for analysis.
[0016] The equalizer 14, such as a dedicated audio processing chip integrated into an in-vehicle infotainment system, an in-vehicle electronic control unit (ECU), or a cloud server connected to the vehicle, is configured to perform the core processing flow of the aforementioned equalizer method. This device receives signals from the audio receiving device 12, analyzes the playing media audio track by track, extracts and determines the gain adjustment amount from the gain difference spectrum for each independent track based on its frequency range, applies the corresponding gain adjustment to tracks that meet the conditions, and finally merges the processed tracks into an audio stream to be output.
[0017] An audio playback device 16, such as the speaker system built into the vehicle 10, is connected to the aforementioned equalizer 14 and is used to receive the processed media audio signal and convert it into sound waves for playback. This device can reproduce the audio content, which has been adaptively adjusted according to ambient noise, in the in-vehicle sound field based on the final signal output by the equalizer 14, thereby achieving intelligent matching for different noise scenarios.
[0018] Figure 2 This is a flowchart illustrating an exemplary embodiment of an in-vehicle sound equalization method. Figure 2 As shown, the method may include the following steps: Step S202: Obtain the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment.
[0019] The background noise signal of the current in-vehicle environment can be acquired by the vehicle's built-in audio acquisition device. This signal is then compared and analyzed with a pre-stored or defined quiet environment reference noise signal to calculate and determine a gain difference spectrum that characterizes the frequency domain energy difference between the two. This spectrum reflects the gain compensation required for the current noise environment at different frequency components relative to an ideal quiet state.
[0020] Next, in order to accurately acquire the background noise signal, the solution in this manual can first acquire the original audio signal that can truly reflect the acoustic environment inside the vehicle, and then process it in a targeted manner to separate the pure background noise component.
[0021] In one embodiment, the system can acquire the mixed raw audio signal inside the vehicle in real time using a microphone array or other audio sensing devices arranged in the vehicle cabin. This signal typically includes background noise generated during vehicle operation, such as engine noise, wind noise, and road noise, as well as media audio played by the vehicle's speaker system, such as music and navigation voice. Echo cancellation processing is then applied to the raw audio signal, using techniques such as adaptive filtering and reference signal comparison to identify and cancel the acoustic components originating from the vehicle's own media playback. This removes interference from the media audio components from the mixed signal, achieving an accurate assessment of the noise environment. After this processing, the remaining signal is primarily the background noise signal corresponding to the current driving state and environment, providing a clean and reliable input for subsequent noise analysis and gain difference spectrum calculation.
[0022] After initially acquiring the background noise signal, further analysis and purification can be performed to improve the accuracy of noise estimation and enhance the system's applicability in complex scenarios. This manual provides a method for identifying the state of the in-vehicle acoustic environment using Voice Activity Detection (VAD) and adopting differentiated processing paths based on different states. The specific steps are as follows: First, the background noise signal initially acquired above is subjected to voice activity detection to analyze whether the signal contains valid human voice components, thereby determining whether the vehicle is in a dialogue scenario with active voices, such as passengers talking or having a phone call. This detection can be achieved by analyzing parameters such as the time-frequency characteristics, energy changes, and zero-crossing rate of the signal, but this specification does not impose any limitations on this.
[0023] If the detection results indicate that the in-vehicle environment is not in a dialogue scenario, it means that the main interference in the current background noise signal is continuous or impulsive noise. To further optimize the noise estimation used for subsequent calculations, non-steady-state noise cancellation processing can be performed on the signal. This suppresses or removes transient, sudden noise components such as sudden honking or window knocking, and obtains a cleaner background noise estimate that is more representative of steady-state environmental noise, which can then be used for subsequent calculations of the gain difference spectrum.
[0024] Conversely, if the detection results indicate that the in-vehicle environment is a conversational scenario, it means that the current passenger's voice interaction is the primary acoustic event. To avoid interference with call clarity or causing auditory discomfort due to audio equalization processing, the system can pause all subsequent audio equalization adjustments in this scenario. That is, it will not apply any gain adjustment based on noise estimation to the media audio signal in this scenario, but will instead directly output the original media audio signal to prioritize the quality and naturalness of in-vehicle voice communication.
[0025] Next, the acquired background noise signal and the reference noise signal can be converted into comparable frequency domain representations, and the above-mentioned gain difference spectrum can be obtained by difference calculation.
[0026] In one embodiment, the acquired background noise signal can be frequency-domain transformed, for example, by a Fast Fourier Transform (FFT), converting it from the time domain to the frequency domain to obtain the corresponding background noise amplitude spectrum. This amplitude spectrum reflects the energy distribution of the current noise environment at various frequency points. Subsequently, a reference noise amplitude spectrum corresponding to the reference noise signal of the aforementioned quiet environment can be obtained. This reference spectrum can represent the noise background pre-collected and stored by the vehicle in a stationary, quiet environment, serving as an ideal noise reference. Finally, the current background noise amplitude spectrum can be subtracted from the reference noise amplitude spectrum point by point at each frequency point. This subtraction operation directly quantifies the energy increment of the current noise environment relative to the reference in the quiet environment (hereinafter referred to as the quiet reference) at each frequency band, and the subtraction result constitutes the aforementioned gain difference spectrum. This difference spectrum clearly indicates the amount of gain adjustment required at different frequencies to compensate for changes in environmental noise, and is usually expressed in logarithmic units such as decibels, thus providing a direct and quantitative basis for subsequent fine-grained gain control by frequency band.
[0027] After obtaining the gain difference spectrum through spectral subtraction, the spectrum can be smoothed to achieve a more gradual and reliable smoothed gain difference spectrum. This further enhances its stability and continuity, reducing the potential impact of transient noise fluctuations or estimation errors on subsequent gain adjustments. Specifically, the smoothing process employs a time-recursive smoothing algorithm. Its core principle is to weightedly fuse the gain difference spectrum of the current frame with the smoothing result of the previous frame, effectively suppressing short-term random fluctuations and enabling the output spectrum to more robustly reflect the trend changes in the noise environment.
[0028] The above smoothing process can be further illustrated by the following formula example: in, This represents the gain difference spectrum of the current frame, which is the unsmoothed difference spectrum obtained directly by subtracting the current background noise amplitude spectrum from the reference noise amplitude spectrum. This represents the smoothed gain difference spectrum of the previous frame, which is the result obtained after the previous frame has undergone the same smoothing process. It is used as historical information in the smoothing calculation of the current frame. This indicates the smoothed gain difference spectrum obtained after smoothing the current frame, which will be used for subsequent processing. Known as the smoothing coefficient, it is a configurable constant, typically ranging from 0 to 1. The smoothing coefficient α determines the weight distribution between the current frame spectrum and the historical spectrum during the fusion process: the larger the α value, the more sensitive the smoothing result is to the instantaneous changes of the current frame; the smaller the α value, the more the smoothing result depends on historical information, the stronger the smoothing effect, and the slower the response to changes. By adjusting α, the system's ability to track noise changes and the stability of the estimation results can be balanced.
[0029] Step S204: Perform track splitting on the currently playing media audio signal to obtain multiple independent audio tracks corresponding to different frequency ranges.
[0030] Audio track splitting technology can be used to separate mixed media audio signals, such as music or podcasts played in a car, into multiple independent tracks. These tracks can include vocal tracks, bass drum tracks, bass tracks, piano tracks, guitar tracks, and other merged signal tracks that have not been further separated from the original song. Each independent track corresponds to a specific type of sound element or instrument and has its dominant or core frequency range.
[0031] For example, the core frequency band of human vocals is typically between 200Hz and 3kHz, the core frequency band of a bass drum is concentrated between 60Hz and 100Hz, the core frequency band of a bass is approximately 80Hz to 200Hz, and the core frequency band of a piano, due to its wide range, can cover approximately 27.5Hz to 4kHz, encompassing the low, mid, and high registers, while the core frequency band of a guitar is roughly in the range of 100Hz to 500Hz. Those skilled in the art will understand that the core frequency bands corresponding to different tracks can be set and adjusted according to the actual audio content, instrument characteristics, or auditory optimization goals. The core purpose is to provide an independent signal basis for subsequent targeted and differentiated gain adjustments for different sound elements.
[0032] In this specification, after obtaining the aforementioned gain difference spectrum, before track splitting, an effective frame determination and noise spectrum replacement mechanism can be further introduced to improve the reliability of the system's judgment of environmental noise status and the rationality of sound effect processing decisions. This mechanism first assesses the significance of the current gain difference spectrum, and then decides whether to activate the backup noise spectrum based on the assessment results, to ensure that the noise information used in subsequent sound effect equalization processing has sufficient validity and reliability.
[0033] In one embodiment, the amplitude corresponding to each frequency point in the gain difference spectrum is first compared with a preset amplitude threshold. Then, the total number of frequency points whose amplitude exceeds the threshold is counted. At this point, the counted number of out-of-limit frequency points is compared with another preset quantity threshold. If the number of out-of-limit frequency points is greater than the quantity threshold, the current frame is determined to be a valid frame, indicating a significant and widespread frequency domain difference between the current noise environment and a quiet baseline. The system will continue to execute subsequent track-separation and gain adjustment processes for the media audio signal, which will not be detailed here. Conversely, if the number of out-of-limit frequency points does not exceed the quantity threshold, the current frame is determined to be an invalid frame, indicating that the current noise difference is not significant or that the noise estimation may be abnormal.
[0034] For frames deemed invalid, the system can further check whether the gain difference spectrum meets a preset noise replacement condition. This condition can be used to identify scenarios where the current noise estimation may be unreliable. For example, in one embodiment, the condition can be set to determine whether the average amplitude of all frequency points of the gain difference spectrum is less than 0 dB. If the replacement condition is met, for example, if the average amplitude is less than 0 dB, it indicates that the average energy of the currently estimated noise spectrum is even lower than the quiet reference. In this case, the system replaces the current background noise amplitude spectrum used to calculate the gain difference spectrum with the aforementioned reference noise amplitude spectrum.
[0035] For example, the above decision logic can be clearly defined using the following formula: in, This represents the total number of frequency points whose amplitude exceeds the preset amplitude threshold C. This indicates a preset quantity threshold. When... Greater than If the current frame is valid, it is determined to be a valid frame; otherwise, it is determined to be an invalid frame. This formalized judgment criterion provides the system with a clear and executable basis for state transitions, ensuring the certainty and consistency of subsequent processing decisions.
[0036] It is worth mentioning that the above replacement operation is equivalent to reverting to a known, stable, quiet environmental noise benchmark when the noise estimation reliability is low. This avoids inappropriate gain adjustments when noise information is unreliable, thus enhancing the robustness of the system. If the replacement conditions are not met, the system can maintain the current noise spectrum or adopt other preset strategies.
[0037] Of course, in addition to the average amplitude condition mentioned above, the noise replacement conditions may also include one or more combinations of the following: the signal-to-noise ratio of the current background noise signal is lower than a preset threshold; the difference between the current gain difference spectrum and the historical frame spectrum exceeds a sudden change threshold; the spectral similarity between the current background noise amplitude spectrum and the reference noise amplitude spectrum is lower than a preset similarity threshold; or the temporal energy of the current frame is lower than an absolute energy threshold. These conditions can be used individually or in combination to more comprehensively identify abnormal noise estimation scenarios, thereby further ensuring the reliability of the system's decision-making in complex environments.
[0038] Step S206: For each independent audio track, determine the corresponding gain adjustment amount from the gain difference spectrum according to its frequency range, and apply the corresponding gain adjustment to the independent audio track if the gain adjustment amount meets the gain condition.
[0039] For each of the separated individual audio tracks, a representative gain adjustment amount within that frequency band can be extracted from the gain difference spectrum based on its corresponding core frequency range. The system in this manual has a preset gain condition; only when the gain adjustment amount corresponding to the audio track meets this condition will a corresponding gain boost be applied to that individual audio track, thereby ensuring the necessity and specificity of the gain adjustment.
[0040] To correlate the gain difference spectrum, which reflects noise variations across the entire frequency band, with the acoustic characteristics of each track, a representative gain adjustment needs to be extracted for each individual track from its corresponding core frequency range. This process is a crucial step connecting noise analysis with specific track gain adjustments.
[0041] Specifically, firstly, for each individual audio track, its corresponding core frequency range can be determined. As mentioned earlier, different sound elements have their dominant energy distribution frequency bands; for example, human voices are mainly concentrated in the 200Hz-3kHz range. The system can preset or dynamically assign a core frequency range for each type of audio track. Then, from the aforementioned gain difference spectrum, the statistical values of the amplitudes of each frequency point within the core frequency range are extracted as the gain adjustment amount corresponding to that individual audio track. Specifically, the system locates all frequency points in the gain difference spectrum that belong to the core frequency range of the audio track, and then calculates a certain statistical characteristic of the amplitudes of these frequency points.
[0042] In one embodiment, the statistical value can be the arithmetic mean of the amplitudes at these frequency points, reflecting the overall gain requirement of the frequency band; in other embodiments, the median, weighted average, or other statistical measures that can represent the central trend can also be used. In this way, a broad spectrum difference information is transformed into a single gain adjustment recommendation value for a specific audio track, thus providing a direct basis for subsequent conditional judgments and gain application.
[0043] In addition, the system introduces preset gain conditions as adjustment trigger thresholds to ensure that gain adjustment is performed only when necessary, avoiding unnecessary amplification of audio tracks with insignificant noise masking effects or sufficient loudness.
[0044] Specifically, for each individual audio track, the system can compare its corresponding gain adjustment with a preset gain threshold. This gain threshold can be set independently based on auditory psychoacoustics, track type, or user preference. For example, different thresholds can be set for vocal tracks and bass tracks.
[0045] If the aforementioned gain adjustment is greater than the gain threshold, it can be determined that the noise masking of the current audio track in its core frequency band is relatively significant, and gain adjustment is necessary. At this time, the system applies a gain to the independent audio track corresponding to the aforementioned gain adjustment. Here, "corresponding" means that the applied gain value is based on or equal to the gain adjustment, thereby achieving precise compensation.
[0046] Conversely, if the gain adjustment is less than or equal to the gain threshold, it indicates that the required compensation for the track is not significant or is within an acceptable range under the current environment. In this case, the system does not apply any gain to the independent track, maintaining its original level. This conditional judgment mechanism ensures the targeted and economical nature of gain adjustment, effectively enhancing key sound elements masked by noise while avoiding unnecessary increases in overall volume or potential distortion.
[0047] Finally, the system incorporates an overall gain processing mechanism. This mechanism performs coordinated signal processing before determining the gain adjustment amounts for each track and during the final output stage. This pre-compensates for the overall masking effect caused by environmental noise before fine-tuning the track-by-track gain, ensuring that the final output signal has a reasonable overall loudness.
[0048] In one embodiment, during the process of determining the corresponding gain adjustment amount from the gain difference spectrum based on the frequency range, the system can first determine an overall gain value based on the minimum gain value of all frequency points in the gain difference spectrum. This minimum value represents the minimum gain compensation required across the entire frequency band for the current noise environment relative to the aforementioned quiet baseline. After determining the overall gain value, the system subtracts this overall gain value from the gain value of each frequency point in the gain difference spectrum, thereby obtaining an overall down-shifted, adjusted gain difference spectrum. This operation is equivalent to removing the noise-induced floor masking, allowing the subsequent gain adjustment amounts for each individual track determined based on this adjusted gain difference spectrum to focus more on the additional gain requirements of each track's core frequency band relative to the floor, which is beneficial for more refined differential control.
[0049] Subsequently, in the signal synthesis output stage shown in the next step S208, i.e., when outputting as the processed media audio signal, the system needs to apply the aforementioned overall gain value to the merged signal to generate the final output signal. After merging all signals with individually applied track gains into a complete audio stream, the system applies the previously determined overall gain value to the merged signal. This operation restores the subtracted "baseline" compensation, thereby ensuring that the final output audio matches the ambient noise in overall loudness, while preserving the clarity and balance improvements brought about by the fine adjustments within each track. Through this subtraction-then-addition overall gain processing, the decoupling and synergy between overall loudness compensation and local detail enhancement are achieved.
[0050] To achieve accurate calculation and coordinated processing of the overall gain, the specific steps and mathematical definitions are detailed below. The process first determines a reference compensation value applicable across the entire frequency band, i.e., the overall gain value. Then, this value is used to adjust the overall gain difference spectrum. Finally, the reference compensation is restored in the output stage.
[0051] At this point, the system can first iterate through the gain difference spectrum in the above example. For all frequency points, find the smallest gain difference, denoted as . This value represents the minimum compensation required relative to a quiet baseline across the entire frequency range for the current noise environment. The system then determines... Is it greater than 0dB? If it is greater than 0dB, then adjust the overall gain value. Set as If it is not greater than 0dB, then Set to 0dB. This rule ensures that the overall gain is always non-negative, and reference compensation is only applied when the overall ambient noise is higher than the quiet reference. OK Then, the system performs a spectral shift operation: smoothing the gain difference spectrum. Subtract each frequency value The adjusted gain difference spectrum was obtained. This operation removes the noise "floor" mask, allowing subsequent gain adjustments for each track to focus more on the additional needs of its core frequency bands relative to this floor. Finally, after merging all the individual track signals, the system applies a gain adjustment to the merged signal. This is done to restore the overall reference loudness. The process of determining the overall gain value and adjusting the gain difference spectrum can be defined by the following formula: in, That is, smoothing the gain difference spectrum At all frequencies The minimum value on. And, = - S′′ represents the adjusted gain difference spectrum, which will serve as the direct basis for extracting the gain adjustment amount of each independent audio track in the subsequent process.
[0052] Step S208: The individual audio track signals are merged to output the processed media audio signal.
[0053] After conditional gain adjustment of each independent audio track, all processed independent audio track signals can be merged back into a single complete audio signal. This merged signal is the media audio signal processed by this adaptive equalization method, and is finally sent to the in-vehicle audio playback device for output, thereby automatically optimizing the loudness and clarity of different audio components in a complex and changing in-vehicle noise environment.
[0054] The following is based on Figure 3 Taking this as an example, the specific process for determining in-car sound equalization is introduced, such as... Figure 3 As shown, this method can be executed by the vehicle's sound equalization system, and specifically includes the following steps: Step S302: Collect in-vehicle audio signals and perform preprocessing.
[0055] In one embodiment, the system acquires a mixed raw signal containing background noise and media audio via an in-vehicle microphone, and then performs echo cancellation processing to remove media audio components such as music and navigation voice played by the speakers, resulting in a preliminarily purified in-vehicle audio signal. For example, a vehicle is traveling at 80 km / h on a highway, and music is playing in the car along with navigation prompts. The signal acquired by the microphone includes wind noise, tire noise, music, and navigation voice. The echo cancellation module uses the playing audio as a reference signal to estimate and filter out the music and navigation voice components in the signal in real time, ultimately outputting a preliminarily purified signal dominated by wind noise and tire noise.
[0056] Step S304: Perform scene assessment and noise reduction.
[0057] In one embodiment, the system performs voice activity detection on the preprocessed signal. If it determines that the scenario is a non-conversational / calling scene inside the vehicle, it performs non-steady-state noise cancellation on the signal to further filter out transient interference such as horn honking, thereby obtaining an in-vehicle background noise signal that characterizes the current steady-state environment. For example, if there is no obvious human voice dialogue in the preprocessed signal, the system determines it to be a "non-conversational scene." At this time, a sudden burst of rapid horn honking is heard outside the vehicle. The non-steady-state noise cancellation module identifies this sudden high-energy transient component and suppresses it from the signal, ultimately obtaining a stable steady-state background noise estimate mainly composed of engine and road rolling noise.
[0058] Step S306: Calculate the gain difference spectrum and determine the validity of the frame.
[0059] In one embodiment, the system performs a frequency domain transformation on the current in-vehicle background noise signal to obtain its amplitude spectrum, and subtracts it from a pre-stored quiet environment reference noise amplitude spectrum to obtain a gain difference spectrum. Subsequently, the system counts the number of valid frequency points based on the comparison results of the amplitude values of each frequency point in this spectrum with a threshold, and determines whether the current frame is a valid frame accordingly. For example, the system subtracts the current noise amplitude spectrum from the reference spectrum collected when the vehicle is stationary in the garage, and finds that the gain difference in the mid-to-low frequency band (e.g., 200-800Hz) generally exceeds 3dB. Statistics show that more than 70% of the frequency point differences are greater than a preset threshold, far exceeding the quantity threshold, thus determining the current frame as a "valid frame," indicating that the environmental noise has changed significantly and audio equalization needs to be activated.
[0060] Step S308: Separate the media audio into tracks and determine the gain strategy.
[0061] In one embodiment, when step S306 determines that a frame is valid, the system performs track splitting on the playing media audio signal to obtain multiple independent audio tracks, such as vocals and instruments. For each independent audio track, the corresponding gain adjustment amount is extracted from the gain difference spectrum based on its core frequency range, and it is determined whether the conditions for applying independent gain are met; simultaneously, the overall gain value required for the output signal is calculated based on the gain difference spectrum. For example, the currently playing song contains vocals, bass, and drums. After track splitting, the system extracts gain adjustment amounts for the vocal track (core frequency band 200Hz-3kHz) and the bass track (core frequency band 80Hz-200Hz), respectively, and finds that the vocal frequency band needs an average compensation of 4.2dB, and the bass frequency band needs a compensation of 5.8dB. At the same time, the minimum gain difference across the entire frequency band is calculated to be 2.5dB, therefore the overall gain is set to 2.5dB.
[0062] Step S310: Apply independent gain and overall gain and output signal.
[0063] In one embodiment, the system applies its corresponding independent gain to each independent audio track that meets the conditions, and then merges all the tracks into a single signal. Finally, the overall gain value determined in step S308 is applied to the merged signal to generate an enhanced audio signal that is ultimately adapted to the current noisy environment and outputs it. For example, since the gain adjustment of each track is greater than its respective gain threshold, the system boosts the vocal track by 4.2dB and the bass track by 5.8dB, and then merges all the tracks. The merged signal is then boosted by an overall 2.5dB due to the aforementioned overall gain. Ultimately, the vocals and bass in the output signal are clearly distinguishable in noisy environments, and the overall loudness matches the ambient noise well, without requiring the user to manually increase the volume.
[0064] Figure 4This is a schematic structural diagram of an electronic device according to an exemplary embodiment. Please refer to... Figure 4 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile storage, and may also include other necessary hardware. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it, forming a vehicle-based audio equalization device at the logical level. Of course, this specification does not exclude other implementation methods besides software implementation, such as logic devices or a combination of hardware and software, etc. In other words, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0065] Figure 5 This specification illustrates a block diagram of an in-vehicle sound equalization device. Please refer to... Figure 5 The device includes: The gain difference spectrum determination unit 502 is used to acquire the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment. Track splitting processing unit 504 is used to split the currently playing media audio signal into multiple independent audio tracks corresponding to different frequency ranges. Gain adjustment unit 506 is used to determine the corresponding gain adjustment amount from the gain difference spectrum according to the frequency range of each independent audio track, and to apply the corresponding gain adjustment to the independent audio track when the gain adjustment amount meets the gain condition. The audio track merging unit 508 is used to merge the signals of each independent audio track so as to output the processed media audio signal.
[0066] Optionally, the gain difference spectrum determination unit 502 is specifically used for: The background noise signal is subjected to frequency domain transformation to obtain the corresponding background noise amplitude spectrum; Obtain the reference noise amplitude spectrum corresponding to the reference noise signal, and subtract the background noise amplitude spectrum from the reference noise amplitude spectrum to obtain the gain difference spectrum.
[0067] Optionally, the device further includes: A smoothing processing unit is used to smooth the gain difference spectrum to obtain a smoothed gain difference spectrum; wherein, the smoothing processing is achieved by weighting the gain difference spectrum of the current frame with the gain difference spectrum of the previous frame.
[0068] Optionally, the device further includes: The frame determination unit is used to compare the amplitude of each frequency point in the gain difference spectrum with a preset threshold, and count the number of frequency points whose amplitude exceeds the threshold. If the number of frequency points exceeds a preset threshold, the current frame is determined to be a valid frame, and the track splitting process of the media audio signal continues. If the number of frequency points does not exceed a preset threshold, the current frame is determined to be an invalid frame, and when the gain difference spectrum meets the preset noise replacement condition, the background noise amplitude spectrum is replaced with the reference noise amplitude spectrum.
[0069] Optionally, the gain adjustment unit 506 is specifically used for: For each individual audio track, determine its corresponding core frequency range; From the gain difference spectrum, the statistical values of the amplitude of each frequency point within the core frequency range are extracted as the gain adjustment amount corresponding to the independent audio track.
[0070] Optionally, the gain adjustment unit 506 is specifically used for: For each individual audio track, the corresponding gain adjustment amount is compared with a preset gain threshold. If the gain adjustment amount is greater than the gain threshold, then a gain corresponding to the gain adjustment amount is applied to the independent audio track; If the gain adjustment is less than or equal to the gain threshold, no gain is applied to the independent audio track.
[0071] Optionally, the gain adjustment unit 506 is specifically used for: The overall gain value is determined based on the minimum gain value of all frequency points in the gain difference spectrum; the overall gain value is subtracted from the gain value of each frequency point in the gain difference spectrum to obtain the adjusted gain difference spectrum, and the gain adjustment amount of each independent audio track is determined based on the adjusted gain difference spectrum; The audio track merging unit 508 is specifically used for: The overall gain value is applied to the merged signal to generate the final output signal.
[0072] Optionally, the track splitting processing unit 504 is specifically used for: Collect the raw audio signals of the in-vehicle environment; The original audio signal is subjected to echo cancellation processing to remove the media audio signal contained therein, which is then used as the background noise signal.
[0073] Optionally, after acquiring the background noise signal, the device further includes: A voice detection unit is used to detect voice activity in the background noise signal; If the detection results indicate that the in-vehicle environment is not in a dialogue scenario, non-steady-state noise cancellation is performed on the background noise signal. If the detection results indicate that the in-vehicle environment is in a dialogue scenario, the media audio signal is not adjusted and is directly output.
[0074] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0075] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0076] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0077] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0078] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0079] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GPS receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0080] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0081] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0082] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0083] Therefore, specific embodiments of the subject matter have been described. Furthermore, the processes depicted in the figures are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0084] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. A method for equalizing in-vehicle sound effects, characterized in that, include: Acquire the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment; The currently playing media audio signal is processed into multiple independent audio tracks corresponding to different frequency ranges. For each independent audio track, the corresponding gain adjustment amount is determined from the gain difference spectrum according to its frequency range, and the corresponding gain adjustment is applied to the independent audio track when the gain adjustment amount meets the gain condition. The individual audio tracks are merged to output the processed media audio signal.
2. The method according to claim 1, characterized in that, Determining the gain difference spectrum of the background noise signal relative to a reference noise signal in a quiet environment includes: The background noise signal is subjected to frequency domain transformation to obtain the corresponding background noise amplitude spectrum; Obtain the reference noise amplitude spectrum corresponding to the reference noise signal, and subtract the background noise amplitude spectrum from the reference noise amplitude spectrum to obtain the gain difference spectrum.
3. The method according to claim 2, characterized in that, The method further includes: The gain difference spectrum is smoothed to obtain a smoothed gain difference spectrum; wherein, the smoothing is achieved by weighting the gain difference spectrum of the current frame with the gain difference spectrum of the previous frame.
4. The method according to claim 2, characterized in that, The method further includes: The amplitude of each frequency point in the gain difference spectrum is compared with a preset threshold, and the number of frequency points whose amplitude exceeds the threshold is counted. If the number of frequency points exceeds a preset threshold, the current frame is determined to be a valid frame, and the track splitting process of the media audio signal continues. If the number of frequency points does not exceed a preset threshold, the current frame is determined to be an invalid frame, and when the gain difference spectrum meets the preset noise replacement condition, the background noise amplitude spectrum is replaced with the reference noise amplitude spectrum.
5. The method according to claim 1, characterized in that, Determining the corresponding gain adjustment amount from the gain difference spectrum based on its frequency range includes: For each individual audio track, determine its corresponding core frequency range; From the gain difference spectrum, the statistical values of the amplitude of each frequency point within the core frequency range are extracted as the gain adjustment amount corresponding to the independent audio track.
6. The method according to claim 1, characterized in that, The step of applying a corresponding gain adjustment to the independent audio track when the gain adjustment amount meets the gain condition includes: For each individual audio track, the corresponding gain adjustment amount is compared with a preset gain threshold. If the gain adjustment amount is greater than the gain threshold, then a gain corresponding to the gain adjustment amount is applied to the independent audio track; If the gain adjustment is less than or equal to the gain threshold, no gain is applied to the independent audio track.
7. The method according to claim 1, characterized in that, Determining the corresponding gain adjustment amount from the gain difference spectrum based on its frequency range includes: The overall gain value is determined based on the minimum gain value of all frequency points in the gain difference spectrum; the overall gain value is subtracted from the gain value of each frequency point in the gain difference spectrum to obtain the adjusted gain difference spectrum, and the gain adjustment amount of each independent audio track is determined based on the adjusted gain difference spectrum; The output of the processed media audio signal includes: The overall gain value is applied to the merged signal to generate the final output signal.
8. The method according to claim 1, characterized in that, The acquisition of the background noise signal of the current in-vehicle environment includes: Collect the raw audio signals of the in-vehicle environment; The original audio signal is subjected to echo cancellation processing to remove the media audio signal contained therein, which is then used as the background noise signal.
9. The method according to claim 1, characterized in that, After acquiring the background noise signal, the method further includes: Speech activity detection is performed on the background noise signal; If the detection results indicate that the in-vehicle environment is not in a dialogue scenario, non-steady-state noise cancellation is performed on the background noise signal. If the detection results indicate that the in-vehicle environment is in a dialogue scenario, the media audio signal is not adjusted and is directly output.
10. An in-vehicle sound equalization device, characterized in that, include: The gain difference spectrum determination unit is used to acquire the background noise signal of the current in-vehicle environment and determine the gain difference spectrum of the background noise signal relative to the reference noise signal of the quiet environment. The track splitting unit is used to split the currently playing media audio signal into multiple independent audio tracks corresponding to different frequency ranges. The gain adjustment unit is used to determine the corresponding gain adjustment amount from the gain difference spectrum according to the frequency range of each independent audio track, and to apply the corresponding gain adjustment to the independent audio track when the gain adjustment amount meets the gain condition. The audio track merging unit is used to merge the signals of each independent audio track so as to output the processed media audio signal.
11. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 9.