Audio processing method, electronic equipment and storage medium

By establishing a sound masking curve based on the noise source signal during vehicle operation and dynamically adjusting the energy of the target audio, the problem of poor audio quality caused by noise masking is solved, and audio details are restored and user experience is improved.

CN120877756APending Publication Date: 2025-10-31ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510835975.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

During driving, wind noise, tire noise, and other noises mask some details in the audio, resulting in poor audio quality and affecting the user experience.

Method used

A sound masking curve is established based on the noise source signal, and the energy of the target audio is dynamically adjusted. By separating sound elements and adjusting the energy of each frequency band according to the masking curve, the proportion of each band is equal to that in the quiet state.

Benefits of technology

It improves the effectiveness of audio during vehicle operation, enhancing audio detail recovery and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877756A_ABST
    Figure CN120877756A_ABST
Patent Text Reader

Abstract

The invention relates to an audio processing method, electronic equipment and a storage medium, and the method comprises the steps: obtaining a noise source signal in a vehicle driving process; determining a sound masking curve according to the noise source signal, wherein the sound masking curve is used for representing the frequency band distribution of the noise source and the intensity of each frequency band; and dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve. According to the technical scheme, the sound masking curve is established based on the noise source signal, the energy of the target audio is dynamically adjusted in the driving process, and the use effect of the audio can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and in particular to an audio processing method, electronic device, and storage medium. Background Technology

[0002] During driving, due to the sound masking effect, noise such as wind noise and tire noise can mask some details in the audio, resulting in poor audio performance and affecting the user experience. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide an audio processing method, electronic device and storage medium that establishes a sound masking curve based on the noise source signal and dynamically adjusts the energy of the target audio during driving, thereby improving the audio performance.

[0004] To achieve the above objectives, this application provides an audio processing method, comprising: Acquire noise source signals while the vehicle is in motion; A sound masking curve is determined based on the noise source signal. The sound masking curve is used to characterize the frequency band distribution of the noise source and the intensity of each frequency band. Based on the sound masking curve, the energy of the target frequency band of the target audio is dynamically adjusted.

[0005] In some embodiments, when the target audio is the audio to be played, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: The target audio is subjected to sound element separation; Based on the sound masking curve, the energy of the target frequency band in each of the sound elements is dynamically adjusted so that the proportion of each sound element is restored to the proportion in the quiet state.

[0006] In some embodiments, dynamically adjusting the energy of the target frequency band in each of the sound elements based on the sound masking curve, so that the proportion of each of the sound elements is equal to the proportion in the quiet state, includes: Based on the sound masking curve, the dominant frequency band of the noise source signal is determined; The target frequency band for each of the sound elements is determined based on the main frequency band. The energy of the target frequency band in each of the aforementioned sound elements is dynamically adjusted so that the proportion of each of the aforementioned sound elements is equal to the proportion in the quiet state.

[0007] In some embodiments, when the target audio is the audio to be played, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: Fit the target noise signal from the noise source signal into the target audio; The sound elements are separated from the fitted target audio. Based on the sound masking curve, a first target frequency band and a second target frequency band are determined for each sound element, wherein the first target frequency band is the target frequency band that needs to be enhanced by energy, and the second target frequency band is the target frequency band that needs to be attenuated by energy. The energy of the first target frequency band and the second target frequency band of each sound element are dynamically adjusted so that the proportion of each sound element is equal to the proportion in the quiet state.

[0008] In some embodiments, when the target audio includes the audio to be played, after dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, the method further includes: Adjust the output phase of multiple target channels to align them.

[0009] In some embodiments, when the target audio is currently received voice data, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: Separate the target human voice from the target audio; Based on the sound masking curve and the frequency band corresponding to the target human voice, the target frequency band in the target audio is identified; The energy of the target frequency band is dynamically enhanced.

[0010] In some embodiments, the method further includes: While dynamically amplifying the energy of the target frequency band, the energy of the non-target frequency bands of the target audio is reduced.

[0011] In some embodiments, acquiring the noise source signal during vehicle operation includes: Acquire a signal from at least one continuous noise source, wherein the at least one continuous noise source includes at least one of wind noise, tire noise, and cabin ambient noise; Predicting transient noise signals based on road surface data; The noise source signal is obtained by processing the signal from the at least one continuous noise source and the transient noise signal.

[0012] This application also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of any of the methods described above.

[0013] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0014] Based on the above, this application provides an audio processing method, electronic device, and storage medium. The method includes: acquiring a noise source signal during vehicle operation; determining a sound masking curve based on the noise source signal, the sound masking curve being used to characterize the frequency band distribution of the noise source and the intensity of each frequency band; and dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve. The technical solution of this application, by establishing a sound masking curve based on the noise source signal and dynamically adjusting the energy of the target audio during vehicle operation, can improve the audio performance. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating an audio processing method according to one embodiment.

[0017] Figure 2 This is another schematic diagram of an audio processing method provided according to an embodiment.

[0018] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment.

[0019] Reference numerals: 310 - Processor; 311 - Memory; 312 - Network interface; 313 - Bus system. Detailed Implementation

[0020] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of this application. Based on the description of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0021] In the description of this application, unless otherwise expressly specified and limited, the terms "set," "install," "connect," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms based on the specific circumstances.

[0022] The terms “first,” “second,” “third,” etc., are used merely to distinguish numerical values ​​or elements with similar properties, rather than to indicate or imply relative importance or a specific order.

[0023] The terms “include,” “comprising,” or any other variation thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0024] Figure 1 This is a flowchart illustrating an audio processing method according to one embodiment. Figure 1 As shown, an audio processing method according to this application includes the following steps: S1, acquire noise source signals during vehicle operation; S2, determine the sound masking curve based on the noise source signal. The sound masking curve is used to characterize the frequency band distribution of the noise source and the intensity of each frequency band. S3 dynamically adjusts the energy of the target frequency band of the target audio based on the sound masking curve.

[0025] Among them, noise source signals refer to noise signals that interfere with useful sound signals during driving, affecting the output and / or recognition effect of useful sound signals. Noise source signals can be directly detected noise signals, noise signals obtained by processing based on models and detected vehicle driving data, or noise signals obtained by processing detected vehicle driving data and noise signals.

[0026] Sound masking curves characterize the frequency distribution and intensity of noise source signals. Frequency distribution refers to the frequency range in which the noise source signal is distributed, and intensity refers to the intensity distribution of the noise source signal within each frequency band. Based on the sound masking effect, when the frequencies of two sound signals are close (< critical bandwidth), or there is a period of time before / after a loud noise (e.g., 0.1-200ms), or the sound pressure level difference between two sound signals is greater than a preset threshold (e.g., >15dB), noise will mask the desired sound, causing loss of detail and affecting the output and / or recognition of the sound signal. After determining the sound masking curve, based on the frequency distribution and intensity of the noise source signal, the frequency bands in the audio signal affected by the masking effect can be analyzed. Furthermore, targeted adjustments to the sound energy can be made to enhance the recognizability and / or sound quality of the audio signal.

[0027] In some embodiments, the target audio includes the audio to be played and / or the currently received voice data. The audio to be played may be, for example, audio used for sound playback such as navigation output voice, music, video, or externally played voice call data. If the audio to be played is affected by noise, occupants may not be able to hear or clearly hear the played sound, affecting the passenger experience. The currently received voice data may be, for example, voice data received for voice recognition, or voice call data collected during calls via the in-vehicle hands-free system. If the currently received voice data is affected by noise, voice commands may not be recognized or may be misrecognized, or the call may be unable to proceed, affecting the passenger experience.

[0028] The target frequency band of the target audio refers to the frequency band affected by noise. During driving, as the noise source signal changes, the target frequency band that needs to be processed may change, and the energy required to adjust the target frequency band will also change. Therefore, by dynamically adjusting the energy of the target frequency band of the target audio, the dynamic requirements for audio processing in driving scenarios can be better met, resulting in a better user experience.

[0029] In some embodiments, step S1, acquiring noise source signals during vehicle operation, includes: Acquire a signal from at least one continuous noise source, the at least one continuous noise source including at least one of wind noise, tire noise, and cabin ambient noise; Predicting transient noise signals based on road surface data; The noise source signal is obtained by processing based on at least one continuous signal and a transient noise signal.

[0030] Continuous noise refers to noise that lasts for a relatively long time during driving, such as wind noise, tire noise, and cabin ambient noise, which usually accompany the entire driving process. Transient noise refers to noise that lasts for a short time during driving, such as the noise generated when the vehicle goes over potholes or speed bumps, which disappears after the vehicle has passed over the potholes or speed bumps.

[0031] The higher the vehicle speed, the greater the wind noise. Vehicle speed can be acquired in real time using a vehicle speed sensor, and the dominant frequency band of wind noise at the current speed can be predicted using a CFD (Computational Fluid Dynamics) model to obtain the wind noise signal. Vibration acceleration sensors in the vehicle chassis collect signals to provide feedback on road noise and tire noise. The vibration acceleration sensors can provide immediate feedback of road information to the processor, offering advanced capabilities for capturing transient noise. Furthermore, data collected by forward-facing cameras or radar can be analyzed to identify road surface data such as potholes and speed bumps, allowing for the prediction of upcoming transient noise signals. An in-vehicle microphone array picks up ambient noise from the cabin. This noise may include voices from occupants that could interfere with audio reception or analysis, as well as voice commands being input. These various noise signals can be fused to fit the noise source signal. By combining continuous noise signals and predicting transient noise, and by fully integrating multiple noise sources, the high dynamic requirements of driving scenarios can be better met, and the audio processing effect can be improved.

[0032] In some embodiments, road surface data can be obtained by analyzing images captured by cameras or radar point clouds. The type of obstacle, such as potholes, sharp objects, speed bumps, or steps, can be determined based on the road surface data. Transient noise signals can be predicted by combining this data with vehicle driving data, such as vehicle speed, braking signals, and acceleration. This allows for the prediction of the timing, frequency, and intensity of transient noise. Furthermore, prediction can be made by combining user behavior models and preset associated data of the currently driven vehicle. For example, if a user habitually reduces their speed to a target speed range when passing speed bumps, the timing of transient noise can be predicted based on the current speed and the target speed range. Simultaneously, the preset associated data of the user's currently driven vehicle can include the frequency and intensity of transient noise corresponding to different obstacle types, thus ultimately achieving transient noise prediction. By predicting transient noise, a sound masking curve can be determined before transient noise occurs, allowing for timely energy compensation of the target audio when transient noise occurs, improving processing efficiency.

[0033] In some embodiments, when the target audio is the audio to be played, step S3, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, includes: Separate the sound elements from the target audio; Based on the sound masking curve, the energy of the target frequency band in each sound element is dynamically adjusted so that the proportion of each sound element is equal to the proportion in the quiet state.

[0034] The audio to be played can be composed of various sound elements. For example, a single frame of audio (with a wavelength of 0.02-0.05 seconds) can be composed of different sound elements such as a whistle, guitar, drums, or birdsong. Different sound elements have different dominant frequency bands. When some components of a sound element are masked, the proportion of sound elements becomes unbalanced, resulting in audio distortion. The proportion refers to the ratio of energy between different frequency bands within a sound element. For example, if a sound element's frequency band is 500-600Hz, the maximum energy in the 500-550Hz range is 50dB, and the maximum energy in the 550-600Hz range is 100dB. This allows us to determine the proportion of energy between different frequency bands. When a sound element is masked, some frequency band details are lost, thus changing the overall proportion.

[0035] The technical solution of this application separates the sound elements of the target audio, allowing each sound element to be processed independently. By dynamically adjusting the energy of the target frequency bands in each sound element, the masked components of each sound element can be restored, making the proportion of each sound element equal to that in the quiet state. Therefore, the final output audio achieves a sound effect close to or the same as that under quiet (static) conditions, resulting in better audio processing. It can be understood that "equal to the proportion in the quiet state" can mean completely equal to the proportion in the quiet state, or within a preset deviation range of the proportion in the quiet state. In some embodiments, the sound element separation of the target audio can be based on the NMF (Nonnegative Matrix Factorization) algorithm.

[0036] After separating the sound elements of the target audio, the target frequency bands in each sound element can be determined based on the sound masking curve, and the energy of the target frequency bands of each sound element can be adjusted accordingly. Compared to directly processing the target audio without separating the sound elements, this method can better achieve the processing of audio details and obtain better sound quality fidelity.

[0037] In some embodiments, based on a sound masking curve, the energy of the target frequency band in each sound element is dynamically adjusted so that the proportion of each sound element is equal to the proportion in a quiet state, including: Based on the sound masking curve, the dominant frequency band of the noise source signal is determined; Determine the target frequency band for each sound element based on the main frequency band; The energy of the target frequency band in each sound element is dynamically adjusted so that the proportion of each sound element is equal to that in the quiet state.

[0038] This involves identifying the dominant frequency band of the noise source signal. Within this dominant frequency band, the sound masking effect is significant. By compensating the energy of frequency bands in the sound elements that are identical to this dominant frequency band, the components of the sound elements can be recovered. For example, some frequency bands in a drum sound are masked; through energy compensation, the complete frequency band of the drum sound can be recovered so that it can be heard by the human ear without distortion. Taking the 2-4kHz frequency band dominated by wind noise as an example, the energy of the sound signal in this frequency band is increased by 3-6dB to make it audible to the human ear.

[0039] In some embodiments, when the target audio is the audio to be played, step S3, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, includes: Fit the target noise signal from the noise source signal into the target audio; Separate the sound elements from the fitted target audio; Based on the sound masking curve, the first target frequency band and the second target frequency band in each sound element are determined. The first target frequency band is the target frequency band that needs to be enhanced by energy, and the second target frequency band is the target frequency band that needs to be attenuated by energy. The energy of the first target frequency band and the second target frequency band of each sound element are dynamically adjusted so that the proportion of each sound element is equal to that in the quiet state.

[0040] Among them, the target noise signal in the noise source signal can be wind noise, tire noise, etc. These noise sources are relatively stable. By fitting them to the audio to be played, the energy of the target audio can be initially compensated.

[0041] After fitting the target noise signal from the noise source signal to the audio to be played, overcompensation or undercompensation of the masked frequency bands may occur. Therefore, for the fitted target audio, based on the sound masking curve, the first and second target frequency bands of each sound element are determined, that is, the frequency bands of the sound element that need energy compensation or energy attenuation are determined. For example, if the noise signal is 500Hz, 500Hz is an externally introduced noise, and the sound signal energy at this frequency needs to be weakened. The frequency band near 500Hz is largely masked by noise due to the masking effect, so more compensation is needed for this component. Thus, the components of the sound element can be restored, making the proportion of each sound element equal to that in the quiet state. In this way, the noise energy can be used to compensate for the target audio to eliminate the masking effect of noise, simplifying the energy compensation process.

[0042] In some embodiments, when the target audio includes the audio to be played, after dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, the method further includes: Adjust the output phase of multiple target channels to align them.

[0043] In this process, after adjusting the energy of the sound signal in the target frequency band, the energy of some frequency bands will be significantly increased relative to other frequency bands. When the audio is output in a multi-channel configuration, a relative deviation will occur between the output phases of different channels, causing sound image shift and resulting in sound distortion. For example, if the energy of the drum sound originally output from the left speaker is amplified, while the energy of the piano sound output from the left speaker is not amplified, the perceived position of the piano sound will be affected, resulting in a sound field deviation. In this case, appropriately enhancing the energy of the piano sound can align the output phases between channels, eliminate the sound field deviation, enhance sound field positioning calibration, maintain a stereo sound field consistent with a quiet state, and further improve the sound effect. In some embodiments, the energy of the output audio of each channel can be adjusted by adjusting the input parameters of the speaker equalizer, thereby achieving the adjustment of the output phase.

[0044] In some embodiments, when the target audio is currently received voice data, step S3, dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, includes: Extract the target human voice from the target audio; Based on the sound masking curve and the frequency band corresponding to the target human voice, the target frequency band in the target human voice is identified. Dynamically enhance the energy of the target frequency band.

[0045] The target voice is the useful human voice that needs to be recognized, typically the voice closest to the microphone array and emitted from the driver's position. After separating the target voice, the masking curve is used to determine the masked portion and degree of masking, and the separated voice is then amplified to improve the speech recognition rate and accuracy.

[0046] In some embodiments, the method further includes: While dynamically enhancing the energy of the target frequency band, the energy of the non-target frequency bands of the target audio is reduced.

[0047] The non-target frequency bands are those in the target audio that do not belong to the sound components required for speech recognition. By emphasizing the separated human voice while weakening other sound components, the speech recognition rate and accuracy can be further improved.

[0048] Please refer to Figure 3 This is a schematic diagram illustrating a specific process according to an embodiment of this application. The process involves: acquiring ambient noise via a microphone array signal; fitting wind noise using vehicle speed signals; determining tire noise using vibration acceleration sensor signals; and predicting transient noise based on camera data. These noise signals are then fitted to obtain the noise source signal, and a sound masking curve is determined. Next, the purpose of the target audio is identified. When it is voice data for speech recognition, the masked components in the human voice signal are determined based on the sound masking curve, and the masked components are compensated to improve the speech recognition rate and accuracy. When it is audio to be played for sound playback, sound elements are separated, and the masked components in each sound element are compensated, thereby restoring the proportion of each sound element to be equal to that in a quiet state, achieving the same sound effect as in a quiet state. In this way, during driving, the currently playing audio, the currently received voice data, or both can be processed based on the sound masking curve, effectively improving the user experience.

[0049] As described above, the audio processing method of this application acquires noise source signals during vehicle operation; determines a sound masking curve based on the noise source signals, the sound masking curve being used to characterize the frequency band distribution of the noise source and the intensity of each frequency band; and dynamically adjusts the energy of the target frequency band of the target audio based on the sound masking curve. The technical solution of this application, by establishing a sound masking curve based on the noise source signals and dynamically adjusting the energy of the target audio during driving, can improve the audio performance.

[0050] This application also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in any of the above embodiments.

[0051] like Figure 3 As shown, the electronic device includes: a processor 310 and a memory 311 storing a computer program; wherein, Figure 3 The processor 310 shown in the diagram does not indicate that there is only one processor 310, but only indicates the positional relationship of the processor 310 relative to other devices. In practical applications, there can be one or more processors 310; similarly, Figure 3 The memory 311 shown in the diagram has the same meaning, that is, it is only used to indicate the positional relationship of memory 311 relative to other devices. In practical applications, there can be one or more memories 311. When the processor 310 runs the computer program, the above-described vehicle control method is implemented.

[0052] The electronic device may also include at least one network interface 312. The various components of the electronic device are coupled together via a bus system 313. It is understood that the bus system 313 is used to implement communication between these components. In addition to a data bus, the bus system 313 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general designated all buses as Bus System 313.

[0053] The memory 311 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 311 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0054] The memory 311 in this embodiment of the invention is used to store various types of data to support the operation of the electronic device. Examples of this data include: any computer programs used to operate on the electronic device, such as operating systems and applications; contact data; phonebook data; messages; pictures; videos, etc. The operating system includes various system programs, such as a framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications, such as media players, browsers, etc., used to implement various application services. Here, the program implementing the method of this embodiment of the invention can be included in the application.

[0055] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in any of the above embodiments.

[0056] Computer-readable storage media can be magnetic random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. When the computer program stored in the computer-readable storage medium is executed by a processor, it implements a control method for the vehicle applied to the aforementioned electronic device. For the specific steps implemented when the computer program is executed by the processor, please refer to [reference needed]. Figure 1 The description of the illustrated embodiments will not be repeated here.

[0057] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. An audio processing method, characterized in that, include: Acquire noise source signals while the vehicle is in motion; A sound masking curve is determined based on the noise source signal. The sound masking curve is used to characterize the frequency band distribution of the noise source and the intensity of each frequency band. Based on the sound masking curve, the energy of the target frequency band of the target audio is dynamically adjusted.

2. The method according to claim 1, characterized in that, When the target audio is the audio to be played, the step of dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: The target audio is subjected to sound element separation; Based on the sound masking curve, the energy of the target frequency band in each of the sound elements is dynamically adjusted so that the proportion of each sound element is equal to the proportion in the quiet state.

3. The method according to claim 2, characterized in that, The step of dynamically adjusting the energy of the target frequency band in each sound element based on the sound masking curve, so that the proportion of each sound element is equal to the proportion in the quiet state, includes: Based on the sound masking curve, the dominant frequency band of the noise source signal is determined; The target frequency band for each of the sound elements is determined based on the main frequency band. The energy of the target frequency band in each of the aforementioned sound elements is dynamically adjusted so that the proportion of each of the aforementioned sound elements is equal to the proportion in the quiet state.

4. The method according to claim 1, characterized in that, When the target audio is the audio to be played, the step of dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: Fit the target noise signal from the noise source signal into the target audio; The sound elements are separated from the fitted target audio. Based on the sound masking curve, a first target frequency band and a second target frequency band are determined for each sound element, wherein the first target frequency band is the target frequency band that needs to be enhanced by energy, and the second target frequency band is the target frequency band that needs to be attenuated by energy. The energy of the first target frequency band and the second target frequency band of each sound element are dynamically adjusted so that the proportion of each sound element is equal to the proportion in the quiet state.

5. The method according to any one of claims 1 to 4, characterized in that, When the target audio includes the audio to be played, after dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve, the method further includes: Adjust the output phase of multiple target channels to align them.

6. The method according to claim 1, characterized in that, When the target audio is currently received voice data, the step of dynamically adjusting the energy of the target frequency band of the target audio based on the sound masking curve includes: Separate the target human voice from the target audio; Based on the sound masking curve and the frequency band corresponding to the target human voice, the target frequency band in the target human voice is identified. The energy of the target frequency band is dynamically enhanced.

7. The method according to claim 6, characterized in that, The method further includes: While dynamically amplifying the energy of the target frequency band, the energy of the non-target frequency bands of the target audio is reduced.

8. The method according to claim 1, characterized in that, The acquisition of noise source signals during vehicle operation includes: Acquire a signal from at least one continuous noise source, wherein the at least one continuous noise source includes at least one of wind noise, tire noise, and cabin ambient noise; Predicting transient noise signals based on road surface data; The noise source signal is obtained by processing the signal from the at least one continuous noise source and the transient noise signal.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • High-quality vehicle-mounted audio and video call method and system based on mobile internet

    CN121864771A