Audio signal processing method and processing device
By performing frame processing and feature extraction of audio signals in the cockpit, determining processing rules based on signal levels, dynamic compression or gain compensation is performed, the problem of large volume differences is solved and the user experience is improved.
Patent Information
- Application Number
- CN202510385035.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-13
AI Technical Summary
When playing music or videos in the cockpit, due to the differences in audio recording or encoding levels of different contents, the volume difference is significant, affecting the user experience.
By performing frame-based processing of the audio signal, the characteristic signal value of each frame of signal frame is extracted, the signal level is determined based on the characteristic signal value, the target signal processing rules are determined according to the signal level, and dynamic compression or gain compensation is performed so that the difference between the characteristic signal values of any two frames of signal frames is within the preset range.
It effectively reduces the volume difference of the audio signal, improves the user experience, and ensures the stability and consistency of the audio signal.
Smart Images

Figure CN120148544A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio processing, and particularly relates to a method and a device for processing audio signals. Background Art
[0002] The cockpit of a vehicle usually can be equipped with a full liquid crystal instrument, a head-up display, and an in-vehicle entertainment system, and can enable the rear seat passengers to obtain an entertainment experience through the screen carried on the rear seat. For example, user-related audio can be played in the cockpit, such as music, navigation, etc.
[0003] When playing music or video in the cockpit, due to the differences in audio recording or encoding levels of different contents, there may be obvious volume differences. For example, the sound of music or video suddenly becomes extremely loud or extremely small, affecting the user experience. Therefore, when playing music or video in the cockpit, the audio in the audio or video will be processed so that the corresponding music or video sound will not suddenly become louder or smaller.
[0004] Currently, the main way to process audio is to process the audio signal after file decoding and audio decoding, which is difficult to fundamentally solve the problem of unstable audio sound. Summary of the Invention
[0005] In view of this, this application provides a method and a device for processing audio signals, which solve or improve the technical problem that in the prior art, there are differences in audio recording or encoding levels of different contents in the intelligent cockpit, resulting in obvious volume differences and thus affecting the user experience.
[0006] As the first aspect of this application, this application provides a method for processing audio signals, including:
[0007] Performing frame division processing on the audio signal to obtain multiple frame signal frames;
[0008] Extracting the characteristic signal value of each frame signal frame;
[0009] Determining the signal level of each frame signal frame according to the characteristic signal value of each frame signal frame and the preset characteristic signal range;
[0010] Determining the corresponding target signal processing rule according to the signal level of each frame signal frame, wherein when the signal levels of the signal frames are different, the corresponding target signal processing rules of the signal frames are different;
[0011] Processing each frame signal frame based on the corresponding target signal processing rule of each frame signal frame so that the difference between the characteristic signal values of any two frame signal frames is within the preset range.
[0012] In one embodiment of the present application, determining the signal level of each signal frame according to the characteristic signal value of each signal frame and the preset characteristic signal range includes:
[0013] When the characteristic signal value of the signal frame is greater than the maximum characteristic signal value in the preset characteristic signal range, determining that the signal level of the signal frame is the first signal level;
[0014] When the characteristic signal value of the signal frame is greater than or equal to the minimum characteristic signal value in the preset characteristic signal range and the characteristic signal value of the signal frame is less than or equal to the maximum characteristic signal value in the preset characteristic signal range, determining that the signal level of the signal frame is the second signal level;
[0015] When the characteristic signal value of the signal frame is less than the minimum characteristic signal value in the preset characteristic signal range, determining that the signal level of the signal frame is the third signal level.
[0016] In one embodiment of the present application, the characteristic signal includes: the spectral level of the signal frame, and / or the overall loudness of the signal frame, and / or the maximum amplitude value of the signal frame.
[0017] In one embodiment of the present application, determining the corresponding target signal processing rule according to the signal level of each signal frame includes:
[0018] When the signal level of the signal frame is the first signal level, determining that the target signal processing rule of the signal frame is dynamic compression;
[0019] When the signal level of the signal frame is the third signal level, determining that the target signal processing rule of the signal frame is gain compensation.
[0020] In one embodiment of the present application, when the target signal processing rule corresponding to the signal frame is dynamic compression,
[0021] wherein, processing each signal frame based on the target signal processing rule corresponding to each signal frame includes:
[0022] Performing dynamic compression on the signal frames with the first signal level according to the preset compression threshold, preset attack time, preset release time, and preset compression ratio.
[0023] In one embodiment of the present application, the preset compression ratio is 4:1.
[0024] In one embodiment of the present application, when the target signal processing rule corresponding to the signal frame is gain compensation,
[0025] wherein, processing each signal frame based on the target signal processing rule corresponding to each signal frame includes:
[0026] Obtain the characteristic signal of the signal frame as the spectral level value of the spectral level, and calculate the initial gain compensation amount according to the spectral level value and the target spectral level;
[0027] Use a low-pass filter to smooth the initial gain compensation amount to determine the gain compensation amount of the signal frame;
[0028] Perform gain compensation on the spectral level value of the signal frame based on the gain compensation amount of the signal frame.
[0029] In an embodiment of the present application, the performing gain compensation on the spectral level value of the signal frame based on the gain compensation amount of the signal frame includes:
[0030] When the gain compensation amount is greater than the preset gain compensation threshold, perform gain compensation on the spectral level value of the signal frame based on the preset gain compensation threshold;
[0031] When the gain compensation amount is less than or equal to the preset gain compensation threshold, perform gain compensation on the spectral level value of the signal frame based on the gain compensation amount.
[0032] In an embodiment of the present application, after processing each signal frame based on the target signal processing rule corresponding to each signal frame, the audio signal processing method further includes:
[0033] Perform smoothing processing and channel balance processing on the audio signal of the signal frame processed by the target signal processing rule.
[0034] As a second aspect of the present application, the present application further provides a processing device for an audio signal, including:
[0035] A frame splitting module, configured to perform frame splitting processing on the audio signal to obtain multiple signal frames;
[0036] A feature extraction module, configured to extract the characteristic signal value of each signal frame;
[0037] A signal level classification module, configured to determine the signal level of each signal frame according to the characteristic signal value of each signal frame and the preset characteristic signal range;
[0038] A signal processing module, configured to determine the corresponding target signal processing rule according to the signal level of each signal frame, and process each signal frame based on the target signal processing rule corresponding to each signal frame, so that the difference between the characteristic signal values of any two signal frames is within the preset range; wherein, when the signal levels of the signal frames are different, the target signal processing rules corresponding to the signal frames are different.
[0039] A method for processing an audio signal provided by the present application frames the audio signal into multiple multi-frame signal frames of a certain length, extracts the characteristic signal values of each frame of the signal frame, determines the signal level of each frame of the signal frame according to the characteristic signal values, determines the corresponding target signal processing rule according to the signal level, and finally processes each frame according to the target signal processing rule, so that the characteristic signal values of the multi-frame signal frames are stable and consistent, thereby making the volume difference of the audio signal smaller and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0041] Figure 1 The figure shows a schematic flowchart of a method for processing an audio signal provided by an embodiment of the present application.
[0042] Figure 2 The figure shows a schematic flowchart of a method for processing an audio signal provided by another embodiment of the present application.
[0043] Figure 3 The figure shows a schematic flowchart of a method for processing an audio signal provided by another embodiment of the present application.
[0044] Figure 4 The figure shows a schematic flowchart of a method for processing an audio signal provided by another embodiment of the present application.
[0045] Figure 5 The figure shows a block diagram of a structure of an audio signal processing device provided by an embodiment of the present application.
[0046] Figure 6 The figure shows a block diagram of a structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In the description of this application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined. In the embodiments of this application, all directional indications (such as up, down, left, right, front, back, top, bottom...) are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally also include unlisted steps or units, or may optionally also include other steps or units inherent to these processes, methods, products, or devices.
[0048] In addition, referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0049] Application Overview
[0050] As described in the background art, user-related audio can be played in the cockpit, such as music, navigation, etc. Among them, the principle of audio playback in the vehicle cockpit is as follows:
[0051] File decoding: Decode the audio file of the music to be played (such as MP3, WAV, etc. formats) into raw audio data.
[0052] Audio decoding: The raw audio data after file decoding needs to be further decoded into an audio signal, and the audio signal can be converted into a format that can be processed by an audio digital signal processor (ADSP, Audio Digital Signal Processor).
[0053] Audio transmission: Transmit the audio signal after audio decoding through an audio transmission protocol.
[0054] Signal processing: Use an audio digital signal processor (ADSP) to digitally process the received audio signal after audio decoding, such as equalization, dynamic range control, etc., to optimize the sound quality.
[0055] Digital-to-analog conversion: Convert the digital audio signal after being processed by the audio digital signal processor into an analog signal to be converted into an analog signal required by an audio output device (such as a speaker).
[0056] Amplification of audio signal: Amplify the analog audio signal after digital-to-analog conversion to drive output devices such as speakers or headphones.
[0057] Playback: Play the amplified analog audio signal through a speaker or other audio output device.
[0058] During the research process, the inventor found that when playing music or videos in the cockpit, there are obvious volume differences. The main reason for the inconsistent volume is that the original levels after decoding different audio files are inconsistent. Therefore, this application provides a method for processing audio signals, which divides the audio signal into multiple signal frames of a certain length, extracts the characteristic signal values of each signal frame, determines the signal level of each signal frame according to the characteristic signal values, determines the corresponding target signal processing rules according to the signal level, and finally processes each frame of signal according to the target signal processing rules, so that the characteristic signal values of the multiple signal frames are stable and consistent, thereby making the volume difference of the audio signal smaller and improving the user experience.
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0060] Exemplary method
[0061] As a first aspect of the present application, the present application provides a method for processing an audio signal. Figure 1 The following shows a schematic flowchart of a method for processing an audio signal provided by an embodiment of the present application. As Figure 1 shown, a method for processing an audio signal includes the following steps:
[0062] S1: Perform frame division on the audio signal to obtain multiple signal frames;
[0063] Specifically, the audio signal is a pulse modulation signal obtained after decoding compressed audio, that is, a PCM (Pulse Code Modulation) format signal. PCM decoding can decode compressed audio, sample it at fixed time intervals, and quantize the sampled values into a group of binary digits to convert the compressed audio into a pulse modulation signal. Among them, the sampling rate during PCM decoding is 44.1 kHz (that is, 44,100 sampling points per second).
[0064] Framing the audio signal means: dividing the continuous audio signal into a series of signal frames, each signal frame having a fixed length, the lengths of each signal frame being the same, and the number of sampling points included in each signal frame being the same, for example, 256 sampling points.
[0065] Specifically, the frame length of each signal frame is: 10 ms to 50 ms (for example, 20 ms), and the number of sampling points corresponding to each signal frame is: frame length (ms) × sampling rate (kHz). For example, when the frame length is 20 ms and the sampling rate is 44.1, the number of sampling points for each corresponding signal frame is 882 points.
[0066] Optionally, when framing the audio signal, there is a certain frame overlap rate between two adjacent signal frames to avoid signal distortion or signal discontinuity. For example, in this application, the frame overlap rate between two adjacent signal frames is 50%, that is, the starting point of each signal frame will overlap half of the ending point of the previous signal frame. In this way, when the audio signal is reconstructed, these overlapping parts help to reduce the discontinuity and distortion between frames.
[0067] S2: Extract the characteristic signal value of each signal frame;
[0068] Specifically, the characteristic signal value can be:
[0069] The overall loudness of the signal frame: that is, the RMS value. The RMS value is a root mean square value that can reflect the energy level of the audio signal over a period of time and reflect the overall loudness value of the audio signal; and / or
[0070] The maximum amplitude of the signal frame: The maximum amplitude of the signal frame refers to taking the maximum amplitude in each signal frame, which is an instantaneous value and can be used to determine whether there is a sudden high energy in the signal of each signal frame.
[0071] The spectral level value of the signal frame: The spectral level value refers to the amplitude magnitude of the audio signal at different frequency components. Specifically, perform a fast Fourier transform (FFT) on each signal frame to calculate the spectral energy distribution of the signal frame, and then use a window function (such as a Hanning window) to weight the samples of the signal frame to reduce the spectral leakage of the signal frame.
[0072] The characteristic signal value of the signal frame can be at least one or more than two of the above overall loudness, the maximum amplitude of the signal frame, and the spectral level value.
[0073] Optionally, when extracting the characteristic signal value of the signal frame, the total delay should be controlled within 1 frame length, that is, about 20 ms, to meet the real-time requirement.
[0074] S3: Determine the signal level of each signal frame according to the characteristic signal value of each signal frame and the preset characteristic signal range;
[0075] Each characteristic signal value corresponds to a preset characteristic signal range. For example, when the characteristic signal value is the spectrum level value, then the spectrum level value corresponds to a preset level range. For example, the preset level range can be: -234 LUFS to 241 LUFS.
[0076] Based on the characteristic signal value of each signal frame and the corresponding preset characteristic signal range, the signal level of each signal frame can be determined. Specifically, the signal level of the signal frame can be divided into three levels: the first signal level (which can also be called the too-high signal level), the second signal level (the normal signal level), and the third signal level (the too-low signal level). Among them, the signal level of the signal frame with the characteristic signal value higher than the preset characteristic signal range is classified as the first signal level, the signal level of the signal frame with the characteristic signal value within the preset characteristic signal range is classified as the second signal level, and the signal level of the signal frame with the characteristic signal value lower than the preset characteristic signal range is classified as the third signal level.
[0077] Optionally, the signal level of the signal frame can not only be divided into the above three signal levels, but also the multiple signal frames that meet the preset characteristic signal range can be further divided into multiple signal levels.
[0078] Also optionally, when classifying the signal level of the signal frame, the judgment factor for classification can include only one characteristic signal value. For example, the signal level of each signal frame is classified only according to the spectrum level value of the signal frame and the preset level range. Also for example, the signal level of each signal frame is classified only according to the overall loudness of the signal frame and the preset overall loudness range. The judgment factor for classification can also include two or more characteristic signal values. For example, the signal level of each signal is classified jointly according to the spectrum level value of the signal frame and the preset level range, the overall loudness of the signal frame, and the preset overall loudness range.
[0079] S4: Determine the corresponding target signal processing rule according to the signal level of each signal frame. Among them, when the signal levels of the signal frames are different, the corresponding target signal processing rules of the signal frames are different;
[0080] Specifically, the signal level of the signal frame and the corresponding target signal processing rule can be stored in the database in advance. After determining the signal level of the signal frame, the target signal processing rule corresponding to the signal level can be queried in the database according to the signal level of the signal frame. For example, the target signal processing rule corresponding to the signal frame of the high-level signal level can be dynamic compression, and the target signal processing rule corresponding to the signal frame of the low-level signal level can be gain compensation.
[0081] S5: Process each signal frame based on the target signal processing rule corresponding to each signal frame, so that the difference between the characteristic signal values of any two signal frames is within a preset range.
[0082] After determining the target signal processing rule of the signal frame, the signal frame can be processed according to the target signal processing rule, so that the difference between the characteristic signal values of any two signal frames is within a preset range, and the characteristic signal values of multiple signal frames are stably not very different, so that the characteristic signal values of the processed audio signal are consistent.
[0083] A processing method for an audio signal provided by this application frames the audio signal into multiple signal frames of a certain length, extracts the characteristic signal value of each signal frame, determines the signal level of each signal frame according to the characteristic signal value, determines the corresponding target signal processing rule according to the signal level, and finally processes each frame of signal according to the target signal processing rule, so that the characteristic signal values of multiple signal frames are stably consistent, thereby making the volume difference of the audio signal smaller and improving the user experience.
[0084] In an embodiment of this application, as Figure 2 shown, the specific determination method for determining the signal level of the signal frame, that is, S3 (determine the signal level of each signal frame according to the characteristic signal value of each signal frame and the preset characteristic signal range) specifically includes the following steps:
[0085] S31: When the characteristic signal value of the signal frame is greater than the maximum characteristic signal value in the preset characteristic signal range, determine that the signal level of the signal frame is the first signal level;
[0086] The characteristic signal value of the signal frame being greater than the maximum characteristic signal value in the preset characteristic signal range indicates that the characteristic signal value of the signal frame is too high. Therefore, it can be determined that the signal level of the signal frame is the first signal level (too high signal level).
[0087] Specifically, the characteristic signal value can be the spectral level value of the signal frame, and the corresponding preset characteristic signal range can be: -234 LUFS to 241 LUFS. Then the signal level of the signal frame with a spectral level value greater than 241 LUFS is the too high signal level.
[0088] S32: When the characteristic signal value of the signal frame is greater than or equal to the minimum characteristic signal value in the preset characteristic signal range and the characteristic signal value of the signal frame is less than or equal to the maximum characteristic signal value in the preset characteristic signal range, determine that the signal level of the signal frame is the second signal level;
[0089] Similarly, if the characteristic signal value of the signal frame is less than or equal to the maximum characteristic signal value in the preset characteristic signal range and greater than or equal to the minimum characteristic signal value in the preset characteristic signal range, it indicates that the characteristic signal value of the signal frame is normal. Therefore, the signal level of the signal frame can be determined to be the second signal level (normal signal level).
[0090] Specifically, the characteristic signal value can be the spectral level value of the signal frame, and the corresponding preset characteristic signal range can be: -234 LUFS to 241 LUFS. Then, the signal level of a signal frame with a spectral level value less than -234 LUFS is the too-low signal level.
[0091] S33: When the characteristic signal value of the signal frame is less than the minimum characteristic signal value in the preset characteristic signal range, determine that the signal level of the signal frame is the third signal level.
[0092] Similarly, if the characteristic signal value of the signal frame is less than the minimum characteristic signal value in the preset characteristic signal range, it indicates that the characteristic signal value of the signal frame is too low. Therefore, the signal level of the signal frame can be determined to be the third signal level (too-low signal level).
[0093] Specifically, the characteristic signal value can be the spectral level value of the signal frame, and the corresponding preset characteristic signal range can be: -234 LUFS to 241 LUFS. Then, the signal level of a signal frame with a spectral level value less than -234 LUFS is the too-low signal level.
[0094] In an embodiment of the present application, as Figure 3 shown, the specific determination method for determining the target signal processing rule of each signal frame, that is, S4 (determining the corresponding target signal processing rule according to the signal level of each signal frame) specifically includes the following steps:
[0095] S41: When the signal level of the signal frame is the first signal level, determine that the target signal processing rule of the signal frame is dynamic compression;
[0096] For a signal frame with a too-high signal level, the target signal processing rule of the signal frame is dynamic compression to prevent discomfort to the listening sensation or damage to the device caused by a too-high level signal.
[0097] Specifically, when the target signal processing rule of the signal frame is dynamic compression, then correspondingly, S5 (processing each signal frame based on the target signal processing rule corresponding to each signal frame) specifically includes the following steps:
[0098] S51: Perform dynamic compression on the signal frame with the first signal level according to the preset compression threshold, preset attack time, preset release time, and preset compression ratio.
[0099] Specifically, the compression method of dynamic compression can be as follows:
[0100] When the characteristic signal value of the signal frame with the first signal level is greater than the preset maximum compression threshold, the compressed characteristic signal value = preset maximum compression threshold + (characteristic signal value before compression - preset maximum compression threshold) / preset compression ratio;
[0101] When the characteristic signal value of the signal frame with the first signal level is greater than or equal to the preset minimum compression threshold and less than or equal to the preset maximum compression threshold, the compressed characteristic signal value = preset minimum compression threshold + preset compression ratio * (characteristic signal value before compression - preset minimum compression threshold);
[0102] Optionally, by setting the preset compression threshold, the signals with amplitudes lower than the compression threshold in the signal frame remain unchanged, and the signals with amplitudes higher than the preset compression threshold in the signal frame are compressed according to the preset compression ratio.
[0103] Optionally, the preset compression ratio is 4:1. The audio signal compressed by this preset compression ratio has the least noise and strong signal stability, so that the sound distortion rate corresponding to the audio is low.
[0104] By reasonably setting the compression threshold, the dynamic range of the audio signal can be effectively controlled, making the audio more balanced and stable, and at the same time avoiding sound distortion or unnaturalness caused by excessive compression.
[0105] Specifically, the preset attack time in the dynamic compression process refers to: the time required from startup to full effect. For example, the preset attack time can be 10 ms.
[0106] The preset release time refers to: the time required from full effect to stop. For example, the preset release time can be 100 ms.
[0107] In an embodiment of the present application, when the signal level of the signal frame is the second signal level, that is, the characteristic signal value of the signal frame satisfies the preset characteristic signal range. At this time, no processing is performed on the signal frame with the second signal level.
[0108] In an embodiment of the present application, as Figure 3 shown, S4 (determining the corresponding target signal processing rule according to the signal level of each frame of signal frame) further includes the following steps:
[0109] S43: When the signal level of the signal frame is the third signal level, determine that the target signal processing rule of the signal frame is gain compensation.
[0110] For signal frames with too low signal levels, the target signal processing rule for the signal frames is gain compensation, which boosts the too low-level signals to ensure that the audio energy always remains stable within the target range, so as to prevent discomfort to the listening experience or damage to the device caused by too high-level signals.
[0111] Correspondingly, when the target signal processing rule for the signal frames is gain compensation, then S5 (processing each signal frame based on the target signal processing rule corresponding to each signal frame) specifically includes the following steps:
[0112] S52: Obtain the spectral level value of the signal frame, and calculate the initial gain compensation amount according to the spectral level value and the target spectral level value;
[0113] Specifically, the specific calculation formula for the initial gain compensation amount can be:
[0114]
[0115] In the formula, gain is the initial gain compensation amount, target dB is the target spectral level value, and current dB is the spectral level value.
[0116] S53: Smooth the initial gain compensation amount using a low-pass filter to determine the gain compensation amount of the signal frame;
[0117] Use a low-pass filter (LPF) to smooth the initial gain compensation amount and smooth the gain change of the signal, thereby avoiding the abruptness or unnatural transition caused by the gain change.
[0118] Specifically, the formula for the smoothing process can be:
[0119] gain smooth [n] = α · gain[n] + (1 - α) · gain smooth [n - 1]
[0120] In the formula, gain smooth [n] is the smoothed gain compensation amount at time point n, a is the smoothing coefficient, gain smooth [n - 1] is the smoothed gain compensation amount at time point n - 1, and gain[n] is the original gain compensation amount at time point n;
[0121] Optionally, the smoothing coefficient corresponding to the smoothing of the initial gain compensation amount by the low-pass filter can be 0.9, that is, a is 0.9.
[0122] S54: Perform gain compensation on the spectral level value of the signal frame based on the gain compensation amount of the signal frame.
[0123] Optionally, S54 (performing gain compensation on the spectral level value of the signal frame based on the gain compensation amount of the signal frame) specifically includes the following steps:
[0124] S541: When the gain compensation amount is greater than the preset gain compensation threshold, perform gain compensation on the spectral level value of the signal frame based on the preset gain compensation threshold;
[0125] By adopting the preset gain compensation threshold, background noise amplification can be avoided.
[0126] Optionally, the preset gain compensation threshold is +12dB to avoid amplifying background noise.
[0127] S542: When the gain compensation amount is less than or equal to the preset gain compensation threshold, perform gain compensation on the spectral level value of the signal frame based on the gain compensation amount.
[0128] In this application, by setting the preset gain compensation threshold, when the calculated gain compensation amount is greater than the preset gain compensation threshold, the preset gain compensation threshold is directly determined as the gain compensation. When the calculated gain compensation amount is less than or equal to the preset gain compensation threshold, the calculated gain compensation amount can be used as the gain compensation. By setting the preset maximum gain compensation threshold, background noise amplification can be avoided.
[0129] In an embodiment of this application, as Figure 4 shown, after S5 (processing each signal frame based on the target signal processing rule corresponding to each signal frame), the audio signal processing method further includes the following steps:
[0130] S6: Perform smoothing processing and channel balance processing on the audio signal of the signal frame processed by the target signal processing rule.
[0131] Optionally, the smoothing processing can perform cross-fading on the signal frame in a cross-fading smoothing processing manner to reduce the abruptness during the adjustment process. Specifically, the time window corresponding to the cross-fading smoothing processing manner is 10ms - 20ms.
[0132] Specifically, channel balance: In a multi-channel audio system (such as stereo or surround sound), the matching degree of the signal intensity (volume) and phase between each channel can ensure the uniform distribution of audio among different channels, avoiding the signal of a certain channel being too strong or too weak, thus affecting the overall auditory effect. Therefore, this application supports multi-channel audio processing and can be applied to scenarios such as in-vehicle and home audio.
[0133] When the smoothing processing and channel balance processing are completed, the processed audio signal can be output, and the delay is controlled within 50ms to meet the real-time requirement.
[0134] As a second aspect of the present application, the present application further provides a processing device for audio signals, as Figure 5 shown, a processing device 100 for audio signals provided by the present application includes:
[0135] A framing module 101, configured to perform framing processing on the audio signal to obtain multiple frame signals;
[0136] That is, the framing module 101 is configured to execute S1 (performing framing processing on the audio signal to obtain multiple frame signals) in the above-mentioned audio signal processing method.
[0137] A feature extraction module 102, configured to extract the feature signal value of each frame signal;
[0138] That is, the feature extraction module 102 is configured to execute S2 (extracting the feature signal value of each frame signal) in the above-mentioned audio signal processing method.
[0139] A signal level classification module 103, configured to determine the signal level of each frame signal according to the feature signal value of each frame signal and a preset feature signal range;
[0140] That is, the signal level classification module 103 is configured to execute S3 (determining the signal level of each frame signal according to the feature signal value of each frame signal and a preset feature signal range) in the above-mentioned audio signal processing method.
[0141] A signal processing module 104, configured to determine a corresponding target signal processing rule according to the signal level of each frame signal; process each frame signal based on the corresponding target signal processing rule of each frame signal so that the difference between the feature signal values of any two frame signals is within a preset range; wherein, when the signal levels of the frame signals are different, the corresponding target signal processing rules of the frame signals are different.
[0142] That is, the signal processing module 104 is configured to execute S4 (determining a corresponding target signal processing rule according to the signal level of each frame signal) and S5 (processing each frame signal based on the corresponding target signal processing rule of each frame signal) in the above-mentioned audio signal processing method.
[0143] The processing device for audio signals provided by the present application frames the audio signal into multiple frame signals of a certain length, extracts the feature signal value of each frame signal, determines the signal level of each frame signal according to the feature signal value, determines the corresponding target signal processing rule according to the signal level, and finally processes each frame signal according to the target signal processing rule, so that the feature signal values of the multiple frame signals are stable and consistent, thereby making the volume difference of the audio signal smaller and improving the user experience.
[0144] Exemplary electronic device
[0145] Next, as the third aspect of the present application, the present application further provides an electronic device. Refer to Figure 6 to describe the electronic device according to an embodiment of the present application.
[0146] Figure 6 The block diagram of the electronic device according to an embodiment of the present application is illustrated.
[0147] As Figure 6 shown, the electronic device 60 includes one or more processors 601 and a memory 602.
[0148] The processor 601 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 60 to perform desired functions.
[0149] The memory 602 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 601 may run the program instructions to implement a method for processing an audio signal and / or other desired functions of various embodiments of the present application described above.
[0150] In one example, the electronic device 60 may further include: an input device 603 and an output device 604, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0151] When the electronic device is a stand-alone device, the input device 603 may be a communication network connector for receiving the collected input signals from the first device and the second device.
[0152] In addition, the input device 603 may further include, for example, a keyboard, a mouse, and so on.
[0153] The output device 604 may output various information to the outside, including the determined distance information, direction information, etc. The output device 604 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0154] Of course, for simplicity, Figure 6Only some of the components related to this application in the electronic device 60 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 60 may further include any other appropriate components.
[0155] As a fifth aspect of this application, this application provides a computer-readable storage medium storing a computer program for performing the steps in a method for processing an audio signal in each of the above embodiments.
[0156] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0157] In addition to the above methods and devices, embodiments of this application may also be a computer program product, which includes computer program information that, when run by a processor, causes the processor to perform the steps in a method for processing an audio signal in various embodiments of this application.
[0158] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0159] The basic principles of this application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in this application are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of this application. In addition, the above-disclosed specific details are only for illustrative and easy-to-understand purposes and are not limitations. The above details do not limit this application to necessarily implement using the above specific details.
[0160] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0161] It should also be noted that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.
Claims
1. A method for processing an audio signal, characterized in that: include: Performing frame processing on the audio signal to obtain multiple signal frames; Extracting characteristic signal values of each signal frame; Determine the signal level of each signal frame according to the characteristic signal value of each signal frame and the preset characteristic signal range; Determine a corresponding target signal processing rule according to the signal level of each signal frame, wherein when the signal levels of the signal frames are different, the target signal processing rules corresponding to the signal frames are different; Each signal frame is processed based on a target signal processing rule corresponding to each signal frame, so that the difference between the characteristic signal values of any two signal frames is within a preset range.
2. The method for processing an audio signal according to claim 1, characterized in that: The step of determining the signal level of each signal frame according to the characteristic signal value of each signal frame and a preset characteristic signal range includes: When the characteristic signal value of the signal frame is greater than the maximum characteristic signal value in the preset characteristic signal range, determining that the signal level of the signal frame is a first signal level; When the characteristic signal value of the signal frame is greater than or equal to the minimum characteristic signal value in the preset characteristic signal range, and the characteristic signal value of the signal frame is less than or equal to the maximum characteristic signal value in the preset characteristic signal range, determining that the signal level of the signal frame is the second signal level; When the characteristic signal value of the signal frame is less than the minimum characteristic signal value in the preset characteristic signal range, the signal level of the signal frame is determined to be a third signal level.
3. The method for processing an audio signal according to claim 2, characterized in that: The characteristic signal value includes: a spectrum level value of a signal frame, and / or an overall loudness of a signal frame, and / or a maximum amplitude value of a signal frame.
4. The method for processing an audio signal according to claim 2, characterized in that: The determining of the corresponding target signal processing rule according to the signal level of each signal frame includes: When the signal level of the signal frame is a first signal level, determining that the target signal processing rule of the signal frame is dynamic compression; When the signal level of the signal frame is the third signal level, it is determined that the target signal processing rule of the signal frame is gain compensation.
5. The method for processing an audio signal according to claim 4, characterized in that: When the target signal processing rule corresponding to the signal frame is dynamic compression, The processing of each signal frame based on the target signal processing rule corresponding to each signal frame includes: The signal frame with the signal level of the first signal level is dynamically compressed according to the preset compression threshold, the preset attack time, the preset release time and the preset compression ratio.
6. The method for processing an audio signal according to claim 5, characterized in that: The preset compression ratio is 4:
1.
7. The method for processing an audio signal according to claim 4, characterized in that: When the target signal processing rule corresponding to the signal frame is gain compensation, The processing of each signal frame based on the target signal processing rule corresponding to each signal frame includes: Acquire a spectrum level value of a spectrum level whose characteristic signal of the signal frame is a spectrum level, and calculate an initial gain compensation amount according to the spectrum level value and a target spectrum level; The initial gain compensation amount is smoothed by using a low-pass filter to determine the gain compensation amount of the signal frame; Gain compensation is performed on the spectrum level value of the signal frame based on the gain compensation amount of the signal frame.
8. The method for processing an audio signal according to claim 7, characterized in that: The step of performing gain compensation on the spectrum level value of the signal frame based on the gain compensation amount of the signal frame comprises: When the gain compensation amount is greater than a preset gain compensation threshold, gain compensation is performed on the spectrum level value of the signal frame based on the preset gain compensation threshold; When the gain compensation amount is less than or equal to a preset gain compensation threshold, gain compensation is performed on the spectrum level value of the signal frame based on the gain compensation amount.
9. The method for processing an audio signal according to claim 7, characterized in that: After processing each signal frame based on the target signal processing rule corresponding to each signal frame, the audio signal processing method further includes: The audio signal of the signal frame processed by the target signal processing rule is smoothed and channel balanced.
10. An audio signal processing device, characterized in that: include: A framing module, used for framing the audio signal to obtain multiple signal frames; A feature extraction module is used to extract the feature signal value of each signal frame; A signal level classification module, used to determine the signal level of each signal frame according to the characteristic signal value of each signal frame and a preset characteristic signal range; A signal processing module is used to determine the corresponding target signal processing rule according to the signal level of each signal frame, and process each signal frame based on the target signal processing rule corresponding to each signal frame, so that the difference between the characteristic signal values of any two signal frames is within a preset range; wherein, when the signal levels of the signal frames are different, the target signal processing rules corresponding to the signal frames are different.