Audio loudspeaker control method and device, equipment, storage medium and product
By performing frequency division processing, virtual bass processing and multi-section dynamic range compression on the audio signal, compressed audio signals are generated and audio playback is solved, and the problems of high power consumption and speaker damage in the prior art are achieved, and efficient audio playback effect is achieved.
Patent Information
- Application Number
- CN202510211644.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, using equalizers to improve low-frequency components requires very high energy, has a large power consumption and is prone to damage to the speakers, and has poor audio output effect.
By performing frequency division processing on the audio signal to be processed, virtual bass processing and multi-section dynamic range compression are performed, compressed audio signals are generated, and audio playback is performed based on the signal.
Improve the low-frequency components of the audio with lower power consumption, reduce speaker damage caused by the use of the equalizer, and effectively improve the audio output effect.
Smart Images

Figure CN120151744A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of audio processing, and in particular, to an audio external playback control method, device, equipment, storage medium and product. Background Art
[0002] In live streaming and short video applications, the speakers of devices such as mobile phones and tablets commonly used by users generally have frequency response defects, such as low-frequency loss and weak mid-low frequency energy, which may cause a large difference between the sound heard by users during audio external playback and the original audio. For example, due to the small size of the mobile phone speaker itself, the attenuation of low-frequency signals is serious, which may lead to insufficient perception of bass and drum sounds in music, and the external playback sound sounds thin and lacks a sense of power; when there is too much high frequency, the external playback sound sounds harsh. Therefore, in order to improve the user's audio experience, it is necessary to optimize the audio signal in the external playback scenario.
[0003] In digital audio processing, an equalizer (EQ, Equalizer) is usually used to optimize audio. The equalizer can enhance or weaken the sound in a specific frequency range to achieve the purpose of improving sound quality, adapting to different playback environments or creating a specific sound effect style. It provides flexibility in adjusting audio and fine control of sound details, and has important applications in music production and mixing. However, for the external playback scenario, due to reasons such as the size limitation of the speaker, using an equalizer to enhance the low-frequency component requires a high amount of energy, consumes a large amount of power and is likely to damage the speaker, and the audio external playback effect is poor. Summary of the Invention
[0004] The embodiments of the present application provide an audio external playback control method, device, equipment, storage medium and product, so as to solve the technical problem in the related art that using an equalizer to enhance the low-frequency component requires a high amount of energy, consumes a large amount of power and is likely to damage the speaker, and the audio external playback effect is poor, and can improve the audio low-frequency component with lower power consumption, reduce speaker damage, and effectively improve the audio external playback effect.
[0005] In a first aspect, the embodiments of the present application provide an audio external playback control method, including:
[0006] Performing frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mixing the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal;
[0007] The virtual bass signal is frequency-divided to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and the second low-pass signal, the third low-pass signal, and the second high-pass signal are respectively subjected to dynamic range compression processing to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal;
[0008] The second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal are mixed to obtain a compressed audio signal;
[0009] Based on the compressed audio signal, audio is externally played.
[0010] In a second aspect, an embodiment of the present application provides an audio external playback control device, including a virtual bass module, a frequency division and compression module, a signal mixing module, and an audio external playback module, where:
[0011] The virtual bass module is configured to perform frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, perform virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mix the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal;
[0012] The frequency division and compression module is configured to perform frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and respectively perform dynamic range compression processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal;
[0013] The signal mixing module is configured to mix the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal to obtain a compressed audio signal;
[0014] The audio external playback module is configured to perform audio external playback based on the compressed audio signal.
[0015] In a third aspect, an embodiment of the present application provides an audio external playback control device, including: a memory and one or more processors;
[0016] The memory is used to store one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio external playback control method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a non-volatile storage medium storing computer-executable instructions, and the computer-executable instructions are used to execute the audio external playback control method as described in the first aspect when executed by a computer processor.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product. The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program, so that the device executes the audio external playback control method as described in the first aspect.
[0020] In the embodiment of the present application, the to-be-processed audio signal is frequency-divided to obtain a first low-pass signal and a first high-pass signal. The first low-pass signal is subjected to virtual bass processing to obtain a first virtual bass low-pass signal. The first virtual bass low-pass signal and the first high-pass signal are mixed to obtain a virtual bass signal. The virtual bass signal is frequency-divided to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal. The second low-pass signal, the third low-pass signal, and the second high-pass signal are respectively subjected to dynamic range compression processing to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal. The second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal are mixed to obtain a compressed audio signal. And audio external playback is performed based on the compressed audio signal. The virtual bass processing simulates the feeling of bass without directly generating low-frequency components, and optimizes the audio external playback quality through fine dynamic compression of different frequency bands, thereby improving the audio low-frequency components with lower power consumption, reducing the situation of speaker damage caused by using an equalizer to enhance the low-frequency components, and effectively improving the audio external playback effect. Description of the Drawings
[0021] Figure 1 is a flowchart of an audio external playback control method provided by an embodiment of the present application;
[0022] Figure 2 is a flowchart of another audio external playback control method provided by an embodiment of the present application;
[0023] Figure 3 is a schematic structural diagram of an audio external playback control device provided by an embodiment of the present application;
[0024] Figure 4 is a schematic structural diagram of an audio external playback control device provided by an embodiment of the present application. Detailed Embodiments
[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following provides a more detailed description of specific embodiments of this application with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. Additionally, it should be noted that for ease of description, only parts related to this application rather than all content are shown in the drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the above-mentioned process can be terminated, but there may also be additional steps not included in the drawings. The above-mentioned process can correspond to methods, functions, procedures, subroutines, subprograms, and so on.
[0026] The audio external playback control method provided by this application can be applied to audio external playback scenarios such as mobile live streaming and group chats. Its purpose is to perform virtual bass processing and band-by-band dynamic range compression on the audio to be processed. The virtual bass processing simulates the feeling of bass without directly generating low-frequency components, and the fine dynamic compression of different frequency bands can optimize the quality of audio external playback, thereby enhancing the low-frequency components of the audio with lower power consumption and effectively improving the audio external playback effect.
[0027] In existing audio external playback control solutions, the optimization points for audio external playback mainly focus on aspects such as stereo dual speakers and spatial audio (such as surround sound and panoramic sound), pursuing immersive experiences and novel gameplay. For the basic sound quality of external playback, simple equalizer (EQ, Equalizer) and dynamic range compressor (DRC, Dynamic Range Compression) solutions are mostly adopted. However, due to the small cavity size of the speakers in devices such as mobile phones and tablets, if one wants to restore the mid-low frequency listening experience, using an equalizer to boost the mid-low frequencies requires a signal with a large amount of energy, which not only easily causes distortion but also easily leads to heat dissipation problems, and may even damage the speakers in severe cases. The dynamic range compressor focuses on providing an overall volume balance and is difficult to achieve fine control over different frequency bands. For example, when processing complex music signals, it is necessary to enhance the clarity of the vocals or reduce the muddiness of the instruments, which requires different flexible treatments for different frequency bands. Based on this, an audio external playback control method according to an embodiment of this application is provided to solve the problems in existing audio external playback control solutions that using an equalizer to boost the low-frequency components requires a high amount of energy, has a large power consumption, is prone to damaging the speakers, and has a poor audio external playback effect.
[0028] Figure 1The flowchart of an audio external playback control method provided by an embodiment of the present application is given. The audio external playback control method provided by the embodiment of the present application can be executed by an audio external playback control device, and the audio external playback control device can be implemented in a hardware and / or software manner and integrated in an audio external playback control device.
[0029] The following describes the audio external playback control method executed by the audio external playback control device as an example. Refer to Figure 1 , the audio external playback control method includes:
[0030] S110: Perform frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, perform virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mix the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal.
[0031] The audio signal to be processed provided by this solution can be each audio frame of the audio to be processed. For example, in scenarios such as mobile live broadcast and multi-person chat, each audio frame of the streaming audio data received by the audio external playback control device.
[0032] Exemplarily, obtain the audio signal to be processed that needs to be optimized for external playback, and perform frequency division processing on the audio signal to be processed according to the preset frequency division position to obtain a first low-pass signal and a first high-pass signal. Perform virtual bass (VB, Virtual Bass) processing on the first low-pass signal to obtain a first virtual bass low-pass signal. Optionally, the first low-pass signal can be subjected to virtual bass processing based on a non-linear device (NLD, Non-Linear Device) in the time domain and / or a phase vocoder (PV, PhaseVocoder) in the frequency domain to obtain a first virtual bass low-pass signal. The first low-pass signal includes the fundamental frequency (basic frequency, the lowest-frequency sine wave component in the emitted sound) of the audio signal to be processed, and the fundamental frequency can be used for subsequent harmonic generation of virtual bass (for example, generating a preset sub-harmonic corresponding to the fundamental frequency (for example, one or a combination of second harmonic, third harmonic, fourth harmonic, etc.)). Virtual bass processing can simulate the feeling of bass without directly generating low-frequency components by detecting the pitch (fundamental frequency) position of the first low-pass signal and generating harmonics.
[0033] In one embodiment, after performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, the first virtual bass low-pass signal can be mixed with the first high-pass signal (for example, adding the first virtual bass low-pass signal and the first high-pass signal), and the mixing result is used as the virtual bass signal.
[0034] S120: Perform frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and perform dynamic range compression processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal respectively to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal.
[0035] Exemplarily, perform frequency division processing on the virtual bass signal according to a preset frequency division position to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal. Optionally, the frequency division processing of the virtual bass signal can be to divide the virtual bass signal into a second low-pass signal, a third low-pass signal, and a second high-pass signal at one time, or to separate the second low-pass signal and the remaining high-pass signal from the virtual bass signal, and then separate the third low-pass signal and the second high-pass signal from the remaining high-pass signal.
[0036] In one embodiment, perform dynamic range compression (DRC, Dynamic Range Compression) processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal respectively to obtain a second compressed low-pass signal corresponding to the second low-pass signal after dynamic range compression processing, a third compressed low-pass signal corresponding to the third low-pass signal after dynamic range compression processing, and a second compressed high-pass signal corresponding to the second high-pass signal after dynamic range compression processing. After obtaining the second low-pass signal, the third low-pass signal, and the second high-pass signal by performing frequency division processing on the virtual bass signal, this solution performs dynamic range compression processing in different frequency bands to achieve multi-band dynamic range compression (MBDRC, Multi-Band Dynamic Range Compression) processing on the virtual bass signal, and optimizes the user's external listening experience through fine dynamic compression and gain compensation in different frequency bands.
[0037] In one embodiment, since signals below the first frequency (e.g., 150 Hz - 200 Hz) cannot be played by the speaker and the effect of processing this part of the signal is relatively small, in order to increase the boosting space for mid-frequency signals, before performing frequency division processing on the virtual bass signal, the virtual bass signal can be pre-filtered to delete the signals below the first frequency in the virtual bass signal, and then frequency division processing is performed on the pre-filtered virtual bass signal to improve the data processing efficiency. For speakers of mobile phones, tablets, etc., the roll-off below 200 Hz is particularly severe, often with an attenuation of more than 40 dB. Moreover, the lower the frequency, the larger the amplitude of the speaker diaphragm at the same sound pressure level, the more power consumption, the greater the non-linear effect, and the greater the damage to the speaker. Boosting this part of the energy in audio external playback optimization will additionally occupy the signal boosting space, increase power consumption, and even cause non-linear noise and damage to the speaker. And this part of the signal cannot be played by speakers of mobile phones, tablets, etc. and does not need to be retained in the external playback state. Therefore, this solution filters out this part of the low-frequency signal before performing frequency division processing on the virtual bass signal to obtain more digital signal energy space and speaker protection.
[0038] S130: Mix the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal to obtain a compressed audio signal.
[0039] Exemplarily, after implementing multi-band dynamic range compression on the virtual bass signal, the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal obtained by band-by-band dynamic range compression are mixed (e.g., adding the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal) to obtain a compressed audio signal.
[0040] S140: Perform audio external playback based on the compressed audio signal.
[0041] Exemplarily, after obtaining the compressed audio signal, audio external playback can be performed based on the compressed audio signal. Among them, the compressed audio signal uses the phenomenon of missing fundamental tones in psychoacoustics. Without directly generating low-frequency components, after frequency division of the audio signal to be processed, the feeling of bass is simulated by detecting the fundamental tone position of the low-frequency signal and generating harmonics, and different bands of the audio signal are finely dynamically compressed and gain compensated through multi-band dynamic range compression, effectively optimizing the user's external playback listening experience while meeting the real-time requirements of online live broadcast and chat scenarios.
[0042] As described above, by performing frequency division processing on the audio signal to be processed, a first low-pass signal and a first high-pass signal are obtained. The first low-pass signal is subjected to virtual bass processing to obtain a first virtual bass low-pass signal, and the first virtual bass low-pass signal is mixed with the first high-pass signal to obtain a virtual bass signal. The virtual bass signal is subjected to frequency division processing to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal. The second low-pass signal, the third low-pass signal, and the second high-pass signal are respectively subjected to dynamic range compression processing to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal. The second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal are mixed to obtain a compressed audio signal, and the compressed audio signal is used for audio playback. The virtual bass processing simulates the feeling of bass without directly generating low-frequency components, and optimizes the audio playback quality through fine dynamic compression in different frequency bands, thereby improving the low-frequency components of the audio with lower power consumption and reducing the situation of speaker damage caused by using an equalizer to enhance the low-frequency components, effectively improving the audio playback effect.
[0043] Based on the above embodiments, Figure 2 a flowchart of another audio playback control method provided by an embodiment of the present application is given. This audio playback control method is a specific implementation of the above audio playback control method. Refer to Figure 2 , this audio playback control method includes:
[0044] S210: Perform frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal. Perform virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mix the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal.
[0045] In a possible embodiment, the audio playback control method provided by this solution performs virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, including:
[0046] S211: Perform transient and steady-state detection on the first low-pass signal to obtain a transient and steady-state detection result, and determine the virtual bass processing method according to the transient and steady-state detection result.
[0047] S212: Perform virtual bass processing on the first low-pass signal according to the virtual bass processing method to obtain a first virtual bass low-pass signal.
[0048] Exemplarily, transient and steady-state detection is performed on the first low-pass signal to obtain the transient and steady-state detection result corresponding to the first low-pass signal. The transient and steady-state detection result is used to indicate whether the first low-pass signal is a transient signal or a steady-state signal. For example, it is detected whether the first low-pass signal is mainly a transient signal (the proportion of the transient signal is greater than the proportion of the steady-state signal) or mainly a steady-state signal (the proportion of the steady-state signal is greater than the proportion of the transient signal). When the first low-pass signal is mainly a transient signal, the transient and steady-state detection result is that the first low-pass signal is a transient signal. When the first low-pass signal is mainly a steady-state signal, the transient and steady-state detection result is that the first low-pass signal is a steady-state signal. Optionally, the first low-pass signal of the continuous audio signal to be processed can be extracted, the characteristic information (such as energy, spectrum, etc.) of each first low-pass signal can be extracted, and the characteristic information of adjacent first low-pass signals can be compared. When the change in the characteristic information is less than the preset change threshold, this part of the first low-pass signal is considered a steady-state signal. When the change in the characteristic information reaches the preset change threshold, this part of the first low-pass signal is considered a transient signal.
[0049] In one embodiment, the virtual bass processing method can be determined according to the transient and steady-state detection result, and the first low-pass signal is subjected to virtual bass processing according to the virtual bass processing method to obtain the first virtual bass low-pass signal. Optionally, the virtual bass processing method can include virtual bass processing based on a non-linear device and virtual bass processing based on a phase vocoder. Different virtual bass processing methods can be determined in advance for different transient and steady-state detection results. This solution can effectively reduce the delay and overhead caused by transient and steady-state separation of the signal and improve the audio external playback optimization efficiency by performing transient and steady-state detection on the first low-pass signal and performing targeted virtual bass processing on the first low-pass signal according to the virtual bass processing method corresponding to the transient and steady-state detection result.
[0050] In one embodiment, the audio external playback control method provided by this solution performs transient and steady-state detection on the first low-pass signal to obtain the transient and steady-state detection result, which may include: determining a candidate detection result obtained by performing transient and steady-state detection on the first low-pass signal; performing smoothing processing on the current candidate detection result and the historical transient and steady-state results, and determining the transient and steady-state detection result according to the smoothing processing result.
[0051] Exemplarily, the transient and steady-state detection results of the first low-pass signals corresponding to a continuous plurality of audio signals to be processed are determined and recorded, where the transient and steady-state detection result corresponding to the currently processed first low-pass signal is the current candidate detection result, and the transient and steady-state detection results corresponding to the previous preset number of first low-pass signals are the historical transient and steady-state results. Smoothing processing (such as EMA smoothing processing) is performed on the current candidate detection result and the historical transient and steady-state results, and the transient and steady-state detection result is determined according to the smoothing processing result.
[0052] For example, when the candidate detection result is a steady-state signal, and the previous preset number of consecutive historical transient and steady-state results are all steady-state signals, it can be considered that the confidence of the steady-state signal is sufficient, and the transient and steady-state detection result is a steady-state signal, while when the majority of the previous preset number of consecutive historical transient and steady-state results are transient signals, it can be considered that the confidence of the steady-state signal is insufficient, and the transient and steady-state detection result remains unchanged. Similarly, when the candidate detection result is a transient signal, and the previous preset number of consecutive historical transient and steady-state results are all transient signals, it can be considered that the confidence of the transient signal is sufficient, and the transient and steady-state detection result is a transient signal, while when the majority of the previous preset number of consecutive historical transient and steady-state results are steady-state signals, it can be considered that the confidence of the transient signal is insufficient, and the transient and steady-state detection result remains unchanged. Assuming that 1 represents a transient signal and 0 represents a steady-state signal, and a number represents the transient and steady-state detection result corresponding to the first low-pass signal of a frame of audio signal to be processed, then for the transient and steady-state detection result sequence 0000001000100000100000, the confidence of 1 is not enough, and it can be considered that the transient and steady-state detection results of this section of the first low-pass signal are all steady-state signals. For the transient and steady-state detection result sequence 00000100111110001111110000000, the confidence of 1 is sufficient, and there are two times in the middle that it will switch to a transient signal. This solution determines the transient and steady-state detection results by smoothing the steady-state results, avoiding the discontinuity in the listening experience caused by frequent switching of the virtual bass processing method, and will not immediately switch the virtual bass processing method once the appearance of a transient signal or a steady-state signal is detected, effectively improving the naturalness of the audio output.
[0053] In a possible embodiment, the audio speaker control method provided by the present solution determines a virtual bass processing method according to a transient and steady-state detection result, including: when the transient and steady-state detection result is a transient audio signal, determining the virtual bass processing method to be a virtual bass processing based on a nonlinear device; when the transient and steady-state detection result is a steady-state audio signal, determining the virtual bass processing method to be a virtual bass processing based on a phase vocoder.
[0054] Exemplarily, when the transient steady-state detection result is a transient audio signal, the virtual bass processing method is determined to be virtual bass processing based on nonlinear devices, and virtual bass processing based on nonlinear devices can effectively reduce transient blurring in audio. When the transient steady-state detection result is a steady-state audio signal, the virtual bass processing method is determined to be virtual bass processing based on a phase vocoder, and virtual bass processing based on a phase vocoder can effectively reduce intermodulation distortion in audio and improve the quality of virtual bass processing.
[0055] Among them, the virtual bass processing based on non-linear devices and the virtual bass processing based on the phase vocoder are respectively applicable to the processing of transient signals and steady-state signals. Using only one of these two schemes alone is likely to introduce artificial errors in the signals. When using the time-frequency domain (NLD-PV) hybrid processing scheme, both the overhead and latency are relatively large, making it difficult to meet the real-time requirements of online live streaming and chat scenarios. Therefore, this scheme adopts a hybrid processing method based on transient and steady-state signal detection, and performs single-branch processing according to the detection results, avoiding the overhead caused by transient and steady-state separation. At the same time, a sliding window mechanism is adopted to avoid frequent switching of the virtual bass processing method, improving the naturalness of audio playback.
[0056] In a possible embodiment, after performing frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, the fundamental frequency of the first low-pass signal can be determined, and a preset multiple (such as 2 times, 3 times, etc.) of the fundamental frequency can be determined as the cut-off frequency. Then, the first low-pass signal is downsampled to the cut-off frequency. Subsequently, virtual bass processing can be performed on the downsampled first low-pass signal to obtain a first virtual bass low-pass signal. After performing virtual bass processing on the first low-pass signal to obtain the first virtual bass low-pass signal, upsampling processing can be performed on the first virtual bass low-pass signal. Correspondingly, mixing the first virtual bass low-pass signal and the first high-pass signal to obtain a virtual bass signal can be achieved by mixing the upsampled first virtual bass low-pass signal and the first high-pass signal, which can reduce the data processing volume while ensuring the quality of virtual bass processing, effectively saving the data processing overhead.
[0057] S220: Perform dry-wet mixing processing on the audio signal to be processed and the virtual bass signal according to the first dry-wet ratio.
[0058] Exemplarily, after obtaining the virtual bass signal, dry-wet mixing processing can be performed on the audio signal to be processed and the virtual bass signal according to the first dry-wet ratio. Among them, the audio signal to be processed is the dry signal, and the virtual bass signal is the wet signal. For example, adding the product of the first dry-wet ratio and the audio signal to be processed to the product of the difference between the first preset reference parameter and the first dry-wet ratio and the virtual bass signal, and the addition result is the dry-wet mixing processing result of the audio signal to be processed and the virtual bass signal (i.e., the virtual bass signal after dry-wet mixing processing). Optionally, the first preset reference parameter can be configured as 1. Correspondingly, performing frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal can be achieved by performing frequency division processing on the virtual bass signal after dry-wet mixing processing. This scheme realizes a smooth transition of virtual bass processing by performing dry-wet mixing processing on the audio signal to be processed and the virtual bass signal, effectively improving the naturalness of audio playback.
[0059] Among them, the first dry-wet ratio provided by this solution can be determined according to a combination of one or more of the audio scene, device model, and room information. For example, the corresponding first dry-wet ratio is determined according to different combinations of one or more of different audio scenes, different device models, and different room information. The corresponding first dry-wet ratio can be determined according to the current audio scene, device model, and room information. In one embodiment, when determining the new first dry-wet ratio, a weighted method of each sampling point can be used to transition from the current first dry-wet ratio to the target first dry-wet ratio in a smoother manner, thereby reducing the sense of discontinuity in the audio output and improving the listening experience of the audio output.
[0060] S230: Perform frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal and a second high-pass signal, and perform dynamic range compression processing on the second low-pass signal, the third low-pass signal and the second high-pass signal to obtain a second compressed low-pass signal, a third compressed low-pass signal and a second compressed high-pass signal.
[0061] In a possible embodiment, the audio speaker control method provided by the present solution can also perform phase synchronization processing on the second low-pass signal to align the phase with the third low-pass signal and / or the second high-pass signal after frequency division processing is performed on the virtual bass signal to obtain the second low-pass signal, the third low-pass signal and the second high-pass signal.
[0062] Exemplarily, after the virtual bass signal is frequency-divided to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, the second low-pass signal may be phase-synchronized to be aligned with the third low-pass signal and / or the second high-pass signal. Accordingly, the second low-pass signal is subjected to dynamic range compression to obtain a second compressed low-pass signal, which may be obtained by subjecting the second low-pass signal after the phase synchronization processing to dynamic range compression. This solution performs phase synchronization processing on the second low-pass signal to synchronize the phases of the second low-pass signal with the third low-pass signal and / or the second high-pass signal, so that the divided signal can be accurately reconstructed, ensuring correct mixing to obtain a compressed audio signal and ensuring the quality of audio speaker control.
[0063] S240: Mix the second compressed low-pass signal, the third compressed low-pass signal and the second compressed high-pass signal to obtain a compressed audio signal.
[0064] S250: performing dry-wet mixing processing on the virtual bass signal and the compressed audio signal according to the second dry-wet ratio.
[0065] Exemplarily, after obtaining the compressed audio signal, the virtual bass signal and the compressed audio signal can be subjected to dry-wet mixing processing according to the second dry-wet ratio. Among them, the virtual bass signal is the dry signal, and the compressed audio signal is the wet signal. For example, the product of the second dry-wet ratio and the virtual bass signal is added to the product of the difference between the second preset reference parameter and the second dry-wet ratio and the compressed audio signal. The addition result is the dry-wet mixing processing result of the virtual bass signal and the compressed audio signal (i.e., the compressed audio signal after dry-wet mixing processing). Optionally, the second preset reference parameter can be configured as 1. Correspondingly, based on the compressed audio signal for audio playback, it can be based on the compressed audio signal after dry-wet mixing processing for audio playback. This solution realizes a smooth transition of multi-stage dynamic range compression by performing dry-wet mixing processing on the virtual bass signal and the compressed audio signal, and can effectively improve the naturalness of audio playback.
[0066] Among them, the second dry-wet ratio provided by this solution can be determined according to one or more combinations of audio scenarios, device models, room information, and device volume. For example, the corresponding second dry-wet ratio can be determined according to different combinations of different audio scenarios, different device models, different room information, and device volume. The corresponding second dry-wet ratio can be determined according to the current audio scenario, device model, room information, and device volume.
[0067] In one embodiment, when determining the new second dry-wet ratio, a weighted method of each sampling point can be adopted to transition from the current second dry-wet ratio to the target second dry-wet ratio in a smoother manner, reducing the generation of discontinuous feelings in audio playback and improving the listening experience of audio playback. Optionally, due to the physical limitations of the mobile phone speaker cavity, when the system volume is too large, the diaphragm and driver of the speaker will exhibit non-linear responses, resulting in non-linear distortion. At this time, if the audio playback control processing is superimposed, the distortion may be aggravated. This solution can detect the system volume. When the system volume increases to the set volume, the volume control is triggered. By setting the new target dry-wet ratio (the proportion of the audio signal to be processed increases, and the proportion of the virtual bass signal decreases), the effectiveness of the algorithm is gradually weakened until it slowly disappears. When the volume decreases, the target dry-wet ratio can be restored according to the same ratio, thus avoiding the negative impact caused by non-linear distortion when the speaker enters the non-linear region at high volume.
[0068] Optionally, the initial values of the dry-wet ratio (including the first dry-wet ratio and the second dry-wet ratio) can be set. The dry signal and the wet signal can be mixed according to the dry-wet ratio. During the audio external playback control process, different target dry-wet ratios can be set according to different audio scenarios, device models, room information, and device volume, etc. When it is detected that the current dry-wet ratio is different from the target dry-wet ratio, the processing effect can be applied at each sampling point to achieve a smooth transition of the sound effect. For example, in a strong noise environment, the effect of enhancing the external playback can be controlled not to take effect; in a scene dominated by music, the effect of the algorithm can be made more significant, while in a voice-dominated situation, the algorithm can take effect normally or appropriately suppress the low-frequency signal to avoid the feeling of muffled human voices. Therefore, different target dry-wet ratios can be set for the external playback algorithm according to the energy distribution of speech, music, and noise in the current scene. When it is detected that the scene changes, the algorithm will gradually take effect according to the target dry-wet ratio, achieving a smooth transition at each sampling point and optimizing the audio external playback effect. Since there are a variety of mobile phone models, different grades and models of mobile phones often have different audio processing capabilities at the underlying hardware level. High-end mobile phone chips have stronger and more diverse processing capabilities, and the system's own effects already include certain optimizations. In contrast, low-end mobile phones have relatively fewer functions, and the optimization space for software algorithms is relatively large. The target audiences of different room types have different characteristics and different focuses in gameplay, and different attentions can be given in audio optimization. This solution can set different external playback optimization parameters for different grades of mobile phones and different online room types based on the optimization methods of grading by model and differentiating room types, maximizing the effect of the algorithm.
[0069] S260: Perform audio external playback based on the compressed audio signal.
[0070] In a possible embodiment, after the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal are mixed to obtain the compressed audio signal, the audio external playback control method provided in this solution can also perform audio signal peak level control on the compressed audio signal. Correspondingly, performing audio external playback based on the compressed audio signal includes: performing audio external playback based on the compressed audio signal after audio signal peak level control.
[0071] Exemplarily, a preset audio signal peak level control tool (such as a limiter Limiter) can be used. By performing audio signal peak level control on the compressed audio signal, it can effectively prevent the audio signal from exceeding a specific level threshold, thereby protecting the audio system from audio signal clipping and distortion.
[0072] This solution uses a combined external speaker optimization method that includes virtual bass processing, multi-band dynamic range compression, and audio signal peak level control. By leveraging the pitch disappearance phenomenon in psychoacoustics, it can simulate the feeling of bass without directly generating low-frequency components. Specifically, it detects the pitch position through virtual bass processing and generates harmonics to create the bass sensation. Multi-band dynamic range compression is used to perform fine dynamic compression and gain compensation on different frequency bands to optimize the user's external speaker listening experience, and audio signal peak level control is used to avoid clipping.
[0073] Specifically, the audio signal to be processed is frequency-divided to obtain a first low-pass signal and a first high-pass signal. The first low-pass signal is subjected to virtual bass processing to obtain a first virtual bass low-pass signal, and the first virtual bass low-pass signal is mixed with the first high-pass signal to obtain a virtual bass signal. The virtual bass signal is frequency-divided to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal. The second low-pass signal, the third low-pass signal, and the second high-pass signal are respectively subjected to dynamic range compression processing to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal. The second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal are mixed to obtain a compressed audio signal, and audio is externally played based on the compressed audio signal. The virtual bass processing simulates the feeling of bass without directly generating low-frequency components, and optimizes the audio external play quality through fine dynamic compression of different frequency bands, thereby improving the audio low-frequency components with lower power consumption and reducing the situation of speaker damage caused by using an equalizer to enhance the low-frequency components, effectively improving the audio external play effect. At the same time, by performing wet-dry mixing processing on the audio signal to be processed and the virtual bass signal, the smooth transition of virtual bass processing is achieved, and by performing wet-dry mixing processing on the virtual bass signal and the compressed audio signal, the smooth transition of multi-band dynamic range compression is achieved, which can effectively improve the naturalness of audio external play.
[0074] Figure 3 It is a schematic structural diagram of an audio external play control device provided by an embodiment of the present application. Refer to Figure 3 As shown in the figure, the audio external play control device includes a virtual bass module 31, a frequency division and compression module 32, a signal mixing module 33, and an audio external play module 34.
[0075] Among them, the virtual bass module 31 is configured to perform frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, perform virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mix the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal; the frequency division compression module 32 is configured to perform frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and perform dynamic range compression processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal respectively to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal; the signal mixing module 33 is configured to mix the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal to obtain a compressed audio signal; the audio playback module 34 is configured to perform audio playback based on the compressed audio signal.
[0076] As described above, by performing frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mixing the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal, performing frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, performing dynamic range compression processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal respectively to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal, mixing the second compressed low-pass signal, the third compressed low-pass signal, and the second compressed high-pass signal to obtain a compressed audio signal, and performing audio playback based on the compressed audio signal, the virtual bass processing simulates the feeling of bass without directly generating low-frequency components, and optimizes the audio playback quality through fine dynamic compression of different frequency bands, thereby improving the low-frequency components of the audio with lower power consumption, reducing the situation of speaker damage caused by using an equalizer to enhance the low-frequency components, and effectively improving the audio playback effect.
[0077] In a possible embodiment, the audio playback control device further includes a first mixing module, and the first mixing module is configured to: perform dry-wet mixing processing on the audio signal to be processed and the virtual bass signal according to a first dry-wet ratio, and the first dry-wet ratio is determined according to a combination of one or more of the audio scene, the device model, and the room information;
[0078] Correspondingly, the frequency division compression module 32 performs frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and is configured to: perform frequency division processing on the virtual bass signal after dry-wet mixing processing to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal.
[0079] In a possible embodiment, the audio external playback control device further includes a second mixing module, which is configured to perform wet-dry mixing processing on the virtual bass signal and the compressed audio signal according to a second wet-dry ratio, and the second wet-dry ratio is determined according to a combination of one or more of the audio scenario, device model, room information, and device volume;
[0080] Correspondingly, the audio external playback module performs audio external playback based on the compressed audio signal, and is configured to perform audio external playback based on the compressed audio signal after wet-dry mixing processing.
[0081] In a possible embodiment, the audio external playback control device further includes a peak control module, which is configured to perform peak level control of the audio signal on the compressed audio signal;
[0082] Correspondingly, the audio external playback module performs audio external playback based on the compressed audio signal, and is configured to perform audio external playback based on the compressed audio signal after peak level control of the audio signal.
[0083] In a possible embodiment, the virtual bass module 31 performs virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and is configured to perform transient-steady state detection on the first low-pass signal to obtain a transient-steady state detection result, and determine a virtual bass processing method according to the transient-steady state detection result; perform virtual bass processing on the first low-pass signal according to the virtual bass processing method to obtain a first virtual bass low-pass signal.
[0084] In a possible embodiment, the virtual bass module 31 performs transient-steady state detection on the first low-pass signal to obtain a transient-steady state detection result, and is configured to determine a candidate detection result for performing transient-steady state detection on the first low-pass signal; perform smoothing processing on the current candidate detection result and the historical transient-steady state result, and determine the transient-steady state detection result according to the smoothing processing result.
[0085] In a possible embodiment, the virtual bass module 31 determines a virtual bass processing method according to the transient-steady state detection result, and is configured to determine that the virtual bass processing method is virtual bass processing based on a non-linear device when the transient-steady state detection result is a transient audio signal; determine that the virtual bass processing method is virtual bass processing based on a phase vocoder when the transient-steady state detection result is a steady-state audio signal.
[0086] In a possible embodiment, the audio external playback control device further includes a downsampling module and an upsampling module. The downsampling module is configured to determine the fundamental frequency of the first low-pass signal, determine a preset multiple of the fundamental frequency as the cut-off frequency; downsample the first low-pass signal to the cut-off frequency; the upsampling module is configured to perform upsampling processing on the first virtual bass low-pass signal;
[0087] Correspondingly, the virtual bass module 31 mixes the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal, configured to: mix the upsampled first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal.
[0088] In a possible embodiment, the audio external playback control device further includes a phase alignment module, configured to: perform phase synchronization processing on the second low-pass signal to align the phase with the third low-pass signal and / or the second high-pass signal;
[0089] Correspondingly, the frequency division compression module 32 performs dynamic range compression processing on the second low-pass signal to obtain a second compressed low-pass signal, configured to: perform dynamic range compression processing on the phase-synchronized second low-pass signal to obtain a second compressed low-pass signal.
[0090] It should be noted that in the embodiments of the above audio external playback control device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present application.
[0091] The embodiments of the present application further provide an audio external playback control device, and the audio external playback control device may integrate the audio external playback control device provided by the embodiments of the present application. Figure 4 It is a schematic structural diagram of an audio external playback control device provided by the embodiments of the present application. Refer to Figure 4 , the audio external playback control device includes: an input device 43, an output device 44, a memory 42, and one or more processors 41; the memory 42 is used to store one or more programs; when the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the audio external playback control method provided by the above embodiments. The above provided audio external playback control device, device and computer can be used to execute the audio external playback control method provided by any of the above embodiments, and have the corresponding functions and beneficial effects.
[0092] The embodiments of the present application further provide a non-volatile storage medium storing computer-executable instructions, and the computer-executable instructions are used to execute the audio external playback control method provided in the above embodiments when executed by a computer processor. Of course, for the non-volatile storage medium storing computer-executable instructions provided in the embodiments of the present application, the computer-executable instructions are not limited to the audio external playback control method provided above, and can also execute relevant operations in the audio external playback control methods provided in any embodiment of the present application. The audio external playback control device, equipment and storage medium provided in the above embodiments can execute the audio external playback control method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, reference can be made to the audio external playback control method provided in any embodiment of the present application.
[0093] On the basis of the above embodiments, the embodiments of the present application further provide a computer program product. Essentially, or the part that contributes to the prior art, or all or part of the technical solution of the present application can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions for causing a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the audio external playback control method provided in each embodiment of the present application.
Claims
1. A method for controlling an audio speaker, characterized in that: include: Performing frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mixing the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal; Performing frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal, and a second high-pass signal, and performing dynamic range compression processing on the second low-pass signal, the third low-pass signal, and the second high-pass signal to obtain a second compressed low-pass signal, a third compressed low-pass signal, and a second compressed high-pass signal; Mixing the second compressed low-pass signal, the third compressed low-pass signal and the second compressed high-pass signal to obtain a compressed audio signal; Audio is played out based on the compressed audio signal.
2. The audio speaker control method according to claim 1, characterized in that: After the first virtual bass low-pass signal is mixed with the first high-pass signal to obtain a virtual bass signal, the method further includes: Performing dry-wet mixing processing on the audio signal to be processed and the virtual bass signal according to a first dry-wet ratio, wherein the first dry-wet ratio is determined according to a combination of one or more of an audio scene, a device model, and room information; Correspondingly, the frequency division processing of the virtual bass signal to obtain a second low-pass signal, a third low-pass signal and a second high-pass signal includes: The virtual bass signal after the dry-wet mixing process is subjected to frequency division processing to obtain a second low-pass signal, a third low-pass signal and a second high-pass signal.
3. The audio speaker control method according to claim 1, characterized in that: After the second compressed low-pass signal, the third compressed low-pass signal and the second compressed high-pass signal are mixed to obtain a compressed audio signal, the method further includes: performing dry-wet mixing processing on the virtual bass signal and the compressed audio signal according to a second dry-wet ratio, wherein the second dry-wet ratio is determined according to a combination of one or more of an audio scene, a device model, room information, and a device volume; Correspondingly, the performing audio playback based on the compressed audio signal includes: Audio is played out based on the compressed audio signal after the dry-wet mixing process.
4. The audio speaker control method according to claim 1, characterized in that: After the second compressed low-pass signal, the third compressed low-pass signal and the second compressed high-pass signal are mixed to obtain a compressed audio signal, the method further includes: Performing audio signal peak level control on the compressed audio signal; Correspondingly, the performing audio playback based on the compressed audio signal includes: Audio is played out based on the compressed audio signal after the peak level of the audio signal is controlled.
5. The audio speaker control method according to claim 1, characterized in that: The performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal includes: Performing transient state detection on the first low-pass signal to obtain a transient state detection result, and determining a virtual bass processing method according to the transient state detection result; The first low-pass signal is subjected to virtual bass processing according to the virtual bass processing method to obtain a first virtual bass low-pass signal.
6. The audio speaker control method according to claim 5, characterized in that: The performing transient and steady-state detection on the first low-pass signal to obtain a transient and steady-state detection result includes: Determine that the first low-pass signal is subjected to transient steady-state detection to obtain a candidate detection result; The current candidate detection result and the historical transient and steady-state results are smoothed, and the transient and steady-state detection result is determined according to the smoothing result.
7. The audio speaker control method according to claim 5, characterized in that: The determining of the virtual bass processing method according to the transient steady-state detection result includes: When the transient-steady-state detection result is a transient audio signal, determining that the virtual bass processing method is a virtual bass processing based on a nonlinear device; When the transient steady-state detection result is a steady-state audio signal, the virtual bass processing method is determined to be a phase vocoder-based virtual bass processing.
8. The audio speaker control method according to claim 1, characterized in that: After the audio signal to be processed is subjected to frequency division processing to obtain the first low-pass signal and the first high-pass signal, the method further includes: Determine a fundamental frequency of the first low-pass signal, and determine a preset multiple of the fundamental frequency as a cutoff frequency; downsampling the first low-pass signal to the cutoff frequency; After performing virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, the method further includes: performing up-sampling processing on the first virtual bass low-pass signal; Correspondingly, the step of mixing the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal includes: The first virtual bass low-pass signal after upsampling is mixed with the first high-pass signal to obtain a virtual bass signal.
9. The audio speaker control method according to claim 1, characterized in that: After the virtual bass signal is subjected to frequency division processing to obtain a second low-pass signal, a third low-pass signal and a second high-pass signal, the method further includes: Processing the second low-pass signal in phase synchronization to be aligned with the third low-pass signal and / or the second high-pass signal; Correspondingly, performing dynamic range compression processing on the second low-pass signal to obtain a second compressed low-pass signal includes: Dynamic range compression is performed on the second low-pass signal after the phase synchronization processing to obtain a second compressed low-pass signal.
10. An audio external speaker control device, characterized in that: It includes a virtual bass module, a frequency division compression module, a signal mixing module and an audio amplifier module, among which: The virtual bass module is configured to perform frequency division processing on the audio signal to be processed to obtain a first low-pass signal and a first high-pass signal, perform virtual bass processing on the first low-pass signal to obtain a first virtual bass low-pass signal, and mix the first virtual bass low-pass signal with the first high-pass signal to obtain a virtual bass signal; The frequency division compression module is configured to perform frequency division processing on the virtual bass signal to obtain a second low-pass signal, a third low-pass signal and a second high-pass signal, and perform dynamic range compression processing on the second low-pass signal, the third low-pass signal and the second high-pass signal to obtain a second compressed low-pass signal, a third compressed low-pass signal and a second compressed high-pass signal; The signal mixing module is configured to mix the second compressed low-pass signal, the third compressed low-pass signal and the second compressed high-pass signal to obtain a compressed audio signal; The audio player module is configured to play audio based on the compressed audio signal.
11. An audio speaker control device, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the audio speaker control method as described in any one of claims 1-9.
12. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions are used to execute the audio speaker control method as described in any one of claims 1-9 when executed by a computer processor.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the audio speaker control method according to any one of claims 1 to 9 is implemented.