Audio processing method, audio processing device and computer storage medium
Through dual speaker design and high-bag crossover design, combined with ambient noise monitoring and intelligent gain adjustment, the problem of insufficient sound quality and noise suppression in the bathroom audio solution is solved, achieving rich sound effects and high-quality auditory experience.
Patent Information
- Application Number
- CN202510873067.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing bathroom audio solutions have shortcomings in sound quality performance, noise suppression and intelligent adjustment, especially in special bathroom environments, which leads to poor auditory experience.
It adopts dual speaker design and high-bass crossover design, combined with the ambient noise monitoring module, collects mixed audio signals through the microphone, dynamically adjusts the gain, and uses auditory masking thresholds and adaptive delay estimation algorithm to realize multi-band noise estimation and echo cancellation, providing intelligent noise reduction and adaptive adjustment.
It realizes rich and delicate sound performance in the bathroom environment, dynamically suppresses noise interference, ensures music and voice quality, and improves listening experience.
Smart Images

Figure CN120390183A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of sound signal processing, and particularly to an audio processing method, an audio processing device, and a computer storage medium. Background Art
[0002] Existing bathroom audio solutions mostly use a single speaker or a simple Bluetooth speaker to play music by connecting to a mobile phone or other playback devices via Bluetooth. Some high-end products may have basic waterproof functions, but there is still much room for improvement in terms of sound quality performance, noise suppression, intelligent adjustment, etc.
[0003] In addition, these solutions usually lack acoustic optimization design for the special bathroom environment. Summary of the Invention
[0004] To solve the above technical problems, this application proposes an audio processing method, an audio processing device, and a computer storage medium.
[0005] To solve the above technical problems, this application proposes an audio processing method, which is applied to a bathroom audio system; the audio processing method includes: Obtain the original playback audio and the mixed audio signal collected by the microphone; Extract the residual noise signal based on the original playback audio and the mixed audio signal; Obtain the dynamic gain based on the residual noise signal and a preset auditory masking threshold; Process the original playback audio according to the dynamic gain, and obtain and output the gain playback audio.
[0006] Wherein, after extracting the residual noise signal based on the original playback audio and the mixed audio signal, the audio processing method further includes: Divide the residual noise signal into several band noise signals according to frequency; The obtaining the dynamic gain based on the residual noise signal and a preset auditory masking threshold includes: Obtain the dynamic gain of each band based on each band noise signal and a preset auditory masking threshold.
[0007] Wherein, after obtaining the original playback audio and the mixed audio signal collected by the microphone, the audio processing method further includes: Obtain the predicted bathroom space echo signal according to the original playback audio; Subtract the predicted bathroom space echo signal from the mixed audio signal to obtain a pure audio signal.
[0008] Among them, the processing of the original played audio according to the dynamic gain includes: Obtain the key frequency band noise signal and its key frequency band dynamic gain in the plurality of frequency band noise signals; Process the key frequency band dynamic gain according to the strength of the key frequency band noise signal to generate a reserved frequency band dynamic gain; Process the key frequency band audio in the original played audio according to the reserved frequency band dynamic gain.
[0009] Among them, the bathroom sound system includes a first speaker, a second speaker, a power amplifier module, a sound effect enhancement module, a sound card module, and a microphone; the first speaker and the second speaker are symmetrically installed on both sides of the central area at the top of the bathroom, and the first speaker and / or the second speaker are internally provided with high and low sound units of an independent power amplifier; Among them, the sound effect enhancement module collects the mixed audio signal through the sound card module and the microphone, and compares and processes it with the original played audio to generate a dynamic gain; the audio enhancement module processes the original played audio according to the dynamic gain to achieve gain compensation.
[0010] Among them, the audio processing method further includes: Extract the left residual noise signal according to the left microphone audio signal in the mixed audio signal and the original played audio; Extract the right residual noise signal according to the right microphone audio signal in the mixed audio signal and the original played audio; Compare the signal magnitudes of the left residual noise signal and the right residual noise signal; Enhance the dynamic gain of the side with the larger residual noise signal, and / or reduce the dynamic gain of the side with the smaller residual noise signal.
[0011] Among them, after obtaining the original played audio and the mixed audio signal collected by the microphone, the audio processing method further includes: Calculate the time delay difference between the original played audio and the mixed audio signal through an adaptive time delay estimation algorithm; Perform dynamic alignment on the original played audio and the mixed audio signal according to the time delay difference.
[0012] Among them, the audio processing method further includes: Respond to the user setting instruction to determine the current equalizer mode; Output the original played audio according to the current equalizer mode; Among them, the current equalizer mode is used to provide an initial gain.
[0013] To solve the above technical problems, the present application also proposes an audio processing device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the audio processing method as described above.
[0014] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the above audio processing method.
[0015] Compared with the prior art, the beneficial effects of the present application are as follows: The audio processing device acquires the original playback audio and the mixed audio signal collected by the microphone; based on the original playback audio and the mixed audio signal, the residual noise signal is extracted; based on the residual noise signal and the preset auditory masking threshold, the dynamic gain is obtained; the original playback audio is processed according to the dynamic gain, and the gain playback audio is obtained and output. Through the above audio processing method, dynamic gain compensation can be realized for the bathroom scenario, the audio output can be dynamically adjusted according to the environmental noise, intelligent noise reduction and adaptive adjustment can be achieved, and rich and delicate sound effects can be ensured for the bathroom audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them: Figure 1 is a schematic framework diagram of a bathroom audio system provided by the prior art; Figure 2 is a schematic flowchart of the first embodiment of the audio processing method provided by the present application; Figure 3 is a schematic framework diagram of a bathroom audio system provided by the present application; Figure 4 is a schematic flowchart of the second embodiment of the audio processing method provided by the present application; Figure 5 is a schematic flowchart of the third embodiment of the audio processing method provided by the present application; Figure 6 is a schematic flowchart of the fourth embodiment of the audio processing method provided by the present application; Figure 7 is a schematic flowchart of the fifth embodiment of the audio processing method provided by the present application; Figure 8 is a schematic structural diagram of an embodiment of the audio processing device provided by the present application; Figure 9 It is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. Specific embodiments
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0018] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0019] As Figure 1 shown, the existing bathroom audio scheme generally only includes one audio source and one speaker, and can only simply realize sound playback in the bathroom environment. However, considering the frequent occurrence of noise in the bathroom environment, the existing bathroom audio direction can only ensure that the audio signal of the audio source can be clearly heard by manually adjusting the speaker volume to increase, but this also greatly reduces the auditory experience. Therefore, the present application provides an audio processing method and a bathroom audio system using the audio processing method based on this.
[0020] Specifically, please refer to Figure 2 and Figure 3 , Figure 2 is a schematic flowchart of the first embodiment of the audio processing method provided by the present application, Figure 3 is a schematic framework diagram of the bathroom audio system provided by the present application.
[0021] The audio processing method of the present application is applied to an audio processing device. Herein, the audio processing device of the present application can be a server, a terminal device, or a system formed by the cooperation of a server and a terminal device. Correspondingly, each part included in the audio processing device, such as each unit, subunit, module, and submodule, can be all arranged in the server, all arranged in the terminal device, or respectively arranged in the server and the terminal device.
[0022] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module, and no specific limitation is made herein.
[0023] It should be noted that the audio processing device of the present application can specifically be a bathroom audio system different from Figure 1 . Specifically, as Figure 3 shown, the bathroom audio system provided by the present application specifically includes a first speaker, a second speaker, a power amplifier module, a sound effect enhancement module, a sound card module, and a microphone. The first speaker and the second speaker are symmetrically installed on both sides of the central area at the top of the bathroom, and the first speaker and / or the second speaker are internally provided with high and low frequency units with independent power amplifiers.
[0024] Among them, the sound effect enhancement module collects the mixed audio signal through the sound card module and the microphone, and compares and processes it with the original playback audio to generate a dynamic gain; the audio enhancement module processes the original playback audio according to the dynamic gain to achieve gain compensation.
[0025] Traditional bathroom audio equipment usually cannot provide high-quality stereo effects, and the sound quality performance is mediocre, especially in the performance of high frequencies and low frequencies. Therefore, the bathroom audio system of the present application adopts a dual-speaker design and a high and low frequency crossover design to achieve richer audio levels and a wider sound field, ensuring clear high frequencies and thick low frequencies, thereby greatly improving the sound quality.
[0026] In a traditional bathroom audio environment, background noises such as the sound of running water and electrical appliances in the bathroom will seriously affect the music playback effect, resulting in a poor listening experience.
[0027] Therefore, the bathroom audio system of the present application integrates an environmental noise monitoring module and uses intelligent noise compensation technology to automatically adjust the volume size and sound quality parameters to ensure the best listening effect under different environmental noises.
[0028] Specifically, in a bathroom scenario, the ambient noise (such as water flow sound, ventilation fan sound, echo) changes dynamically and has a complex frequency band. It is difficult for traditional fixed-gain compensation techniques to stably maintain the clarity of audio playback. This application proposes an algorithm based on multi-band noise estimation and dynamic gain compensation. By analyzing the reference signal (original audio) and the mixed signal collected by the microphone in real time, the gain of the playback audio is dynamically adjusted to ensure that the perceived voice / music quality by the user is not interfered by noise, and at the same time, bathroom echo is suppressed.
[0029] As Figure 2 shown, the specific steps are as follows: Step S11: Obtain the original playback audio and the mixed audio signal collected by the microphone.
[0030] In the embodiment of this application, the original playback audio is determined from the audio source input of the audio source, and then the mixed audio signal of the environment is collected by the microphone. Among them, the generation process of the mixed audio signal is: the original playback audio is played in the bathroom environment through the power amplifier module and the dual speakers, and the microphone also deployed in the bathroom environment starts to collect the mixed audio signal in the environment at the same time, which mainly consists of the original playback audio, the ambient noise signal, and the echo signal of the original playback audio.
[0031] After obtaining the original playback audio and the mixed audio signal, this application also needs to perform signal alignment and delay compensation, specifically to perform time synchronization on the reference signal (original playback audio) and the mixed audio signal collected by the microphone (playback audio + ambient noise + echo).
[0032] This application mainly calculates the delay difference Δt between the original playback audio and the mixed audio signal through the adaptive delay estimation algorithm, and then compensates the timestamp of the original playback audio or the timestamp of the mixed audio signal according to the delay difference to achieve the dynamic alignment of the original playback audio and the mixed audio signal.
[0033] In a specific implementation manner, this application can calculate the delay difference Δt between the original playback audio and the mixed audio signal through the normalized cross-correlation method. The specific process is as follows: The original playback audio and the mixed audio signal are frame-processed. The frame length N is usually 10 - 50 ms. For example, at a sampling rate of 441 Hz, N = 2048 points, ensuring the preliminary alignment of the two signal time axes, which can be achieved through rough delay estimation or synchronization markers, etc.
[0034] Calculate the cross-correlation function of the original playback audio and the mixed audio signal :
[0035] where m is the number of delay offset points, and the search range ; is the original played audio; is the signal of the mixed audio signal time-shifted by m units based on discrete time points ; is the discrete time point.
[0036] To avoid the influence of signal amplitude, the cross-correlation function is normalized:
[0037] where, after normalization , 1 represents a perfect match, is the cross-correlation parameter the result after normalization processing.
[0038] Find the global maximum value of the normalized cross-correlation function:
[0039] where, is the global maximum value of the normalized cross-correlation parameter.
[0040] To improve the accuracy, perform quadratic interpolation (such as parabolic fitting) on the points near the peak:
[0041] where, is the number of offset points, representing the exact peak position (non-integer) calculated by interpolation, that is, the optimal time shift amount for aligning the two signals; .
[0042] Convert the number of offset points to a time difference:
[0043] where, represents the sampling frequency of the signal.
[0044] Due to the hardware delay difference (usually 10 - 50 ms) between the playback and recording of bathroom devices, the noise estimation will be inaccurate. In this application, the signals are dynamically aligned through the above method to eliminate the time delay error and provide accurate input for subsequent noise estimation.
[0045] Step S12: Extract the residual noise signal based on the original played audio and the mixed audio signal.
[0046] In the embodiment of this application, the mixed audio signal and the dynamically aligned original played audio are subjected to frequency domain subtraction to extract the residual noise signal, and the noise energy En(f) and the signal-to-noise ratio SNR(f) are calculated.
[0047] Step S13: Obtain the dynamic gain based on the residual noise signal and the preset auditory masking threshold.
[0048] In the embodiment of the present application, according to the noise energy En(f) of the residual noise signal and a preset auditory masking threshold, the gain G(f) required for dynamic compensation is calculated. The calculation formula is as follows:
[0049] Wherein, is the target perception ability; is the smoothing factor; is the anti-zero parameter.
[0050] Furthermore, the present application also needs to limit the dynamic range of the above dynamic gain, such as -6dB to +12dB, to avoid over-amplification or compression.
[0051] Among them, the auditory masking threshold ( ) is a key concept in psychoacoustics, which refers to the minimum energy threshold required for a certain sound (masking sound) to completely mask another sound (masked sound) at a specific frequency and time. Its core function is to determine the audio signal components that cannot be perceived by the human ear, and it is commonly used in scenarios such as audio compression (such as MP3), noise suppression, and virtual sound field expansion.
[0052] Since the fixed gain rule in the traditional bathroom audio scheme cannot adapt to dynamic noise, such as the sudden sound of water flow will cause the gain effect to deteriorate, the present application dynamically adjusts the gain of each frequency band based on the auditory masking effect to improve the gain effect.
[0053] Step S14: Process the original played audio according to the dynamic gain, and obtain and output the gain-played audio.
[0054] In the embodiment of the present application, the original played audio is compensated according to the dynamic gain determined in step S13 to dynamically adjust the signal energy to achieve goals such as noise suppression, sound field control, or auditory masking.
[0055] In a specific implementation manner, the action mode of the above dynamic gain is as follows: (1) Direct multiplication (frequency domain / time domain): Time-frequency domain processing: Multiply the gain coefficient with the signal spectrum to suppress the noise-dominated frequency band:
[0056] Wherein, is the spectrum of the original played audio in the time-frequency domain; is the spectrum of the output signal after the dynamic gain compensation process.
[0057] Time domain implementation: Adjust the gain of a specific frequency band through an FIR / IIR filter bank.
[0058] (2) Dynamic range control: Noise threshold: When it is the case, set to completely suppress this frequency band. Among them, is the time-frequency domain energy estimation of the noise signal.
[0059] Gradual attenuation: A smooth transition (such as a cosine curve) is adopted near the noise threshold to avoid auditory abruptness.
[0060] In this application, the audio processing device acquires the original playback audio and the mixed audio signal collected by the microphone; based on the original playback audio and the mixed audio signal, extracts the residual noise signal; based on the residual noise signal and the preset auditory masking threshold, obtains the dynamic gain; processes the original playback audio according to the dynamic gain, and obtains and outputs the gain playback audio. Through the above audio processing method, dynamic gain compensation can be achieved for the bathroom scenario, the audio output can be dynamically adjusted according to the environmental noise, intelligent noise reduction and adaptive adjustment can be realized, and it is ensured that the bathroom sound system provides rich and delicate sound effects.
[0061] Furthermore, the bathroom noise frequency band is complex, and single-band compensation will cause speech distortion or noise residue. For this reason, this application also provides a multi-band noise estimation scheme. For details, please refer to Figure 4 , Figure 4 which is the schematic flowchart of the second embodiment of the audio processing method provided by this application.
[0062] As Figure 4 shown, the specific steps are as follows: Step S21: Acquire the original playback audio and the mixed audio signal collected by the microphone.
[0063] Step S22: Based on the original playback audio and the mixed audio signal, extract the residual noise signal.
[0064] Step S23: Divide the residual noise signal into several band noise signals according to frequency.
[0065] In the embodiment of this application, the noise signal is divided into multiple frequency bands through a Mel-scale filter bank, including but not limited to: such as low-frequency exhaust fan sound, medium-frequency water flow sound, high-frequency water droplet sound, etc., and the noise energy En(f) and signal-to-noise ratio SNR(f) of each frequency band are calculated.
[0066] Step S24: Based on each band noise signal and the preset auditory masking threshold, obtain the dynamic gain of each band.
[0067] In the embodiment of this application, based on the auditory masking effect, the gain of each frequency band is dynamically adjusted.
[0068] Step S25: Process the original played audio according to the dynamic gain of each frequency band, and obtain and output the gain-played audio.
[0069] In the embodiment of the present application, the original played audio is processed by frequency band according to the frequency band of the residual noise signal, so as to adjust the signal energy of the original played audio of each frequency band according to the dynamic gain of each frequency band. After adjustment, the audio of all frequency bands is recombined into a full-frequency band signal and output to the speaker for output. Adjust the signal energy of the original played audio of each frequency band. After adjustment, recombine the audio of all frequency bands into a full-frequency band signal and output it to the speaker for output.
[0070] Furthermore, since global gain adjustment will cause distortion in the voice frequency band, on the one hand, the present application realizes independent compensation for each frequency band, and on the other hand, it can also keep the audio signals in the key voice frequency bands (such as 500 Hz - 2000 Hz) without gain adjustment.
[0071] Specifically, in audio processing, the frequency band of 500 Hz - 2000 Hz is regarded as the key frequency band, mainly because this range concentrates the core information of human speech and music, and is also the key clue for sound source localization. And the above-mentioned retention of the key voice frequency band is not completely unprocessed, but is carefully processed compared with other frequency band signals, and the following conditions need to be met: Priority protection: Avoid excessive attenuation (for example, the gain should not be lower than 0.7).
[0072] Targeted compensation: Only adjust when the noise significantly interferes, and it needs to meet the auditory masking threshold.
[0073] Among them, for the key voice frequency band with slight noise (SNR > 15 dB), the gain can be set to ≈1.0, that is, it tends to retain the integrity of the original signal. For significant noise (SNR < 6 dB), the gain can be set to = 0.6 - 0.9 + smooth transition, that is, to balance noise reduction and signal loss.
[0074] The present application analyzes each frequency band independently, calculates the gain independently, compensates the gain independently, and accurately locates the noise-dominated frequency band.
[0075] Furthermore, due to the serious echo in the closed bathroom space and the inability of traditional algorithms to distinguish noise from echo, the present application also provides an echo cancellation scheme. For details, please refer to Figure 5 , Figure 5 which is the schematic flowchart of the third embodiment of the audio processing method provided by the present application.
[0076] As Figure 5 shown, the specific steps are as follows: Step S31: Obtain the predicted bathroom space echo signal according to the original played audio.
[0077] In the embodiments of the present application, an adaptive filtering algorithm (such as NLMS, Normalized Least Mean Squares Algorithm) is used to predict the echo path of the bathroom space.
[0078] Step S32: Subtract the predicted bathroom space echo signal from the mixed audio signal to obtain a pure audio signal.
[0079] In the embodiments of the present application, before calculating the dynamic gain, the present application subtracts the predicted echo of the bathroom scene from the mixed audio signal collected by the microphone, so as to obtain a pure noise signal, which is beneficial to improving the accuracy of gain calculation.
[0080] The present application is beneficial to optimizing the filter convergence speed by combining acoustic characteristic modeling (such as tile reflection coefficient).
[0081] Please continue to refer to Figure 6 , Figure 6 which is a schematic flowchart of the fourth embodiment of the audio processing method provided by the present application.
[0082] As Figure 6 shown, the specific steps are as follows: Step S41: Extract the left residual noise signal according to the left microphone audio signal and the original playback audio in the mixed audio signal.
[0083] Step S42: Extract the right residual noise signal according to the right microphone audio signal and the original playback audio in the mixed audio signal.
[0084] In the embodiments of the present application, the left microphone audio signal and the right microphone audio signal are distinguished in the mixed audio signal, and the difference between the two audio signals characterizes the audio effects at different positions in the bathroom environment.
[0085] Step S43: Compare the signal magnitudes of the left residual noise signal and the right residual noise signal.
[0086] Step S44: Enhance the dynamic gain of the side with the larger residual noise signal, and / or reduce the dynamic gain of the side with the smaller residual noise signal.
[0087] In the embodiments of the present application, the left and right microphone noise differences are calculated, and the virtual sound source gain is increased on the side with stronger noise (such as the left ear) to compensate for the auditory balance. Among them, is the noise signal estimation of the left microphone channel in the time-frequency domain, is the noise signal estimation of the right microphone channel in the time-frequency domain.
[0088] By refining the independent gain of the left and right microphone audio signals, this application can achieve the audio processing effect at any position in the bathroom scenario, improving the audio experience of the bathroom sound system.
[0089] Furthermore, the audio processing solution of this application can also expand the sound field through virtual stereo effect technology, enabling an open soundscape to be experienced even in a small space.
[0090] Among them, the virtual stereo effect technology expands the sound field perception range of the original audio through a series of signal processing algorithms and principles of acoustic psychology, enabling it to exceed the actual position limitations of physical speakers. The implementation process of the virtual stereo effect technology involved in this application is as follows: receiving the input audio signal and extracting the sound source azimuth parameters; loading the pre-stored HRTF filter bank and selecting the corresponding ITD / IID parameters according to the target sound image azimuth; applying frequency domain equalization (EQ) and phase modulation to the left and right channels to simulate the pinna filtering effect; processing the speaker output signal through an adaptive crosstalk cancellation matrix.
[0091] Specifically, this application can use independent component analysis (ICA) or non-negative matrix factorization (NMF) to separate sound sources from mixed audio without relying on deep learning. In addition, this application can also roughly estimate the sound source direction through the azimuth estimation based on Gammatone filters, mimicking the frequency response of the human ear basilar membrane and using the time-frequency energy distribution.
[0092] Furthermore, existing devices require multiple steps to adjust the volume, switch songs or modes, which is cumbersome and not convenient. Moreover, different users have different preferences for timbre, but existing devices often lack flexible personalized setting options. In response to this, this application also provides a human-computer interaction solution. For details, please refer to Figure 7 , Figure 7 which is the flowchart of the fifth embodiment of the audio processing method provided by this application.
[0093] As Figure 7 shown, the specific steps are as follows: Step S51: Respond to the user setting instruction and determine the current equalizer mode.
[0094] In the embodiment of this application, the user can specify any equalizer (EQ) mode or customize a user equalizer (EQ) mode by inputting a user setting instruction in the APP operation interface of the bathroom sound system.
[0095] Step S52: Output the original playback audio according to the current equalizer mode.
[0096] In an embodiment of the present application, the original playback audio determined by the above-mentioned audio source is mainly output according to the current equalizer mode specified by the user, and the initial gain can also be determined by the current equalizer mode specified by the user.
[0097] Furthermore, the bathroom sound system of the present application can also connect to a mobile phone app via Bluetooth, simplifying the control panel or app interface design and providing intuitive and easy-to-use control options, including preset and custom EQ modes and a one-click beautiful sound function, to facilitate users' daily operations. The bathroom sound system of the present application also provides a variety of EQ tuning functions, allowing users to adjust the tone balance according to their personal preferences, and save and call custom settings through the mobile phone app to meet the auditory needs of different users.
[0098] The audio processing method of this application provides a richer and more delicate sound performance through the high and low frequency division design and dual speaker layout; it can dynamically adjust the output according to the ambient noise to ensure that the music is always clearly audible; the sound settings can be easily adjusted through the mobile phone APP, including preset and custom EQ modes, and one-click beautiful sound function, which greatly improves the flexibility and satisfaction of use.
[0099] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0100] To implement the above audio processing method, this application also proposes an audio processing device, please refer to Figure 8 , Figure 8 It is a structural diagram of an embodiment of an audio processing device provided by this application.
[0101] The audio processing device 700 of this embodiment includes a processor 71 , a memory 72 , an input / output device 73 , and a bus 74 .
[0102] The processor 71 , the memory 72 , and the input / output device 73 are respectively connected to a bus 74 . The memory 72 stores program data, and the processor 71 is used to execute the program data to implement the audio processing method described in the above embodiment.
[0103] In an embodiment of the present application, the processor 71 may also be referred to as a CPU (Central Processing Unit). The processor 71 may be an integrated circuit chip with the ability to process signals. The processor 71 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 71 may also be any conventional processor, etc.
[0104] The present application also provides a computer storage medium. Please continue to refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the audio processing method in the above embodiment.
[0105] When the embodiments of the present application are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0106] The above description is only the embodiments of the present application and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. An audio processing method, characterized in that, The audio processing method is applied to a bathroom audio system; the audio processing method includes: Obtain the original playback audio and the mixed audio signal collected by a microphone; Extract the residual noise signal based on the original playback audio and the mixed audio signal; Obtain the dynamic gain based on the residual noise signal and a preset auditory masking threshold; Process the original playback audio according to the dynamic gain to obtain and output the gain playback audio.
2. The audio processing method according to claim 1, wherein after extracting the residual noise signal based on the original playback audio and the mixed audio signal, the audio processing method further includes: Divide the residual noise signal into several band noise signals according to frequency; The obtaining of the dynamic gain based on the residual noise signal and a preset auditory masking threshold includes: Obtain the dynamic gain of each band based on each band noise signal and a preset auditory masking threshold.
3. The audio processing method according to claim 1 or 2, wherein after obtaining the original playback audio and the mixed audio signal collected by a microphone, the audio processing method further includes: Obtain the predicted bathroom space echo signal according to the original playback audio; Subtract the predicted bathroom space echo signal from the mixed audio signal to obtain a pure audio signal.
4. The audio processing method according to claim 2, wherein the processing of the original playback audio according to the dynamic gain includes: Obtain the key band noise signal and its key band dynamic gain in the several band noise signals; Process the key band dynamic gain according to the strength of the key band noise signal to generate the reserved band dynamic gain; Process the key band audio in the original playback audio according to the reserved band dynamic gain.
5. The audio processing method according to claim 1, wherein the bathroom audio system includes a first speaker, a second speaker, a power amplifier module, a sound effect enhancement module, a sound card module, and a microphone; the first speaker and the second speaker are symmetrically installed on both sides of the central area at the top of the bathroom, and the first speaker and / or the second speaker has built-in high and low sound units of an independent power amplifier; wherein, the sound effect enhancement module collects the mixed audio signal through the sound card module and the microphone, and compares and processes it with the original playback audio to generate the dynamic gain; the audio enhancement module processes the original playback audio according to the dynamic gain to achieve gain compensation.
6. The audio processing method according to claim 5, wherein the audio processing method further includes: Extract the left residual noise signal according to the left microphone audio signal in the mixed audio signal and the original playback audio; Extract the right residual noise signal according to the right microphone audio signal in the mixed audio signal and the original playback audio; Compare the signal magnitudes of the left residual noise signal and the right residual noise signal; Increase the dynamic gain of the side with a larger residual noise signal, and / or decrease the dynamic gain of the side with a smaller residual noise signal.
7. The audio processing method according to claim 1, wherein: After obtaining the original played audio and the mixed audio signal collected by the microphone, the audio processing method further includes: Calculating the time delay difference between the original played audio and the mixed audio signal through an adaptive time delay estimation algorithm; Dynamically aligning the original played audio and the mixed audio signal according to the time delay difference.
8. The audio processing method according to claim 1, wherein: The audio processing method further includes: Responding to a user setting instruction to determine the current equalizer mode; Outputting the original played audio according to the current equalizer mode; wherein the current equalizer mode is used to provide an initial gain.
9. An audio processing device, characterized in that, The audio processing device includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the audio processing method according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the audio processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multifunctional bathroom cell phone bracket provided with loudspeaker boxes
CN104539772A
Intelligent control switch for bathroom
CN108375923A
Volume adjustment method and device, sound equipment and storage medium
CN113194381A
Noise-dependent gain method and device, vehicle-mounted system, electronic equipment and storage medium
CN115862657A
Active noise control system
CN117095665A
Cited By
Wireless microphone system and sound mixing method
CN120897139A