An audio processing method, an audio processing device, and a computer storage medium
Through dual-speaker design and high-bass frequency division design, combined with environmental noise monitoring and dynamic gain compensation technology, the problems of insufficient sound quality and noise interference in bathroom audio solutions are solved, and high-quality sound performance and intelligent adjustment are achieved.
Patent Information
- Application Number
- CN202510873067.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Existing bathroom audio solutions have shortcomings in sound quality, noise suppression and intelligent adjustment, and lack acoustic optimization design for the special environment of the bathroom.
It adopts a dual-speaker design and a high-bass frequency division design, combined with an environmental noise monitoring module, collects mixed audio signals through a microphone, dynamically adjusts the volume and sound quality parameters, and uses an adaptive delay estimation algorithm and auditory masking threshold for dynamic gain compensation, suppressing echoes and noise, and providing rich sound effects.
It achieves high-quality stereo effects in bathroom environments, suppresses background noise interference, ensures music and voice quality, provides intelligent noise reduction and adaptive adjustment, and enhances the listening experience.
Smart Images

Figure CN120390183B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sound signal processing, in particular to an audio processing method, an audio processing device and a computer storage medium. BACKGROUND
[0002] The existing bathroom sound system mostly uses a single speaker or a simple Bluetooth sound box to play music through Bluetooth connection with a mobile phone or other playback devices. Some high-end products may have basic waterproof function, but there is still a large space for improvement in sound quality performance, noise suppression, intelligent adjustment, etc.
[0003] In addition, these solutions usually lack acoustic optimization design for the special environment of the bathroom. SUMMARY
[0004] To solve the above technical problems, the present application provides an audio processing method, an audio processing device and a computer storage medium.
[0005] To solve the above technical problems, the present application provides an audio processing method, which is applied to a bathroom sound system; the audio processing method comprises:
[0006] obtaining original playback audio and mixed audio signals collected through a microphone;
[0007] extracting residual noise signals based on the original playback audio and the mixed audio signals;
[0008] obtaining a dynamic gain based on the residual noise signals and a preset auditory masking threshold;
[0009] processing the original playback audio according to the dynamic gain to obtain and output gain playback audio.
[0010] After the residual noise signals are extracted based on the original playback audio and the mixed audio signals, the audio processing method further comprises:
[0011] dividing the residual noise signals into a plurality of frequency band noise signals according to frequency;
[0012] The obtaining of the dynamic gain based on the residual noise signals and the preset auditory masking threshold comprises:
[0013] obtaining a dynamic gain of each frequency band based on each frequency band noise signal and the preset auditory masking threshold.
[0014] After the original playback audio and the mixed audio signals collected through the microphone are obtained, the audio processing method further comprises:
[0015] According to the original playing audio, a predicted bathroom space echo signal is obtained;
[0016] The predicted bathroom space echo signal is subtracted from the mixed audio signal to obtain a pure audio signal.
[0017] The processing of the original playing audio according to the dynamic gain comprises:
[0018] A key frequency band noise signal and a key frequency band dynamic gain in the plurality of frequency band noise signals are obtained.
[0019] The key frequency band dynamic gain is processed according to the strength of the key frequency band noise signal to generate a reserved frequency band dynamic gain.
[0020] The key frequency band audio in the original playing audio is processed according to the reserved frequency band dynamic gain.
[0021] The bathroom sound system comprises a first loudspeaker, a second loudspeaker, a power amplifier module, an audio effect enhancement module, a sound card module, and a microphone. The first loudspeaker and the second loudspeaker are symmetrically installed on both sides of the center area of the top of the bathroom, and the first loudspeaker and / or the second loudspeaker are built-in independent power amplifier bass units.
[0022] The audio effect enhancement module acquires the mixed audio signal through the sound card module and the microphone, and compares and processes the original playing audio to generate a dynamic gain. The audio enhancement module processes the original playing audio according to the dynamic gain to achieve gain compensation.
[0023] The audio processing method further comprises:
[0024] According to the left microphone audio signal in the mixed audio signal and the original playing audio, a left residual noise signal is extracted.
[0025] According to the right microphone audio signal in the mixed audio signal and the original playing audio, a right residual noise signal is extracted.
[0026] The signal sizes of the left residual noise signal and the right residual noise signal are compared.
[0027] The dynamic gain of the side with the larger residual noise signal is enhanced, and / or the dynamic gain of the side with the smaller residual noise signal is reduced.
[0028] After the original playing audio and the mixed audio signal acquired by the microphone are obtained, the audio processing method further comprises:
[0029] calculating a time delay difference of the original playing audio and the mixed audio signal by an adaptive time delay estimation algorithm;
[0030] dynamically aligning the original playing audio and the mixed audio signal according to the time delay difference.
[0031] The audio processing method further includes:
[0032] determining a current equalizer mode in response to a user setting instruction;
[0033] outputting the original playing audio according to the current equalizer mode;
[0034] The current equalizer mode is used to provide an initial gain.
[0035] To solve the above technical problems, the present application further provides an audio processing device, which comprises a memory and a processor coupled with the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to realize the audio processing method as described above.
[0036] To solve the above technical problems, the present application further provides a computer storage medium, which is used to store program data, and the program data is used to realize the audio processing method as described above when executed by a computer.
[0037] Compared with the prior art, the present application has the beneficial effects that: the audio processing device acquires original playing audio and mixed audio signals collected by a microphone; based on the original playing audio and the mixed audio signals, a residual noise signal is extracted; based on the residual noise signal and a preset auditory masking threshold, a dynamic gain is acquired; the original playing audio is processed according to the dynamic gain to acquire and output gain playing audio. Through the above audio processing method, dynamic gain compensation can be realized for a bathroom scene, audio output can be dynamically adjusted according to environmental noise, intelligent noise reduction and adaptive adjustment are realized, and it is ensured that the bathroom sound system provides rich and delicate sound performance. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:
[0039] Figure 1 is a frame schematic diagram of a bathroom sound system provided by the prior art;
[0040] Figure 2is a flowchart of a first embodiment of an audio processing method provided by the present application;
[0041] Figure 3 is a frame diagram of a bathroom sound system provided by the present application;
[0042] Figure 4 is a flowchart of a second embodiment of an audio processing method provided by the present application;
[0043] Figure 5 is a flowchart of a third embodiment of an audio processing method provided by the present application;
[0044] Figure 6 is a flowchart of a fourth embodiment of an audio processing method provided by the present application;
[0045] Figure 7 is a flowchart of a fifth embodiment of an audio processing method provided by the present application;
[0046] Figure 8 is a structure diagram of an embodiment of an audio processing device provided by the present application;
[0047] Figure 9 is a structure diagram of an embodiment of a computer storage medium provided by the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0049] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0050] As Figure 1As shown, the existing bathroom sound scheme generally only includes one audio source and one speaker, and can only simply realize sound playing in the bathroom environment. However, considering that noise frequently occurs in the bathroom environment, the existing bathroom sound direction can only be manually adjusted to increase the speaker volume to ensure that the audio signal of the audio source can be clearly heard, but this also greatly reduces the auditory experience. Therefore, the present application provides an audio processing method and a bathroom sound system using the same.
[0051] Specifically refer to Figure 2 and Figure 3 , Figure 2 is a flowchart of the first embodiment of the audio processing method provided by the present application, Figure 3 is a framework diagram of the bathroom sound system provided by the present application.
[0052] The audio processing method of the present application is applied to an audio processing device. The audio processing device of the present application can be a server, a terminal device, or a system comprising a server and a terminal device. Accordingly, each part of the audio processing device, such as each unit, subunit, module, and sub-module, can be provided in the server, the terminal device, or both.
[0053] Further, the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster comprising multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module, without specific limitation.
[0054] It should be noted that the audio processing device of the present application can be a bathroom sound system that is different from Figure 1 . As shown in Figure 3 , the bathroom sound system provided by the present application specifically comprises a first speaker, a second speaker, a power amplifier module, an audio effect enhancement module, a sound card module, and a microphone. The first speaker and the second speaker are symmetrically installed on both sides of the center area of the bathroom roof, and the first speaker and / or the second speaker is built-in with an independent power amplifier bass unit.
[0055] The audio effect enhancement module acquires the mixed audio signal through the sound card module and the microphone, and compares and processes the original playing audio to generate a dynamic gain. The audio enhancement module processes the original playing audio according to the dynamic gain to realize gain compensation.
[0056] Traditional bathroom audio systems typically fail to provide high-quality stereo sound, resulting in mediocre sound quality, particularly in the high and low frequencies. Therefore, the bathroom audio system proposed in this application utilizes a dual-speaker design and a high-bass crossover design to achieve richer audio layers and a wider soundstage, ensuring clear high frequencies and powerful low frequencies, thereby significantly improving sound quality.
[0057] In the traditional bathroom sound environment, background noise such as the sound of running water and electrical appliances in the bathroom will seriously affect the effect of music playback, resulting in a poor listening experience.
[0058] Therefore, the bathroom audio system of the present application integrates an environmental noise monitoring module and uses intelligent noise compensation technology to automatically adjust the volume and sound quality parameters to ensure the best listening effect under different environmental noises.
[0059] Specifically, in bathroom scenarios, ambient noise (such as running water, exhaust fans, and echoes) varies dynamically and has complex frequency bands, making it difficult for traditional fixed gain compensation techniques to maintain stable audio playback clarity. This application proposes an algorithm based on multi-band noise estimation and dynamic gain compensation. By analyzing the mixed signal of the reference signal (original audio) and the microphone capture in real time, it dynamically adjusts the gain of the played audio, ensuring that the user's perceived voice / music quality is not affected by noise interference while also suppressing bathroom echoes.
[0060] like Figure 2 As shown, the specific steps are as follows:
[0061] Step S11: obtaining the original playback audio and the mixed audio signal collected by the microphone.
[0062] In this embodiment of the present application, the original playback audio is determined from the audio source input, and then a microphone is used to collect a mixed audio signal from the environment. The mixed audio signal is generated as follows: the original playback audio is played in the bathroom environment through the power amplifier module and dual speakers, and the microphone also deployed in the bathroom environment simultaneously begins to collect the mixed audio signal from the environment, which is mainly composed of the original playback audio, the ambient noise signal, and the echo signal of the original playback audio.
[0063] After obtaining the original playback audio and mixed audio signals, this application also needs to perform signal alignment and delay compensation, specifically to synchronize the reference signal (original playback audio) and the mixed audio signal (playback audio + ambient noise + echo) collected by the microphone.
[0064] This application mainly calculates the delay difference Δt between the original playback audio and the mixed audio signal through an adaptive delay estimation algorithm, and then compensates the timestamp of the original playback audio or the timestamp of the mixed audio signal according to the delay difference to achieve dynamic alignment of the original playback audio and the mixed audio signal.
[0065] In a specific embodiment, the application can calculate the time delay difference Δt of the original play audio and mixed audio signals by the normalized cross-correlation method. The specific process is as follows:
[0066] Frame the original play audio and mixed audio signals, and the frame length N is usually 10-50 ms, such as N=2048 points under 441 Hz sampling rate, to ensure the preliminary alignment of the time axis of the two signals, which can be achieved by coarse time delay estimation or synchronization markers, etc.
[0067] Calculate the cross-correlation function of the original play audio and mixed audio signals :
[0068]
[0069] Where m is the time delay offset point number, and the search range ; is the original play audio; is the signal of the mixed audio signal after time shifting m units based on the discrete time point ; is the discrete time point.
[0070] In order to avoid the influence of signal amplitude, the cross-correlation function is normalized:
[0071]
[0072] Where the normalized , 1 represents complete matching, is the cross-correlation parameter of the normalized processing result.
[0073] Find the global maximum value of the normalized cross-correlation function:
[0074]
[0075] Where is the global maximum value of the normalized cross-correlation parameter.
[0076] In order to improve the accuracy, the points near the peak value are twice interpolated (such as parabolic fitting):
[0077]
[0078] Where is the offset point number, which represents the accurate peak value position (non-integer) calculated by interpolation, that is, the best time shift amount of the alignment of the two signals; .
[0079] Convert the offset point number to time difference:
[0080]
[0081] in, Indicates the sampling frequency of the signal.
[0082] Since the hardware delay difference between bathroom equipment playback and recording (usually 10-50ms) can lead to inaccurate noise estimation, this application dynamically aligns the signals through the above method to eliminate delay errors and provide accurate input for subsequent noise estimation.
[0083] Step S12: extracting a residual noise signal based on the original playback audio and the mixed audio signal.
[0084] In the embodiment of the present application, frequency domain subtraction is performed on the mixed audio signal and the original playback audio after dynamic alignment, the residual noise signal is extracted, and the noise energy En(f) and the signal-to-noise ratio SNR(f) are calculated.
[0085] Step S13: obtaining a dynamic gain based on the residual noise signal and a preset auditory masking threshold.
[0086] In the embodiment of the present application, the gain G(f) required for dynamic compensation is calculated based on the noise energy En(f) of the residual noise signal and the preset auditory masking threshold. The calculation formula is as follows:
[0087]
[0088] in, Target perception ability; is the smoothing factor; To prevent zero parameters.
[0089] Furthermore, the present application also requires a dynamic range limit for the above dynamic gain, such as -6dB to +12dB, to avoid excessive amplification or compression.
[0090] Among them, the auditory masking threshold ( ) is a key concept in psychoacoustics. It refers to the minimum energy threshold required for one sound (the masking sound) to completely obscure another sound (the masked sound) at a specific frequency and time. Its core function is to identify audio signal components that are imperceptible to the human ear. It is commonly used in audio compression (such as MP3), noise suppression, and virtual sound field expansion.
[0091] Since the fixed gain rule used in traditional bathroom audio solutions cannot adapt to dynamic noise, such as the sudden sound of flowing water, which will cause the gain effect to deteriorate, this application dynamically adjusts the gain of each frequency band based on the auditory masking effect to improve the gain effect.
[0092] Step S14: Process the original playback audio according to the dynamic gain, and obtain and output the gained playback audio.
[0093] In the embodiment of the present application, the original playing audio is compensated by the dynamic gain determined in step S13 to dynamically adjust the signal energy to achieve the goals of noise suppression, sound field control, or auditory masking.
[0094] In a specific implementation, the dynamic gain works as follows:
[0095] (1) Direct multiplication (frequency domain / time domain): Time-frequency domain processing: multiply the gain coefficient with the signal spectrum to suppress the noise dominant frequency band:
[0096]
[0097] wherein, is the spectrum of the original playing audio in the time-frequency domain; is the output signal spectrum after the dynamic gain compensation processing.
[0098] Time domain implementation: gain adjustment is performed on the specific frequency band through the FIR / IIR filter set.
[0099] (2) Dynamic range control:
[0100] Noise threshold: when , set to completely suppress the frequency band. Wherein, is the time-frequency energy estimation of the noise signal.
[0101] Gradual attenuation: smooth transition (such as cosine curve) is adopted near the noise threshold to avoid auditory abruptness.
[0102] In the present application, the audio processing device acquires the original playing audio and the mixed audio signal collected through the microphone; based on the original playing audio and the mixed audio signal, the residual noise signal is extracted; based on the residual noise signal and the preset auditory masking threshold, the dynamic gain is acquired; the original playing audio is processed according to the dynamic gain to acquire and output the gain playing audio. Through the above-mentioned audio processing method, dynamic gain compensation can be realized for the bathroom scene, the audio output is dynamically adjusted according to the environmental noise, intelligent noise reduction and adaptive adjustment are realized, and it is ensured that the bathroom sound provides rich and delicate sound performance.
[0103] Further, the bathroom noise frequency band is complex, and single frequency band compensation may cause voice distortion or noise residue. To this end, the present application further provides a multi-frequency band noise estimation scheme. For details, please refer to Figure 4 , Figure 4 is the flowchart of the second embodiment of the audio processing method provided by the present application.
[0104] AsFigure 4 As shown, the specific steps are as follows:
[0105] Step S21: obtaining the original playback audio and the mixed audio signal collected by the microphone.
[0106] Step S22: extracting a residual noise signal based on the original playback audio and the mixed audio signal.
[0107] Step S23: Divide the residual noise signal into several frequency band noise signals according to the frequency.
[0108] In an embodiment of the present application, a Mel-scale filter bank is used to divide the noise signal into multiple frequency bands, including but not limited to low-frequency ventilation fan sound, medium-frequency water flow sound, high-frequency water drop sound, etc., and the noise energy En(f) and signal-to-noise ratio SNR(f) of each frequency band are calculated.
[0109] Step S24: obtaining a dynamic gain of each frequency band based on the noise signal of each frequency band and a preset auditory masking threshold.
[0110] In the embodiment of the present application, the gain of each frequency band is dynamically adjusted based on the auditory masking effect.
[0111] Step S25: Process the original playback audio according to the dynamic gain of each frequency band, and obtain and output the gained playback audio.
[0112] In the embodiment of the present application, the original playback audio is processed according to the frequency band of the residual noise signal, so as to obtain the dynamic gain of each frequency band. Adjust the signal energy of the original audio played in each frequency band. After adjustment, all frequency band audios are resynthesized into a full-band signal and output to the speaker for output.
[0113] Furthermore, since global gain adjustment will cause distortion in the voice band, the present application realizes independent compensation of each frequency band on the one hand, and on the other hand, it can also retain the audio signal of the key voice frequency band (such as 500Hz-2000Hz) without gain adjustment.
[0114] Specifically, in audio processing, the 500Hz-2000Hz frequency band is considered a critical frequency band, primarily because it concentrates the core information of human speech and music and is also a key clue for sound source localization. While preserving the aforementioned critical speech frequency band does not mean completely eliminating it, it is treated with greater caution than signals in other frequency bands, and the following conditions must be met:
[0115] Priority protection: avoid excessive attenuation (e.g. gain should not be lower than 0.7).
[0116] Targeted compensation: Adjust only when noise is significantly interfering and must meet the auditory masking threshold.
[0117] wherein, for the voice key band with slight noise (SNR>15dB), the gain can be set as ≈1.0, i.e. biased to preserve the original signal integrity. For the voice key band with significant noise (SNR<6dB), the gain can be set as =0.6-0.9+smooth transition, i.e. balance noise reduction and signal loss.
[0118] The present application performs multi-band independent analysis, independent gain calculation and independent gain compensation, and accurately locates the dominant noise band.
[0119] Further, due to the serious echo in the closed bathroom space, the traditional algorithm cannot distinguish noise from echo. In this regard, the present application also provides an echo cancellation scheme. Please refer to Figure 5 , Figure 5 is a flowchart of the third embodiment of the audio processing method provided by the present application.
[0120] As shown in Figure 5 , the specific steps are as follows:
[0121] Step S31: obtaining a predicted bathroom space echo signal according to the original playback audio.
[0122] In the embodiment of the present application, an adaptive filtering algorithm (such as NLMS, Normalized Least Mean Squares Algorithm) is used to predict the bathroom space echo path.
[0123] Step S32: subtracting the predicted bathroom space echo signal from the mixed audio signal to obtain a pure audio signal.
[0124] In the embodiment of the present application, before dynamic gain calculation, the present application subtracts the predicted echo of the bathroom scene from the mixed audio signal collected by the microphone, thereby obtaining a pure noise signal, which is conducive to improving the accuracy of gain calculation.
[0125] The present application is conducive to optimizing the filter convergence speed by combining acoustic characteristic modeling (such as tile reflection coefficient).
[0126] Please continue to refer to Figure 6 , Figure 6 is a flowchart of the fourth embodiment of the audio processing method provided by the present application.
[0127] As shown in Figure 6 , the specific steps are as follows:
[0128] Step S41: extracting a left residual noise signal according to the left microphone audio signal in the mixed audio signal and the original playback audio.
[0129] Step S42: extracting a right residual noise signal according to the right microphone audio signal in the mixed audio signal and the original playing audio.
[0130] In the embodiment of the present application, the left microphone audio signal and the right microphone audio signal are distinguished in the mixed audio signal, and the difference between the two audio signals represents the audio effect at different positions in the bathroom environment.
[0131] Step S43: comparing the signal size of the left residual noise signal and the right residual noise signal.
[0132] Step S44: enhancing the dynamic gain of the side with the larger residual noise signal and / or reducing the dynamic gain of the side with the smaller residual noise signal.
[0133] In the embodiment of the present application, the left microphone noise difference is calculated, and the virtual sound source gain is increased on the side with stronger noise (e.g. the left ear) to compensate for the auditory balance. Wherein, is the noise signal estimation of the left microphone channel in the time-frequency domain, is the noise signal estimation of the right microphone channel in the time-frequency domain.
[0134] Through the independent gain of the left and right microphone audio signals, the present application can achieve the audio processing effect at any position in the bathroom scene and improve the audio experience of the bathroom sound system.
[0135] Further, the audio processing scheme of the present application can also expand the sound field through virtual stereo sound effect technology, so that an open sound landscape can be experienced even in a small space.
[0136] The virtual stereo sound effect technology expands the sound field perception range of the original audio beyond the actual position limit of the physical loudspeaker through a series of signal processing algorithms and acoustic psychology principles. The implementation process of the virtual stereo sound effect technology involved in the present application is as follows: receiving an input audio signal and extracting a sound source orientation parameter; loading a pre-stored HRTF filter set and selecting corresponding ITD / IID parameters according to the target sound image orientation; applying frequency domain equalization (EQ) and phase modulation to the left and right channels to simulate the pinna filtering effect; and processing the loudspeaker output signal through an adaptive crosstalk cancellation matrix.
[0137] Specifically, the present application can separate sound sources from mixed audio using independent component analysis (ICA) or non-negative matrix factorization (NMF) without relying on deep learning. In addition, the present application can also estimate the orientation based on the Gammatone filter to simulate the human ear basilar membrane frequency response and roughly estimate the sound source direction through time-frequency energy distribution.
[0138] Further, the existing device needs to adjust the volume, switch songs or modes through multiple steps, which is cumbersome and inconvenient. Different users have different preferences for tone, but the existing device often lacks flexible personalized setting options. To this end, the present application also provides a human-computer interaction scheme. For details, please refer to Figure 7 , Figure 7 is a flowchart of the fifth embodiment of the audio processing method provided by the present application.
[0139] As Figure 7 shown, the specific steps are as follows:
[0140] Step S51: In response to a user setting instruction, determine the current equalizer mode.
[0141] In the embodiments of the present application, the user can specify any equalizer (EQ) mode or customize a user equalizer (EQ) mode in the APP operation interface of the bathroom sound system by inputting a user setting instruction.
[0142] Step S52: Output the original playing audio according to the current equalizer mode.
[0143] In the embodiments of the present application, the original playing audio determined by the audio source is mainly output according to the current equalizer mode specified by the user, and the initial gain can also be determined by the current equalizer mode specified by the user.
[0144] Further, the bathroom sound system of the present application can also connect the mobile phone APP through Bluetooth, simplify the control panel or APP interface design, and provide intuitive and easy-to-use control options, including preset and custom EQ mode, one-key beautiful sound function, etc., to facilitate the user's daily operation. The bathroom sound system of the present application also provides diversified EQ tuning functions, allowing users to adjust the tone balance according to personal preferences, and saves and calls custom settings through the mobile phone APP, meeting the hearing needs of different users.
[0145] The audio processing method of the present application provides more rich and delicate sound effects through high and low frequency division design and double speaker layout; can dynamically adjust the output according to the environmental noise to ensure that the music is always clear and audible; through the mobile phone APP, the sound effect setting can be easily adjusted, including preset and custom EQ mode, and one-key beautiful sound function, which greatly improves the flexibility and satisfaction of use.
[0146] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0147] To realize the audio processing method, the application further provides an audio processing device, please refer to Figure 8 , Figure 8 is a structural schematic diagram of an embodiment of the audio processing device provided by the application.
[0148] The audio processing device 700 of the embodiment comprises a processor 71, a memory 72, an input and output device 73 and a bus 74.
[0149] The processor 71, the memory 72 and the input and output device 73 are connected with the bus 74 respectively, the memory 72 stores program data, and the processor 71 is used to execute the program data to realize the audio processing method described in the above embodiments.
[0150] In the embodiment of the application, the processor 71 can also be called a CPU (Central Processing Unit). The processor 71 can be an integrated circuit chip with signal processing capability. The processor 71 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 71 can also be any conventional processor.
[0151] The application further provides a computer storage medium, please continue to refer to Figure 9 , Figure 9 is a structural schematic diagram of an embodiment of the computer storage medium provided by the application. The computer storage medium 600 stores a computer program 61. The computer program 61 is used to realize the audio processing method of the above embodiments when executed by a processor.
[0152] The embodiments of the present application are realized in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0153] The above is only the embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. An audio processing method, characterized in that: The audio processing method is applied to a bathroom sound system; the audio processing method includes: Get the original playback audio and the mixed audio signal collected by the microphone; Extracting a residual noise signal based on the original playback audio and the mixed audio signal; Obtaining a dynamic gain based on the residual noise signal and a preset auditory masking threshold; Processing the original playback audio according to the dynamic gain, obtaining and outputting the gained playback audio; After extracting the residual noise signal based on the original played audio and the mixed audio signal, the audio processing method further includes: Dividing the residual noise signal into a plurality of frequency band noise signals according to frequency; The obtaining of a dynamic gain based on the residual noise signal and a preset auditory masking threshold comprises: Obtaining the dynamic gain of each frequency band based on the noise signal of each frequency band and the preset auditory masking threshold; The processing of the original playback audio according to the dynamic gain includes: Acquire a key frequency band noise signal and a key frequency band dynamic gain thereof from the plurality of frequency band noise signals; Processing the dynamic gain of the key frequency band according to the strength of the noise signal in the key frequency band to generate the dynamic gain of the retained frequency band to balance the noise reduction and loss signal; wherein the key frequency band is the core frequency band of human speech and music; The key frequency band audio in the original playback audio is processed according to the retained frequency band dynamic gain.
2. The audio processing method according to claim 1, wherein: After obtaining the original playback audio and the mixed audio signal collected by the microphone, the audio processing method further includes: Obtaining a predicted bathroom space echo signal based on the original playback audio; The predicted bathroom space echo signal is subtracted from the mixed audio signal to obtain a pure audio signal.
3. The audio processing method according to claim 1, wherein: The bathroom audio system includes a first speaker, a second speaker, an amplifier module, a sound enhancement module, a sound card module, and a microphone; the first speaker and the second speaker are symmetrically mounted on both sides of the center area of the bathroom ceiling, and the first speaker and / or the second speaker have built-in high and low frequency units with independent amplifiers; Among them, the sound enhancement module collects the mixed audio signal through the sound card module and the microphone, and compares it with the original playback audio to generate a dynamic gain; the sound enhancement module processes the original playback audio according to the dynamic gain to achieve gain compensation.
4. The audio processing method according to claim 3, characterized in that: The audio processing method further includes: Extracting a left residual noise signal according to the left microphone audio signal and the original playback audio in the mixed audio signal; Extracting a right residual noise signal according to the right microphone audio signal and the original playback audio in the mixed audio signal; comparing the signal magnitudes of the left residual noise signal and the right residual noise signal; The dynamic gain of the side with a larger residual noise signal is increased, and / or the dynamic gain of the side with a smaller residual noise signal is decreased.
5. The audio processing method according to claim 1, wherein: After obtaining the original playback audio and the mixed audio signal collected by the microphone, the audio processing method further includes: Calculating the delay difference between the original played audio and the mixed audio signal by an adaptive delay estimation algorithm; The original playback audio and the mixed audio signal are dynamically aligned according to the time delay difference.
6. The audio processing method according to claim 1, wherein: The audio processing method further includes: In response to a user setting instruction, determining a current equalizer mode; Outputting the original playback audio according to the current equalizer mode; The current equalizer mode is used to provide an initial gain.
7. An audio processing device, characterized in that: The audio processing device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the audio processing method according to any one of claims 1 to 6.
8. A computer storage medium, characterized in that The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the audio processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Volume adjustment method and device, sound equipment and storage medium
CN113194381A
Active noise control system
CN117095665A
Audio signal processing method and device, equipment and storage medium
CN117392994A
Cited By
Audio processing system and method based on active regulation and control of sound field
CN122138117A