An audio enhancement method, apparatus, electronic device, and storage medium
Patent Information
- Application Number
- CN202510229269.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2026-08-28
AI Technical Summary
这就导致增强后的音频的音质效果较差,用户体验感较低
[0026] Fourthly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the audio enhancement method provided in the first aspect above.
Smart Images

Figure CN122658331A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal equipment technology, and in particular to an audio enhancement method, apparatus, electronic device and storage medium. Background Technology
[0002] With advancements in audio technology, portable electronic devices can now provide a relatively complete and immersive sound experience even without external speakers, thanks to their built-in audio components. This is because audio, as an electronic signal, can be captured, stored, and processed using various technologies, enabling electronic devices to accurately reproduce rich sound effects. However, despite these significant advancements, many portable electronic products, due to size, hardware limitations, and space constraints, often struggle to deliver high-quality sound and fail to fully meet users' audio experience needs. To address this, audio signal enhancement is often necessary.
[0003] Currently, audio enhancement techniques such as compressor technology or virtual bass technology are commonly used. However, these enhancement methods mainly optimize and enhance the overall audio. This results in poor sound quality and a lower user experience after enhancement. Summary of the Invention
[0004] To address the aforementioned issues, embodiments of this application provide an audio enhancement method, apparatus, electronic device, and storage medium, which can perform targeted audio enhancement based on timbre, thereby improving the user experience.
[0005] To achieve the above objectives, in a first aspect, embodiments of this application provide an audio enhancement method, comprising: acquiring a first frequency domain signal and a first time domain signal corresponding to an original audio; determining at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and a first spectral feature corresponding to at least one timbre template, wherein the at least one timbre template contains at least one target timbre; extracting a second frequency domain signal corresponding to the target timbre from the first frequency domain signal; adjusting the second frequency domain signal based on the target spectral feature corresponding to the target timbre to obtain a third frequency domain signal; wherein the target spectral feature is the spectral feature corresponding to the target timbre among at least one first spectral feature; acquiring a second time domain signal corresponding to the third frequency domain signal; enhancing the second time domain signal based on the sound effect parameters corresponding to the target timbre to obtain a third time domain signal; and mixing the third time domain signal with the first time domain signal to obtain a first target audio signal.
[0006] The audio enhancement method provided in this application analyzes the time-domain and frequency-domain signals of the original audio, precisely extracts and adjusts the target timbre that matches the timbre template, and makes targeted adjustments to the target timbre to obtain an optimized first target audio signal. This allows for precise modification of the timbre quality of a specific timbre in the complex original audio, avoiding impact on other timbres. This makes the audio enhancement process more personalized and customized, ensuring that the optimized first audio signal better meets expectations and significantly improving the overall audio quality.
[0007] In one feasible implementation, obtaining a first frequency domain signal and a first time domain signal corresponding to the original audio includes: obtaining a first original time domain signal corresponding to the original audio; segmenting the first original time domain signal into frames to obtain a second original time domain signal after segmentation; suppressing the second original time domain signal based on a spectrum leakage suppression strategy to obtain a first time domain signal; and performing frequency domain transformation on the first time domain signal to obtain a first frequency domain signal. By using the above method, the first original time domain signal corresponding to the original audio is obtained, and the first original time domain signal is processed by framing and a spectrum leakage suppression strategy. The processed first time domain signal is then transformed to obtain a first frequency domain signal. In this way, by framing and suppressing spectrum leakage, the transformed signal is more accurate, and the optimized frequency domain signal makes the frequency components clearer, providing a more stable foundation for subsequent audio enhancement.
[0008] In one feasible implementation, based on the matching result of a first frequency domain signal and a first spectral feature corresponding to at least one timbre template, at least one target timbre in the original audio is determined. This includes: acquiring a second spectral feature of the first frequency domain signal; the second spectral feature being used to characterize the timbre feature corresponding to the first frequency domain signal; and, if the second spectral feature successfully matches the first spectral feature corresponding to at least one timbre template, acquiring at least one target timbre in the original audio. Using the above method, based on the first spectral feature corresponding to the timbre template, the timbre that needs to be enhanced in the original audio, i.e., the target timbre, is extracted. This allows for accurate identification of the target timbre that needs enhancement in the original audio, making the audio enhancement process more targeted, ensuring the enhancement effect is natural and conforms to the desired timbre feature, thereby improving audio quality and listening experience.
[0009] In one feasible implementation, when each second spectral feature successfully matches a first spectral feature corresponding to at least one timbre template, at least one target timbre contained in the original audio is determined. This includes: determining a matching confidence level based on the first and second spectral features; the matching confidence level is used to characterize the similarity between the spectral features of the target timbre and the spectral features of the timbre template; and determining that at least one target timbre contained in the original audio is present when the matching confidence level is greater than a preset threshold. Using the above method, by comparing the similarity between the second spectral feature corresponding to the first frequency domain signal and the first spectral feature corresponding to the timbre template, and determining that the similarity reaches a preset threshold after calculating the similarity, the target timbre is identified. Thus, by setting a preset threshold, target timbres that meet the requirements can be filtered out, thereby facilitating subsequent targeted enhancement of the original audio.
[0010] In one feasible implementation, determining the matching confidence based on a first spectral feature and a second spectral feature includes: determining a first divergence based on the first and second spectral features; the first divergence characterizes the degree of deviation between the spectral features of the target timbre and the spectral features of the timbre template; and normalizing the first divergence to obtain the matching confidence. Normalization is used to determine the matching confidence of the first and second spectral features. This eliminates the differences in dimensions and orders of magnitude between different features, ensuring that the first and second spectral features are on the same baseline when compared, thereby improving the accuracy and reliability of the matching confidence calculation.
[0011] In one feasible implementation, a third frequency domain signal is obtained by adjusting the second frequency domain signal based on the target spectral features corresponding to the target timbre. This includes: acquiring the third spectral features of the second frequency domain signal; using the third spectral features to characterize the timbre features corresponding to the second frequency domain signal; and replacing the third spectral features in the second frequency domain signal with the target spectral features to obtain the third frequency domain signal, wherein the timbre corresponding to the third spectral features is the same as the timbre corresponding to the target spectral features. By replacing the third spectral features corresponding to each target timbre in the original audio with the corresponding target spectral features, a replaced third frequency domain signal is obtained. This reduces interference from irrelevant spectral components, such as external noise, improves the clarity of the timbre in the original audio, makes the processed audio more consistent with the expected timbre features, and enhances the overall sound quality and listening experience.
[0012] In one feasible implementation, obtaining the second time-domain signal corresponding to the third frequency-domain signal includes: performing a time-domain transformation on the third frequency-domain signal to obtain the second time-domain signal. By using the above method, the second time-domain signal is obtained through transformation, thereby restoring the optimized target timbre's performance in the time domain. This not only reduces noise and impurities to a certain extent but also preserves the frequency-domain optimized timbre characteristics in the time domain, making the final output audio more natural and smooth, and improving the overall sound quality and listening experience.
[0013] In one feasible implementation, a second time-domain signal is enhanced based on sound effect parameters corresponding to the target timbre to obtain a third time-domain signal. This includes: obtaining sound effect parameters corresponding to each target spectral feature from a preset timbre database; and enhancing the second time-domain signal based on each sound effect parameter to obtain the third time-domain signal. By using this method, corresponding sound effect parameters are obtained from a pre-set timbre database, and the second time-domain signal is enhanced based on these parameters. In this way, the sound effect parameters not only enhance the detail of the audio but also make the final third time-domain signal more suitable for user needs, improving the overall listening experience.
[0014] In one feasible implementation, mixing the third time-domain signal with the first time-domain signal to obtain the first target audio signal includes: determining a first weight corresponding to the first time-domain signal and a second weight corresponding to the third time-domain signal based on a preset strategy; wherein the preset strategy includes the number of timbres in the original audio; and mixing the first time-domain signal and the third time-domain signal based on the first weight and the second weight to obtain the first target audio signal. By using the above method, the third time-domain signal corresponding to the target timbre in the first audio signal is mixed with the first time-domain signals corresponding to other timbres in the first audio signal through weighting to obtain the mixed first target audio signal. This effectively balances the timbre characteristics of the timbres, ensuring that the enhancement effect of the target timbre is natural and layered, thereby improving the overall performance and listening experience of the audio.
[0015] In one feasible implementation, a first time-domain signal and a third time-domain signal are mixed based on a first weight and a second weight to obtain a first target audio signal. This includes: mixing the first time-domain signal and the third time-domain signal based on the first weight and the second weight to obtain a mixed fourth time-domain signal; and adjusting the fourth time-domain signal based on a gain suppression strategy to obtain the first target audio signal. Using this method, the signals are first mixed using weights to obtain a mixed fourth time-domain signal, and then processed based on a gain suppression strategy to obtain the first target audio signal. This eliminates distortion caused by excessive gain exceeding a preset decibel, avoids clipping or sound quality degradation in the audio signal, ensures a natural and clear enhancement effect on the audio signal corresponding to the target timbre, and improves the overall audio quality and listening experience.
[0016] In one feasible implementation, before determining at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and the first spectral features corresponding to at least one timbre template, the method further includes: acquiring sample timbres corresponding to at least one sample audio; the sample timbres include user-preset personalized timbres and non-personalized timbres; and constructing a preset timbre database based on at least one sample timbre. By acquiring sample timbres corresponding to at least one sample audio and constructing a preset timbre database based on the sample frequency domain signals corresponding to these sample timbres, the introduction of personalized sample timbres allows for customization according to user preferences, enhancing the personalization of the audio experience. Simultaneously, including non-personalized sample timbres expands the diversity of the timbre database, thereby improving the efficiency of target timbre recognition.
[0017] In one feasible implementation, a preset timbre database is constructed based on at least one sample timbre, including: acquiring a first sample spectral feature and a second sample spectral feature of the sample frequency domain signal corresponding to each sample timbre; wherein, the first sample spectral feature is used to characterize the timbre feature corresponding to the sample frequency domain signal, and the second sample spectral feature is used to characterize the timbre intensity change of the sample frequency domain signal in different time ranges; the first sample spectral feature and the second sample spectral feature belong to the spectral features of the same timbre; each first sample spectral feature is separated to obtain at least one first spectral feature; the first spectral feature is used to characterize the spectral feature corresponding to each single timbre extracted from the first sample feature; based on each second sample spectral feature, sample sound effect parameters are obtained; the sample sound effect parameters are the parameters corresponding to the timbre in the sample audio; and each first spectral feature and sample sound effect parameter are stored in the preset timbre database. Using the above method, by acquiring the first sample spectral feature and the second sample spectral feature of each sample frequency domain signal in advance, and setting the sample sound effect parameters based on the second spectral feature, a timbre database is constructed. Thus, by pre-storing the timbre features (first sample spectral features) and the corresponding sample sound effect parameters, the target timbre can be quickly identified and enhanced, improving processing efficiency. Meanwhile, the rich content of the timbre database provides a more accurate basis for subsequent audio enhancement, ensuring that the final audio effect meets expectations and can be personalized according to different timbre characteristics, enhancing the user's listening experience.
[0018] In one feasible implementation, obtaining the first and second sample spectral features of the sample frequency domain signal corresponding to each sample timbre includes: when the timbre in the sample audio is a single timbre, obtaining the starting point of each sample frequency domain signal; determining the time interval of each sample frequency domain signal based on the starting point; segmenting each sample frequency domain signal based on each time interval and a preset frame shift time; obtaining the segmented sample frequency domain signal; and decomposing each segmented sample frequency domain signal to obtain the first and second sample spectral features of each segmented sample frequency domain signal. Using the above method, in scenarios where the sample audio is a single timbre, segmenting each sample frequency domain signal through time intervals and frame shift times better captures the detailed changes in the audio signal. Furthermore, the first and second sample spectral features obtained by decomposing the segmented sample frequency domain signals can effectively characterize the timbre features. This contributes to the accuracy of subsequent audio processing and timbre enhancement, improving the overall audio quality and user experience.
[0019] In one feasible implementation, obtaining the first and second sample spectral features of each sample frequency domain signal includes: when the sample audio has multiple timbres, obtaining the starting point of each sample frequency domain signal; determining the time interval of each sample frequency domain signal based on the starting point; segmenting each sample frequency domain signal based on each time interval and a preset frame shift time; obtaining the segmented sample frequency domain signals; and clustering each segmented sample frequency domain signal to obtain the first and second sample spectral features of each clustered sample frequency domain signal. Using this method, in scenarios where the sample audio has multiple timbres, segmenting each sample frequency domain signal through time intervals and frame shift times better captures the detailed changes in the audio signal. Furthermore, clustering the segmented sample frequency domain signals yields the first and second sample spectral features, which effectively characterize timbre features. This improves the accuracy of subsequent audio processing and timbre enhancement, enhancing the overall audio quality and user experience.
[0020] In one feasible implementation, each segmented sample frequency domain signal is clustered to obtain a first sample spectral feature and a second sample spectral feature for each clustered sample frequency domain signal. This includes: using each segmented sample frequency domain signal as input to a first model, and using the first model to obtain the first sample spectral feature and the second sample spectral feature for each clustered sample frequency domain signal; wherein the first model is a neural network model or a machine learning model. Using the above model, the first sample spectral feature and the second sample spectral feature of the sample frequency domain signal can be automatically extracted. This improves the accuracy of audio processing and provides a foundation for subsequent timbre enhancement and audio optimization.
[0021] In one feasible implementation, obtaining the sample timbre corresponding to at least one sample audio audio includes: in response to a user's selection operation on a first selection control of an audio selection interface of a first application, displaying a timbre extraction interface, the timbre extraction interface including a feature selection area; the feature selection area being used to characterize an audio segment in the original sample audio selected by the user; in response to a user's input operation on the feature selection area, displaying a timbre storage interface, the timbre storage interface including at least one second selection control, the second selection control being used to characterize at least one timbre in the sample audio; and in response to a user's selection operation on at least one second selection control, obtaining the sample timbre corresponding to at least one sample audio audio. Using the above method, audio features can be flexibly selected and extracted through user interaction, thus providing stronger personalized control, allowing users to accurately select audio segments and timbres according to their needs, thereby improving the accuracy and flexibility of audio processing.
[0022] In one feasible implementation, the audio effect parameters include at least one of the following: the gain of the frequency domain signal corresponding to the target timbre, the frequency of the frequency domain signal corresponding to the target timbre, and the smoothing time of the gain change of the frequency domain signal corresponding to the target timbre. These audio effect parameters provide a comprehensive audio adjustment strategy, making the audio signal corresponding to the target timbre more accurate and natural during enhancement, while also taking into account user preferences and different listening needs, ultimately improving the overall audio performance and listening experience.
[0023] In one feasible implementation, the method further includes: if the second spectral feature fails to match the first spectral feature corresponding to at least one timbre template, obtaining a fifth time-domain signal corresponding to the first frequency domain signal; and mixing the fifth time-domain signal and the first time-domain signal to obtain a second target audio signal. By employing this method, when each second spectral feature fails to match the timbre template, obtaining the fifth time-domain signal corresponding to the first frequency domain signal and mixing it with the first time-domain signal can effectively address the situation of timbre matching failure. This ensures the balance and continuity of the audio signal, avoids audio quality degradation due to matching failure, and thus guarantees the overall performance of the audio.
[0024] To achieve the above objectives, in a second aspect, this application provides an audio enhancement device that has the function of implementing the electronic device behavior in the audio enhancement method of the first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0025] Thirdly, this application provides an electronic device, including: a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the audio enhancement method provided in the first aspect above.
[0026] Fourthly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the audio enhancement method provided in the first aspect above.
[0027] Fifthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the audio enhancement method as described in the first aspect above.
[0028] It is understood that the beneficial effects that the technical solutions provided in the second to fifth aspects described above can be achieved by referring to the beneficial effects of the first aspect and any feasible implementation thereof, which will not be repeated here. Attached Figure Description
[0029] Figure 1 This is a comparative diagram of audio enhancement before and after, provided in an embodiment of this application;
[0030] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0031] Figure 3 This is a schematic diagram of the layered architecture of the software system of the electronic device provided in the embodiments of this application;
[0032] Figure 4 This is a first flowchart illustrating an audio enhancement method provided in an embodiment of this application;
[0033] Figure 5 This is the first flowchart of an audio enhancement method provided in an embodiment of this application;
[0034] Figure 6 This is a schematic diagram of a frequency domain conversion provided in an embodiment of this application;
[0035] Figure 7 This is a schematic diagram of a process for determining a target timbre provided in an embodiment of this application;
[0036] Figure 8 This is a schematic diagram of a decomposed spectral feature provided in an embodiment of this application;
[0037] Figure 9 This is a schematic diagram illustrating an embodiment of obtaining spectral features provided in this application;
[0038] Figure 10 This is a schematic diagram of audio signal mixing provided in an embodiment of this application;
[0039] Figure 11 This is a second flowchart illustrating an audio enhancement method provided in an embodiment of this application;
[0040] Figure 12 This is a flowchart illustrating a method for constructing a timbre database according to an embodiment of this application;
[0041] Figure 13 This is a schematic diagram of the first interface for obtaining personalized sample timbre provided in an embodiment of this application;
[0042] Figure 14 This is a schematic diagram of a second interface for obtaining personalized sample timbre provided in an embodiment of this application;
[0043] Figure 15 This is a schematic diagram of the first process for obtaining spectral features provided in this embodiment;
[0044] Figure 16 This is a schematic diagram of a starting point detection provided in this embodiment;
[0045] Figure 17 This is a schematic diagram of the second process for obtaining spectral features provided in this embodiment;
[0046] Figure 18 This is a schematic diagram illustrating the separation of spectral features provided in an embodiment of this application;
[0047] Figure 19 This is a schematic diagram of the structure of an audio enhancement device provided in an embodiment of this application;
[0048] Figure 20 This is a schematic diagram of another audio enhancement device provided in an embodiment of this application;
[0049] Figure 21 This is a schematic diagram of the chip system provided in the embodiments of this application. Detailed Implementation
[0050] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.
[0051] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0052] Furthermore, in this application, directional terms such as "upper," "lower," "inner," and "outer" are defined relative to the indicated placement of the components in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the placement of the components in the accompanying drawings.
[0053] To facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art, the technical terms involved in the embodiments of this application will be explained below.
[0054] Timbre refers to the unique quality or characteristic of a sound that allows us to distinguish sounds from different sources, even if they have the same pitch and volume. For example, the different timbres of a piano and a guitar.
[0055] Audio signals are the electronic representation of sound, referring to sound signals that can be perceived by the human ear. As a physical signal, sound signals originate from the propagation of sound waves in the air. These sound waves contain various vibrational components from low to high frequencies (i.e., bass, mid-range, and treble). When these sound waves are received by devices such as microphones, they are converted into audio signals. Low-frequency (bass) and high-frequency (treble) components occupy an important position in audio signals, determining the sound's depth, clarity, and spatiality.
[0056] Time-domain signals refer to the representation of audio signals in the time domain. For audio signals, the time domain is the dimension that describes the changes of audio signals over time, while time-domain signals describe the waveform of audio signals through changes on the time axis, showing how the intensity (amplitude) of audio signals changes over time.
[0057] Frequency domain signal refers to the representation of an audio signal in the frequency domain. For audio signals, the frequency domain is a holistic dimension describing the various frequency components of the audio signal, and the frequency domain signal is a representation obtained after frequency analysis of the audio signal. It shows the intensity distribution of the signal at different frequencies, reflecting the energy distribution and importance of each frequency component in the audio signal.
[0058] Spectral features refer to the features extracted by analyzing the components of an audio signal in the frequency domain (i.e., the frequency distribution of the signal). They reveal the intensity, distribution, and other important information of different frequency components in the audio signal.
[0059] Spectral leakage refers to the phenomenon where frequency components "leak" into adjacent frequency ranges due to the truncation of an audio signal into finite-length segments during spectral analysis using Discrete Fourier Transform or Fast Fourier Transform. In other words, when analyzing a non-periodic signal and truncating it into finite-length segments, repeatedly piecing the signal back together can lead to discontinuities at the signal edges. This artificially introduced discontinuity causes the frequency components of the original signal to spread out in the frequency domain, rather than concentrating at their intended frequency points, thus creating spectral leakage.
[0060] The Fast Fourier Transform (FFT) is an efficient algorithm used to transform audio signals from the time domain (or spatial domain) to the frequency domain. It is essentially a fast implementation of the Discrete Fourier Transform (DFT), significantly reducing the number of computations required to calculate the DFT.
[0061] Non-negative matrix factorization (NMF) is a mathematical method that decomposes a non-negative matrix into two non-negative matrices. For example, a non-negative matrix V is decomposed into the product of two non-negative matrices G and B. In the audio domain, NMF is used to decompose the overall spectral characteristics of an audio signal into a set of basis vectors and weights. For example, the basis matrix identifies the spectral characteristics of each timbre, and the activation matrix represents the temporal weights of the basis matrices.
[0062] The embodiments of this application will now be described with reference to the accompanying drawings.
[0063] Currently, many portable electronic products, due to size, hardware limitations, and space constraints, often struggle to provide high-quality sound and fail to fully meet users' audio experience needs. To address this issue, audio signal enhancement is often necessary.
[0064] For example, the enhancement effect of the audio signal can be shown by the waveform of the audio file. The waveform can help us clearly see the changes in different frequency bands of the signal.
[0065] Figure 1 This is a schematic diagram showing a comparison of audio enhancement before and after, provided in an embodiment of this application.
[0066] For example, such as Figure 1 As shown in (a) above, the amplitude distribution of the audio signal waveform is relatively uniform without any processing, the waveform edges are unclear, and there is little detail. Figure 1 As shown in (b), after audio enhancement processing, the waveform presents a highly structured cluster of vertical white lines with sharp waveform edges. High-frequency details are clearly expressed through dense white stripes, making the details clearer.
[0067] In some embodiments, audio is often compressed using compressor technology. Specifically, the original audio is enhanced to obtain an enhanced audio signal, and this enhanced audio signal is then subjected to multi-band compression to obtain a compressed enhanced audio signal. However, in other embodiments, the audio signal is first processed by a bass enhancement circuit to generate an enhanced bass signal. Then, a sensing circuit generates statistics based on the processed signal, and a bass parameter controller uses these statistics to generate appropriate parameters to optimize the bass enhancement. Finally, the enhanced signal is delivered to the speaker to produce audio output, demonstrating the bass enhancement effect.
[0068] Both of the above embodiments belong to indiscriminate enhancement, meaning that all parts of the audio signal are uniformly enhanced, including those parts that do not need enhancement. Furthermore, this holistic audio enhancement is usually accompanied by the application of overall gain, and the increase in gain often requires smoothing. This can suppress transient signal changes in the audio, thereby weakening the audio's impact. For example, the impact of drumbeats, the detail of instrumental performance, or the subtle fluctuations in vocals. These transient signals are often key parts of audio expressiveness and emotional transmission; once smoothed, the audio may lose its original dynamic range and sense of layering, thus affecting the overall listening experience.
[0069] In conclusion, this enhancement method may result in poor sound quality and negatively impact the user's listening experience.
[0070] To address the aforementioned issues, this application provides an audio enhancement method that can precisely adjust the sound quality of a specific timbre in the original audio, avoiding any impact on other timbres. This makes the audio enhancement process more personalized and customized, ensuring that the optimized audio signal better meets expectations and significantly improving the overall audio quality.
[0071] The image display method provided in this embodiment can be applied to electronic devices. In some embodiments, the electronic device may be a mobile phone, tablet computer, handheld computer, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, etc. This application embodiment does not impose any special limitations on the specific type of electronic device.
[0072] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0073] like Figure 2As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a Universal Serial Bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antennas 01 and 02, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, buttons 190, a motor 191, a camera 192, a display screen 193, and a Subscriber Identification Module (SIM) card interface 194, etc. The sensor module 180 may include a touch sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a geomagnetic sensor 180D, etc.
[0074] Processor 110 may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0075] The controller can generate operation control signals based on the instruction opcode and timing signals, thereby controlling the process of acquiring and executing instructions.
[0076] A GPU is a hardware device used to process graphics and image data. Its core function is graphics rendering, that is, converting graphics data into pixels on the screen.
[0077] DSPs are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency point, a digital signal processor is used to perform Fourier transforms on the frequency point energy, etc.
[0078] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0079] In some embodiments, the processor 110 may include one or more interfaces. These interfaces may include an Inter-Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, or a Universal Asynchronous Receiver / Transmitter (UART) interface. The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0080] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0081] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable audio playback through Bluetooth headphones.
[0082] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0083] In some embodiments, processor 110 may include one or more interfaces.
[0084] The external memory interface 120 can be used to connect to an external non-volatile memory, thereby expanding the phone's storage capacity. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be saved in the external non-volatile memory.
[0085] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0086] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0087] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 193, camera 192, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0088] The wireless communication function of electronic device 100 can be implemented through antenna 01, antenna 02, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0089] Antennas 01 and 02 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 01 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0090] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 01, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 01. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0091] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The wireless communication module 160 receives electromagnetic waves via antenna 02, modulates and filters the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, modulate and amplify them, and then convert them into electromagnetic waves for radiation via antenna 02.
[0092] In some embodiments, antenna 01 of electronic device 100 is coupled to mobile communication module 110, and antenna 02 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. Wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. GNSS can include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0093] Electronic device 100 implements display functions through a GPU, a display screen 193, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 193 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0094] The display screen 193 is used to display images, videos, etc. In this embodiment, the display screen 193 can be used to display an interface, such as an audio selection interface. The electronic device 100 can implement the shooting function through an ISP, camera 192, video codec, GPU, display screen 193, and application processor.
[0095] Camera 192 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 192, where N is a positive integer greater than 1.
[0096] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170 and application processor.
[0097] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0098] Touch sensor 180A, also known as a "touch device," can be disposed on display screen 193. The touch sensor 180A and display screen 193 together form a touchscreen, also known as a "touchscreen." Touch sensor 180A is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 193. In other embodiments, touch sensor 180A may also be disposed on the surface of electronic device 100, in a different location than display screen 193.
[0099] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100.
[0100] The 180C barometric pressure sensor is used to measure barometric pressure.
[0101] The geomagnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the geomagnetic sensor 180D to detect the opening and closing of the flip cover.
[0102] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0103] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 193. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0104] The SIM card interface 194 is used to connect the SIM card.
[0105] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the mobile phone. In other embodiments of this application, the mobile phone may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0106] Of course, it is understandable that the above... Figure 2 The illustration shown is merely an example of an electronic device in the form of a mobile phone. If the electronic device is a tablet, handheld computer, PC, PDA, wearable device (such as a smartwatch, smart bracelet), or other device form factor, the structure of the electronic device may include more advanced features. Figure 2 The fewer structures shown can also include more than Figure 2 The structures shown are not limited here.
[0107] It is understandable that, generally speaking, the implementation of electronic device functions requires not only hardware support but also software cooperation. The software system of electronic devices can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture... Taking the system as an example, the software structure of the electronic device is illustrated.
[0108] Figure 3 This is a schematic diagram of the layered architecture of the software system of the electronic device provided in this embodiment of the application.
[0109] In some examples, refer to Figure 3As shown in this embodiment, the software of the electronic device is divided into five layers, from top to bottom: the application layer, the framework layer (or application framework layer), the system library and Android runtime, the HAL layer (hardware abstraction layer), and the driver layer (or kernel layer). The system library and Android runtime can also be referred to as the native framework layer or the native layer.
[0110] The application layer can include a series of applications. For example... Figure 3 As shown, the application layer can include applications (APPs) such as camera, gallery, calendar, map, WLAN, music, SMS, call, video, and audio extraction.
[0111] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer.
[0112] The application framework layer includes some predefined functions or services. For example, the application framework layer may include an activity manager, window manager, content provider, view system, phone manager, resource manager, notification manager, camera service, audio service, etc., and this application embodiment does not impose any limitations on this.
[0113] The system library can include multiple functional modules, such as the surface manager and media libraries. The surface manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The media libraries support playback and recording of various common audio and video formats, as well as still image files. The media libraries support multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0114] The Hardware Abstraction Layer (HAL) is the interface layer between the operating system kernel and the hardware circuitry, designed to abstract the hardware. It hides the platform-specific hardware interface details, providing the operating system with a virtual hardware platform that is hardware-independent and portable across multiple platforms. The HAL provides a standard interface that exposes device hardware functionality to the higher-level Java API framework (i.e., the framework layer). The HAL contains multiple library modules, each implementing an interface for a specific type of hardware component. Examples include: memory allocation HAL, Bluetooth module, camera HAL (also known as camera HAL or camera hardware abstraction module), audio processing module, etc. The camera HAL includes a thumbnail-based image processing module and an image processing module, as well as a sensor HAL (or isensor service).
[0115] The driver layer is the layer between hardware and software. The driver layer includes at least display drivers, camera drivers, audio drivers, sensor drivers, battery drivers, etc., but this application does not limit this.
[0116] The hardware layer includes displays, audio digital processors, speakers, etc.
[0117] The audio enhancement method provided in the embodiments of this application will be further described below with reference to the accompanying drawings.
[0118] Figure 4 This is a first flowchart illustrating an audio enhancement method provided in an embodiment of this application.
[0119] Figure 5 This is the first flowchart of an audio enhancement method provided in the embodiments of this application.
[0120] like Figure 4 and Figure 5 As shown, the audio enhancement method may include the following steps:
[0121] Step S1: Obtain the first frequency domain signal and the first time domain signal corresponding to the original audio.
[0122] In some embodiments, step S1 includes steps S10-S14.
[0123] Step S10: Obtain the first original time-domain signal corresponding to the original audio.
[0124] For example, the first raw time-domain signal corresponding to 0.1 seconds of raw audio is obtained.
[0125] Step S11: Frame the first original time domain signal to obtain the second original time domain signal after framing.
[0126] In audio signal processing, framing refers to dividing a continuous audio signal into a series of shorter segments to facilitate the capture of signal changes and details.
[0127] In one implementation, the first original time-domain signal is divided into frames based on the length of each frame.
[0128] Continuing with the example above, if each frame is 2ms long, the first original time-domain signal of 0.1 seconds is divided into frames to obtain a second original time-domain signal of 50 frames.
[0129] It should be noted that the length of each frame can be adjusted to obtain a second original time-domain signal after framing; no specific limitations are made here.
[0130] Optionally, to ensure signal continuity and reduce information loss, the distance or overlap between adjacent frames, known as "frame shift," can be set based on the determined frame length. For example, if a 50% overlap rate is used, the actual movement distance of each frame will be half the frame length. Continuing with the example above, if the overlap rate is set to 50%, the first original time-domain signal of 0.1 seconds is divided into 100 frames to obtain the second original time-domain signal, where each frame is 1ms long.
[0131] Step S12: Based on the spectrum leakage suppression strategy, the second original time-domain signal is suppressed to obtain the first time-domain signal.
[0132] In one implementation, the second original time-domain signal is windowed to obtain the first time-domain signal. Windowing is a commonly used strategy for suppressing spectral leakage. The windowing function reduces the discontinuities in the second original time-domain signal.
[0133] Continuing with the example above, the first time-domain signal is obtained by windowing each frame of the second original time-domain signal using the "Hanning window" function.
[0134] In another implementation, the second original time-domain signal is processed by increasing the frame length to obtain the first time-domain signal. After framing, the second original time-domain signal may exhibit discontinuities between frames, especially at frame boundaries. Increasing the frame length reduces the number of frames, thereby reducing the frequency of discontinuities.
[0135] Continuing with the example above, the second original time-domain signal is processed by increasing the frame length, for example, increasing the length of each frame from the original 2 milliseconds to 4 milliseconds, thereby obtaining the first time-domain signal.
[0136] It should be noted that spectrum leakage suppression strategies can also include adjusting frame shift and overlap rate or selecting an appropriate sampling rate, etc., which are not specifically limited here.
[0137] Step S13: Perform frequency domain conversion on the first time domain signal to obtain the first frequency domain signal.
[0138] In one implementation, the first time-domain signal is transformed into the frequency domain using a Fast Fourier Transform (FFT) to obtain the frequency domain representation of the audio signal, i.e., the first frequency domain signal. In other words, by converting the original audio's first time-domain signal into a first frequency domain signal using FFT, the original spectral characteristics corresponding to the first frequency domain signal can be obtained. Optionally, the original spectral characteristics are typically represented in the form of a non-negative matrix.
[0139] For example, the spectral characteristics can be represented by formula (1):
[0140]
[0141] Where K represents the spectral resolution, N represents the number of signal frames, and V K (t N ) represents the amplitude value of the Nth frame signal at the Kth frequency point.
[0142] Figure 6 This is a schematic diagram of a frequency domain conversion provided in an embodiment of this application.
[0143] like Figure 6 As shown, Figure 6 Figure (a) shows the audio signal in the time domain, i.e., the first time-domain signal. This figure depicts the amplitude of the audio signal as it changes over time, reflecting the variation of the spectral amplitude of the audio signal at different time points. The first time-domain signal is then transformed into the frequency domain using a Fast Fourier Transform (FFT). After FFT processing, the following figure is obtained: Figure 6 The first frequency domain signal is shown in (b) above. The first frequency domain signal demonstrates the characteristics of the audio signal in the frequency domain, revealing the variation of the spectral amplitude of the audio signal at different frequencies.
[0144] Step S2: Based on the matching result of the first frequency domain signal and the first spectral feature corresponding to at least one timbre template, determine at least one target timbre contained in the original audio, wherein at least one timbre template contains at least one target timbre.
[0145] Figure 7 This is a schematic diagram of a process for determining a target timbre provided in an embodiment of this application.
[0146] In some embodiments, combined with Figure 7 As shown, step S2 includes steps S21-S22.
[0147] Step S21: Obtain the second spectral feature of the first frequency domain signal; the second spectral feature is used to characterize the timbre feature corresponding to the first frequency domain signal.
[0148] It should be noted that the timbre characteristics reflected in the second spectral feature are a set of spectral features that exist independently for each timbre, that is, they are not spectral features that are mixed together.
[0149] In one example, combining Figure 8 The second spectral characteristics of the first frequency domain signal are explained.
[0150] Figure 8 This is a schematic diagram of a decomposed spectral feature provided in an embodiment of this application.
[0151] like Figure 8 As shown, at least one second spectral feature of the first frequency domain signal is obtained through nonnegative matrix factorization (NMF).
[0152] Frequency domain signals typically correspond to a spectral feature. In audio signal processing, the spectral feature is the feature extracted after the time-domain signal is converted to the frequency domain. The spectral feature can also be called a spectrogram. In this embodiment, the spectral feature is used as an example for explanation. The spectral feature usually exists in the form of a non-negative matrix. Decomposing the spectral feature using non-negative matrix factorization (NMF) typically yields a basis matrix and an activation matrix. The basis matrix represents the spectral feature of each timbre, and the activation matrix represents the weight of the basis matrix over time, i.e., the change in timbre intensity over different time ranges.
[0153] For example, the basis matrix and activation matrix of the spectral features can be determined by formula (2):
[0154] V = B·G (Formula 2)
[0155] Where V is the overall spectral feature, B is the basis matrix, and G is the activation matrix.
[0156] In this embodiment, the first original spectral feature V corresponding to the first frequency domain signal is decomposed by nonnegative matrix factorization (NMF) to obtain the second spectral feature W (basis matrix). Continuing with... Figure 8 As shown, the original audio contains three timbres: "guitar sound," "piano sound," and "child's voice." The first original spectral feature V of the first frequency domain signal is decomposed to obtain the second spectral feature W (basis matrix). The second spectral feature W includes three columns, each corresponding to the spectral feature of a timbre. For example, the spectral feature P1 corresponds to "piano sound," the spectral feature P2 corresponds to "child's voice," and the spectral feature P3 corresponds to "guitar sound."
[0157] Step S22: If the second spectral feature successfully matches the first spectral feature corresponding to at least one timbre template, obtain at least one target timbre contained in the original audio.
[0158] The first spectral feature is usually extracted in advance from known timbre samples, representing a relatively standard or typical timbre characteristic. The second spectral feature is the spectral feature corresponding to the original audio acquired in real time, reflecting the instantaneous state or changes of sound in the actual environment.
[0159] It should be noted that a match is considered successful if any timbre in the second spectral feature matches a corresponding first spectral feature in multiple timbre templates. In other words, even if the second spectral feature contains multiple timbres, a match is considered successful as long as any one of these timbres can find a corresponding timbre template, regardless of whether the other timbres are in multiple timbre templates. For example, if the second spectral feature includes "guitar sound," "piano sound," and "child's voice," a match is confirmed as successful if one or more of these three timbres can find a corresponding match in multiple timbre templates.
[0160] Optionally, the first spectral feature corresponding to the timbre template can be stored in a pre-set timbre database. For example, suppose the timbre database pre-stores the first spectral features corresponding to 5 timbre templates, namely the first spectral feature Q1 corresponding to the first timbre "violin sound", the first spectral feature Q2 corresponding to the second timbre "flute sound", the first spectral feature Q3 corresponding to the third timbre "piano sound", the first spectral feature Q4 corresponding to the fourth timbre "child's voice", and the first spectral feature Q5 corresponding to the fifth timbre "saxophone sound".
[0161] In one implementation, step S22 includes steps S221-S222.
[0162] Step S221: Determine the matching confidence based on the first spectral feature and the second spectral feature; the matching confidence is used to characterize the similarity between the spectral features of the target timbre and the spectral features of the timbre template.
[0163] In one example, a first divergence is determined based on a first spectral feature and a second spectral feature; the first divergence is used to characterize the degree of deviation between the spectral features of the target timbre and the spectral features of the timbre template; the first divergence is normalized to obtain the matching confidence.
[0164] For example, the first divergence can be determined with reference to the following formula (3):
[0165]
[0166] Among them, D IS B is the first divergence. P B is the second spectral characteristic of the target timbre. Q The first spectral feature corresponding to the timbre template. This represents the summation of all terms from i=1 to i=T, where i refers to the current time point, and B... P (i) The spectral characteristics of the target timbre at the i-th time point, B Q (i) Spectral characteristics of the timbre template at time point i.
[0167] Following the example above, the matching confidence score can be determined using the following formula (4):
[0168]
[0169] For example, the matching confidence of the first and second spectral features is determined according to formulas (3) and (4). Continuing to combine... Figure 8 As shown, the original audio includes three timbres: the second spectral feature P1 corresponding to "piano sound," the second spectral feature P2 corresponding to "child's voice," and the second spectral feature P3 corresponding to "guitar sound." Specifically, if the matching confidence levels of the second spectral feature P1 corresponding to "piano sound" and the five timbres are respectively... in, It is the matching confidence of the second spectral feature P1 corresponding to the "piano sound" and the first spectral feature Q1 corresponding to the first timbre "violin sound". It is the matching confidence of the second spectral feature P1 corresponding to the "piano sound" and the first spectral feature Q2 corresponding to the second timbre "flute sound". It is the matching confidence of the second spectral feature P1 corresponding to the "piano sound" and the first spectral feature Q3 corresponding to the third timbre "piano sound". It is the matching confidence of the second spectral feature P1 corresponding to the "piano sound" and the first spectral feature Q4 corresponding to the fourth timbre "child's voice". It is the matching confidence of the second spectral feature P1 corresponding to the "piano sound" and the first spectral feature Q5 corresponding to the fifth timbre, the "saxophone sound". The other two timbres are similar.
[0170] Step S222: If the matching confidence is greater than a preset threshold, determine at least one target timbre contained in the original audio.
[0171] Continuing with the example above, if the confidence level of the match between the second spectral feature P1 corresponding to the "piano sound" in the original audio and the first spectral feature Q3 corresponding to the third timbre "piano sound" is... The matching confidence of the second spectral feature P2 corresponding to the "child's voice" and the first spectral feature Q4 corresponding to the fourth timbre "child's voice". If both exceed the preset threshold, determine the two target timbres contained in the original audio, namely, "piano sound" and "child's voice".
[0172] Step S3: Extract the second frequency domain signal corresponding to the target timbre from the first frequency domain signal.
[0173] Continuing with the example above, there are three corresponding timbre features in the first frequency domain signal: "guitar sound", "piano sound" and "child's voice". The target timbres are "piano sound" and "child's voice". Therefore, the frequency domain signals corresponding to these two target timbres will be filtered and extracted.
[0174] Step S4: Based on the target spectral features corresponding to the target timbre, adjust the second frequency domain signal to obtain the third frequency domain signal; the target spectral features are at least one of the first spectral features that correspond to the target timbre.
[0175] In some embodiments, step S4 includes steps S41-S43.
[0176] Step S41: Obtain the third spectral feature of the second frequency domain signal and the target spectral feature corresponding to the target timbre; the third spectral feature is used to characterize the timbre feature corresponding to the second frequency domain signal.
[0177] It should be noted that the timbre characteristics reflected by the third spectral feature are a set of spectral features that exist independently for each timbre, that is, they are not spectral features that are mixed together.
[0178] In one example, combined Figure 9 To explain, Figure 9 This is a schematic diagram of obtaining spectral features according to an embodiment of this application. Step S41 includes steps S411 and S412.
[0179] Step S411: Obtain the third spectral characteristics of the second frequency domain signal.
[0180] The specific details of step S411 can be found in step S21 above, and will not be repeated here.
[0181] Continuing with the example above, the second frequency domain signal is a frequency domain signal that includes both "piano sound" and "child's voice". Therefore, the third frequency domain feature (basis matrix) can be obtained by decomposing the second original spectral feature V' corresponding to the second frequency domain signal. For example, the spectral features include P1' for "piano sound" and P2' for "child's voice".
[0182] Step S412: Obtain the target spectral feature corresponding to the target timbre from at least one first spectral feature included in the timbre database.
[0183] Following the example of step S22 above, the timbre database pre-stores the first spectral features corresponding to 5 timbre templates and the two target timbres, "piano sound" and "child's voice", included in the original audio. Further, the first spectral feature Q3 corresponding to "piano sound" is taken as the target spectral feature Q3' and the first spectral feature Q4 corresponding to "child's voice" is taken as the target spectral feature Q4'.
[0184] Step S42: Replace the third spectral feature in the second frequency domain signal with the target spectral feature to obtain the third frequency domain signal, wherein the timbre corresponding to the third spectral feature is the same as the timbre corresponding to the target spectral feature.
[0185] Continuing with the example above, based on formula (2) and Figure 9 As shown, the second original spectral feature V' corresponding to the second frequency domain signal is decomposed into a third spectral feature (basis matrix) B' and a fourth spectral feature (activation matrix) G' corresponding to the third spectral feature (basis matrix). Further, the third spectral feature is replaced with the target spectral feature to obtain the third original spectral feature V”. That is, V’ = B' × G' is replaced with V” = B” × G', thus obtaining the third frequency domain signal corresponding to the third original spectral feature V”.
[0186] Step S5: Obtain the second time domain signal corresponding to the third frequency domain signal.
[0187] In one implementation, step S5 includes step S51.
[0188] Step S51: Perform time-domain conversion on the third frequency domain signal to obtain the second time domain signal.
[0189] Optionally, the second time-domain signal can be obtained by performing an inverse fast Fourier transform (iFFT) on the third frequency-domain signal.
[0190] Step S6: Based on the sound effect parameters corresponding to the target timbre, enhance the second time domain signal to obtain the third time domain signal.
[0191] Optionally, the sound effect parameters are pre-configured. The sound effect parameters may include at least one of the following: the gain of the frequency domain signal corresponding to the target timbre, the frequency of the frequency domain signal corresponding to the target timbre, and the smoothing time of the gain change of the frequency domain signal corresponding to the target timbre.
[0192] In one implementation, step S6 includes steps S61-S62.
[0193] Step S61: Obtain the sound effect parameters corresponding to each target spectral feature from the preset timbre database.
[0194] Continuing with the example above, the timbre database can pre-store multiple sample sound effect parameters corresponding to the first spectral feature: sample sound effect parameter X1 corresponding to the first spectral feature Q1, sample sound effect parameter X2 corresponding to the first spectral feature Q2, sample sound effect parameter X3 corresponding to the first spectral feature Q3, sample sound effect parameter X4 corresponding to the first spectral feature Q4, and sample sound effect parameter X5 corresponding to the first spectral feature Q5.
[0195] In this embodiment of the application, the target spectral features include the following: target spectral feature Q3' and target spectral feature Q4'. Further, the sound effect parameters corresponding to the target spectral features are obtained from multiple sample sound effect parameters in a preset timbre database, namely sample sound effect parameter X3 and sample sound effect parameter X4, and the sample sound effect parameter X3 and sample sound effect parameter X4 are used as the sound effect parameters X3' and X4' corresponding to the target spectral features.
[0196] Step S62: Based on each sound effect parameter, enhance the second time domain signal to obtain the third time domain signal.
[0197] Continuing with the example above, the "piano sound" is enhanced based on the sound effect parameter X3', and the "child's voice" is enhanced based on the sound effect parameter X4', thus obtaining the enhanced third time-domain signal.
[0198] Step S7: Mix the third time domain signal with the first time domain signal to obtain the first target audio signal.
[0199] The third time-domain signal is the time-domain signal in the original audio that requires timbre enhancement. The first time-domain signal is the time-domain signal in the original audio without any enhancement processing.
[0200] In one embodiment, step S7 includes steps S71-S72.
[0201] Step S71: Based on a preset strategy, determine the first weight corresponding to the first time domain signal and the second weight corresponding to the third time domain signal; wherein, the preset strategy includes the number of timbres in the original audio.
[0202] The preset strategy can be set according to the number of timbres in the original audio. For example, when the original audio consists of one or a few timbres, a higher first weight can be assigned to the first time-domain signal to preserve its natural and harmonious characteristics, while an appropriate second weight is assigned to the third time-domain signal for subtle optimization and enhancement. For instance, the first weight could be set to 0.7 and the second weight to 0.3. When the original audio consists of multiple timbres, a higher second weight might be assigned to the third time-domain signal to better highlight the characteristics and clarity of each timbre. Correspondingly, the first weight could be reduced to avoid interference that might affect the overall quality of the original audio; for example, the first weight could be set to 0.4 and the second weight to 0.6.
[0203] It should be noted that the dynamic settings of the first and second weights mentioned above can take into account not only the number of timbres in the original audio, but also other influencing factors, such as specific scenario requirements, which are not specifically limited here.
[0204] Continuing with the example above, based on the preset strategy, the first weight corresponding to the first time domain signal is determined to be 0.4, and the second weight corresponding to the third time domain signal is determined to be 0.6.
[0205] Step S72: Based on the first weight and the second weight, mix the first time domain signal and the third time domain signal to obtain the first target audio signal.
[0206] In one implementation, step S72 includes step S720.
[0207] Step S720: Based on the first weight and the second weight, the first time domain signal and the third time domain signal are mixed to obtain the first target audio signal.
[0208] In the embodiments of this application, mixing mainly refers to a linear combination in the gain (i.e., signal strength) dimension, but it is not just a simple gain adjustment; it involves processing and optimization in multiple dimensions. Optionally, it may also involve noise suppression, that is, for background noise or other unwanted sound components present in the audio signal, noise reduction and other factors can be applied during the mixing process.
[0209] Continuing with the example above, with the first weight being 0.4 and the second weight being 0.6, the first time-domain signal and the third time-domain signal are mixed to obtain the first target audio signal.
[0210] Optionally, to ensure a good effect for the first target audio signal, appropriate adjustments can be made during the mixing process.
[0211] In another implementation, Figure 10 This is a schematic diagram of audio signal mixing provided in an embodiment of this application. Figure 10As shown, step S72 includes steps S721-S722.
[0212] Step S721: Based on the first weight and the second weight, the first time domain signal and the third time domain signal are mixed to obtain the mixed fourth time domain signal.
[0213] The specific details of step S721 can be found in step S720 above, and will not be repeated here.
[0214] Step S722: Based on the gain suppression strategy, adjust the fourth time domain signal to obtain the first target audio signal.
[0215] Gain suppression strategies are designed to ensure that the amplitude of the audio signal remains within a preset range, meaning it does not exceed the maximum threshold that electronic devices can process, thus preventing distortion.
[0216] In one example, a device that controls the signal amplitude, such as a compressor, can be used to suppress the gain of the fourth time-domain signal. For example, the fourth time-domain signal is input into a compressor to obtain an output signal, which is then used as the first target audio signal.
[0217] In summary, by matching the first spectral feature corresponding to the preset timbre template, the timbre that needs to be enhanced (target timbre) in the original audio can be obtained. The timbre that needs to be enhanced is enhanced in a targeted manner, avoiding the impact on other timbres. This makes the audio enhancement process more personalized and customized. At the same time, since only a portion of the timbre in the original audio is enhanced, there will be no overall gain smoothing phenomenon, and instantaneous volume changes in the audio will not be suppressed, thus ensuring the dynamics of the audio and significantly improving the overall audio quality.
[0218] Corresponding to the embodiments of the aforementioned audio enhancement methods, this application also provides another audio enhancement method. The following describes... Figure 5 and Figure 11 An example is provided.
[0219] Figure 11 This is a second flowchart illustrating an audio enhancement method provided in an embodiment of this application.
[0220] like Figure 5 and Figure 11 As shown, the audio enhancement method may include the following steps:
[0221] Step S100: Obtain the first frequency domain signal and the first time domain signal corresponding to the original audio.
[0222] Step S200: If the matching result between the first frequency domain signal and the first spectral feature corresponding to at least one timbre template fails, obtain the fifth time domain signal.
[0223] In one implementation, step S200 includes steps S200a-S200b.
[0224] Step S200a: Obtain the second spectral characteristics of the first frequency domain signal.
[0225] The above step S200a can refer to the specific content of step S21, and is not specifically limited here.
[0226] Step S200b: If the second spectral feature fails to match the first spectral feature corresponding to at least one timbre template, obtain the fifth time domain signal corresponding to the first frequency domain signal.
[0227] For example, if the second spectral feature fails to match the first spectral feature corresponding to at least one timbre template, the first frequency domain signal is subjected to inverse fast Fourier transform (iFFT) to obtain the fifth time domain signal.
[0228] The fifth time-domain signal can be an audio signal that has undergone simple noise reduction processing.
[0229] Step S300: Mix the fifth time domain signal and the first time domain signal to obtain the second target audio signal.
[0230] The above step S300 can refer to the specific content of step S7, and is not specifically limited here.
[0231] If the matching result between the first frequency domain signal and the first spectral feature corresponding to at least one timbre template fails, it means that the original audio cannot be enhanced directly using preset sound effect parameters. The fifth time domain signal, which has undergone simple noise reduction but has not been enhanced with specific timbre, is mixed with the first time domain signal to obtain the second target audio signal.
[0232] In summary, if the first frequency domain signal fails to match the first spectral characteristics corresponding to at least one timbre template, a second target audio signal can be obtained by mixing the fifth time domain signal and the first time domain signal. This process not only provides flexibility but also ensures high-quality audio output even in the absence of precise timbre enhancement.
[0233] In the embodiments of the two audio enhancement methods described above, the first spectral features and sound effect parameters corresponding to at least one timbre template are pre-set and are generally stored in a database for easy integration, such as the timbre database in the embodiments of this application. Combined with Figure 12 As shown in the embodiments of this application, a method for constructing a timbre database is also provided.
[0234] Figure 12This is a flowchart of a method for constructing a timbre database provided in an embodiment of this application.
[0235] like Figure 12 As shown, the method for constructing a timbre database may include the following steps:
[0236] Step S01: Obtain the sample timbre corresponding to at least one sample audio; the sample timbre includes user-preset personalized timbre and non-personalized timbre.
[0237] Personalized timbres can be user-defined timbre templates for a specific sound segment, such as a user using their own recorded voice as a timbre template. Non-personalized sample timbres are general, standard timbre templates, such as instrument sounds or wind sounds.
[0238] The following explains how to obtain sample timbres. Specific examples of obtaining non-personalized sample timbres will not be provided; instead, we will use the method of obtaining personalized sample timbres as an example.
[0239] In one implementation, the electronic device 100, in response to a user's selection operation on a first selection control of the audio selection interface of a first application, displays a timbre extraction interface, the timbre extraction interface including a feature selection area; the feature selection area is used to characterize an audio segment in the original sample audio selected by the user; in response to a user's input operation on the feature selection area, displays a timbre storage interface, the timbre storage interface including at least one second selection control, the second selection control being used to characterize at least one timbre in the sample audio; in response to a user's selection operation on at least one second selection control, obtains a personalized sample timbre corresponding to at least one sample audio.
[0240] Alternatively, there are various ways to obtain personalized sample timbres, which will be illustrated below with two specific examples.
[0241] Figure 13 This is a schematic diagram of the first interface for obtaining personalized sample timbre provided in an embodiment of this application.
[0242] like Figure 13 As shown, in one example, the electronic device 100 can display an audio extraction application icon 10 on the main screen interface 1. In response to a user's click operation 001 on the audio extraction application 10 on the main screen interface 1, the electronic device 100 launches the audio extraction application, and the interface displayed by the electronic device 100 changes from... Figure 13 The main screen interface 1 shown in (a) switches to the one provided by [the original text]. Figure 13The audio selection interface 2 is shown in (b) above. The audio selection interface 2 includes at least one first selection control for selecting different original audio files. Further, in response to a user's click operation 002 on the first selection control corresponding to the selected original audio file, the interface displayed by the electronic device 100 is... Figure 13 The audio selection interface 2 shown in (b) switches to the one provided by [the original text]. Figure 13 The timbre extraction interface 3 is shown in (c) above. The timbre extraction interface 3 includes a feature selection area 31, a sliding control 32, and an extraction control 33. The feature selection area 31 represents the user's input to select a specific original audio segment from the original audio. The sliding control 32 represents the user's selection of a progress bar from the original audio, thereby obtaining a specific original audio segment. Furthermore, in response to the user's operation on the feature selection area 31 or the sliding control 32, such as a sliding operation 003 on the sliding control 32 and a clicking operation 004 on the extraction control 33, the interface displayed by the electronic device 100 changes. Figure 13 The timbre extraction interface shown in (c) 3 switches to the one provided by Figure 13 The timbre storage interface 4 is shown in (d) above. The timbre storage interface 4 includes at least one second selection control, which is used to characterize at least one timbre in the sample audio. The electronic device 100, in response to a user's click operation 005 on the second selection control, acquires a personalized sample timbre.
[0243] Figure 14 This is a schematic diagram of a second interface for obtaining personalized sample timbre provided in an embodiment of this application.
[0244] Continue as Figure 13 and Figure 14 As shown, in another example, the main screen interface 1 of the electronic device 100 also displays a music player application icon 11. In response to a user's click operation 006 on the music player application icon 11, the electronic device 100 launches the music player application, and the interface displayed by the electronic device 100 is... Figure 14 The main screen interface 1 shown in (a) switches to the one provided by [the original text]. Figure 14 The music playback interface 5 is shown in (b) above. The music playback interface 5 includes a tone selection control 51. In response to a user's click operation 007 on the tone selection control 51 in the music playback interface 5, the interface displayed by the electronic device 100 changes from... Figure 14 The music playback interface shown in (a) 5 switches to the one provided by Figure 13The timbre extraction interface 3 is shown in (c). In response to user operations on the feature selection area 31 or the sliding control 32 on the timbre extraction interface 3, such as a sliding operation 003 on the sliding control 32 and a clicking operation 004 on the extraction control 33, the interface displayed by the electronic device 100 is as follows: Figure 13 The timbre extraction interface shown in (c) 3 switches to the one provided by Figure 13 The tone storage interface 4 is shown in (d) in the diagram. Further, the electronic device 100, in response to a user's click operation 005 on the second selection control in the tone storage interface 4, acquires a personalized sample tone.
[0245] It should be noted that when the electronic device 100 responds to the user's click operation 007 on the timbre selection control 51 in the music playback interface 5, the electronic device 100 can directly obtain the specific original audio segment in the original audio in the music playback interface 5. That is to say, there is no need to generate a new timbre extraction interface 3. The timbre can be obtained in the existing music playback interface 5 (not shown here), and then the personalized sample frequency domain signal can be obtained.
[0246] Step S02: Construct a preset timbre database based on at least one sample timbre.
[0247] In one implementation, step S02 includes steps S021-S024.
[0248] Step S021: Obtain the first sample spectral feature and the second sample spectral feature of the sample frequency domain signal corresponding to each sample timbre; wherein, the first sample spectral feature is used to characterize the timbre feature corresponding to the sample frequency domain signal, and the second sample spectral feature is used to characterize the timbre intensity change of the sample frequency domain signal in different time ranges; the first sample spectral feature and the second sample spectral feature belong to the spectral features of the same timbre.
[0249] In other words, the first sample spectral feature is the basis matrix of the sample frequency domain signal, and the second sample spectral feature is the activation matrix of the sample frequency domain signal.
[0250] The sample timbre may be a single timbre or multiple timbres. For different timbre scenarios, the methods for obtaining the first sample spectral features and the second sample spectral features of the sample frequency domain signal are different. The following will explain this with specific examples.
[0251] In one example, in a monophonic scenario, combined with Figure 15 and Figure 16 As shown, Figure 15 This is a schematic diagram of the first process for obtaining spectral features according to an embodiment of this application. Figure 16 This is a schematic diagram of a starting point detection provided in an embodiment of this application.
[0252] like Figure 15 and Figure 16 As shown in (a), step S021 includes steps S021a-S021e.
[0253] Step S021a: When the sample audio is monophonic, obtain the starting point of the frequency domain signal of each sample.
[0254] For example, electronic device 100 can obtain the starting point of each sample frequency domain signal by onset detection.
[0255] It should be noted that the sample frequency domain signal refers to the frequency domain transformed signal corresponding to the original audio segment selected by the user. Specifically, after the user completes the selection of the original audio segment in the timbre extraction interface 3, the electronic device 100 will perform the following processing flow: first, perform time-frequency transformation processing based on the selected audio segment to generate the corresponding sample frequency domain signal; furthermore, obtain the starting point of the sample frequency domain signal.
[0256] Step S021b: Based on the starting point, determine the time interval for each sample frequency domain signal.
[0257] Step S021c: Based on each time interval and the preset frame shift time, segment (slice) each sample frequency domain signal; obtain the segmented sample frequency domain signal.
[0258] Step S021d: Decompose each segmented sample frequency domain signal to obtain the first sample spectral features and the second sample spectral features of each segmented sample frequency domain signal.
[0259] It should be noted that each sample frequency domain signal after segmentation includes the first sample spectral features and the second sample spectral features, which correspond to the front-end and back-end interaction logic of the timbre storage interface 4.
[0260] In another example, in a multi-timbre scenario, combining Figure 16 and Figure 17 As shown, Figure 17 This is a second flowchart illustrating an embodiment of the present application for obtaining spectral features.
[0261] like Figure 16 (b) and Figure 17 As shown, the multiple timbres include timbre 1, timbre 2, and timbre 3. Step S021 includes steps S021e-S021h.
[0262] Step S021e: When the sample audio has multiple timbres, obtain the starting point of the frequency domain signal of each sample.
[0263] Step S021f: Based on the starting point, determine the time interval for each sample frequency domain signal.
[0264] Step S021g: Based on each time interval and a preset frame shift time, segment each sample frequency domain signal; obtain the segmented sample frequency domain signal.
[0265] The specific content of steps S021e-S021g can be referred to steps S021a-S021c above, and is not specifically limited here.
[0266] Step S021h: Cluster each segmented sample frequency domain signal to obtain the first sample spectral features and the second sample spectral features of each clustered sample frequency domain signal.
[0267] For example, step S021h includes step S021h1.
[0268] Step S021h1: Use each segmented sample frequency domain signal as input to the first model, and use the first model to obtain the first sample spectral features and the second sample spectral features of each clustered sample frequency domain signal; wherein, the first model is a neural network model or a machine learning model.
[0269] Step S022: Separate each first sample spectral feature to obtain at least one first spectral feature; the first spectral feature is used to characterize the spectral feature corresponding to each single timbre extracted from the first sample features.
[0270] Figure 18 This is a schematic diagram of the separation of spectral features provided in an embodiment of this application.
[0271] For example, such as Figure 18 As shown, if the sample audio corresponding to the first sample spectral feature includes three timbres, such as "guitar sound", "piano sound", and "child's voice", the first sample spectral feature is separated to obtain the first spectral feature corresponding to each single timbre, such as the first spectral feature P1 corresponding to "piano sound", the first spectral feature P2 corresponding to "child's voice", and the first spectral feature P3 corresponding to "guitar sound".
[0272] Step S023: Based on the spectral features of each second sample, obtain the sample sound effect parameters; the sample sound effect parameters are the parameters corresponding to the timbre in the sample audio.
[0273] The sample sound effect parameters are user-defined settings.
[0274] Step S024: Store each first spectral feature and sample sound effect parameter into a preset timbre database.
[0275] In summary, by predefining the first spectral features and sample audio parameters, a second spectral signal that matches the first frequency domain signal can be directly obtained in real-time audio enhancement scenarios, which helps to accurately locate and extract the target timbre that needs to be enhanced.
[0276] It should be noted that the embodiments of this application are applicable to all frequency components of audio, including low frequency, mid frequency, and high frequency. In particular, this embodiment provides a stronger enhancement effect for the low frequency component.
[0277] Figure 19 This is a schematic diagram of the structure of an audio enhancement device provided in an embodiment of this application.
[0278] like Figure 19 As shown, the audio enhancement device 200 provided in this application embodiment can be applied to the electronic device in the above embodiment. The audio enhancement device 200 may include: a display screen 201, a memory 202, a processor 203, and a communication module 204. These devices can be connected via one or more communication buses 205. The display screen 201 may include a display panel 2011 and a touch sensor 2012. The display panel 2011 is used to display images, and the touch sensor 2012 can transmit detected touch operations to the application processor to determine the touch event type, providing visual output related to the touch operation through the display panel 2011. The processor 203 may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. Different processing units may be independent devices or integrated into one or more processors. The memory 202 is coupled to the processor 203 and is used to store various software programs and / or computer instructions. The memory 202 may include volatile memory and / or non-volatile memory. When the processor executes computer instructions, the electronic device can perform the various functions or steps performed in the above method embodiments.
[0279] Figure 20 This is a schematic diagram of another audio enhancement device provided in the embodiments of this application.
[0280] like Figure 20As shown, in some other embodiments, this application also provides an audio enhancement device 300, which can be applied to the electronic device in the above embodiments. The audio enhancement device includes an acquisition module 301 configured to: acquire a first frequency domain signal and a first time domain signal corresponding to the original audio; a determination module 302 configured to: determine at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and a first spectral feature corresponding to at least one timbre template, wherein the at least one timbre template contains at least one target timbre; and an extraction module 303 configured to: extract the first spectral feature corresponding to the target timbre from the first frequency domain signal. The second frequency domain signal; adjustment module 304 is configured to: adjust the second frequency domain signal based on the target spectral features corresponding to the target timbre to obtain a third frequency domain signal; the target spectral features are at least one of the first spectral features corresponding to the target timbre; acquisition module 301 is further configured to: acquire the second time domain signal corresponding to the third frequency domain signal; enhancement module 305 is configured to: enhance the second time domain signal based on the sound effect parameters corresponding to the target timbre to obtain a third time domain signal; mixing module 306 is configured to: mix the third time domain signal with the first time domain signal to obtain a first target audio signal. This application also provides an electronic device, which includes a processor and a memory coupled together. The memory stores program instructions, which, when executed by the processor, cause the processor to perform the various functions or steps described in the above method embodiments.
[0281] Figure 21 This is a schematic diagram of the chip system provided in the embodiments of this application.
[0282] like Figure 21 As shown, this application embodiment also provides a chip system 400, such as a SoC, which includes at least one processor 401 and at least one interface circuit 402. The processor 401 and the interface circuit 402 can be interconnected via lines. For example, the interface circuit 402 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 402 can be used to send signals to other devices (e.g., the processor 401 or the touchscreen of an electronic device). Exemplarily, the interface circuit 402 can read instructions stored in the memory and send the instructions to the processor 401. When the instructions are executed by the processor 401, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this application embodiment does not specifically limit this.
[0283] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0284] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0285] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0286] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0287] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0288] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0289] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0290] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An audio enhancement method, characterized in that, include: Obtain the first frequency domain signal and the first time domain signal corresponding to the original audio; Based on the matching result of the first frequency domain signal and the first spectral feature corresponding to at least one timbre template, at least one target timbre contained in the original audio is determined, and the at least one timbre template contains the at least one target timbre; Extract the second frequency domain signal corresponding to the target timbre from the first frequency domain signal; Based on the target spectral features corresponding to the target timbre, the second frequency domain signal is adjusted to obtain a third frequency domain signal; the target spectral features are at least one of the first spectral features that corresponds to the target timbre. Obtain the second time domain signal corresponding to the third frequency domain signal; Based on the sound effect parameters corresponding to the target timbre, the second time-domain signal is enhanced to obtain the third time-domain signal; The third time-domain signal is mixed with the first time-domain signal to obtain the first target audio signal.
2. The audio enhancement method according to claim 1, characterized in that, The step of obtaining the first frequency domain signal and the first time domain signal corresponding to the original audio includes: Obtain the first original time-domain signal corresponding to the original audio; The first original time-domain signal is divided into frames to obtain the second original time-domain signal after framing. Based on the spectrum leakage suppression strategy, the second original time-domain signal is suppressed to obtain the first time-domain signal; The first time-domain signal is converted to the frequency domain to obtain the first frequency-domain signal.
3. The audio enhancement method according to claim 1 or 2, characterized in that, The step of determining at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and the first spectral features corresponding to at least one timbre template includes: The second spectral feature of the first frequency domain signal is obtained; the second spectral feature is used to characterize the timbre feature corresponding to the first frequency domain signal. If the second spectral feature successfully matches the first spectral feature corresponding to the at least one timbre template, at least one of the target timbres contained in the original audio is obtained.
4. The audio enhancement method according to claim 3, characterized in that, The step of determining at least one target timbre contained in the original audio when the second spectral feature successfully matches the first spectral feature corresponding to the at least one timbre template includes: Based on the first spectral feature and the second spectral feature, a matching confidence level is determined; the matching confidence level is used to characterize the similarity between the spectral features of the target timbre and the spectral features of the timbre template. If the matching confidence level is greater than a preset threshold, at least one target timbre contained in the original audio is determined.
5. The audio enhancement method according to claim 4, characterized in that, The step of determining the matching confidence based on the first spectral feature and the second spectral feature includes: Based on the first spectral feature and the second spectral feature, a first divergence is determined; the first divergence is used to characterize the degree of deviation between the spectral features of the target timbre and the spectral features of the timbre template. The first divergence is normalized to obtain the matching confidence.
6. The audio enhancement method according to claim 3, characterized in that, The step of adjusting the second frequency domain signal based on the target spectral features corresponding to the target timbre to obtain the third frequency domain signal includes: Obtain the third spectral feature of the second frequency domain signal and the target spectral feature corresponding to the target timbre; the third spectral feature is used to characterize the timbre feature corresponding to the second frequency domain signal; The third spectral feature in the second frequency domain signal is replaced with the target spectral feature to obtain the third frequency domain signal, wherein the timbre corresponding to the third spectral feature is the same as the timbre corresponding to the target spectral feature.
7. The audio enhancement method according to claim 1, characterized in that, The step of obtaining the second time-domain signal corresponding to the third frequency-domain signal includes: The third frequency domain signal is converted into the second time domain signal by time domain transformation.
8. The audio enhancement method according to claim 1, characterized in that, The step of enhancing the second time-domain signal based on the sound effect parameters corresponding to the target timbre to obtain the third time-domain signal includes: Obtain the sound effect parameters corresponding to each target spectral feature from a preset timbre database; Based on each of the aforementioned sound effect parameters, the second time-domain signal is enhanced to obtain the third time-domain signal.
9. The audio enhancement method according to claim 1, characterized in that, The step of mixing the third time-domain signal with the first time-domain signal to obtain the first target audio signal includes: Based on a preset strategy, a first weight corresponding to the first time-domain signal and a second weight corresponding to the third time-domain signal are determined; wherein, the preset strategy includes the number of timbres in the original audio. Based on the first weight and the second weight, the first time-domain signal and the third time-domain signal are mixed to obtain the first target audio signal.
10. The audio enhancement method according to claim 9, characterized in that, The step of mixing the first time-domain signal and the third time-domain signal based on the first weight and the second weight to obtain the first target audio signal includes: Based on the first weight and the second weight, the first time-domain signal and the third time-domain signal are mixed to obtain a mixed fourth time-domain signal; Based on the gain suppression strategy, the fourth time-domain signal is adjusted to obtain the first target audio signal.
11. The audio enhancement method according to claim 1, characterized in that, Before determining at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and the first spectral features corresponding to at least one timbre template, the method further includes: Obtain at least one sample timbre corresponding to a sample audio; the sample timbre includes user-preset personalized timbre and non-personalized timbre; A preset timbre database is constructed based on at least one of the sample timbres.
12. The audio enhancement method according to claim 11, characterized in that, The step of constructing a preset timbre database based on at least one of the sample timbres includes: Obtain the first sample spectral feature and the second sample spectral feature of the sample frequency domain signal corresponding to each sample timbre; wherein, the first sample spectral feature is used to characterize the timbre feature corresponding to the sample frequency domain signal, and the second sample spectral feature is used to characterize the timbre intensity change of the sample frequency domain signal in different time ranges; the first sample spectral feature and the second sample spectral feature belong to the spectral features of the same timbre; Each of the first sample spectral features is separated to obtain at least one first spectral feature; the first spectral feature is used to characterize the spectral feature corresponding to each single timbre extracted based on the first sample features; Based on the spectral features of each second sample, sample sound effect parameters are obtained; the sample sound effect parameters are the parameters corresponding to the sample timbre. Each of the first spectral features and the sample sound effect parameters is stored in the preset timbre database.
13. The audio enhancement method according to claim 12, characterized in that, The process of obtaining the first sample spectral features and the second sample spectral features of the sample frequency domain signal corresponding to each sample timbre includes: If the timbre in the sample audio is a single timbre, obtain the starting point of the frequency domain signal of each sample. Based on the starting point, the time interval for each of the sample frequency domain signals is determined; Based on each time interval and a preset frame shift time, each sample frequency domain signal is segmented to obtain the segmented sample frequency domain signal; Each of the segmented sample frequency domain signals is decomposed to obtain the first sample spectral feature and the second sample spectral feature of each segmented sample frequency domain signal.
14. The audio enhancement method according to claim 12, characterized in that, The process of obtaining the first sample spectral features and the second sample spectral features of each sample frequency domain signal includes: In the case where the timbre in the sample audio is multi-timbre, the starting point of the frequency domain signal of each sample is obtained; Based on the starting point, the time interval for each of the sample frequency domain signals is determined; Based on each time interval and a preset frame shift time, each sample frequency domain signal is segmented to obtain the segmented sample frequency domain signal; Each segmented sample frequency domain signal is clustered to obtain the first sample spectral feature and the second sample spectral feature of each clustered sample frequency domain signal.
15. The audio enhancement method according to claim 14, characterized in that, The step of clustering each segmented sample frequency domain signal to obtain the first sample spectral feature and the second sample spectral feature of each clustered sample frequency domain signal includes: Each segmented sample frequency domain signal is used as input to a first model. Using the first model, the first sample spectral features and the second sample spectral features of each clustered sample frequency domain signal are obtained; wherein, the first model is a neural network model or a machine learning model.
16. The audio enhancement method according to claim 11, characterized in that, The step of obtaining the sample timbre corresponding to at least one sample audio includes: In response to a user's selection operation on the first selection control of the audio selection interface of the first application, a timbre extraction interface is displayed, the timbre extraction interface including a feature selection area; the feature selection area is used to characterize the audio segment in the original audio sample selected by the user. In response to the user's input operation on the feature selection area, a timbre storage interface is displayed, the timbre storage interface including at least one second selection control, the second selection control being used to characterize at least one timbre in the sample audio; In response to the user's selection operation on the at least one second selection control, the sample timbre corresponding to the at least one sample audio is obtained.
17. The audio enhancement method according to any one of claims 1-14, characterized in that, The sound effect parameters include at least one of the following: the gain of the frequency domain signal corresponding to the target timbre, the frequency of the frequency domain signal corresponding to the target timbre, and the smoothing time of the gain change of the frequency domain signal corresponding to the target timbre.
18. The audio enhancement method according to claim 3, characterized in that, Also includes: If the second spectral feature fails to match the first spectral feature corresponding to the at least one timbre template, the fifth time domain signal corresponding to the first frequency domain signal is obtained. The second target audio signal is obtained by mixing the fifth time-domain signal and the first time-domain signal.
19. An audio enhancement device, characterized in that, include: The acquisition module is configured to acquire the first frequency domain signal and the first time domain signal corresponding to the original audio. The determining module is configured to: determine at least one target timbre contained in the original audio based on the matching result of the first frequency domain signal and the first spectral feature corresponding to at least one timbre template, wherein the at least one timbre template contains the at least one target timbre; The extraction module is configured to: extract the second frequency domain signal corresponding to the target timbre from the first frequency domain signal; The adjustment module is configured to: adjust the second frequency domain signal based on the target spectral features corresponding to the target timbre to obtain a third frequency domain signal; the target spectral features are the spectral features corresponding to the target timbre among the at least one of the first spectral features. The acquisition module is further configured to: acquire the second time domain signal corresponding to the third frequency domain signal; The enhancement module is configured to enhance the second time-domain signal based on the sound effect parameters corresponding to the target timbre to obtain a third time-domain signal; The mixing module is configured to mix the third time-domain signal with the first time-domain signal to obtain a first target audio signal.
20. An electronic device, characterized in that, include: A memory and one or more processors; the memory is coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the audio enhancement method as described in any one of claims 1-18.
21. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed electronically, cause the electronic device to perform the audio enhancement method as described in any one of claims 1-18.
22. A computer program product, characterized in that, When the computer program product is run on a computer, the computer performs the audio enhancement method as described in any one of claims 1-18.