Audio quality improvement method and device, electronic equipment and storage medium

By constructing a non-negative matrix to separate audio signals, identify the environment type, and adjust the audio parameters, the problem of interference in complex environments is solved, and the user's auditory experience is improved.

CN120279936APending Publication Date: 2025-07-08BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510568628.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional audio processing methods fail to effectively deal with interference in complex environments, affecting users' auditory experience.

Method used

By constructing a non-negative matrix to separate the target audio signals, identify the environment type in combination with frequency and time characteristics, and adjust the audio parameters to suit different environments.

Benefits of technology

Improve the adaptability of audio signals in different environments, ensuring that users get the best auditory experience in various occasions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279936A_ABST
    Figure CN120279936A_ABST
Patent Text Reader

Abstract

The invention provides an audio quality improvement method and device, electronic equipment and a storage medium, and belongs to the technical field of audio processing, and the method comprises the steps: constructing a non-negative matrix based on the spectrum characteristics of target mixed audio signals, separating the target mixed audio signals to obtain a plurality of target audio signals, the target mixed audio signal is a mixed audio signal in a target environment; determining the environment type of the target environment based on the frequency characteristics and the time characteristics of the plurality of target audio signals; and adjusting parameters of an output audio signal based on the environment type of the target environment, wherein the output audio signal is an audio signal output by the target equipment. According to the audio quality improvement method and device, the electronic equipment and the storage medium provided by the invention, the auditory experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of audio processing, and more specifically, relates to a method and device for improving audio quality, an electronic device, and a storage medium. Background Art

[0002] With the continuous development of audio technology and the increasing demand for audio quality by people, how to ensure the clarity and audibility of audio in various complex environments has become an urgent problem to be solved. Traditional audio processing methods often ignore the impact of the environment on audio quality, resulting in the audio signal being easily interfered in noisy or special environments, thus affecting the user's auditory experience. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a method and device for improving audio quality, an electronic device, and a storage medium to improve the user's auditory experience.

[0004] In the first aspect of the embodiments of the present disclosure, a method for improving audio quality is provided, including: Constructing a non - negative matrix based on the spectral characteristics of the target mixed audio signal, and separating the target mixed audio signal to obtain multiple target audio signals, where the target mixed audio signal is a mixed audio signal in a target environment; Determining the environmental type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals; Adjusting the parameters of the output audio signal based on the environmental type of the target environment, where the output audio signal is the audio signal output by the target device.

[0005] In the second aspect of the embodiments of the present disclosure, a device for improving audio quality is provided, including: An audio separation module, configured to construct a non - negative matrix based on the spectral characteristics of the target mixed audio signal, and separate the target mixed audio signal to obtain multiple target audio signals, where the target mixed audio signal is a mixed audio signal in a target environment; An environment recognition module, configured to determine the environmental type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals; An audio adjustment module, configured to adjust the parameters of the output audio signal based on the environmental type of the target environment, where the output audio signal is the audio signal output by the target device.

[0006] In the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above - mentioned method for improving audio quality are implemented.

[0007] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps of the above-described method for improving audio quality.

[0008] The beneficial effects of the method, apparatus, electronic device, and storage medium for improving audio quality provided by the embodiments of the present disclosure are as follows: By constructing a non-negative matrix based on the spectral characteristics of the target mixed audio signal and separating multiple target audio signals, the embodiments of the present disclosure can clearly disassemble different components in complex mixed audio. Secondly, determining the target environment type based on the frequency and time characteristics of multiple target audio signals can make the environment judgment more in line with the actual audio situation, enhancing the accuracy of the judgment. Furthermore, more targeted countermeasures can be taken. Finally, adjusting the parameters of the output audio signal according to the characteristics of the target environment improves the adaptability of the audio in different environments, ensuring that users can obtain the best auditory experience in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0010] Figure 1 is a schematic flowchart of a method for improving audio quality provided by an embodiment of the present disclosure; Figure 2 is a structural block diagram of an apparatus for improving audio quality provided by an embodiment of the present disclosure; Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0012] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0013] Please refer to Figure 1 , Figure 1A flowchart of a method for improving audio quality provided by an embodiment of the present disclosure is provided, and the method includes: S101: constructing a non-negative matrix based on the frequency spectrum characteristics of a target mixed audio signal, and separating the target mixed audio signal to obtain a plurality of target audio signals, where the target mixed audio signal is a mixed audio signal in a target environment.

[0014] In this embodiment, Fourier transform is performed on the target mixed audio signal to obtain the frequency spectrum characteristics of the target mixed audio signal.

[0015] The spectrum characteristics can reflect key information such as the energy distribution of the audio signal at different frequencies. By performing Fourier transform on the target mixed audio signal, the target mixed audio signal can be converted from the time domain to the frequency domain to obtain the spectrum characteristics.

[0016] The amplitude and other information of the spectral features are non-negative values. A non-negative matrix is ​​constructed based on the spectral features of the target mixed audio signal. The constructed non-negative matrix is ​​processed using a non-negative matrix decomposition algorithm to decompose the target mixed audio signal into multiple target audio signals. Multiple target audio signals are individual components separated from the original mixed audio. For example, in a mixed audio containing speech and background music, separate speech signals and background music signals can be separated.

[0017] S102: Determine an environment type of a target environment based on frequency characteristics and time characteristics of a plurality of target audio signals.

[0018] In this embodiment, the frequency characteristics of each target audio signal include the distribution and intensity of different frequency components, for example, a speech signal has a specific energy distribution pattern in a certain frequency range. The time characteristics involve the temporal variation of the audio signal, such as the duration, periodicity, and intermittency of the sound.

[0019] The frequency characteristics and time characteristics of multiple target audio signals may be matched with the preset environment audio to obtain the type of the target environment.

[0020] For example, if the separated audio signal contains more low-frequency machine sounds with a certain periodicity, and also contains some audio components of human voice communication, it can be judged that the target environment is a factory workshop; if it is mainly natural sounds such as birdsong and wind, and the sound is relatively soothing, it can be judged as a natural environment.

[0021] S103: adjusting parameters of the output audio signal based on the environment type of the target environment, the output audio signal being an audio signal output by the target device.

[0022] In this embodiment, the target device may be a wearable device, such as smart glasses, smart headphones, etc.

[0023] After determining the environmental type of the target environment, the parameters of the audio signal output by the target device can be adjusted according to the characteristics of this environmental type. The parameters of the output audio signal can include volume, equalizer settings, audio encoding format, channel mode, etc.

[0024] Exemplarily, it is assumed that in a bar, it is necessary to increase the volume of the output audio to ensure that the user can clearly hear the audio signal output by the target device; it is assumed that in a library, the volume can be reduced and the equalizer can be adjusted to make the audio sound softer without affecting the surrounding environment.

[0025] It can be concluded from the above that in this embodiment, by constructing a non - negative matrix based on the spectral characteristics of the target mixed audio signal and separating multiple target audio signals, different components in the complex mixed audio can be clearly disassembled. Secondly, by determining the target environmental type according to the frequency and time characteristics of multiple target audio signals, the environmental judgment can be more in line with the actual audio situation, enhancing the accuracy of the judgment. Furthermore, more targeted countermeasures can be taken. Finally, by adjusting the parameters of the output audio signal according to the characteristics of the target environment, the adaptability of the audio in different environments is improved, ensuring that the user can obtain the best auditory experience in different occasions.

[0026] In an embodiment of the present disclosure, constructing a non - negative matrix based on the spectral characteristics of the target mixed audio signal and separating the target mixed audio signal to obtain multiple target audio signals includes: Constructing a non - negative matrix based on the spectral characteristics of the target mixed audio signal to obtain the product of the spectral matrix and the activation intensity matrix of the target audio signal; Obtaining multiple target audio signals based on the product of the spectral matrix and the activation intensity matrix of the target audio signal.

[0027] In this embodiment, by performing a Fourier transform on the target mixed audio signal, the spectral characteristics of the target mixed audio signal can be obtained. The spectral characteristics reflect the energy distribution of the audio signal at different frequencies.

[0028] Based on the spectral characteristics, a non - negative matrix is constructed. The non - negative matrix can be regarded as the product of the spectral matrix and the activation intensity matrix of the target audio signal. Among them, the spectral matrix is the "template" of each source audio signal in the frequency domain, reflecting the inherent characteristics of different source audios in terms of frequency components, such as the energy distribution law of the speech signal in a specific frequency band. The activation intensity matrix represents the "activity level" of these source audio signals at different moments in the mixed audio, that is, the contribution of a certain source audio to the mixed audio at a certain instant.

[0029] Restore multiple target audio signals based on the product of the spectral matrix and the activation intensity matrix of the target audio signal. Process this product through the non - negative matrix factorization algorithm to separate the mixed audio signals according to the rules contained in the spectral matrix and the activation intensity matrix. Such as speech, background music, environmental noise, etc., can be clearly separated to achieve the purpose of audio separation.

[0030] Exemplarily, assume that the mixed audio signal can be represented by a non - negative matrix V. Each column of the non - negative matrix can represent the spectral characteristics of the mixed audio signal at a specific time point, and the rows can represent different frequency components. The goal of non - negative matrix factorization is to decompose V into two non - negative matrices W and H, that is, V≈WH.

[0031] In the context of audio separation, W can be regarded as the spectral matrix of each source audio signal, and H can be regarded as the activation intensity matrix of these source audio signals at each time point. For example, if there are two source audio signals (such as speech and background music), the number of columns r of W = 2. The two columns of W represent the spectral templates of speech and background music respectively, and the two rows of H represent the activation intensities of speech and background music at each time point.

[0032] First, the mixed audio signal can be frame - processed and converted into a suitable matrix. Each frame of the audio signal is subjected to a fast Fourier transform to obtain its spectral representation. For example, for an audio of length T seconds, frame it with a frame length of t milliseconds. After performing a fast Fourier transform on each frame, arrange the spectral coefficients into the columns of a matrix. To ensure that the elements in the matrix are non - negative, it can be achieved by taking the absolute value of the result of the fast Fourier transform.

[0033] Randomly initialize the spectral matrix W (m•r) and the activation intensity matrix H (r•n), where m is the number of frequency components, n is the number of frames, and r is the pre - estimated number of source audio signals. For example, for a simple case of mixing speech and background music, r = 2. The elements of the spectral matrix W can be randomly initialized in the interval [0, 1], and the elements of the activation intensity matrix H can also be in the interval [0, 1].

[0034] For each source audio signal s i (i = 1, …, r), its spectrum at each time point can be reconstructed according to the spectral template of W and the activation intensity of H. For example, the spectral estimate of the i - th source audio signal at the j - th frame is , where represents the i - th column of W, and then the spectrum is converted back to the time - domain audio signal through the inverse Fourier transform, thus achieving the separation of the audio signal.

[0035] As can be seen from the above, in this embodiment, a non - negative matrix is constructed based on the spectral characteristics of the target mixed audio signal and decomposed into the product of the spectral matrix and the activation intensity matrix of the target audio signal, and then multiple clear and independent target audio signals are separated. This improves the accuracy and efficiency of audio signal separation, enabling the effective extraction of the required audio components even in complex environments. At the same time, this embodiment retains the spectral characteristics of the audio signal, providing a richer and more accurate information basis for subsequent audio processing, thus bringing a better audio processing effect and user experience.

[0036] In an embodiment of the present disclosure, it further includes: Cluster the target mixed audio signal to obtain multiple clusters; Determine the spectral matrix of the target audio signal based on the center of each cluster; Determine the activation intensity matrix of the target audio signal based on the distance between the data points and the center in each cluster.

[0037] In this embodiment, the spectral matrix and the activation intensity matrix are the initialized spectral matrix and the initialized activation intensity matrix.

[0038] Non - negative matrix factorization is an iterative optimization process, and its goal is to find appropriate \(W\) and \(H\) such that \(V\approx WH\). Before starting the iteration, initial values need to be provided for \(W\) and \(H\) so that non - negative matrix factorization can have a starting point for subsequent calculations. If no initialization is performed, the values of \(W\) and \(H\) are uncertain. For example, if the initial values are all zero, when using iterative algorithms such as the multiplicative update rule, \(W\) and \(H\) may always remain zero and no meaningful decomposition result can be obtained.

[0039] In this embodiment, when clustering the target mixed audio signal, the K - Means clustering can be used to divide the target mixed audio signal into multiple clusters according to its similarity in the spectral feature space. The audio data within each cluster has a certain correlation.

[0040] Based on the center of each cluster to determine the spectral matrix of the target audio signal, regarding the typical features in terms of frequency reflected by the cluster center as the initial content of the spectral matrix. The cluster center can represent the commonalities of the audio signals within this cluster in terms of frequency distribution, etc., and can be regarded as a kind of spectral "template".

[0041] Determine the activation intensity matrix of the target audio signal according to the distance between the data points and the center in each cluster. The data points with a short distance indicate that the corresponding audio components are more active in the audio source represented by this cluster, while those with a long distance are relatively inactive. In this way, the activation intensity matrix is constructed to reflect the differences in the activity levels of different audio components.

[0042] Exemplarily, assume there is a target mixed audio signal recorded at an outdoor music festival site, which contains the electric guitar sound, drum sound, lead singer's voice played by a rock band on stage, as well as the cheers, conversations of the audience below the stage, and the wind sound at the scene, etc., with many sounds intertwined together.

[0043] First, use the K-Means clustering algorithm to divide according to the similarity of these audio signals in the spectral feature space. For example, the high-frequency string sound of the electric guitar and the high-frequency overtone part of the lead singer's voice are relatively similar in spectral features and will be clustered into one cluster; the strong low-frequency rhythm sound of the drums is classified into another cluster; the continuous conversation sound of the audience forms a cluster because its frequency is relatively concentrated and stable.

[0044] Next, taking the cluster where the electric guitar and the lead singer's voice are located as an example, the typical frequency features such as the prominent energy in the 500Hz - 2000Hz frequency band presented by the cluster center become the initial content of the corresponding part in the spectral matrix of the target audio signal and serve as the spectral "template" for subsequent decomposition.

[0045] Finally, in this cluster, the audio segments corresponding to the data points close to the center, such as the segments where the lead singer sings passionately, indicate that their audio components are more active and are assigned higher values at the corresponding positions in the activation intensity matrix; while the relatively weak harmony segments that are far from the center and appear occasionally are assigned lower values. The activation intensity matrix constructed in this way clearly reflects the differences in the activity levels of different audio components and lays a foundation for subsequent non-negative matrix factorization iterative optimization and accurate separation of audio signals.

[0046] It can be concluded from the above that clustering the target mixed audio signal in this embodiment can group it according to the similarity of audio features, making the chaotic audio data initially orderly. Determining the spectral matrix based on the cluster center accurately captures the typical frequency features of each cluster of audio and provides a reliable template for subsequent decomposition. And using the distance between the data points in the cluster and the center to determine the activation intensity matrix reflects the differences in the activity levels of audio components.

[0047] In an embodiment of the present disclosure, it further includes: Calculating the partial derivatives corresponding to the spectral matrix and the activation intensity matrix based on the objective function; Updating the spectral matrix and the activation intensity matrix based on the learning efficiency and partial derivatives corresponding to the spectral matrix and the activation intensity matrix to obtain the target spectral matrix and the target activation intensity matrix; Obtaining multiple target audio signals based on the product of the target spectral matrix and the target activation intensity matrix of the target audio signal.

[0048] In practical applications, such as audio separation, V represents the non - negative matrix of the mixed audio signal, which contains the complex information after mixing multiple source audio signals. By continuously updating the spectral matrix W and the activation intensity matrix H, the decomposed matrix can better fit the actual audio signal structure. For example, during the iteration process, the spectral matrix W can be gradually adjusted to a template that can accurately represent the spectral characteristics of each source audio signal, and the activation intensity matrix H can be adjusted to a matrix that correctly reflects the activation situation of each source audio signal at different time points.

[0049] In this embodiment, the objective function is expressed as:

[0050] Where, represents the objective function, represents the non - negative matrix of the mixed audio signal, represents the spectral matrix, represents the activation intensity matrix, represents the Frobenius norm.

[0051] The objective function is used to measure the difference between the current decomposition result and the original mixed audio signal. By minimizing this objective function, the decomposed matrix can be closer to the actual audio signal structure.

[0052] In this embodiment, calculate the partial derivatives of the objective function with respect to the spectral matrix W and the activation intensity matrix H.

[0053] The objective function The partial derivative with respect to the spectral matrix W is:

[0054] Where, represents the partial derivative of the objective function with respect to the spectral matrix W, represents the time length.

[0055] The objective function The partial derivative with respect to the activation intensity matrix H is:

[0056] Where, represents the partial derivative of the objective function with respect to the activation intensity matrix H.

[0057] In traditional non - negative matrix update rules, the update step size for each iteration is relatively fixed. However, during the iteration process, the requirements for the step size are different at different stages. In the initial stage, a larger step size may be needed to quickly approach the optimal solution range. When approaching the optimal solution, a smaller step size is required for fine - tuning to avoid missing the optimal solution.

[0058] In this embodiment, the spectral matrix and the activation intensity matrix are updated based on the learning efficiency and partial derivatives corresponding to them.

[0059] Introduce an adaptive learning rate and , the update rule for the spectral matrix W is:

[0060] where, represents the element value of the \(i\) - th row and \(a\) - th column of the updated spectral matrix W, represents the element value of the \(i\) - th row and \(a\) - th column of the spectral matrix W, represents the learning efficiency corresponding to the spectral matrix W.

[0061] The update rule for the activation intensity matrix H is:

[0062] where, represents the element value of the \(a\) - th row and \(j\) - th column of the updated activation intensity matrix H, represents the element value of the \(a\) - th row and \(j\) - th column of the activation intensity matrix H, represents the learning efficiency corresponding to the activation intensity matrix H.

[0063] In this embodiment, it can be determined according to the current iteration number \(t\) and the error change.

[0064] The learning efficiency corresponding to the spectral matrix W is expressed as:

[0065] The learning efficiency corresponding to the activation intensity matrix H is expressed as:

[0066] where, , , , can all be parameters determined based on experience or experiments, and \(t\) represents the iteration number.

[0067] In this embodiment, compared with the multiplication update rule with a traditional fixed update step size, the adaptive learning rate can converge to a better solution faster. For the complex problem of mixing signal separation, this embodiment can achieve a separation effect equivalent to or even better than that of traditional methods within fewer iterations. Moreover, it can better adapt to different types of mixed signals and data characteristics.

[0068] As can be seen from the above, the partial derivative calculation in this embodiment indicates the direction for matrix update and can ensure approximation towards the optimal solution. The adaptive learning efficiency can dynamically adjust the step size according to the iteration process, with a large step size in the initial stage for rapid approximation and a small step size in the later stage for fine adjustment to avoid missing the optimal solution. Thus, the target matrix can be accurately obtained, multiple target audio signals can be precisely separated, and the quality and efficiency of audio processing can be improved.

[0069] In an embodiment of the present disclosure, determining the environmental type of the target environment based on the frequency characteristics and time characteristics of multiple target audio signals includes: Extracting the frequency characteristics of multiple target audio signals based on a convolutional neural network; Extracting the time characteristics of multiple target audio signals based on a recurrent neural network; Fusing the frequency characteristics and time characteristics to determine the environmental type of the target environment.

[0070] In this embodiment, the energy distribution or spectral shape of different frequency bands of multiple target audio signals can be extracted based on a convolutional neural network. Audio signals generated in different environments often have differences in frequency composition; the characteristic change amplitude or periodicity of multiple target audio signals at different time points can be extracted based on a recurrent neural network. The time characteristics reflect features such as the occurrence order and duration of sound events in the environment.

[0071] Performing weighted summation on the frequency characteristics and time characteristics to obtain the target environmental characteristics, and inputting the target environmental characteristics into the trained decision tree model to obtain the environmental type of the target environment.

[0072] Exemplarily, assume that we have a mixed audio recorded in a city park. After preprocessing, multiple target audio signals such as bird calls, children's laughter, and the sound of the breeze blowing through the leaves are separated.

[0073] Using a convolutional neural network to extract frequency characteristics. For bird calls, the convolutional neural network can capture the unique energy distribution in its high-frequency band, presenting a sharp and regular spectral shape, which is a specific frequency manifestation determined by the structure of the bird's vocal organs; children's laughter has relatively concentrated energy in the mid-low frequency band, and the spectral shape is relatively broad, reflecting the frequency characteristics of human voices.

[0074] Use a recurrent neural network to extract temporal features. For example, for the sound of a gentle breeze, the recurrent neural network can analyze that its feature variation amplitude is extremely small on a long time scale, has a stable periodicity, and almost continues without interruption; while the sound of children playing is intermittent, sometimes strong and sometimes weak, with a large feature variation amplitude at different time points, and shows a short-time concentrated burst, reflecting the dynamic changes in the sound when children are playing.

[0075] Sum the two types of features according to certain weights. For example, a higher weight for the feature of high-frequency bird chirping is beneficial for distinguishing the outdoor natural environment, and the target environmental feature is obtained. Finally, input it into the trained decision tree model. The model judges that this is an urban park environment based on various environmental audio feature patterns learned in the past, and accurately completes the identification of the environmental type.

[0076] It can be concluded from the above that in this embodiment, the convolutional neural network accurately captures the frequency features of the target audio signal, distinguishes the differences in the spectrum of different sound sources, and provides key frequency clues for environmental judgment. The recurrent neural network deeply explores the temporal features of the audio, presenting the variation law of the sound over time. By fusing the frequency features and temporal features, that is, by analyzing the spatio-temporal characteristics of the audio, the accuracy of environmental type determination is improved.

[0077] In an embodiment of the present disclosure, based on the environmental type of the target environment, adjust the parameters of the output audio signal. The output audio signal is the audio signal output by the target device, including: Obtain the acoustic features of the initial output audio signal. The initial output audio signal is the audio signal output before adjustment based on the environmental type of the target environment; Input the acoustic features into the bioacoustic perception model to obtain the adjustment parameters of the initial output audio signal.

[0078] In this embodiment, the acoustic features of the initial output audio signal may include spectral features, such as the energy distribution corresponding to different frequency components, the peak value and bandwidth of the spectrum, etc.; and also include temporal features, such as the duration of the audio, rhythm changes, the start and end times of the sound, etc.

[0079] Construct a bioacoustic perception model based on factors such as the sound perception law of the human auditory system and the sound propagation and influence effects in different environments. The bioacoustic perception model can simulate the subjective feelings and actual auditory effects of humans in a specific environment when hearing audio with different acoustic features. For example, in a noisy factory environment, humans are relatively more sensitive to low-frequency sounds and are more likely to recognize them, while in a quiet library environment, subtle high-frequency sound changes are also easily noticed.

[0080] Input the acoustic features into the bioacoustic perception model to obtain the adjustment parameters of the initial output audio signal. The adjustment parameters may include adjustments to the volume size, gain adjustment of each frequency band of the equalizer, switching of the channel mode, etc.

[0081] As can be seen from the above, in this embodiment, by first obtaining the acoustic features of the audio and then analyzing with the bioacoustic perception model, the adjustment parameters of the audio signal are intelligently determined according to the target environment type, so that the audio output can better meet the requirements of different environments.

[0082] In an embodiment of the present disclosure, it further includes: Determining an adjustment step size corresponding to the adjustment parameter based on the correlation degree between the adjustment parameter and the environment type; Adjusting the adjustment parameter based on the adjustment step size.

[0083] In this embodiment, the correlation coefficient between the adjustment parameter and the environment type is calculated, and the correlation degree between the adjustment parameter and the environment type is determined based on the correlation coefficient.

[0084] For example, in a quiet library environment, the volume adjustment parameter is highly correlated with this environment type because the volume needs to be turned very low to avoid disturbing others, while parameters such as the channel mode are relatively less correlated with this environment type. By calculating the correlation coefficient, the correlation degree between each adjustment parameter and the specific environment type is quantified to obtain a correlation degree value, and the adjustment step size corresponding to the adjustment parameter is determined based on the correlation degree between the adjustment parameter and the environment type.

[0085] A high-correlation adjustment parameter indicates that it plays a key role in adapting to the target environment and improving audio quality. Therefore, a relatively large adjustment step size can be set for it. In this way, when adjusting the audio parameters, a large change can be quickly made in the direction that meets the requirements of this environment, and finally, the adjustment parameter is adjusted according to the determined adjustment step size.

[0086] Exemplarily, assume that in a bustling outdoor market environment, the target device is smart glasses.

[0087] First, the audio information of the surrounding environment is collected through sensors, and after processing, multiple target audio signals are obtained, such as the noisy conversations of the crowd, the shouts of the vendors at the stalls, and the horn sounds of the occasional passing vehicles. Based on the frequency characteristics and time characteristics of these target audio signals, it is determined that the environment is of the outdoor market type.

[0088] Then, start adjusting the parameters of the audio signal output by the smart glasses. It is known that parameters such as the volume adjustment parameter, the equalizer low-frequency band gain adjustment parameter, and the channel mode switching parameter need to be adjusted. First, calculate the correlation coefficients of each adjustment parameter with the outdoor market environment type: Since the market is noisy and a larger volume is required to ensure audibility, the correlation coefficient of the volume adjustment parameter with the environment type reaches 0.8 after calculation; for the equalizer low-frequency band gain adjustment parameter, considering that there is more low-frequency noise in the market environment, appropriately increasing the low-frequency gain can make the music more penetrating, and the correlation coefficient is 0.6; the channel mode switching parameter relatively helps less in adapting to this environment, and the correlation coefficient is only 0.3.

[0089] Based on these correlation coefficients, determine the association degree values. The volume adjustment parameter has a high association degree, and set its adjustment step size to increase by 2 volume units each time; the equalizer low-frequency band gain adjustment parameter has the second-highest association degree, and the step size is set to increase by a gain value of 0.3 each time; the channel mode switching parameter has a low association degree, and the step size is set to consider switching only once every 5 detections (to avoid frequent and unnecessary switching).

[0090] Finally, adjust the adjustment parameters according to their respective adjustment step sizes. The smart glasses can quickly increase the volume in steps, and at the same time, the equalizer low-frequency band gain increases steadily. After several adjustments, the audio output perfectly adapts to the outdoor market environment, making the music clearly audible and full of a live feeling, and allowing users to enjoy high-quality audio without frequent manual adjustment.

[0091] It can be concluded from the above that the adjustment method of setting the step size through the association degree in this embodiment can make the parameters of the audio output more accurately and efficiently adapt to the target environment type, and improve the audio quality and the user's auditory experience in the corresponding environment.

[0092] Corresponding to the audio quality improvement method in the above embodiment, Figure 2 is a structural block diagram of an audio quality improvement device provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 As shown in the figure, the audio quality improvement device 20 includes: an audio separation module 21, an environment recognition module 22, and an audio adjustment module 23.

[0093] Among them, the audio separation module 21 is used to construct a non-negative matrix based on the spectral characteristics of the target mixed audio signal, and separate the target mixed audio signal into multiple target audio signals. The target mixed audio signal is the mixed audio signal in the target environment; The environment recognition module 22 is used to determine the environment type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals; The audio adjustment module 23 is used to adjust the parameters of the output audio signal based on the environment type of the target environment. The output audio signal is the audio signal output by the target device.

[0094] In one embodiment of the present disclosure, the audio separation module 21 is specifically configured to: Construct a non - negative matrix based on the spectral characteristics of the target mixed audio signal to obtain the product of the spectral matrix and the activation intensity matrix of the target audio signal; Obtain multiple target audio signals based on the product of the spectral matrix and the activation intensity matrix of the target audio signal.

[0095] In one embodiment of the present disclosure, the audio separation module 21 is further specifically configured to: Cluster the target mixed audio signal to obtain multiple clusters; Determine the spectral matrix of the target audio signal based on the center of each cluster; Determine the activation intensity matrix of the target audio signal based on the distance between the data points and the center in each cluster.

[0096] In one embodiment of the present disclosure, the audio separation module 21 is further specifically configured to: Calculate the partial derivatives corresponding to the spectral matrix and the activation intensity matrix based on the objective function; Update the spectral matrix and the activation intensity matrix based on the learning efficiency and partial derivatives corresponding to the spectral matrix and the activation intensity matrix to obtain the target spectral matrix and the target activation intensity matrix; Obtain multiple target audio signals based on the product of the target spectral matrix and the target activation intensity matrix of the target audio signal.

[0097] In one embodiment of the present disclosure, the environment recognition module 22 is specifically configured to: Extract the frequency features of multiple target audio signals based on a convolutional neural network; Extract the time features of multiple target audio signals based on a recurrent neural network; Fuse the frequency features and the time features to determine the environment type of the target environment.

[0098] In one embodiment of the present disclosure, the audio adjustment module 23 is specifically configured to: Obtain the acoustic features of the initial output audio signal, where the initial output audio signal is the audio signal output before being adjusted based on the environment type of the target environment; Input the acoustic features into the bio - acoustic perception model to obtain the adjustment parameters of the initial output audio signal.

[0099] In one embodiment of the present disclosure, the audio adjustment module 23 is further specifically configured to: Determine the adjustment step size corresponding to the adjustment parameter based on the correlation degree between the adjustment parameter and the environment type; Adjust the adjustment parameter based on the adjustment step size.

[0100] See Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As shown in Figure 3 , the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned device embodiments, for example Figure 2 the functions of the modules 21 to 23 shown

[0101] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0102] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0103] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0104] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the audio quality improvement method provided in the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0105] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0106] The computer-readable storage medium can be the internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0107] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0109] In several embodiments provided by the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, or can also be electrical, mechanical or other forms of connection.

[0110] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0111] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0112] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. An audio quality improvement method, characterized in that Including: Construct a non - negative matrix based on the spectral characteristics of the target mixed audio signal, and separate the target mixed audio signal to obtain multiple target audio signals, where the target mixed audio signal is a mixed audio signal in a target environment; Determine the environmental type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals; Adjust the parameters of the output audio signal based on the environmental type of the target environment, where the output audio signal is the audio signal output by the target device.

2. The audio quality improvement method according to claim 1, wherein The constructing a non - negative matrix based on the spectral characteristics of the target mixed audio signal and separating the target mixed audio signal to obtain multiple target audio signals includes: Construct a non - negative matrix based on the spectral characteristics of the target mixed audio signal to obtain the product of the spectral matrix and the activation intensity matrix of the target audio signal; Obtain multiple target audio signals based on the product of the spectral matrix and the activation intensity matrix of the target audio signal.

3. The audio quality improvement method according to claim 2, wherein Also including: Cluster the target mixed audio signal to obtain multiple clusters; Determine the spectral matrix of the target audio signal based on the center of each cluster; Determine the activation intensity matrix of the target audio signal based on the distance between the data points and the center in each cluster.

4. The audio quality improvement method according to claim 2, wherein Also including: Calculate the partial derivatives corresponding to the spectral matrix and the activation intensity matrix based on the objective function; Update the spectral matrix and the activation intensity matrix based on the learning efficiency corresponding to the spectral matrix and the activation intensity matrix and the partial derivatives to obtain the target spectral matrix and the target activation intensity matrix; Obtain multiple target audio signals based on the product of the target spectral matrix and the target activation intensity matrix of the target audio signal.

5. The audio quality improvement method according to claim 1, wherein Determining the environmental type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals includes: Extract the frequency characteristics of the multiple target audio signals based on a convolutional neural network; Extract the time characteristics of the multiple target audio signals based on a recurrent neural network; Fuse the frequency characteristics and the time characteristics to determine the environmental type of the target environment.

6. The audio quality improvement method according to claim 1, characterized in that, Adjusting the parameters of the output audio signal based on the environmental type of the target environment, where the output audio signal is the audio signal output by the target device, includes: Obtain the acoustic characteristics of the initial output audio signal, where the initial output audio signal is the audio signal output before being adjusted based on the environmental type of the target environment; Input the acoustic characteristics into a bio - acoustic perception model to obtain the adjustment parameters of the initial output audio signal.

7. The audio quality improvement method according to claim 6, characterized in that, Also including: Determine the adjustment step size corresponding to the adjustment parameter based on the correlation degree between the adjustment parameter and the environmental type; Adjust the adjustment parameter based on the adjustment step size.

8. An audio quality improvement device, characterized in that, Including: An audio separation module, which is used to construct a non - negative matrix based on the spectral characteristics of the target mixed audio signal and separate the target mixed audio signal to obtain multiple target audio signals, where the target mixed audio signal is a mixed audio signal in a target environment; An environment recognition module, which is used to determine the environmental type of the target environment based on the frequency characteristics and time characteristics of the multiple target audio signals; An audio adjustment module, configured to adjust parameters of an output audio signal based on an environmental type of the target environment, where the output audio signal is an audio signal output by a target device.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.