Audio enhancement method and system, wearable device and storage medium

Through the audio clip matching technology based on the big model, the working mode of the target audio clip is determined, which solves the problem of improving audio quality in the existing technology, and realizes efficient and personalized audio processing, which significantly improves the audio quality.

CN120148533APending Publication Date: 2025-06-13BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335411.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve audio quality, especially in environments with severe noise interference, where audio clarity and intelligibility are affected.

Method used

Through large-model-based audio clip matching technology, the working mode of the target audio clip is determined. The method includes determining the working mode of the standard audio clip based on the big model, matching the target characteristics of the target audio clip with the standard characteristics of the standard audio clip, and determining the working mode of the target audio clip based on the matching result.

Benefits of technology

It significantly improves the accuracy and efficiency of audio processing, and can personalize the processing according to the characteristics of different audio clips, improving the quality and effect of the overall audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148533A_ABST
    Figure CN120148533A_ABST
Patent Text Reader

Abstract

The invention provides an audio enhancement method and system, wearable equipment and a storage medium, and belongs to the technical field of audio enhancement, and the method comprises the steps: determining a working mode of a standard audio clip based on a large model; the target feature of the target audio clip is matched with the standard feature of the standard audio clip, a matching result is determined, and the target audio clip is one clip in the played audio; and determining the working mode of the target audio clip based on the matching result. According to the audio enhancement method and system, the wearable device and the storage medium provided by the invention, the efficiency and reliability of the audio enhancement method can be improved, and the audio quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of audio enhancement, and more specifically, relates to an audio enhancement method and system, a wearable device, and a storage medium. Background Art

[0002] Audio enhancement plays an extremely important role in the current fields of multimedia and digital signal processing. With the continuous progress of communication technologies and the increasing popularity of multimedia applications, audio quality has become one of the key factors affecting user experience. In a real environment, audio signals are often interfered by various noises, such as background noise, device noise, etc., and these factors will seriously affect the clarity and intelligibility of the audio.

[0003] Therefore, there is an urgent need for an efficient and reliable audio enhancement method to improve audio quality. Summary of the Invention

[0004] The purpose of the present disclosure is to provide an audio enhancement method and system, a wearable device, and a storage medium, so as to improve the efficiency and reliability of the audio enhancement method and improve audio quality.

[0005] In the first aspect of the embodiments of the present disclosure, an audio enhancement method is provided, including: Determining the working mode of a standard audio segment based on a large model; Matching the target features of a target audio segment with the standard features of the standard audio segment to determine a matching result, where the target audio segment is a segment in the played audio; Determining the working mode of the target audio segment based on the matching result.

[0006] In the second aspect of the embodiments of the present disclosure, an audio enhancement system is provided, including: A working mode determination module, configured to determine the working mode of a standard audio segment based on a large model; A matching module, configured to match the target features of a target audio segment with the standard features of the standard audio segment to determine a matching result, where the target audio segment is a segment in the played audio; A working mode selection module, configured to determine the working mode of the target audio segment based on the matching result.

[0007] In the third aspect of the embodiments of the present disclosure, a wearable device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned audio enhancement method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above-described audio enhancement method.

[0009] The beneficial effects of the audio enhancement method, system, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows: The present disclosure can accurately capture and learn high-quality audio features. By matching the target audio segment with the target features and the standard features of the standard audio segment, it can efficiently extract key information from a large amount of audio data, thereby significantly improving the accuracy and efficiency of audio processing. At the same time, determining the working mode of the target audio segment based on the matching result makes the audio processing more targeted and adaptable, and can perform personalized processing according to the characteristics of different audio segments, thereby improving the overall quality and effect of the audio. The present disclosure not only improves the accuracy and efficiency of audio processing, but also provides reliable technical support for the overall improvement of audio quality, and has broad application prospects and practical value. Therefore, the present disclosure can improve the efficiency and reliability of the audio enhancement method and improve the audio quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0011] Figure 1 It is a schematic flowchart of an audio enhancement method provided by an embodiment of the present disclosure; Figure 2 It is a block diagram of the structure of an audio enhancement system provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0013] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the drawings.

[0014] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of an audio enhancement method provided by an embodiment of the present disclosure. The method includes: S101: Determine the working mode of the standard audio segment based on the large model.

[0015] In this embodiment, the large model is a deep learning model with a large number of parameters. It can be trained with a large amount of data and can learn complex patterns and feature representations. The large model uses its powerful learning and generalization capabilities to analyze the input standard audio segment, mine the potential patterns and features in the audio, and thus determine the working mode of the standard audio segment. The standard audio segment is a pre-set audio segment, serving as a reference benchmark for subsequent audio analysis and comparison. The standard audio segment provides a standard for judging the working mode of the target audio segment.

[0016] The working mode is a mode for audio processing of the standard audio segment, which may include a first working mode, a second working mode, and a third working mode. The first working mode may be an audio noise reduction mode, the second working mode may be an audio noise reduction mode with a first ratio and an audio enhancement mode with a second ratio, and the third working mode may be an audio enhancement mode.

[0017] In this embodiment, the large model is used to analyze the selected standard audio segment. Through mechanisms such as feature extraction and pattern recognition inside the model, it is determined which specific working mode the standard audio segment belongs to.

[0018] S102: Match the target features of the target audio segment with the standard features of the standard audio segment to determine the matching result. The target audio segment is a segment in the playing audio.

[0019] In this embodiment, the playing audio includes multiple segments. The target audio segment is a segment intercepted from the overall playing audio. The target audio segment can be divided based on time or audio content. By analyzing the relationship between the target audio segment and the standard audio segment in this embodiment, the corresponding processing method for the target audio segment can be determined to achieve audio enhancement.

[0020] The target features are a set of information that can characterize the characteristics of the target audio segment extracted from the target audio segment. The target features can be extracted from multiple dimensions. For example, noise features, time-domain features, frequency-domain features, and timbre features, etc. The target features are used to compare with the standard features of the standard audio segment to judge the similarity degree between the target audio segment and the standard audio segment, providing a basis for determining the matching result. The standard features are the features extracted from the standard audio segment, corresponding to the target features.

[0021] In this embodiment, the result reflecting the similarity between the target audio segment and the standard audio segment obtained through the matching process can be represented by a numerical value, such as a similarity score; a category, such as highly matched, moderately matched, or lowly matched; or other forms. Based on the matching result, the characteristics of the target audio segment can be determined, and then the working mode to be adopted for it can be decided to achieve the purpose of audio enhancement.

[0022] S103: Determine the working mode of the target audio segment based on the matching result.

[0023] In this embodiment, the matching result is the information reflecting the similarity between the two obtained after comparing the target features of the target audio segment with the standard features of the standard audio segment. The matching result can be a specific numerical value, such as a similarity score; or it can be a classification description, such as highly matched, moderately matched, unmatched, etc. The matching result essentially quantifies or qualitatively shows the fit of the target audio segment with each standard audio segment at the feature level.

[0024] The matching result is used to determine which working mode should be adopted for the target audio segment for audio processing. Different working modes correspond to different audio processing operations, aiming to solve different problems in the audio or achieve different audio enhancement effects. For example, the audio noise reduction mode mainly removes the noise in the audio; the audio enhancement mode focuses on improving the quality of the audio in terms of clarity, volume, timbre, etc.

[0025] Exemplarily, assume that this embodiment is processing a podcast audio recorded in a noisy coffee shop, aiming to improve its audio quality.

[0026] First, determine the standard audio segment and its working mode.

[0027] Collect the standard audio segments: Standard audio segment A: A 10 - second human voice segment recorded in a professional recording studio, clear and without background noise.

[0028] Standard audio segment B: A 10 - second human voice segment recorded in a mildly noisy environment (similar to the atmosphere of a library).

[0029] Standard audio segment C: A 10 - second human voice segment recorded in a noisy street environment.

[0030] Second, determine the working mode based on the large - model.

[0031] Use a pre-trained large audio processing model, such as a model based on the Transformer architecture and trained on a vast amount of audio data. Input the standard audio clip A into the large model, and the model analyzes its audio features, such as a stable spectrum, clear fundamental frequency, and formants, etc., and determines its working mode as the third working mode, that is, the audio enhancement mode, because the main requirement of this clip is to further enhance the audio expressiveness, such as optimizing the timbre and enhancing the dynamic range.

[0032] For the standard audio clip B, the large model analyzes that there is a certain amount of low-frequency background noise and the human voice is relatively clear, and determines its working mode as the second working mode. Assuming that according to the ratio of noise to useful audio, the proportion of the audio noise reduction mode is 30% and the proportion of the audio enhancement mode is 70%.

[0033] In the standard audio clip C, the background noise is strong and seriously interferes with the human voice, and the large model determines its working mode as the first working mode, that is, the audio noise reduction mode.

[0034] Third, process the target audio clip.

[0035] Extract the target audio clip: From the podcast audio recorded in a café, select an audio clip with a duration of 5 seconds as the target audio clip.

[0036] Extract the target features: Use the Mel-Frequency Cepstral Coefficients (MFCC) algorithm to extract the features of the target audio clip. MFCC can effectively simulate the human ear's perception characteristics of sound frequencies, and the extracted feature vectors contain important feature information of the audio, such as features reflecting the fundamental frequency and formants of the human voice.

[0037] Fourth, feature matching and determination of the working mode.

[0038] Feature matching: Match the MFCC feature vectors of the target audio clip with the MFCC feature vectors of the three standard audio clips. Use the cosine similarity algorithm to calculate the similarity scores.

[0039] The similarity score with the standard audio clip A is 0.45; the similarity score with the standard audio clip B is 0.7; the similarity score with the standard audio clip C is 0.55.

[0040] Determine the working mode: Since the similarity between the target audio clip and the standard audio clip B is the highest, the working mode of the target audio clip is determined to be the second working mode. According to the working mode setting corresponding to the standard audio clip B, the processing of this target audio clip will adopt 30% audio noise reduction mode and 70% audio enhancement mode.

[0041] As can be seen from the above, the present disclosure can accurately capture and learn high-quality audio features. By matching the target audio segment with the target features and the standard features of the standard audio segment, key information can be efficiently extracted from a large amount of audio data, thereby significantly improving the accuracy and efficiency of audio processing. At the same time, determining the working mode of the target audio segment based on the matching result makes the audio processing more targeted and adaptable, and can perform personalized processing according to the characteristics of different audio segments, thereby improving the overall quality and effect of the audio. The present disclosure not only improves the accuracy and efficiency of audio processing, but also provides reliable technical support for the overall improvement of audio quality, and has broad application prospects and practical value. Therefore, the present disclosure can improve the efficiency and reliability of the audio enhancement method and improve the audio quality.

[0042] In an embodiment of the present disclosure, before determining the working mode of the standard audio segment based on the large model, it further includes: Determining a first audio spectrogram based on the acoustic wave characteristics of the preset audio; Determining a second audio spectrogram based on the electromagnetic signal characteristics of the preset audio; Fusing the first audio spectrogram and the second audio spectrogram to determine a standard audio spectrogram; Dividing the standard audio based on the frequency characteristics of the standard audio spectrogram to determine multiple standard audio segments.

[0043] In this embodiment, the preset audio is a piece of audio set in advance, and the acoustic wave characteristics are the characteristics possessed by the preset audio signal when it propagates in the form of acoustic waves. The above characteristics include, but are not limited to, the frequency, amplitude, phase, etc. of the acoustic wave. By analyzing the acoustic wave characteristics in this embodiment, the basic attributes of the audio signal can be understood, and then a spectrogram that can reflect information such as the frequency distribution of the audio, that is, the first audio spectrogram, can be generated.

[0044] In this embodiment, the acoustic wave of the preset audio can be analyzed, and the preset audio can be transformed from the time domain to the frequency domain by using the Fourier transform to graphically display the energy distribution of the audio at different frequencies, and obtain the first spectrogram.

[0045] When the audio signal is transmitted by electromagnetic means (such as radio broadcast, Bluetooth transmission, etc.), it will carry some information related to electromagnetic characteristics. For example, the carrier frequency of the signal, the modulation method, and the distribution of the signal strength in different frequency bands, etc. The above characteristics are correlated with the content of the audio itself, and relevant information of the audio can also be obtained by analyzing the electromagnetic signal characteristics. Analyze and process the signal characteristics of the preset audio during electromagnetic transmission, and convert it into a frequency domain representation to obtain the second audio spectrogram.

[0046] Fuse the first audio spectrogram and the second audio spectrogram to determine a standard audio spectrogram, including: performing weighted calculation on the energy values of corresponding frequency points in the first audio spectrogram and the energy values of corresponding frequency points in the second audio spectrogram based on weighted average to obtain the standard audio spectrogram. The standard audio spectrogram not only contains details at the sound wave level but also integrates frequency feature information related to electromagnetic transmission.

[0047] In this embodiment, analyze the change of frequency features in the standard audio spectrogram, and divide the entire standard audio into multiple segments with different frequency characteristics according to the above changes. Each segment is a standard audio segment.

[0048] It can be concluded from the above that in this embodiment, the standard audio spectrogram is constructed by fusing the sound wave features and electromagnetic signal features of the preset audio, and the standard audio segments are divided accordingly, significantly enhancing the comprehensiveness and accuracy of audio analysis.

[0049] In an embodiment of the present disclosure, there are multiple working modes, and each working mode includes audio noise reduction and / or audio enhancement; Audio noise reduction is to perform noise reduction processing on the standard audio segment based on the wavelet transform algorithm; Audio enhancement is to perform enhancement processing on the standard audio segment based on the independent component analysis algorithm.

[0050] In this embodiment, audio noise reduction is used to reduce or eliminate the noise mixed in the audio signal, making the audio clearer. The noise can come from the background environment. The wavelet transform algorithm decomposes the audio signal of the standard audio segment into components with different frequencies and time scales. In audio noise reduction, by analyzing these components, the wavelet coefficients corresponding to the noise are identified, and then the above coefficients are processed to remove or weaken the noise components. Finally, the audio is reconstructed through inverse wavelet transform to achieve noise reduction of the standard audio segment.

[0051] Audio enhancement can make the audio sound more full and rich by improving the clarity of the audio, enhancing the timbre, and expanding the dynamic range, etc. When the quality of the audio itself is poor or the expressiveness needs to be further improved, audio enhancement can make the audio reach a better auditory effect, such as making the music audio more infectious. The independent component analysis algorithm assumes that the mixed audio signal is composed of multiple independent source signals, and a separation matrix is found to separate the mixed signal into independent components. In audio enhancement, this algorithm is used to separate the different components in the standard audio segment, then the components to be enhanced are processed separately, and finally the processed components are recombined to achieve audio enhancement.

[0052] There are multiple different audio processing strategies for the working mode. Each strategy can only include an audio noise reduction operation, or only include an audio enhancement operation, or can also include both audio noise reduction and audio enhancement operations at the same time.

[0053] As can be seen from the above, the application of multiple working modes can flexibly meet the audio processing requirements in different scenarios. This embodiment not only improves the audio quality, but also enhances the adaptability and practicality of audio processing, bringing a better auditory experience to users.

[0054] In an embodiment of the present disclosure, the wavelet transform algorithm includes wavelet coefficients; Performing noise reduction processing on a standard audio segment based on the wavelet transform algorithm includes: Adjusting the weight of the standard audio segment corresponding to the wavelet coefficient based on the user's basic information; Reconstructing the standard audio segment based on the wavelet coefficient and the weight of the standard audio segment corresponding to the wavelet coefficient to determine the denoised standard audio segment.

[0055] In this embodiment, the wavelet coefficient is the coefficient obtained by performing a convolution operation on the audio signal of the standard audio segment and the wavelet function during the wavelet transform process. The above coefficients reflect the similarity degree between the audio signal and the wavelet function at different scales (corresponding to different frequency ranges) and positions (corresponding to different time points), and contain rich information of the audio signal. Audio components with different frequency and time characteristics correspond to different wavelet coefficient values.

[0056] The standard audio segment can be a low-frequency segment, a mid-frequency segment or a high-frequency segment.

[0057] Adjusting the weight of the standard audio segment corresponding to the wavelet coefficient based on the user's basic information includes: In response to the user's basic information being children and adolescents, the weight of the low-frequency segment corresponding to the wavelet coefficient is the first proportion, the weight of the mid-frequency segment corresponding to the wavelet coefficient is the second proportion, and the weight of the high-frequency segment corresponding to the wavelet coefficient is the third proportion. The sum of the first proportion, the second proportion and the third proportion is 1, the second proportion is greater than the third proportion, and the third proportion is greater than the first proportion; In response to the user's basic information being adults, increasing the weight of the low-frequency segment corresponding to the wavelet coefficient by the first step length, that is, the weight of the fourth proportion, decreasing the weight of the mid-frequency segment corresponding to the wavelet coefficient by the second step length, that is, the weight of the fifth proportion, and decreasing the weight of the high-frequency segment corresponding to the wavelet coefficient by the third step length, that is, the weight of the sixth proportion. The sum of the fourth proportion, the fifth proportion and the sixth proportion is 1, the fifth proportion is greater than the fourth proportion, and the fourth proportion is greater than the sixth proportion; The first step length, the second step length and the third step length can be set according to experience; In response to the user's basic information indicating that the user is an elderly person, increase the weight of the low-frequency segment corresponding to the wavelet coefficients by the fourth step length, i.e., the weight of the seventh proportion, decrease the weight of the mid-frequency segment corresponding to the wavelet coefficients by the fifth step length, i.e., the weight of the eighth proportion, and decrease the weight of the high-frequency segment corresponding to the wavelet coefficients by the sixth step length, i.e., the weight of the ninth proportion. The sum of the seventh proportion, the eighth proportion, and the ninth proportion is 1, the seventh proportion is greater than the eighth proportion, and the eighth proportion is greater than the ninth proportion. The fourth step length, the fifth step length, and the sixth step length can be set according to experience.

[0058] The user's basic information may include the user's age, hearing condition, usage scenario preference, etc. According to the above information, different weights are assigned to the standard audio segments corresponding to the wavelet coefficients. For example, if the user is an elderly person and is more sensitive to low-frequency sounds, the weight of the standard audio segment associated with the wavelet coefficients corresponding to the low-frequency part will be appropriately increased.

[0059] After adjusting the weights of the standard audio segments corresponding to the wavelet coefficients, use the inverse wavelet transform to recombine the above wavelet coefficients with the new weights to restore them into an audio signal. In this process, since the wavelet coefficients corresponding to the noise are weakened, the reconstructed audio signal is the denoised standard audio segment and meets the personalized needs of the user based on the basic information.

[0060] It can be concluded from the above that in this embodiment, by flexibly adjusting the weights of the standard audio segments corresponding to the wavelet coefficients, customized noise reduction can be performed according to the personalized needs of different users, improving the pertinence and effectiveness of the noise reduction process. Secondly, by combining weight adjustment with wavelet coefficients to reconstruct the standard audio segments, not only the important features of the audio signal are retained, but also the noise interference is effectively removed, significantly improving the quality of the denoised audio segment. This embodiment not only enhances the clarity of the audio but also optimizes the user's auditory experience.

[0061] In an embodiment of the present disclosure, the standard audio segment is enhanced based on the independent component analysis algorithm, including: Determine the initial separation matrix; Update the initial separation matrix based on the iterative formula to obtain the target separation matrix; Determine the independent audio components based on the target separation matrix; Perform inverse processing on the independent audio components to obtain the enhanced audio signal frames; Overlap and add adjacent enhanced audio signal frames to obtain the enhanced standard audio segment.

[0062] At the beginning of the independent component analysis algorithm, an initial matrix needs to be set for the separation process, that is, the separation matrix is initialized. This matrix is the starting point for subsequent iterative updates, and the values of its elements will affect the convergence speed and final result of the algorithm. The separation matrix is initialized randomly, but it can also be set based on some prior knowledge or experience.

[0063] Using the iterative formula of the independent component analysis algorithm, the initialized separation matrix is adjusted repeatedly. In each iteration, according to the state of the current separation matrix and the characteristics of the mixed audio signal, the updated values of the matrix elements are calculated, and the separation matrix is gradually optimized in the direction of accurately separating the independent audio components. After multiple iterations, when a certain convergence condition is met, such as the change in matrix elements is less than a certain threshold, the target separation matrix is obtained.

[0064] Multiply the mixed standard audio segment by the target separation matrix. Through the above matrix operation, the mixed audio signal is decomposed into multiple independent audio components according to the transformation relationship defined by the target separation matrix. The above components can respectively correspond to different sound sources, such as speech, the sounds of different musical instruments, etc.

[0065] After separating the independent audio components, corresponding enhancement processing is performed on each component, such as increasing the volume of the speech component and optimizing the timbre of the background music. Then, the above enhanced components are converted back to the time domain through inverse transformation to obtain the enhanced audio signal frames.

[0066] Since the audio signal is usually divided into multiple frames during processing, in order to ensure the continuity and smoothness of the audio, the overlapping parts of adjacent enhanced audio signal frames are added. That is, there is a certain proportion of overlapping areas between adjacent frames. In these overlapping areas, the audio sample values at the corresponding positions of the two frames are added, and then combined into a complete audio segment to obtain the enhanced standard audio segment.

[0067] It can be concluded from the above that in this embodiment, by initializing the separation matrix and gradually optimizing it to the target matrix, the independent components in the audio can be accurately separated, the noise and interference can be effectively removed, and the audio quality can be improved. After inverse processing of the independent components, clear and pure enhanced audio signal frames can be obtained, and through overlapping addition, the coherence and naturalness of the audio segment are ensured, avoiding auditory discomfort caused by frame discontinuity.

[0068] In an embodiment of the present disclosure, the target features of the target audio segment and the standard features of the standard audio segment are matched to determine the matching result, including: Determine the target noise feature and target timbre feature of the target audio segment; Determine the standard noise feature and standard timbre feature of the standard audio segment; Determine the noise similarity based on the target noise feature and the standard noise feature; Determine the timbre similarity based on the target timbre feature and the standard timbre feature; Among them, both the noise similarity and the timbre similarity are matching results.

[0069] In this embodiment, the target noise feature is the characteristic information about noise extracted from the target audio segment, such as the frequency distribution of the noise, the intensity level, the type of the noise, etc. The target timbre feature reflects the characteristic information of the sound feature in the target audio segment, such as the fundamental frequency, the harmonic structure, the formant frequency and its intensity, etc. The above features determine the unique quality of the sound, such as distinguishing the sound features of different musical instruments or different people speaking.

[0070] The standard noise feature is the characteristic information about noise that the standard audio segment has, and also covers aspects such as the frequency distribution, intensity, and type of the noise, serving as a reference standard for measuring the noise feature of the target audio segment.

[0071] The standard timbre feature represents the characteristic information of the sound feature of the standard audio segment, including relevant information such as the fundamental frequency, harmonics, and formants, and is the reference basis for judging the timbre feature of the target audio segment.

[0072] The target noise feature is used to compare with the standard noise feature of the standard audio segment to measure the similarity degree of the target audio segment and the standard audio segment in terms of noise. The target timbre feature is compared with the standard timbre feature of the standard audio segment to determine the similarity degree of the target audio segment and the standard audio segment in terms of timbre, which helps to judge the sound type of the target audio segment and provides a reference for audio enhancement and other processing.

[0073] Calculate the noise similarity based on the first formula; The feature vector of the target noise feature is , and the feature vector of the standard noise feature is ; The first formula is:

[0074] Among them, represents the noise similarity, represents the norm of the feature vector of the target noise feature, represents the norm of the feature vector of the standard noise feature.

[0075] Calculate the timbre similarity based on the second formula; Extract the target fundamental frequency feature of the target timbre feature , the target harmonic amplitude feature , ; Extract the standard fundamental frequency feature of the standard timbre feature , Standard harmonic amplitude feature , ; Determine the fundamental frequency distance ; Determine the harmonic amplitude distance ; Among them, represents the number of harmonics; The second formula is: ; ; Among them, represents the weight of the fundamental frequency distance, represents the weight of the harmonic amplitude distance.

[0076] Determine the working mode of the target audio segment based on the matching result, including: In response to the noise similarity being greater than or equal to the first noise similarity threshold and the timbre similarity being less than the first timbre similarity threshold, use the audio noise reduction mode as the working mode of the target audio frequency band; In response to the noise similarity being greater than or equal to the first noise similarity threshold and the timbre similarity being greater than or equal to the first timbre similarity threshold, combine the audio noise reduction mode with the first weight and the audio enhancement mode with the second weight as the working mode of the target audio frequency band; In response to the noise similarity being less than the first noise similarity threshold and the timbre similarity being greater than or equal to the first timbre similarity threshold, use the audio enhancement mode as the working mode of the target audio frequency band; In response to the noise similarity being less than the first noise similarity threshold and the timbre similarity being less than the first timbre similarity threshold, combine the audio noise reduction mode with the third weight and the audio enhancement mode with the fourth weight as the working mode of the target audio frequency band.

[0077] Among them, the degrees of the first weight, the second weight, the third weight, and the fourth weight are different and can be set according to experience.

[0078] It can be concluded from the above that in this embodiment, by separately extracting the noise and timbre characteristics of the target audio segment and the target audio segment, a refined analysis of the audio characteristics is realized. This embodiment not only improves the accuracy of audio matching, but also effectively reduces the subjectivity of human judgment and enhances the reliability and stability of the matching.

[0079] In an embodiment of the present disclosure, the audio enhancement method further includes: switching the working mode in response to the number of targets of the target audio segment within the first time period being greater than the first number.

[0080] In this embodiment, the working modes include a first working mode, a second working mode, and a third working mode; the first working mode is audio noise reduction, the second working mode is audio noise reduction and audio enhancement, and the third working mode is audio enhancement; In response to the number of target audio segments within a first time period being greater than a first number, switching the working mode includes: If the current working mode is audio noise reduction, in response to the number of target audio segments within the first time period being greater than the first number and the timbre similarity being less than a first threshold, switch the first working mode to the second working mode; If the current working mode is audio noise reduction and audio enhancement, in response to the number of target audio segments within the first time period being greater than the first number and the timbre similarity being greater than or equal to the first threshold, switch the second working mode to the third working mode; If the current working mode is audio enhancement, in response to the number of target audio segments within the first time period being greater than the first number and the noise similarity being less than a second threshold, switch the third working mode to the second working mode; If the current working mode is audio noise reduction and audio enhancement, in response to the number of target audio segments within the first time period being greater than the first number and the noise similarity being greater than or the second threshold, switch the second working mode to the first working mode; Wherein, the first number, the first threshold, and the second threshold can be set according to experience.

[0081] The target number can be divided according to the duration of the played audio and the duration of the first time period.

[0082] When processing audio in this embodiment, it will monitor the number of target audio segments within the first time period. When the target number exceeds the first number, the system will switch to different working modes according to the current working mode, timbre similarity, and noise similarity according to the above rules, so as to process the audio more flexibly and effectively.

[0083] It can be concluded from the above that when it is detected that the number of target audio segments within the first time period exceeds the preset first number, the working mode can be automatically switched to effectively meet the high-density audio processing requirements, ensure the efficiency and quality of audio enhancement processing, and improve the user experience.

[0084] Corresponding to the audio enhancement method in the above embodiment, Figure 2 is a structural block diagram of an audio enhancement system provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The audio enhancement system 20 includes: a working mode determination module 21, a matching module 22, and a working mode selection module 23.

[0085] Among them, the working mode determination module 21 is used to determine the working mode of the standard audio segment based on the large model; The matching module 22 is used to match the target features of the target audio segment and the standard features of the standard audio segment to determine the matching result. The target audio segment is a segment in the played audio; The working mode selection module 23 is used to determine the working mode of the target audio segment based on the matching result.

[0086] In an embodiment of the present disclosure, the audio enhancement system 20 further includes: a standard audio segment determination module; The standard audio segment determination module is used to determine the first audio spectrogram based on the acoustic wave features of the preset audio; Determine the second audio spectrogram based on the electromagnetic signal features of the preset audio; Fuse the first audio spectrogram and the second audio spectrogram to determine the standard audio spectrogram; Divide the standard audio based on the frequency features of the standard audio spectrogram to determine multiple standard audio segments.

[0087] In an embodiment of the present disclosure, there are multiple working modes, and each working mode includes audio noise reduction and / or audio enhancement; Audio noise reduction is to perform noise reduction processing on the standard audio segment based on the wavelet transform algorithm; Audio enhancement is to perform enhancement processing on the standard audio segment based on the independent component analysis algorithm.

[0088] In an embodiment of the present disclosure, the wavelet transform algorithm includes wavelet coefficients; The audio enhancement system 20 further includes: an audio noise reduction module; The audio noise reduction module is used to adjust the weight of the standard audio segment corresponding to the wavelet coefficients based on the basic information of the user; Reconstruct the standard audio segment based on the wavelet coefficients and the weight of the standard audio segment corresponding to the wavelet coefficients to determine the denoised standard audio segment.

[0089] In an embodiment of the present disclosure, the audio enhancement system 20 further includes: an audio enhancement module; The audio enhancement module is used to determine the initial separation matrix; Update the initial separation matrix based on the iterative formula to obtain the target separation matrix; Determine the independent audio components based on the target separation matrix; Perform inverse processing on the independent audio components to obtain the enhanced audio signal frames; Overlap and add adjacent enhanced audio signal frames to obtain the enhanced standard audio segment.

[0090] In an embodiment of the present disclosure, the matching module 22 is specifically configured to determine the target noise feature and the target timbre feature of the target audio segment; determine the standard noise feature and the standard timbre feature of the standard audio segment; determine the noise similarity based on the target noise feature and the standard noise feature; determine the timbre similarity based on the target timbre feature and the standard timbre feature; wherein both the noise similarity and the timbre similarity are matching results.

[0091] In an embodiment of the present disclosure, the audio enhancement system 20 further includes: a working mode switching module; The working mode switching module is configured to switch the working mode in response to the target number of target audio segments in the first time period being greater than the first number.

[0092] See Figure 3 , Figure 3 is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. As Figure 3 shown, the wearable device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of the various modules / units in the above system embodiments, for example Figure 2 the functions of the modules 21 to 23 shown.

[0093] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0094] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the orientation information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0095] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may further include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0096] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first embodiment and the second embodiment of the audio enhancement method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the wearable device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0097] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the foregoing embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0098] The computer-readable storage medium may be an internal storage unit of the wearable device in any of the foregoing embodiments, such as the hard disk or memory of the wearable device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, etc. equipped on the wearable device. Further, the computer-readable storage medium may further include both an internal storage unit and an external storage device of the wearable device. The computer-readable storage medium is used to store the computer program and other programs and data required by the wearable device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0100] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the wearable device and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0101] In several embodiments provided in this application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, or can be electrical, mechanical, or other forms of connection.

[0102] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.

[0103] In addition, the functional units in each embodiment of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0104] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. An audio enhancement method, characterized in that: include: Determine the working mode of standard audio clips based on the large model; Matching a target feature of a target audio segment with a standard feature of the standard audio segment to determine a matching result, wherein the target audio segment is a segment of the played audio; The working mode of the target audio segment is determined based on the matching result.

2. The audio enhancement method according to claim 1, characterized in that: Before the working mode of determining the standard audio segment based on the large model, the method further includes: Determine a first audio spectrogram based on the sound wave characteristics of the preset audio; Determine a second audio spectrum graph based on the electromagnetic signal characteristics of the preset audio; The first audio spectrum graph and the second audio spectrum graph are merged to determine the standard audio spectrum graph; The standard audio is divided based on the frequency characteristics of the standard audio spectrum to determine a plurality of standard audio segments.

3. The audio enhancement method according to claim 1, characterized in that: There are multiple working modes, each of which includes audio noise reduction and / or audio enhancement; The audio noise reduction is to perform noise reduction processing on the standard audio segment based on a wavelet transform algorithm; The audio enhancement is to perform enhancement processing on the standard audio segment based on an independent component analysis algorithm.

4. The audio enhancement method according to claim 3, characterized in that: The wavelet transform algorithm includes wavelet coefficients; The performing noise reduction processing on the standard audio segment based on the wavelet transform algorithm comprises: Adjusting the weight of the standard audio segment corresponding to the wavelet coefficient based on the basic information of the user; The standard audio segment is reconstructed based on the wavelet coefficients and the weights of the standard audio segment corresponding to the wavelet coefficients to determine a standard audio segment after noise reduction.

5. The audio enhancement method according to claim 3, characterized in that: The step of performing enhancement processing on the standard audio segment based on the independent component analysis algorithm includes: Determine the initialization separation matrix; Update the initial separation matrix based on an iterative formula to obtain a target separation matrix; determining independent audio components based on the target separation matrix; Performing inverse processing on the independent audio components to obtain an enhanced audio signal frame; Adjacent enhanced audio signal frames are overlapped and added to obtain enhanced standard audio segments.

6. The audio enhancement method according to claim 1, wherein: The step of matching the target feature of the target audio segment with the standard feature of the standard audio segment to determine a matching result includes: Determining a target noise characteristic and a target timbre characteristic of the target audio segment; Determining standard noise characteristics and standard timbre characteristics of the standard audio segment; Determining noise similarity based on the target noise feature and the standard noise feature; Determining timbre similarity based on the target timbre feature and the standard timbre feature; The noise similarity and the timbre similarity are both the matching results.

7. The audio enhancement method according to claim 1, characterized in that: Also includes: In response to a target number of target audio segments within a first time period being greater than a first number, switching the working mode.

8. An audio enhancement system, characterized in that: include: A working mode determination module, used for determining the working mode of the standard audio clip based on the large model; A matching module, used for matching a target feature of a target audio segment with a standard feature of the standard audio segment to determine a matching result, wherein the target audio segment is a segment of the played audio; A working mode selection module is used to determine the working mode of the target audio segment based on the matching result.

9. A wearable device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.