Audio signal enhancement method, device, apparatus and readable recording medium

The proposed audio signal enhancement method addresses the issue of signal compression by using a trained classifier to isolate and enhance specific audio signals in games, thereby improving the overall audio experience without affecting other audio signals.

JP7681699B2Active Publication Date: 2025-05-22AAC TECHNOLOGIES (NANJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023532254
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-03-16
Publication Date
2025-05-22
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

Existing audio signal enhancement methods in games often compress other audio signals when enhancing weak signals, affecting the overall audio experience.

Method used

An audio signal enhancement method that involves obtaining audio features from actual audio signals, classifying them using a trained classifier, and performing enhancement processing on target audio signals based on their type, thereby isolating and improving the enhancement of specific audio signals without affecting others.

Benefits of technology

Effectively enhances target audio signals while maintaining the quality of other audio signals, improving the accuracy and precision of audio signal enhancement in gaming environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681699000001
    Figure 0007681699000001
  • Figure 0007681699000002
    Figure 0007681699000002
  • Figure 0007681699000003
    Figure 0007681699000003
Patent Text Reader

Abstract

According to the present invention, there is provided an audio signal enhancement method, an apparatus, a device and a readable recording medium. First, a first audio feature corresponding to an actual audio signal is obtained, then the first audio feature is input to a trained classifier for classification and identification to obtain audio type representation data corresponding to the actual audio signal, and finally, referring to the audio type representation data, an enhancement process is performed on a target audio signal that matches a target audio type in the actual audio signal to obtain an enhanced audio signal. According to the embodiment of the present invention, the target audio signal can be effectively enhanced and the enhancement accuracy of the target audio signal can be improved by using the trained classifier to classify and identify the actual audio signal and enhance the target audio signal that matches the target audio type.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of audio processing technology, and in particular to an audio signal enhancement method, device, apparatus and readable recording medium. [Background technology]

[0002] As more and more affluent games at home and abroad have attracted the attention of the public, playing games using electronic devices has become a part of popular culture. Game sounds are reproduced by micro-speakers built into electronic devices such as mobile phones, but due to their weak output, the reproduction effect of weak audio signals during games, such as footsteps, is poor. In the prior art, fixed gain equalizers (EQ) or dynamic range controls (DRC) have been commonly used to enhance weak audio signals in games, but this compresses other audio signals such as gunshots and propeller sounds, or affects the tone of other audio signals when tuning footsteps. Summary of the Invention [Problem to be solved by the invention]

[0003] The present invention aims to provide an audio signal enhancement method, device, equipment and readable recording medium, and at least solves the problem in the prior art that enhancing a target weak audio signal affects the effect of other audio signals. [Means for solving the problem]

[0004] According to a first embodiment of the present invention, there is provided an audio signal enhancement method, the audio signal enhancement method comprising: obtaining a first audio feature corresponding to a real audio signal; inputting the first audio feature into a trained classifier for classification and identification, and obtaining audio type signature data corresponding to the actual audio signal; Referring to the audio type characterization data, performing enhancement processing on a target audio signal that matches the target audio type in the actual audio signal, and obtaining an enhanced audio signal.

[0005] According to a second embodiment of the present invention, an audio signal enhancement device is provided. This audio signal enhancement device The acquisition module is an acquisition module that acquires a first audio feature corresponding to an actual audio signal, The classification module is a classification module that inputs the first audio feature into a trained classifier for classification and identification, and acquires audio type characterization data corresponding to the actual audio signal. The enhancement module is an enhancement module that refers to the audio type characterization data, performs enhancement processing on a target audio signal that matches the target audio type in the actual audio signal, and obtains an enhanced audio signal.

[0006] According to a third embodiment of the present invention, an electronic device is provided. This electronic device includes a memory and a processor. The memory records information including program instructions, and the processor executes the program recorded in the memory. When the processor executes the program, it executes each step in the audio signal enhancement method described in the first embodiment of the present invention.

[0007] According to a fourth embodiment of the present invention, a computer-readable recording medium on which a program is recorded is provided. When the program is executed by a processor, it executes each step in the audio signal enhancement method described in the first embodiment of the present invention.

[0008] As described above, the audio signal enhancing method, device, apparatus and readable recording medium provided by the present invention obtain a first audio feature corresponding to an actual audio signal, input the first audio feature into a trained classifier for classification and identification, obtain audio type representation data corresponding to the actual audio signal, refer to the audio type representation data, and perform an enhancement process on a target audio signal that matches the target audio type in the actual audio signal to obtain an enhanced audio signal. According to the embodiment of the present invention, the target audio signal can be effectively enhanced and the enhancement accuracy of the target audio signal can be improved by using the trained classifier to classify and identify the actual audio signal and enhance the target audio signal that matches the target audio type. [Brief description of the drawings]

[0009] [Figure 1] 1 is a schematic diagram showing a basic flow of an audio signal enhancing method according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is a schematic diagram illustrating a framing process provided by a first embodiment of the present invention. [Diagram 3] FIG. 2 is a waveform diagram showing an input audio provided by the first embodiment of the present invention. [Figure 4] FIG. 2 is a waveform diagram showing output audio provided by the first embodiment of the present invention. [Diagram 5] FIG. 4 is a schematic diagram illustrating a detailed flow of an audio signal enhancement method provided by a second embodiment of the present invention. [Figure 6] FIG. 11 is a schematic diagram illustrating program modules of an audio signal enhancing apparatus provided by a third embodiment of the present invention. [Figure 7] FIG. 13 is a schematic diagram showing a configuration of an electronic device provided according to a fourth embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] In order to make the objectives, features and advantages of the present invention clearer and easier to understand, the following will clearly and in detail describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Of course, the embodiments described below are only some of the embodiments of the present invention, and are not limited thereto. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are included in the protection scope of the present invention.

[0011] In order to solve the problem in the related art that enhancing a target weak audio signal affects the effect of other audio signals, the first embodiment of the present invention provides an audio signal enhancement method. Figure 1 is a basic flowchart of the audio signal enhancement method according to this embodiment. The audio signal enhancement method includes the following steps:

[0012] In step 101, a first audio feature corresponding to a real audio signal is obtained.

[0013] Specifically, in this embodiment, the real audio signals may be various kinds of audio signals in practical scenarios such as games, for example, the audio signals of character footsteps, gunshots, or propeller sounds in a game.

[0014] In some aspects of this embodiment, before the step of obtaining a first audio feature corresponding to the actual audio signal, the method further includes the following steps: performing a framing process on the actual audio signal according to the framing indicator to obtain a second frame signal; extracting audio features in each second frame signal respectively to obtain a combination of second audio features, where the audio features include at least one of a time domain feature, a frequency domain feature, and a time-frequency domain feature; and performing average and variance operations on the target audio feature in the combination of the second audio features to obtain the first audio feature, or performing average and variance operations on the combination of the second audio features of the actual audio signal and the past audio signal to obtain the first audio feature, where the signal collection time point of the past audio signal is earlier than the actual audio signal.

[0015] Specifically, in this embodiment, the framing index includes the unit length of a data frame and the overlap length (i.e., frame shift) of adjacent data frames. As shown in FIG. 2, in this embodiment, framing is preferably performed using overlap framing. Through overlap framing, the transition between frames can be smoothed so as to maintain continuity. The unit frame length is 20 ms, and the frame overlap length is 1 / 2 of the unit frame length, i.e., 10 ms. It should be understood that the specific values ​​of the unit frame length and the frame overlap length in this embodiment are merely typical examples and do not constitute any inherent limitation to this embodiment. After obtaining the frame signals, audio features are extracted from each frame signal. Here, the audio features may be time domain features, frequency domain features, or time-frequency domain features, for example, the frequency domain features may be MFCC (Mel Frequency Cepstrum Coefficient), LPCC (Linear Prediction Cepstral Coefficient). In addition, the extracted audio features are combined to obtain a combination of audio features. Then, to improve the robustness of the audio features, average and variance operations are performed on the combination of audio features of multiple adjacent frame signals. For example, when MFCCs are used as audio features, a set of 40-dimensional MFCC coefficients is extracted every second, and then average and variance operations are performed to obtain an 80-dimensional feature vector every second, which can effectively improve the robustness of the audio features. In addition, to reduce the amount of calculation, the number of adjacent frames used for the average and variance operations can be appropriately reduced, and in a scenario where real-time detection is performed, the average and variance operations may be performed on the combination of audio features of the currently collected audio signal and the previously collected audio signal.

[0016] In step 102, the first audio feature is input to a trained classifier for classification and identification to obtain audio type signature data corresponding to the actual audio signal.

[0017] Specifically, in this embodiment, after obtaining the audio features, the audio features are classified and identified using a trained classifier, and audio type representative data corresponding to the actual audio signal is output. In addition, in this embodiment, the audio types may be represented using 0 and 1, where 1 represents a target audio signal such as footsteps, and 0 represents a non-target audio signal such as non-footsteps.

[0018] In some aspects of this embodiment, before the step of inputting the first audio feature into the trained classifier to classify and identify it, the method further includes the steps of obtaining a predetermined audio signal sample set, obtaining second audio features corresponding to each of the multiple audio signal samples in the audio signal sample set to obtain an audio feature sample set, and training a predetermined classifier model based on the audio feature sample set to obtain a trained classifier.

[0019] Specifically, in this embodiment, the predetermined audio signal sample set includes a target audio signal set (e.g., footstep audio data set) and a non-target audio signal set (e.g., non-footstep audio data set), where the target audio signal set and the non-target audio signal set respectively include the target audio signal (e.g., footstep) and the non-target audio signal (e.g., non-footstep) of each scene, and these two signal sets are used to obtain a classifier, so the two signal sets are of equal size. For example, the footstep audio data set is 1 hour, and the non-footstep audio data set is also 1 hour, and includes audio signals of as many scenes as possible. The audio features of the audio signal samples in the audio signal sample set are respectively extracted to obtain an audio feature sample set, and the audio feature sample set is divided into a training set and a test set, and a pre-prepared classifier model is trained according to the training set in the audio feature sample set and the classification method of machine learning to obtain a classifier that can correctly distinguish between the target audio signal and the non-target audio signal. As a classification method, a general machine learning classification method such as a Support Vector Machine (SVM), a Gaussian Mixture Model (GMM), or a Convolutional Neural Networks (CNN) model may be used.

[0020] In addition, in some aspects of the present embodiment, before the step of respectively obtaining second audio features corresponding to a plurality of audio signal samples in the audio signal sample set, the method further includes the steps of: performing a framing process on each audio signal sample in the audio signal sample set according to a predetermined framing index to obtain a first frame signal, where the framing index includes a data frame unit length and an overlap length of adjacent data frames; extracting audio features in each first frame signal respectively to obtain a combination of first audio features, where the audio features include at least one of a time domain feature, a frequency domain feature, and a time-frequency domain feature; and performing average and variance operations on target audio features in the combination of first audio features to obtain a second audio feature.

[0021] Specifically, in this embodiment, the extraction and dimensions of audio features in the audio feature sample set are the same as those in the actual audio signal, but the number of adjacent frame signals used in performing the operation on the combination of audio features in the audio feature sample set is larger. In addition, the predetermined framing index includes a unit length of a data frame and an overlap length of the data frame, and further performs framing by overlap framing. The unit frame length is 10ms to 20ms, and the overlap length of the frames is 1 / 2 of the unit frame length. After obtaining the frame signals, audio features are extracted from each frame signal. The audio features may be time domain features, frequency domain features, or time-frequency domain features. In addition, the extracted audio features are combined to obtain a combination of audio features. In addition, an average operation and a variance operation are performed on the combination of audio features of multiple adjacent frame signals to obtain the audio features in the audio feature sample set.

[0022] In step 103, referring to the audio type representation data, an enhancement process is performed on the target audio signal that matches the target audio type in the actual audio signal to obtain an enhanced audio signal.

[0023] Specifically, in this embodiment, by referring to the result of the classification output of the classifier, enhancement processing can be performed only on the target audio signal that matches the target audio type in the actual audio signal, and an enhanced audio signal can be obtained.

[0024] In addition, in some aspects of this embodiment, the step of referring to the audio type signature data and performing an enhancement process on the target audio signal in the actual audio signal that matches the target audio type to obtain an enhanced audio signal includes the steps of performing median filtering on the audio type signature data a predetermined number of times to obtain audio type signature data without outliers, and if the audio type signature data without outliers corresponds to the target audio type, performing gain and / or dynamic range enhancement on the target audio signal of a different frequency band that matches the target audio type in the actual audio signal to obtain an enhanced audio signal.

[0025] Specifically, in this embodiment, after the classifier outputs an audio type signature data 0 / 1 signal, the median filter performs median filtering on the 0 / 1 signal, and the median filtering may be performed once or twice, to remove outliers and obtain a square wave signal. The window length of the median filter used in this embodiment is 3. When the audio type signature data is 1, the EQ / DRC performs gain and / or dynamic range enhancement on the target audio signal in different frequency bands. Also, when the audio type signature data is 0, the enhancement process by the EQ / DRC is not performed. Here, the EQ is used for gain on the target audio signal in different frequency bands, and a peak filter is usually used. The DRC may be multi-band, and is used for dynamic compression or enhancement process of different parameters on the target audio signal in different frequency bands, to obtain an enhanced audio signal.

[0026] In addition, in some aspects of this embodiment, the step of performing gain and / or dynamic range enhancement on the target audio signals of different frequency bands in the actual audio signal that match the target audio type includes performing gain and / or dynamic range enhancement on the target audio signals of different frequency bands in the actual audio signal that match the target audio type with reference to a predetermined equalizer fade-in / fade-out time and / or with reference to a predetermined dynamic range control time parameter.

[0027] Specifically, in this embodiment, the enhancement process is performed only on the target audio signal, and not on the non-target audio signal, so that in a hard enhancement method of switching between enhancement and non-enhancement, the sound may become louder or quieter, or a POP sound (level jump) may occur, so the enhancement process may be performed on the target audio signal such as footsteps by setting a fade-in time and a fade-out time and adjusting the gain of the EQ, or the dynamic range enhancement may be performed on the target audio signal such as footsteps by adjusting the time parameter of the DRC. According to such a soft enhancement method, the parameters can be smoothly switched between footsteps and non-footsteps, and the overall playback effect of the target audio signal such as a footstep sound source in an actual scene can be improved.

[0028] In addition, in some aspects of this embodiment, after the step of referring to audio type representation data, performing an enhancement process on a target audio signal that matches the target audio type in the actual audio signal, and obtaining an enhanced audio signal, a step of performing a clipping process on the enhanced audio signal and obtaining an enhanced audio signal without clipping is included.

[0029] Specifically, in this embodiment, a clipping process is performed on the enhanced audio signal by a limiter to prevent the clipping of the enhanced audio signal from becoming too large, and an enhanced audio signal without clipping is obtained. The waveform of the input audio signal is shown in Fig. 3, and the waveform of the output audio signal that has been subjected to the enhancement process and the limiter process is shown in Fig. 4. The horizontal axis of the waveforms shown in Fig. 3 and Fig. 4 represents time in units of s, and the vertical axis represents the sound intensity of the audio signal, i.e., sound pressure, in units of V.

[0030] According to the above technical solution of the embodiment of the present invention, a first audio feature corresponding to an actual audio signal is obtained, the first audio feature is input to a trained classifier for classification and identification, audio type representation data corresponding to the actual audio signal is obtained, and an enhancement process is performed on a target audio signal that matches a target audio type in the actual audio signal by referring to the audio type representation data, to obtain an enhanced audio signal. According to the embodiment of the present invention, the target audio signal is effectively enhanced by using the trained classifier to classify and identify the actual audio signal, and the target audio signal that matches the target audio type is enhanced, so that the target audio signal can be effectively enhanced and the enhancement accuracy of the target audio signal can be improved.

[0031] The method shown in Fig. 5 is a detailed audio signal enhancement method according to a second embodiment of the present invention. The audio signal enhancement method includes the following steps:

[0032] In step 501, a first audio feature corresponding to a real audio signal is obtained.

[0033] Specifically, in this embodiment, the real audio signals may be various kinds of audio signals in practical scenarios such as games, for example, audio signals such as character footsteps, gunshots, propeller sounds, etc. in a game.

[0034] In step 502, a pre-defined classifier model is trained based on a sample set of audio features to obtain a trained classifier.

[0035] Specifically, in this embodiment, the predetermined audio signal sample set includes a target audio signal set (e.g., a footstep audio data set) and a non-target audio signal set (e.g., a non-footstep audio data set), where the target audio signal set and the non-target audio signal set include the target audio signal (e.g., footstep) and the non-target audio signal (e.g., non-footstep) of each scene, respectively. The audio features of the audio signal samples in the audio signal sample set are extracted to obtain an audio feature sample set, and the audio feature sample set is divided into a training set and a test set. A pre-prepared classifier model is trained based on the training set in the audio feature sample set and a machine learning classification method to obtain a classifier that can correctly distinguish between the target audio signal and the non-target audio signal. In addition, the classifier model may be trained using a general machine learning classification method, such as a Support Vector Machine (SVM), a Gaussian Mixture Model (GMM), or a Convolutional Neural Networks (CNN) model.

[0036] In step 503, the first audio feature is input to a trained classifier for classification and identification to obtain audio type signature data corresponding to the actual audio signal.

[0037] Specifically, in this embodiment, after obtaining the audio features, the audio features are classified and identified using a trained classifier, and audio type representative data corresponding to the actual audio signal is output. In addition, in this embodiment, the audio types may be represented using 0 and 1, where 1 represents a target audio signal such as footsteps, and 0 represents a non-target audio signal such as non-footsteps.

[0038] In step 504, median filtering is performed a predetermined number of times on the audio type signature data to obtain audio type signature data without outliers.

[0039] Specifically, in this embodiment, after the classifier outputs the audio type representative data 0 / 1 signal, the median filter performs median filtering on the 0 / 1 signal, the median filtering can be performed once or twice, and the outlier is removed to obtain a square wave signal. The window length of the median filter used in this embodiment is 3.

[0040] In step 505, if the outlier-free audio type signature data corresponds to the target audio type, gain and / or dynamic range enhancement is performed on the target audio signal in different frequency bands that match the target audio type in the actual audio signal to obtain an enhanced audio signal.

[0041] When the audio type signature data is 1, the EQ / DRC performs gain and / or dynamic range enhancement on the target audio signal in different frequency bands. When the audio type signature data is 0, the EQ / DRC does not perform enhancement processing. Here, the EQ is used for gain on the target audio signal in different frequency bands, and typically uses a peak filter. The DRC may be multi-band, and is used for dynamic compression or enhancement processing of different parameters on the target audio signal in different frequency bands to obtain an enhanced audio signal.

[0042] In step 506, a clipping process is performed on the augmented audio signal to obtain a clip-free augmented audio signal.

[0043] Specifically, in this embodiment, in order to prevent the clipping of the enhanced audio signal from becoming too large, a clipping process is performed on the enhanced audio signal by a limiter to obtain an enhanced audio signal without clipping.

[0044] In addition, the magnitude of the codes in each step in this embodiment does not mean the order in which the steps are executed, and the order in which each step is executed should be determined by its function and inherent logic, and does not constitute an inherent limitation on the implementation process of the embodiment of the present invention.

[0045] An embodiment of the present invention provides an audio signal enhancement method, which includes obtaining a first audio feature corresponding to an actual audio signal, inputting the first audio feature into a trained classifier for classification and identification, obtaining audio type representation data corresponding to the actual audio signal, and referring to the audio type representation data, performing an enhancement process on a target audio signal that matches a target audio type in the actual audio signal to obtain an enhanced audio signal. According to the embodiment of the present invention, the target audio signal can be effectively enhanced and the enhancement accuracy of the target audio signal can be improved by using the trained classifier to classify and identify the actual audio signal and enhance the target audio signal that matches the target audio type.

[0046] Fig. 6 is a diagram illustrating an audio signal enhancing device provided by the third embodiment of the present invention. The audio signal enhancing method in the above embodiment can be realized by the audio signal enhancing device. As shown in Fig. 6, the audio signal enhancing device is configured as follows.

[0047] The acquiring module 601 acquires a first audio feature corresponding to a real audio signal. The classification module 602 inputs the first audio feature into a trained classifier for classification and identification, and obtains audio type signature data corresponding to the actual audio signal. The enhancement module 603 refers to the audio type representation data, and performs enhancement processing on the target audio signal that matches the target audio type in the actual audio signal to obtain an enhanced audio signal.

[0048] In some aspects of the present embodiment, the audio signal enhancing apparatus further includes a first calculation module, which performs a framing process on the actual audio signal according to the framing index to obtain second frame signals, respectively extracts audio features in each second frame signal, and obtains a combination of second audio features, where the audio features include at least one of a time domain feature, a frequency domain feature, and a time-frequency domain feature, and performs an average operation and a variance operation on a target audio feature in the combination of second audio features to obtain a first audio feature, or performs an average operation and a variance operation on a combination of the second audio features of the actual audio signal and the past audio signal to obtain the first audio feature, where a signal collection time point of the past audio signal is earlier than the actual audio signal.

[0049] In some aspects of the present embodiment, the audio signal enhancing apparatus further includes a training module, which is configured to obtain a predetermined audio signal sample set, obtain second audio features corresponding to the plurality of audio signal samples in the audio signal sample set, obtain an audio feature sample set, train a predetermined classifier model based on the audio feature sample set, and obtain a trained classifier.

[0050] In some aspects of the present embodiment, the audio signal enhancing apparatus further includes a second calculation module, which performs a framing process on each audio signal sample in the audio signal sample set according to a predetermined framing index to obtain a first frame signal, where the framing index includes a data frame unit length and an overlap length of adjacent data frames, respectively extracts audio features in each first frame signal to obtain a first audio feature combination, where the audio features include at least one of a time domain feature, a frequency domain feature, and a time-frequency domain feature, and performs an average operation and a variance operation on a target audio feature in the first audio feature combination to obtain a second audio feature.

[0051] In addition, in some aspects of the present embodiment, specifically, the enhancement module 603 is used to perform median filtering on the audio type signature data a predetermined number of times to obtain outlier-free audio type signature data, and if the outlier-free audio type signature data corresponds to the target audio type, perform gain and / or dynamic range enhancement on the target audio signal of a different frequency band that matches the target audio type in the actual audio signal to obtain an enhanced audio signal.

[0052] In addition, in some aspects of this embodiment, the enhancement module 603 is adapted to perform gain and / or dynamic range enhancement with reference to predefined equalizer fade-in / fade-out times and / or predefined dynamic range control time parameters for the target audio signal in different frequency bands that match the target audio type in the actual audio signal.

[0053] In some aspects of the present embodiment, the audio signal enhancing apparatus further comprises a clipping module, the clipping module being adapted to perform a clipping process on the enhanced audio signal to obtain a clipping-free enhanced audio signal.

[0054] It should be noted that the audio signal enhancing method in the first embodiment and the second embodiment can both be implemented based on the audio signal enhancing device provided in this embodiment, and those skilled in the art can clearly understand it. In addition, for convenience and conciseness of description, the specific working process of the audio signal enhancing device in this embodiment can refer to the corresponding process in the method embodiment, and detailed description will not be repeated here.

[0055] According to the audio signal enhancing device provided in this embodiment, a first audio feature corresponding to an actual audio signal is obtained, the first audio feature is input to a trained classifier for classification and identification, audio type representation data corresponding to the actual audio signal is obtained, and an enhancement process is performed on a target audio signal that matches a target audio type in the actual audio signal by referring to the audio type representation data, to obtain an enhanced audio signal. According to the embodiment of the present invention, the target audio signal is effectively enhanced by using the trained classifier to classify and identify the actual audio signal, and the target audio signal that matches the target audio type is enhanced, so that the target audio signal can be effectively enhanced and the enhancement accuracy of the target audio signal can be improved.

[0056] Referring to Fig. 7, Fig. 7 is a diagram showing an electronic device provided by a fourth embodiment of the present invention. According to this electronic device, the audio signal enhancement method in the above embodiment can be realized. As shown in Fig. 7, this electronic device includes a memory 701, a processor 702, and a program 703 recorded in the memory 701 and executed by the processor 702. When the program 703 is executed by the processor 702, the audio signal enhancement method in the above embodiment can be realized. Here, the number of processors may be one or more.

[0057] The memory 701 may be a high-speed random access memory (RAM) memory or a non-volatile memory such as a disk memory. The memory 701 is used to store executable program code, and the processor 702 is coupled to the memory 701.

[0058] Furthermore, the present invention provides a computer-readable recording medium. The computer-readable recording medium may be provided in the electronic device in each of the above-described embodiments. The computer-readable recording medium may be the memory in the embodiment shown in FIG. 7.

[0059] When the computer-readable recording medium is executed by a processor, the computer-readable recording medium performs the audio signal enhancing method of the embodiment, and may be various recording media capable of storing program code, such as a USB memory, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a disk, a CD-ROM, etc.

[0060] In some embodiments provided by the present invention, the disclosed devices and methods may be implemented in other forms. For example, the embodiments of the above devices are merely schematic. For example, the module division that is merely a logical function division can be divided in other forms when actually implemented. For example, a plurality of modules or components can be combined, or integrated into another system, or some features can be ignored or not implemented. Also, the illustrated or discussed mutual coupling, direct coupling, or communication connection may be an indirect coupling or communication connection via any interface, device, or module that may be electrical, mechanical, or other methods.

[0061] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be arranged in one place or dispersed among a plurality of network modules. Some or all of these modules can be selected according to practical needs to achieve the objectives of this embodiment.

[0062] Also, each functional module in each embodiment of the present invention may be integrated into one processing module, each module may physically exist separately, or two or more modules may be integrated into one module. The above integrated module may be realized in the form of hardware or in the form of a software functional module.

[0063] When the integrated module is realized as a software function module and sold or used as an independent product, it can be stored in a computer-readable recording medium. Based on this understanding, the technical solution in the present invention can essentially be embodied in the form of a software product, either the part that contributes to the prior art or all or part of the technical solution. This computer software product is stored in a readable recording medium and includes some instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps in each embodiment of the present invention.

[0064] It should be noted that while the above embodiments are described as a series of operations for the sake of simplicity, those skilled in the art should understand that the present invention is not limited by the sequence of operations described, as some steps may be performed in other sequences or simultaneously according to the present invention. It should also be understood by those skilled in the art that the embodiments described herein are preferred embodiments, and that the operations or modules of the present invention are not necessarily required by the present invention.

[0065] In the above embodiments, the description of each embodiment is focused on each, and the details of an embodiment that are not described in detail can be referred to the relevant descriptions of other embodiments.

[0066] The above describes the audio signal enhancement method, device, equipment, and readable recording medium provided by the present invention. However, those skilled in the art should understand that there may be changes in specific implementation and application scope based on the ideas of the embodiments of the present application, and in general, the contents of this specification should not be construed as limiting the present invention.

Claims

1. 1. A method for audio signal enhancement, comprising: obtaining a first audio feature corresponding to a real audio signal; inputting the first audio feature into a trained classifier for classification and identification, and obtaining audio type signature data corresponding to the actual audio signal; and performing an enhancement process on a target audio signal that matches a target audio type in the actual audio signal by referring to the audio type representation data to obtain an enhanced audio signal; The step of obtaining an enhanced audio signal by referring to the audio type representation data and performing an enhancement process on a target audio signal that matches a target audio type in the actual audio signal, comprises: performing median filtering on the audio type signature data a predetermined number of times to obtain audio type signature data without outliers; if the outlier-free audio type signature data corresponds to the target audio type, performing gain and / or dynamic range enhancement on the target audio signal in different frequency bands that match the target audio type in the actual audio signal to obtain an enhanced audio signal; 13. A method for enhancing an audio signal, comprising:

2. prior to the step of inputting the first audio feature to a trained classifier for classification, obtaining a predetermined audio signal sample set; obtaining second audio features corresponding to the audio signal samples in the audio signal sample set, respectively, to obtain an audio feature sample set; training a pre-defined classifier model based on the audio feature sample set to obtain a trained classifier.

2. The method of claim 1, wherein the audio signal is enhanced.

3. prior to the step of obtaining second audio features corresponding to a plurality of audio signal samples in the audio signal sample set, performing a framing process on each audio signal sample in the audio signal sample set according to a predetermined framing indicator to obtain a first frame signal; Wherein, the framing indicator includes a data frame unit length and an overlap length of adjacent data frames; extracting audio features in each of the first frame signals to obtain a combination of first audio features; Wherein the audio features include at least one of time domain features, frequency domain features, and time-frequency domain features; performing an average operation and a variance operation on the target audio feature in the combination of the first audio features to obtain the second audio feature.

3. The method of claim 2, wherein the audio signal is enhanced.

4. prior to the step of obtaining a first audio feature corresponding to the actual audio signal, performing a framing process on the actual audio signal according to the framing indicator to obtain a second frame signal; extracting audio features in each of the second frame signals to obtain a combination of second audio features; Wherein the audio features include at least one of time domain features, frequency domain features, and time-frequency domain features; performing an average operation and a variance operation on a target audio feature in the combination of the second audio features to obtain the first audio feature; or performing an average operation and a variance operation on a combination of the second audio features of the actual audio signal and the past audio signal to obtain the first audio feature; wherein a signal collection time point of the past audio signal is earlier than that of the actual audio signal; Further comprising:

4. The method of claim 3, wherein the audio signal is enhanced.

5. 2. The step of performing gain and / or dynamic range enhancement on the target audio signal in different frequency bands that match the target audio type in the actual audio signal, comprising: performing gains on the target audio signal in different frequency bands corresponding to a target audio type in the actual audio signal with reference to a predetermined equalizer fade-in / fade-out time and / or performing dynamic range enhancement with reference to a predetermined dynamic range control time parameter; 2. The method of claim 1, wherein the audio signal is enhanced.

6. After the step of referring to the audio type representation data, and performing an enhancement process on a target audio signal that matches a target audio type in the actual audio signal to obtain an enhanced audio signal, performing a clipping process on the enhanced audio signal to obtain a clipping-free enhanced audio signal; 2. The method of claim 1, wherein the audio signal is enhanced.

7. 1. An audio signal enhancement device, comprising: an acquisition module for acquiring a first audio feature corresponding to a real audio signal; a classification module for inputting the first audio feature into a trained classifier for classification and identification, and obtaining audio type signature data corresponding to the actual audio signal; an enhancement module for referring to the audio type signature data, and performing an enhancement process on a target audio signal in the actual audio signal that matches a target audio type, to obtain an enhanced audio signal, the enhancement module performing a median filtering on the audio type signature data a predetermined number of times to obtain outlier-free audio type signature data, and when the outlier-free audio type signature data corresponds to the target audio type, performing a gain and / or dynamic range enhancement on the target audio signal of a different frequency band that matches the target audio type in the actual audio signal, to obtain an enhanced audio signal.

1. An audio signal enhancing device comprising:

8. An electronic device, A memory and a processor, the memory stores information including program instructions; The processor executes a program stored in the memory, When the processor executes the program, the processor executes the steps of the method according to any one of claims 1 to 6.

1. An electronic device comprising:

9. A computer-readable recording medium on which a program is recorded, When the program is executed by a processor, the program executes the steps of the method according to any one of claims 1 to 6. A computer-readable recording medium comprising:

Citation Information

Patent Citations

  • Audio classification model training method, audio classification method, device and equipment

    CN111369982A

  • Audio category determination method and device, storage medium and electronic device

    CN113593603A

  • Digital compressor for compressing an audio signal

    JP2016530765A

  • Method and apparatus for dynamic volume adjustment via audio classification

    JP2021536705A

  • Data driven audio enhancement

    US20190392852A1