Speech enhancement method and device
By adding a feature memory system to the feature extraction process of voice signals, recording the voice characteristics of people familiar with the user, and using these features for voice enhancement when the voice signal quality is poor, the problem of poor voice enhancement effect in the prior art is solved, and the voice enhancement effect is improved in noisy environments.
Patent Information
- Application Number
- CN202111368857.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Existing voice enhancement technologies are difficult to effectively improve voice enhancement effects when voice signal quality is poor, especially in noisy environments.
By adding a feature memory system to the feature extraction process of the voice signal, the voice characteristics of people familiar with the user are recorded, and these characteristics are used for speech enhancement when the voice signal is poor.
The speech enhancement effect is improved, especially in noisy environments, where the currently received speech signal is enhanced by utilizing known speech features.
Smart Images

Figure CN113921026B_ABST
Abstract
Description
Technical Field
[0001] This application relates to audio processing technologies, and more particularly, to a voice enhancement method and apparatus. Background Art
[0002] Hearing aids (also known as "hearing assistive devices") are widely used for hearing compensation of hearing-impaired patients. They can amplify sounds that the hearing-impaired patients cannot hear, and then utilize their residual hearing to send the sounds to the auditory center of the brain, so that the patients can perceive the sounds.
[0003] Voice enhancement devices such as hearing aids generally need to use voice enhancement technologies to amplify voice signals in sounds. Existing voice enhancement technologies mainly use single-time voice enhancement algorithms, that is, every time a segment of sound is input, the hearing aid directly runs the relevant voice enhancement algorithm to process the sound input. To reduce latency, most real-time voice enhancement algorithms (especially deep learning-based algorithms) use a system design of feature extraction - model operation. However, in some cases, the quality of the voice signals in the input sound is poor, and it is possible that the voice enhancement system cannot extract sufficient features for voice enhancement. If voice enhancement is performed based on such sound features, it is often difficult to obtain a satisfactory voice enhancement effect.
[0004] Therefore, it is necessary to provide a new voice enhancement method to solve the problems existing in the prior art. Summary of the Invention
[0005] An object of this application is to provide a voice enhancement method, apparatus, and storage medium that can improve the voice enhancement effect when the quality of voice signals is poor.
[0006] The inventors of this application found that many systems using voice enhancement are related to people familiar to the user. For example, the scenarios of voice calls mostly involve calls with family members, colleagues, and friends. Therefore, if a feature memory system can be added during the feature extraction of voice signals, especially by extracting and recording the voice features of people familiar to the patient, then these extracted features will help the voice enhancement algorithm better improve the voice enhancement effect. For example, during a call between the user of a communication device and these familiar people, if the voice features of these people in a quiet environment are already known, then when these people enter a noisy environment and talk to the user, the voice enhancement algorithm adopted by the communication device can utilize the voice features extracted in the previous quiet environment, which helps to improve the voice enhancement effect.
[0007] In one aspect of the present application, a voice enhancement method is provided, the method comprising: receiving a current audio input signal having a voice part and a non-voice part; determining a voice feature of the voice part in the current audio input signal; determining the voice quality of the current audio input signal; evaluating whether the voice quality meets a predetermined voice quality requirement; and in response to the voice quality meeting the predetermined voice quality requirement, creating or updating a reference voice feature with the voice feature, wherein the reference voice feature is used to enhance the voice part in the audio input signal.
[0008] In some embodiments, determining the voice quality of the current audio input signal comprises: determining a voice signal-to-noise ratio of the current audio input signal, the voice signal-to-noise ratio representing a ratio of the power of the voice part and the power of the non-voice part.
[0009] In some embodiments, evaluating whether the voice quality meets a predetermined voice quality requirement comprises: comparing the voice signal-to-noise ratio with a predetermined voice signal-to-noise ratio threshold; and in response to the voice signal-to-noise ratio being greater than the predetermined voice signal-to-noise ratio threshold, determining that the voice quality meets the predetermined voice quality requirement.
[0010] In some embodiments, the method further comprises: obtaining one or more pre-stored reference voice features; and retrieving a reference voice feature matching the voice feature from the one or more pre-stored reference voice features.
[0011] In some embodiments, the method further comprises: in response to no reference voice feature matching the voice feature being retrieved, creating a new reference voice feature with the voice feature of the current audio input signal; and using the voice feature of the voice part in the current audio input signal to enhance the voice part in the current audio input signal.
[0012] In some embodiments, the method further comprises: comparing the duration of the current audio input signal with a predetermined duration threshold; and in response to the duration of the current audio input signal being greater than the predetermined duration threshold, creating a reference voice feature with the voice feature of the current audio input signal.
[0013] In some embodiments, the method further comprises: in response to a reference voice feature matching the voice feature being retrieved, comparing the voice quality of the current audio input signal with the voice quality corresponding to the matching reference voice feature; in response to the voice quality of the current audio input signal being better than the voice quality corresponding to the matching reference voice feature, updating the matching reference voice feature with the voice feature of the current audio input signal; and using the voice feature of the voice part in the current audio input signal to enhance the voice part in the current audio input signal.
[0014] In some embodiments, the method further includes: in response to the voice quality of the current audio input signal being not better than the voice quality corresponding to the matched reference voice feature, using the voice feature of the voice part in the current audio input signal and the matched reference voice feature to enhance the voice part in the current audio input signal.
[0015] In some embodiments, the method further includes: in response to no reference voice feature matching the voice feature being retrieved and the voice quality not meeting the predetermined voice quality requirement, using the voice feature of the voice part in the current audio input signal to enhance the voice part in the current audio input signal.
[0016] In some embodiments, the method further includes: in response to a reference voice feature matching the voice feature being retrieved and the voice quality not meeting the predetermined voice quality requirement, using the voice feature of the voice part in the current audio input signal and the matched reference voice feature to enhance the voice feature.
[0017] In some embodiments, the voice feature includes a fundamental period or Mel cepstral coefficients.
[0018] In some embodiments, determining the voice feature of the voice part in the current audio input signal includes: determining a voice enhancement feature and a voice comparison feature of the voice part in the current audio input signal, where the voice enhancement feature contains more feature information than the voice comparison feature.
[0019] In another aspect of the present application, there is also provided a voice enhancement device, which includes a non-transitory computer storage medium, on which one or more executable instructions are stored, and after being executed by a processor, the one or more executable instructions perform the processing steps of the above aspect.
[0020] In yet another aspect of the present application, there is also provided a non-transitory computer storage medium, on which one or more executable instructions are stored, and after being executed by a processor, the one or more executable instructions perform the processing steps of the above aspect.
[0021] The above is an overview of the present application. There may be simplification, generalization, and omission of details. Therefore, those skilled in the art should recognize that this part is only illustrative and is not intended to limit the scope of the present application in any way. This overview part is neither intended to identify the key features or essential features of the claimed subject matter nor intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other features of the present application will be more fully and clearly understood by reference to the following description and the appended claims, in conjunction with the accompanying drawings. It is to be understood that these drawings only depict several embodiments of the present application and should not be considered as limiting the scope of the present application. By using the accompanying drawings, the present application will be more clearly and detailedly described.
[0023] Figure 1 FIG. 100 is a flowchart showing a voice enhancement method according to an embodiment of the present application;
[0024] Figure 2 FIG. 200 is a flowchart showing a voice enhancement method according to an embodiment of the present application. Detailed Description of the Invention
[0025] In the following detailed description, reference is made to the accompanying drawings which form a part hereof. In the drawings, like reference numerals generally represent like components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, the drawings, and the claims are not intended to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter of the present application. It will be understood that various configurations, substitutions, combinations, and designs of the various aspects of the present application generally described herein and illustrated in the accompanying drawings can be made, all of which are expressly incorporated as part of the present application.
[0026] Figure 1 FIG. 100 is a flowchart showing a voice enhancement method according to an embodiment of the present application. It is to be understood that the voice enhancement method 100 of the present application can be used in various audio devices and implemented as a voice enhancement device coupled to or integrated in an audio device. The audio device may be, for example, a hearing aid or an electronic device such as a headphone or a mobile communication terminal having audio acquisition and / or audio output functions.
[0027] As Figure 1As shown, at step 101, an audio input signal is received by the sound input of the voice enhancement device. For example, the voice enhancement device can be set or integrated in a voice processing device with a microphone such as a Bluetooth headset, a hearing aid, a headset, etc., so that the microphone of these devices can be used to collect ambient sound and generate an audio input signal. These audio input signals can then be provided to the sound input of the voice enhancement device. In some other examples, the sound input of the voice enhancement device can also be communicatively coupled to another voice device, such as a separate microphone or a microphone, in a wired or wireless manner, and receive an audio input signal from these voice devices. Depending on the environment in which the audio input signal is collected, the audio input signal can include a voice part composed of voice and a non-voice part composed of background sound, and the intensity ratio of these two parts may vary. Background sound is usually some sounds in the environment, and they may be sounds that are not expected to be amplified or enhanced; while voice is the sound emitted by a certain person or certain people, which is usually the sound that is expected to be amplified. It can be understood that the voice enhancement method of the embodiments of the present application is to enhance the voice part in the audio input signal, that is, the human voice.
[0028] Generally speaking, in addition to common voice characteristics such as sound intensity, loudness, pitch, etc., the voices emitted by different people have different characteristics, or rather, different voice characteristics. Therefore, the voice signal will include voice characteristics, and these voice characteristics can be characterized by different parameters. For example, the pitch period and the Mel-scale Frequency Cepstral Coefficients (MFCC) can usually be used as voice characteristics to characterize the voices of different people. Specifically, the pitch period reflects the time interval or the opening and closing frequency between two adjacent openings and closings of the glottis, so it is an important characteristic for describing the voice excitation source. The shape of the vocal tract can accurately represent the phonemes it generates, and the shape of the vocal tract is manifested in the form of the envelope of the short-time power spectrum, and the MFCC can represent this envelope, so the MFCC can also be used as a voice characteristic to distinguish different voices. Those skilled in the art can understand that the voice characteristics described herein can also adopt other suitable characteristic parameters, or combinations of these characteristic parameters.
[0029] Correspondingly, at step 102, the voice enhancement feature extraction unit extracts and determines the voice enhancement features of the voice part in the audio input signal. The voice enhancement feature extraction unit can be coupled to the sound input to receive the audio input signal.
[0030] In some embodiments, a deep learning algorithm can be used to extract speech enhancement features of the speech part in an audio input signal. For example, a neural network model can be constructed and trained, and the pitch period and / or speech features such as MFCC of the speech signal in the audio input signal can be extracted through this neural network model. In other embodiments, the audio input signal can also be processed in other ways to extract speech enhancement features. For example, to extract MFCC coefficients, the original audio input signal can be pre-emphasized by a high-pass filter to enhance the high-frequency part, and then framed, windowed, and fast Fourier transformed to obtain the power spectrum of each frame signal; then, Mel filtering, logarithmic energy operation, and discrete cosine transform (DCT) processing can be adopted to obtain the required MFCC coefficients. It can be understood that the above algorithms for extracting speech enhancement features are merely exemplary, and those skilled in the art can select different feature extraction methods according to the characteristics of the speech features to be extracted and the available hardware resources.
[0031] It can be understood that the speech enhancement features extracted by the speech enhancement feature extraction unit will be used for speech signal enhancement subsequently, so preferably it can include more feature information.
[0032] Still referring to Figure 1 , at step 103, the speech quality of the audio input signal is determined by the speech quality prediction unit. Similarly, this speech quality prediction unit can be coupled to the sound input to receive the audio input signal.
[0033] As described in the background section of this application, speech quality also has a significant impact on speech recognition. Poor-quality speech signals may be difficult to extract sufficient features for speech enhancement. Therefore, the speech enhancement method of the embodiments of this application will further determine the quality of the speech signal.
[0034] Speech quality can be characterized by various suitable parameters. In one embodiment, the speech quality can be determined by determining the speech signal-to-noise ratio p-SNR of the audio input signal. Specifically, the speech signal-to-noise ratio represents the ratio of the average power of the speech part to the average power of the non-speech part. In one embodiment, a prediction method based on energy, a prediction method based on cepstrum, or a deep learning method can be used to predict the speech signal-to-noise ratio p-SNR of the audio input signal.
[0035] Further, the speech signal-to-noise ratio p-SNR of the audio input signal can be compared with a preset speech signal-to-noise ratio threshold t-SNR to evaluate the speech quality. Specifically, if the speech signal-to-noise ratio p-SNR exceeds the predetermined threshold t-SNR, it can be considered that the audio input signal contains enough or strong enough speech parts to meet the predetermined speech quality requirements; otherwise, it is considered that the audio input signal does not meet the predetermined speech quality requirements. After determining whether the speech signal-to-noise ratio exceeds the speech signal-to-noise ratio threshold, the audio input signal can be further operated (such as reading or storing, etc.), specifically referring to the description of step 105 below. In one embodiment, the preset speech signal-to-noise ratio threshold t-SNR can be 0.5, but those skilled in the art can set t-SNR to other values according to actual needs, such as 0.3 to 0.6, and the present application does not limit this.
[0036] It can be understood that although the speech quality is determined by exemplarily determining the speech signal-to-noise ratio of the audio input signal in step 103, in other embodiments, other parameters can also be used to evaluate and determine the speech quality, such as using parameters like speech recognition rate. In addition, in other embodiments, in addition to using the relative intensity of the speech part relative to the non-speech part (speech signal-to-noise ratio), the speech quality of the audio input signal can also be determined by determining the absolute intensity of the speech part in the audio input signal, or the absolute intensity of the non-speech signal, and those skilled in the art can make adjustments according to the actual situation.
[0037] As described above, the inventors of the present application found that during the speech enhancement process, if known speech features (usually extracted in a relatively quiet or ideal environment) can be used to assist the speech enhancement process of the currently received speech signal, the speech enhancement effect will be significantly improved. In order to be able to determine which known speech feature information the current speech signal corresponds to, the speech enhancement method 100 further includes step 104 for extracting comparison features.
[0038] Specifically, at step 104, the speech comparison feature extraction unit extracts the speech comparison features of the speech part in the audio input signal for comparison with one or more pre-stored reference speech features at step 105.
[0039] Similarly, similar to extracting the voice enhancement features of the voice part in the audio input signal in step 102, in step 104, a deep neural network model can also be constructed and trained using, for example, a deep learning algorithm, and the voice comparison features in the voice part can be determined by extracting parameters such as the fundamental period or MFCC, Filter Bank, etc., for subsequent voice feature comparison in step 105. In one embodiment, the voice enhancement features extracted at step 102 and the voice comparison features extracted at step 104 can be at least partially different. In a certain embodiment, in order to save processing resources, the voice comparison features can have less feature information, for example, less than the feature information included in the voice enhancement features. For example, the voice comparison features extracted in step 104 can be the voice features used in speaker identification (or voice ID) technology, that is, voiceprint features; while the voice enhancement features can include features such as MFCC, Filter Bank, etc. in addition to Voice ID. In other embodiments, the voice comparison features and the voice enhancement features have the same features. In addition, the voice comparison features and the voice enhancement features can also have completely different features. For example, the voice enhancement features can be MFCC, while the voice comparison features can be, for example, the vector I-vector (Identity Vector) mapped on the total factor matrix. The voice enhancement features usually do not include the vector I-vector.
[0040] It should be noted that although Figure 1 shows two independent steps 102 (extracting voice enhancement features for the voice enhancement algorithm) and step 104 (extracting voice comparison features for voice comparison) to extract the feature information in the human voice, those skilled in the art can understand that the voice enhancement features extracted at step 102 for the voice enhancement algorithm and the voice comparison features extracted at step 104 for voice comparison can also be the same. In other words, in some embodiments, step 102 and step 104 can be the same step, and the voice enhancement feature extraction unit and the voice comparison feature extraction unit can also be the same unit.
[0041] In some embodiments, the voice comparison features can be represented as a vector of length N. This feature vector can be compared with the vector of the reference voice features pre-stored in the database in the subsequent step 105, where these two types of vectors can have the same or similar formats.
[0042] Specifically, in step 105, the voice feature comparison unit compares the voice comparison feature vector extracted in step 104 with one or more pre-stored reference voice feature vectors. Further, based on the comparison result of these two types of vectors, it can be determined whether the person represented by the voice comparison feature vector is a known person in the database.
[0043] In one embodiment, a similarity calculation algorithm such as the cosine distance algorithm can be used to compare the voice comparison feature vector and the reference voice feature vector. That is, by means of retrieval, the reference voice feature vector stored in a predetermined database with the shortest distance (i.e., the highest similarity) to the extracted voice comparison feature vector is matched. Specifically, the cosine distance algorithm first calculates the cosine value (cosine similarity) between the voice comparison feature vector to be compared and the reference voice feature vector, which is represented by equation (1) for example.
[0044]
[0045] Wherein, cos(θ) represents the cosine value of the two feature vectors, A represents the voice comparison feature vector to be compared, B represents the reference voice feature vector, and n is the dimension of the two feature vectors (a natural number). Then, the cosine distance is obtained through 1 - cos(θ).
[0046] It can be understood that in practical applications, multiple reference voice feature vectors can be respectively compared with the voice comparison feature vector, and the reference voice feature vector with the closest distance is used as the matching vector of the voice comparison feature vector, and at the same time, this minimum distance is set as d-cos. Although the above takes the cosine distance as an example to represent the method of calculating the similarity between two vectors, those skilled in the art can also use other suitable similarity calculation methods, such as the Euclidean Distance, etc., to calculate the similarity of the feature vectors, and the present application does not limit this.
[0047] Further, the determined minimum distance d-cos after comparison can be compared with a preset distance threshold t-cos, where the preset distance threshold t-cos can be a value determined according to experience or historical data and is used to determine whether the voice comparison feature and the reference voice feature are from the same person. In one embodiment, when the minimum distance d-cos is less than or equal to the distance threshold t-cos, it indicates that the similarity between the voice comparison feature and a certain reference voice feature is relatively high, and it can be determined that there is a reference voice feature in the database that matches the voice signal in the audio input signal, that is, the emitter of the voice signal in the currently received audio input signal has been recorded in the existing database; on the contrary, when the minimum distance d-cos is greater than the distance threshold t-cos, it indicates that the similarity between the voice comparison feature and all reference voice features is not high enough, and it is determined that there is no reference voice feature in the preset database that matches the voice signal in the audio input signal, that is, the emitter of the voice signal in the currently received audio input signal has not been recorded in the existing database.
[0048] Continue to refer to Figure 1 , in step 105, the voice quality evaluation unit evaluates the voice quality (e.g., voice signal-to-noise ratio p-SNR) determined in step 103 and generates a quality evaluation result. As described before, the voice signal-to-noise ratio p-SNR of the audio input signal can be compared with a preset voice signal-to-noise ratio threshold t-SNR to evaluate the voice quality. Then, the quality evaluation result and the feature comparison result generated by the voice feature comparison unit can be provided to the enhanced feature selection unit for subsequent operations. According to the feature comparison result and the voice quality evaluation result, step 105 generally includes 4 different situations and processing methods.
[0049] In the first case, if the voice signal-to-noise ratio p-SNR of the audio input signal determined in step 103 is greater than the preset voice signal-to-noise ratio threshold t-SNR and the minimum distance d-cos determined in step 105 is greater than the minimum distance threshold t-cos (i.e., p-SNR>t-SNR and d-cos>t-cos), then it can be considered that the current audio input signal includes a strong enough human voice, its voice quality can meet the preset voice quality requirements, and there are not enough approximate reference voice features in the database of the voice enhancement device. This indicates that the currently input audio input signal may contain a new human voice. In this case, a new reference voice feature can be created in the database for subsequent comparison and enhancement of the input audio input signal.
[0050] In some embodiments, the storage format of the reference speech features stored in the speech enhancement device in its database may be: [timestamp; ID; p-SNR; speech comparison feature vector; speech enhancement feature vector]. Among them, the timestamp represents the storage time of this reference speech feature in the database; ID represents the number of this reference speech feature; p-SNR represents the speech signal-to-noise ratio of this reference speech feature; the speech comparison feature vector represents the speech feature vector for comparison of this reference speech feature; the speech enhancement feature vector represents the speech feature vector for speech enhancement of this reference speech feature. In some examples, the storage length of each reference speech feature may be the length of the speech comparison feature vector + the length of the speech enhancement feature vector + 3 (the unit is byte, and the extra 3 bytes can be used to store the data of the timestamp, ID, and p-SNR). It can be understood that the reference speech features can also adopt other storage formats. For example, when the speech comparison feature vector and the speech enhancement feature vector are the same vector, it may only include the information of the timestamp, ID, p-SNR, and the speech comparison (enhancement) feature vector.
[0051] Continue to refer to the appendix Figure 1 In the first case, in step 106, the enhancement feature selection unit may select the speech enhancement features extracted in step 102 and input them to the speech enhancement algorithm unit, and use the speech enhancement features to perform speech enhancement on the audio input signal. That is to say, since there is no corresponding reference speech feature for the audio input signal in the database before, and its own speech quality also meets the requirements, the speech enhancement features extracted in step 102 can be used as the reference speech features required for speech enhancement. At the same time, the speech enhancement features can also be stored in the database for use as reference speech features during subsequent processing. In one embodiment, the speech enhancement algorithm used by the speech enhancement algorithm unit may adopt the form of a neural network, such as a convolutional neural network (CNN), a recurrent neural network (RNN), a neural network combining convolution and recurrence (CRN), etc. Those skilled in the art can understand that various feature-based speech enhancement algorithms can be used, and this application does not limit this.
[0052] In the second case, if the speech signal-to-noise ratio p-SNR of the audio input signal determined in step 103 is greater than the preset speech signal-to-noise ratio threshold t-SNR and the minimum distance d-cos determined in step 105 is less than the minimum distance threshold t-cos (i.e., p-SNR > t-SNR, d-cos < t-cos), then it can be considered that the current audio input signal includes strong enough human voices, and the existing database also includes reference speech features with a high enough similarity to the audio input signal. Therefore, it is possible to use the reference speech features stored in the database for speech enhancement.
[0053] In this case, the p-SNR of the audio input signal can be further compared with the p-SNR of the matching reference speech feature in the database. If the p-SNR of the audio input signal is less than the p-SNR of the matching reference speech feature in the database, then it can be considered that the speech quality of the current speech input signal is inferior to the speech quality of the matching reference speech feature when it was stored. Accordingly, in some embodiments, the speech enhancement feature in the matching reference speech feature can be read in step 106, and combined with the speech enhancement feature of the audio input signal extracted in step 102, and the two speech enhancement features can be used together to enhance the currently processed speech input signal; in other embodiments, the currently processed speech input signal can also be enhanced only with the speech enhancement feature in the matching reference speech feature read in step 106 from the database. On the contrary, if the p-SNR of the audio input signal is greater than the p-SNR of the matching reference speech feature in the database, then it can be considered that the speech quality of the current speech input signal is better than the speech quality of the matching reference speech feature when it was stored. Accordingly, in step 105, the speech feature of the current audio input signal can be used to update the matching reference speech feature in the database for subsequent speech feature matching and enhancement. In some embodiments, additionally or alternatively, the duration of the current audio input signal can also be compared with a predetermined duration threshold, and the reference speech feature can be updated according to the duration comparison result and / or the quality assessment result. And, in step 106, the enhancement feature selection unit can directly use the speech enhancement feature of the audio input signal extracted in step 102 to enhance the current speech input signal. After that, in step 107, the enhanced audio output signal can be output, for example, played out through a microphone.
[0054] In the third case, if it is determined in step 103 that the speech signal-to-noise ratio p-SNR of the audio input signal is less than a preset speech signal-to-noise ratio threshold t-SNR and it is determined in step 105 that the minimum distance d-cos is greater than the minimum distance threshold t-cos (i.e., p-SNR < t-SNR, d-cos > t-cos), then it can be considered that the current audio input signal does not include a strong enough human voice, and the existing database does not include a sufficiently approximate reference speech feature. In this case, in step 106, only the speech enhancement feature of the audio input signal extracted in step 102 is used to enhance the current speech input signal. After that, in step 107, the enhanced audio output signal can be output.
[0055] In the fourth case, if it is determined in step 103 that the voice signal-to-noise ratio p-SNR of the audio input signal is less than a preset voice signal-to-noise ratio threshold t-SNR and it is determined in step 105 that the minimum distance d-cos is less than the minimum distance threshold t-cos (i.e., p-SNR < t-SNR, d-cos < t-cos), then it can be considered that the current audio input signal does not include a strong enough human voice, but the existing database includes sufficiently approximate reference voice features. In this case, in step 106, the voice enhancement features of the reference voice features matched in the database can be read, and combined with the voice enhancement features in the audio input signal extracted in step 102, and the current processed voice input signal can be enhanced jointly by these two voice enhancement features; in some other embodiments, the current voice input signal can also be enhanced only by the voice enhancement features of the audio input signal extracted in step 102. After that, in step 107, the enhanced audio output signal can be output.
[0056] It can be seen that based on the above method, when it is necessary to enhance the voice in the voice input signal, the voice part of the current voice input signal can be enhanced according to the reference voice features matched in the existing database, and these matched reference voice features are often those obtained by collection in a relatively quiet environment. Therefore, the method of the present application can effectively improve the voice enhancement effect.
[0057] In addition, during the actual use process, the reference feature data in the database can also be continuously updated as the usage time increases, and the feature data collected in an ideal environment can be stored in the database, so that the database can often store the feature data with relatively high voice quality. This also further improves the subsequent voice enhancement effect.
[0058] Figure 2 FIG. 200 shows a flowchart of a voice enhancement method according to an embodiment of the present application. It can be understood that one or more steps in flowchart 200 can be implemented in a manner similar to the same or similar steps shown in Figure 1 method 100, and can be executed by a processing device. Among them, the processing device can be an electronic device with voice signal processing capabilities, such as a hearing aid or a headset with a processor.
[0059] As Figure 2As shown, the method starts from step 201, where the processing device can receive a current audio input signal having a speech part and a non-speech part. Then, at step 202, the processing device can determine the speech features of the speech part in the current audio input signal, and at step 203, the processing device can determine the speech quality of the current audio input signal. Thus, at step 204, the processing device can evaluate whether the speech quality meets a predetermined speech quality requirement; then, at step 205, in response to the evaluation result in step 204, that is, the speech quality meets the predetermined speech quality requirement, the processing device can create or update a reference speech feature with the speech features, where the reference speech feature is used to enhance the speech part in the audio input signal.
[0060] It can be seen that in the above manner, the reference speech features stored in the speech feature database can be created or updated, so that as the actual usage time increases, the reference speech features with better quality are retained.
[0061] In some embodiments, after step 205, the speech enhancement method 200 further includes a step of using the reference speech feature to perform speech enhancement processing on the currently input audio input signal. For example, the reference speech feature matching the speech features of the speech part in the current audio input signal can be retrieved from one or more pre-stored reference speech features, so that one or both of the speech features of the speech part in the current audio input signal and the matching reference speech feature can be used to enhance the speech part in the current audio input signal. In particular, when the speech quality of the current audio input signal is not better than the speech quality corresponding to the matching reference speech feature, the matching reference speech feature can be used to enhance the speech part in the current audio input signal; or when the speech quality of the current audio input signal is better than the speech quality corresponding to the matching reference speech feature, while updating the matching reference speech feature in the database with the speech features of the speech part in the current audio input signal, the updated reference speech feature can be used to perform speech enhancement on the current audio input signal.
[0062] In some embodiments, the present application also provides some computer program products, which include a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium includes computer-executable code for executing Figure 1 or Figure 2 the steps in the method embodiments shown. In some embodiments, the computer program product can be stored in a hardware device, such as an audio device.
[0063] Embodiments of the present invention can be implemented through hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable logic devices such as field programmable gate arrays, or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0064] It should be noted that although several steps or modules of the voice enhancement method, device, and storage medium are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0065] Those of ordinary skill in the art can understand and implement other changes to the disclosed embodiments by studying the specification, the disclosed content, the drawings, and the appended claims. In the claims, the term "comprising" does not exclude other elements and steps, and the terms "a" and "one" do not exclude a plurality. In the actual application of the present application, a component may perform the functions of multiple technical features recited in the claims. Any reference numerals in the claims should not be construed as limiting the scope.
Claims
1. A speech enhancement method, characterized in that: The method comprises: receiving a current audio input signal having a speech portion and a non-speech portion; Determining speech features of a speech portion in the current audio input signal; Determining the voice quality of the current audio input signal; evaluating whether the voice quality meets a predetermined voice quality requirement; Retrieving a reference speech feature that matches the speech feature from one or more pre-stored reference speech features; and In response to the evaluation result of the predetermined speech quality requirement and the matching result of the one or more reference speech features, enhancing the speech portion in the current audio input signal using one or both of the speech features of the speech portion in the current audio input signal and the matched reference speech features; The matching result of the one or more reference voice features indicates whether the person represented by the voice feature is one of the persons represented by the one or more reference voice features.
2. The method according to claim 1, in response to the speech quality meeting the predetermined speech quality requirement, creating or updating a reference speech feature using the speech feature.
3. The method according to claim 1, characterized in that Wherein determining the voice quality of the current audio input signal comprises: A speech signal-to-noise ratio of the current audio input signal is determined, where the speech signal-to-noise ratio represents a ratio of power of the speech portion to power of the non-speech portion.
4. The method according to claim 3, characterized in that The evaluating whether the voice quality meets the predetermined voice quality requirement includes: comparing the speech signal-to-noise ratio with a predetermined speech signal-to-noise ratio threshold; and In response to the speech signal-to-noise ratio being greater than the predetermined speech signal-to-noise ratio threshold, it is determined that the speech quality meets a predetermined speech quality requirement.
5. The method according to claim 2, characterized in that: The method further comprises: In response to not retrieving a reference speech feature that matches the speech feature, creating a new reference speech feature using the speech feature of the current audio input signal; and The speech part in the current audio input signal is enhanced by using the speech feature of the speech part in the current audio input signal.
6. The method according to claim 5, characterized in that The method further comprises: Comparing the duration of the current audio input signal with a predetermined duration threshold; In response to the duration of the current audio input signal being greater than the predetermined duration threshold, a reference speech feature is created using the speech feature of the current audio input signal.
7. The method according to claim 1, characterized in that The method further comprises: In response to retrieving a reference speech feature that matches the speech feature, comparing the speech quality of the current audio input signal with the speech quality corresponding to the matched reference speech feature; In response to the voice quality of the current audio input signal being better than the voice quality corresponding to the matched reference voice feature, updating the matched reference voice feature using the voice feature of the current audio input signal; and The speech part in the current audio input signal is enhanced by using the speech feature of the speech part in the current audio input signal.
8. The method according to claim 7, characterized in that The method further comprises: In response to the speech quality of the current audio input signal being not better than the speech quality corresponding to the matched reference speech feature, the speech feature of the speech portion in the current audio input signal and the matched reference speech feature are used to enhance the speech portion in the current audio input signal.
9. The method according to claim 1, characterized in that: The method further comprises: In response to not retrieving a reference speech feature matching the speech feature and the speech quality not meeting the predetermined speech quality requirement, enhancing the speech portion in the current audio input signal using the speech feature of the speech portion in the current audio input signal.
10. The method according to claim 1, characterized in that The method further comprises: In response to retrieving a reference speech feature that matches the speech feature and the speech quality not meeting the predetermined speech quality requirement, the speech feature is enhanced using the speech feature of the speech portion in the current audio input signal and the matched reference speech feature.
11. The method according to claim 1, characterized in that: The speech features include pitch period or Mel-cepstral coefficients.
12. The method according to claim 1, characterized in that The step of determining the speech feature of the speech part in the current audio input signal comprises: Determining a speech enhancement feature and a speech comparison feature of a speech part in the current audio input signal; And wherein the reference speech feature includes a reference speech enhancement feature and a reference speech comparison feature, the speech enhancement feature and the reference speech enhancement feature are used to enhance the speech part of the audio input signal, and the speech comparison feature is used to match the reference speech comparison feature.
13. A speech enhancement device, characterized in that: The apparatus includes a non-transitory computer storage medium having one or more executable instructions stored thereon, wherein the one or more executable instructions are executed by a processor to perform the following steps: receiving a current audio input signal having a speech portion and a non-speech portion; Determining speech features of a speech portion in the current audio input signal; Determining the voice quality of the current audio input signal; evaluating whether the voice quality meets a predetermined voice quality requirement; Retrieving a reference speech feature that matches the speech feature from one or more pre-stored reference speech features; as well as In response to the evaluation result of the predetermined speech quality requirement and the matching result of the one or more reference speech features, enhancing the speech portion in the current audio input signal using one or both of the speech features of the speech portion in the current audio input signal and the matched reference speech features; The matching result of the one or more reference voice features indicates whether the person represented by the voice feature is one of the persons represented by the one or more reference voice features.
14. A non-transitory computer storage medium having one or more executable instructions stored thereon, wherein the one or more executable instructions are executed by a processor to perform the following steps: receiving a current audio input signal having a speech portion and a non-speech portion; Determining speech features of a speech portion in the current audio input signal; Determining the voice quality of the current audio input signal; evaluating whether the voice quality meets a predetermined voice quality requirement; Retrieving a reference speech feature that matches the speech feature from one or more pre-stored reference speech features; as well as In response to the evaluation result of the predetermined speech quality requirement and the matching result of the one or more reference speech features, enhancing the speech portion in the current audio input signal using one or both of the speech features of the speech portion in the current audio input signal and the matched reference speech features; The matching result of the one or more reference voice features indicates whether the person represented by the voice feature is one of the persons represented by the one or more reference voice features.
Citation Information
Patent Citations
Method, device and terminal equipment for processing voice
CN103971696A
Voice enhancement method and device, electronic equipment and storage medium
CN112201247A