Voice enhancement method, electronic device, storage medium and chip system

By collecting users' facial and lip movement images and combining them with speech features to train a speech enhancement network in real time, the problem of inaccurate user voice recognition and enhancement in audio and video calls by terminal devices has been solved, achieving higher accuracy and quality.

CN116072136BActive Publication Date: 2026-02-06HUAWEI DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111279132.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-31
Publication Date
2026-02-06
Estimated Expiration
2041-10-31

AI Technical Summary

Technical Problem

Existing terminal devices are not accurate enough in processing audio data during audio and video calls, especially in recognizing and enhancing user voices due to errors and noise interference.

Method used

By collecting users' facial images and lip movement sequence images, and combining them with speech features, a speech enhancement network is trained in real time to improve the accuracy of user voice recognition and enhancement. This includes training the network in a quiet environment and adjusting it to reduce false cancellation.

Benefits of technology

It improves the accuracy of user voice recognition and enhancement in audio data, reduces false muting, and enhances the quality of audio and video calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116072136B_ABST
    Figure CN116072136B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of audio technology, and provides a speech enhancement method, an electronic device, a storage medium and a chip system. The method comprises the following steps: collecting a first face image of a first user; if the first face image does not match stored face data, obtaining a voice feature of the first user; storing first face data of the first user; collecting a second face image and first audio data of the first user; if the second face image matches the stored first face data, enhancing the voice of the first user in the first audio data based on the voice feature of the first user and then outputting the voice. The second face image is collected, so that the accuracy of identifying the first user can be improved. In combination with the voice feature of the first user, the accuracy of enhancing the voice of the first user in the first audio data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of audio technology, and in particular to a voice enhancement method, an electronic device, a storage medium and a chip system. BACKGROUND

[0002] With the continuous development of terminal devices, the functions of terminal devices are increasing, and the scenarios in which terminal devices need to use audio and video call functions are increasing.

[0003] When the terminal device is in the scenario of audio and video call, the terminal device can collect the sound emitted by the user in the current scenario to obtain audio data. However, the current scenario can include various sounds, and the audio data collected by the terminal device includes not only the sound emitted by the user, but also other sounds in the current scenario, that is, noise. In order to improve the quality of audio and video call, the terminal device can process the audio data, for example, denoising, or enhancing specific sounds (such as the sound emitted by the user). However, the effect of the current terminal device in processing the audio data is not accurate enough and needs to be improved. SUMMARY

[0004] The present application provides a voice enhancement method, an electronic device, a storage medium and a chip system, which solves the problem that the effect of voice enhancement on user sound is not accurate enough in the process of optimizing audio data by the terminal device in the prior art.

[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0006] In a first aspect, a voice enhancement method is provided, and the method comprises:

[0007] Collecting a first face image of a first user;

[0008] If the first face image does not match the stored face data, obtaining a voice feature of the first user;

[0009] Storing first face data of the first user;

[0010] Collecting a second face image and first audio data of the first user;

[0011] If the second face image matches the stored first face data, enhancing the voice of the first user in the first audio data based on the voice feature of the first user and outputting.

[0012] The terminal device collects the second face image and the first audio data, can determine the first user who is speaking in the current scene, and combines the voice feature of the first user obtained in the terminal device, so that the terminal device can enhance the voice of the first user in the first audio data. By collecting the second face image, the accuracy of identifying the first user can be improved, and by combining the voice feature of the first user, the accuracy of enhancing the voice of the first user in the first audio data can be improved.

[0013] In a first possible implementation manner of the first aspect, the method further includes:

[0014] collecting second audio data of the first user;

[0015] if the first face image does not match the stored face data, outputting third audio data according to the second audio data, wherein the voice of the first user in the third audio data is not enhanced.

[0016] In the case that the terminal device determines that the voice enhancement network does not learn the voice feature of the first user, the terminal device does not further denoise the collected second audio data through the voice enhancement network, thereby avoiding the case that the voice enhancement network causes false de-noising of the second audio data, and improving the reliability of the terminal device.

[0017] In any of the possible implementation manners of the first aspect, in a second possible implementation manner of the first aspect, if the second face image matches the stored first face data, the first audio data is output after the voice of the first user in the first audio data is enhanced based on the voice feature of the first user, and the output includes:

[0018] detecting whether the lips of the first user move;

[0019] if the second face image matches the stored first face data and the lips of the first user move, the first audio data is output after the voice of the first user in the first audio data is enhanced based on the voice feature of the first user.

[0020] Before enhancing the first audio data, the terminal device can first determine whether the lips of the user move. If the lips of the user do not move, it indicates that the user is not currently speaking, and the voice of the first user in the first audio data does not need to be enhanced. If the lips of the user move, it indicates that the user is currently speaking, and the terminal device can determine the speech of the user through lip-reading technology according to the dynamic change process of the lips of the user, and further improve the accuracy of enhancing the voice of the user in the second audio data by combining the obtained voice feature of the first user.

[0021] In a third possible implementation of the first aspect, in the second possible implementation of the first aspect, the detecting whether the lips of the first user move comprises:

[0022] obtaining a first lip movement sequence image of the first user;

[0023] detecting whether the lips in the first lip movement sequence image move;

[0024] the outputting, based on the voice feature of the first user, the enhanced voice of the first user in the first audio data comprises:

[0025] outputting, based on the first lip movement sequence image, the second face image and the first audio data, the enhanced voice of the first user in the first audio data by the voice enhancement network.

[0026] By detecting whether the lips of the first user move, the enhanced voice of the first user in the first audio data can be obtained when the lips of the user move, i.e., when the user is speaking, in combination with the first lip movement sequence image, so that the accuracy of enhancing the voice of the first user can be improved.

[0027] In a fourth possible implementation of the first aspect, after the outputting, based on the voice feature of the first user, the enhanced voice of the first user in the first audio data, the method further comprises:

[0028] obtaining the enhanced first audio data;

[0029] detecting whether the enhanced first audio data has a muting phenomenon;

[0030] if the enhanced first audio data has the muting phenomenon, obtaining the voice feature of the first user again;

[0031] if the enhanced first audio data does not have the muting phenomenon, continuing to output, based on the voice feature of the first user, the enhanced voice of the first user in the audio data collected again.

[0032] By detecting whether the enhanced first audio data has the muting phenomenon, the reliability of the terminal device in enhancing the voice of the first user can be improved.

[0033] In a fifth possible implementation of the first aspect, the obtaining the voice feature of the first user comprises:

[0034] collect fourth audio data of the first user and a first sequence image, the first sequence image comprising: face information and lip information;

[0035] obtain a voice feature of the first user based on the first sequence image and the fourth audio data.

[0036] In a sixth possible implementation of the first aspect, the obtaining the voice feature of the first user comprises:

[0037] learning the voice feature of the first user through a speech enhancement network.

[0038] In a seventh possible implementation of the first aspect, the learning the voice feature of the first user through the speech enhancement network comprises:

[0039] obtaining a third face image and a second lip movement sequence image according to the first sequence image;

[0040] inputting the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learning the voice feature of the first user through the speech enhancement network.

[0041] In an eighth possible implementation of the first aspect, before the inputting the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learning the voice feature of the first user through the speech enhancement network, the method further comprises:

[0042] determining whether the lips of the first user move according to the second lip movement sequence image;

[0043] determining whether the current scene is a quiet environment according to the fourth audio data;

[0044] the inputting the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learning the voice feature of the first user through the speech enhancement network comprises:

[0045] if the current scene is a quiet environment and the lips of the first user move, inputting the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learning the voice feature of the first user through the speech enhancement network.

[0046] By determining whether the second lip movement sequence image and the fourth audio data collected by the terminal device satisfy the voice learning condition of the first user, and in a case where the voice learning condition of the first user is satisfied, learning the voice of the first user according to the second face sequence image and the fourth audio data, the accuracy of the voice feature of the first user learned can be improved.

[0047] In a ninth possible implementation of the first aspect, according to the eighth possible implementation of the first aspect, the determining whether the current scene is a quiet environment according to the fourth audio data comprises:

[0048] inputting the fourth audio data into the speech enhancement network to obtain first denoising data;

[0049] comparing the fourth audio data and the first denoising data;

[0050] if the similarity between the fourth audio data and the first denoising data is greater than or equal to a similarity threshold, determining that the current scene is a quiet environment;

[0051] if the similarity between the fourth audio data and the first denoising data is less than the similarity threshold, determining that the current scene is not a quiet environment.

[0052] By collecting the fourth audio data in a quiet environment, the collected fourth audio data can be used as label data for training the speech enhancement network, simplifying the process of training the speech enhancement network, thereby improving the efficiency of training the speech enhancement network and improving the timeliness of the terminal device calling the speech enhancement network.

[0053] In a tenth possible implementation of the first aspect, according to any one of the seventh to ninth possible implementations of the first aspect, the inputting the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learning the voice feature of the first user by the speech enhancement network comprises:

[0054] mixing the fourth audio data with pre-stored noise data to obtain mixed audio data;

[0055] inputting the mixed audio data, the third face image and the second lip movement sequence image into the speech enhancement network to obtain second denoising data;

[0056] adjusting the speech enhancement network according to the second denoising data and the fourth audio data, so that the speech enhancement network learns the voice feature of the first user.

[0057] The different second de-noising data are obtained through continuous training, and the speech enhancement network is adjusted according to the comparison result of the continuously obtained second de-noising data and the collected first audio data, so that the voice feature of the first user is obtained, which can further improve the accuracy of the voice feature of the first user.

[0058] In a second aspect, another speech enhancement method is provided, which includes:

[0059] Collecting first audio data of a first user and first sequence images, the first sequence images including face information and lip information of the first user;

[0060] If it is determined according to the first sequence images that the face information of the first user matches stored face data and the lips of the first user move, then the voice of the first user in the first audio data is enhanced according to the voice feature of the first user and then outputted.

[0061] In a first possible implementation manner of the second aspect, the method further includes:

[0062] If it is determined according to the first sequence images that the face information of the first user does not match stored face data, or if it is determined according to the first sequence images that the lips of the first user do not move, then second audio data is outputted according to the first audio data, and the voice of the first user in the second audio data is not enhanced.

[0063] Based on any one of the possible implementation manners of the second aspect, in a second possible implementation manner of the second aspect, the voice of the first user in the first audio data is enhanced according to the voice feature of the first user and then outputted, which includes:

[0064] According to the first sequence images, a first face image and a first lip movement sequence image are extracted;

[0065] According to the first audio data, the first face image and the first lip movement sequence image, the voice of the first user in the first audio data is enhanced through the speech enhancement network and then outputted.

[0066] In a third possible implementation manner of the second aspect, before the voice of the first user in the first audio data is enhanced according to the voice feature of the first user and then outputted, the method further includes:

[0067] The first sequence images are extracted to obtain a first lip movement sequence image;

[0068] According to the first lip movement sequence image, it is determined whether the lips of the first user move.

[0069] In a fourth possible implementation manner of the second aspect, before the outputting the first audio data after enhancing the voice of the first user in the first audio data according to the voice feature of the first user, the method further includes:

[0070] extracting the first sequence of images to obtain first face information;

[0071] determining whether each of the stored face data includes face data matching the first face information.

[0072] In a fifth possible implementation manner of the second aspect, after the outputting the first audio data after enhancing the voice of the first user in the first audio data according to the voice feature of the first user, the method further includes:

[0073] obtaining the enhanced first audio data;

[0074] detecting whether the enhanced first audio data has a muting phenomenon;

[0075] if the enhanced first audio data has the muting phenomenon, obtaining the voice feature of the first user again;

[0076] if the enhanced first audio data does not have the muting phenomenon, continuing to output the first audio data after enhancing the voice of the first user in the audio data collected again based on the voice feature of the first user.

[0077] In a sixth possible implementation manner of the second aspect, the collecting the first audio data and the first sequence of images of the first user includes:

[0078] in response to an operation for calling a voice enhancement network, collecting the first audio data and the first sequence of images of the first user.

[0079] In a third aspect, a voice enhancement device is provided, and the device includes:

[0080] a collecting module configured to collect a first face image of a first user;

[0081] a first obtaining module configured to, if the first face image does not match stored face data, obtain a voice feature of the first user;

[0082] a storage module configured to store first face data of the first user;

[0083] The collection module is further configured to collect a second facial image and first audio data of the first user.

[0084] The output module is configured to, if the second facial image matches the stored first facial data, output the voice of the first user in the first audio data after enhancing the voice based on a voice feature of the first user.

[0085] In a first possible implementation manner of the third aspect, the collection module is further configured to collect second audio data of the first user.

[0086] The output module is further configured to, if the first facial image does not match the stored facial data, output third audio data according to the second audio data, wherein the voice of the first user in the third audio data is not enhanced.

[0087] In the second possible implementation manner of the third aspect, based on any one of the foregoing possible implementation manners of the third aspect, the output module is specifically configured to detect whether the lips of the first user move; if the second facial image matches the stored first facial data and the lips of the first user move, output the voice of the first user in the first audio data after enhancing the voice based on a voice feature of the first user.

[0088] In the third possible implementation manner of the third aspect, based on the second possible implementation manner of the third aspect, the output module is specifically configured to acquire a first lip movement sequence image of the first user; detect whether the lips in the first lip movement sequence image move; and the outputting the voice of the first user in the first audio data after enhancing the voice based on a voice feature of the first user comprises: acquiring a first lip movement sequence image of the first user; detecting whether the lips in the first lip movement sequence image move; and outputting the voice of the first user in the first audio data after enhancing the voice based on a voice feature of the first user according to the first lip movement sequence image, the second facial image and the first audio data through the voice enhancement network.

[0089] In the fourth possible implementation manner of the third aspect, based on any one of the foregoing possible implementation manners of the third aspect, the device further comprises:

[0090] The second acquisition module is configured to acquire the enhanced first audio data.

[0091] The detection module is configured to detect whether the enhanced first audio data has a muting phenomenon.

[0092] The first acquisition module is further configured to, if the enhanced first audio data has the muting phenomenon, acquire the voice feature of the first user again.

[0093] The output module is further configured to, if the enhanced first audio data does not have the muting phenomenon, continue to enhance and output the first user's voice in the re-collected audio data based on the voice feature of the first user.

[0094] In a fifth possible implementation of the third aspect, the first obtaining module is specifically configured to collect fourth audio data of the first user and a first sequence image, the first sequence image including face information and lip information; and obtain the voice feature of the first user based on the first sequence image and the fourth audio data.

[0095] Based on the fifth possible implementation of the third aspect, in a sixth possible implementation of the third aspect, the first obtaining module is further specifically configured to learn the voice feature of the first user through a speech enhancement network.

[0096] Based on the sixth possible implementation of the third aspect, in a seventh possible implementation of the third aspect, the first obtaining module is further specifically configured to obtain a third face image and a second lip movement sequence image according to the first sequence image; input the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learn the voice feature of the first user through the speech enhancement network.

[0097] Based on the seventh possible implementation of the third aspect, in an eighth possible implementation of the third aspect, the device further includes:

[0098] A first judging module configured to determine whether the first user's lips move according to the second lip movement sequence image;

[0099] A second judging module configured to determine whether the current scene is a quiet environment according to the fourth audio data;

[0100] The first obtaining module is further specifically configured to, if the current scene is a quiet environment and the first user's lips move, input the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learn the voice feature of the first user through the speech enhancement network.

[0101] In a ninth possible implementation of the third aspect, based on the eighth possible implementation of the third aspect, the second determining module is specifically configured to input the fourth audio data into the speech enhancement network to obtain first de-noised data; compare the fourth audio data and the first de-noised data; if a similarity between the fourth audio data and the first de-noised data is greater than or equal to a similarity threshold, determine that the current scene is a quiet environment; and if the similarity between the fourth audio data and the first de-noised data is less than the similarity threshold, determine that the current scene is not a quiet environment.

[0102] In a tenth possible implementation of the third aspect, based on any one of the seventh to ninth possible implementations of the third aspect, the first obtaining module is further configured to mix the fourth audio data with pre-stored noise data to obtain mixed audio data; input the mixed audio data, the third face image, and the second lip movement sequence image into the speech enhancement network to obtain second de-noised data; and adjust the speech enhancement network according to the second de-noised data and the fourth audio data, so that the speech enhancement network learns the voice feature of the first user.

[0103] In a fourth aspect, another speech enhancement apparatus is provided, and the apparatus includes:

[0104] An acquisition module is configured to acquire first audio data and first sequence images of a first user, the first sequence images including face information and lip information of the first user.

[0105] An output module is configured to, if it is determined according to the first sequence images that the face information of the first user matches stored face data and the lips of the first user move, output, according to a voice feature of the first user, the voice of the first user in the first audio data after the voice of the first user is enhanced.

[0106] In a first possible implementation of the fourth aspect, the output module is further configured to, if it is determined according to the first sequence images that the face information of the first user does not match the stored face data, or if it is determined according to the first sequence images that the lips of the first user do not move, output second audio data according to the first audio data, the voice of the first user in the second audio data not being enhanced.

[0107] In a second possible implementation form of the fourth aspect, based on any of the preceding possible implementation forms of the fourth aspect, the output module is configured to: extract a first face image and a first lip movement sequence image according to the first sequence images; and output the first user's voice in the first audio data after enhancing the first user's voice by the speech enhancement network according to the first audio data, the first face image, and the first lip movement sequence image.

[0108] In a third possible implementation form of the fourth aspect, the apparatus further includes:

[0109] a first extraction module configured to extract the first sequence images to obtain a first lip movement sequence image;

[0110] a first judgment module configured to judge whether the first user's lips move according to the first lip movement sequence image.

[0111] In a fourth possible implementation form of the fourth aspect, the apparatus further includes:

[0112] a second extraction module configured to extract the first sequence images to obtain first face information;

[0113] a second judgment module configured to traverse each of the stored face data to determine whether the stored face data includes face data matching the first face information.

[0114] In a fifth possible implementation form of the fourth aspect, based on any of the preceding possible implementation forms of the fourth aspect, the apparatus further includes:

[0115] a first acquisition module configured to acquire the enhanced first audio data;

[0116] a detection module configured to detect whether the enhanced first audio data has a mute phenomenon;

[0117] a second acquisition module configured to, if the enhanced first audio data has the mute phenomenon, acquire the first user's voice feature again;

[0118] the output module is further configured to, if the enhanced first audio data does not have the mute phenomenon, continue to output the first user's voice in the re-collected audio data after enhancing the first user's voice based on the first user's voice feature.

[0119] In a sixth possible implementation form of the fourth aspect, based on any of the preceding possible implementation forms of the fourth aspect, the collection module is configured to collect the first audio data and the first sequence images of the first user in response to an operation for calling a speech enhancement network.

[0120] In a fifth aspect, a voice enhancement method is provided, which is applied in a scenario of voice or video call composed of a first terminal device and a second terminal device, and the method comprises:

[0121] When the first terminal device detects an operation of triggering initiation of voice or video call to the second terminal device, or when the first terminal device receives a request of voice or video call initiated by the second terminal device, a first face image of a first user is collected;

[0122] If the first face image does not match stored face data, the first terminal device acquires a voice feature of the first user;

[0123] The first terminal device stores first face data of the first user;

[0124] The first terminal device collects a second face image and first audio data of the first user;

[0125] If the second face image matches the stored first face data, the first terminal device outputs, after enhancing the voice of the first user in the first audio data based on the voice feature of the first user, to the second terminal device.

[0126] In a first possible implementation manner of the fifth aspect, the method further comprises:

[0127] The first terminal device collects second audio data of the first user;

[0128] If the first face image does not match the stored face data, the first terminal device outputs third audio data according to the second audio data, and the voice of the first user in the third audio data is not enhanced.

[0129] In a sixth aspect, an electronic device is provided, comprising a processor configured to execute a computer program stored in a memory to implement the voice enhancement method according to any one of the first aspect or the second aspect.

[0130] In a seventh aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the voice enhancement method according to any one of the first aspect or the second aspect.

[0131] In an eighth aspect, a chip system is provided, which comprises a memory and a processor, and the processor executes a computer program stored in the memory to implement the voice enhancement method according to any one of the first aspect or the second aspect.

[0132] It can be understood that the beneficial effects of the above-mentioned second aspect to the eighth aspect can be referred to the relevant description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0133] Figure 1A A terminal device provided by an embodiment of the present application adopts a multi-modal speech enhancement technology to enhance audio data, and a framework schematic diagram of the terminal device is provided;

[0134] Figure 1A Another terminal device provided by an embodiment of the present application adopts a multi-modal speech enhancement technology to enhance audio data, and a framework schematic diagram of the terminal device is provided;

[0135] Figure 2 A scene schematic diagram of a speech enhancement scene provided by an embodiment of the present application is provided;

[0136] Figure 3 A flow schematic diagram of a terminal device reminding a user to input audio data and a face sequence image provided by an embodiment of the present application is provided;

[0137] Figure 4 An illustrative flowchart of a speech enhancement method provided by an embodiment of the present application is provided;

[0138] Figure 5 A framework schematic diagram of a terminal device online training a speech enhancement network provided by an embodiment of the present application is provided;

[0139] Figure 6 A framework schematic diagram of a speech enhancement network performing speech enhancement provided by an embodiment of the present application is provided;

[0140] Figure 7 A structural block diagram of a speech enhancement device provided by an embodiment of the present application is provided;

[0141] Figure 8 A structural block diagram of another speech enhancement device provided by an embodiment of the present application is provided;

[0142] Figure 9 A structural block diagram of still another speech enhancement device provided by an embodiment of the present application is provided;

[0143] Figure 10 A structural block diagram of still another speech enhancement device provided by an embodiment of the present application is provided;

[0144] Figure 11 A structural block diagram of still another speech enhancement device provided by an embodiment of the present application is provided;

[0145] Figure 12 A structural block diagram of still another speech enhancement device provided by an embodiment of the present application is provided

[0146] Figure 13 A structural block diagram of yet another voice enhancement device provided for an embodiment of the present application

[0147] Figure 14 A structural schematic diagram of a terminal device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0148] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, voice enhancement techniques, audio mixing techniques, and terminal devices are omitted so as not to obscure the description of the present application.

[0149] The terminology used in the following description merely for the purpose of describing particular embodiments of the present application and is not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The term "comprising" means "including, but not limited to."

[0150] The terminal device can use voice enhancement techniques to denoise the collected audio data to obtain denoised audio data, and enhance the sound emitted by the user in the audio data. In a multi-person scenario, the terminal device cannot accurately identify the sound emitted by a certain user from the sounds of multiple users. Therefore, the multi-modal voice enhancement technique can be used to denoise the collected audio data in combination with the face sequence images of multiple users in the scene where the terminal device is located, to obtain the sound emitted by a certain user.

[0151] The face sequence image is an image set arranged in order by multiple face images. For example, the face sequence image of the user can be an image set arranged in time sequence by multiple frames of face images corresponding to the user in the video data collected by the terminal device.

[0152] Referring to Figure 1A , Figure 1A A schematic diagram of a framework in which a terminal device uses a multi-modal voice enhancement technique to enhance audio data is shown.

[0153] Specifically, after collecting the audio data, the terminal device can first perform short time fourier transform (STFT) on the audio data to obtain the amplitude spectrum and the phase spectrum of the audio data, and then perform feature extraction on the amplitude spectrum of the audio data through a two-dimensional convolutional neural network (CNN) to obtain speech features. Moreover, the terminal device can also perform feature extraction on the collected lip sequence images through a three-dimensional CNN to obtain visual features.

[0154] Then, the terminal device can combine the extracted speech features and visual features, and input the combined features into a speech enhancement network to obtain a mask corresponding to the phase spectrum of the audio data, so that an enhanced amplitude spectrum can be obtained according to the mask and the amplitude spectrum of the audio data.

[0155] Finally, the terminal device can input the enhanced amplitude spectrum and the phase spectrum of the audio data into another speech enhancement network to obtain an enhanced phase spectrum and a re-enhanced amplitude spectrum, and then perform inverse short time fourier transform (ISTFT) on the enhanced phase spectrum and the re-enhanced amplitude spectrum to obtain enhanced audio data.

[0156] Referring to Figure 1B , Figure 1B Another framework diagram for enhancing audio data by a terminal device using a multi-modal speech enhancement technology is shown.

[0157] Specifically, the terminal device can receive input audio data and face sequence images, perform feature extraction on the audio data to obtain real and imaginary parts of the audio data, and then perform feature extraction on the real and imaginary parts through a two-dimensional CNN to obtain speech features.

[0158] The terminal device can also perform feature extraction on a plurality of face sequence images corresponding to a plurality of users through a pre-set face feature extraction network (facenet), and then perform further extraction on the extracted features through a two-dimensional CNN network to obtain visual features corresponding to the plurality of users.

[0159] Then, the terminal device can combine the speech features and the visual features, and input the combined features into a speech enhancement network to obtain a mask corresponding to the real and imaginary parts of the audio data, so as to obtain the product of the mask and the real and imaginary parts of the audio data, that is, the enhanced real and imaginary parts of the audio data. Finally, the enhanced real and imaginary parts of the audio data are subjected to ISTFT to obtain speech-enhanced audio data.

[0160] However, in the multi-modal speech enhancement technology shown in the above Figure 1A In the multi-modal speech enhancement technology shown in the above

[0161] In the multi-modal speech enhancement technology shown in the above Figure 1B Compared with the multi-modal speech enhancement technology shown in the above Figure 1A Although the information of the specific speaker is encoded, the speech enhancement network is not trained for the specific speaker, and the speech enhancement network cannot focus on the lips in the face sequence image, which also affects the speech enhancement effect.

[0162] Therefore, the present application proposes a training method of a speech enhancement network, which acquires audio data of a user, and combines the face sequence image of the user to acquire the voice feature of the user, and performs online training on the speech enhancement network in real time, so that when the terminal device performs denoising and enhancement by using the trained speech enhancement network, the user's lip movement can be accurately identified based on the specific face feature and the corresponding voice feature of the user, the phenomenon of false muting is reduced, and the accuracy of speech enhancement by the speech enhancement network is improved.

[0163] For example, in the scenario of video call, the terminal device can determine the target user who is speaking in the current scene according to the collected face image, and combine the voice feature of the target user acquired by the terminal device, so that the terminal device can enhance the voice of the target user in the collected audio data, thereby improving the accuracy of identifying the target user, and further improving the accuracy of enhancing the voice of the target user.

[0164] In a multi-person scene, the audio data collected by the terminal device can include the voices of multiple users, and the terminal device needs to enhance the voice of the holder of the terminal device, so the terminal device can determine the voice feature of the holder according to the face image of the holder collected, so as to identify the voice of the holder in the collected audio data, and further enhance the voice of the holder, thereby improving the accuracy of the terminal device in enhancing the voice of the specified user.

[0165] In noisy environments, if a video call is in progress, the audio data captured by the terminal device will contain a significant amount of noise, resulting in low clarity of the user's voice. The terminal device can determine the user's vocal characteristics based on the captured facial image, and then enhance the unclear voice in the audio data based on these characteristics, thereby improving the effectiveness of the terminal device's voice enhancement.

[0166] Of course, the embodiments of this application are not limited to the above-mentioned scenarios, and can also be applied to other scenarios, such as video conferencing, live streaming and instant messaging. The embodiments of this application do not limit the application scenarios of the speech enhancement method.

[0167] The speech enhancement method provided in this application can be applied in various scenarios, each of which can include multiple terminal devices, and each terminal device can train the speech enhancement network. For example, in scenarios such as video calls, video conferencing, and live streaming, the terminal devices can train the speech enhancement network for specific users in the scenario, resulting in a speech enhancement network trained on specific facial features, enabling the speech enhancement network to more accurately recognize the voices emitted by specific users.

[0168] See Figure 2 , Figure 2 The illustration shows a voice enhancement scenario using a video call between two terminal devices as an example. The voice enhancement scenario may include: a first terminal device 201 and a second terminal device 202, which are located in the same network. The owner of the first terminal device 201 is a first user, and the owner of the second terminal device 202 is a second user.

[0169] During a video call, the first terminal device 201 can acquire a first sequence of images of the first user and audio data of the scene where the first terminal device 201 is located. Similarly, the second terminal device 202 can acquire a second sequence of images of the second user and audio data of the scene where the second terminal device 202 is located. If, after acquiring the first sequence of images, the first terminal device 201 determines, based on any face image in the first sequence of images, that the first user's first facial features are not included in its face database, it indicates that the speech enhancement network has not been trained on the first user's first facial features. In this case, the first terminal device 201 can learn the first user's voice features in real time based on the first user's audio data and the first sequence of images, thus completing online training of the speech enhancement network and improving the speech enhancement effect of the terminal device.

[0170] The voice enhancement network is trained for each sample face feature in the face library. For example, the terminal device learns the voice feature of a user according to the face sequence image of the user, and after the voice enhancement network is trained, the terminal device can add the specific face feature of the user to the face library. The voice feature of the user can include the timbre, frequency and voiceprint of the user's voice, and the embodiments of the present application do not limit the information included in the voice feature of the user.

[0171] Moreover, the first sequence image collected by the first terminal device can include image information and lip movement information, the image information extracted from the first sequence image by the first terminal device can include a face image and a face sequence image of the first user, and the lip movement information extracted from the first sequence image by the first terminal device can be a lip movement sequence image of the first user. Moreover, the face image and the lip movement sequence image can be extracted from the face sequence image.

[0172] In addition, the face information, the face data and the face feature can all be used to represent the face of the user, and they can be in the same form or different forms. The face information can be face data, or the face data can be extracted from the face information. For example, the face image contains the face of the first user, the face displayed in the face image can be face information, and the face data extracted from the face image can be face data. The embodiments of the present application do not make specific limitations in this regard.

[0173] Similarly, if the second terminal device 202 determines that the voice enhancement network of the second terminal device 202 does not include the second face feature of the second user after collecting the second sequence image of the second user, the second terminal device 202 can also use a similar method to the above process to perform online training on the voice enhancement network of the second terminal device 202.

[0174] It should be noted that the terminal device can also remind the user to enter the audio data and the face sequence image in a quiet environment when it is first required to perform denoising enhancement through the voice enhancement network, so that the voice enhancement network of the terminal device can obtain the voice feature of the user.

[0175] For example, referring to Figure 3, a schematic diagram of reminding a user to input audio data and a face sequence image by a terminal device is shown. According to an operation triggered by the user, the terminal device first detects that a voice enhancement network needs to be called. The terminal device can display a reminder interface and display reminder information "In order for you to better experience video call, please input sound in a quiet environment and face the mobile phone" to remind the user to input audio data and a face sequence image so that the user can better experience the voice enhancement function. Meanwhile, the terminal device can also provide a confirmation option and a cancellation option for the user to select. If the terminal device detects that the user triggers an operation on the confirmation option, the terminal device can switch to an audio collection interface, which can display a piece of text to remind the user to read the text and provide options such as re-recording, confirmation and cancellation. Meanwhile, the terminal device can also call the camera to collect the face sequence image of the user so that after the collection is completed, the terminal device can learn the voice characteristics of the user according to the collected audio data and face sequence image to complete the training of the voice enhancement network. If the terminal device detects that the user triggers an operation on the cancellation option in the reminder interface, the terminal device can switch to a video call interface, denoises and enhances the audio data through the default voice enhancement network, and reminds the user to input audio data and a face sequence image again when the voice enhancement network needs to be called next time.

[0176] It should be noted that the first face feature, the second face feature and the specific face feature in the embodiments of the present application are all used to indicate the face feature of the user holding the terminal device, that is, the face feature of the user currently using the terminal device, and the description of the face feature of the user holding the terminal device in the embodiments of the present application is not limited.

[0177] The above describes the terminal device can train the voice enhancement network for the user holding the terminal device in real time when the two terminal devices perform video call, thereby improving the effect of voice enhancement. In actual application, the terminal device can also apply the voice enhancement network in other scenarios. For example, when multiple terminal devices perform video conference, the terminal device can enhance the voice of the speaker (the user currently speaking) from the audio data in combination with the face sequence images of multiple users respectively; or, in the process of network live broadcast of a terminal device to other terminal devices, the terminal device can extract the voice of the user from the audio data and remove the noise in the environment, thereby realizing voice enhancement; or, in the scenario of voice interaction between the user and the vehicle machine, the vehicle machine can collect the voice of the user to obtain audio data, and extract the voice of the user through the voice enhancement network, then convert the extracted audio data to obtain text information, and finally execute the instruction corresponding to the text information, that is, the instruction issued by the user.

[0178] Figure 4is a schematic flowchart of a voice enhancement method provided by an embodiment of the present application. By way of example and without limitation, the method can be applied to any of the terminal devices described above, see Figure 4 The method comprises the following steps.

[0179] Step 401: Obtain specific facial features according to the collected facial sequence images.

[0180] In a scenario where the terminal device needs to call a voice enhancement network to perform voice enhancement on audio data, the terminal device can first collect audio data. Meanwhile, the terminal device can also collect facial sequence images of a user (hereinafter referred to as a user) currently using the terminal device, so that the terminal device can extract features from the facial sequence images through a facial feature extraction network to obtain specific facial features corresponding to the user.

[0181] Specifically, when the terminal device detects that the voice enhancement network needs to be called, the terminal device can collect facial sequence images in real time and traverse each image in the collected facial sequence images in the order of collection time to determine whether the facial sequence images include a facial part of the user, obtaining a facial image in the facial sequence images. When the terminal device detects that any image of the facial sequence images includes the facial part of the user, the terminal device can input the facial image into a pre-set facial feature extraction network to extract features from the facial image through the facial feature extraction network, obtaining specific facial features.

[0182] Further, after the terminal device detects the facial image including the facial part, the terminal device can continue to collect facial sequence images and continue to identify whether other images in the facial sequence images include facial parts. If the other images also include facial parts, the terminal device can also input the facial images including the facial parts into the facial feature extraction network to obtain specific facial features.

[0183] In addition, after the terminal device collects the facial sequence images, the terminal device can extract a lip movement sequence image of the user according to the facial part in the facial sequence images, so that in subsequent steps, the terminal device can determine whether the user is speaking according to the lip movement sequence image, or find a user currently speaking from facial sequence images of multiple users.

[0184] Step 402: Compare the specific facial features with sample facial features in a facial library.

[0185] The face library is generated according to each face feature used for training the speech enhancement network. For example, after the terminal device completes training of the speech enhancement network for a certain face feature, the terminal device can add the face feature to the face library as a sample face feature, so that whether the speech enhancement network is trained for a certain face feature can be determined according to each sample face feature in the face library, that is, whether the voice feature of the user is learned in the speech enhancement network can be determined according to each sample face feature in the face library.

[0186] Correspondingly, after the terminal device extracts the specific face feature, the terminal device can compare the specific face feature with each sample face feature in the face library to determine whether the face library includes a sample face feature that is the same as or similar to the specific face feature.

[0187] If any sample face feature in the face library is the same as or similar to the specific face feature, it indicates that the terminal device has possibly trained the speech enhancement network for the specific face feature, and the speech enhancement network has learned the voice feature of the user. The terminal device can perform step 403 to enhance the voice of the user in the collected audio data by using the speech enhancement network to obtain enhanced audio data.

[0188] If the face library does not include a sample face feature that is the same as or similar to the specific face feature, it indicates that the terminal device has not trained the speech enhancement network for the specific face feature, and the speech enhancement network has not learned the voice feature of the user. The terminal device can perform step 404 to temporarily not call the speech enhancement network to denoise and enhance the audio data until the speech enhancement network is trained for the specific face feature.

[0189] It should be noted that in a multi-user scenario, the face sequence image collected by the terminal device in step 401 can extract specific face features of multiple users. If the terminal device detects that specific face features of multiple users are collected, the terminal device can further detect the lip features of each specific face feature to determine whether the user is speaking, thereby determining the speaker who is speaking, and comparing the specific face feature of the speaker with each sample face feature in the face library.

[0190] Of course, the plurality of specific face features acquired by the terminal device according to the face sequence images can also correspond to only one user. Correspondingly, in actual application, due to the influence of environmental factors and light, the specific face features extracted by the terminal device and the sample face features stored in the face library can have certain differences for the face features of the same user. Then, the terminal device can compare the plurality of specific face features extracted from the plurality of images with the sample face features, determine whether the differences between the specific face features and the sample face features are within an error range, and thus determine whether the face library includes the sample face features identical to or similar to the specific face features.

[0191] In step 403, if any sample face feature in the face library is identical to the specific face feature, the terminal device performs de-noising enhancement on the collected audio data through the voice enhancement network.

[0192] After the terminal device determines that the specific face feature is identical to a sample face feature in the face library, it indicates that the voice enhancement network has been trained for the user currently using the terminal device. The terminal device can call the voice enhancement network to enhance the collected audio data.

[0193] Before the terminal device enhances the audio data through the voice enhancement network, the terminal device can determine whether the user is currently speaking according to the action of the user's lips in the face sequence images. If it is detected that the user's lips have obvious actions, it indicates that the user is speaking. The terminal device can obtain the lip movement sequence image of the user from the face sequence images, that is, the sequence image of the user's lip movement when speaking. Then, the terminal device can input the extracted lip movement sequence image, the collected audio data, and the face image in the face sequence image into the voice enhancement network, and perform multi-modal voice enhancement processing through the voice enhancement network to obtain de-noised and enhanced audio data.

[0194] In the process of obtaining the lip movement sequence image from the face sequence image, the terminal device can first recognize the face included in each image of the face sequence image, then recognize the user's lips in the recognized face region, and retain the region where the lips are located in each image and remove other regions, thereby obtaining the lip movement sequence image of the user.

[0195] It should be noted that the terminal device can also obtain the lip movement sequence image of the user in other ways. For example, when the terminal device includes a plurality of cameras, the terminal device can use different cameras to focus on different parts to collect the face sequence image and the lip movement sequence image of the user. Of course, other ways of obtaining the lip movement sequence image of the user can also be used, which are not limited by the embodiments of the present application.

[0196] In addition, in the process of performing the multi-modal speech enhancement processing, the speech enhancement network can extract a specific face feature according to the input face sequence image, find the specific face feature and the corresponding voice according to the specific face feature and the voice feature of the user, and thus can enhance the specific face feature and the corresponding voice in the collected audio data, so as to realize speech enhancement on the currently collected audio data and obtain the audio data after speech enhancement. Other processes of the speech enhancement network in the multi-modal speech enhancement processing can refer to the prior art, and details are not described herein.

[0197] If the terminal device does not detect that the lips of the user have obvious movements according to the collected face sequence image, the terminal device can continue to detect the lips of the user until the lips of the user have obvious movements, that is, the user starts to speak. Then, the terminal device can perform the multi-modal speech enhancement processing in the above manner to obtain the audio data after denoising.

[0198] It should be noted that, in order to improve the reliability of the terminal device in denoising the audio data, the terminal device can detect whether the audio data output by the speech enhancement network has muting after the speech enhancement network denoises the audio data. If the audio data output by the speech enhancement network still has muting, it means that the audio data output by the speech enhancement network still has errors, and then the terminal device does not execute step 403 of calling the speech enhancement network for denoising and enhancement, but executes step 404 and step 405, and outputs the currently collected audio data and re-trains the speech enhancement network to obtain the voice feature of the user, so as to reduce the muting of the speech enhancement network and improve the accuracy and reliability of the speech enhancement network.

[0199] If the audio data output by the speech enhancement network does not have muting, it means that the speech enhancement network can achieve good denoising and enhancement effect on the collected audio data, and then the terminal device can continue to execute step 403 of calling the speech enhancement network for denoising and enhancement until the terminal device stops calling the speech enhancement network according to the operation triggered by the user.

[0200] Step 404: If each sample face feature in the face library is different from the specific face feature, the currently collected audio data is output.

[0201] The terminal device determines that the voice enhancement network has not been trained for the user currently using the terminal device and has not learned the voice feature of the user after determining that the face library does not include a sample face feature identical to the specific face feature. The terminal device can not call the voice enhancement network to denoise and enhance the audio data, but output the audio data that has not been processed, that is, output the currently collected audio data, to avoid the phenomenon of false de-noising caused by the voice enhancement network.

[0202] It should be noted that the terminal device can also execute step 405 simultaneously in the process of executing step 404, that is, the terminal device can perform online training on the voice enhancement network in real time according to the currently collected sequence images and audio data.

[0203] Step 405, training the voice enhancement network according to the face sequence images and the audio data.

[0204] After determining that the voice enhancement network has not been trained for the current user, the terminal device can perform online training on the voice enhancement network in real time according to the face sequence images collected in step 401 in combination with the collected audio data, to obtain the voice feature of the user until a pre-set training condition is met. The training condition can be that the number of training times reaches a number threshold, or that the error output by the voice enhancement network is less than or equal to an error threshold.

[0205] Referring to Figure 5 , Figure 5 A framework diagram for online training of a voice enhancement network by a terminal device provided by an embodiment of the present application, the process of online training of the voice enhancement network in step 405 can include the following steps:

[0206] Step 4051, denoising and enhancing the audio data A collected in combination with the lip movement sequence images B extracted from the face sequence images through the voice enhancement network to obtain first denoised data C.

[0207] Step 4052, comparing the audio data A and the first denoised data C to determine whether the audio data A and the first denoised data C are similar.

[0208] Step 4053, determining whether the user is currently speaking according to the lip movement sequence images B.

[0209] Step 4054, determining whether the face library includes a specific face feature extracted from the face sequence images.

[0210] Step 4055, if the audio data A and the first de-noising data C are similar, the user is currently speaking, and the specific face feature is included in the face library, the pre-stored noise data D is obtained, the noise data D is mixed with the audio data A to obtain mixed data.

[0211] Step 4056, the mixed data is input into the speech enhancement network, the speech enhancement network is trained in combination with the audio data A to obtain the trained speech enhancement network.

[0212] Step 4057, the specific face feature is added to the face library as a sample face feature.

[0213] In step 402, it has been determined that the specific face feature is not included in the face library, and when step 4054 is executed, it is not necessary to determine again whether the specific face feature is included in the face library.

[0214] Moreover, the terminal device extracts the lip movement sequence image B in step 4051, which is similar to the process of extracting the lip movement sequence image by the terminal device in step 403, and will not be described here. In addition, after the face sequence image is extracted in step 401, the terminal device can extract the lip movement sequence image B according to the face sequence image, so that the reliability of the speech enhancement network can be determined according to the extracted lip movement sequence image B in step 403, and the speech enhancement network can be trained in step 405. That is, after step 401 is executed, step 4051 can be executed.

[0215] Then, the terminal device can input the audio data A, the lip movement sequence image B and the face image in the face sequence image into the speech enhancement network, de-noise the audio data A through the speech enhancement network to obtain the first de-noising data C. Then, the terminal device can compare the audio data A and the first de-noising data C to determine the similarity between the audio data A and the first de-noising data C, so as to determine the environment in which the terminal device is currently located and whether the user is using the terminal device in a quiet environment.

[0216] If the similarity between the audio data A and the first de-noising data C is greater than or equal to the similarity threshold, it indicates that the terminal device is currently in a quiet environment, and the terminal device can train the speech enhancement network according to the currently collected audio data A. If the similarity between the audio data A and the first de-noising data C is less than the similarity threshold, it indicates that the terminal device is not currently in a quiet environment, and the terminal device needs to continue to collect audio data until the terminal device is in a quiet environment. Further, if the terminal device is not currently in a quiet environment, the terminal device can remind the user to go to a quiet environment so as to train and use the speech enhancement network.

[0217] For example, the time length of the audio data A and the first de-noising data C is 10 seconds (s), the terminal device can compare the audio data A and the first de-noising data C at each time (at 1s, 2s, 3s, 4s, 5s, 6s, 7s, 8s, 9s and 10s), and obtain the data difference (such as frequency and amplitude) between the audio data A and the first de-noising data C at each time, so as to obtain the similarity between the two. If the similarity is 1 and the similarity threshold is 0.9, when the similarity is greater than or equal to 0.9, it can be determined that the terminal device is currently in a quiet environment; if the similarity is less than 0.9, it can be determined that the terminal device is not currently in a quiet environment.

[0218] While determining the environment of the terminal device, the terminal device can also detect whether the user's lips have moved according to the lip movement sequence image B, that is, whether the user is speaking. If it is detected according to the lip movement sequence image B that the user's lips have moved, it means that the user is speaking, and the terminal device can train the speech enhancement network according to the lip movement sequence image B. However, if the terminal device does not detect that the user's lips have moved, that is, the user does not speak, the terminal device can remind the user to speak in order to collect the audio data A and the lip movement sequence image B and complete the training of the speech enhancement network.

[0219] Since the terminal device has determined in step 402 that the specific face feature is not in the face library, when it is determined that the current environment is quiet and the user's lips have moved, the terminal device can train the speech enhancement network according to the collected audio data A, face sequence image and lip movement sequence image B, in combination with the pre-stored noise data D, to obtain the speech enhancement network for the current user.

[0220] In the process of training the speech enhancement network, the terminal device can obtain the pre-stored noise data D and mix the noise data D with the collected audio data A to obtain mixed audio data with noise. Then, the terminal device can input the mixed audio data, the lip movement sequence image B and the face image in the face sequence image into the speech enhancement network, and use the audio data A collected in the quiet environment as the label data for training, so as to compare the second de-noising data output by the speech enhancement network with the collected audio data A, adjust the speech enhancement network according to the comparison result, and repeat the above process until the speech enhancement network that meets the training condition is trained.

[0221] The training condition can be that the number of times of training the speech enhancement network reaches a number threshold, or that the similarity between the audio data output by the speech enhancement network and the collected audio data reaches a similarity threshold, or other conditions. The embodiments of the present application do not limit the training condition for training the speech enhancement network.

[0222] Further, the terminal device can add the specific face feature of the current user to the face library, take the specific face feature as a sample face feature in the face library, complete updating of the face library, and avoid the need for training the speech enhancement network again when the user corresponding to the specific face feature uses the terminal device to call the speech enhancement network.

[0223] It should be noted that, in the embodiments of the present application, steps 4053 and 4054 are taken as examples for illustration after step 4051, and in actual application, steps 4053 and 4054 can be executed simultaneously with step 4051 or before step 4051, and the embodiments of the present application do not limit this.

[0224] Step 406, calling the speech enhancement network to denoise and enhance the collected audio data.

[0225] After the terminal device completes training of the speech enhancement network, the terminal device can start the speech enhancement function, call the trained speech enhancement network to denoise and enhance the currently collected audio data, thereby enhancing the voice of the user, and making the voice of the user clearer.

[0226] It should be noted that, in step 401, the terminal device can identify one or more face images including face parts from the collected face sequence images, and the following process of speech enhancement of the terminal device is described by taking the terminal device identifying one face image as an example.

[0227] Referring to Figure 6 , Figure 6 The speech enhancement network is shown in a schematic diagram of a framework for speech enhancement, and the process of the terminal device calling the multi-modal speech enhancement network to perform speech enhancement on the collected audio data can include:

[0228] The terminal device can first collect a face sequence image, identify the face sequence image, obtain a face image including a face part in the face sequence image, and then extract a face feature of the user from the face image.

[0229] At the same time, the terminal device can also continue to identify the face sequence image, and extract a lip sequence image based on the lips of the face part collected from the face sequence image. Moreover, the terminal device can collect audio data at the same time of collecting the face sequence image.

[0230] Then, the terminal device can input the obtained face feature, lip sequence image and audio data into the speech enhancement network.

[0231] Correspondingly, the speech enhancement network can first perform feature extraction on the lip sequence image and the audio data respectively to obtain visual features corresponding to the lip sequence image and audio features corresponding to the audio data. Moreover, in order to improve the effect of speech enhancement, the terminal device can also align the visual features and the audio features to obtain aligned visual features and aligned audio features.

[0232] Subsequently, the speech enhancement network can fuse the face features, the aligned visual features and the aligned audio features to obtain fused data, and perform feature extraction on the fused data by using a self-attention model to obtain initial features. Then, feature extraction is performed on the initial features to obtain fused audio features. Finally, the speech enhancement network performs ISTFT on the fused audio features to obtain enhanced audio data.

[0233] In addition, it should be noted that the above embodiment only takes one user as an example to illustrate the process of training the speech enhancement network. In actual application, the terminal device can also train the speech enhancement network in a manner similar to the above training process to obtain the trained speech enhancement network in a multi-person scene.

[0234] In summary, the speech enhancement method provided by the embodiments of the present application, the terminal device acquires a first face image of a user; if the first face image does not match the stored face data, acquires a voice feature of the user, and stores the first face data of the user; acquires a second face image and first audio data of the user; if the second face image matches the stored first face data, enhances the voice of the user in the first audio data based on the voice feature of the first user and outputs. The terminal device can learn the voice feature of the first user in combination with the first face feature of the first user, and can improve the accuracy of recognizing the first user by acquiring the second face image, and can improve the accuracy of enhancing the voice of the first user in the first audio data in combination with the voice feature of the first user (for example, the average 1-second false de-voice phenomenon occurs within 1 minute, and the signal-to-noise ratio between the audio data output by the speech enhancement network and the pure stream audio data reaches 8.8).

[0235] Moreover, before enhancing the first audio data, the terminal device can first determine whether the lips of the user have moved. If the lips of the user have not moved, it means that the user is not currently speaking, and there is no need to enhance the voice of the first user in the first audio data. If the lips of the user have moved, it means that the user is currently speaking, and the terminal device can determine the speech of the user by lip-reading technology according to the dynamic change process of the lips of the user, and can further improve the accuracy of enhancing the voice of the user in the second audio data in combination with the acquired voice feature of the first user.

[0236] In addition, the terminal device inputs the face image of the user to the voice enhancement network, and can perform de-noising and enhancement on the collected audio data in combination with the voice feature of the first user learned by the voice enhancement network, so as to improve the de-noising effect of the voice enhancement network and the effect of enhancing the voice of the user.

[0237] In addition, when it is determined that the voice enhancement network does not learn the voice feature of the first user, the terminal device no longer performs de-noising and enhancement on the collected audio data through the voice enhancement network, but outputs the collected audio data, thereby avoiding the case that the voice enhancement network causes false de-sound to the audio data, and improving the reliability of the terminal device.

[0238] Further, by collecting the audio data in a quiet environment, the collected audio data can be used as label data for training the voice enhancement network, thereby simplifying the process of training the voice enhancement network, improving the efficiency of training the voice enhancement network, and improving the timeliness of calling the voice enhancement network by the terminal device.

[0239] Further, during the running process of the terminal device, whether the specific face feature of the user currently using the terminal device is included in the face library can be determined according to the collected face sequence image and the audio data, so as to determine whether the voice enhancement network can be called to perform de-noising and enhancement on the audio data, thereby improving the accuracy of calling the voice enhancement network and the accuracy of de-noising and enhancement by the voice enhancement network.

[0240] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0241] According to the voice enhancement method described in the above embodiments, Figure 7 is a structural block diagram of a voice enhancement device provided by the embodiments of the present application, and only the parts related to the embodiments of the present application are shown for ease of description.

[0242] Referring to Figure 7 The device comprises:

[0243] The acquisition module 701 is configured to acquire a first face image of a first user.

[0244] The first acquisition module 702 is configured to acquire the voice feature of the first user if the first face image does not match the stored face data.

[0245] The storage module 703 is configured to store the first face data of the first user.

[0246] The collection module 701 is further configured to collect a second facial image and first audio data of the first user.

[0247] The output module 704 is configured to, if the second facial image matches the stored first facial data, perform voice enhancement on the voice of the first user in the first audio data based on the voice feature of the first user and then output the first audio data.

[0248] Optionally, the collection module 701 is further configured to collect second audio data of the first user.

[0249] The output module 704 is further configured to, if the first facial image does not match the stored facial data, output third audio data according to the second audio data, wherein the voice of the first user in the third audio data is not enhanced.

[0250] Optionally, the output module 704 is specifically configured to detect whether the lips of the first user move; if the second facial image matches the stored first facial data and the lips of the first user move, the voice of the first user in the first audio data is enhanced based on the voice feature of the first user and then output.

[0251] Optionally, the output module 704 is specifically configured to acquire a first lip movement sequence image of the first user, detect whether the lips in the first lip movement sequence image move, and perform the voice enhancement on the voice of the first user in the first audio data based on the voice feature of the first user and then output, including: acquiring the first lip movement sequence image, the second facial image and the first audio data, and performing the voice enhancement on the voice of the first user in the first audio data through the voice enhancement network and then output.

[0252] Optionally, referring to Figure 8 , the device further includes:

[0253] The second acquisition module 705 is configured to acquire the enhanced first audio data.

[0254] The detection module 706 is configured to detect whether the enhanced first audio data has a muting phenomenon.

[0255] The first acquisition module 702 is further configured to, if the enhanced first audio data has the muting phenomenon, acquire the voice feature of the first user again.

[0256] The output module 704 is further configured to, if the enhanced first audio data does not have the muting phenomenon, continue to perform the voice enhancement on the voice of the first user in the re-collected audio data based on the voice feature of the first user and then output.

[0257] Optionally, the first acquisition module 702 is specifically configured to collect the fourth audio data and a first sequence image of the first user, the first sequence image comprising face information and lip information; and acquire a voice feature of the first user based on the first sequence image and the fourth audio data.

[0258] Optionally, the first acquisition module 702 is further configured to learn the voice feature of the first user through a speech enhancement network.

[0259] Optionally, the first acquisition module 702 is further configured to acquire a third face image and a second lip movement sequence image according to the first sequence image; input the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learn the voice feature of the first user through the speech enhancement network.

[0260] Optionally, referring to Figure 9 , the apparatus further comprises:

[0261] a first judgment module 707 configured to determine whether the lips of the first user move according to the second lip movement sequence image;

[0262] a second judgment module 708 configured to determine whether the current scene is a quiet environment according to the fourth audio data;

[0263] The first acquisition module 702 is further configured to, if the current scene is a quiet environment and the lips of the first user move, input the third face image, the second lip movement sequence image and the fourth audio data into the speech enhancement network, and learn the voice feature of the first user through the speech enhancement network.

[0264] Optionally, the second judgment module 708 is specifically configured to input the fourth audio data into the speech enhancement network to obtain first denoising data; compare the fourth audio data and the first denoising data; if the similarity between the fourth audio data and the first denoising data is greater than or equal to a similarity threshold, it is determined that the current scene is a quiet environment; if the similarity between the fourth audio data and the first denoising data is less than the similarity threshold, it is determined that the current scene is not a quiet environment.

[0265] Optionally, the first acquisition module 702 is also specifically configured to mix the fourth audio data with pre-stored noise data to obtain mixed audio data; input the mixed audio data, the third face image and the second lip movement sequence image into the speech enhancement network to obtain second de-noising data; and adjust the speech enhancement network according to the second de-noising data and the fourth audio data, so that the speech enhancement network learns the voice feature of the first user.

[0266] Figure 10 is another structural block diagram of a speech enhancement device provided by the embodiment of the application. For ease of illustration, only parts related to the embodiment of the application are shown.

[0267] With reference to Figure 10 , the device comprises:

[0268] The acquisition module 1001 is configured to acquire first audio data and first sequence images of a first user, wherein the first sequence images comprise face information and lip information of the first user.

[0269] The output module 1002 is configured to, if it is determined according to the first sequence images that the face information of the first user matches stored face data and the lips of the first user move, output the first audio data after the voice of the first user is enhanced according to the voice feature of the first user.

[0270] Optionally, the output module 1002 is also configured to, if it is determined according to the first sequence images that the face information of the first user does not match the stored face data, or if it is determined according to the first sequence images that the lips of the first user do not move, output second audio data according to the first audio data, wherein the voice of the first user in the second audio data is not enhanced.

[0271] Optionally, the output module 1002 is specifically configured to extract a first face image and a first lip movement sequence image according to the first sequence images; and enhance the voice of the first user in the first audio data through the speech enhancement network according to the first audio data, the first face image and the first lip movement sequence image, and then output.

[0272] Optionally, with reference to Figure 11 , the device further comprises:

[0273] The first extraction module 1003 is configured to extract the first sequence images to obtain a first lip movement sequence image.

[0274] The first judgment module 1004 is configured to determine whether the lips of the first user move according to the first lip movement sequence image.

[0275] Optionally, referring to Figure 12 , the apparatus further comprises:

[0276] The second extraction module 1005 is configured to extract the first sequence of images to obtain the first facial information.

[0277] The second judgment module 1006 is configured to traverse each of the stored facial data to determine whether each of the stored facial data includes facial data matching the first facial information.

[0278] Optionally, referring to Figure 13 , the apparatus further comprises:

[0279] The first acquisition module 1007 is configured to acquire the enhanced first audio data.

[0280] The detection module 1008 is configured to detect whether the enhanced first audio data has a muting phenomenon.

[0281] The second acquisition module 1009 is configured to reacquire the voice feature of the first user if the enhanced first audio data has a muting phenomenon.

[0282] The output module 1002 is further configured to continue to enhance and output the voice of the first user in the reacquired audio data based on the voice feature of the first user if the enhanced first audio data does not have a muting phenomenon.

[0283] Optionally, the acquisition module 1001 is specifically configured to acquire the first audio data and the first sequence of images of the first user in response to an operation for calling a voice enhancement network.

[0284] In summary, the voice enhancement apparatus provided by the embodiments of the present application is configured to enable a terminal device to acquire a first facial image of a user; acquire a voice feature of the user and store a first facial data of the user if the first facial image does not match stored facial data; acquire a second facial image and first audio data of the user; and enhance and output the voice of the user in the first audio data based on the voice feature of the first user if the second facial image matches the stored first facial data. The terminal device can learn the voice feature of the first user in combination with the first facial feature of the first user, improve the accuracy of recognizing the first user by acquiring the second facial image, and improve the accuracy of enhancing the voice of the first user in the first audio data in combination with the voice feature of the first user (for example, the average muting phenomenon of 1 second occurs within 1 minute, and the signal-to-noise ratio between the audio data output by the voice enhancement network and the pure stream audio data reaches 8.8).

[0285] The following describes a terminal device related to an embodiment of the present application. Please refer to Figure 14 , Figure 14 is a structural schematic diagram of a terminal device provided by an embodiment of the present application.

[0286] The terminal device can include a processor 1410, an external memory interface 1420, an internal memory 1421, a universal serial bus (USB) interface 1430, a charging management module 1440, a power management module 1441, a battery 1442, an antenna 1, an antenna 2, a mobile communication module 1450, a wireless communication module 1460, an audio module 1470, a loudspeaker 1470A, a receiver 1470B, a microphone 1470C, a headset interface 1470D, a sensor module 1480, a key 1490, a motor 1491, an indicator 1492, a camera 1493, a display screen 1494, and a subscriber identification module (SIM) card interface 1495, etc. The sensor module 1480 can include a pressure sensor 1480A, a gyroscope sensor 1480B, a barometric pressure sensor 1480C, a magnetic sensor 1480D, an acceleration sensor 1480E, a distance sensor 1480F, a proximity light sensor 1480G, a fingerprint sensor 1480H, a temperature sensor 1480J, a touch sensor 1480K, an ambient light sensor 1480L, a bone conduction sensor 1480M, etc.

[0287] It can be understood that the structure shown in the embodiment of the present application does not constitute a specific limitation on the terminal device. In another embodiment of the present application, the terminal device can include more or fewer components than the diagram, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0288] The processor 1410 can include one or more processing units, such as: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video code, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors.

[0289] The controller can be the nerve center and command center of the terminal device. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0290] The memory in the processor 1410 can also be configured to store instructions and data. In some embodiments, the memory in the processor 1410 is a cache memory. The memory can save instructions or data that the processor 1410 has just used or repeatedly uses. If the processor 1410 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 1410, thereby improving the efficiency of the system.

[0291] In some embodiments, the processor 1410 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0292] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 1410 can contain multiple sets of I2C bus. The processor 1410 can be coupled to the touch sensor 1480K, the charger, the flash, the camera 1493, etc. through different I2C bus interfaces respectively. For example, the processor 1410 can be coupled to the touch sensor 1480K through an I2C interface, so that the processor 1410 and the touch sensor 1480K communicate through the I2C bus interface, and the touch function of the terminal device is realized.

[0293] The I2S interface can be used for audio communication. In some embodiments, the processor 1410 can contain multiple sets of I2S bus. The processor 1410 can be coupled to the audio module 1470 through the I2S bus, and communication between the processor 1410 and the audio module 1470 is realized. In some embodiments, the audio module 1470 can deliver audio signals to the wireless communication module 1460 through the I2S interface, and the function of answering the phone through the Bluetooth earphone is realized.

[0294] The PCM interface can also be used for audio communication, which samples, quantizes and encodes analog signals. In some embodiments, the audio module 1470 and the wireless communication module 1460 can be coupled through the PCM bus interface. In some embodiments, the audio module 1470 can also deliver audio signals to the wireless communication module 1460 through the PCM interface, and the function of answering the phone through the Bluetooth earphone is realized. The I2S interface and the PCM interface can both be used for audio communication.

[0295] The UART interface is a universal serial data bus, which is used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 1410 and the wireless communication module 1460. For example, the processor 1410 communicates with the Bluetooth module in the wireless communication module 1460 through the UART interface, and the Bluetooth function is realized. In some embodiments, the audio module 1470 can deliver audio signals to the wireless communication module 1460 through the UART interface, and the function of playing music through the Bluetooth earphone is realized.

[0296] The MIPI interface can be used to connect the processor 1410 and the display screen 1494, the camera 1493, and other peripheral devices. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like. In some embodiments, the processor 1410 and the camera 1493 communicate through the CSI interface to implement the photographing function of the terminal device. The processor 1410 and the display screen 1494 communicate through the DSI interface to implement the display function of the terminal device.

[0297] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 1410 and the camera 1493, the display screen 1494, the wireless communication module 1460, the audio module 1470, the sensor module 1480, and the like. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, and the like.

[0298] The USB interface 1430 is an interface that conforms to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, or the like. The USB interface 1430 can be used to connect a charger to charge the terminal device, and can also be used to transmit data between the terminal device and a peripheral device. The interface can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other terminal devices, such as AR devices, and the like.

[0299] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the terminal device. In some other embodiments of the present application, the terminal device can also use different interface connection methods or a combination of multiple interface connection methods.

[0300] The charging management module 1440 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 1440 can receive charging input from a wired charger through the USB interface 1430. In some wireless charging embodiments, the charging management module 1440 can receive wireless charging input through a wireless charging coil of the terminal device. The charging management module 1440 can charge the battery 1442 while also supplying power to the terminal device through the power management module 1441.

[0301] The power management module 1441 is configured to connect the battery 1442 and the charging management module 1440 to the processor 1410. The power management module 1441 receives input power from the battery 1442 and / or the charging management module 1440, and supplies power to the processor 1410, the internal memory 1421, the external memory, the display 1494, the camera 1493, the wireless communication module 1460, and the like. The power management module 1441 can also be configured to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), and the like. In some embodiments, the power management module 1441 can also be disposed in the processor 1410. In some embodiments, the power management module 1441 and the charging management module 1440 can also be disposed in the same device.

[0302] The wireless communication function of the terminal device can be implemented by the antenna 1, the antenna 2, the mobile communication module 1450, the wireless communication module 1460, the modem processor, and the baseband processor, and the like.

[0303] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic waves. Each antenna in the terminal device can be configured to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some embodiments, the antennas can be used in combination with a tuning switch.

[0304] The mobile communication module 1450 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied to the terminal device. The mobile communication module 1450 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like. The mobile communication module 1450 can receive electromagnetic waves from the antenna 1, filter, amplify, and the like the received electromagnetic waves, and transmit the processed signals to the modem processor for demodulation. The mobile communication module 1450 can also amplify the signals modulated by the modem processor, and radiate the signals as electromagnetic waves through the antenna 1. In some embodiments, at least part of the function modules of the mobile communication module 1450 can be disposed in the processor 1410. In some embodiments, at least part of the function modules of the mobile communication module 1450 and at least part of the modules of the processor 1410 can be disposed in the same device.

[0305] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 1470A, a microphone 1470B, etc.), or displays an image or a video through a display screen 1494. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 1410, and can be disposed in the same device as the mobile communication module 1450 or other functional modules.

[0306] The wireless communication module 1460 can provide a wireless communication solution applied to the terminal device, including wireless local area networks (WLAN) (such as a wireless fidelity (Wi-Fi) network), Bluetooth (BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), and the like. The wireless communication module 1460 can be one or more devices integrated with at least one communication processing module. The wireless communication module 1460 receives electromagnetic waves via an antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 1410. The wireless communication module 1460 can also receive a signal to be transmitted from the processor 1410, perform frequency modulation and amplification on the signal, and radiate the signal as an electromagnetic wave via the antenna 2.

[0307] In some embodiments, antenna 1 of the terminal device is coupled with the mobile communication module 1450, and antenna 2 is coupled with the wireless communication module 1460, so that the terminal device can communicate with the network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0308] The terminal device implements display functions through the GPU, the display screen 1494, and the application processor, etc. The GPU is a microprocessor for image processing, connected with the display screen 1494 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 1410 can include one or more GPUs that execute program instructions to generate or change display information.

[0309] The display screen 1494 is configured to display images, videos, and the like. The display screen 1494 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), or the like. In some embodiments, the terminal device can include one or N display screens 1494, where N is a positive integer greater than 1.

[0310] The terminal device can implement the photographing function through the ISP, the camera 1493, the video codec, the GPU, the display screen 1494, and the application processor, and the like.

[0311] The ISP is configured to process the data fed back by the camera 1493. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the algorithm for the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 1493.

[0312] The camera 1493 is configured to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or the like format image signal. In some embodiments, the terminal device can include one or N cameras 1493, where N is a positive integer greater than 1.

[0313] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the terminal device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0314] The video codec is used to compress or decompress digital video. The terminal device can support one or more video codecs. In this way, the terminal device can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0315] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also continuously self-learn. Through the NPU, intelligent cognitive applications of the terminal device can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0316] The external memory interface 1420 can be used to connect an external memory card, such as a Micro SD card, to realize the expansion of the storage capacity of the terminal device. The external memory card communicates with the processor 1410 through the external memory interface 1420 to realize the data storage function. For example, music, video, etc. Files are saved in the external memory card.

[0317] The internal memory 1421 can be used to store computer executable program codes, which include instructions. The processor 1410 executes various function applications and data processing of the terminal device by running the instructions stored in the internal memory 1421. The internal memory 1421 can include a storage program area and a storage data area. The storage program area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The storage data area can store data created during the use of the terminal device (such as audio data, a phone book, etc.), etc. In addition, the internal memory 1421 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0318] The terminal device can realize audio functions through an audio module 1470, a speaker 1470A, a receiver 1470B, a microphone 1470C, a headset interface 1470D, and an application processor, etc. For example, music playing, recording, etc.

[0319] The audio module 1470 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 1470 can also be configured to encode and decode audio signals. In some embodiments, the audio module 1470 can be disposed in the processor 1410, or some functional modules of the audio module 1470 can be disposed in the processor 1410.

[0320] The speaker 1470A, also referred to as a “loudspeaker”, is configured to convert an audio electrical signal into a sound signal. The terminal device can listen to music or listen to a hands-free call through the speaker 1470A.

[0321] The receiver 1470B, also referred to as a “earpiece”, is configured to convert an audio electrical signal into a sound signal. When the terminal device answers a call or a voice message, the user can listen to the voice by holding the receiver 1470B close to the ear.

[0322] The microphone 1470C, also referred to as a “microphone”, “sound collector”, is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by holding the mouth close to the microphone 1470C, and input the sound signal into the microphone 1470C. The terminal device can be provided with at least one microphone 1470C. In other embodiments, the terminal device can be provided with two microphones 1470C, in addition to collecting sound signals, the noise reduction function can also be realized. In other embodiments, the terminal device can also be provided with three, four or more microphones 1470C, to realize the collection of sound signals, noise reduction, and also to identify the source of the sound, to realize the directional recording function, etc.

[0323] The earphone interface 1470D is configured to connect a wired earphone. The earphone interface 1470D can be a USB interface 1430, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0324] The pressure sensor 1480A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 1480A can be disposed on the display screen 1494. The pressure sensor 1480A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 1480A, the capacitance between the electrodes changes. The terminal device determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen 1494, the terminal device detects the intensity of the touch operation based on the pressure sensor 1480A. The terminal device can also calculate the position of the touch based on the detection signal of the pressure sensor 1480A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view a short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.

[0325] The gyroscope sensor 1480B can be configured to determine the motion attitude of the terminal device. In some embodiments, the angular velocity of the terminal device around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 1480B. The gyroscope sensor 1480B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 1480B detects the angle of shaking of the terminal device, calculates the distance that the lens module needs to compensate based on the angle, and lets the lens offset the shaking of the terminal device by reverse movement to achieve anti-shake. The gyroscope sensor 1480B can also be used for navigation and motion sensing game scenarios.

[0326] The barometric pressure sensor 1480C is configured to measure air pressure. In some embodiments, the terminal device calculates the altitude, assists positioning and navigation based on the air pressure value measured by the barometric pressure sensor 1480C.

[0327] The magnetic sensor 1480D includes a Hall sensor. The terminal device can detect the opening and closing of a flip cover by using the magnetic sensor 1480D. In some embodiments, when the terminal device is a flip phone, the terminal device can detect the opening and closing of the flip cover based on the magnetic sensor 1480D. Then, based on the detected opening and closing state of the cover or the flip cover, the terminal device can set features such as automatic unlocking of the flip cover.

[0328] The acceleration sensor 1480E can detect the acceleration of the terminal device in various directions (generally three axes). When the terminal device is stationary, the acceleration sensor 1480E can detect the magnitude and direction of gravity. The acceleration sensor 1480E can also be used to identify the attitude of the terminal device and applied to applications such as landscape / portrait screen switching and pedometers.

[0329] Distance sensor 1480F is used to measure distance. The terminal device can measure distance by infrared or laser. In some embodiments, the terminal device can take a picture of a scene and use the distance sensor 1480F to measure the distance to achieve fast focus.

[0330] Proximity light sensor 1480G can include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode can be an infrared light emitting diode. The terminal device emits infrared light outwardly through the light emitting diode. The terminal device detects infrared reflected light from nearby objects using the photodiode. When sufficient reflected light is detected, it can be determined that there is an object near the terminal device. When insufficient reflected light is detected, the terminal device can determine that there is no object near the terminal device. The terminal device can use the proximity light sensor 1480G to detect when a user is holding the terminal device close to the ear for a call, so as to automatically turn off the screen to achieve power saving. The proximity light sensor 1480G can also be used for automatic unlocking and locking of the screen in a holster mode or a pocket mode.

[0331] Ambient light sensor 1480L is used to sense ambient light brightness. The terminal device can adaptively adjust the display screen 1494 brightness according to the sensed ambient light brightness. The ambient light sensor 1480L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 1480L can also cooperate with the proximity light sensor 1480G to detect whether the terminal device is in the pocket to prevent accidental touch.

[0332] Fingerprint sensor 1480H is used to collect fingerprints. The terminal device can use the collected fingerprint characteristics to implement fingerprint unlocking, access application lock, fingerprint photograph, fingerprint answer incoming call, etc.

[0333] Temperature sensor 1480J is used to detect temperature. In some embodiments, the terminal device uses the temperature detected by the temperature sensor 1480J to implement temperature processing strategies. For example, when the temperature reported by the temperature sensor 1480J exceeds a threshold, the terminal device reduces the performance of the processor located near the temperature sensor 1480J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the terminal device heats the battery 1442 to avoid abnormal shutdown of the terminal device caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, the terminal device performs voltage boosting on the output voltage of the battery 1442 to avoid abnormal shutdown caused by low temperature.

[0334] Touch sensor 1480K, also called "touch panel". Touch sensor 1480K can be disposed on display screen 1494, and touch screen, also called "touch panel", can be formed by touch sensor 1480K and display screen 1494. Touch sensor 1480K is used to detect touch operation acting on or near it. Touch sensor can transmit detected touch operation to application processor to determine touch event type. Visual output related to touch operation can be provided through display screen 1494. In other embodiments, touch sensor 1480K can also be disposed on the surface of terminal device, which is different from the position of display screen 1494.

[0335] Bone conduction sensor 1480M can obtain vibration signal. In some embodiments, bone conduction sensor 1480M can obtain vibration signal of human body sound vibration bone block. Bone conduction sensor 1480M can also contact human body pulse to receive blood pressure pulsation signal. In some embodiments, bone conduction sensor 1480M can also be disposed in earphone to form bone conduction earphone. Audio module 1470 can analyze voice signal based on vibration signal of sound vibration bone block obtained by bone conduction sensor 1480M to realize voice function. Application processor can analyze heart rate information based on blood pressure pulsation signal obtained by bone conduction sensor 1480M to realize heart rate detection function.

[0336] Keys 1490 include power on key, volume key, etc. Keys 1490 can be mechanical keys. They can also be touch keys. Terminal device can receive key input to generate key signal input related to user settings and function control of terminal device.

[0337] Motor 1491 can generate vibration prompt. Motor 1491 can be used for incoming call vibration prompt, and also can be used for touch vibration feedback. For example, touch operation acting on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. Touch operation acting on different regions of display screen 1494 can also correspond to different vibration feedback effects of motor 1491. Different application scenarios (such as time reminder, receiving information, alarm, game, etc.) can also correspond to different vibration feedback effects. Touch vibration feedback effect can also support customization.

[0338] Indicator 1492 can be indicator light, which can be used to indicate charging state, power change, and also can be used to indicate message, missed call, notification, etc.

[0339] The SIM card interface 1495 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 1495 to realize contact and separation with the terminal device. The terminal device can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 1495 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. The same SIM card interface 1495 can simultaneously insert multiple cards. The types of the multiple cards can be the same or different. The SIM card interface 1495 can also be compatible with different types of SIM cards. The SIM card interface 1495 can also be compatible with external storage cards. The terminal device interacts with the network through the SIM card to realize functions such as calling and data communication. In some embodiments, the terminal device uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal device and cannot be separated from the terminal device.

[0340] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0341] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0342] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0343] In the embodiments of the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described system embodiment is merely illustrative. For example, the division of the modules or units is merely logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0344] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0345] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present alone, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0346] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium at least includes any entity or device capable of carrying the computer program code to the terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium can not be an electrical carrier signal and a telecommunication signal.

[0347] Finally, it should be noted that the above only describes specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed by the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A speech enhancement method, characterized in that, The method includes: Collect the first facial image of the first user; If the first face image does not match the stored face data, then the voice features of the first user are obtained; Store the first user's first facial data; Collect the second facial image and first audio data of the first user; If the second face image matches the stored first face data, then the voice of the first user in the first audio data is enhanced and output based on the voice characteristics of the first user.

2. The method according to claim 1, characterized in that, The method further includes: Collect the second audio data of the first user; If the first face image does not match the stored face data, then the third audio data is output based on the second audio data, in which the voice of the first user is not enhanced.

3. The method according to claim 1, characterized in that, If the second face image matches the stored first face data, then based on the voice characteristics of the first user, the voice of the first user in the first audio data is enhanced and output, including: Detect whether the first user's lips are moving; If the second face image matches the stored first face data and the first user's lips move, then the first user's voice in the first audio data is enhanced and output based on the first user's voice characteristics.

4. The method according to claim 3, characterized in that, The detection of whether the first user's lips have moved includes: Obtain the first lip movement sequence image of the first user; Detect whether the lips in the first lip movement sequence image move; The step of enhancing and outputting the first user's voice in the first audio data based on the first user's voice characteristics includes: Based on the first lip movement sequence image, the second face image, and the first audio data, the voice of the first user in the first audio data is enhanced and then output through a speech enhancement network.

5. The method according to any one of claims 1 to 4, characterized in that, After enhancing and outputting the voice of the first user in the first audio data based on the voice characteristics of the first user, the method further includes: Obtain the enhanced first audio data; Detect whether the enhanced first audio data exhibits muting. If the enhanced first audio data is muted, then the voice characteristics of the first user are acquired again. If the enhanced first audio data does not exhibit any muting, then based on the voice characteristics of the first user, the voice of the first user in the newly acquired audio data is enhanced and output.

6. The method according to claim 1, characterized in that, The step of obtaining the voice features of the first user includes: Collect the fourth audio data and the first sequence of images of the first user, wherein the first sequence of images includes: facial information and lip information; Based on the first sequence of images and the fourth audio data, the voice characteristics of the first user are obtained.

7. The method according to claim 6, characterized in that, The step of obtaining the voice features of the first user includes: The voice enhancement network learns the voice characteristics of the first user.

8. The method according to claim 7, characterized in that, The step of learning the voice features of the first user through a speech enhancement network includes: A third face image and a second lip movement sequence image are obtained based on the first sequence image; The third face image, the second lip movement sequence image, and the fourth audio data are input into the speech enhancement network, which learns the voice features of the first user.

9. The method according to claim 8, characterized in that, Before inputting the third face image, the second lip movement sequence image, and the fourth audio data into the speech enhancement network to learn the voice features of the first user through the speech enhancement network, the method further includes: Based on the second lip movement sequence image, determine whether the first user's lips have moved; Determine whether the current scene is a quiet environment based on the fourth audio data; The step of inputting the third face image, the second lip movement sequence image, and the fourth audio data into the speech enhancement network, and learning the voice features of the first user through the speech enhancement network, includes: If the current scene is a quiet environment and the first user's lips move, then the third face image, the second lip movement sequence image, and the fourth audio data are input into the speech enhancement network, and the speech enhancement network learns the voice features of the first user.

10. The method according to claim 9, characterized in that, The step of determining whether the current scene is a quiet environment based on the fourth audio data includes: The fourth audio data is input into the speech enhancement network to obtain the first denoised data; Compare the fourth audio data with the first denoised data; If the similarity between the fourth audio data and the first denoised data is greater than or equal to the similarity threshold, then the current scene is determined to be a quiet environment. If the similarity between the fourth audio data and the first denoised data is less than the similarity threshold, then the current scene is determined to be not a quiet environment.

11. The method according to any one of claims 8 to 10, characterized in that, The step of inputting the third face image, the second lip movement sequence image, and the fourth audio data into the speech enhancement network, and learning the voice features of the first user through the speech enhancement network, includes: The fourth audio data is mixed with pre-stored noise data to obtain mixed audio data; The mixed audio data, the third face image, and the second lip movement sequence image are input into the speech enhancement network to obtain the second denoised data. Based on the second denoised data and the fourth audio data, the speech enhancement network is adjusted so that it learns the voice features of the first user.

12. A speech enhancement method, characterized in that, The method includes: Collect first audio data and first sequence of images from a first user, wherein the first sequence of images includes: the first user's facial information and lip information; If it is determined from the first sequence of images that the facial information of the first user matches the stored facial data, and the lips of the first user move, then the voice of the first user in the first audio data is enhanced and output based on the voice characteristics of the first user; the stored facial data is used to indicate that the voice characteristics of the user corresponding to the facial data have been obtained.

13. The method according to claim 12, characterized in that, The method further includes: If it is determined from the first sequence of images that the face information of the first user does not match the stored face data, or if it is determined from the first sequence of images that the lips of the first user have not moved, then the second audio data is output based on the first audio data, wherein the voice of the first user in the second audio data is not enhanced.

14. The method according to claim 12, characterized in that, The step of enhancing and outputting the voice of the first user in the first audio data based on the voice characteristics of the first user includes: Based on the first sequence of images, a first face image and a first lip movement sequence image are extracted; Based on the first audio data, the first face image, and the first lip movement sequence image, the voice of the first user in the first audio data is enhanced and then output through a speech enhancement network.

15. The method according to claim 12, characterized in that, Before enhancing and outputting the voice of the first user in the first audio data based on the voice characteristics of the first user, the method further includes: Extracting from the first sequence image yields the first lip movement sequence image; Based on the first lip movement sequence image, determine whether the first user's lips have moved.

16. The method according to claim 12, characterized in that, Before enhancing and outputting the voice of the first user in the first audio data based on the voice characteristics of the first user, the method further includes: Extracting the first sequence of images yields the first face information; Iterate through each of the stored face data sets to determine whether each of the stored face data sets includes face data that matches the first face information.

17. The method according to claim 12, characterized in that, After enhancing and outputting the voice of the first user in the first audio data based on the voice characteristics of the first user, the method further includes: Obtain the enhanced first audio data; Detect whether the enhanced first audio data exhibits muting. If the enhanced first audio data is muted, then the voice characteristics of the first user are acquired again. If the enhanced first audio data does not exhibit any muting, then based on the voice characteristics of the first user, the voice of the first user in the newly acquired audio data is enhanced and output.

18. The method according to any one of claims 12 to 17, characterized in that, The acquisition of the first user's first audio data and first sequence of images includes: In response to an operation that invokes the speech enhancement network, the first audio data and the first sequence of images of the first user are acquired.

19. A speech enhancement method, characterized in that, The method, applicable to voice or video calls consisting of a first terminal device and a second terminal device, includes: When the first terminal device detects that an operation to initiate a voice or video call to the second terminal device has been triggered, or when the first terminal device receives a request for a voice or video call initiated by the second terminal device, it collects the first face image of the first user. If the first face image does not match the stored face data, the first terminal device acquires the voice features of the first user; The first terminal device stores the first user's first facial data; The first terminal device collects the second facial image and first audio data of the first user; If the second face image matches the stored first face data, the first terminal device enhances the first user's voice in the first audio data based on the first user's voice characteristics and then outputs it to the second terminal device.

20. The method according to claim 19, characterized in that, The method further includes: The first terminal device collects the second audio data of the first user; If the first face image does not match the stored face data, the first terminal device outputs third audio data based on the second audio data, wherein the first user's voice in the third audio data is not enhanced.

21. An electronic device, characterized in that, include: A processor for running a computer program stored in a memory to implement the speech enhancement method as claimed in any one of claims 1 to 11 or 12 to 18.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the speech enhancement method as claimed in any one of claims 1 to 11 or 12 to 18.

23. A chip system, characterized in that, The chip system includes a memory and a processor, the processor executing a computer program stored in the memory to implement the speech enhancement method as claimed in any one of claims 1 to 11 or 12 to 18.

Citation Information

Patent Citations

  • Intelligent noise reduction method and device in multi-person teleconference

    CN112543302A