Voiceprint noise reduction control method and device, electronic equipment and storage medium

By obtaining the matching results between facial information and target facial information, the voiceprint noise reduction function is automatically controlled, which solves the poor user experience problem caused by manual control in the existing technology and realizes automated and high-precision voiceprint noise reduction control.

CN120833772APending Publication Date: 2025-10-24BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410491619.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing voiceprint noise reduction technology cannot achieve adaptive optimal algorithms, which requires users to manually control the voiceprint noise reduction function, resulting in poor user experience and reduced product competitiveness.

Method used

By obtaining the matching results between facial information and target facial information, the voiceprint noise reduction function is automatically turned on or off. The low-power, always-on image sensor is used to obtain facial information in real time, and the voiceprint noise reduction status is adjusted based on the matching results.

Benefits of technology

Automatic voiceprint noise reduction control is achieved without manual user operation, which improves the user experience, reduces interactive operations, and improves the accuracy and adaptability of the voiceprint noise reduction function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833772A_ABST
    Figure CN120833772A_ABST
Patent Text Reader

Abstract

The invention relates to a voiceprint noise reduction control method and device, electronic equipment and a storage medium. The voiceprint noise reduction control method comprises the following steps: acquiring face information, and determining a matching result of the face information and target face information; and based on the matching result, controlling a voiceprint noise reduction state to be an open or closed state, the voiceprint noise reduction being used for performing noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information. According to the method and the device, the voiceprint noise reduction of the electronic equipment can be determined to be opened or closed based on the face information and the target face information, so that the flexibility of controlling the voiceprint noise reduction state is improved, and the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer vision and audio processing, and particularly relates to a voiceprint noise reduction control method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Voiceprint noise reduction technology is a method of using voiceprint recognition technology to reduce noise in a speech signal. Voiceprint recognition technology refers to a technology of identifying the identity of a speaker by analyzing the speech characteristics of a person, such as pitch, speech rate, tone, etc. In voiceprint noise reduction technology, the identity of the speaker needs to be determined first by voiceprint recognition of the speech signal, and then the speech signal is processed for noise reduction according to the voiceprint characteristics of the identity to improve the quality and clarity of the speech signal. However, in the related art, the traditional voiceprint noise reduction scheme cannot achieve an adaptive optimal algorithm. SUMMARY

[0003] To overcome the problems in the related art, the present disclosure provides a voiceprint noise reduction control method, device, electronic equipment and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, a voiceprint noise reduction control method is provided, comprising: obtaining face information; determining a matching result of the face information and target face information; and based on the matching result, controlling a state of voiceprint noise reduction to be an open state or a closed state, the voiceprint noise reduction being used to perform noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information.

[0005] In an implementation, based on the matching result, controlling the state of voiceprint noise reduction to be the open state or the closed state comprises:

[0006] In response to the matching result representing that the obtained face information matches the target face information, controlling the state of the voiceprint noise reduction to be the open state; and in response to the matching result representing that the obtained face information does not match the target face information, controlling the state of the voiceprint noise reduction to be the closed state.

[0007] In an implementation, the face information is obtained by a low-power always-on image sensor; the obtaining of the face information comprises: obtaining face information at a preset time interval; and before the controlling of the state of the voiceprint noise reduction to be the open state or the closed state based on the matching result, the method further comprises: determining that a face recognition result determined based on the matching result is inconsistent with a face recognition result determined last time, the face recognition result comprising correct or incorrect.

[0008] In an implementation, the method further comprises: if it is determined that the face recognition result determined based on the matching result is consistent with the face recognition result determined last time, returning to execute the process of obtaining the face information.

[0009] In an implementation, before the state of the voiceprint noise reduction is controlled to be in the open state or the closed state based on the matching result, the method further includes: determining that there is an active audio call.

[0010] In an implementation, the method further includes: in response to there being no active audio call, saving the matching result, and when it is monitored that there is an active audio call, controlling the state of the voiceprint noise reduction to be in the open state or the closed state based on the matching result.

[0011] In an implementation, the determining the matching result of the face information and the target face information includes: in response to the number of people in the face information being consistent with the number of people in the target face information and the face being consistent, determining that the face information and the target face information match; and in response to the number of people in the face information being inconsistent with the number of people in the target face information and / or the face being inconsistent, determining that the face information and the target face information do not match.

[0012] According to a second aspect of the embodiments of the present disclosure, a voiceprint noise reduction control apparatus is provided, including: an acquisition unit configured to acquire face information; a matching unit configured to determine a matching result of the face information and target face information; and a processing unit configured to control a state of voiceprint noise reduction to be in an open state or a closed state based on the matching result, the voiceprint noise reduction being used to perform noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information.

[0013] In an implementation, the processing unit controls the state of the voiceprint noise reduction to be in the open state or the closed state based on the matching result in the following manner: in response to the matching result indicating that the acquired face information matches the target face information, the state of the voiceprint noise reduction is controlled to be in the open state; and in response to the matching result indicating that the acquired face information does not match the target face information, the state of the voiceprint noise reduction is controlled to be in the closed state.

[0014] In an implementation, the face information is acquired by a low-power always-on image sensor; the acquisition unit acquires the face information in the following manner: acquiring the face information at a preset time interval; and the processing unit is further configured to: before controlling the state of the voiceprint noise reduction to be in the open state or the closed state based on the matching result, determining that a face recognition result determined based on the matching result is inconsistent with a face recognition result determined last time, the face recognition result including being correct or being incorrect.

[0015] In an implementation, the matching unit is further configured to: if it is determined that the face recognition result determined based on the matching result is consistent with the face recognition result determined last time, return to execute the process of acquiring the face information.

[0016] In an implementation, before the processing unit controls the state of the voiceprint noise reduction to be turned on or turned off based on the matching result, the processing unit is further configured to: determine that there is an active audio call.

[0017] In an implementation, the processing unit is further configured to: in response to there being no active audio call, save the matching result, and in response to monitoring that there is an active audio call, control the state of the voiceprint noise reduction to be turned on or turned off based on the saved matching result.

[0018] In an implementation, the matching unit determines the matching result of the face information and the target face information in the following manner: in response to the number of people in the face information being consistent with the number of people in the target face information, and the face being consistent, determining that the face information and the target face information match; in response to the number of people in the face information being inconsistent with the number of people in the target face information, and / or the face being inconsistent, determining that the face information and the target face information do not match.

[0019] According to a third aspect of embodiments of the present disclosure, an electronic device is provided, comprising:

[0020] a processor; a memory for storing the processor-executable instructions; wherein the processor is configured to perform the voiceprint noise reduction control method in the first aspect or any one of the implementation manners of the first aspect.

[0021] According to a fourth aspect of embodiments of the present disclosure, a storage medium is provided, the storage medium storing instructions, when the instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to perform the method in the first aspect or any one of the implementation manners of the first aspect.

[0022] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects: based on the matching result of the obtained face information and the target face information, the voiceprint noise reduction function is controlled to be turned on or turned off. The voiceprint noise reduction function is automatically controlled to be turned on or turned off based on the obtained face information without manual operation of the user, the accuracy of turning on or turning off the voiceprint noise reduction function is improved, and the experience of the user when using the voiceprint noise reduction function is enhanced.

[0023] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0025] Figure 1is a flowchart of a voiceprint noise reduction control method according to an example embodiment.

[0026] Figure 2 is a flowchart of a voiceprint noise reduction function state control method according to an example embodiment.

[0027] Figure 3 is a flowchart of a face information acquisition method according to an example embodiment.

[0028] Figure 4 is a flowchart of a face information acquisition method according to an example embodiment.

[0029] Figure 5 is a flowchart of a noise reduction control method according to an example embodiment.

[0030] Figure 6 is a flowchart of a matching result determination method according to an example embodiment.

[0031] Figure 7 is a block diagram of a voiceprint noise reduction control method according to an example embodiment.

[0032] Figure 8 is a schematic diagram of a voiceprint noise reduction control method flow according to an example embodiment.

[0033] Figure 9 is a block diagram of a voiceprint noise reduction control device according to an example embodiment.

[0034] Figure 10 is a block diagram of a device for voiceprint noise reduction control according to an example embodiment. DETAILED DESCRIPTION

[0035] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements or details in which the illustrative example embodiments described below are not meant to be limiting. Embodiments are provided as examples consistent with the disclosure and that work well in the practices disclosed herein. Deviations in form can be possible in the interests of promoting clarity and understanding.

[0036] In the drawings, like reference numerals refer to same or similar elements throughout. The described embodiments are part of the disclosure, but not all embodiments. The embodiments described below with reference to the drawings are examples and are intended to explain the disclosure and should not be understood as limiting the disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the disclosure without creative labor are within the scope of the disclosure. The embodiments of the disclosure are described in detail below with reference to the drawings.

[0037] The voiceprint noise reduction control method provided by the embodiments of the present disclosure is applied to the scenario of remote call of an electronic device. The voiceprint noise reduction control method provided by the embodiments of the present disclosure is mainly used for automatically controlling the voiceprint noise reduction function in the call process. The voiceprint noise reduction technology is a method of using voiceprint recognition technology to reduce noise of a voice signal. The voiceprint recognition technology refers to a technology of identifying the identity of a speaker by analyzing the voice characteristics of a person, such as pitch, speech rate, and intonation. In the voiceprint noise reduction technology, voiceprint recognition needs to be performed on the voice signal first to determine the identity of the speaker, and then noise reduction processing is performed on the voice signal according to the voiceprint characteristics of the identity to improve the quality and clarity of the voice signal.

[0038] In the related art, the switch for controlling the voiceprint noise reduction function is mainly set in the call interface and manually controlled by the user to realize the opening or closing of the voiceprint noise reduction function. The voiceprint noise reduction capability is to only transmit the voice of the inputter, and all other sounds are muted, but in the call or conference scenario, such as using a remote conference application conference, there is often a need to pick up the voice of one or more people. The current design of all voiceprint noise reduction solutions cannot achieve an adaptive optimal algorithm. If the electronic device is not held by the person, the non-inputter cannot hear the sound even if he or she is facing the sound collecting device. If you want to pick up the sound of other people in the screen of the electronic device, you need to manually set the sound pickup function to a non-voiceprint mode through the user interface (UI). If you want to pick up only the voice of the inputter, you also need to turn on the voiceprint mode through the UI.

[0039] However, in the related art, since the user needs to manually control the switch of the voiceprint noise reduction function, there is a problem of mistakenly muting the sound of the non-input but speaking person. The user will realize that the voiceprint noise reduction is turned on only when the other party cannot hear the sound, at which time the user has already missed part of the message after manually turning off the voiceprint noise reduction, which brings a poor user experience. Since the related art solution cannot achieve an adaptive optimal algorithm, if the electronic device is in a review state, the non-inputter cannot hear the sound even if he or she is facing the sound collecting device. At the same time, there is a situation that although the electronic device is used by the person, the voiceprint noise reduction function is forgotten to be turned on due to the complex function and the switch hidden deep in the interface, which reduces the product competitiveness.

[0040] Therefore, the embodiments of the present disclosure provide a voiceprint noise reduction control method, which determines the state of voiceprint noise reduction based on the matching result of the obtained target image and target face information, reduces the interactive operation required by the user to use the voiceprint noise reduction function, and improves the user experience.

[0041] For example, in the scenario of a remote conference, since the voiceprint noise reduction function in the related art can only be manually turned on or off, when the terminal is passed to a non-recording user, and the non-recording user uses the terminal to speak, the terminal will filter out the voice of the non-recording user, resulting in that the voice of the non-recording user cannot be obtained. In the embodiment of the present disclosure, when the terminal is passed to a non-recording user, the terminal automatically obtains the face information of the non-recording user, and automatically turns off the voiceprint noise reduction function based on the face information, so as to realize normal recording of the voice of the non-recording user.

[0042] Figure 1 is a flowchart of a voiceprint noise reduction control method according to an exemplary embodiment, as shown in Figure 1 , comprising the following steps.

[0043] In step S11, face information is obtained.

[0044] In the embodiment of the present disclosure, the obtained face information can be a picture obtained by an acquisition device, or video information, feature data or point cloud data containing face information, etc. It should be understood that the above face information is only used for exemplary description, and the specific content and format of the face information are not limited in the embodiment of the present disclosure.

[0045] In step S12, the matching result of the face information and the target face information is determined.

[0046] In the embodiment of the present disclosure, the matching result of the face information and the target face information can be determined based on the similarity of the face features in the face information and the target face information. If the similarity is greater than a preset similarity threshold, it is determined that the face information and the target face information match, otherwise it is determined that the face information and the target face information do not match. It should be understood that a pre-trained similarity matching model can also be used to determine the matching result of the face information and the target face information. The above method of determining the matching result of the face information and the target face information is only used for exemplary description, and the matching result of the face information and the target face information is not specifically limited in the embodiment of the present disclosure.

[0047] In step S13, based on the matching result, the state of voiceprint noise reduction is controlled to be in an open or closed state.

[0048] In the embodiments of the present disclosure, the voiceprint noise reduction is a method for reducing noise of an audio signal based on a voiceprint corresponding to target face information. The target face information is information containing a face of a target user, which can be an image containing the face of the target user, or feature information extracted based on the face of the user, etc. For example, it can be an image of the face of the target user, or point cloud information of the face of the target user. The voiceprint is a sound wave spectrum converted from audio information input by a user, which can be obtained by collecting a speech recording and analyzing the frequency of the sound concentration area. The voiceprint corresponding to the target face information can be obtained by pre-recording audio information by an electronic device through obtaining the target face information, so as to establish a corresponding relationship between the target face information and the voiceprint. The voiceprint recognition technology refers to a technology for identifying the identity of a speaker by analyzing the voice characteristics of a person, such as pitch, speech rate, tone, etc. In the voiceprint noise reduction, the voiceprint of the speech signal needs to be recognized first to determine the identity of the speaker, and then the voice signal is processed based on the voiceprint characteristics of the identity to improve the quality and clarity of the voice signal.

[0049] It should be understood that the operations of acquiring face information, determining the matching result of the face information and the target face information, and controlling the state of the voiceprint noise reduction to be in an open or closed state in the voiceprint noise reduction control method involved in the embodiments of the present disclosure can be implemented by one electronic device, or can be implemented by at least two electronic devices. For example, one terminal can be used to acquire the face information, determine the matching result of the face information and the target face information, and control the state of the voiceprint noise reduction to be in an open or closed state based on the matching result. Alternatively, a camera can be used to acquire the face information, a conference software server can be used to determine the matching result of the face information and the target face information, and an instruction can be sent to a microphone with noise reduction function to control the state of the voiceprint noise reduction of the microphone to be in an open or closed state.

[0050] In the embodiments of the present disclosure, the operations of acquiring face information, determining the matching result of the face information and the target face information, and controlling the state of the voiceprint noise reduction to be in an open or closed state are exemplarily described using one electronic device.

[0051] In the embodiments of the present disclosure, by controlling the voiceprint noise reduction to be in an open or closed state based on the matching result of the target image and the target face information, the automatic opening or closing of the voiceprint noise reduction can be realized without user operation, the interactive operation of the user is reduced, and the use experience of the user is improved.

[0052] In the embodiments of the present disclosure, the adjustment of the voiceprint noise reduction function mainly includes the opening and closing of the voiceprint noise reduction function.

[0053] Figure 2is a flowchart of a voiceprint noise reduction state control method according to an example embodiment, as shown in Figure 2 includes the following steps.

[0054] Figure 2 The steps in step S21 in Figure 1 are the same as step S11 in , which will not be repeated here, and can be referred to in the above description of the embodiments. Only the differences will be described below.

[0055] In step S22a, in response to the matching result representing that the face information matches the target face information, the state of voiceprint noise reduction is controlled to be an open state.

[0056] In the embodiments of the present disclosure, when it is determined that the face information matches the target face information, the electronic device is controlled to adjust the state of voiceprint noise reduction to an open state, i.e., the voiceprint noise reduction function is turned on, so as to filter out noise other than the audio information corresponding to the target face information, and reduce the influence of noise generated by the external environment or other people's voices on the voice of the target user.

[0057] In step S22b, in response to the matching result representing that the face information does not match the target face information, the state of voiceprint noise reduction is controlled to be a closed state.

[0058] In the embodiments of the present disclosure, if it is detected that the matching result is that the face information does not match the target face information, the electronic device is controlled to adjust the state of voiceprint noise reduction to a closed state, i.e., the voiceprint noise reduction function is turned off, so that the electronic device can present all the received audio information, and avoid that other desired audio information cannot be normally obtained due to turning on the voiceprint noise reduction function.

[0059] In the embodiments of the present disclosure, the image acquisition device used to acquire the face information is a low-power always-on image sensor.

[0060] Figure 3 is a flowchart of a face information acquisition method according to an example embodiment, as shown in Figure 3 includes the following steps.

[0061] In step S31, face information is acquired at a preset time interval.

[0062] In the embodiments of the present disclosure, the low-power always-on image sensor can always remain in a working state and acquire images at a fixed frequency, and in combination with existing face detection technology, real-time acquisition of face information is realized. It should be understood that the technology of using a low-power always-on image sensor to acquire face information in real time is mainly used in technical scenarios in which an electronic device detects face information and performs unlocking or monitors the device to respond to acquired face information to record videos or interact with other smart home devices.

[0063] In the embodiments of the present disclosure, the electronic device acquires the face information through the low-power always-on image sensor at fixed time intervals, which can improve the real-time performance of the acquired target image and reduce the power consumption required for real-time acquisition of the image.

[0064] In step S32, it is determined that the face recognition result determined based on the matching result is inconsistent with the face recognition result determined last time, and the state of controlling the voiceprint noise reduction based on the face recognition result is the opening or closing state.

[0065] In the embodiments of the present disclosure, the face recognition result includes correct or incorrect. It should be understood that the face recognition result determined last time is the last face recognition result with the shortest time from the current face recognition result. For example, if the face recognition result is determined at a fixed time interval of 1 second, for the face recognition result corresponding to the 10th second, the face recognition result determined last time is the face recognition result corresponding to the 9th second. If the electronic device determines that the matching result of the face information and the target face information is the same as the last matching result determined, that is, the face recognition result obtained and the last face recognition result are both correct or both incorrect.

[0066] In the embodiments of the present disclosure, when it is detected that the current face recognition result is inconsistent with the face recognition result determined last time, the voiceprint noise reduction function of the electronic device is controlled to be opened or closed based on the current face recognition result. For example, if the face recognition result is set to correct to correspond to the opening of the voiceprint noise reduction function, and the face recognition result is incorrect to correspond to the closing of the voiceprint noise reduction function, when it is determined that the current face recognition result is correct and the face recognition result determined last time is incorrect, the voiceprint noise reduction function is opened; when it is determined that the current face recognition result is incorrect and the face recognition result determined last time is correct, the voiceprint noise reduction function is closed.

[0067] It should be understood that the face information can be acquired by the image acquisition device corresponding to the electronic device, for example, can be acquired by the camera module in the electronic device, or can be acquired by the camera remotely connected to the electronic device. The face information can be a picture, a video or point cloud data acquired by a 3D structured light sensor. In the embodiments of the present disclosure, the face information is acquired by the low-power always-on image sensor only for exemplary description, and the specific format of the target image is not limited.

[0068] In the embodiments of the present disclosure, the target face information can be the face information input in advance by the user using the electronic device. It should be understood that the target face information can also be face data information acquired remotely by the electronic device through the cloud. The above-mentioned manner of acquiring the target image and the target face information is only for exemplary description, and the acquisition manner of the target image and the target face information is not limited in the embodiments of the present disclosure.

[0069] In the embodiments of the present disclosure, when it is detected that the current face recognition result is consistent with the last determined face recognition result, the voiceprint noise reduction function is not controlled, the redundant operation is reduced, and the performance loss of the system is saved.

[0070] Figure 4 is a flowchart of a face information acquisition method according to an example embodiment, as shown in Figure 4 , comprising the following steps.

[0071] Figure 4 The steps in step S41 in Figure 3 are the same as step S31 in , and will not be described here. Please refer to the above description of the embodiments, and only the differences will be described below.

[0072] In step S42, if it is determined that the face recognition result determined based on the matching result is consistent with the last determined face recognition result, the flow of acquiring face information is returned.

[0073] In the embodiments of the present disclosure, if it is detected that the current face recognition information determined based on the matching result is consistent with the last determined face recognition result, that is, the face recognition result has not changed from the last time, the voiceprint noise reduction function does not need to be controlled, and the next flow of acquiring face information is continued, so as to avoid repeated control operation of the voiceprint noise reduction function and unnecessary waste of resources caused by redundant operation. For example, if the face recognition result is set to correctly open the voiceprint noise reduction function, and the face recognition result is set to incorrectly close the voiceprint noise reduction function, when it is determined that the current face recognition result is correct and the last determined face recognition result is also correct, the voiceprint noise reduction function is not controlled.

[0074] Figure 5 is a flowchart of a noise reduction control method according to an example embodiment, as shown in Figure 5 , comprising the following steps.

[0075] Figure 5 The steps in steps S51 and S52 in Figure 1 are the same as steps S11 and S12 in , and will not be described here. Please refer to the above description of the embodiments, and only the differences will be described below.

[0076] In step S53a, in response to determining that there is an active audio call, the state of voiceprint noise reduction is controlled to be in an open or closed state.

[0077] In the embodiments of the present disclosure, the audio call function is a function of realizing real-time voice communication between two or more parties in a communication application. The audio call is applied to application scenarios such as personal calls, work meetings, online education, and voice live broadcast that need to obtain and send audio information. The voiceprint noise reduction function is used to process the obtained audio information when there is an active audio call, and send the processed audio information, so that the sent audio information meets the requirements of clarity and recognition. Therefore, the opening and closing of the voiceprint noise reduction function need to be controlled when there is an active audio call.

[0078] In step S53b, in response to the absence of an active audio call, the matching result is saved, and when it is monitored that the audio call function is in an active state, the state of the voiceprint noise reduction is controlled to be in an open or closed state based on the matching result.

[0079] In the embodiments of the present disclosure, when there is no active audio call, the opening and closing of the voiceprint noise reduction function do not need to be controlled, and then the matching result of the current determined face information and the target face information can be temporarily stored. When it is monitored that there is an active audio call function, the temporarily stored matching result is used to control the state of the voiceprint noise reduction to be in an open or closed state, so as to avoid the situation that when the audio call changes from an inactive state to an active state, the voiceprint noise reduction function is set to be in an open or closed state by default, resulting in abnormal target audio acquisition.

[0080] In the embodiments of the present disclosure, the matching result of the face information and the target face information can be determined based on the number of people and the face information and the target face information.

[0081] Figure 6 is a flowchart of a matching result determination method according to an example embodiment, as shown in Figure 6 , comprising the following steps.

[0082] Figure 6 The steps in step S61 in Figure 1 are the same as step S11 in , and will not be described here. Please refer to the above description of the embodiments, and only the differences will be described below.

[0083] In step S62a, in response to the number of people in the face information and the target face information being consistent, and the faces being consistent, it is determined that the face information and the target face information match.

[0084] In the embodiments of the present disclosure, the consistency of the number of people and the face in the face information and the target face information is detected by first detecting whether the number of people contained in the face information and the target face information is consistent, and then further detecting whether the faces in the face information and the target face information are the same person. When it is detected that the number of people and the face in the face information are consistent with the number of people and the face in the pre-stored target face information, it is determined that the matching result of the face information and the target face information is matched.

[0085] In the embodiments of the present disclosure, the method for detecting whether the faces in the face information and the target face information are consistent can be similarity matching using a pre-trained model, or can be obtaining similarity information based on the extracted feature values using a pre-set algorithm, and if the similarity is greater than a similarity threshold, it is determined that the faces are consistent. It should be understood that the above method for detecting whether the faces in the face information and the target face information are consistent is only used for illustrative description, and the method for detecting whether the faces in the face information and the target face information are consistent is not specifically limited in the embodiments of the present disclosure.

[0086] In step S62b, in response to the number of people in the face information and the target face information being inconsistent, and / or the faces being inconsistent, it is determined that the face information and the target face information are not matched.

[0087] In the embodiments of the present disclosure, if it is detected that the number of people or the face in the face information and the target face information is inconsistent, or neither the number of people nor the face is consistent, that is, when it is detected that the user currently using the electronic device is a user other than the user of the pre-stored target face information or the user currently using the electronic device is multiple people, it is determined that the face information and the target face information are not matched. It should be understood that if the number of people contained in the face information and the target face information is first detected to be consistent, and then the faces in the face information and the target face information are further detected to be the same person, when it is detected that the number of people in the face information and the target face information is inconsistent, it is directly determined that the face information and the target face information are not matched, and there is no need to continue to detect whether the faces in the face information and the target face information are consistent, so as to save the time loss of determining whether the face information and the target face information are matched.

[0088] In the embodiments of the present disclosure, if it is detected that the target image and the target face information are not matched, that is, the user currently using the electronic device is a user other than the user of the pre-stored target face information, at this time, in order to obtain the audio information of the user, it is necessary to close the voiceprint noise reduction function, and all the audio information obtained by the sound collecting device is transmitted, so when it is detected that the target image and the target face information are not matched, the voiceprint noise reduction is closed and all the audio information is picked up.

[0089] In an exemplary embodiment, a block diagram of the electronic device for implementing the voiceprint noise reduction control method is as shown in Figure 7 Figure 7 ​is a block diagram illustrating a voiceprint noise reduction control method according to an example embodiment.

[0090] In Figure 7 , the directed lines between modules or services are used to represent the data flow between modules or services. The always-on camera technology is implemented through an always-on camera detection module, which is included in an advanced digital signal processor (ADSP). The ADSP uses the always-on camera detection module to perform always-on camera (AON) algorithm detection and reports the results to the sensor service in the system framework layer of the electronic device. The AON algorithm detection is a technology that uses a camera for continuous monitoring and detection. The always-on camera captures image or video streams at fixed time intervals and applies a series of algorithms to analyze and process the content of the images or videos. For example, it can recognize face information appearing in the images or videos. The audio service in the system framework listens for changes in the captured information in the sensor service. When the sensor service captures information, the captured information is sent to the audio hardware abstraction layer (Audio HAL) and the current call state is obtained. Finally, the voiceprint detection module in the ADSP is called to implement the function of voiceprint detection adaptive opening. This not only allows the voiceprint noise reduction to be turned on or off at the beginning of each communication, but also enables automatic switching when the personnel in front of the screen changes during communication.

[0091] In an example embodiment, the electronic device implements the voiceprint noise reduction control method based on Figure 7 the content in the block diagram. Figure 8 Figure 8 is a schematic diagram of the flow of the voiceprint noise reduction control method according to an example embodiment.

[0092] In Figure 8 ​In the method, the AON module is started, a low-power consumption captured picture is obtained, and the picture is uploaded to an ADSP neural processing unit (enpu) in real time for algorithm analysis. If the number of people in front of the screen is not one, the algorithm calculates whether the number of people in front of the screen is one, and if the number of people in front of the screen is one, it is determined whether the face has been recorded, and if the face has been recorded, the output is yes, otherwise, the output is no. It is determined whether the result is the same as the last result, and if the result is the same as the last result, the result is not reported, and if the result is different from the last result, the result is reported to the AON service of the sensor service. The Audio Service creates a listener of the AON event, which can listen to the face information in real time. After receiving the face information change report, it is detected whether the system is in a voice over Internet Protocol (voip) active state, and if the system is in the voip active state, it is determined whether the received message is a single person and a voiceprint recorder, and if so, a message is sent to the voiceprint noise reduction module in the ADSP to turn on the function. If it is not a single person or not a voiceprint recorder, a message is sent to the voiceprint noise reduction module in the ADSP to turn off the function. If the system is in a voip inactive state after receiving the information report, the face information state is recorded, and when voip is created next time, it is first determined whether the last face information is a single person and a voiceprint recorder. If so, the AP side sends the state to the voiceprint noise reduction module in the ADSP to automatically turn on the voiceprint noise reduction module; if not, the voiceprint noise reduction module is turned off, and the normal voip processing is adopted. The correctly processed uplink data audio record is transmitted to an application (APP), and the function of automatically turning on and off the voiceprint noise reduction in the voip call is realized, and the user has the best experience.

[0093] In the embodiments of the present disclosure, the face information is obtained by controlling the electronic device, and the voiceprint noise reduction function is turned on or off based on the matching result of the face information and the target face information. The mapping based on the face information and the voiceprint information is realized, and it is automatically determined whether to turn on the voiceprint noise reduction mode, and the intelligent noise reduction is realized. By default, the voiceprint noise reduction function is turned on or off based on the matching result of the face information and the target face information, the voiceprint noise reduction function is turned on or off based on the algorithm, the scene in which the uplink sound is incorrectly muted due to manual judgment is avoided, the user cannot find the switch of the voiceprint noise reduction mode due to the difficulty in finding the switch of the intelligent noise reduction function, and thus the competitiveness of the product is reduced.

[0094] Based on the same concept, the present disclosure also provides a voiceprint noise reduction control device.

[0095] It can be understood that the voiceprint noise reduction control apparatus provided by the embodiments of the present disclosure includes hardware structures and / or software modules corresponding to the implementation of each function in order to achieve the above functions. In combination with the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of the present disclosure.

[0096] Figure 9 is a block diagram of a voiceprint noise reduction control apparatus 100 according to an exemplary embodiment. Referring to Figure 9 The apparatus includes an acquisition unit 101, a matching unit 102, and a processing unit 103.

[0097] The acquisition unit 101 is configured to acquire face information. The matching unit 102 is configured to determine a matching result of the face information and target face information. The processing unit 103 is configured to control a state of voiceprint noise reduction to be an open state or a closed state based on the matching result, the voiceprint noise reduction being configured to perform noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information.

[0098] In an embodiment, the processing unit 103 controls the state of the voiceprint noise reduction to be the open state or the closed state based on the matching result in the following manner: in response to the matching result indicating that the face information matches the target face information, the processing unit 103 controls the state of the voiceprint noise reduction to be the open state; and in response to the matching result indicating that the face information does not match the target face information, the processing unit 103 controls the state of the voiceprint noise reduction to be the closed state.

[0099] In an embodiment, the face information is acquired by a low-power always-on image sensor. The acquisition unit 101 acquires the face information in the following manner: the acquisition unit 101 acquires the face information at a preset time interval. Before the processing unit 103 controls the state of the voiceprint noise reduction to be the open state or the closed state based on the matching result, the matching unit 102 is further configured to: determine that a face recognition result determined based on the matching result is inconsistent with a face recognition result determined most recently, the face recognition result including correct or incorrect.

[0100] In an embodiment, the matching unit 102 is further configured to: if it is determined that the face recognition result determined based on the matching result is consistent with the face recognition result determined most recently, return to execute the process of acquiring the face information.

[0101] In an embodiment, the processing unit 103 is further configured to: before controlling the state of the voiceprint noise reduction to be the open state or the closed state based on the matching result, determine that there is an active audio call.

[0102] In an embodiment, the processing unit 103 is further configured to, in response to the absence of the active audio call, save the matching result, and in response to monitoring that the audio call function is in the active state, control the voiceprint noise reduction based on the matching result.

[0103] In an embodiment, the matching unit 102 is configured to determine the matching result of the obtained face information and the target face information in the following manner: in response to the number of people in the obtained face information being consistent with the number of people in the target face information, and the face being consistent, determining that the obtained face information matches the target face information; in response to the number of people in the obtained face information being inconsistent with the number of people in the target face information, and / or the face being inconsistent, determining that the obtained face information does not match the target face information.

[0104] With regard to the apparatus in the above-described embodiments, specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.

[0105] Figure 10 is a block diagram of an apparatus 200 for voiceprint noise reduction control according to an example embodiment. The apparatus 200 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like, for example.

[0106] Referring to Figure 10 The apparatus 200 can include one or more of the following components: a processing component 202, a memory 204, a power supply component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.

[0107] The processing component 202 generally controls the overall operations of the apparatus 200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 202 can include one or more processors 220 to execute instructions to complete all or part of steps of the methods described above. In addition, the processing component 202 can include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 can include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.

[0108] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phonebook data, messages, pictures, videos, and the like. The memory 204 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0109] The power component 206 provides power to the various components of the device 200. The power component 206 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.

[0110] The multimedia component 208 includes a screen providing an output interface between the device 200 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, swiping, and gestures on the touch panel. The touch sensors can not only sense a boundary of a touching or swiping action, but also detect duration and pressure related to the touching or swiping action. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. The front and / or rear camera can receive external multimedia data when the device 200 is in an operation mode, such as a shooting mode or a video mode. Each of the front and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0111] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.

[0112] The I / O interface 212 provides an interface between the processing component 202 and peripheral interface modules, which can be a keyboard, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0113] The sensor component 214 includes one or more sensors to provide status assessments for various aspects of the device 200. For example, the sensor component 214 can detect an open / closed status of the device 200, relative positioning of components, such as a display and keypad of the device 200, a change in position of the device 200 or a component of the device 200, presence or absence of user contact with the device 200, orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor component 214 can include proximity sensor(s) configured to detect presence of nearby objects without any physical contact. The sensor component 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0114] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and another device. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technology.

[0115] In an exemplary embodiment, the device 200 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic elements, to perform the above-described methods.

[0116] In an exemplary embodiment, a non-transitory computer-readable storage medium, such as the memory 204 including instructions, is also provided, which can be executed by the processor 220 of the device 200 to perform the above-described methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.

[0117] It can be understood that, in the present disclosure, "multiple" refers to two or more, and other quantifiers are similar. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. The singular form "a", "said" and "the" are also intended to include the plural form, unless the context clearly indicates otherwise.

[0118] It can be further understood that the terms "first", "second" and the like are used to describe various information, but these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a particular order or importance. In fact, the expressions of "first", "second" and the like can be completely interchangeable. For example, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information without departing from the scope of the present disclosure.

[0119] It can be further understood that, unless otherwise specified, "connection" includes direct connection between the two without other components, and also includes indirect connection between the two with other elements.

[0120] It can be further understood that, although the operations in the embodiments of the present disclosure are described in a specific order in the accompanying drawings, it should not be understood as requiring the specific order or serial order shown, or requiring all the shown operations to be performed to obtain the desired results. In a specific environment, multi-tasking and parallel processing can be advantageous.

[0121] Other embodiments of the present disclosure will be apparent to those skilled in the art upon consideration of the specification and practice of the disclosed application. The present application is intended to cover any variations, uses or adaptive changes of the present disclosure following the general principles of the present disclosure and including common knowledge or conventional technical means in the art which are not disclosed in the present disclosure.

[0122] It should be understood that the present disclosure is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A voiceprint noise reduction control method, characterized in that, The method comprises: obtaining face information; determining a matching result of the face information and target face information; based on the matching result, controlling a voiceprint noise reduction state to be in an open state or a closed state, the voiceprint noise reduction being used for performing noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information.

2. The voiceprint noise reduction control method of claim 1, wherein, The method comprises: in response to the matching result representing that the face information matches the target face information, controlling the voiceprint noise reduction state to be in the open state; in response to the matching result representing that the face information does not match the target face information, controlling the voiceprint noise reduction state to be in the closed state.

3. The voiceprint noise reduction control method according to claim 1 or 2, characterized in that, The face information is obtained by a low-power always-on image sensor. The method comprises: obtaining face information at a preset time interval; before the step of controlling the voiceprint noise reduction state to be in the open state or the closed state based on the matching result, the method further comprises: determining that a face recognition result determined based on the matching result is inconsistent with a face recognition result determined last time, the face recognition result comprising correct or incorrect.

4. The method of claim 3, wherein, The method further comprises: if it is determined that the face recognition result determined based on the matching result is consistent with the face recognition result determined last time, returning to execute the process of obtaining face information.

5. The voiceprint noise reduction control method of claim 1, wherein, before the step of controlling the voiceprint noise reduction state to be in the open state or the closed state based on the matching result, the method further comprises: determining that there is an active audio call.

6. The voiceprint noise reduction control method according to claim 5, characterized in that, The method further comprises: in response to there being no active audio call, saving the matching result, and based on the matching result, controlling the voiceprint noise reduction state to be in the open state or the closed state when it is monitored that there is an active audio call function.

7. The voiceprint noise reduction control method of claim 1, wherein, The method comprises: in response to the number of people in the face information being consistent with the number of people in the target face information and the face being consistent, determining that the face information matches the target face information; in response to the number of people in the face information being inconsistent with the number of people in the target face information and / or the face being inconsistent, determining that the face information does not match the target face information.

8. A voiceprint noise reduction control device, characterized in that, The device comprises: an obtaining unit, configured to obtain face information; a matching unit, configured to determine a matching result of the face information and target face information; a processing unit, configured to control a voiceprint noise reduction state to be in an open state or a closed state based on the matching result, the voiceprint noise reduction being used for performing noise reduction processing on an audio signal based on a voiceprint corresponding to the target face information.

9. The voiceprint noise reduction control device of claim 8, wherein, The processing unit controls the voiceprint noise reduction state to be in the open state or the closed state based on the matching result in the following manner: in response to the matching result representing that the obtained face information matches the target face information, controlling the voiceprint noise reduction state to be in the open state; in response to the matching result representing that the obtained face information does not match the target face information, controlling the voiceprint noise reduction state to be in the closed state.

10. The voiceprint noise reduction control device according to claim 8 or 9, characterized in that, The face information is obtained by a low-power always-on image sensor. The obtaining unit obtains face information in the following manner: obtaining face information at a preset time interval. The processing unit is further configured to: before the voiceprint noise reduction state is controlled to be turned on or turned off based on the matching result, determine whether a face recognition result determined based on the matching result is inconsistent with a face recognition result determined last time, the face recognition result including correct or incorrect.

11. The apparatus of claim 10, wherein, The matching unit is further configured to: If it is determined that the face recognition result determined based on the matching result is consistent with the face recognition result determined last time, return to execute the process of obtaining face information.

12. The voiceprint noise reduction control device of claim 8, wherein, Before the voiceprint noise reduction state is controlled to be turned on or turned off based on the matching result, the processing unit is further configured to: Determine that there is an active audio call.

13. The voiceprint noise reduction control device of claim 12, wherein, The processing unit is further configured to: In response to there being no active audio call, save the matching result, and when it is monitored that there is an active audio call function, control the voiceprint noise reduction state to be turned on or turned off based on the matching result.

14. The voiceprint noise reduction control device of claim 8, wherein, The matching unit determines the matching result of the face information and the target face information in the following manner: In response to the number of people in the face information being consistent with the number of people in the target face information, and the faces being consistent, determine that the face information and the target face information match; In response to the number of people in the face information being inconsistent with the number of people in the target face information, and / or the faces being inconsistent, determine that the face information and the target face information do not match.

15. An electronic device, comprising: Comprise: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the method of any one of claims 1 to 7.

16. A storage medium, characterized by The storage medium has instructions stored therein, and when the instructions in the storage medium are executed by a processor, the method of any one of claims 1 to 7 can be executed.