Audio source enhancement facilitated using video data

By combining video and audio data, the target audio source is identified and audio processing is performed, solving the problem of target audio signal quality degradation in noisy environments and achieving high-quality audio enhancement and noise reduction.

CN112151063BActive Publication Date: 2025-12-09SYNAPTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010587240.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-27
Filing Date
2020-06-24
Publication Date
2025-12-09
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

In noisy environments, the quality of the target audio signal is easily degraded, and existing technologies struggle to effectively enhance the target audio and reduce noise.

Method used

By combining video and audio data, using video and audio subsystems, the target audio source is identified and audio processing is controlled based on the video signal, thereby enhancing the target audio signal and suppressing noise.

Benefits of technology

It delivers high-quality target audio enhancement even in noisy environments, automatically identifies and controls voice application sessions, reduces noise, and enters sleep mode when the target audio source is absent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112151063B_ABST
    Figure CN112151063B_ABST
Patent Text Reader

Abstract

Systems and methods for audio signal enhancement facilitated using video data are provided. In one example, a method includes receiving a multi-channel audio signal including audio input detected by a plurality of audio input devices. The method also includes receiving an image captured by a video input device. The method further includes determining a first signal based at least in part on the image. The first signal is indicative of a likelihood associated with a target audio source. The method also includes determining a second signal based at least in part on the multi-channel audio signal and the first signal. The second signal is indicative of a likelihood associated with an audio component attributed to the target audio source. The method further includes processing the multi-channel audio signal based at least in part on the second signal to generate an output audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] According to one or more embodiments, the present application relates generally to audio signal processing, and more particularly to audio source enhancement facilitated, for example, using video data. BACKGROUND

[0002] In recent years, audio and video conferencing systems have become popular. In the presence of noise and / or other interfering audio sounds, the quality of a target audio signal degrades. Such audio quality degradation can be easily noticed especially in crowded public environments such as office areas, call centers, cafeteria, and so on. Thus, an audio enhancement solution that enables higher audio quality of a target audio signal even in a noisy environment is desirable. SUMMARY

[0003] According to various embodiments also discussed herein, systems and methods for audio signal enhancement using video data are provided. In some embodiments, such systems and methods can provide a supervised audio / video architecture that allows enhancement of a target audio (e.g., speech of one or more target audio sources) even in a noisy environment. In some aspects, such systems and methods can be used to provide audio signals, and in some cases video signals, for use in voice applications such as voice over internet protocol applications.

[0004] In one or more embodiments, a method includes receiving a multi-channel audio signal including audio input detected by a plurality of audio input devices. The method also includes receiving an image captured by a video input device. The method further includes determining a first signal based at least in part on the image. The first signal is indicative of a likelihood associated with a target audio source. The method also includes determining a second signal based at least in part on the multi-channel audio signal and the first signal. The second signal is indicative of a likelihood associated with an audio component attributed to the target audio source. The method further includes processing the multi-channel audio signal based at least in part on the second signal to generate an output audio signal.

[0005] In one or more embodiments, the system includes a video subsystem and an audio subsystem. The video subsystem is configured to receive images captured by a video input device. The video subsystem includes an identification component configured to determine a first signal based at least partially on the image. The first signal indicates the likelihood of association with a target audio source. The audio subsystem is configured to receive a multi-channel audio signal comprising audio input detected by a plurality of audio input devices. The audio subsystem includes a logic component configured to determine a second signal based at least partially on the multi-channel audio signal and the first signal. The second signal indicates the likelihood of association with an audio component attributed to the target audio source. The audio subsystem also includes an audio processing component configured to process the multi-channel audio signal at least partially based on the second signal to generate an output audio signal.

[0006] The scope of this disclosure is defined by the claims, which are incorporated herein by reference. A more complete understanding of this disclosure, and the implementation of its additional advantages, will be provided to those skilled in the art through consideration of the following detailed description of one or more embodiments. Reference will be made to the accompanying drawings, which will first be briefly described. Attached Figure Description

[0007] A better understanding of the aspects and advantages of this disclosure can be obtained by referring to the following accompanying drawings and the detailed description that follows. It should be recognized that the same reference numerals are used to identify one or more of the same elements illustrated in the drawings, which are shown for the purpose of illustrating embodiments of the disclosure and not for the purpose of limiting the embodiments of the disclosure. The components in the drawings are not necessarily to scale, but rather the emphasis is on clearly illustrating the principles of the disclosure.

[0008] Figure 1 The illustration depicts an example operating environment according to one or more embodiments of the present disclosure, in which a system can operate to facilitate audio source enhancement.

[0009] Figure 2 A high-level diagram of an audio / video processing system for facilitating audio source enhancement, according to one or more embodiments of the present disclosure, is illustrated.

[0010] Figure 3 An example system including a video subsystem and an audio subsystem is illustrated according to one or more embodiments of the present disclosure.

[0011] Figure 4A An example of an input video frame is shown.

[0012] Figure 4B The illustration shows a process according to one or more embodiments of the present disclosure. Figure 4A An example of an output video frame obtained from the background of an input video frame.

[0013] Figure 5 FIG. illustrates an example system including a video subsystem and an audio subsystem to support multiple target audio sources, in accordance with one or more embodiments of the present disclosure.

[0014] Figure 6 FIG. illustrates a flow diagram of an example process for audio source enhancement facilitated using video data, in accordance with one or more embodiments of the present disclosure.

[0015] Figure 7 FIG. illustrates an example electronic system to implement audio source enhancement, in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] The detailed description set forth below is intended as a description of various configurations of the subject technology and is not intended to represent the only configurations in which the subject technology can be practiced. The appended drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details for the purpose of providing a thorough understanding of the subject technology. However, it will be apparent to those skilled in the art that the subject technology can be practiced without

[0017] Various techniques are provided herein to provide audio source enhancement facilitated using video data. In some embodiments, a supervised audio / video system architecture is provided herein to facilitate audio channel noise reduction using video data. In this regard, audio and video modalities are leveraged together to facilitate selective audio source enhancement. Using various embodiments, higher quality of target audio (e.g., speech of one or more target audio sources) can be provided even in noisy environments relative to cases where only audio modalities are leveraged. In some aspects, the audio / video system can authenticate a certain user (e.g., a target audio source) and automatically control the flow of a voice application session (e.g., a call), supervise audio noise reduction to enhance only this authenticated user and remove unwanted (e.g., associated with other talkers) ambient noise, and automatically set the voice application session into a sleep mode when the authenticated user is not present or participating in the call.

[0018] Audio source enhancement techniques can be implemented in single microphone or multi-microphone environments. Such techniques are often used to enhance a target audio source and / or reduce or remove noise. In some cases, such techniques can enhance a target audio source and / or reduce or remove noise by making assumptions about noise spatial or spectral statistics. As an example, for a conference application in general, audio source enhancement can be performed to enhance only the speech from a primary conference user, thereby suppressing all remaining sounds. In some cases, speech from multiple users (e.g., each identified as a primary conference user) can be enhanced while all remaining sounds are suppressed.

[0019] Although the present disclosure is primarily described in association with voice applications, such as Voice over Internet Protocol (VoIP) applications, various embodiments can be utilized to facilitate audio source enhancement in any application in which audio source enhancement can be desired. Further, although the present disclosure is generally described for multi-channel audio implementations, embodiments of the present disclosure can be applied to single-channel audio implementations in some embodiments.

[0020] Figure 1 An example operating environment 100 is illustrated in which a system 105 can operate to facilitate audio source enhancement in accordance with one or more embodiments of the present disclosure. The operating environment 100 includes the system 105, a target audio source 110 (e.g., a user's voice), and noise sources 115A-C. The system 105 includes an audio / video (A / V) processing system 120, audio input devices 125A-D (e.g., microphones), a video input device 130 (e.g., a camera), audio output devices 135A and 135B (e.g., speakers), and a video output device 140 (e.g., a display). In Figure 1 In the example illustrated in FIG. 1, the operating environment 100 is illustrated as the interior of a room 145 (e.g., a conference room, a room of a home), although it is contemplated that the operating environment 100 can include other areas, such as the interior of a vehicle, an outdoor stadium, or an airport.

[0021] Note that although the system 105 is depicted as including four audio input devices, one video input device, two audio output devices, and one video output device, the system 105 can include more or fewer audio input devices, video input devices, audio output devices, and / or video output devices than those illustrated in FIG. 1. Figure 1Fewer or more audio input devices, video input devices, audio output devices, and / or video output devices can be shown in the middle. Also, while the system 105 is depicted as surrounding various of these audio and video devices, various devices can be provided in separate housings and / or as part of separate systems, with the audio / video processing system 120 separate from and communicatively coupled to the audio input devices 125A-D, the video input device 130, the audio output devices 135A and 135B, and / or the video output device 140. In this regard, in some aspects, the audio input devices 125A-D, the video input device 130, the audio output devices 135A and 135B, and / or the video output device 140 can be part of and / or otherwise communicatively coupled to the audio / video processing system 120.

[0022] The audio / video processing system 120 can receive audio signals from the audio input devices 125A-D and video signals (e.g., video frames) from the video input device 130. The audio input devices 125A-D can capture (e.g., detect, sense) audio signals. In some cases, the audio signals can be referred to as forming a multi-channel audio signal, with each channel associated with one of the audio input devices 125A-D. The video input device 130 can capture (e.g., detect, sense) video signals. The video signals can be referred to as video frames or images. The audio / video processing system 120 can process the audio signals using audio processing techniques to detect and enhance target audio 150 produced by the target audio source 110. The target audio 150 is an audio component of the multi-channel audio signal. The target audio 150 can be enhanced (e.g., increased in amplitude and / or clarity) and / or any sound other than the target audio 150 can be suppressed (e.g., decreased in amplitude) by the enhancement. The audio / video processing system 120 can provide the audio signals to the audio output devices 135A and / or 135B, and the video signals (e.g., still images or video) to the video output device 140. The audio output devices 135A and / or 135B can output the audio signals, and the video output device 140 can output the video signals for consumption by one or more users.

[0023] Target audio source 110 can be a person whose speech is to be enhanced by audio / video processing system 120. In embodiments, target audio source 110 can be a person who is participating (e.g., engaged) in a voice application. For example, the person can be participating in a VoIP call. Target audio source 110 can be referred to as an authorized user or an authenticated user (e.g., at least for purposes of the VoIP call). Target audio source 110 produces target audio 150 (e.g., speech) that is to be enhanced by audio / video processing system 120. In addition to target audio source 110, other audio sources in operating environment 100 include noise sources 115A-C. In various embodiments, all audio other than target audio 150 is treated as noise. In Figure 1 In the example illustrated in FIG. 1, noise sources 115A, 115B, and 115C include a speaker playing music, a television playing a television program, and a non-target speaker with a conversion, respectively. It will be appreciated that other noise sources can be present in various operating environments.

[0024] Audio / video processing system 120 can process a multi-channel audio signal to generate an enhanced audio signal. In generating the enhanced audio signal, audio / video processing system 120 takes into account that target audio 150 and noise (e.g., produced by noise sources 115A-C) can arrive at audio input devices 125A-D of system 105 from different directions, the location of each audio source can change over time, and target audio 150 and / or noise can reflect off stationary objects (e.g., walls) within room 145. For example, noise sources 115A-C can produce noise at different locations within room 145, and / or target audio source 110 can speak while walking around room 145. In some embodiments, processing of the multi-channel audio input to obtain the enhanced audio signal can be facilitated by using a video signal from video input device 130, as further described herein.

[0025] As an example, audio / video processing system 120 can include a spatial filter (e.g., a beamformer) that receives the audio signal, identifies the direction of target audio 150 produced by target audio source 110, and outputs an enhanced audio signal (e.g., also referred to as an enhanced target signal) that enhances target audio 150 (e.g., speech or other sound of interest) produced by target audio source 110 using constructive interference and noise cancellation techniques. The operation of the spatial filter to detect the signal and / or enhance the signal can be facilitated by using a video signal (e.g., data derived from the video signal).

[0026] Audio / video processing system 120 can provide enhanced audio signals for use in voice applications (such as voice recognition engines or voice command processors) or as input signals to VoIP applications during VoIP calls. As an example, for illustrative purposes only, consider a VoIP application. In various embodiments, to facilitate the transmission side, audio / video processing system 120 can be used to facilitate VoIP communication across networks (e.g., for conferencing applications). VoIP communication may consist of only voice (e.g., only audio signals) or may include both voice and video. In some cases, audio / video processing system 120 can process images from video input device 130 (e.g., blurring the image) and provide the blurred image for use in VoIP calls. Processed images can be provided for VoIP calls. To facilitate the receiving side, audio / video processing system 120 can receive signals (e.g., audio signals, and in some cases, video signals) from remote devices (e.g., directly or via a network) and output the received signals for use in VoIP communication. For example, the received audio signal can be output via audio output device 135A and / or 135B, and the received video signal can be output via video output device 140.

[0027] On the transmission side, one or more analog-to-digital converters (ADCs) can be used to digitize analog signals (e.g., audio signals, video signals) from one or more input devices (e.g., audio input devices, video input devices), and on the receiving side, one or more digital-to-analog converters (DACs) can be used to generate analog signals (e.g., audio signals, video signals) from digital signals that will be provided by one or more output devices (e.g., audio output devices, video input devices).

[0028] Figure 2 A high-level diagram of an audio / video processing system 200 for facilitating audio source enhancement according to one or more embodiments of the present disclosure is illustrated. However, not all depicted components may be necessary, and one or more embodiments may include additional components not shown in the figures. Variations in the arrangement and type of components, including additional components, different components, and / or fewer components, may be made without departing from the scope of the claims set forth herein. In embodiments, the audio / video processing system 200 may be… Figure 1 The audio / video processing system 120 may include Figure 1 The audio / video processing system 120 or may be Figure 1 It is part of the audio / video processing system 120. For illustrative purposes, regarding Figure 1 The operating environment 100 describes the audio / video processing system 200, although the audio / video processing system 200 can be used in other operating environments.

[0029] The audio / video system 200 includes a video subsystem 205 and an audio subsystem 210. The video subsystem 205 receives an input video frame c(l) (e.g., an image) as input from a video input device 220 (such as a camera), and generates an output video frame ĉ(l) and a monitoring signal (in...). Figure 2 (The term is denoted as "supervision"). Video subsystem 205 provides (e.g., transmits) an output video frame K(l) for use in voice application 215 (such as a VoIP application), and provides (e.g., transmits) a supervision signal to audio subsystem 210. The output video frame K(l) can be a video input frame c(l) or a processed version thereof. In one aspect, the input video frame c(l) can be blurred to obtain the output video frame K(l). For example, a portion of the input video frame c(l) that does not include the target audio source can be blurred.

[0030] The audio subsystem 210 receives a monitoring signal and a multi-channel audio input signal as input. The multi-channel audio input signal is M audio signals x1(l), ..., x2 detected by an array of audio input devices in the operating environment. M The set of (l) is formed, where l represents a time sample. Each audio signal can be provided by a corresponding audio input device and can be associated with an audio channel (e.g., also simply referred to as a channel). Figure 2 In the middle, the audio input device 225A provides an audio signal x l (l), and the audio input device 225B provides an audio signal x M (l). The ellipsis between audio input devices 225A and 225B may indicate one or more additional audio input devices, or no additional input devices (e.g., M=2). For illustrative purposes, audio input devices 225A and 225B are microphones (e.g., forming a microphone array), and audio signals x1(l) and x M (l) is a microphone signal, although in other embodiments, audio input devices 225A, 225B and / or other audio input devices may be other types of audio input devices for providing audio signals to audio subsystem 210.

[0031] In some aspects, M can be at least two to facilitate spatial audio processing to enhance the target audio. When multiple audio input devices are available, they can be utilized to perform spatial processing to improve the performance of speech enhancement techniques. Such spatial diversity can be used in beamforming and / or other methods to better detect / extract the desired source signal (e.g., speech from the target audio source) and suppress interfering source signals (e.g., noise and / or speech from other people). In other aspects, M can be one (e.g., a single microphone) with appropriate single audio input processing to enhance the target audio.

[0032] The audio subsystem 210 can include a multi-channel noise reduction component and a gate component. The multi-channel noise reduction component can facilitate enhancement of audio signals provided by a speaker of interest (e.g., enhancement of speech of such a target audio source). In embodiments, the multi-channel noise reduction component can be controlled by an external voice activity detection (VAD). In some cases, the multi-channel noise reduction component can be configured to be geometrically free (e.g., a user can be anywhere in a 360° space). The gate component can mute signals sent to the voice application 215 (e.g., generate muted audio). For example, the gate component can mute signals sent to the voice application 215 when a target audio source is not in the field of view of the video input device 220 and / or is not engaged with the voice application 215. The selective muting can be controlled based on data (e.g., one or more state variables) provided and continuously updated by the video subsystem 205.

[0033] The multi-channel noise reduction component and the gate component can operate based at least in part on a multi-channel audio input signal and a supervisory signal. For each time sample / , the audio subsystem 210 generates an output audio signal s(l) (e.g., an enhanced audio signal) and provides (e.g., transmits) the output audio signal s(l) for use in the voice application 215. The output audio signal s(l) can enhance audio components of the multi-channel audio input signal that are associated with target audio (e.g., speech) produced by a target audio source. In this regard, the audio subsystem 210 can analyze each of the audio signals (e.g., analyze each audio channel) and utilize data from the video subsystem 205, such as the supervisory signal, to determine whether such audio components associated with a target audio source are present and to process the audio components to obtain the output audio signal s(l).

[0034] In some embodiments, the audio / video processing system 200 can be used to direct the flow of voice application sessions (e.g., meetings, VoIP calls). In an aspect, if it is determined that a target audio source is not in the field of view of a video input device or otherwise not participating in a voice application session, the audio / video processing system 200 can turn on or off one or more of the video input device (e.g., camera) and / or the audio input device (e.g., microphone), reduce playback sound, and / or other operations (e.g., without requiring manual operation by a user). In some cases, the voice application session can be set (e.g., automatically set) into a sleep mode when the target audio source is not present or participating in the session.

[0035] For example, if the target audio source's gaze is directed towards video input device 220 and / or the target audio source is within a threshold distance of video input device 220, it can be determined that the target audio source is participating in the session. In some cases, whether the target audio source participates may depend on characteristics of the target audio source, such as historical data and / or preferences regarding the target audio source's behavior relative to video input device 220. For example, such historical data and / or preferences may indicate whether the target audio source has a habit of being outside the field of view of video input device 220 while speaking (or otherwise participating in the session) and / or whether the target audio source gazes at the video input device while speaking (or otherwise participating in the session).

[0036] In various embodiments, the audio / video processing system 200 can authenticate a user (e.g., specify / identify a target audio source) and automatically control the voice application session. Audio noise reduction can be monitored to enhance the authenticated user and remove any ambient noise, including noise from any unauthorized speakers that may be attributable to the video input device 220, either outside or inside its field of view. In some cases, the voice application session can be set (e.g., automatically set) to sleep mode when the target audio source is absent or not participating in the session.

[0037] Each of the video subsystem 205 and the audio subsystem 210 may include appropriate input / interface circuitry to receive and process video and audio signals, respectively. Such input / interface circuitry can be used to implement anti-aliasing filtering, analog-to-digital conversion, and / or other processing operations. Note that... Figure 2 The diagram illustrates the transmission side of the audio / video processing system 200. In some cases, the audio / video processing system 200 also includes a receiving side to receive audio signals and / or video signals and provide the received signals to an output device.

[0038] Figure 3 An example system 300 including a video subsystem 305 and an audio subsystem 310 according to one or more embodiments of the present disclosure is illustrated. However, not all of the depicted components may be necessary, and one or more embodiments may include additional components not shown in the figures. Variations in the arrangement and type of components, including additional components, different components, and / or fewer components, may be made without departing from the scope of the claims set forth herein. In an embodiment, the video subsystem 305 may be Figure 2 The video subsystem 205 or a part thereof may include Figure 2 The video subsystem 205 or a part thereof may be Figure 2 The video subsystem 205 or a part thereof, or may be implemented in other ways. Figure 2 The video subsystem 205 or a portion thereof. In an embodiment, the audio subsystem 310 may be...Figure 2 The audio subsystem 210 or a portion thereof can include Figure 2 The audio subsystem 210 or a portion thereof can be Figure 2 A portion of the audio subsystem 210 or a portion thereof can be part of Figure 3 The audio subsystem 210 or a portion thereof.

[0039] The video subsystem 305 includes a face detection component 315, a face recognition component 320, a lip motion detection component 325, and a video processing component 330. The face detection component 315 (e.g., also referred to as and / or implemented by a face detector) receives an input video frame c(l) from a video input device (e.g., a camera). In this regard, the video input device can capture the input video frame c(l) and provide the input video frame c(l) to the face detection component 315. The input video frame c(l) includes image data within a field of view (e.g., also referred to as a field of vision) of the video input device.

[0040] For the input video frame c(l), the face detection component 315 detects faces in the input video frame c(l) and generates a face detection signal for each detected face in the input video frame c(l). If no face is detected in the input video frame c(l), the face detection signal generated by the face detection component 315 can indicate that no face is detected in the input video frame c(l). In Figure 3 In this regard, the face detection component 315 detects N faces in the input video frame c(l) and generates the face detection signal b n (l), where n = 1,..., N and each face detection signal is associated with a respective face detected in the input video frame c(l). In this regard, the face detection component 315 provides a face detection signal for each speaker present in the field of view of the video input device. Thus, the face detection signal b n (l) can be referred to as a detected face or as corresponding to a detected face. For example, b 1 (l) is a face detection signal associated with (e.g., corresponding to) a first speaker, b 2 (l) is a face detection signal associated with a second speaker, and so on. Note that the indices / identifiers associated with each speaker (e.g., first, second) can generally be arbitrary and are used to facilitate identification of different speakers. The face detection component 315 provides the face detection signal b n (l) to the face recognition component 320.

[0041] The face detection component 315 can determine a location of any face in the input video frame c(l). The face detection signal b n(l) can be or can include data indicative of a location of a detected face. By way of non-limiting example, the face detection component 315 can utilize a histogram of gradients method, a Viola Jones method, a convolutional neural network (CNN) method (e.g., such as a multi-task CNN (MTCNN) method), and / or any other method generally suitable to facilitate face detection. In some cases, each of these methods can model a human face using a set of generic patterns that output a high response if applied to a face image at the correct location and correct scale. In an aspect, the face detection signal b n (l) is a bounding box (e.g., also referred to as a face box) representing a location and size of a detected face in the input video frame c(l). For example, the location and / or size of a detected face can be represented as coordinates of the input video frame c(l). In some cases, the input video frame c(l) can be visually adjusted such that each detected face in the input video frame c(l) has a bounding box drawn around it.

[0042] In some aspects, in addition to location and size, the face detection component 315 and / or other detection components can identify features of a detected face, such as facial landmarks. In one example, an MTCNN-based face detector can output coordinates of the approximate locations of two eye, nose, and mouth extrema for each detected face. These facial landmarks can be used to align / constrain the face to a generic frontal face, which generally facilitates face recognition (e.g., making face recognition easier). In an aspect, the face detection component 315 can include a face detector for outputting bounding boxes and one or more landmark detectors for identifying facial landmarks.

[0043] The face recognition component 320 (e.g., also referred to as a recognition component, a discrimination component, or a face recognizer) receives the face detection signals b n (l) from the face detection component 315 and processes the face detection signals b n (l) to determine whether any of the face detection signals b n (l) are associated with a target audio source (e.g., an authorized user). The target audio source can be a user that is using the audio / video processing system 300 (such as for a conferencing application). In this regard, in embodiments, the target audio source is a user whose target audio (e.g., speech) is to be enhanced by the audio / video processing system 300.

[0044] Based on the face detection signals b nThe face recognition component 320 generates a face detection signal b(l) and a face detection state Fd(l) based on the determination of whether any of the one or more faces in (1) is associated with the target audio source. In some cases, the face recognition component 320 can also generate a signal d(l) based on the determination. The signal d(l) can include data that facilitates processing of the input video frame c(l), such as a bounding box and / or facial landmark detection. The face recognition component 320 can determine the face detection signal b n is most likely associated with the target audio source. The face detection signal can be provided as the face detection signal b(l). The face recognition component 320 transmits the face detection signal b(l) to the lip motion detection component 325. For example, if the face associated with the face detection signal b 3 is determined to have the highest likelihood (e.g., compared to the remaining face detection signals) of being the target audio source, the face recognition component 320 sets b(l) = b 3 (l) and transmits the face detection signal b(l) to the lip motion detection component 325. In some cases, the face recognition component 320 can determine that none of the detected faces can be associated with the target audio source (e.g., none of the detected faces have at least a minimum threshold likelihood of being the target audio source).

[0045] The face detection state Fd(l) generated by the face recognition component 320 can indicate whether an authorized user is determined to be present in the input video frame c(l). In this regard, the face detection state is a signal that indicates a likelihood (e.g., a probability, a confidence score) that the audio source identified by the face detection state Fd(l) is the target audio source. In one aspect, the face detection state Fd(l) can be a binary signal. For example, in these cases, the face detection state Fd(l) can be 1 only when the target audio source is detected in the field of view of the video input device (e.g., the target audio source is determined to be in the field of view of the video input device), and 0 in other cases. In some cases, the face detection state Fd(l) can take into account whether the target audio source is determined to be engaged with the voice application. In these cases, the face detection state Fd(l) can be 1 only when the target audio source is detected in the field of view of the video input device and is engaged with the voice application, and 0 in other cases. For example, the target audio source can be determined to be engaged based on a direction of gaze of the target audio source and / or an estimated distance between the target audio source and the video input device. In another aspect, the face detection state Fd(l) is not binary and can be a likelihood (e.g., between 0 and 1) that the audio source identified by the face detection state Fd(l) is the target audio source.

[0046] In some aspects, to make the determination, the face recognition component 320 can determine whether the face detection signal bn (1) whether any of the associated detected faces are close enough to a prior face identifier (denoted as prior ID in Figure 2 ) also referred to as a predefined face identifier. The prior face identifier can be a face of a target audio source (e.g., an authorized / authenticated user) of the audio / video processing system 300 or can be associated with a face of a target audio source of the audio / video processing system 300. In an aspect, the prior face identifier can be data (such as an image) of the target audio source, which can be compared to the detected faces in the input video frame c(l).

[0047] As one example, the prior face identifier can be determined during an active enrollment / registration phase. For example, in some cases, a person intending to use the audio / video processing system 305 and / or other components associated with facilitating a voice application can need to subscribe or otherwise register to use the associated equipment and / or software. The prior face identifier can be a pre-registered face. In this regard, the user pre-registers himself as an authorized user of the audio / video processing system 300 (e.g., at least for the purpose of using a voice application such as the voice application 215 of the Figure 4A As another example, the prior face identifier can be determined at the start of a voice application session (e.g., a call) by assuming that the target audio source is a main frontal face in the field of view of a video input device. In this regard, the audio / video processing system 305 identifies a user in front of a video input device communicatively coupled to the audio / video processing system 305 as the target audio source. In some cases, the face can be determined / identified as being associated with an authorized user based on the size of the face and / or the direction of the gaze. For example, if a person's gaze is away from the video capture device (e.g., the person is not engaged with the video capture device) or if the person walks by the video capture device, a person with the largest face in front of the field of view of the video capture device can be determined to not be an authorized user. In some cases, whether a person intending to use the audio / video processing system 305 to facilitate an application (e.g., a voice application) without prior enrollment / registration can depend on settings from an owner and / or manufacturer of the audio / video processing system 305 and / or other components associated with facilitating the application, on settings (e.g., security settings, privacy settings) from a provider of the application, and / or on other entities and / or factors.

[0048] In some aspects, the discerning / identifying of the user does not involve determining the actual identity of the user, and does not involve storing data of the user (e.g., biometrics such as facial landmark characteristics). In this regard, discerning / identifying the user can involve being able to distinguish a certain user from other users during one or more sessions (e.g., based on facial characteristics and / or without determining any actual identity), with data obtained from analyzing images containing faces and / or analyzing audio (e.g., speech) being utilized to make such distinctions.

[0049] In some aspects, the deep video embedding can be used as a face detection signal b n (l) or as part of the processing of the face detection signal b n (l) to determine whether a face (e.g., facial landmarks) is close enough to a previous face identifier. The face recognition component 320 can use a deep convolutional neural network (DCNN)-based approach to recognize a face, such as a face of a target audio source. In such an approach, the face recognition component 320 can receive facial landmarks (e.g., locations, sizes, and / or shapes of a person's lips, nose, eyes, forehead, etc.) in the input video frame c(l). In some cases, the facial landmarks can be received by the face recognition component 320 from the face detection component 315. The DCNN can be trained to embed (e.g., map) a given face image patch into a D-dimensional vector f. The DCNN maps face images of the same individual to the same or similar vector f, independent of environmental condition variations and / or minor pose variations affecting the face image. The similarity between any two faces (e.g., a first face with embedding vector f1 and a second face with embedding vector f2) can be determined (e.g., computed, represented) via a measure (such as L2 similarity or cosine similarity) between their corresponding embedding vectors f1 and f2. To avoid false positives, the similarity between face vectors of two different individuals is preferably large enough (e.g., the similarity between the face vectors is above a threshold).

[0050] To train such a network, assume the availability of a face dataset. In some cases, the face dataset can include face images of available individuals with varying poses, lighting, makeup, and other real-world conditions (e.g., MS-Celeb-1M, CASIA-Webface). Each training batch of the DCNN can include data triplets sampled from the face dataset. Each data triplet can include a face image of an individual (e.g., referred to as an anchor (a)), another face image of the same individual with some real-world variation (e.g., referred to as a positive (p)), and a face image of a different individual (e.g., referred to as a negative (n)). To start the training process, the weights of the DCNN can be randomly initialized. The randomly initialized DCNN can be used to determine a face vector for each of the three face images of a given triplet in order to minimize a triplet loss. The triplet loss can require the DCNN to be penalized if the distance between the anchor and the positive face vector is large, or, conversely, if the distance between the anchor and the negative face vector is small.

[0051] In some aspects, alternatively or in addition to the foregoing methods, the face recognition component 320 can utilize other methods to facilitate detection of the target audio source. The face recognition component 320 can perform face recognition using an eigenface method (e.g., involving a learned classifier in addition to eigen vectors of a covariance matrix of a set of face images), and / or can compute line edge maps for all faces of a dataset and utilize a classifier to distinguish an incoming face image. Various methods can utilize user faces that have been previously enrolled (e.g., registered for purposes of using a voice application or other application previously).

[0052] The lip motion detection component 325 receives the face detection signal b(l) and detects any lip motion associated with the detected face (e.g., determined to be the target audio source). Whether the target audio source is speaking can be based at least in part on any detected lip motion. The lip motion component 325 generates a lip motion detection state Lp(l) and transmits the lip motion detection state Lp(l) to the audio supervision logic component 340. The lip motion detection state Lp(l) indicates a probability (e.g., likelihood, confidence score) that the lips of the target audio source are moving. In some cases, the lip motion detection state Lp(l) indicates a probability (e.g., likelihood, confidence score) that the target audio source is speaking.

[0053] To detect lip motion, the lip motion detection component 325 can identify (e.g., place, position) a plurality of landmarks on the lips of the detected face associated with the face detection signal b(l). In some cases, for a given face, a relative distance between the upper lip and the lower lip can be determined (e.g., estimated) to determine whether the lips are open or closed. If the relative distance changes across frames (e.g., captured by the video input device) enough (e.g., changes more than a threshold amount), the lip motion detection component 325 can determine that the lips are moving.

[0054] The video processing component 330 can receive as input a face detection output including a bounding box and face landmark detection. As an example, in one embodiment, the video processing component 330 is implemented as a background blur component. In such an embodiment, such information (generally represented as signal d(l)) can be used to define a mask around the face. The mask identifies / represents the portion of the input video frame c(l) that is to be blurred by the background blur component. Whether using a bounding box or a convex hull polygon for the face landmarks, a morphological dilation of the detected face region can be performed so that the person's hair and neck are not blurred. The blur itself can be a Gaussian blur, a box blur, or generally any other type of blur. The blur can remove high frequency information from the input video frame c(l) so that if there are other people in the input video frame c(l), their faces cannot be discerned after the blur is applied. In some cases, the entire background region can be replaced by a single color. The single color can be the average background of the scene. In some cases, a user-selected static background or a user-selected moving background can be utilized to replace the background region. As an example, independent of the actual location of the authorized user, the background region can be replaced with an office background or a naturally inspired background (e.g., selected by the authorized user). In some cases, removing, replacing, and / or blurring the background region can enhance privacy (e.g., of the target audio source, other people, and / or location).

[0055] Based on the signal d(l), the background blur component can blur any region around the detected face of the authorized user. In one aspect, the signal d(l) provides a mask region that identifies a region of the input video frame c(l) surrounding the detected face of the authorized user. Alternatively, the signal d(l) provides a region of the face so that the background blur component blurs any region outside of the face region. In some cases, the blurring can provide privacy (e.g., for the authorized user and / or the authorized user's surroundings) and / or facilitate detection of the target audio source (e.g., when other aspects of the input video frame are blurred). In some cases, if no target audio source is detected, the entire input video frame is blurred or blanked.

[0056] Figure 4A and 4BAn example of an input video frame c(l) (labeled 405) and an output video frame ĉ(l) (labeled 410) obtained by processing the background of the input video frame c(l) is illustrated in accordance with one or more embodiments of the present disclosure. In Figure 4B the input video frame 405 includes a person 415, a stereo system 420, a person 425, and a person 430 determined to be a target audio source (e.g., by the facial recognition component 320). As shown in Figure 2 the input video frame 405 is processed such that the output video frame 410 includes the person 415, and the rest of the input video frame 405 (e.g., its background) is replaced with a diagonal background. Note that in some cases, the video subsystem 305 can include an object detection component (e.g., also referred to as an object detector) to detect objects in the input video frame that can be noise sources, such as the stereo system 420. The detected objects can be identified and used to facilitate audio noise reduction.

[0057] Because the background blur component receives facial detection input at each frame, the background blur component can implement blurring of the background that is consistent with (e.g., tracks) the movement of the authorized user. For example, the blurring of the background can follow the target audio source as the target audio source stands up, moves his or her head, and so on. In some cases, the entire body of the target audio source captured in the video frame by the video input device can be segmented such that the hands and / or other body parts of the target audio source are not blurred. For example, by not blurring the body parts of the authorized user, the authorized user can use body language and hand gestures to communicate data. The segmentation can be performed using DCNN-based semantic segmentation or body pose estimation (e.g., OpenPose based on DCNN).

[0058] Although the foregoing describes embodiments in which the video processing component 330 applies a blur to the input video frame c (1), the video processing component 330 can process the input video frame c(l) in other ways that apply a blur or in addition to applying a blur. As one example, a filter can be applied to the input video frame c(l) to enhance the visibility of the target audio source. As another example, in certain applications, a filter can be applied to the input video frame c(l) to adjust the appearance of the target audio source, such as for privacy considerations and / or based on the preferences of the target audio source. In some cases, the video processing component 330 is optional. For example, in some cases, no processing component is utilized such that the output video frame ĉ(l) can be the same as the input video frame c(l).

[0059] Turning now to the audio subsystem 310, the audio subsystem 310 includes an audio VAD component 335, an audio supervision logic component 340, and an audio processing component 345. The audio VAD component 335 receives (from the audio inputs xl(l),...,x M The audio VAD component 335 can be an external (e.g., neural network inference based) audio based VAD that receives a multi-channel audio signal (formed from the audio inputs xl(l),...,x

[0060] In some aspects, the audio VAD component 335 can be used to determine whether the audio input is speech or non-speech, and the video subsystem 305 (e.g., Lp(l) and Fd(l) provided by the video subsystem 305) can be used to determine whether the activity is target audio (e.g., target speech) or interfering audio (e.g., interfering speech). In this regard, the audio VAD component 335 is not used in some cases to distinguish between two (or more) speakers. For example, the VAD signal a(l) can indicate a probability (e.g., likelihood, confidence score) that a person is speaking. False positives associated with identifying that a target audio source is speaking can occur when the audio modality is utilized alone, and similarly can occur when the video modality is utilized alone, when the target audio source is not speaking. For example, for the video modality, the lip motion detection state Lp(l) can sometimes produce false positives. As an example, during a conversation, a speaker can produce movement of the lips without making a sound. Using the various embodiments, false detections associated with identifying that a target audio source is speaking when the target audio source is not actually speaking can be reduced by combining the audio and video modalities together. In one case, the audio and video modalities can be combined by taking the minimum value of a(l) and Lp(l) (e.g., the smaller of a(l) and Lp(l)) to reduce false detections of each modality, as discussed with respect to the audio supervision logic component 340.

[0061] The audio supervision logic 340 generates an audio-video VAD supervision signal p(l) and a hard gating signal g(l). The signal p(l) and g(l) are generated based at least in part on the face detection state Fd(l), the lip motion detection state Lp(l), and the VAD signal a(l). In some cases, the audio supervision logic 340 can apply a non-linear combination of the face detection state Fd(l), the lip motion detection state Lp(l), and the VAD signal a(l) to generate the signal p(l) and g(l). The face detection state Fd(l) and the lip motion detection state Lp(l) can collectively provide Figure 3 supervision signal illustrated in the middle. In this regard, the face detection state Fd(l) and the lip motion detection state Lp(l) provide data that facilitates audio processing by the audio subsystem 310.

[0062] As an example, (for purposes of explanation only) assume that all state variables (e.g., Lp(l), Fd(l), a(l), and / or others) are binary or limited to a range between 0 and 1, p(l) can be defined as the minimum between a(l) and Lp(l) (e.g., p(l) = min(a(l), Lp(l)). In this example case, when utilizing the "min" combination, it can be assumed that each modality (e.g., audio and video) is designed to produce target speech detection with more false positives than false negatives. Similarly, as an example, g(l) can be defined as the minimum between a(l) and Fd(l) (e.g., g(l) = min(a(l), Fd(l)). In some cases, for g(l), temporal smoothing can be applied to prevent the gating from producing jarring rapid discontinuities.

[0063] In some aspects, such data from the video subsystem 305 can facilitate utilization of a VAD (such as a neural network based VAD) that is typically used to identify portions of a signal in the presence of high confidence that observes interference noise in isolation, even in cases where the noise includes speech produced by the interfering speaker(s). In such cases, noise reduction can be facilitated by utilizing the audio modality as well as the video modality (e.g., through supervision by the video subsystem 305) rather than exclusively utilizing the audio modality.

[0064] The audio-video VAD supervisory signal p(l) can control the estimation of noise and speech statistics of the adaptive multi-channel filter. The audio-video VAD supervisory signal p(l) can indicate a probability (e.g., a likelihood, a confidence score) that an audio component of the multi-channel audio signal actually belongs to the target audio source (e.g., to perform enhancement on the correct audio component). The hard gate signal g(l) can be used to mute or unmute the output signal. For example, the hard gate signal g(l) can be used to mute the output signal when there is a high probability (e.g., based at least in part on the values of Fd(l) and Lp(l)) that the target audio source is not in the field of view of the video capture device or participating in the call. In an aspect, the audio supervisory logic 340 and the audio processing logic 345 can collectively implement the multi-channel noise reduction logic and the gate logic of the audio subsystem 310.

[0065] In some embodiments, the audio / video processing system 300 can be used to direct the flow of a voice application session (e.g., a conference, a VoIP call). In an aspect, if it is determined that the target audio source is not in the field of view of the video input device or otherwise not participating in the voice application session, the audio / video processing system 300 can turn on or off one or more of the video input device (e.g., a camera) and / or the audio input device (e.g., a microphone), reduce playback sound, and / or other operations (e.g., without requiring manual operation by a user). In some cases, the voice application session can be set (e.g., automatically set) into a sleep mode when the target audio source is not present or participating in the session. In one case, when the face detection state Fd(l) has a state (e.g., a value) that indicates that the target audio source is not in the field of view of the video input device, the audio / video processing system 300 can mute the audio playback (e.g., set the output audio signal s(l) to zero). Muting the audio playback can also improve privacy in the downlink of the voice application session.

[0066] Each of the video subsystem 305 and the audio subsystem 310 can include appropriate input / interface circuitry to receive and process video signals and audio signals, respectively. Such input / interface circuitry can be used to implement anti-aliasing filtering, analog-to-digital conversion, and / or other processing operations. Note that, Figure 5 The transmit side of the audio / video processing system 300 is illustrated. In some cases, the audio / video processing system 300 also includes a receive side to receive audio signals and / or video signals and provide the received signals to an output device.

[0067] Therefore, in various embodiments, the generation of enhanced audio signals (e.g., s(1)) from multi-channel audio signals is facilitated by utilizing video signals (e.g., c(1)). Identifying / recognizing users from video input signals (e.g., c(l)) and audio input signals (e.g., multi-channel audio signals) and generating appropriate output video signals (e.g., ĉ(l)) and output audio signals (e.g., s(l)) can involve the ability to distinguish one user from other users during one or more sessions of an application (e.g., a voice application). The distinction between one user and other users can be represented as a probability (e.g., likelihood, confidence score) and can be based at least in part on the output signals, such as b... n (l), b(l), d(l), Lp(l), Fd(l), a(l), p(l) and g(1) are obtained by appropriate analysis of the video signal by the video subsystem 305 and appropriate analysis of the audio signal and output signal (e.g., Lp(1), Fd(1)) of the video subsystem 305 by the audio subsystem 310.

[0068] Figure 2 An example system 500 including a video subsystem 505 and an audio subsystem 510 according to one or more embodiments of the present disclosure is illustrated. However, not all of the depicted components may be necessary, and one or more embodiments may include additional components not shown in the figures. Variations in the arrangement and type of components, including additional components, different components, and / or fewer components, may be made without departing from the scope of the claims set forth herein. In an embodiment, the video subsystem 505 may be Figure 2 The video subsystem 205 or a part thereof may include Figure 2 The video subsystem 205 or a part thereof may be Figure 2 The video subsystem 205 or a part thereof, or may be implemented in other ways. Figure 2 The video subsystem 205 or a portion thereof. In an embodiment, the audio subsystem 510 may be... Figure 2 The audio subsystem 210 or a part thereof may include Figure 2 The audio subsystem 210 or a part thereof may be Figure 3 The audio subsystem 210 or a portion thereof, or may be implemented in other ways. Figure 5 The audio subsystem 210 or a part thereof.

[0069] The video subsystem 505 includes a face detection unit 515, a face recognition unit 520, a lip movement detection unit 525, and a video processing unit 530. The audio subsystem 510 includes an audio VAD unit 535, an audio supervision logic unit 540, and an audio processing unit 545. Figure 3The descriptions of Figure 5 , for clarity, provide examples and other descriptions of the differences between Figure 5 and Figure 3 . In this regard, Figure 5 The components of the audio / video processing system 500 can be implemented in the same or similar manner as the various corresponding components of the audio / video processing system 300. Figure 3

[0070] In Figure 5 , the audio / video processing system 500 can be used to facilitate audio signal enhancement (e.g., simultaneous audio signal enhancement) for multiple target audio sources. In this regard, an enhanced audio stream can be generated for multiple target audio sources. As an example, for the mthtarget audio source (e.g., the mthauthenticated user), the facial recognition component 520 can provide a face detection signal b m (l), the signal d m (l), and the face detection state Fd m (l); the lip motion detection component 525 can provide a lip motion detection state Lp m (l); the audio VAD component 535 can provide a VAD signal a m (l); the audio supervision logic component 540 can provide an audio-video VAD supervision signal p m (l) and the hard gate signal g m (l); the video processing component 530 can provide an output video frame ĉ m (l); and the audio processing component 545 can provide an output audio signal s m (l). The facial recognition component 520 can associate each detected face with one of the multiple target audio sources based at least in part on a plurality of prior facial recognition identifiers (denoted as prior IDs). Figure 6 FIGURE 3 illustrates an example case of Figure 3 , in which the audio / video processing system 300 accommodates a single target audio source.

[0071] Figure 7 FIGURE 6 illustrates a flowchart of an example process 600 for audio source enhancement facilitated using video data, in accordance with one or more embodiments of the present disclosure. For explanatory purposes, the example process 600 is described herein with reference to the audio / video processing system 300 of Figure 1 , although the example process 600 can be utilized with other systems. It is noted that one or more operations can be combined, omitted, and / or performed in a different order as desired.

[0072] ​At block 605, the video subsystem 305 receives an image (e.g., input video frame c(l)) captured by a video input device (e.g., a camera). At block 610, the audio subsystem 310 receives a multi-channel audio signal comprising audio inputs (e.g., xl(l),...,xM(l)) detected by a plurality of audio input devices (e.g., microphones). M At block 615, the video subsystem 305 determines, based at least in part on the image, a first signal indicative of a likelihood (e.g., a probability, a confidence score) associated with the target audio source. In some aspects, the first signal can be indicative of a likelihood that a detected face in the image is a face of the target audio source. In some cases, the first signal can be a face detection state Fd(l) generated by the face recognition component 320. The face detection state Fd(l) can be a binary signal or a non-binary signal.

[0073] At block 615, the video subsystem 305 determines, based at least in part on the image, a first signal indicative of a likelihood (e.g., a probability, a confidence score) associated with the target audio source. In some aspects, the first signal can be indicative of a likelihood that a detected face in the image is a face of the target audio source. In some cases, the first signal can be a face detection state Fd(l) generated by the face recognition component 320. The face detection state Fd(l) can be a binary signal or a non-binary signal.

[0074] At block 620, the audio subsystem 310 determines a second signal indicative of a likelihood associated with audio attributed to the target audio source. The second signal can be determined based at least in part on the first signal generated by the video subsystem 305 at block 615. In some cases, the second signal can be determined further based on detected lip motion (e.g., a lip motion detection state Lp(l)) and an audio VAD signal (e.g., a(l)). In some aspects, the second signal can be indicative of a likelihood that an audio component detected in the multi-channel audio signal belongs to the target audio source. In some cases, the second signal can be an audio-video VAD supervision signal p(l) generated by the audio supervision logic component 340.

[0075] At block 625, the audio subsystem 310 processes the multi-channel audio signal based at least in part on the second signal to generate an output audio signal (e.g., an enhanced audio signal s(l)). At block 630, the video subsystem 305 processes the image to generate an output video signal (e.g., ĉ(l)). In an aspect, the video subsystem 305 can apply a blur to the image. At block 635, the audio / video processing system 300 transmits the output audio signal (e.g., for use in a voice application). At block 640, the audio / video processing system 300 transmits the output video signal (e.g., for use in a voice application). In some cases, such as when the voice application involves a voice-only call, blocks 630 and 640 can be optional.

[0076] Figure 1An example electronic system 700 for implementing audio source enhancement is illustrated in accordance with one or more embodiments of the present disclosure. However, not all of the depicted components can be required, and one or more embodiments can include additional components not pictured. Variations in the arrangement and type of the components can be made without departing from the scope of the claims as set forth herein.

[0077] The electronic system 700 includes one or more processors 705, memory 710, input components 715, output components 720, and communication interfaces 725. The various components of the electronic system 700 can interface and communicate over buses or other electronic communication interfaces. The electronic system 700, for example, can be or can be coupled to a mobile phone, a tablet computer, a laptop computer, a desktop computer, a car, a personal digital assistant (PDA), a television, a speaker (e.g., a conference speaker with image capture capabilities), or any electronic device that typically receives audio and video signals (e.g., from audio and video input devices) and transmits the signals directly or via a network to other devices.

[0078] The processor(s) 705 can include one or more of a processor, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device (PLD) (e.g., a field-programmable gate array (FPGA)), a digital signal processing (DSP) device, or other device that can be configured to perform various operations discussed herein for audio source enhancement, either by way of hardwiring, executing software instructions, or a combination of the two. In this regard, the processor(s) 705 can be operable to execute instructions stored in the memory 710 and / or other memory components. In embodiments, the processor(s) 705 can perform operations of various components of the audio / video processing systems 120, 200, 300, and 500 of FIGS. 1, 2, 3, and 5, respectively. As an example, the processor(s) 705 can receive multi-channel audio input signals from audio input devices (e.g., 125A-D of FIG. 1) and images from video input devices (e.g., 130 of FIG. 1) and process these audio and video signals. Figure 1 , 2 ​ ​

[0079] ​​​Memory 710 can be implemented as one or more memory devices operable to store data, including audio data, video data, and program instructions. Memory 710 can include one or more types of memory devices, including volatile and non-volatile memory devices, such as random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory, a hard disk drive, and / or other types of memory.

[0080] Input component 715 can include one or more devices for receiving input. In an aspect, input component 715 can include a touch screen, a touch pad display, a keypad, one or more buttons, a dial or a knob, and / or other components operable to enable a user to interact with electronic system 700. In some cases, input component 715 can include an audio input device(s) (e.g., a microphone) or a video input device(s) (e.g., a camera). For example, input component 715 can provide input audio signals and input video signals to processor(s) 705. In other cases, input component 715 does not include an audio input device(s) and / or a video input device(s) that provide input audio signals and input video signals to processor(s) 705 for the purposes of audio source enhancement. Output component 720 can include one or more devices for emitting audio and / or video output. In some cases, output component 720 can include an audio output device(s) (e.g., a speaker) or a video input device(s) (e.g., a display).

[0081] Communication interface 725 facilitates communications between electronic system 700 and networks and external devices. For example, communication interface 725 can enable Wi-Fi (e.g., IEEE 802.11) or Bluetooth connectivity between electronic system 700 and one or more local devices, such as external device 730, or enable a connection to a wireless router to provide network access to external device 735 via network 740. In various embodiments, communication interface 725 can include wired and / or other wireless communication components for facilitating direct or indirect communication between electronic system 700 and other devices. As an example, a user(s) of external device 735 can engage in a VoIP call with a user(s) of electronic system 700 via wireless communication between electronic system 700 and network 740, and between network 740 and external device 735.

[0082] Where applicable, various embodiments provided by the present disclosure can be implemented using hardware, software, or combinations of hardware and software. Moreover, where applicable, the various hardware components and / or software components set forth herein can be combined into a composite component comprising software, hardware, and / or both, where applicable. It will depend on the particular application and the overall design constraints imposed on the overall system. In addition, where applicable, the various hardware and / or software components set forth herein can be separated into sub-components not explicitly described, where applicable. In addition, where applicable, the various connections between system components set forth herein can also be implemented as a software program, hardware, and / or both. Where applicable, embodiments of the present disclosure can be implemented using a general purpose computer, a special purpose computer, a microprocessor, or any other programmable data processing apparatus to produce a machine, such that the various steps of the disclosed implementations can be performed.

[0083] According to the present disclosure, software, such as program code and / or data, can be stored on one or more computer-readable media. It is also contemplated that software identified herein can be implemented using one or more general purpose or special purpose computers and / or computer systems, networked and / or otherwise. Where applicable, the order of the steps described herein can be changed, combined into composite steps, and / or separated into sub-steps to provide features described herein.

[0084] The foregoing disclosure is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Thus, various alternatives and modifications for various embodiments of the present disclosure described herein will be apparent to those skilled in the art. Accordingly, the various embodiments described herein should be construed as being merely illustrative and not restrictive. Embodiments of the present disclosure have been described as they can come to be implemented, and the person of ordinary skill in the art will recognize that changes can be made to form and details without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

Claims

1. A method comprising: Receives multi-channel audio signals including audio input detected by multiple audio input devices; Receives images captured by a video input device; Detect at least one face in the image; Identifying one of the at least one faces as an authenticated user is based at least in part on a predefined facial identifier; The first signal is determined at least in part based on the image, wherein the first signal indicates the likelihood of association with the authenticated user; The second signal is determined at least in part based on the multi-channel audio signal and the first signal, wherein the second signal indicates the probability of being associated with an audio component attributed to the user of the authentication. as well as The multi-channel audio signal is processed at least in part based on the second signal to generate an output audio signal, wherein the processing includes: Determine whether the authenticated user exists or is participating in the session; When the authenticated user does not exist or is not participating in the session, a silent audio is generated; When the authenticated user is present and participates in the session, the audio component attributed to the authenticated user is enhanced and the audio component associated with other speakers is removed; Determining whether the authenticated user exists or participates in the session includes: Whether a certified user participates in a session is determined based on historical data and / or the certified user's preferences regarding the certified user's behavior relative to the video input device.

2. The method of claim 1, wherein the plurality of audio input devices comprises an array of microphones.

3. The method according to claim 1, further comprising: Receive multiple images; Identify the audio source in the plurality of images as the authenticated user; as well as Lip motion detection is performed on the audio source based at least in part on the plurality of images, wherein the second signal is also based on the lip motion detection.

4. The method of claim 1, wherein processing the multichannel audio signal comprises processing the multichannel audio signal to generate silent audio based at least in part on: whether the authenticated user is identified in the image, the position of the authenticated user relative to the video input device, the direction of the authenticated user's gaze, and / or whether lip movements of the authenticated user are detected.

5. The method of claim 1, wherein the first signal is a binary signal, and wherein the binary signal is in a first state at least in part based on the authentication of the user being identified as being in the image.

6. The method of claim 1, further comprising performing audio speech activity detection (VAD) on the multi-channel audio signal to generate a VAD signal, wherein the second signal is determined at least in part based on the VAD signal.

7. The method according to claim 1, further comprising: Determine the location of the authenticated user in the image; as well as The image is processed at least in part based on the location to generate an output video signal.

8. The method of claim 7 further includes transmitting the output audio signal and the output video signal to an external device via a network.

9. The method of claim 7, wherein processing the image includes blurring a portion of the image at least partially based on the location to generate the output video signal.

10. The method of claim 7, wherein if it is determined that the authenticated user is not in the image, the output video signal comprises a completely blurred image or a completely blanked image.

11. The method of claim 1, further comprising determining the direction of the gaze of the authenticated user based at least in part on the image, wherein the first signal and / or the second signal is also based on the direction of the gaze.

12. The method of claim 1, further comprising transmitting the output audio signal for use in a voice application via Internet Protocol.

13. The method of claim 12, further comprising setting the session of the Internet Protocol voice application to sleep mode based at least on the location of the authenticated user relative to the video input device.

14. A system comprising: A video subsystem configured to receive images captured by a video input device, the video subsystem comprising: A recognition component configured to: detect at least one face in the image; identify one of the at least one faces as an authenticated user based at least in part on a predefined facial identifier; determine a first signal based at least in part on the image, wherein the first signal indicates the probability of association with the authenticated user; and An audio subsystem configured to receive multi-channel audio signals including audio inputs detected by a plurality of audio input devices, the audio subsystem comprising: A logic component configured to determine a second signal based at least in part on the multi-channel audio signal and the first signal, wherein the second signal indicates the probability of association with an audio component attributed to the authenticated user; and An audio processing unit configured to process the multi-channel audio signal at least in part based on the second signal to generate an output audio signal, wherein the processing includes: Determine whether the authenticated user exists or is participating in the session; When the authenticated user does not exist or is not participating in the session, a silent audio is generated; When the authenticated user is present and participates in the session, the audio component attributed to the authenticated user is enhanced and the audio component associated with other speakers is removed; Determining whether the authenticated user exists or participates in the session includes: Whether a certified user participates in a session is determined based on historical data and / or the certified user's preferences regarding the certified user's behavior relative to the video input device.

15. The system of claim 14, wherein the video subsystem further comprises a video processing unit configured to process the image to generate an output video signal based at least in part on the location of the authenticated user in the image.

16. The system of claim 15, wherein the video processing component includes a background blurring component configured to blur a portion of the image at least partially based on the location to generate the output video signal.

17. The system according to claim 14, wherein: The recognition component is also configured to identify audio sources in multiple images as the authenticated user; The video subsystem further includes a lip motion detection component configured to perform lip motion detection on the audio source based at least in part on the plurality of images; as well as The second signal is also based on the lip movement detection.

18. The system of claim 14, wherein the audio subsystem further comprises an audio speech activity detection (VAD) component configured to perform VAD on the multi-channel audio signal to generate a VAD signal, wherein the second signal is determined at least in part based on the VAD signal.

19. The system of claim 14, wherein the audio processing unit is configured to process the multichannel audio signal to generate silent audio based at least in part on: whether the authenticated user is identified in the image, the position of the authenticated user relative to the video input device, the direction of the authenticated user's gaze, and / or whether lip movements of the authenticated user are detected.

Citation Information

Patent Citations

  • Microphone array voice reinforcing system and microphone array voice reinforcing method with combination of audio information and video information

    CN106328156A