Voice direction-based microphone control

By analyzing the audio signal spectrum captured by the microphone and applying machine learning model, it automatically identifies whether the user faces the microphone and unmutes it, solving the problem of delay caused by the user's unaware of the microphone mute, and improving the communication efficiency of telecommunications applications.

CN120544562APending Publication Date: 2025-08-26HEWLETT PACKARD DEVELOPMENT COMPANY LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510696400.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2018-12-17
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In telecommunications applications, the user's microphone is often muted and the user is not aware of it, resulting in delays or repetitions during speeches. The prior art is difficult to automatically identify whether the user faces the microphone in order to automatically unmute.

Method used

By analyzing the spectrum or frequency content of the audio signal captured by the microphone, combined with machine learning models such as a fully connected neural network (FCNN) or a convolutional neural network (CNN), determine whether the user is facing the microphone and automatically or prompt the user to unmute.

Benefits of technology

It realizes automatic or promptly demutes the microphone when the user speaks, reduces the repeated processing of audio capture and improves communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544562A_ABST
    Figure CN120544562A_ABST
Patent Text Reader

Abstract

The invention relates to microphone control based on voice direction. According to an example, an apparatus may include a processor and a non-transitory computer readable medium storing instructions thereon, the processor may execute the instructions to access an audio signal of a user's speech captured by a microphone when the microphone is in a mute state. The processor may also execute instructions to analyze spectral or frequency content of the accessed audio signal to determine whether the user faces the microphone as it speaks. In addition, based on a determination that the user is facing the microphone while the user speaks, the processor may execute the instructions to demute the microphone.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Telecommunications applications such as teleconferencing and videoconferencing applications can facilitate communications between multiple remotely located users to communicate with each other over an Internet Protocol network, over a land-based telephone network, and / or over a cellular network. In particular, telecommunications applications can enable audio to be captured locally for each user and transmitted to other users so that the user can hear the other users' voices via these networks. Some telecommunications applications can also enable still and / or video images of users to be captured locally and transmitted to other users so that the user can view the other users via these networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Features of the present disclosure are illustrated by way of example and not limitation in the following figure(s), in which like numerals indicate like elements, and in which:

[0003] Figure 1 A block diagram illustrating an example apparatus that can automatically control the unmuting of a microphone based on whether a user is likely to be facing the microphone when the user is speaking;

[0004] Figure 2A Shows that it can include Figure 1 A block diagram of an example system featuring the example apparatus depicted in;

[0005] Figure 2B An example process block diagram illustrating operations that may be performed during a training phase and an inference phase of a captured audio signal;

[0006] Figure 3 A block diagram illustrating an example apparatus that can automatically control the unmuting of a microphone based on whether a user is likely to be facing the microphone when the user is speaking;

[0007] Figure 4 and Figure 5 Depicted are example methods for automatically unmuting a microphone based on a determination as to whether a user is facing the microphone when the user is speaking; and

[0008] Figure 6 A block diagram is shown of an example non-transitory computer-readable medium that may have machine-readable instructions stored thereon that, when executed by a processor, may cause the processor to prompt a user to unmute a microphone based on a determination that the user may be facing the microphone when the user is speaking. DETAILED DESCRIPTION

[0009] For the purpose of simplicity and illustration, the principles of the present disclosure are described by primarily referring to the examples of the present disclosure. In the following description, many specific details are set forth to provide an understanding of the examples. However, it will be clear to those skilled in the art that the examples can be implemented without being limited to these specific details. In some cases, well-known methods and / or structures are not described in detail to avoid unnecessarily obscuring the description of the examples. In addition, the examples can be used together in various combinations.

[0010] Throughout this disclosure, the terms "a" and "an" are intended to mean one of a particular element or a plurality of a particular element. As used herein, the term "include" means including but not limited to, and the term "comprising" means including but not limited to. The term "based on" may mean based in part on.

[0011] When an audio conferencing application is activated, the microphone may start out muted. Often, users may not realize their microphone is muted and may begin speaking before unmuting their microphone. This can lead to confusion at the start of a conference call. This can also occur when a user intentionally mutes their microphone during an audio conference or other application and forgets to unmute it before speaking again.

[0012] Disclosed herein are apparatus, systems, and methods for automatically unmuting a microphone based on a determination that the user intended the user's voice to be captured. For example, a processor may determine whether the user is facing a muted microphone when the user speaks and automatically unmute the microphone based on the determination. The processor may make this determination by analyzing the spectrum or frequency content of the audio signal captured by the microphone. Additionally or alternatively, the processor may make this determination by applying a machine learning model to the captured audio signal. In some examples, the processor may implement voice activity detection techniques to determine whether the captured audio signal includes the user's voice. In some examples, the determination of whether the user is facing a muted microphone may be based on training a fully connected neural network (FCNN) or a convolutional neural network (CNN) to identify the directionality of speech.

[0013] In some examples, the second audio signal captured by the second microphone can be used to analyze the characteristics of the second audio signal captured by the second microphone to determine whether the user is likely facing the microphone and the second microphone when the user is speaking. In these examples, the processor can determine whether to unmute the microphone and the second microphone based on the determination of whether the user is facing the microphone discussed above and the determination based on the analysis of the characteristics of the audio signal and the second audio signal.

[0014] By implementing the devices, systems, and methods disclosed herein, a microphone can be automatically unmuted and / or a user can be prompted to unmute the microphone based on a determination that the user is facing a muted microphone when the user speaks. Thus, for example, the user's speech can be directed to an application for analysis, storage, translation, or the like. As another example, the user's speech can be directed to a communication interface for output during an audio conference. In any aspect, audio captured while the microphone is muted can be stored and used in an application and / or audio conference, which can reduce the additional processing that may be performed to capture, analyze, and store audio that may be repeated in the event that previously captured audio is lost or discarded.

[0015] First reference Figure 1 、 2A and 2B. Figure 1 A block diagram illustrating an example apparatus 100 that may automatically control the unmuting of a microphone based on whether a user is likely to be facing the microphone when the user is speaking. Figure 2A Shows that it can include Figure 1 1. Block diagram of an example system 200 illustrating features of the example apparatus 100 depicted in FIG. Figure 2B An example process block diagram 250 illustrates operations that may be performed during the training phase and the inference phase of the captured audio signal 222. It should be understood that Figure 1 、 2A The example apparatus 100, example system 200, and / or example process block diagram 250 depicted in and 2B may include additional components and some of the components described herein may be removed and / or modified without departing from the scope of the example apparatus 100, example system 200, and / or example process block diagram 250 disclosed herein.

[0016] Apparatus 100 may be a computing device or other electronic device, such as a personal computer, laptop computer, tablet computer, smartphone, or the like, that may facilitate automatic unmuting of microphone 204 based on a determination that user 220 is facing microphone 204 while user 220 is speaking. That is, apparatus 100 may capture an audio signal 222 of a user's voice while microphone 204 is muted, and may automatically unmute microphone 204 based on a determination that user 220 is facing microphone 204 while user 220 is speaking. Additionally, based on a determination that user 220 is facing microphone 204 while user 220 is speaking, apparatus 100 may store the captured audio signal 222, may activate a voice dictation application, may transmit the captured audio signal 222 to a remotely located system 240, for example, via network 230, and / or the like.

[0017] According to an example, the processor 102 may selectively transmit an audio signal of the captured audio 222, such as a data file including the audio signal, via the communication interface 208. The communication interface 208 may include software and / or hardware components through which the device 100 may transmit and / or receive data files. For example, the communication interface 208 may include a network interface of the device 100. The data file may include audio and / or video signals, such as data packets corresponding to the audio and / or video signals.

[0018] According to an example, the device 100, and more specifically, the processor 102 of the device 100, may determine whether the audio signal 222 includes audio that the user 220 intends to transmit to another user, for example, via execution of an audio or video conferencing application, and may transmit the audio signal based on the determination that the user 220 intends to transmit the audio to the other user. However, based on the determination that the user may not intend to transmit audio, the processor 102 may not transmit the audio signal. The processor 102 may determine the user's intention regarding whether to transmit audio in various ways as discussed herein.

[0019] like Figure 1 As shown in , the device 100 may include a processor 102 that can control the operation of the device 100. The processor 102 can be a semiconductor-based microprocessor, a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), and / or other hardware devices. The device 100 may also include a non-transitory computer-readable medium 110, which may have machine-readable instructions 112-118 (which may also be referred to as computer-readable instructions) stored thereon and executable by the processor 102. The non-transitory computer-readable medium 110 can be an electronic, magnetic, optical, or other physical storage device that includes or stores executable instructions. The non-transitory computer-readable medium 110 can be, for example, a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, and the like. The term "non-transitory" does not include transient propagation signals.

[0020] like Figure 2A As shown in FIG, system 200 may include Figure 1 1 and 10. The system 200 may also include a data store 202, a microphone 204, an output device (or multiple output devices) 206, and a communication interface 208. Electrical signals may be transmitted between some or all of the components 102, 110, 202-208 of the system 200 via a link 210, which may be a communication bus, a wire, and / or the like.

[0021] The processor 102 may execute or otherwise implement a telecommunications application to facilitate a telephone conference or video conference in which the user 220 may be a participant. The processor 102 may also or alternatively implement another type of application that can use and / or store the user's voice. In any respect, the microphone 204 may capture audio (or equivalently, sound, audio signal, etc.), and in some examples, the captured audio 222 may be transmitted over a network 230 via the communication interface 208. The network 230 may be an IP network, a telephone network, and / or a cellular network. In addition, the captured audio 222 may be transmitted across the network 230 to a remote system 240 so that the captured audio 222 can be output at the remote system 240. The captured audio 222 may be converted and / or stored in a data file, and the communication interface 208 may transmit the data file over the network 230.

[0022] In operation, microphone 204 can capture audio 222 and transmit the captured audio 222 to data storage 202 and / or processor 102. Additionally, microphone 204 or another component can convert captured audio 222 or store captured audio 222 in a data file. For example, captured audio 222 can be stored or encapsulated in an IP packet. In some examples, microphone 204 can capture audio signals 222 while microphone 204 is in a muted state. That is, when in a muted state, microphone 204 can continue to capture audio signals 222, and processor 102 can continue to process captured audio signals 222, but may not automatically send captured audio signals 222 to communication interface 208. When unmuted, microphone 204 can capture audio signals 222, and processor 102 can process captured audio 222 and send captured audio 222 to communication interface 208 for transmission via network 230.

[0023] The processor 102 may retrieve, decode, and execute the instructions 112 to access the audio signal 222 of the user's 220 speech captured by the microphone 204 while the microphone 204 is in a muted state. As discussed herein, when the microphone 204 is in a muted state, the microphone 204 may capture the audio signal 222 and may store the captured audio signal 222 in the data store 202. Thus, for example, the processor 102 may access the captured audio signal 222 from the data store 202.

[0024] Processor 102 may retrieve, decode, and execute instructions 114 to analyze the spectrum or frequency content of the accessed audio signal 222 to determine the direction in which user 220 was located when user 220 was speaking. For example, processor 102 may perform spectrum and / or frequency content analysis on the accessed audio signal 222 to determine whether user 220 was facing the microphone when user 220 was speaking. For example, when user 220 is facing away from microphone 204, the captured audio 222 may have lower intensity in the high frequency range due to high-frequency roll-off. By training a classifier using training data corresponding to speech samples from different directions, such as the direction of the user during the user's speech, and training data corresponding to speech samples from different users, the direction of user 220's speech may be classified as being toward microphone 204 or away from microphone 204, such as being toward one side of microphone 204. That is, the ML model can be trained using voice samples of a user facing the microphone and voice samples of a user not facing the microphone, and the ML model can capture the difference in the spectral and / or frequency content of the voice samples to be able to distinguish whether the captured audio signal 222 includes spectral and / or frequency content consistent with the voice of the user facing the microphone 204. This can be particularly useful when the ML model is deployed during inference during the start of a conference call on a Voice over IP (VoIP) system. The ML model can also be used to switch between muting and unmuting during a conference call.

[0025] The ML model may also use input from a voice activity detector (VAD), which can detect the presence or absence of human speech in an audio signal. The ML model may employ manually designed features on each frame, such as spectral roll-off above a threshold frequency, an average level in a measured spectrum, a difference spectrum across frames, and / or the like. Alternatively, a deep learning model employing a deep neural network (DNN) such as a convolutional neural network (CNN), a long short-term memory (LSTM) cascaded with a fully connected neural network (FCNN), or the like may be used to automatically extract deep features to train a machine learning model to classify between front-facing and side-facing (using head motion) profiles.

[0026] Processor 102 may retrieve, decode, and execute instructions 116 to unmute microphone 204 based on a determination that user 220 is facing microphone 204 while user 220 is speaking. Processor 102 may unmute microphone 204 based on a determination that user 220 is facing microphone 204 while the user is speaking, because user 220 facing microphone 204 while the user is speaking may be an indication that user 220 intends for the user's voice to be captured. In some situations, such as at the beginning of a conference call, microphone 204 may default to a muted state, and user 220 may begin speaking without first changing the microphone to an unmuted state. As a result, user 220 may need to repeat what user 220 said, which user 220 may find uneconomical. By implementing instructions 112-116, when user 220 faces microphone 204, the voice of user 220 captured when microphone 204 is muted may still be used when it is determined that the voice may have been uttered, which can enable user 220 to continue speaking without having to repeat the earlier voice.

[0027] In some examples, processor 102 may be remote from microphone 204. In these examples, processor 102 may access captured audio signal 222 from a remotely located electronic device that may be connected to microphone 204 via network 230. Additionally, processor 102 may also output an instruction to the remotely located electronic device to unmute microphone 204 via network 230. In response to receiving the instruction, the remotely located electronic device may unmute microphone 204.

[0028] The output device(s) 206 shown in the system 200 may include, for example, a speaker, a display, and the like. The output device(s) 206 may output, for example, audio received from the remote system 240. The output device(s) 206 may also output images and / or video received from the remote system 240.

[0029] Now go to Figure 2B, illustrates an example process block diagram 250 that processor 102 may perform to determine the direction of a user's speech. As shown, processor 102 may operate during a training phase 252 and an inference phase 254. Specifically, audio 222 captured by microphone 204 may be converted from an analog signal to a digital signal, and the digital signal may be filtered 260. During training phase 252, which may be, for example, a machine learning model training phase 252, feature extraction 262 may be applied to the converted and filtered signal. The extracted features of the converted and filtered signal may be used to generate a speaker model 264. Speaker model 264 may capture differences in the spectral and / or frequency content of the converted and filtered signal to enable differentiation between whether the captured audio signal 222 includes spectral and / or frequency content consistent with the speech of a user facing microphone 204 or includes spectral and / or frequency content consistent with the speech of a user not facing microphone 204. Thus, for example, multiple converted and filtered signals may be used to generate speaker model 264 during training phase 252.

[0030] During the inference phase 254, features of the transformed and filtered signal can be extracted 266. Additionally, a deployed directionality model 268 can be applied to the extracted features. Deployed directionality model 268 can be generated using speaker model 264 and can be used to determine the direction from which user 220 was speaking when audio 222 was captured. Based on the application of deployed directionality model 268, a determination 270 can be made regarding the direction 272 of the user's speech, e.g., whether the user was facing microphone 204 when audio 222 was captured. Additionally, direction 272 of the user's speech can be output, e.g., to control the operation of microphone 204. As discussed herein, direction 272 of the user's speech can be used to determine whether to unmute a muted microphone 204.

[0031] Now refer to Figure 1-3 . Figure 3 1 is a block diagram illustrating an example apparatus 300 that can automatically control the unmuting of the microphone 204 based on whether the user 220 is likely to be facing the microphone 204 when the user 220 is speaking. Figure 3 The example apparatus 300 depicted in FIG. 3 may include additional components and some of the components described herein may be removed and / or modified without departing from the scope of the example apparatus 300 disclosed herein.

[0032] The apparatus 300 may be similar to Figure 1, and thus may include a processor 302, which may be similar to the processor 102, and a non-transitory computer-readable medium 310, which may be similar to the non-transitory computer-readable medium 110. The computer-readable medium 310 may have machine-readable instructions 312-322 (also referred to as computer-readable instructions) stored thereon that may be executed by the processor 302.

[0033] The processor 302 can retrieve, decode, and execute instructions 312 to access the audio signal 222 of the user's 220 voice captured by the microphone 204. As discussed herein, the microphone 204 can capture the audio signal 222 while the microphone 204 is in a muted state. Additionally, the microphone 204 can capture the audio signal 222 and can store the captured audio signal 222 in the data store 202. Thus, for example, the processor 302 can access the captured audio signal 222 from the data store 202. In other examples where the processor 302 is remote from the microphone 204, the processor 302 can access the audio signal 222 via the network 230.

[0034] Processor 302 may retrieve, decode, and execute instructions 314 to determine whether user 220 was facing microphone 204 when user 220 was speaking, e.g., to generate captured audio 222. As discussed herein, processor 302 may perform spectral and / or frequency content analysis of accessed audio signal 222 to determine whether user 220 was facing microphone 204 when user 220 was speaking. Additionally or alternatively, processor 302 may apply a machine learning model, as discussed herein, to captured audio signal 222 to determine whether user 220 was likely facing microphone 204 when user 220 was speaking.

[0035] In some examples, processor 302 may determine whether microphone 204 is in a muted state when microphone 204 captures audio signal 222 of the voice of user 220. In these examples, processor 302 may determine, based on the determination that microphone 204 is in a muted state, whether user 220 is facing microphone 204 when user 220 is speaking. Additionally, when microphone 204 captures audio signal 222 of the voice of user 220, processor 302 may, based on the determination that microphone 204 is not in a muted state, e.g., in an unmuted state, output the captured audio signal 222 without analyzing the spectrum or frequency content of the captured audio signal 222.

[0036] Processor 302 may fetch, decode, and execute instructions 316 to unmute microphone 204 based on a determination that user 220 is facing microphone 204 while user 220 is speaking. Processor 302 may unmute microphone 204 based on a determination that user 220 is facing microphone 204 while the user is speaking because user 220 facing microphone 204 while the user is speaking may be an indication that user 220 intends for the user's voice to be captured.

[0037] The processor 302 can retrieve, decode, and execute instructions 318 to output the captured audio signal 222 based on a determination that the user 220 is facing the microphone 204 while the user 220 is speaking. For example, the processor 302 can output the captured audio signal 222 to the communication interface 208 so that the communication interface 208 can output the captured audio signal 222 to the remote system 240 via the network 230. Additionally or alternatively, the processor 302 can output the captured audio signal 222 to an application or device so that the captured audio signal 222 is stored, translated, or the like.

[0038] The processor 302 may retrieve, decode, and execute instructions 320 to maintain the microphone 204 in a muted state and / or discard the captured audio signal 222 based on a determination that the user 220 is not facing the microphone 204 when the user 220 is speaking. That is, for example, in addition to maintaining the microphone 204 in a muted state, the processor 302 may also not output the captured audio signal 222 based on a determination that the user 220 is not facing the microphone 204 when the user 220 is speaking.

[0039] Processor 302 may retrieve, decode, and execute instructions 322 to access second audio signal 224 of the user 220's voice captured by second microphone 226 when second microphone 226 is in a muted state, with second microphone 226 spaced apart from microphone 204. For example, second microphone 226 may be positioned at least several inches from microphone 204 such that, if user 220 is facing one side of one or both of microphone 204 and second microphone 226, sound waves may arrive at second microphone 226 at a different time than at microphone 204. By way of specific example, second microphone 226 may be located, for example, on one side of a laptop computing device and microphone 204 may be located on an opposite side of the laptop computing device.

[0040] The processor 302 may retrieve, decode, and execute instructions 314 to analyze characteristics of the audio signal 222 captured by the microphone 204 and the second audio signal 224 captured by the second microphone 226. For example, the processor 302 may determine the timing at which the microphone 204 captured the audio signal 222 and the timing at which the second microphone 226 captured the second audio signal 224. For example, the processor 302 may implement a time difference of arrival technique to detect the direction of the captured audio 222, 224.

[0041] Processor 302 may also retrieve, decode, and execute instructions 314 to determine, based on the analyzed characteristics, whether user 220 is facing microphone 204 and second microphone 226 when user 220 is speaking. For example, processor 302 may determine that user 220 is facing microphone 204 and second microphone 226 when user 220 is speaking based on a determination that microphone 204 captures audio signal 222 within a predefined time period during which second microphone 226 captures second audio 224. The predefined time period may be based on testing and / or training using various user voices. Additionally, processor 302 may determine that user 220 is not facing microphone 204 and second microphone 226 based on a determination that microphone 204 captures audio signal 222 outside of the predefined time period during which second microphone 226 captures second audio 224.

[0042] The processor 302 can obtain, decode and execute instructions 314 to further determine whether to unmute the microphone 304 and the second microphone 226 based on both the determination that the user 220 is facing the microphone 204 by analyzing the spectrum or frequency content of the accessed audio signal 222 and the determination that the user 220 is facing the microphone 204 and the second microphone 226 based on the analyzed characteristics.

[0043] about Figure 4 and Figure 5 The method 400 depicted in FIG. 4 discusses various ways in which the apparatus 100, 300 may be implemented in more detail. In particular, Figure 4 and Figure 5 Depicted are example methods 400 and 500, respectively, for automatically unmuting microphone 204 based on a determination as to whether user 200 is facing microphone 204 when user 200 is speaking. It will be apparent to one of ordinary skill in the art that example methods 400 and 500 may represent generalized illustrations and that other operations may be added or existing operations may be removed, modified, or rearranged without departing from the scope of methods 400 and 500.

[0044] For illustrative purposes, refer to Figure 1-3Methods 400 and 500 are described with reference to the apparatuses 100 and 300 shown in FIG. It should be understood that apparatuses having other configurations may be implemented to perform methods 400 and / or 500 without departing from the scope of methods 400 and / or 500.

[0045] At block 402, the processor 102, 302 may access the audio signal 222 of the voice of the user 220 captured by the microphone 204. At block 404, the processor 102, 302 may determine whether the microphone 204 was in a muted state when the microphone 204 captured the audio signal 222 of the voice of the user 220. Based on a determination that the microphone 204 was not in a muted state when the microphone 204 captured the audio signal 222 of the voice of the user 220, at block 406, the processor 102, 302 may output the captured audio signal 222. The processor 102, 302 may output the captured audio signal 222 in any manner described herein.

[0046] However, based on a determination that microphone 204 was muted when microphone 204 captured audio signal 222 of the user's voice, at block 408, processor 102, 302 may apply a machine learning model to the captured audio signal 222 to determine whether user 220 was likely facing microphone 204 when user 220 was speaking. Based on a determination that user 220 was likely facing microphone 204 when user 220 was speaking, processor 102, 302 may unmute microphone 204 at block 412. Additionally, processor 102, 302 may output the captured audio signal at block 406. However, based on a determination that user 220 was likely not facing microphone 204 when user 220 was speaking, processor 102, 302 may discard captured audio signal 414 at block 414.

[0047] Now go to Figure 5 At block 502, the processor 102, 302 may access audio signals captured by the microphone 204 and the second microphone 226. That is, the processor 102, 302 may access the audio signal captured by the microphone 204, such as an audio file containing the audio signal, and the second audio signal 224 captured by the second microphone 226. In some examples, the audio signals 222, 224 may be captured while each of the microphone 204 and the second microphone 226 is in a muted state. Additionally, as discussed herein, the second microphone 226 may be spaced apart from the microphone 204.

[0048] At block 504, the processor 102, 302 may analyze characteristics of the audio 222 captured by the microphone 204 and the second audio 224 captured by the second microphone 226. For example, the processor 102, 302 may analyze the captured audio 222, 224 to determine when to capture the audio signals 222, 224.

[0049] At block 506, the processor 102, 302 may determine based on the analyzed characteristics whether the user 220 is likely facing the microphone 204 and the second microphone 226 when the user 220 is speaking. For example, the processor 102, 302 may determine based on timing within a predefined time period that the user is likely facing the microphone 204 and the second microphone 226.

[0050] At block 508, the processor 102, 302 may determine whether to place the microphone 204 and the second microphone 226 in an unmuted state based on both a determination that the user 220 is facing the microphone 204 through application of the machine learning model and a determination that the user 220 is facing the microphone 204 and the second microphone 226 based on the analyzed characteristics. That is, the processor 102, 302 may determine that the user 220 is facing the microphone 204 and the second microphone 226 when the user 220 has been determined to have likely faced the microphone 204 while the user was speaking through application of the machine learning model and analysis of the audio signals 222 and 224. However, the processor 102, 302 may determine that the user 220 is not facing the microphone 204 or the second microphone 226 when the user 220 has not been determined to have likely faced the microphone 204 while the user was speaking through application of the machine learning model and analysis of the audio signals 222 and 224.

[0051] Based on a determination that the user 220 may be facing the microphone 204 and the second microphone 226 when the user 220 is speaking, the processor 102, 302 may unmute the microphone 204 and the second microphone 226 at block 510. However, based on a determination that the user 220 may not be facing the microphone 204 and the second microphone 226 when the user 220 is speaking, the processor 102, 302 may discard the captured audio signal 222 and the second captured audio signal 224 at block 512.

[0052] Some or all of the operations described in methods 400 and / or 500 may be included as utilities, programs, or subroutines in any desired computer-accessible medium. In addition, some or all of the operations described in methods 400 and / or 500 may be embodied by computer programs, which may exist in a variety of forms, both active and inactive. For example, they may exist as machine-readable instructions, including source code, object code, executable code, or other formats. Any of the foregoing may be embodied on a non-transitory computer-readable storage medium. Examples of non-transitory computer-readable storage media include computer system RAM, ROM, EPROM, EEPROM, and magnetic or optical disks or tapes. Therefore, it should be understood that any electronic device capable of performing the functions described above may perform those functions listed above.

[0053] Now go to Figure 6 , shows a block diagram of an example non-transitory computer-readable medium 600 that may have machine-readable instructions stored thereon that, when executed by a processor, may cause the processor to prompt the user 220 to unmute the microphone 204 based on a determination that the user 220 may be facing the microphone 204 when the user is speaking. It should be understood that Figure 6 The non-transitory computer readable medium 600 depicted in FIG may include additional instructions, and some of the instructions described herein may be removed and / or modified without departing from the scope of the non-transitory computer readable medium 600 disclosed herein. For illustrative purposes, reference is made to FIG. Figure 1-3 The non-transitory computer-readable medium 600 will be described with reference to the apparatuses 100 and 300 shown in FIG.

[0054] The non-transitory computer readable medium 600 may have stored thereon machine readable instructions 602-608, such as Figure 1 A processor such as processor 102 described in

[0044] can execute these instructions. Non-transitory computer-readable medium 600 can be an electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Non-transitory computer-readable medium 600 can be, for example, a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, and the like. The term "non-transitory" does not include transitory propagating signals.

[0055] The processor may retrieve, decode, and execute instructions 602 to access an audio file of the user's speech captured by microphone 204. The processor may retrieve, decode, and execute instructions 604 to determine whether microphone 204 was muted when microphone 204 captured the user's speech. The processor may retrieve, decode, and execute instructions 606 to apply a machine learning model to the captured user's speech to determine whether the user was likely facing microphone 204 when the user spoke, based on the determination that microphone 204 was muted when microphone 204 captured the user's speech. The machine learning model is generated using a classifier trained using training data corresponding to the user's orientation during the user's speech. Additionally, the processor may retrieve, decode, and execute instructions 608 to output an indication to user 220 to unmute microphone 204 based on the determination that user 220 was likely facing microphone 204 when the user 220 spoke.

[0056] although Figure 6 204 , the user 220 is facing the microphone 204, and the second microphone 226 is facing the microphone 204. In addition, or alternatively, the non-transitory computer-readable medium may further include instructions that cause the processor to access second audio of the user's voice captured by the second microphone 226 when the second microphone 226 is in a muted state, the second microphone 226 being spaced apart from the microphone 204, analyze characteristics of the audio captured by the microphone 204 and the second audio captured by the second microphone 226, determine whether the user 220 is facing the microphone 204 and the second microphone 226 when the user 220 is speaking based on the analyzed characteristics, and determine whether the microphone 204 is to be placed in an unmuted state or remain in a muted state based on both the determination that the user 220 is facing the microphone 204 by applying the machine learning model and the determination that the user 220 is facing the microphone 204 and the second microphone 226 based on the analyzed characteristics.

[0057] While described in detail throughout the entirety of this disclosure, representative examples of the disclosure have utility across a wide range of applications, and the above discussion is not intended to and should not be construed as limiting, but rather is provided as an illustrative discussion of various aspects of the disclosure.

[0058] What has been described and shown herein are examples of the present disclosure along with some variations thereof. The terms, descriptions, and figures used herein are set forth by way of illustration only and are not meant to be limiting. Many variations are possible within the scope of the present disclosure, the scope of which is intended to be defined by the following claims and their equivalents, in which all terms are to be interpreted in their broadest reasonable sense unless otherwise indicated.

Claims

1. A device comprising: processor; as well as A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to: accessing an audio signal of a user's voice captured by a microphone while the microphone is muted; analyzing the spectrum or frequency content of the accessed audio signal to determine the direction the user is facing when the user speaks; unmuting the microphone based on a determination that the user is facing the microphone while the user is speaking; as well as After the microphone is unmuted, the stored user's speech is used so that the user can continue speaking without having to repeat earlier speech. The user's voice stored therein is the user's voice captured when the microphone is muted.

2. The apparatus according to claim 1, further comprising: Communication interface; as well as Among other things, the instructions cause the processor to: Based on a determination that the user is facing the microphone while the user is speaking, the captured audio signal is output through the communication interface.

3. The device according to claim 1, wherein The instruction also causes the processor to: Based on a determination that the user was not facing the microphone when the user spoke, the microphone is maintained in a muted state.

4. The device according to claim 1, wherein The instruction also causes the processor to: Based on a determination that the user was not facing the microphone when the user spoke, the captured audio signal is discarded.

5. The device according to claim 1, wherein To access the audio signal captured by the microphone, the instructions also cause the processor to: receiving a data file including the captured audio signal from a remotely located electronic device via a network; and An instruction is output to the remotely located electronic device via the network to unmute the microphone.

6. The device according to claim 1, wherein The instruction also causes the processor to: accessing a second audio signal of a user's speech captured by a second microphone while the second microphone is in a muted state, the second microphone being spaced apart from the microphone; analyzing characteristics of an audio signal captured by the microphone and a second audio signal captured by a second microphone; determining, based on the analyzed characteristics, whether the user is facing the microphone and the second microphone when the user speaks; as well as Determining whether to unmute the microphone and the second microphone is based on both a determination that the user is facing the microphone by analyzing the spectrum or frequency content of the accessed audio signal and a determination that the user is facing the microphone and the second microphone based on the analyzed characteristics.

7. The device according to claim 1, wherein The instruction also causes the processor to: Determine if the microphone is muted; analyzing a spectrum or frequency content of the accessed audio signal to determine, based on a determination that the microphone is muted, whether the user is facing the microphone when the user speaks; as well as Based on a determination that the microphone is not in a muted state, the captured audio signal is output without analyzing the spectrum or frequency content of the captured audio signal.

8. The device according to claim 1, wherein The processor accesses audio captured by the microphone during the conference call.

9. A method comprising: accessing, by the processor, an audio signal of the user's voice captured by the microphone; determining, by the processor, whether the microphone is in a muted state when the microphone captures an audio signal of a user's voice; Based on a determination that the microphone was in a muted state when the microphone captured an audio signal of the user's voice, applying, by the processor, a machine learning model to the captured audio signal to determine whether the user was likely facing the microphone when the user spoke; placing, by the processor, the microphone in an unmuted state based on a determination that the user may be facing the microphone while the user is speaking; as well as After the microphone is unmuted, the stored user's voice is used so that the user can continue speaking without having to repeat earlier voices. The user's voice stored therein is the user's voice captured when the microphone is muted.

10. The method according to claim 9, further comprising: Based on a determination that the user is likely facing the microphone when the user speaks, a captured audio signal of the user's voice is output.

11. The method according to claim 9, further comprising: receiving a data file including the captured audio signal from a remotely located electronic device via a network; as well as An instruction is output to the remotely located electronic device via the network to unmute the microphone.

12. The method according to claim 9, further comprising: accessing second audio of a user's voice captured by a second microphone while the second microphone is in a muted state, the second microphone being spaced apart from the microphone; analyzing characteristics of the audio captured by the microphone and second audio captured by the second microphone; determining, based on the analyzed characteristics, whether the user is likely to be facing the microphone and the second microphone when the user speaks; as well as Based on both a determination that the user is facing the microphone by applying a machine learning model and a determination that the user is facing the microphone and the second microphone based on the analyzed characteristics, it is determined whether the microphone and the second microphone are to be placed in an unmuted state.

13. A non-transitory computer-readable medium having machine-readable instructions stored thereon that, when executed by a processor, cause the processor to: Accessing an audio file of the user's voice captured by a microphone; Determine whether the microphone is muted when it captures the user's voice; applying a machine learning model to the captured speech of the user to determine whether the user was likely facing the microphone when the user spoke, based on a determination that the microphone was muted when the microphone captured the user's speech, the machine learning model generated using a classifier trained using training data corresponding to the user's orientation during the user's speech; Based on a determination that the user is likely facing the microphone when the user speaks, outputting an instruction to the user to place the microphone in an unmuted state; and After the microphone is unmuted, the stored user's speech is used so that the user can continue speaking without having to repeat earlier speech. The user's voice stored therein is the user's voice captured when the microphone is muted.

14. The non-transitory computer-readable medium of claim 13, wherein: The instruction also causes the processor to: Based on a determination that the user is likely facing the microphone when the user speaks, the captured audio file is output through the communication interface.

15. The non-transitory computer-readable medium of claim 13, wherein: The instruction also causes the processor to: accessing second audio of a user's speech captured by a second microphone while the second microphone is in a muted state, the second microphone being spaced apart from the microphone; analyzing characteristics of the audio captured by the microphone and second audio captured by the second microphone; determining, based on the analyzed characteristics, whether the user is facing the microphone and the second microphone when the user speaks; as well as Based on both the determination that the user is facing the microphone through application of the machine learning model and the determination that the user is facing the microphone and the second microphone based on the analyzed characteristics, it is determined whether the microphone is to be placed in an unmuted state or to remain in a muted state.