Compensation for Facial Coverings in Captured Audio

By adjusting the speech frequency amplitude to compensate for the attenuation of the facial covering, the speech attenuation problem caused by the facial covering is solved and the voice comprehension effect is improved.

CN115331685BActive Publication Date: 2025-07-18AVAYA MANAGEMENT LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210436522.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-26
Filing Date
2022-04-25
Publication Date
2025-07-18
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

Facial coverings lead to attenuation of speech frequency, which is difficult to effectively compensate for in the prior art, making speech difficult to understand.

Method used

The speech frequency amplitude is adjusted by compensator, and the affected frequency is selectively amplified to compensate for the attenuation according to the type and attenuation of the face covering.

Benefits of technology

Improves the voice comprehension effect when wearing facial coverings, making the voice sound more natural and close to the level when there is no facial covering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331685B_ABST
    Figure CN115331685B_ABST
Patent Text Reader

Abstract

The present invention discloses compensation for a facial covering in captured audio. The techniques disclosed herein enable compensation for attenuation caused by a facial covering in captured audio. In a particular embodiment, a method includes determining that a facial covering is positioned to cover the mouth of a user of a user system. The method further includes receiving audio including speech from the user and adjusting the amplitude of frequencies in the audio to compensate for the facial covering.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Worldwide, facial coverings, such as masks positioned over people's mouths, have been widely used during the global pandemic to prevent the spread of viruses and other infections. During normal (non-pandemic) times, facial coverings are still used in many situations to protect individuals and others. For example, facial coverings are common in medical settings and in other workplaces to prevent harmful airborne contamination (e.g., harmful dust particles). Facial coverings tend to block portions of the audio spoken by the wearer, making them more difficult to understand. The blocked components of the speech are not linear and cannot be simply restored by increasing the speech level using common means such as speaking louder, turning up the volume of a voice or video call, or getting closer in a face-to-face conversation. Summary of the Invention

[0002] The techniques disclosed herein enable compensation for attenuation caused by a facial covering in captured audio. In a particular embodiment, a method includes determining that a facial covering is positioned to cover a mouth of a user of a user system. The method further includes receiving audio including speech from the user and adjusting an amplitude of frequencies in the audio to compensate for the facial covering.

[0003] In some embodiments, the method includes transmitting the audio via a communication session between the user system and another user system after adjusting the frequencies.

[0004] In some embodiments, adjusting the amplitude of the frequencies includes amplifying the frequencies based on attenuation of the frequencies caused by the facial covering. The attenuation may indicate that a first set of frequencies should be amplified by a first amount and a second set of frequencies should be amplified by a second amount.

[0005] In some embodiments, the method includes receiving reference audio including reference speech from the user when the mouth is not covered by the facial covering. In those embodiments, the method may include comparing the reference audio with the audio to determine an amount by which the frequencies have been attenuated by the facial covering. Similarly, in those embodiments, the method may include receiving training audio including training speech from the user when the mouth is covered by the facial covering, where the training speech and the reference speech include words spoken by the user from the same script, and comparing the reference audio with the training audio to determine an amount by which the frequencies have been attenuated by the facial covering.

[0006] In some embodiments, determining that the facial covering is positioned to cover the user's mouth includes receiving a video of the user and using facial recognition to determine that the mouth is covered.

[0007] In some embodiments, adjusting the amplitude of the frequency includes accessing a profile for the facial covering, the profile indicating the amount by which the frequency and the amplitude should be adjusted.

[0008] In some embodiments, the method includes receiving a video of the user and replacing the facial covering in the video with a synthetic mouth of the user.

[0009] In another embodiment, there is provided an apparatus having one or more computer-readable storage media and a processing system operatively coupled to the one or more computer-readable storage media. Program instructions stored on the one or more computer-readable storage media, when read and executed by the processing system, cause the processing system to determine that a facial covering is positioned to cover the mouth of a user of a user system. The program instructions also cause the processing system to receive audio including speech from the user and adjust the amplitude of a frequency in the audio to compensate for the facial covering. Description of the Drawings

[0010] Figure 1 Illustrates an implementation for compensating for a facial covering in captured audio.

[0011] Figure 2 Illustrates an operation for compensating for a facial covering in captured audio.

[0012] Figure 3 Illustrates an operation scenario for compensating for a facial covering in captured audio.

[0013] Figure 4 Illustrates an implementation for compensating for a facial covering in captured audio.

[0014] Figure 5 Illustrates an operation scenario for compensating for a facial covering in captured audio.

[0015] Figure 6 Illustrates a voice frequency spectrogram for compensating for a facial covering in captured audio.

[0016] Figure 7 Illustrates an operation scenario for compensating for a facial covering in captured video.

[0017] Figure 8Illustrated is a computing architecture for compensating for a facial covering in captured audio. Detailed Description

[0018] The examples provided herein enable compensation for the effects of wearing a facial covering (e.g., a face mask, a shield, etc.) when speaking to a user system. Since the effects of the facial covering are non-linear (i.e., all vocal frequencies are not affected by the same amount), simply increasing the volume of the speech captured from a user wearing a facial covering will not account for those effects. Instead, the amplitudes of the frequencies in the speech will increase across the board, even for frequencies in the speech that are not affected (or are negligibly affected) by the facial covering. The compensation described below accounts for the non-linear effects by selectively amplifying the frequencies in the speech based on the degree to which the corresponding frequencies are affected by the facial covering. Advantageously, frequencies that are not affected by the facial covering will not be amplified, while affected frequencies will be amplified by an amount corresponding to the degree to which those frequencies are attenuated by the facial covering.

[0019] Figure 1 Illustrated is implementation 100 for compensating for a facial covering in captured audio. Implementation 100 includes user system 101 having compensator 121 and microphone 122. User system 101 is operated by user 141. User system 101 can be a telephone, a tablet computer, a laptop computer, a desktop computer, a conferencing system, or some other type of computing system. Compensator 121 can be implemented as software instructions executed by user system 101 (e.g., can be a component of a communication client application or other application that captures audio) or as a hardware processing circuit. Microphone 122 captures sound and provides an audio signal representing the sound to user system 101. Microphone 122 can be incorporated into user system 101, can be connected to user system 101 via a wired connection, or can be connected to user system 101 via a wireless connection. In some examples, compensator 121 can be incorporated into microphone 122 or can be connected in the communication path of the audio between microphone 122 and user system 101.

[0020] Figure 2Illustrated is operation 200 for compensating for a facial covering in captured audio. In this example, operation 200 is performed by compensator 121 of user system 101. In other examples, operation 200 may be performed by a compensator in a system remote from user system 101, such as compensator in communication session system 401 in implementation 400 below. In operation 200, compensator 121 determines that a facial covering, in this case facial covering 131, is positioned to cover the mouth of user 141 (201). Facial covering 131 may be a face mask, a face shield, or other type of covering that is intended to prevent particulate matter from being exhaled from the mouth into the surrounding air or inhaled from the surrounding air when positioned to cover the mouth of user 141 (and typically the nose of user 141). By covering their mouth with facial covering 131, user 141 has placed a material (e.g., cloth, paper, plastic in the case of a face shield, or other type of facial covering material) through which sound generated by the voice of user 141 will travel between their mouth and microphone 122.

[0021] Compensator 121 may determine that facial covering 131 is specifically positioned over the mouth of user 141 (as opposed to another facial covering), may determine the type of facial covering 131 (e.g., cloth face mask, paper face mask, plastic face shield, etc.) that is positioned over the mouth of user 141, or may simply determine that a facial covering is positioned over the mouth of user 141 without additional detail. Compensator 121 may receive an input from user 141 indicating that facial covering 131 is being worn, may process captured video of user 141 to determine that the mouth of user 141 is covered by facial covering 131 (e.g., may use a facial recognition algorithm to identify that the mouth of user 141 is covered), may identify a specific attenuation pattern in the audio of user 141 who is speaking that indicates the presence of a facial covering, or may determine in some other way that a facial covering is positioned over the mouth of user 141.

[0022] Compensator 121 receives audio 111 that includes speech from user 141 (202). Audio 111 is received from microphone 122 after being captured by microphone 122. Audio 111 may be audio for transmission in a communication session between user system 101 and another communication system (e.g., another user system operated by another user), may be audio for recording in a memory of user system 101 or elsewhere (e.g., a cloud storage system), or may be audio captured from user 141 for some other reason.

[0023] Since compensator 121 determines that face covering 131 covers the mouth of user 141, compensator 121 adjusts the amplitude of the frequencies in audio 111 to compensate for face covering 131 (203). The presence of face covering 131 between the mouth of user 141 and microphone 122 attenuates the amplitude of at least a portion of the frequencies in the sound generated by the voice of user 141 as the sound passes through face covering 131. As such, audio 111 representing the sound captured by microphone 122 has an amplitude of the corresponding frequencies that is attenuated relative to the amplitude if user 141 were not wearing the face mask. Compensator 121 adjusts the respective amplitudes of the affected frequencies to the levels (or at least close to those levels) at which the amplitudes would be if user 141 were not wearing face covering 131. Compensator 121 can operate on an analog version of audio 111 or a digitized version of audio 111. Compensator 121 can adjust the amplitude in a manner similar to the way an audio equalizer adjusts the power (i.e., amplitude) of frequencies in audio.

[0024] In some examples, the amount by which certain frequencies should be adjusted can be predefined within compensator 121. In those examples, the predefined adjustment amount can be based on a "one size fits all" or "best fit" philosophy, where the adjustment is predefined to account for the attenuation caused by many different types of face coverings (e.g., cloth, paper, plastic, etc.). For example, if a set of frequencies typically attenuates an amount of amplitude that depends on a range of face covering materials, then the predefined adjustment can define an amount that is in the middle of that range. In some examples, if compensator 121 determines the specific type of face covering 131 above, then the predefined adjustment can include an amount for the specific type of face covering. For example, depending on the type of face covering 131, the amount by which the amplitude of a set of frequencies is adjusted can vary from the predefined amount.

[0025] In other examples, compensator 121 can be trained to identify the amount by which the magnitude of a frequency is attenuated such that those frequencies can be amplified by a proportional amount to return the speech of user 141 to a level similar to the level if no face covering 131 was present. Compensator 121 can be specifically trained to account for face covering 131, can be trained to account for a specific type of face covering (e.g., trained for cloth, paper, etc.), can be trained to account for any type of face covering (e.g., the one-size-fits-all approach discussed above), can be trained to account for different types of face coverings depending on determining what user 141 is wearing (e.g., if user 141's face covering 131 is cloth, then trained to account for a cloth face mask, and if user 141 wears a paper face mask at a different time, then trained to account for a paper face mask), can be specifically trained to account for the speech of user 141, can be trained to account for multiple user voices, and / or can be trained in some other manner. In some cases, compensator 121 can analyze the speech in the audio from user 141 when no face covering is present on user 141's mouth to learn over time what to expect from the speech level of user 141 (i.e., the magnitude at the corresponding frequencies). Regardless of what type of face covering 131 ultimately is, compensator 121 can simply amplify the frequencies in audio 111 to a level corresponding to the level that compensator 121 has learned to expect. In some cases, compensator 121 can be able to identify the presence of face covering 131 in the above step based on comparing the levels in audio 111 to the levels that compensator 121 expects from a user 141 without a face mask.

[0026] Advantageously, adjusting the magnitude of the attenuated frequencies in audio 111 to a level close to what would be expected if face covering 131 did not cover user 141's mouth will make the speech from user 141 more intelligible when user 141 is wearing face covering 131. Thus, when played back by user system 101 or some other system (e.g., the other endpoint on a communication session), the speech of user 141 will be more intelligible than if no adjustment had ever been performed, even if the voice of user 141 does not sound exactly like the voice if user 141 was not wearing face covering 131.

[0027] Figure 3Illustrated is an operational scenario 300 for compensating for a facial covering in captured audio. Operational scenario 300 is an example of how compensator 121 can be explicitly trained to compensate for user 141 wearing facial covering 131 to cover their mouth. In this example, compensator 121 receives reference audio 301 from user 141 via microphone 122 at step 1 when user 141 is not wearing any kind of facial covering. Reference audio 301 includes speech from user 141 in which user 141 speaks the script of the words. Compensator 121 can provide the script to user 141 (e.g., direct user system 101 to display the words in the script to user 141) or user 141 can use their own. Compensator 121 then receives training audio 302 from user 141 via microphone 122 at step 2 when user 141 is wearing facial covering 131 to cover their mouth. Training audio 302 includes speech from user 141 in which user 141 speaks the same script of the words as for reference audio 301. Compensator 121 can also direct user 141 to speak the words from the script in the same manner (or as close to the same manner as possible) (e.g., same volume, rhythm, pace, etc.) as user 141 spoke the words to produce reference audio 301 to minimize the number of variables between reference audio 301 and training audio 302 other than the presence of facial covering 131 in training audio 302 and its absence in reference audio 301. Preferably, the script includes words that will capture the full voice frequency range of user 141. Although this example has the receipt of training audio 302 occur after the receipt of reference audio 301, in other examples reference audio 301 can be received after training audio 302.

[0028] Compensator 121 compares reference audio 301 with training audio 302 at step 3 to determine how much the frequencies of user 141's voice are attenuated due to facial covering 131 in training audio 302. Since reference audio 301 and training audio 302 include speech using the same script, the frequencies included therein should have been spoken by user 141 at a similar amplitude. Thus, the difference in amplitude (i.e., attenuation) between the frequencies in reference audio 301 and the corresponding frequencies in training audio 302 can be considered to be caused by facial covering 131. Compensator 121 then uses the difference in amplitude at least over the range of frequencies typical for human speech (e.g., approximately 125 Hz to 8000 Hz) at step 4 to create a profile that user 141 can enable when wearing facial covering 131. This profile instructs compensator 121 as to the frequencies and the amount by which those frequencies should be amplified in order to compensate for user 141 wearing facial covering 131 in subsequently received audio (e.g., audio 111).

[0029] In some examples, user 141 may similarly train compensator 121 while wearing different types of face coverings over their mouth. A separate profile associated with user 141 may be created for each type of face covering. Compensator 121 may then load or otherwise access the appropriate profile for the face covering worn by user 141 after determining the type of face covering being worn. For example, user 141 may indicate that they are wearing a cloth face mask, and in response, compensator 121 loads the profile for user 141 wearing a cloth face mask. In some examples, the face covering profiles generated for user 141 may be stored in a cloud storage system. Even if user 141 is operating a user system other than user system 101, that other user system may load the profile from the cloud to compensate for user 141 wearing the face covering corresponding to the profile.

[0030] Figure 4 Illustrated is an implementation 400 for compensating for a face covering in captured audio. Implementation 400 includes a communication session system 401, user systems 402-405, and a communication network 406. Communication network 406 includes one or more local and / or wide area computing networks, including the Internet, through which communication session system 401 and user systems 402-405 communicate. User systems 402-405 may each include a telephone, laptop computer, desktop workstation, tablet computer, conference room system, or some other type of user-operable computing device. Communication session system 401 may be an audio / video conferencing server, a packet telephony server, a web-based presentation server, or some other type of computing system that facilitates a user communication session between endpoints. User systems 402-405 may each execute a client application that enables user systems 402-405 to connect to communication session system 401 and join a communication session facilitated by communication session system 401.

[0031] In operation, a real-time communication session is established between user systems 402 - 405 operated by respective users 422 - 425. The communication session enables users 422 - 425 to talk to each other in real time via their respective endpoints, i.e., user systems 402 - 405. The communication session system 401 includes a compensator that determines when a user is wearing a facial covering and adjusts the audio received from the user over the communication session to compensate for attenuation caused by the facial covering. The adjusted audio is then sent to the other people on the communication session. In this example, only user 422 wears a facial covering. Thus, only the audio of user 422 from user system 402 is adjusted by the communication network 406 before being sent to user systems 403 - 405 for playback to users 423 - 425, as described below. In other examples, one or more of users 423 - 425 may also wear facial coverings, and the communication session system 401 may similarly adjust the audio of those users that is received.

[0032] Figure 5 An operational scenario 500 for compensating for a facial covering in captured audio is illustrated. In operational scenario 500, user system 402 captures user communication 501 in step 1 for inclusion on the communication session. User communication 501 includes at least the audio of the speaking user 422 that is captured, but may also include other forms of user communication, such as a screen capture video of the display of user system 402 and / or a video of user 422 captured concurrently with the audio. User system 402 transmits user communication 501 to communication session system 401 in step 2 for distribution over the communication session to user systems 403 - 405.

[0033] The communication session system 401 identifies in step 3 that the user 422 is wearing a face covering 431 when generating user communication 501 (i.e., when speaking). The communication session system 401 can identify that the user 422 is wearing a face covering 431 from analyzing the user communication 501. For example, the communication session system 401 can determine that the amplitude of the frequencies in the audio of the user communication 501 indicates that a face covering is being worn, or, if the user communication 501 includes video of the user 422, then the communication session system 401 can use a facial recognition algorithm to determine that the mouth of the user 422 is covered by the face covering 431. In an alternative example, the user system 402 can provide an indication to the communication session system 401 that the user 422 is wearing a face covering 431 outside of the user communication 501. For example, the user interface of a client application executed on the user system 402 can include a toggle for the user 422 to indicate that a face covering 431 is being worn. The user can indicate or the communication session system 401 can otherwise identify specifically which face covering 431 is being worn, what type of face covering 431 (e.g., cloth face mask, paper face mask, face shield, etc.) is being worn, or that a face covering is being worn regardless of type.

[0034] In this example, the communication session system 401 stores a profile of the face covering associated with the user. The profile can be generated by the communication session system 401 performing a training process similar to the training process described in the operation scenario 300, or can be received from the user system performing a training process like the training process described in the operation scenario 300. The communication session system 401 loads in step 4 the profile associated with the user 422 for the face covering 431. The profile can be specifically for the face covering 431 or can be a profile for the type of face covering of which the face covering 431 is a member, depending on the level of specificity of the communication session system 401's identification of the face covering 431 in step 3 or the level of specificity of the profile stored for the user 422 (e.g., a profile can be stored for a specific face mask or for a face mask type). If no profile exists for a specific face covering 431, then the communication session system 401 can determine whether a profile exists for a face covering of the same type as the face covering 431. If no profile still exists (e.g., the user 422 may not have been trained for that type of face covering), then the communication session system 401 can use a default profile for that type of face covering or for face coverings in general. Although the default profile is not customized specifically for the attenuation caused by the user 422's face covering, using the default profile anyway to adjust the audio in the user communication 501 will most likely result in improved speech understanding during playback.

[0035] Communication session system 401 adjusts the audio in user communication 501 according to a configuration file in step 5. In particular, the configuration file indicates the amount by which the amplitude of the corresponding frequencies in the audio should be amplified and communication session system 401 performs those amplifications substantially in real time to minimize the latency of user communication 501 over the communication session. After adjusting the audio, communication session system 401 transmits user communication 501 to each of user systems 403 - 405 in step 6. Upon receiving user communication 501, each of user systems 403 - 405 plays the audio in user communication 501 to the corresponding user 423 - 425. When each of users 423 - 425 hears the played audio, due to the adjustments made by communication session system 401, the audio should sound more like user 422 is not speaking through face covering 431 to them.

[0036] In some examples, step 3 may be performed once and the configuration file determined in step 4 may be used for the remainder of the communication session. In other examples, communication session system 401 may later determine in the communication session that user 422 is no longer wearing a face covering (e.g., an input indicating that face covering 431 has been removed may be received from user 422 or the face covering 431 may no longer be detectable in a captured video of user 422). In those examples, communication session system 401 may stop adjusting the audio in user communication 501 because there is no longer a face covering to compensate for. Similarly, if communication session system 401 identifies that user 422 has re - donned a face covering, face covering 431 or otherwise, then communication session system 401 may then reload the configuration file for that face covering and start adjusting the audio again.

[0037] Figure 6 A voice frequency spectrogram 600 for compensating for a face covering in captured audio is illustrated. Spectrogram 600 is a graph of the amplitude in decibels (dB) of frequencies in hertz (Hz) for a frequency range common to human speech. Spectrogram 600 includes a line representing reference audio 621 and a line representing training audio 622. Reference audio 621 is similar to reference audio 301 from above in that reference audio 621 includes speech received from a user when the user is not wearing a face covering. Similarly, training audio 622 is similar to training audio 302 from above in that training audio 622 includes speech received from a user when the user is wearing a face covering. As is clearly visible from spectrogram 600, the amplitude in training audio 622 is almost entirely lower compared to the amplitude in reference audio 621, and the amount by which the amplitude is lower varies non - linearly with respect to frequency.

[0038] The difference between the reference audio 621 and the training audio 622 at any same frequency can be used to indicate the amount by which the audio should be adjusted at the corresponding frequency when an audio like the training audio 622 is received while the user is wearing the facial covering. For example, based on the information shown in the spectrogram 600, at 4200 Hz, the amplitude of the received audio should be increased by approximately 7 dB, while at 2000 Hz no amplification is required (i.e., the reference audio 621 and the training audio 622 overlap at this point). In some examples, rather than tracking amplitude adjustments for each possible frequency in the speech range as might seem possible from the continuous lines representing the reference audio 621 and the training audio 622 on the spectrogram 600, the amount of adjustment can be divided into frequency groups, each frequency group including a range of frequencies. These groups can have varying sizes or can have a consistent size (e.g., 100 Hz) based on having a similar amount of amplitude adjustment for the frequency range. In an example of varying frequency ranges, one range can be 2000 - 2200 Hz, which corresponds to no change in amplitude, while another range can be 4000 - 4600 Hz, which corresponds to a 7 dB change in amplitude, representing the most suitable change over all frequencies in that range as can be visualized on the spectrogram 600 and determined by the most suitable algorithm of the compensator. Other ranges with corresponding changes in amplitude will also correspond to the remainder of the speech frequency spectrum. In additional examples, the frequency group to be adjusted can simply be that all frequencies above a given frequency should be adjusted. For example, based on the spectrogram 600, the compensator can determine that all frequencies above 3400 Hz should be amplified by 5 dB, while frequencies below 3400 Hz should remain as they are. Adjusting frequencies in this way can work well for a default profile where more specific adjustments have not been determined for a particular user and facial covering combination.

[0039] Figure 7Illustrated is an operational scenario 700 for compensating for a facial covering in a captured video. Operational scenario 700 involves a user system 701, which is an example of the user system 101 from above. A compensator similar to compensator 121 can direct the user system 701 to perform the steps discussed below, or alternatively some other hardware / software element of the user system 701 can direct the user system 701. In this example, user 741 is operating the user system 701 in a real-time video communication session with one or more other endpoints, and in step 1 captures a video 721 that includes a video image of user 741. In this example, user 741 is wearing a facial covering 731 in video 721 and the user system 701 identifies that fact in step 2. The user system 701 can identify the facial covering 731 by processing the video 721 (e.g., using facial recognition), or can identify in some other way (such as the way described in the example above) that user 741 is wearing a facial covering 731.

[0040] After detecting the facial covering 731, the user system 701 edits the video 721 in step 3 to remove the facial covering 731 and replaces the facial covering 731 with a synthetic version of the mouth, nose, cheeks, and any other elements of user 741 that were covered by the facial covering 731. The algorithm used to perform the editing can be pre-trained using videos of user 741 without the facial covering, which allows the algorithm to learn what user 741 looks like under the facial covering 731. The algorithm then replaces the facial covering 731 in the image of video 721 with a synthetic version of the covered portion of user 741's face that the algorithm has learned. In some examples, the algorithm can also be trained to synthesize mouth / facial movements consistent with user 741 speaking specific words, such that user 741 appears in video 721 to be speaking in correspondence with the audio of user 741 actually speaking during the communication session that was captured (e.g., the audio captured and adjusted in the example above). Similarly, the algorithm can be trained to make the synthetic portion of user 741's face express emotions in combination with the expressions made by the portion of user 741's face that can be seen outside of the facial covering 731. In other examples, if the algorithm has not been specifically trained for user 741, the algorithm can be able to estimate what the covered portion of user 741's face looks like based on other people used to train the algorithm and based on what the algorithm can see in video 721 (e.g., skin color, hair color, etc.).

[0041] After editing video 721 to replace face covering 731, video 721 is transmitted via a communication session in step 4. Preferably, the above steps occur substantially in real time to reduce latency on the communication session. In any case, when played at the receiving endpoint, video 721 includes a video image of user 741, while face covering 731 is not visible, and instead a synthetic version of the portion of user 741's face that was covered by face covering 731. Although in this example video 721 is transmitted from user system 701, in other examples video 721 can be used for other purposes, such as posting on a video sharing service or simply saving to memory. Additionally, one or more of the remaining steps can be performed elsewhere, such as at the communication session system, rather than on user system 701 itself, when video 721 is captured at user system 701. In the scenario of adjusting audio according to the above example and editing video according to operating scenario 700, it should appear to the user watching video 721 and listening to the corresponding audio that user 741 is not wearing face covering 731. In some examples, operating scenario 700 can occur to compensate for face covering 731 in the video without simultaneously compensating the corresponding audio.

[0042] Figure 8 Computing architecture 800 is illustrated for compensating for a face covering in captured audio. Computing architecture 800 is an example computing architecture for user systems 101, 402 - 405, 701, and communication session system 401, but those systems can use alternative configurations. Computing architecture 800 includes communication interface 801, user interface 802, and processing system 803. Processing system 803 is linked to communication interface 801 and user interface 802. Processing system 803 includes processing circuitry 805 and memory device 806 that stores operating software 807.

[0043] Communication interface 801 includes components for communicating via a communication link, such as a network card, port, RF transceiver, processing circuitry and software, or some other communication device. Communication interface 801 can be configured to communicate via a metallic, wireless, or optical link. Communication interface 801 can be configured to use TDM, IP, Ethernet, optical networks, wireless protocols, communication signaling, or some other communication format - including combinations thereof.

[0044] User interface 802 includes components for interacting with a user. User interface 802 can include a keyboard, display screen, mouse, touchpad, or some other user input / output device. In some examples, user interface 802 can be omitted.

[0045] Processing circuit 805 includes a microprocessor and other circuitry that retrieves and executes operating software 807 from memory device 806. Memory device 806 includes a computer-readable storage medium, such as a disk drive, flash drive, data storage circuit, or some other memory device. In any example, the storage medium of memory device 806 will not be considered a propagated signal. Operating software 807 includes a computer program, firmware, or some other form of machine-readable processing instructions. Operating software 807 includes compensation module 808. Operating software 807 may also include an operating system, utilities, drivers, network interfaces, applications, or some other type of software. When executed by processing circuit 805, operating software 807 guides processing system 803 to operate the computing architecture 800 as described herein.

[0046] In particular, compensation module 808 guides processing system 803 to determine that a face covering is positioned to cover the mouth of a user of the user system. Compensation module 808 also guides processing system 803 to receive audio including speech from the user and adjust the amplitude of frequencies in the audio to compensate for the face covering.

[0047] The description and drawings included herein depict specific implementations of the claimed invention. For purposes of teaching the principles of the invention, some conventional aspects have been simplified or omitted. Additionally, some variations of these implementations that fall within the scope of the invention may be apparent. It will also be appreciated that the above features may be combined in various ways to form multiple implementations. As a result, the invention is not limited to the above specific implementations, but is only limited by the claims and their equivalents.

Claims

1. A method, comprising: Establishing a communication session between a user system and another user system; Determining that a facial covering is positioned to cover the mouth of a user of the user system; Receiving, via the communication session, video of the user and audio including speech from the user; And Replacing the facial covering in the video with a synthetic mouth of the user, wherein movement of the synthetic mouth is consistent with words in the speech; Adjusting an amplitude of frequencies in the audio to compensate for the facial covering; and After adjusting the amplitude of the frequencies, transmitting the video and the audio via the communication session to the another user system.

2. The method of claim 1, comprising: Determining a facial covering type corresponding to the facial covering, wherein adjusting the amplitude of the frequencies is based on the facial covering type.

3. The method of claim 1, wherein adjusting the amplitude of the frequencies comprises: Amplifying the frequencies based on attenuation of the frequencies caused by the facial covering, wherein the attenuation indicates that a first set of frequencies should be amplified by a first amount and a second set of frequencies should be amplified by a second amount.

4. The method of claim 1, comprising: Receiving reference audio including reference speech from the user when the mouth is not covered by the facial covering; And Comparing the reference audio with the audio to determine an amount by which the frequencies have been attenuated by the facial covering.

5. The method of claim 4, comprising: Receiving training audio including training speech from the user when the mouth is covered by the facial covering, wherein the training speech and the reference speech comprise words spoken by the user from the same script.

6. An apparatus, comprising: One or more computer-readable storage media; A processing system operatively coupled to the one or more computer-readable storage media; And Program instructions stored on the one or more computer-readable storage media, which when read and executed by the processing system cause the processing system to: Establish a communication session between a user system and another user system; Determine that a facial covering is positioned to cover the mouth of a user of the user system; Receive, via the communication session, video of the user and audio including speech from the user; Replace the facial covering in the video with a synthetic mouth of the user, Wherein movement of the synthetic mouth is consistent with words in the speech; Adjust an amplitude of frequencies in the audio to compensate for the facial covering; and After the amplitude of the frequencies is adjusted, transmit the video and the audio via the communication session to the another user system.

7. The apparatus of claim 6, wherein the program instructions cause the processing system to: Determine a facial covering type corresponding to the facial covering, wherein adjusting the amplitude of the frequencies is based on the facial covering type.

8. The apparatus of claim 6, wherein, to adjust the amplitude of the frequencies, the program instructions cause the processing system to: Amplify the frequency based on the attenuation of the frequency caused by the facial covering, where the attenuation indicates that a first set of frequencies should be amplified by a first amount and a second set of frequencies should be amplified by a second amount.

9. The apparatus of claim 6, wherein the program instructions direct the processing system to: Receive reference audio including reference speech from the user when the mouth is not covered by the facial covering; and Compare the reference audio with the audio to determine the amount by which the frequency has been attenuated by the facial covering.

10. The apparatus of claim 9, wherein the program instructions direct the processing system to: Receive training audio including training speech from the user when the mouth is covered by the facial covering, where the training speech and the reference speech include words spoken by the user from the same script.

Citation Information

Patent Citations

  • Respirator mask speech enhancement apparatus and method

    CN104955526A

  • Voice input device and telephone set

    JP2015135358A