De-reverberation of audio signals
By combining microphone arrays and ultrasonic sensors with machine learning models, the dereverberation technology solves the problem of audio signal reverberation in video conferencing systems and improves the quality of audio communication.
Patent Information
- Application Number
- CN201980098084.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-03
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2039-07-03
AI Technical Summary
In video conferencing systems, audio signals are easily affected by reverberation, which reduces speech intelligibility. Existing technologies are difficult to effectively remove reverberation, which affects communication quality.
The system captures audio signals using beamforming technology using a microphone array, combines ultrasonic sensors with machine learning models to determine room properties and occupant locations, and applies dereverberation parameters to reduce reverberation in the audio signal.
It effectively reduces reverberation in audio signals, improves the quality of audio signals, and enhances the clarity and intelligibility of audio communications in video conferencing systems.
Smart Images

Figure CN114026638B_ABST
Abstract
Description
Background Art
[0001] Video conferencing systems can be used for communication between parties in different locations. A video conferencing system at a near end can capture audio-video information at the near end and transmit the audio-video information to a far end. Similarly, a video conferencing system at a far end can capture audio-visual information at the far end and transmit the audio-visual information to the near end. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Figure 1 An example of a video conferencing system in a near-end room including multiple individuals according to the present disclosure is illustrated;
[0003] Figure 2 illustrates an example of a technique for performing dereverberation on an audio signal according to the present disclosure;
[0004] Figure 3 illustrates an example of a video conferencing system and related operations for performing dereverberation according to the present disclosure;
[0005] Figure 4 is a flow chart illustrating an example method of performing dereverberation in a video conferencing system according to the present disclosure;
[0006] Figure 5 is a flowchart illustrating another example method of performing dereverberation in a video conferencing system according to the present disclosure; and
[0007] Figure 6 is a block diagram that provides an example illustration of a computing device that may be employed in the present disclosure. DETAILED DESCRIPTION
[0008] The present disclosure describes a machine-readable storage medium and a method and system for dereverberation of audio signals, such as those applicable in the context of video conferencing systems. Examples of the present disclosure may include a machine-readable storage medium comprising instructions that, when executed by a processor, cause the processor to: determine a position of a person in a room. The instructions, when executed by the processor, may cause the processor to: use beamforming to capture an audio signal received from the position of the person. The instructions, when executed by the processor, may cause the processor to: determine room properties based in part on a signal sweep of the room. The instructions, when executed by the processor, may cause the processor to: determine dereverberation parameters based in part on the position of the person and the room properties. The instructions, when executed by the processor, may cause the processor to: apply the dereverberation parameters to the audio signal. In one example, the instructions cause the processor to: in response to the position of the person satisfying a position criterion, apply the dereverberation parameters to the audio signal to reduce reverberation in the audio signal; and transmit the audio signal with the reduced reverberation. In another example, the instructions cause the processor to: provide the position of the person and the room properties to a machine learning model; and use the machine learning model to determine the dereverberation parameters. In another example, an ultrasonic sensor can be used to perform a signal sweep of the room at an ultrasonic frequency. In yet another example, the room properties include room surface reflectance, room geometry, room boundaries, or a combination thereof. In one example, the instructions cause the processor to: use camera information to determine the position of the person relative to the room boundary. In another example, the instructions cause the processor to: compare the data output from the signal sweep of the room and the camera information with predefined room tags to determine the room properties; or compare the data output from the signal sweep of the room and the camera information with detected room tags to determine the room properties, wherein the detected room tags can be determined using an infrared light emitting diode (IR LED) or a laser emitter and a camera; or provide the data output from the signal sweep of the room and the camera information to a machine learning model to determine the room properties, wherein the machine learning model can be trained to classify the signal sweep data and camera information for determining the room properties.
[0009] Another example of the present disclosure may include a method for dereverberation of an audio signal. The method may include determining a position of a person in a room based in part on camera information. The method may include using beamforming to capture an audio signal received from the position of the person. The method may include determining room properties based in part on an ultrasonic signal sweep of the room. The method may include providing the position of the person and the room properties to a machine learning model. The method may include determining dereverberation parameters based on the machine learning model. The method may include applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal in response to the position of the person satisfying a position criterion. The method may include transmitting the audio signal with reduced reverberation. In one example, the room properties may include room surface reflectivity, room geometry, room boundaries, or a combination thereof. In another example, the method may include training the machine learning model to determine the dereverberation parameters based on relative person position and room properties.
[0010] Another example of the present disclosure may include a system for dereverberation of audio signals. The system may include a camera to capture camera information of a room. The system may include a microphone to capture audio signals received from the location of a person in the room. The system may include an ultrasonic sensor to capture signal sweep information of the room. The system may include a machine-readable storage medium to store a machine learning model. The system may include a processor. The processor may determine the location of the person in the room based in part on the camera information. The processor may use beamforming to capture the audio signals received from the location of the person. The processor may determine room properties based in part on the signal sweep information. The processor may provide the location of the person and the room properties to the machine learning model. The processor may use the machine learning model to determine dereverberation parameters. The processor may apply the dereverberation parameters to the audio signal to reduce reverberation in the audio signal. The processor may transmit the audio signal with reduced reverberation. In one example, the processor may apply the dereverberation parameters to the audio signal when the distance between the microphone and the location of the person is below a defined threshold. In another example, the camera may be a stereo camera, a structured light sensor camera, or a time-of-flight camera. In yet another example, the system may be a video conferencing system.
[0011] In these examples, it is noted that when discussing the storage medium, method, or system, any such discussion can be considered applicable to the other examples, regardless of whether they are explicitly discussed in the context of that example. Thus, for example, when discussing details about audio signals in the context of the storage medium, such discussion also relates to the methods and systems described herein, and vice versa.
[0012] Now turning to the accompanying drawings, Figure 1 The diagram illustrates an example of a video conferencing system 100 in a near-end room 120 including multiple people 110. The video conferencing system 100 may include a camera 102 to capture camera information of the near-end room 120. For example, the camera 102 may capture video of the people 110 in the near-end room 120. The video captured in the near-end room 120 may be converted into a video signal, and the video signal may be transmitted to the far-end room 150. The video conferencing system 100 may include a speaker (or loudspeaker) 104. The speaker 104 may receive audio signals from the far-end room 150 and generate sound based on the audio signals. The video conferencing system 100 may include a microphone 106 to capture audio in the near-end room 120. For example, the microphone 106 may capture audio spoken by the people 110 in the near-end room 120. The audio captured in the near-end room 120 may be converted into an audio signal, and the audio signal may be transmitted to the far-end room 150. Additionally, the video conferencing system 100 may include a display 108 to display video signals received from the remote room 150 .
[0013] In one example, far-end room 150 may include video conferencing system 130. Video conferencing system 130 may include camera 132 to capture camera information of far-end room 150. For example, camera 132 may capture video of person 140 in far-end room 160. The video captured in far-end room 150 may be converted into a video signal, and the video signal may be transmitted to near-end room 120. Video conferencing system 130 may include speaker 134, which may receive audio signals from near-end room 120 and generate sound based on the audio signals. Video conferencing system 130 may include microphone 136 to capture audio in far-end room 150. For example, microphone 136 may capture audio spoken by person 140 in far-end room 150. The audio captured in far-end room 150 may be converted into an audio signal, and the audio signal may be transmitted to near-end room 120. Furthermore, video conferencing system 130 may include display 138 to display the video signal received from near-end room 120.
[0014] exist Figure 1In the illustrated example, video conferencing system 100 in near-end room 120 and video conferencing system 130 in far-end room 150 may enable communication between person 110 in near-end room 120 and person 140 in far-end room 150. For example, person 110 in near-end room 120 may be able to see and hear person 140 in far-end room 150 based on audio-visual information transmitted between video conferencing system 100 in near-end room 120 and video conferencing system 130 in far-end room 150. In this non-limiting example, near-end room 120 may include four people, and far-end room 150 may include two people, although other numbers of people may be present in near-end room 120 and far-end room 150.
[0015] In one example, microphone 106 that captures audio spoken by person 110 in near-end room 120 may be a microphone array. The microphone array may include multiple microphones placed at different spatial locations. The microphone array may use beamforming to capture the audio spoken by person 110 in near-end room 120. The different spatial locations of the microphones in the microphone array that capture the audio spoken by person 110 may generate beamforming parameters. Based on the beamforming parameters, the signal strength of signals emanating from a specific direction in near-end room 120 (such as the location of person 110 in near-end room 120) may be increased. The signal strength of signals emanating from other directions in near-end room 120 (such as locations different from the location of person 110 in near-end room 120, for example, due to noise) may be combined in a benign or destructive manner based on the beamforming parameters, resulting in degradation of signals to and from locations different from the location of person 110 in near-end room 120. Thus, by using the principles of sound propagation, the microphone array may provide the ability to enhance signals emanating from a particular direction in the near-end room 120 based on knowledge of that particular direction.
[0016] In one example, beamforming technology using a microphone array can adaptively track a moving person, listen for sounds in one or more directions of the moving person, and suppress sounds (or noise) from other directions. Beamforming using a microphone array can enhance the sound quality of received speech by increasing the gain of the audio signal in the direction of the moving person and reducing the amount of far-end speaker echo received at the microphone(s) in the microphone array. In other words, by varying the gain and phase delay of a given microphone output in the microphone array, sound signals from a specific direction can be amplified through constructive interference, while sound signals from other directions can be attenuated through destructive interference. The gain and phase delay of the microphone(s) in the microphone array can be considered beamforming parameters. Furthermore, since the gain and phase delay of a given microphone output can vary based on the position of the person 110, the beamforming parameters can also depend on the position of the person 110.
[0017] Furthermore, beamforming techniques using microphone arrays can be categorized as data-independent or fixed, or data-dependent or adaptive. For data-independent or fixed beamforming techniques, the beamforming parameters may be fixed during operation. For data-dependent or adaptive beamforming techniques, the beamforming parameters may be continuously updated based on the received signal. Examples of fixed beamforming techniques may include delay and beamforming, subarray delay and beamforming, super-directional beamforming, or near-field super-directional beamforming. Examples of adaptive beamforming techniques may include generalized sidelobe canceller beamforming, adaptive microphone array noise reduction system (AMNOR) beamforming, or filtered beamforming.
[0018] In one example, person 110 in near-end room 120 may speak, and the corresponding sound may be captured using microphone 106 of video conferencing system 100 in near-end room 120. The sound captured by microphone 106 may be subject to reverberation, which is the persistence of sound after the sound is generated. Reverberation may occur when the sound is reflected, which may cause multiple reflections to accumulate, and then decay as the sound is absorbed by surfaces or objects in near-end room 120 (which may include furniture, people, air, etc.). When the sound from person 110 stops, the effect of reverberation may be noticeable, but the reflections may continue, causing the sound to persist. Reverberation can exist in indoor spaces, but it can also exist in outdoor environments where reflections are present. The level of reverberation may depend in part on the distance between person 110 and microphone 106. For example, an increased distance between person 110 and microphone 106 may result in an increased level of reverberation, while a decreased distance between person 110 and microphone 106 may result in a decreased level of reverberation.
[0019] In one example, the sound (including reverberation) captured by microphone 106 can be transmitted as an audio signal to video conferencing system 130 in far-end room 150. This audio signal can be used to produce sound at speaker 134 of video conferencing system 130 in far-end room 150. However, the sound produced at speaker 134 may include reverberation produced in near-end room 120. Therefore, when person 140 in far-end room 150 hears or hears a sound or speech from person 110 in near-end room 120, the reverberation may reduce the intelligibility of speech in the sound or speech from person 110 in near-end room 120.
[0020] In one example, dereverberation can be used to reduce the reverberation level in an audio signal transmitted from a video conferencing system 100 in a near-end room 120 to a video conferencing system 130 in a far-end room 150. Dedereverberation can remove reverberation effects and reduce contamination in an audio signal after the sound has been picked up or detected by microphone 106 of video conferencing system 100 in near-end room 120. The audio signal transmitted from video conferencing system 100 in near-end room 120 can be a near-end speech signal, which can be derived from an audio signal captured at near-end room 120 using a microphone array using beamforming. Dedereverberation can be applied to the near-end speech signal to remove reverberation from the audio signal. An audio signal including the near-end speech signal (i.e., the audio signal to which dereverberation has been applied) can be transmitted to video conferencing system 130 in far-end room 150.
[0021] In one example, dereverberation can be achieved using various methods. For example, reverberation can be reduced or eliminated using a mathematical model of the acoustic system (or room). After estimating the room acoustic model parameters, an estimate of the original signal can be determined. In another example, reverberation can be suppressed by treating reverberation as a type of (convolutional) noise and performing a denoising process specifically adapted for reverberation. In yet another example, the original dereverberated signal can be estimated from the microphone signal using, for example, a deep neural network machine learning method.
[0022] Figure 2 An example of a technique for performing dereverberation on an audio signal according to the present disclosure is illustrated. Dedereverberation can be performed using a computing device 216 in a near-end room 220. The computing device 216 can be part of a video conferencing system that captures audio and video from the near-end room and transmits the audio and video to a far-end room 230. The computing device 216 can include or be coupled to a speaker 204 (or loudspeaker), a camera 206 (such as a stereo camera, a structured light sensor camera, or a time-of-flight camera), and a microphone array 212. In other words, the speaker 204, camera 206, and microphone array 212 can be integrated with the computing device 216, or can be separate units coupled to the computing device 216.
[0023] In one example, camera 206 can capture camera information of near-end room 200. The camera information can be a digital image and / or digital video of near-end room 200. The camera information can be provided to people detector and tracker unit 208 operating on computing device 216. People detector and tracker unit 208 can analyze the camera information using object detection, which can include facial detection. People detector and tracker unit 208 can also analyze the camera information using depth estimation, which can rely on the relative scale of objects in the image. Based on the camera information, people detector and tracker unit 208 can determine the number of people in near-end room 220 and the locations of people in near-end room 220. The locations of people in near-end room 220 can be used to determine the distance between the people and microphone array 212. The person(s) detected in near-end room 220 based on the camera information can include a person currently speaking or a person not currently speaking (e.g., a person in near-end room 220 who is listening to another person who is speaking).
[0024] In one example, the position of a person may be a relative position relative to the number of people in near-end room 220. The relative position of the person may imply the relative position of one or more people relative to the microphones in microphone array 212. The relative position may be determined based on determining the position of a camera relative to the microphones in microphone array 212. The position of the camera relative to the microphones in microphone array 212 may be determined manually or using object detection. The camera position may be determined once or periodically, as camera 206 and the microphones in microphone array 212 may be fixed or semi-fixed.
[0025] As a non-limiting example, based on camera information captured using camera 206, people detector and tracker unit 208 may detect that there are four people in near-end room 220. Furthermore, based on the camera information, people detector and tracker unit 208 may determine that a first person is at a first location in near-end room 220, a second person is at a second location in near-end room 220, a third person is at a third location in near-end room 220, and a fourth person is at a fourth location in near-end room 220.
[0026] In one example, the person detector and tracker unit 208 can track a person in the near-end room 220 over a period of time. The person detector and tracker unit 208 can operate when the level of change in the incoming video frames is above a defined threshold. For example, the person detector and tracker unit 208 can operate during the start of a video conference call when a person enters and settles down in the near-end room 220, and can operate in a reduced mode when the person is unlikely to move in the near-end room 220 and therefore maintains an orientation relative to the microphone array 212.
[0027] In one example, the person detector and tracker unit 208 may provide person location information to a beamformer 210 operating on a computing device 216. The person location information may indicate the location of a person in the near-end room 220. The beamformer 210 may be a fixed beamformer (e.g., a beamformer that performs delay and beamforming) or an adaptive beamformer. The beamformer 210 may be coupled to a microphone array 212. The beamformer 210 and the microphone array 212 may work together to perform beamforming. The beamformer 210 and the microphone array 212 may capture audio signals received from the location of the person in the near-end room 220. For example, when a person in the near-end room 220 speaks and the person's location is determined based on the person location information, the beamformer 210 and the microphone array 212 may capture audio signals received from the location of the person in the near-end room 220. The audio signals may be captured using beamforming parameters, which may be set based on the location of the person in the near-end room.
[0028] In one example, the audio signal captured at microphone 212 using beamformer 210 may be subject to reverberation. For example, the audio signal captured at microphone 212 may include reverberation due to the persistence of sound in near-end room 220. Reverberation may be generated when the sound is reflected, which may cause multiple reflections to accumulate and then decay as the sound is absorbed by surfaces or objects in near-end room 220. Beamformer 210 may provide the audio signal with reverberation to dereverberation engine 214 operating on computing device 216. In other words, the output of beamformer 210 may be an input to dereverberation engine 214.
[0029] In one example, the dereverberation engine 214 may determine room properties of the near-end room 220. Room properties may include room boundaries / surface reflectivity, room geometry, and the like. The dereverberation engine 214 may determine one or more dereverberation parameters based on person position information indicating the position of a person in the near-end room 220 and the room properties of the near-end room 220. In other words, the dereverberation parameter(s) may be set based on the determined position of the person in the near-end room 220, where the position may indicate the distance between the person and the microphone array 212. Furthermore, the dereverberation parameter(s) may be set based on room properties such as room boundaries / surface reflectivity, room geometry, and the like.
[0030] In one example, the location of the person in the near-end room 220 and the room properties may be provided to a machine learning model, and the machine learning model may be used to determine (one or more) dereverberation parameters. In other words, the location of the person in the near-end room 220 and the room properties may be provided as input to the machine learning model, and (one or more) dereverberation parameters may be output by the machine learning model.
[0031] In one example, dereverberation parameter(s) may be applied to an audio signal received from the location of a person in near-end room 220, thereby producing an audio signal with reduced reverberation. In other words, the dereverberation parameter(s) may be applied to reduce reverberation caused by reflections in near-end room 220, which may produce a resulting audio signal that is less affected by reverberation. This resulting audio signal may be near-end signal 218 transmitted to far-end room 230. Because dereverberation has been applied to near-end signal 218 to reduce reverberation, near-end signal 218 may have improved sound quality.
[0032] Similarly, dereverberation parameter(s) may be determined and applied at far-end room 230. Thus, dereverberation parameter(s) may be applied to far-end signal 202, and then far-end signal 202 may be transmitted to near-end room 220. Because dereverberation has been applied to far-end signal 202 to reduce reverberation, far-end signal 202 may have increased sound quality.
[0033] In one configuration, the dereverberation engine 214 may determine room properties based in part on data output received from an ultrasonic sensor 222 communicatively coupled to the dereverberation engine 214. The ultrasonic sensor 222 may perform a signal sweep of the near-end room 220. For example, the ultrasonic sensor 222 may perform a signal sweep of the near-end room 220 at an ultrasonic frequency. The ultrasonic frequency may be above the upper audible limit of human hearing. As an example, the ultrasonic frequency used by the ultrasonic sensor 222 may be in the range of 20 kilohertz (kHz) to several gigahertz. The ultrasonic sensor 222 may perform a signal sweep of the near-end room 220 to detect objects in the near-end room 220 and measure distances, which may correspond to data output. The data output may be used, along with information from the camera 206, to determine room properties.
[0034] Alternatively, the signal sweep of the near-end room 220 may be an electromagnetic energy sweep using light, radar, sonar, or the like.
[0035] In a specific example, the ultrasonic sensor 222 may include an ultrasonic signal generator and an electronic beam steering block attached to the ultrasonic transmitter array of the ultrasonic sensor 222. The ultrasonic signal generator and the electronic beam steering block attached to the ultrasonic transmitter array may sweep the near-end room 220 at an ultrasonic frequency. The signal sweep may generate a data output that may be used by the dereverberation engine 214 to calculate the room boundary / surface reflectivity of the near-end room 220. Furthermore, the dereverberation engine 214 may provide the person location(s) detected using the camera 206 and the room boundary / surface reflectivity as input to a trained model (such as a machine learning model). The machine learning model may be trained a priori using the room boundary / surface reflectivity and the person location relative to the room boundary. The machine learning model may receive the input and provide an output of estimated dereverberation parameter(s), which may be applied to achieve dereverberation.
[0036] In one example, the dereverberation engine 214 may apply dereverberation parameters to the audio signal when the position of the person in the near-end room 220 satisfies a position criterion. The position criterion may be satisfied when the distance between the position of the person in the near-end room 220 and the microphone array 212 is above a defined threshold. For example, the position criterion may be satisfied when the distance is greater than 10 feet, greater than 15 feet, greater than 20 feet, greater than 25 feet, etc. Thus, the dereverberation parameters may be applied when the person is at an increasing distance from the microphone array 212, and may not be applied when the person is at a decreasing distance from the microphone array 212.
[0037] In one example, the beamformer 210 may operate with N beams or N channels, where N is a positive integer. One channel or one beam may correspond to a person detected using the person detector and tracker unit 208. The dereverberation engine 214 may remove the dereverberation of one channel or one beam corresponding to the detected person.
[0038] As a non-limiting example, the person detector and tracker unit 208 may detect three people in the near-end room 220. In this example, the beamformer 210 may use a first beam or channel to receive an audio signal from the first person in the near-end room 220, a second beam or channel to receive an audio signal from the second person in the near-end room 220, and a third beam or channel to receive an audio signal from the third person in the near-end room 220. The dereverberation engine 214 may determine that the first person is located 15 feet from the microphone array 212 and meets the location criteria, but the second and third people are located 5 feet and 6 feet, respectively, and do not meet the location criteria. The dereverberation engine 214 may perform dereverberation on the first beam or channel, but may not perform dereverberation on the second and third beams or channels.
[0039] De-reverberation can be blindly applied to both people in near-end room 220 and far-end room 230, which will result in an increased number of calculations. Furthermore, the distance of a speaker from the microphone array can be estimated based on speech signal power, where decreasing signal power indicates a speaker at an increasing distance from the microphone array, and increasing signal power indicates a speaker at a decreasing distance from the microphone array. A decision can be made whether to implement dereverberation based on whether the speech signal power is decreasing or increasing. However, this approach will fail if a speaker positioned relatively close to the microphone array speaks at a decreasing volume, resulting in a decreasing signal power (even if the speaker is positioned relatively close to the microphone array). Similarly, this approach will fail if a speaker positioned relatively far from the microphone array speaks at an increasing volume, resulting in an increasing signal power (even if the speaker is positioned relatively far from the microphone array).
[0040] In the present disclosure, camera information can be used to determine the relative distances between people (including speakers) in near-end room 220. This information can be used to more accurately determine the relative distances between people in near-end room 220 than signal power from microphone array 212, which can erroneously identify people as being relatively far away or the camera as being close to the microphone array 212. In the present disclosure, the relative distances between people can be used to determine whether to apply dereverberation. For example, dereverberation can be applied when the relative distance meets a location criterion, and not applied when the relative distance does not meet the location criterion. By selectively applying dereverberation based on relative distance with respect to the location criterion, computational efficiency and voice quality can be increased.
[0041] Figure 3An example of a video conferencing system 300 for performing dereverberation is illustrated. Video conferencing system 300 can be a near-end video conferencing system or a far-end video conferencing system. Video conferencing system 300 can include a camera 310 (such as a stereo camera, a structured light sensor camera, or a time-of-flight camera), a microphone array 320, an ultrasonic sensor 330, a processor 340 for performing dereverberation on an audio signal 322, and a machine-readable storage medium 370 for storing a machine learning model 372. A non-limiting example of processor 340 can be a digital signal processor (DSP).
[0042] In one example, camera 310 can capture camera information 312 of a room. Camera information 312 can include video information of the room, which can include multiple video frames. Camera 310 can operate continuously or intermittently to capture camera information 312 of the room. For example, camera 310 can operate continuously during a video conferencing session, or can operate intermittently during a video conferencing session (e.g., at the beginning of the video conferencing session and at defined periods during the video conferencing session).
[0043] In one example, microphone array 320 can capture audio signals 322 received from the positions of people in a room. Microphone array 320 can include multiple microphones located at different spatial locations. The microphones in microphone array 320 can be omnidirectional microphones, directional microphones, or a combination of omnidirectional and directional microphones.
[0044] In one example, ultrasonic sensor 330 may generate data output 332. Ultrasonic sensor 330 may measure distance using ultrasonic waves. Ultrasonic sensor 330 may include an ultrasonic element that transmits ultrasonic waves, and the ultrasonic element may receive ultrasonic waves reflected from a target. Ultrasonic sensor 330 may measure the distance to the target by measuring the time between the transmission of the ultrasonic wave and the reception of the reflected ultrasonic wave. The distance to the target may be included in data output 332 of ultrasonic sensor 330.
[0045] In one example, processor 340 may include a person location determination module 342. Person location determination module 342 may determine person location(s) 344 based on camera information 312. For example, person location determination module 342 may analyze camera information 312 using techniques such as depth estimation, object detection, and facial recognition to determine the number of people in a room and the locations of specific people within that number of people in the room. Person location(s) 344 may be relative locations relative to the locations of other people in the room.
[0046] In one example, the processor can include a beamforming module 346. The beamforming module 346 can use the microphone array 320 to perform beamforming to capture the audio signal 322 received from the person's location. In one example, the beamforming module 346 can use a fixed beamforming technique, such as delay and beamforming, subarray delay and beamforming, super-directional beamforming, or near-field super-directional beamforming. In another example, the beamforming module 346 can use an adaptive beamforming technique, such as generalized sidelobe canceller beamforming, AMNOR beamforming, or filtered beamforming.
[0047] In one example, beamforming module 346 can use beamforming parameters 348 to capture audio signals 322 received from the person's location, where beamforming parameters 346 can be based on the person's location in the room. In other words, person location 344 can be determined using camera information 312, and person location 344 can be used to set or adjust beamforming parameters 348. Based on beamforming parameters 348, audio signals 322 can be captured from the person's location.
[0048] In one example, processor 340 may include a room property determination module 350 to determine room properties 352 of a room. Room properties 352 may include room boundary / surface reflectivity and / or room geometry. In one example, room boundary / surface reflectivity may indicate the effectiveness of a material surface in reflecting sound, where the surface may be included in the room. The surface may be, but is not limited to, a glass surface, a metal surface, a wood surface, a cotton surface, a carpet surface, a concrete surface, a plastic surface, a paper surface, a ceramic surface, etc. The surface may be a wall or boundary of the room. Thus, room boundary / surface reflectivity may include the reflectivity of a glass surface, the reflectivity of a metal surface, etc. Therefore, the reflectivity may vary depending on the surface type. In another example, room geometry may indicate the shape, size, and relative arrangement of the room. For example, the room geometry may indicate whether the room is rectangular, circular, oval, square, etc. The room geometry may indicate the size of the room, which may indicate whether the room is an office, a conference room, a hall, etc.
[0049] In one example, room property determination module 350 can determine room properties 352 based on camera information 312 and data output 332 received from ultrasonic sensor 330. Camera information 312 can be analyzed using object detection, computer vision (e.g., Harris Corner detection), depth estimation, and the like to detect various objects (such as furniture, windows, etc.), surfaces, people, walls, and the like in the room, as well as the number of people in the room and person location(s) 344. Furthermore, ultrasonic sensor 330 can perform an ultrasonic signal sweep of the room and generate data output 332, which can include distance(s) to various objects, surfaces, people, walls, and the like in the room. Thus, room properties 352, including room boundary / surface reflectivity, can be determined based on the determined distance(s) from the ultrasonic signal sweep in conjunction with camera information 312.
[0050] In another example, infrared time of light sensor(s) can be used as an alternative to ultrasonic sensor 330 for distance estimation of various objects, surfaces, people, walls, etc. in a room. For example, infrared time of light sensor(s) can emit infrared signals in a room, and based on the amount of time it takes for the light to be emitted and then detected, the distance can be estimated. The determined distance can be used to determine room properties 352 including room boundaries / surface reflectivity.
[0051] In one configuration, the dereverberation module 354 can provide the room properties 352 and the person location(s) 344 to the machine learning model 372. The machine learning model 372 can be pre-trained to classify the room properties 352 and the person location(s) 344 relative to the room boundaries / surfaces. For example, the machine learning model 372 can be trained to classify various types of reflective surfaces (e.g., metal, wood, concrete), room geometry, room boundaries / surfaces and acoustics, the orientation or position of speakers relative to the room boundaries and surfaces that reflect sound, and the like. The dereverberation module 354 can use the machine learning model 372 to determine the dereverberation parameter(s) 356 to apply to the audio signal 322 based on the room properties 352 and the person location(s) 344. In other words, the room properties 352 and the person location(s) 344 can be provided as input to the machine learning model 372, and the dereverberation parameter(s) 356 can be output from the machine learning model 372.
[0052] As a non-limiting example, the dereverberation module 354 may provide an input indicating that the room has glass walls and carpeted floors, and the speaker is located near the wall. Based on this input, the dereverberation module 354 may use the machine learning model 372 to determine that specific dereverberation parameters 356 are to be applied to the audio signal 322 to reduce reverberation in the audio signal 322. As another non-limiting example, the dereverberation module 354 may provide an input indicating that the room has concrete walls, and the speaker is located in the center of the room. Based on this input, the dereverberation module 354 may use the machine learning model 372 to determine that specific dereverberation parameters 356 are to be applied to the audio signal 322 to reduce reverberation in the audio signal 322.
[0053] As examples, supervised learning, unsupervised learning, or reinforcement learning can be used to generate the machine learning model 372. The machine learning model 372 can apply feature learning, sparse dictionary learning, anomaly detection, decision trees, association rules, heuristic rules, etc. to improve the performance of the machine learning model 372 over time. In addition, the machine learning model 372 can incorporate statistical models (e.g., regression), principal component analysis, deep neural networks, or a type of artificial intelligence (AI).
[0054] In one example, the dereverberation module 354 may perform dereverberation on the audio signal 322 using dereverberation parameter(s) 356. For example, the dereverberation module 354 may apply the dereverberation parameter(s) 356 to the audio signal 322 to reduce reverberation in the audio signal 322. In one example, the dereverberation module 354 may apply the dereverberation parameter(s) 356 to the audio signal 322 when the location criterion 378 is satisfied. For example, the location criterion 378 may be satisfied when the distance between the person location 344 and the microphone 320 is above a defined threshold, and may not be satisfied when the distance between the person location 344 and the microphone 320 is below the defined threshold.
[0055] In one example, the processor 340 may include an audio signal transmission module 358. The audio signal transmission module 358 may receive the audio signal 322 with reduced reverberation from the dereverberation module 354. The audio signal transmission module 358 may transmit the audio signal with reduced reverberation to, for example, a remote video conferencing system.
[0056] In one configuration, the room property determination module 350 can use predefined room tags 374 stored in the machine-readable storage medium 370 to determine room properties 352, including room surface reflectivity (or room surface reflectivity properties). The predefined room tags 374 can include data samples that have been labeled with a tag. In other words, unlabeled data can be labeled with informative tags to produce labeled data. The predefined room tags 374 can correspond to potential objects or surfaces in the room, such as chairs, tables, mirrors, artwork, rugs, windows, floor tiles, glass windows, concrete walls, etc. The predefined room tag 374 for a given object can correspond to a predetermined room surface reflectivity. The room property determination module 350 can compare the data output 332 from the ultrasound signal sweep of the room and the camera information 312 with the predefined room tags 374 to determine the room properties 352. In other words, the room property determination module 350 may compare or map the predefined room labels 374 (or manual surface texture labels) to room surfaces having acoustic reflectivity parameters (as indicated in the camera information 312 and / or data output 332 ) to determine the room properties 352 .
[0057] In one configuration, the room property determination module 350 can use detected room tags 376 stored in the machine-readable storage medium 370 to determine room properties 352, including room surface reflectivity (or room surface reflectance properties). The detected room tags 376 can be generated using an infrared light emitting diode (IR LED) or a laser emitter and a camera 310 (or a pair of IR LED / laser emitters and a camera). In this case, the detected room tags 376 generated using the LED / laser emitter and camera can be considered the true source. The room property determination module 350 can compare the data output 332 of the ultrasonic signal sweep from the room and the camera information 312 with the detected room tags 376 to determine the room properties 352. In other words, the room property determination module 350 can compare or map the detected room tags 376 (or detected surface texture tags) to room surfaces having acoustic reflectivity parameters (as indicated in the camera information 312 and / or data output 332) to determine the room properties 352.
[0058] In one configuration, the room property determination module 350 can use a separate machine learning model 372 to determine room properties 352 including room surface reflectivity (or room surface reflectance properties). In one example, the separate machine learning model 372 can be a deep learning model that is trained to detect and classify surfaces based on camera data and signal sweep data indicating occupant locations. The room property determination module 350 can provide the data output 332 from the ultrasonic signal sweep of the room and the camera information 312 to the separate machine learning model 372 to determine the room properties 352.
[0059] In one configuration, spatial audio technology can be used to create directional sound at a far-end video conferencing system by collecting information from the near end. The far-end device can be a soundbar or headband headphones for which directional sound can be created. For a soundbar, beamforming can be used to create directional sound. For a headband headphones, head-related transfer functions (HTRF) can be used to create directional sound. The direction of the person at the near end can be estimated by using camera information 312, and the average position of the person can be selected to accommodate slight movements of the person at the near end. Information about the direction of the person and the average position of the person can be transmitted from the near-end video conferencing system 300 to the far-end video conferencing system to enable the creation of directional sound. By selecting the average position of the person, the microphone beamformer or HTRF spatial audio renderer at the far-end video conferencing system does not need to continuously change parameters, thereby saving computation at the far-end video conferencing system.
[0060] Figure 4 4 is a flow chart illustrating an example method 400 for performing dereverberation in a video conferencing system. The method may be executed as instructions on a machine, wherein the instructions may be included on a non-transitory machine-readable storage medium. The method may include determining a location of a person in a room, as in block 410. The method may include using beamforming to capture an audio signal received from the location of the person, as in block 420. The method may include determining room properties based in part on a signal sweep of the room, as in block 430. The method may include determining dereverberation parameters based in part on the location of the person and the room properties, as in block 440. The method may include applying the dereverberation parameters to the audio signal, as in block 450. In one example, method 400 may be performed using video conferencing system 300, but method 400 is not limited to being performed using video conferencing system 300.
[0061] Figure 5is a flow chart illustrating an example method 500 for performing dereverberation in a video conferencing system. The method may be executed as instructions on a machine, wherein the instructions may be included on a non-transitory machine-readable storage medium. The method may include determining a location of a person in a room based in part on camera information, as in block 510. The method may include using beamforming to capture an audio signal received from the location of the person, as in block 520. The method may include determining room properties based in part on an ultrasound signal sweep of the room, as in block 530. The method may include providing the location of the person and the room properties to a machine learning model, as in block 540. The method may include determining dereverberation parameters based on the machine learning model, as in block 550. The method may include applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal, in response to the location of the person satisfying a location criterion, as in block 560. The method may include transmitting the audio signal with reduced reverberation, as in block 570. In one example, method 500 may be performed using video conferencing system 300 , but method 500 is not limited to being performed using video conferencing system 300 .
[0062] Figure 6 A computing device 610 is illustrated on which the modules of the present disclosure may be executed. A computing device 610 is illustrated on which a high-level example of the present disclosure may be executed. The computing device 610 may include a processor(s) 612 in communication with a memory device 620. The computing device may include a local communication interface 618 for components in the computing device. For example, the local communication interface may be a local data bus and / or an associated address or control bus, as may be desired.
[0063] The memory device 620 may contain a module 624 executable by the processor(s) 612 and data for the module 624. The module 624 may perform the functions described earlier, such as: determining a location of a person in a room based in part on camera information; capturing an audio signal received from the location of the person using beamforming; determining room properties based in part on an ultrasound signal sweep of the room; providing the location of the person and the room properties to a machine learning model; determining dereverberation parameters based on the machine learning model; in response to the location of the person satisfying a location criterion, applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal; and transmitting the audio signal with the reduced reverberation.
[0064] A data repository 622 may also be located in the memory device 620 for storing data related to the modules 624 and other applications, along with an operating system executable by the processor(s) 612 .
[0065] Other applications may also be stored in the memory device 620 and executed by the processor(s) 612. The components or modules discussed in this description may be implemented in the form of machine-readable software using a high-level programming language that is compiled, interpreted, or executed using a mixture of these methods.
[0066] The computing device may also have access to I / O (input / output) devices 614 that can be used by the computing device. An example of an I / O device is a display screen that can be used to display output from the computing device. Networking devices 616 and similar communication devices may be included in the computing device. Networking devices 616 may be wired or wireless networking devices that connect to the Internet, a local area network (LAN), a wide area network (WAN), or other computing network.
[0067] The components or modules shown as being stored in the memory device 620 can be executed by the processor 612. The term "executable" can refer to program files in a form that can be executed by the processor 612. For example, a program in a higher-level language can be compiled into machine code that can be loaded into a random access portion of the memory device 620 and executed by the processor 612, or source code can be loaded by another executable program and interpreted to generate instructions in a random access portion of the memory to be executed by the processor. The executable program can be stored in a portion or component of the memory device 620. For example, the memory device 620 can be a random access memory (RAM), a read-only memory (ROM), a flash memory, a solid-state drive, a memory card, a hard drive, an optical disk, a floppy disk, a magnetic tape, or other memory component.
[0068] Processor 612 may represent multiple processors, and memory 620 may represent multiple memory units operating in parallel with the processing circuitry. This can provide parallel processing channels for processes and data in the system. Local interface 618 may function as a network to facilitate communication between multiple processors and multiple memories. Local interface 618 may utilize additional systems designed for coordinated communication, such as load balancing, bulk data transfer, and the like.
[0069] Although the flowchart presented for the present disclosure may imply a specific execution order, the execution order may be different from the illustrated order. For example, the order of the other two frames may be rearranged relative to the illustrated order. In addition, two or more frames shown in succession may be executed in parallel or in a partially parallelized manner. In some configurations, the (one or more) frames shown in the flowchart may be omitted or skipped. For the purpose of enhancing practicality, accounting, performance, measurement, fault diagnosis, or for similar reasons, multiple counters, state variables, warning semaphores, or messages may be added to the logic flow.
[0070] Some functional units described in this specification have been labeled as modules to more specifically emphasize their implementation independence. For example, a module can be implemented as a hardware circuit, including custom very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, or programmable logic devices.
[0071] Modules may also be implemented in machine-readable software for execution by various types of processors. An identified module of executable code may, for example, include blocks of computer instructions that may be organized as objects, procedures, or functions. However, the executable files of the identified modules need not be physically located together, but may include disparate instructions stored in different locations that constitute the module and, when logically joined together, achieve the stated purpose of the module.
[0072] In fact, the module of executable code can be a single instruction, or many instructions, and can even be distributed on several different code segments, distributed in the middle of different programs and distributed across several memory devices.Similarly, operational data can be identified and illustrated in this article in the module, and can be embodied and organized in the data structure of suitable type with suitable form.Operational data can be collected as a single data set, or can be distributed on different locations, including on the different storage devices of distribution.These modules can be passive or active, including the agent that can be operated to perform the desired function.
[0073] The disclosure described herein may also be stored on a computer-readable storage medium, including volatile and nonvolatile, removable and non-removable media implemented with the disclosure, for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may include, but are not limited to, RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory devices, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage devices, cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other computer storage media that can be used to store the desired information and the disclosed content described.
[0074] The devices described herein may also include communication connections or networking devices and networking connections that allow the devices to communicate with other devices. A communication connection may be an example of a communication medium. A communication medium may embody computer-readable instructions, data structures, program modules, and other data in a modulated data signal (such as a carrier wave or other transmission mechanism), and may include information delivery media. By way of example and not limitation, communication media may include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. As used herein, the term computer-readable media may include communication media.
[0075] Reference has been made to the examples illustrated in the drawings, and specific language has been used herein to describe these examples. However, it will be understood that no limitation of the scope of the present disclosure is thereby intended. Alterations and further modifications of the features described herein, as well as additional applications of the examples described herein, are considered to be within the scope of this description.
[0076] In addition, the described features, structures or characteristics can be combined in a suitable manner. In the foregoing description, many specific details are provided, such as examples of various configurations, to provide a comprehensive understanding of the examples of the disclosed content described. The present disclosure can be practiced without some of the specific details or can be practiced using other methods, components, equipment, etc. In other examples, some structures or operations are not shown or described in detail to avoid blurring aspects of the present disclosure.
[0077] Although the subject matter has been described in language specific to structural features and / or operations, it is to be understood that the subject matter defined in the appended claims is not limited to the specific features and operations described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. Many modifications and alternative arrangements may be devised without departing from the scope of the disclosed content.
Claims
1. A machine-readable storage medium comprising instructions that, when executed by a processor, cause the processor to: Determine the location of people in the room; capturing, using beamforming with a microphone array, an audio signal received from the position of the person; determining a room property based in part on a signal sweep of the room; determining dereverberation parameters based in part on the position of the person and the room properties; responsive to the position of the person satisfying a position criterion, applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal; as well as transmitting said audio signal with reduced reverberation, The position criterion is satisfied when the distance between the position of the person in the room and the microphone array is above a defined threshold.
2. The machine-readable storage medium of claim 1 , wherein the instructions cause the processor to: providing the location of the person and the properties of the room to a machine learning model; and The dereverberation parameters are determined using the machine learning model.
3. The machine-readable storage medium of claim 1, wherein the signal sweep of the room is performed at ultrasonic frequencies using an ultrasonic sensor. 4 . The machine-readable storage medium of claim 1 , wherein the room properties include room surface reflectivity, room geometry, room boundaries, or a combination thereof. 5 . The machine-readable storage medium of claim 1 , wherein the instructions cause the processor to use camera information to determine the position of the person relative to a room boundary.
6. The machine-readable storage medium of claim 1 , wherein the instructions cause the processor to: comparing data output from a signal sweep of the room and camera information to predefined room labels to determine properties of the room; or comparing data output from a signal sweep of the room and camera information to a detected room tag determined using an infrared light emitting diode (IR LED) or laser emitter and a camera to determine the nature of the room; or The data output from the signal sweep of the room and the camera information are provided to a machine learning model to determine the room properties, wherein the machine learning model is trained to classify the signal sweep data and the camera information for determining the room properties.
7. A method for dereverberation of an audio signal, comprising: determining the location of people in the room based in part on the camera information; capturing, using beamforming with a microphone array, an audio signal received from the position of the person; determining room properties based in part on a sweep of the ultrasound signal across the room; providing the location of the person and the nature of the room to a machine learning model; determining dereverberation parameters based on the machine learning model; responsive to the position of the person satisfying a position criterion, applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal; as well as transmitting said audio signal with reduced reverberation, The position criterion is met when the distance between the position of the person in the room and the microphone array is above a defined threshold.
8. The method of claim 7, wherein the room properties include room surface reflectivity, room geometry, room boundaries, or a combination thereof.
9. The method according to claim 7, comprising: The machine learning model is trained to determine the dereverberation parameters based on relative occupant positions and room properties.
10. The method according to claim 7, comprising: comparing data output from an ultrasound signal sweep of the room to predefined room labels to determine properties of the room; or comparing data output from an ultrasound signal sweep of the room with a detected room tag determined using an infrared light emitting diode (IR LED) or laser emitter and a camera to determine the room properties; or Data output from the ultrasound signal sweep of the room is provided to a machine learning model to determine the room properties, wherein the machine learning model is trained to classify the ultrasound signal sweep data and determine the room properties.
11. A system for dereverberation of an audio signal, comprising: Camera, used to capture camera information of the room; a microphone for capturing audio signals received from the positions of people in the room; Ultrasonic sensors to capture the signal sweep across the room; a machine-readable storage medium for storing a machine learning model; and The processor is configured to: determining a location of the person in a room based in part on the camera information; capturing the audio signal received from the position of the person using beamforming; determining a room property based in part on the signal sweep information; providing the machine learning model with the location of the person and the properties of the room; determining dereverberation parameters using the machine learning model; responsive to the position of the person satisfying a position criterion, applying the dereverberation parameters to the audio signal to reduce reverberation in the audio signal; as well as transmitting said audio signal with reduced reverberation, The position criterion is met when the distance between the position of the person in the room and the microphone array is above a defined threshold.
12. The system of claim 11, wherein the camera is a stereo camera, a structured light sensor camera, or a time-of-flight camera.
13. The system of claim 11, wherein the system is a video conferencing system.
Citation Information
Patent Citations
Speech processing device, speech processing method, and speech processing system
US20160203828A1
Multi-modal dereverbaration in far-field audio systems
US20190028829A1