Systems and methods for multi-microphone automatic clinical documentation

By using a combination of microphone arrays and mobile devices in a multi-microphone system, along with signal-to-noise ratio weighting and beamforming in a speech processing system, the challenges of signal acquisition and speaker recognition in traditional systems are solved, resulting in more accurate speech transcription.

CN115516553BActive Publication Date: 2025-11-25MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180033186.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-08
Filing Date
2021-05-10
Publication Date
2025-11-25
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

Traditional multi-microphone automated clinical documentation systems face challenges in signal acquisition, signal combination management, speaker location and identity tracking, leading to inaccurate speech transcription.

Method used

By using a multi-microphone system consisting of a microphone array and mobile electronic devices, combined with a speech processing system, signal-to-noise ratio weighting and beamforming of the audio stream are performed. Spatial and spectral information is used for speech activity detection and speaker identification to achieve audio stream alignment.

Benefits of technology

It improves the accuracy and reliability of speech transcription and can effectively manage the signal combination of multiple microphone systems and speaker identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115516553B_ABST
    Figure CN115516553B_ABST
Patent Text Reader

Abstract

A method, computer program product, and computing system for receiving audio consultation information from a first microphone system, thereby defining a first audio stream. Audio consultation information can be received from a second microphone system, thereby defining a second audio stream. Speech activity can be detected in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream. Speech activity can be detected in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream. The first audio stream and the second audio stream can be aligned based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 022,269, filed May 8, 2020, the entirety of which is incorporated by reference herein. BACKGROUND

[0003] Automated clinical documentation (ACD) can be used, for example, to convert transcribed conversations (e.g., of a physician, a patient, and / or other participants, such as a patient’s family member, a nurse, a physician’s assistant, etc.) speech into formatted (e.g., medical) reports. Such reports can be reviewed, for example, to ensure the accuracy of the physician, scribe, etc. report.

[0004] The process of capturing speech can include distant conversation automatic speech recognition (DCASR). DCASR includes multiple microphone systems configured to record and recognize speech of one or more speakers (e.g., an array of microphones and a single mobile microphone carried by a speaker). Conventional DCASR methods are subject to various challenges. For example, acquiring signals through multiple microphone systems (e.g., one or more arrays of microphones and one or more mobile microphones) can be problematic. For example, when using a single microphone for beamforming, there are single-channel noise reduction and dereverberation methods that cannot be applied to multiple microphone systems. In addition, conventional DCASR methods cannot manage the combination of enhanced signals from various devices within a room. Conventional DCASR methods can experience problems in localizing the position of a speaker at any given time, tracking their movements, and then identifying their identity. This can result in less accurate text transcription of the recorded speech. In addition, automatic speech recognition (ASR) systems that convert audio to text can not be able to handle various metadata and audio signals from multiple sources (e.g., multiple microphone systems). SUMMARY

[0005] In one implementation, a computer-implemented method performed by a computer can include, but is not limited to, receiving audio encounter information from a first microphone system, thereby defining a first audio stream. Audio encounter information can be received from a second microphone system, thereby defining a second audio stream. Speech activity can be detected in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream. Speech activity can be detected in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream. The first audio stream and the second audio stream can be aligned based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

[0006] One or more of the following features can be included. The first microphone system can include an array of microphones. The second microphone system can include a mobile electronic device. The first audio stream and the second audio stream can be processed with one or more speech processing systems in response to aligning the first audio stream and the second audio stream. Processing the first audio stream and the second audio stream with one or more speech processing systems can include weighting the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio of the first audio stream and a signal-to-noise ratio of the second audio stream, thereby defining a first audio stream weight and a second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream with a first speech processing system, thereby defining a first speech processing output, processing the second audio stream with a second speech processing system, thereby defining a second speech processing output, and combining the first speech processing output and the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight.

[0007] In another implementation, a computer program product resides on a computer- readable medium and has a plurality of instructions stored on it. The instructions, when executed by a processor, cause the processor to perform operations including, but not limited to, receiving audio consultation information from a first microphone system, thereby defining a first audio stream. Audio consultation information can be received from a second microphone system, thereby defining a second audio stream. Speech activity can be detected in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream. Speech activity can be detected in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream. The first audio stream and the second audio stream can be aligned based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

[0008] One or more of the following features can be included. The first microphone system can include an array of microphones. The second microphone system can include a mobile electronic device. The first audio stream and the second audio stream can be processed with one or more speech processing systems in response to aligning the first audio stream and the second audio stream. Processing the first audio stream and the second audio stream with one or more speech processing systems can include weighting the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio for the first audio stream and a signal-to-noise ratio for the second audio stream, thereby defining a first audio stream weight and a second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream with a first speech processing system, thereby defining a first speech processing output; processing the second audio stream with a second speech processing system, thereby defining a second speech processing output; and combining the first speech processing output and the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight.

[0009] In another implementation, a computing system includes a processor and a memory configured to perform operations including, but not limited to, receiving audio consultation information from a first microphone system, thereby defining a first audio stream. The processor can also be configured to receive audio consultation information from a second microphone system, thereby defining a second audio stream. The processor can also be configured to detect speech activity in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream. The processor can also be configured to detect speech activity in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream. The processor can also be configured to align the first audio stream and the second audio stream based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

[0010] One or more of the following features can be included. The first microphone system can include an array of microphones. The second microphone system can include a mobile electronic device. The first audio stream and the second audio stream can be processed with one or more speech processing systems in response to aligning the first audio stream and the second audio stream. Processing the first audio stream and the second audio stream with one or more speech processing systems can include weighting the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio for the first audio stream and a signal-to-noise ratio for the second audio stream, thereby defining a first audio stream weight and a second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight. Processing the first audio stream and the second audio stream with one or more speech processing systems can include processing the first audio stream with a first speech processing system, thereby defining a first speech processing output; processing the second audio stream with a second speech processing system, thereby defining a second speech processing output; and combining the first speech processing output and the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight.

[0011] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a schematic diagram of an automated clinical documentation computer system and automated clinical documentation process coupled to a distributed computing network;

[0013] Figure 2 is a schematic diagram of a modular ACD system of the automated clinical documentation computer system of Figure 1 ;

[0014] Figure 3 is a schematic diagram of a mixed media ACD device included within the modular ACD system of Figure 2 ;

[0015] Figure 4 is a schematic diagram of various modules included within an ACD computer device included within the modular ACD system of Figure 2 ;

[0016] Figure 5 is a flow diagram of an implementation of the automated clinical documentation process of Figure 1 ;

[0017] Figures 6-7 is a schematic diagram of a modular ACD system for various implementations of the automated clinical documentation process of Figure 1 ;

[0018] Figure 8 is a flowchart of an implementation of an automatic clinical documentation process of Figure 1

[0019] Figures 9-10 is a schematic of audio encounter information received by various microphones of a microphone array for various implementations of an automatic clinical documentation process according to Figure 1

[0020] Figure 11 is a flowchart of an implementation of an automatic clinical documentation process of Figure 1

[0021] Figure 12 is a schematic of audio encounter information received by various microphones of a microphone array for various implementations of an automatic clinical documentation process according to Figure 1

[0022] Figure 13 is a schematic of multiple speaker representations generated for an implementation of an automatic clinical documentation process according to Figure 1

[0023] Figure 14 is a schematic of speaker metadata generated for various implementations of an automatic clinical documentation process according to Figure 1

[0024] Figure 15 is a flowchart of an implementation of an automatic clinical documentation process of Figure 1

[0025] Figure 16 is a schematic of a modular ACD system for an implementation of an automatic clinical documentation process according to Figure 1

[0026] Figure 17 is a flowchart of an implementation of an automatic clinical documentation process of Figure 1

[0027] Figures 18-19 is a schematic of alignment of audio encounter information received by various microphone systems for various implementations of an automatic clinical documentation process according to Figure 1

[0028] Figures 20A-20B is a schematic of processing of audio encounter information received by various microphone systems configured using various speech processing systems for various implementations of an automatic clinical documentation process according to Figure 1

[0029] Like reference numbers in the various figures indicate like elements. DETAILED DESCRIPTION​​​​​​​​​​​

[0030] System Overview:

[0031] Reference Figure 2 An automated clinical documentation process 10 is shown. As will be discussed in greater detail below, the automated clinical documentation process 10 can be configured to automate the collection and processing of clinical encounter information to generate / store / distribute medical records.

[0032] The automated clinical documentation process 10 can be implemented as a server-side process, a client-side process, or a hybrid server-side / client-side process. For example, the automated clinical documentation process 10 can be implemented as a purely server-side process via the automated clinical documentation process 10s. Alternatively, the automated clinical documentation process 10 can be implemented as a purely client-side process via one or more of the automated clinical documentation process 10cl, the automated clinical documentation process 10c2, the automated clinical documentation process 10c3, and the automated clinical documentation process 10c4. Alternatively, the automated clinical documentation process 10 can be implemented as a hybrid server-side / client-side process via the automated clinical documentation process 10s in combination with one or more of the automated clinical documentation process 10cl, the automated clinical documentation process 10c2, the automated clinical documentation process 10c3, and the automated clinical documentation process 10c4.

[0033] Accordingly, the automated clinical documentation process 10 used in this disclosure can include any combination of the automated clinical documentation process 10s, the automated clinical documentation process 10cl, the automated clinical documentation process 10c2, the automated clinical documentation process 10c3, and the automated clinical documentation process 10c4.

[0034] The automated clinical documentation process 10s can be a server application and can reside on and be executed by an automated clinical documentation (ACD) computer system 12, which can be connected to a network 14 (e.g., the Internet or a local area network). The ACD computer system 12 can include various components, examples of which can include, but are not limited to, a personal computer, a server computer, a series of server computers, a minicomputer, a mainframe computer, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, a cloud-based computing system, and a cloud-based storage platform.

[0035] As known in the art, the SAN can include one or more of a personal computer, a server computer, a series of server computers, a minicomputer, a mainframe computer, a RAID device, and a NAS system. The various components of the ACD computer system 12 can execute one or more operating systems, examples of which can include, but are not limited to: for example, Microsoft Windows Server tm , Redhat Linux tm , Unix, or a custom operating system.

[0036] The set of instructions and subroutines of the automated clinical documentation process 10s that can be stored in the storage device 16 coupled to the ACD computer system 12 can be executed by one or more processors (not shown) and one or more memory architectures (not shown) included within the ACD computer system 12. Examples of the storage device 16 can include, but are not limited to: a hard disk drive; a RAID device; a random access memory (RAM); a read-only memory (ROM); and all forms of flash memory devices.

[0037] The network 14 can be connected to one or more auxiliary networks (e.g., the network 18), examples of which can include, but are not limited to: for example, a local area network; a wide area network; or an intranet.

[0038] Various IO requests (e.g., the IO request 20) can be sent from the automated clinical documentation process 10s, the automated clinical documentation process 10cl, the automated clinical documentation process 10c2, the automated clinical documentation process 10c3, and / or the automated clinical documentation process 10c4 to the ACD computer system 12. Examples of the IO request 20 can include, but are not limited to, a data write request (i.e., a request to write content to the ACD computer system 12) and a data read request (i.e., a request to read content from the ACD computer system 12).

[0039] The instruction sets and subroutines of the automated clinical documentation process 10cl, the automated clinical documentation process 10c2, the automated clinical documentation process 10c3, and / or the automated clinical documentation process 10c4, which can be stored on the storage device 20, 22, 24, 26 coupled to the ACD client electronic device 28, 30, 32, 34, respectively, can be executed by one or more processors (not shown) and one or more memory architectures (not shown) incorporated into the ACD client electronic device 28, 30, 32, 34, respectively. The storage device 20, 22, 24, 26 can include, but is not limited to, a hard disk drive; an optical drive; a RAID device; random access memory (RAM); read-only memory (ROM), and all forms of flash memory storage. Examples of the ACD client electronic device 28, 30, 32, 34 can include, but are not limited to, a personal computing device 28 (e.g., a smart phone, a personal digital assistant, a laptop computer, a notebook computer, and a desktop computer), an audio input device 30 (e.g., a hand-held microphone, a lapel microphone, an embedded microphone such as a microphone embedded in eyewear, a smart phone, a tablet computer, and / or a watch, and an audio recording device), a display device 32 (e.g., a tablet computer, a computer monitor, and a smart television), a machine vision input device 34 (e.g., an RGB imaging system, an infrared imaging system, an ultraviolet imaging system, a laser imaging system, a sonar imaging system, a radar imaging system, and a thermal imaging system), a hybrid device (e.g., a single device that includes the functionality of one or more of the aforementioned reference devices; not shown), an audio presentation device (e.g., a speaker system, a headphone system, or an earbud system; not shown), various medical devices (e.g., a medical imaging device, a heart monitor, a weight scale, a thermometer, and a blood pressure machine; not shown), and a dedicated network device (not shown).

[0040] The users 36, 38, 40, 42 can access the ACD computer system 12 directly through the network 14 or through the auxiliary network 18. In addition, the ACD computer system 12 can be connected to the network 14 through the auxiliary network 18, as shown by link line 44.

[0041] Various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) can be directly or indirectly coupled to network 14 (or network 18). For example, personal computing device 28 is shown as being directly coupled to network 14 via a hardwired network connection. In addition, machine vision input device 34 is shown as being directly coupled to network 18 via a hardwired network connection. Audio input device 30 is shown as being wirelessly coupled to network 14 via a wireless communication channel 46 established between audio input device 30 and a wireless access point (i.e., WAP) 48, which is shown as being directly coupled to network 14. WAP 48 can be, for example, an IEEE 802.11a, 802.11b, 802.11g, 802.11n, Wi-Fi, and / or Bluetooth device capable of establishing wireless communication channel 46 between audio input device 30 and WAP 48. Display device 32 is shown as being wirelessly coupled to network 14 by way of a wireless communication channel 50 established between display device 32 and a WAP 52, which is shown as being directly coupled to network 14.

[0042] Various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) can each execute an operating system, examples of which can include, but are not limited to, Microsoft Windows tm , Apple Macintosh tm , Redhat Linux tm , or a custom operating system, where the combination of various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) and ACD computer system 12 can form a modular ACD system 54.

[0043] Referring also to Figure 3 , a simplified example embodiment of a modular ACD system 54 is shown, which is configured to automate clinical documentation. Modular ACD system 54 can include a machine vision system 100 configured to obtain machine vision encounter information 102 regarding a patient encounter, an audio recording system 104 configured to obtain audio encounter information 106 regarding the patient encounter, and a computer system (e.g., ACD computer system 12) configured to receive (respectively) machine vision encounter information 102 and audio encounter information 106 from machine vision system 100 and audio recording system 104. Modular ACD system 54 can also include a display rendering system 108 configured to render visual information 110 and an audio rendering system 112 configured to render audio information 114, where ACD computer system 12 can be configured to provide (respectively) visual information 110 and audio information 114 to display rendering system 108 and audio rendering system 112.

[0044] Examples of the machine vision system 100 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 34, examples of which can include, but are not limited to, RGB imaging systems, infrared imaging systems, ultraviolet imaging systems, laser imaging systems, sonar imaging systems, radar imaging systems, and thermal imaging systems). Examples of the audio recording system 104 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 30, examples of which can include, but are not limited to, handheld microphones, lapel microphones, embedded microphones such as those embedded within eyeglasses, smartphones, tablet computers, and / or watches, and audio recording devices). Examples of the display presentation system 108 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 32, examples of which can include, but are not limited to, tablet computers, computer monitors, and smart televisions). Examples of the audio presentation system 112 can include, but are not limited to, one or more ACD client electronic devices (e.g., the audio presentation device 116, examples of which can include, but are not limited to, speaker systems, headphone systems, and earbud systems).

[0045] As will be discussed in greater detail below, the ACD computer system 12 can be configured to access one or more data sources 118 (e.g., a plurality of separate data sources 120, 122, 124, 126, 128), examples of which can include, but are not limited to, one or more of a user profile data source, a voiceprint data source, a sound characteristic data source (e.g., for adapting automated speech recognition models), a faceprint data source, an anthropomorphic data source, a discourse identifier data source, a wearable token identifier data source, an interaction identifier data source, a medical condition symptom data source, a prescription compatibility data source, a medical insurance coverage data source, and a home health care data source. While five different examples of data sources 118 are shown in this particular example, this is for illustrative purposes only and is not intended to be limiting of the present disclosure, as other configurations are possible and are considered to be within the scope of the present disclosure.

[0046] As will be discussed in greater detail below, the modular ACD system 54 can be configured to monitor a monitored space (e.g., the monitored space 130) in a clinical environment, examples of which can include, but are not limited to, a physician’s office, a medical facility, a medical practice, a medical laboratory, an urgent care facility, a medical clinic, an emergency room, an operating room, a hospital, a long-term care facility, a rehabilitation facility, a nursing home, and a hospice facility. Thus, examples of the aforementioned patient visit can include, but are not limited to, a patient visiting one or more of the aforementioned clinical environments (e.g., a physician’s office, a medical facility, a medical practice, a medical laboratory, an urgent care facility, a medical clinic, an emergency room, an operating room, a hospital, a long-term care facility, a rehabilitation facility, a nursing home, and a hospice facility).

[0047] When the aforementioned clinical environment is larger or requires a higher level of resolution, the machine vision system 100 can include multiple discrete machine vision systems. As mentioned above, examples of the machine vision system 100 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 34, examples of which can include, but are not limited to, an RGB imaging system, an infrared imaging system, an ultraviolet imaging system, a laser imaging system, a sonar imaging system, a radar imaging system, and a thermal imaging system). Thus, the machine vision system 100 can include one or more of each of an RGB imaging system, an infrared imaging system, an ultraviolet imaging system, a laser imaging system, a sonar imaging system, a radar imaging system, and a thermal imaging system.

[0048] When the aforementioned clinical environment is larger or requires a higher level of resolution, the audio recording system 104 can include multiple discrete audio recording systems. As mentioned above, examples of the audio recording system 104 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 30, examples of which can include, but are not limited to, a handheld microphone, a lapel microphone, an embedded microphone (such as a microphone embedded within eyewear, a smartphone, a tablet computer, and / or a watch), and an audio recording device). Thus, the audio recording system 104 can include one or more of each of a handheld microphone, a lapel microphone, an embedded microphone (such as a microphone embedded within eyewear, a smartphone, a tablet computer, and / or a watch), and an audio recording device.

[0049] When the aforementioned clinical environment is larger or requires a higher level of resolution, the display presentation system 108 can include multiple discrete display presentation systems. As mentioned above, examples of the display presentation system 108 can include, but are not limited to, one or more ACD client electronic devices (e.g., the ACD client electronic device 32, examples of which can include, but are not limited to, a tablet computer, a computer monitor, and a smart television). Thus, the display presentation system 108 can include one or more of each of a tablet computer, a computer monitor, and a smart television.

[0050] When the aforementioned clinical environment is larger or requires a higher level of resolution, the audio presentation system 112 can include multiple discrete audio presentation systems. As mentioned above, examples of the audio presentation system 112 can include, but are not limited to, one or more ACD client electronic devices (e.g., the audio presentation device 116, examples of which can include, but are not limited to, a speaker system, a headphone system, or an earbud system). Thus, the audio presentation system 112 can include one or more of each of a speaker system, a headphone system, or an earbud system.

[0051] The ACD computer system 12 can include multiple discrete computer systems. As mentioned above, the ACD computer system 12 can include various components, examples of which can include, but are not limited to, a personal computer, a server computer, a series of server computers, a minicomputer, a mainframe computer, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, a cloud-based computing system, and a cloud-based storage platform. Thus, the ACD computer system 12 can include one or more of each of a personal computer, a server computer, a series of server computers, a minicomputer, a mainframe computer, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, a cloud-based computing system, and a cloud-based storage platform.

[0052] Reference is also made to Figure 2The audio recording system 104 can include a directional microphone array 200 having a plurality of discrete microphone assemblies. For example, the audio recording system 104 can include a plurality of discrete audio capture devices (e.g., audio capture devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that can form the microphone array 200. As will be discussed in greater detail below, the modular ACD system 54 can be configured to form one or more audio recording beams (e.g., audio recording beams 220, 222, 224) via the discrete audio capture devices (e.g., audio capture devices 202, 204, 206, 208, 210, 212, 214, 216, 218) included within the audio recording system 104.

[0053] For example, the modular ACD system 54 can also be configured to steer one or more audio recording beams (e.g., audio recording beams 220, 222, 224) to one or more visit participants (e.g., visit participants 226, 228, 230) of the patient visit described above. Examples of visit participants (e.g., visit participants 226, 228, 230) can include, but are not limited to, medical professionals (e.g., doctors, nurses, physician assistants, laboratory technicians, physical therapists, scribes (e.g., transcriptionists), and / or staff participating in the patient visit), patients (e.g., persons visiting the clinical environment described above for a patient visit), and third parties (e.g., friends of the patient, relatives of the patient, and / or acquaintances of the patient participating in the patient visit).

[0054] Accordingly, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize one or more discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) to form an audio recording beam. For example, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize the audio acquisition device 210 to form an audio recording beam 220, thereby enabling the capture of audio (e.g., speech) produced by the encounter participant 226 (as the audio acquisition device 210 is pointed at (i.e., directed toward) the encounter participant 226). Further, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize the audio acquisition devices 204, 206 to form an audio recording beam 222, thereby enabling the capture of audio (e.g., speech) produced by the encounter participant 228 (as the audio acquisition devices 204, 206 are pointed at (i.e., directed toward) the encounter participant 228). Further, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize the audio acquisition devices 212, 214 to form an audio recording beam 224, thereby enabling the capture of audio (e.g., speech) produced by the encounter participant 230 (as the audio acquisition devices 212, 214 are pointed at (i.e., directed toward) the encounter participant 230). Further, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize null-steering precoding to cancel interference between speakers and / or noise.

[0055] As known in the art, null-steering precoding is a method of spatial signal processing by which a multi-antenna transmitter can null multi-user interference signals in wireless communications, where null-steering precoding can mitigate the effects of background noise and unknown user interference.

[0056] In particular, null-steering precoding can be a method of beamforming for narrowband signals that can compensate for the delay in receiving signals from a particular source at different elements of an antenna array. Generally, to improve the performance of an antenna array, incoming signals can be summed and averaged, where certain signals can be weighted and signal delays can be compensated for.

[0057] The machine vision system 100 and the audio recording system 104 can be independent devices (e.g., as shown in FIG. 1) or can be integrated into a single device (e.g., as shown in FIG. 2). Figure 4Additionally / alternatively, the machine vision system 100 and the audio recording system 104 can be combined into one package to form a hybrid media ACD device 232. For example, the hybrid media ACD device 232 can be configured to be mounted to a structure (e.g., a wall, a ceiling, a beam, a column) within the aforementioned clinical environments (e.g., a physician's office, a medical facility, a medical practice, a medical laboratory, an urgent care facility, a medical clinic, an emergency room, an operating room, a hospital, a long-term care facility, a rehabilitation facility, a nursing home, and a hospice facility), thereby allowing them to be easily installed. Moreover, the modular ACD system 54 can be configured to include multiple hybrid media ACD devices (e.g., the hybrid media ACD device 232) when a larger or higher level of resolution is needed in the aforementioned clinical environments.

[0058] The modular ACD system 54 can also be configured to direct one or more audio recording beams (e.g., the audio recording beams 220, 222, 224) to one or more encounter participants (e.g., the encounter participants 226, 228, 230) of a patient encounter based at least in part on the machine vision encounter information 102. As described above, the hybrid media ACD device 232 (and the machine vision system 100 / audio recording system 104 included therein) can be configured to monitor one or more encounter participants (e.g., the encounter participants 226, 228, 230) of a patient encounter.

[0059] In particular, the machine vision system 100 (as a standalone system or as a component of the hybrid media ACD device 232) can be configured to detect humanoid shapes within the aforementioned clinical environments (e.g., a physician's office, a medical facility, a medical practice, a medical laboratory, an urgent care facility, a medical clinic, an emergency room, an operating room, a hospital, a long-term care facility, a rehabilitation facility, a nursing home, and a hospice facility). And when the machine vision system 100 detects these humanoid shapes, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize one or more discrete audio acquisition devices (e.g., the audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) to form an audio recording beam (e.g., the audio recording beams 220, 222, 224) that is directed at each detected humanoid shape (e.g., the encounter participants 226, 228, 230).

[0060] As described above, the ACD computer system 12 can be configured to receive machine vision session information 102 and audio session information 106 from the machine vision system 100 and the audio recording system 104, respectively; and can be configured to provide visual information 110 and audio information 114 to the display presentation system 108 and the audio presentation system 112, respectively. Depending on how the modular ACD system 54 (and / or the hybrid media ACD device 232) is configured, the ACD computer system 12 can be included within the hybrid media ACD device 232 or external to the hybrid media ACD device 232.

[0061] As described above, the ACD computer system 12 can perform all or a portion of the automated clinical documentation process 10, where the set of instructions and subroutines of the automated clinical documentation process 10 (which can be stored on one or more of, for example, the storage devices 16, 20, 22, 24, 26) can be executed by the ACD computer system 12 and / or one or more ACD client electronic devices 28, 30, 32, 34.

[0062] The automated clinical documentation process:

[0063] In some implementations consistent with the present disclosure, systems and methods for speaker segmentation and diarization and distant speech recognition using spatial and spectral information can be provided. For example, distant conversation automatic speech recognition (DCASR) can include multiple microphone systems (e.g., an array of microphones and a single mobile microphone carried by a speaker) configured to record and recognize speech of one or more speakers. When utilizing signals from multiple microphones, traditional DCASR methods are subject to various challenges. For example, traditional beamforming techniques combine multiple microphone signals through spatial filtering, but these are mostly limited to a single microphone array device (i.e., they generally do not provide a mechanism for combining various microphones that are not part of an array, nor do they provide different microphone arrays deployed in a room as independent devices). Traditional techniques also allow for single-channel noise reduction and dereverberation techniques. However, these techniques do not take into account spatial information that can be available. Further, traditional DCASR methods are unable to manage the combination of enhanced signals from various devices in a room. In some implementations, traditional DCASR methods can encounter problems when localizing the position of a speaker at any given time, tracking their movements, and then recognizing their identity. This can result in less accurate diarized text transcriptions (i.e., text transcriptions with speaker labels to define who said what in a conversation when). Further, ASR systems that convert audio to text can be unable to handle various metadata and audio signals from multiple sources (e.g., multiple microphone systems).

[0064] Accordingly, implementations of the present disclosure can address these challenges experienced by conventional DCASR approaches by predefining beamforming configurations based at least in part on information associated with an acoustic environment; selecting particular beam patterns and / or null patterns for particular locations (e.g., where beams and nulls are selected to point to certain locations in a room, where some locations can have a high probability of being occupied by a particular speaker); using spatial and spectral information for voice activity detection and localization; using spatial and spectral information for speaker identification; and aligning audio streams from multiple microphone systems based at least in part on voice activity detection. For example, within the scope of the present disclosure, the automated clinical documentation process 10 can include various hardware and / or software modules configured to perform various DCASR functions.

[0065] Reference is also made to Figures 5-7 And in some implementations, the machine vision system 100 and the audio recording system 104 can be combined into one package to form a hybrid media ACD device 232. In this example, the ACD computer system 12 can be configured to receive machine vision encounter information 102 and audio encounter information 106 from the machine vision system 100 and the audio recording system 104, respectively. Specifically, the ACD computer system 12 can receive the audio encounter information at a voice activity detection (VAD) module (e.g., VAD module 400), an acoustic beamforming module (e.g., acoustic beamforming module 402), and / or a speaker identification and tracking module (e.g., speaker identification and tracking module 404). In some implementations, the ACD computer system 12 can also receive the machine vision encounter information 102 at the speaker identification and tracking module 404. In some implementations, the VAD module 400 can communicate with a sound localization module 406 to output metadata for the speaker identification and tracking module 404. In some implementations, the acoustic beamforming module 402 can be configured to provide a plurality of beams and / or a plurality of nulls to a beam and null selection module (e.g., beam and null selection module 408). The speaker identification and tracking module 404 can be configured to provide speaker identification and tracking information to the beam and null selection module 408.

[0066] In some implementations, the beam and null selection module 408 can be configured to provide the aligned audio streams from the audio recording system 104 to a device selection and weighting module (e.g., the device selection and weighting module 410). Further, the beam and null selection module 408 can be configured to provide the aligned audio streams to an alignment module (e.g., the alignment module 412). The alignment module 412 can be configured to receive audio encounter information from a VAD module associated with a second microphone system (e.g., the VAD module 414 associated with the mobile electronic device 416). In some implementations, the alignment module 412 can be configured to provide the aligned audio streams from the mobile electronic device 416 to the device selection and weighting module 410. In some implementations, the device selection and weighting module 410 can provide the aligned audio streams from the audio recording system 104, the aligned audio streams from the mobile electronic device 416, and metadata from other modules to one or more speech processing systems (e.g., the speech processing system 418). In some implementations, the speech processing system 418 can be an automatic speech recognition (ASR) system.

[0067] As will be discussed in greater detail below, the combination of the various modules of the ACD computer system 12 / automated clinical documentation process 10 can be configured to improve DCASR by providing aligned audio encounter information from multiple microphone systems. In this manner, the automated clinical documentation process 10 can use spatial and spectral information determined by the various modules of the ACD computer system 12 / automated clinical documentation process 10 to improve distant speech recognition.

[0068] Referring at least to Figure 6 , the automated clinical documentation process 10 can receive 500 information associated with an acoustic environment. A plurality of filters can be predefined 502 to produce a plurality of beams based at least in part on the information associated with the acoustic environment. A plurality of filters can be predefined 504 to produce a plurality of nulls based at least in part on the information associated with the acoustic environment. The plurality of beams and the plurality of nulls produced by the plurality of predefined filters can be used to obtain 506 audio encounter information via one or more microphone arrays.

[0069] As will be discussed in greater detail below, the automated clinical documentation process 10 can predefine or precompute a plurality of filters to produce beams and nulls to target specific talkers and / or noise sources to improve DCASR. As is known in the art, a beamformer can use a set of filters associated with a set of input signals to generate a spatially filtered sensitivity pattern formed by beams and nulls having characteristics determined by the filters. As will be discussed in greater detail below, the automated clinical documentation process 10 can receive 500 information associated with an acoustic environment and can predefine filters to produce beams and / or nulls to target or isolate specific talkers and / or noise sources. In this way, the automated clinical documentation process 10 can utilize acoustic environment information to predefine beams and nulls that can be selected for various situations (i.e., select certain beams and nulls when a patient is talking with a medical professional and select other beams and nulls when a medical professional is talking with a patient).

[0070] In some implementations, the automated clinical documentation process 10 can receive 500 information associated with an acoustic environment. The acoustic environment can represent a layout and acoustic properties of a room or other space in which a plurality of microphone systems can be deployed. For example, the information associated with the acoustic environment can describe a type / size of the room, activity zones in the room in which participants can be or operate, acoustic properties of the room (e.g., expected noise types, reverberation range, etc.), a location of the microphone system(s) in the acoustic environment, etc. In some implementations, the automated clinical documentation process 10 can provide a user interface to receive the information associated with the acoustic environment. Thus, a user and / or the automated clinical documentation process 10 can provide (e.g., via the user interface) the information associated with the acoustic environment. However, it should be appreciated that the information associated with the acoustic environment can be received in various ways (e.g., default information for a default acoustic environment, automatically defined by the automated clinical documentation process 10, etc.).

[0071] In some implementations, the information associated with the acoustic environment can indicate one or more target talker locations within the acoustic environment. Reference is also made to Figure 6 and in some implementations, assume that the acoustic environment 600 includes one or more consultation participants (e.g., consultation participants 226, 228, 230) of the patient consultation described above. As described above, examples of the consultation participants (e.g., consultation participants 226, 228, 230) can include, but are not limited to: a medical professional (e.g., a doctor, a nurse, a doctor’s assistant, a lab technician, a physical therapist, a scribe (e.g., a transcriptionist), and / or a staff member participating in the patient consultation), a patient (e.g., a person visiting the clinical environment described above for a patient consultation), and a third party (e.g., a friend of the patient participating in the patient consultation, a relative of the patient, and / or an acquaintance of the patient).

[0072] In some implementations, the automated clinical documentation process 10 can receive 500 information associated with the acoustic environment 600 that indicates the location of the visit participants 226, 228, 230. In some implementations, the information associated with the acoustic environment 600 can indicate the location within the acoustic environment that one or more target speakers are likely to be when speaking. For example, assume that the automated clinical documentation process 10 determines that the exam table is located at approximately, for example, 45° to the base of the array of microphones 200. In this example, the automated clinical documentation process 10 can determine that the patient is most likely to speak while sitting on or near the exam table at approximately 45° to the base of the array of microphones 200. Further assume that the automated clinical documentation process 10 determines that the doctor’s desk is located at approximately, for example, 90° to the base of the array of microphones 200. In this example, the automated clinical documentation process 10 can determine that the doctor is most likely to speak while sitting on or near the desk at approximately 90° to the base of the array of microphones 200. Additionally, assume that the automated clinical documentation process 10 determines that the waiting area is located at approximately, for example, 120° to the base of the array of microphones 200. In this example, the automated clinical documentation process 10 can determine that other patients or other third parties are most likely to speak from approximately 120° to the base of the array of microphones 200. While three examples of the relative location of target speakers within the acoustic environment have been provided, it should be understood that the information associated with the acoustic environment can include any number of target speaker locations or probability-based target speaker locations for any number of target speakers within the scope of the present disclosure.

[0073] In some implementations, the automated clinical documentation process 10 can predefine 502 a plurality of filters to produce a plurality of beams based at least in part on information associated with the acoustic environment. The beams can generally include a constructive interference pattern between microphones of a microphone array that is produced by modifying a phase and / or amplitude of a signal at each microphone of the microphone array via the plurality of filters. The constructive interference pattern can improve signal processing performance of the microphone array. Predefining 502 the plurality of filters to produce the plurality of beams can generally include defining the plurality of beams at any point in time prior to obtaining 506 audio encounter information via the one or more microphone arrays. The plurality of filters can include a plurality of finite impulse response (FIR) filters configured to adjust the phase and / or amplitude of the signal at each microphone of the microphone array. Accordingly, the automated clinical documentation process 10 can predefine a plurality of filters (e.g., a plurality of FIR filters) to produce a plurality of beams that are configured to“look” or orient in a particular direction based at least in part on information associated with the acoustic environment. In some implementations, the automated clinical documentation process 10 can predefine 502 the plurality of filters to produce the plurality of beams by adjusting the phase and / or amplitude of the signal at each microphone of the microphone array.

[0074] For example, the automated clinical documentation process 10 can receive 500 information associated with the acoustic environment and can determine how acoustic properties of the acoustic environment affect beams deployed within the acoustic environment. For example, assume that the information associated with the acoustic environment indicates that there can be a particular reverberation level at a particular frequency and varying amplitude at different portions of the acoustic environment. In this example, the automated clinical documentation process 10 can predefine 502 a plurality of filters to produce a plurality of beams to account for the layout, acoustic properties, etc. of the acoustic environment. As will be discussed in greater detail below, by predefining 502 the plurality of filters to produce the plurality of beams, the automated clinical documentation process 10 can allow for selection of particular beams for various situations.

[0075] In some implementations, predefining 502 the plurality of filters to produce the plurality of beams based at least in part on information associated with the acoustic environment can include predefining 508 the plurality of filters to produce one or more beams configured to receive audio encounter information from one or more target speaker locations within the acoustic environment. For example, again with reference to Figure 6The automated clinical documentation process 10 can predefine 508 a plurality of filters to produce one or more beams configured to receive audio encounter information from one or more target speaker locations within the acoustic environment. Continuing with the example above, assume that the automated clinical documentation process 10 determines, based on information associated with the acoustic environment 600, that the patient is most likely to speak while seated at or near the exam table that is approximately 45° from the base of the microphone array 200; the doctor is most likely to speak while seated at or near the desk that is approximately 90° from the base of the microphone array 200; and the other patient or other third party is most likely to speak from approximately 120° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can predefine 508 a plurality of filters to produce a beam 220 for receiving audio encounter information from the patient (e.g., the encounter participant 228); a beam 222 for receiving audio encounter information from the doctor (e.g., the encounter participant 226); and a beam 224 for receiving audio encounter information from the other patient / third party (e.g., the encounter participant 230). As will be discussed in greater detail below, the automated clinical documentation process 10 can select a particular beam to receive audio encounter information when a particular participant is speaking.

[0076] In some implementations, predefining 502 a plurality of filters to produce a plurality of beams based at least in part on information associated with the acoustic environment can include predefining 510 a plurality of filters to produce a plurality of frequency-independent beams based at least in part on information associated with the acoustic environment. In some implementations, the spatial sensitivity of a beam to audio encounter information can be frequency-dependent. For example, a sensitivity beam can be narrower for receiving high frequency signals, while a sensitivity beam can be wider for low frequency signals. Accordingly, the automated clinical documentation process 10 can predefine 510 a plurality of filters to produce a plurality of frequency-independent beams based at least in part on information associated with the acoustic environment. For example, and as described above, the automated clinical documentation process 10 can determine, based at least in part on information associated with the acoustic environment, that a speaker can be located at a particular location within the acoustic environment when speaking. Further, the automated clinical documentation process 10 can receive acoustic properties of the acoustic environment. For example, the automated clinical documentation process 10 can determine the location(s) and frequency characteristics of one or more noise sources within the acoustic environment. In this way, the automated clinical documentation process 10 can predefine 510 a plurality of filters to produce a plurality of frequency-independent beams that account for the acoustic properties of the acoustic environment while providing sufficient microphone sensitivity for frequency variations.

[0077] For example, the automated clinical documentation process 10 can predefine 502 a plurality of filters to produce beams 220, 222, 224 by modifying the phase and / or amplitude of the signal of each microphone such that the beams are sensitive enough to receive audio encounter information of a target speaker within the acoustic environment regardless of frequency. In this way, the automated clinical documentation process 10 can predefine 510 a plurality of filters to produce a plurality of frequency-independent beams (e.g., frequency-independent beams 220, 222, 224) based at least in part on information associated with the acoustic environment (e.g., acoustic environment 600).

[0078] In some implementations, the automated clinical documentation process 10 can predefine 504 a plurality of filters to produce a plurality of nulls based at least in part on information associated with the acoustic environment. A null can generally include a pattern of destructive interference between microphones of a microphone array produced by modifying the phase and / or amplitude of the signal at each microphone of the microphone array. The pattern of destructive interference can limit the reception of signals by the microphone array. Predefining 504 a plurality of filters to produce a plurality of nulls can generally include defining the plurality of nulls at any point in time prior to obtaining 506 audio encounter information via the one or more microphone arrays. As described above, the plurality of filters can include a plurality of finite impulse response (FIR) filters configured to adjust the phase and / or amplitude of the signal at each microphone of the microphone array. Accordingly, the automated clinical documentation process 10 can predefine a plurality of filters (e.g., a plurality of FIR filters) to produce a plurality of nulls configured to“look” or direct in a particular direction based at least in part on information associated with the acoustic environment. In contrast to beams, nulls can limit or attenuate reception in a target direction. In some implementations, the automated clinical documentation process 10 can predefine 504 a plurality of filters to produce a plurality of nulls by adjusting the phase and / or amplitude of the signal at each microphone of the microphone array.

[0079] For example, the automated clinical documentation process 10 can receive 500 information associated with the acoustic environment and can determine how the acoustic properties of the acoustic environment affect beams deployed within the acoustic environment. For example and as described above, assume that the information associated with the acoustic environment indicates that a particular noise signal (e.g., the sound of an air conditioning system) can exist at a particular frequency and at varying amplitudes in different portions of the acoustic environment. In this example, the automated clinical documentation process 10 can predefine 504 a plurality of filters to produce a plurality of nulls to limit the reception of the noise signal. By predefining 504 a plurality of filters to produce a plurality of nulls, the automated clinical documentation process 10 can allow for selection of a particular null for various situations, as will be discussed in greater detail below.

[0080] In some implementations, the automated clinical documentation process 10 can predefine 504 a plurality of filters to produce one or more nulls configured to limit receipt of noise signals from one or more noise sources based at least in part on information associated with the acoustic environment. For example, referring again to Figure 6 assuming the automated clinical documentation process 10 determines from information associated with the acoustic environment 600 that a noise source (e.g., the fan 244) is located at approximately, for example, 70° from the base of the microphone array 200, and a second noise source (e.g., the doorway to the busy hallway 246) is located at approximately, for example, 150° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can predefine 512 a plurality of filters to produce a null 248 to limit receipt of noise signals from the first noise source (e.g., the fan 244), and a null 250 to limit receipt of noise signals from the second noise source (e.g., the doorway to the busy hallway 246). While an example has been described of predefining 504 a plurality of filters to produce two nulls for two noise sources, it can be appreciated that within the scope of the present disclosure, the automated clinical documentation process 10 can predefine 504 a plurality of filters to produce any number of nulls for any number of noise sources. In some implementations, the automated clinical documentation process 10 can select particular nulls to limit receipt of noise signals for various situations.

[0081] In some implementations, predefining 504 a plurality of filters to produce a plurality of nulls based at least in part on information associated with the acoustic environment can include predefining 512 a plurality of filters to produce one or more nulls to limit receipt of audio patient encounter information from one or more target speaker locations within the acoustic environment. For example, referring again to Figure 7 the automated clinical documentation process 10 can predefine 512 a plurality of filters to produce one or more nulls to limit receipt of audio patient encounter information from one or more target speaker locations within the acoustic environment. Continuing the example above, assuming the automated clinical documentation process 10 determines from information associated with the acoustic environment 600 that a patient is most likely to speak while seated at or near the exam table at approximately 45° from the base of the microphone array 200; a physician is most likely to speak while seated at or near the desk at approximately 90° from the base of the microphone array 200; and other patients or other third parties are most likely to speak from approximately 120° from the base of the microphone array 200.

[0082] In this example and also referring to Figure 4In another example, assume that the patient (e.g., interviewee 228) and the physician (e.g., interviewer 226) spoke to each other for a short duration (i.e., "total speech"). The automated clinical documentation process 10 can form two output signals, where a first signal is the sum of a beam pattern directed toward the physician (e.g., interviewer 226) and a null directed toward the patient (e.g., interviewee 228); and a second signal is the sum of a beam pattern directed toward the patient (e.g., interviewee 228) and a null directed toward the physician (e.g., interviewer 226). In this way, the combination of the two signals formed by the beam and null defined for each signal can have a better chance of accurately transcribing the words spoken by the physician and the patient.

[0083] In another example, assume that the patient (e.g., interviewee 228) and the physician (e.g., interviewer 226) spoke to each other for a short duration (i.e., "total speech"). The automated clinical documentation process 10 can form two output signals, where a first signal is the sum of a beam pattern directed toward the physician (e.g., interviewer 226) and a null directed toward the patient (e.g., interviewee 228); and a second signal is the sum of a beam pattern directed toward the patient (e.g., interviewee 228) and a null directed toward the physician (e.g., interviewer 226). In this way, the combination of the two signals formed by the beam and null defined for each signal can have a better chance of accurately transcribing the words spoken by the physician and the patient.

[0084] In some implementations, the automated clinical documentation process 10 can obtain 506 audio interview information via one or more microphone arrays using a plurality of beams and a plurality of nulls produced by a plurality of predefined filters. As described above, in the example of an audio signal, the one or more microphone arrays can utilize the beams and nulls to receive audio interview information from a particular speaker and limit the reception of audio interview information from other speakers or other sound sources. For example, the automated clinical documentation process 10 can obtain the audio interview information 106 by combining a plurality of discrete microphone elements (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) in a microphone array (e.g., microphone array 200) with a plurality of predefined filters in such a way that signals of a particular angle experience constructive interference while other signals experience destructive interference.

[0085] In some implementations, obtaining 506 audio encounter information via the one or more microphone arrays using the plurality of beams and the plurality of nulls produced by the plurality of predefined filters can include adaptively steering 514 the plurality of beams and the plurality of nulls. For example, the automated clinical documentation process 10 can allow the plurality of beams and / or the plurality of nulls to be steered to a target speaker within the acoustic environment. Returning to the example above, assume that the patient (e.g., the encounter participant 228) is speaking while moving within the acoustic environment 600. Further assume that the patient (e.g., the encounter participant 228) moves from approximately 45° relative to the base of the microphone array 200 to approximately 35° relative to the base of the microphone array 200, for example. In this example, the automated clinical documentation process 10 can adaptively steer 514 the beam (e.g., the beam 220) to follow the patient (e.g., the encounter participant 228) as the patient moves within the acoustic environment 600. Moreover, assume that the physician (e.g., the encounter participant 226) begins moving within the acoustic environment 600 while the patient (e.g., the encounter participant 228) is speaking. In this example, the automated clinical documentation process 10 can adaptively steer 514 the null (e.g., the null 254) to follow the physician (e.g., the encounter participant 226) to limit any noise signals or audio encounter information from the physician (e.g., the encounter participant 226) while the patient (e.g., the encounter participant 228) is speaking. While an example of adaptively steering one beam and one null has been provided, it will be appreciated that any number of beams and / or nulls can be adaptively steered for various purposes in the present disclosure.

[0086] In some implementations, obtaining 506 audio encounter information from a particular speaker and limiting audio encounter information from other speakers can include one or more of: selecting 516 one or more beams from the plurality of beams; and selecting 518 one or more nulls from the plurality of nulls. For example and as described above, the automated clinical documentation process 10 can predefine a plurality of filters to produce a plurality of beams and / or a plurality of nulls based at least in part on information associated with the acoustic environment. Referring again to Figures 8-10 These actions can generally be performed by or associated with the acoustic beamforming module 402, in this manner, the automated clinical documentation process 10 can provide the predefined filters to the beam and null selection module 508. As will be discussed in greater detail below, the beam and null selection module 508 can be configured to use the plurality of predefined filters to select 516 one or more beams from the plurality of beams and / or to select 518 one or more nulls from the plurality of nulls for various situations.

[0087] For example and as described above, assume that a patient (e.g., the visit participant 228) begins speaking. The automated clinical documentation process 10 can detect the audio visit information of the patient, and can select 516 one or more beams (e.g., the beam 220) to receive the audio visit information from the patient (e.g., the visit participant 228). In addition, the automated clinical documentation process 10 can detect speech from another participant (e.g., the participant 230). In one example, the automated clinical documentation process 10 can select 516 one or more beams (e.g., the beam 224) to receive the audio visit information from the patient 230, or can select 518 one or more nulls (e.g., the null 256) to limit the reception of the audio visit information from the participant 230. In this way, when audio visit information from a particular speaker is detected, the automated clinical documentation process 10 can select the beams and nulls to utilize from a predefined plurality of beams and nulls produced by the filters.

[0088] Referring at least to Figure 4 , the automated clinical documentation process 10 can receive 800 audio visit information from a microphone array. Speech activity within one or more portions of the audio visit information can be identified 802 based at least in part on a correlation between the audio visit information received from the microphone array. Location information for the one or more portions of the audio visit information can be determined 804 based at least in part on a correlation between signals received by each microphone of the microphone array. The one or more portions of the audio visit information can be tagged 806 with the speech activity and the location information.

[0089] In some implementations, the automated clinical documentation process 10 can receive 800 audio visit information from a microphone array. Referring again to Figure 9 , and in some implementations, the microphone array (e.g., the microphone array 200) can include a plurality of microphones (e.g., the audio capture devices 202, 204, 206, 208, 210, 212, 214, 216, 218) configured to receive audio visit information (e.g., the audio visit information 106). As described above, the audio visit information 106 can include speech signals or other signals recorded by the plurality of audio capture devices. In some implementations, the automated clinical documentation process 10 can receive the audio visit information 106 at a voice activity detection (VAD) module (e.g., the VAD module 400) and / or a sound localization module (e.g., the sound localization module 406) of the ACD computer system 12.

[0090] Referring also to Figure 9assuming that the automated clinical documentation process 10 receives audio encounter information 106 at multiple microphones (e.g., audio capture devices 202, 204, 206, 208) of the microphone array 200. While only four audio capture devices are shown, it should be understood that this is for ease of explanation and that any number of audio capture devices of a microphone array can record audio encounter information within the scope of the present disclosure. In this example, assume that each microphone receives an audio signal (e.g., audio encounter information 106). Based on the orientation of each microphone, the properties of each microphone, the properties of the audio signal, etc., each microphone can receive a different version of the signal. For example, the audio encounter information received by microphone 202 can have different signal components (i.e., amplitude and phase) than the audio encounter information received by microphones 204, 206, and / or 208. As will be discussed in greater detail below, the automated clinical documentation process 10 can utilize the correlation between the audio encounter information to detect speech and determine location information associated with the audio encounter information.

[0091] In some implementations, the automated clinical documentation process 10 can identify 802 speech activity within one or more portions of the audio encounter information based at least in part on the correlation between the audio encounter information received by the microphone array. Again referring to the example of Figure 9 In some implementations, the audio encounter information 106 can include multiple portions or frames of audio information (e.g., portions 900, 902, 904, 906, 908, 910, 912, 914, 916, 918, 920, 922, 924, 926), and in one example, each portion or frame can represent a predefined amount of time (e.g., 20 milliseconds) of the audio encounter information 106. While an example of, for example, 14 portions of the audio encounter information 106 has been described, it can be understood that this is for example purposes only and that within the scope of the present disclosure, the audio encounter information 106 can include or be defined as any number of portions.

[0092] In some implementations, the automated clinical documentation process 10 can determine a correlation between audio encounter information received by the microphone array 200. For example, the automated clinical documentation process 10 can compare one or more portions of the audio encounter information 106 (e.g., portions 900, 902, 904, 906, 908, 910, 912, 914, 916, 918, 920, 922, 924, 926) to determine a degree of correlation between audio encounter information present in each of the portions across the plurality of microphones of the microphone array. In some implementations, the automated clinical documentation process 10 can perform various cross-correlation processes known in the art to determine a degree of similarity between one or more portions of the audio encounter information 106 (e.g., portions 900, 902, 904, 906, 908, 910, 912, 914, 916, 918, 920, 922, 924, 926) across the plurality of microphones of the microphone array 200.

[0093] For example, assume that the automated clinical documentation process 10 receives only environmental noise (i.e., no speech and no directional noise source). The automated clinical documentation process 10 can determine that the spectrum observed in each microphone channel is different (i.e., uncorrelated at each microphone). However, assume that the automated clinical documentation process 10 receives speech or other “directional” signals within the audio encounter information. In this example, the automated clinical documentation process 10 can determine that one or more portions of the audio encounter information (e.g., portions of the audio encounter information having speech components) are highly correlated at each microphone of the microphone array.

[0094] In some implementations, the automated clinical documentation process 10 can identify 802 speech activity within one or more portions of the audio encounter information based at least in part on determining a threshold amount or degree of correlation between audio encounter information received from the microphone array. For example and as described above, various thresholds can be defined (e.g., user-defined, default thresholds, automatically defined via the automated clinical documentation process 10, etc.) to determine when portions of audio encounter information are sufficiently correlated. Thus, in response to determining at least a threshold degree of correlation between portions of audio encounter information across a plurality of microphones, the automated clinical documentation process 10 can determine or identify speech activity within one or more portions of the audio encounter information. In conjunction with determining a threshold degree of correlation between portions of audio encounter information, the automated clinical documentation process 10 can use other methods known in the art for voice activity detection (VAD) to identify speech activity, such as filtering, noise reduction, applying classification rules, etc. In this way, conventional VAD techniques can be used in conjunction with determining a threshold correlation between one or more portions of audio encounter information to identify 802 speech activity within one or more portions of the audio encounter information.

[0095] Refer again Figure 6 For example, suppose the automated clinical documentation process 10 determines at least a threshold amount or degree of correlation between portions 900, 902, 904, 906, 908, 910, and 912 of audio consultation information 106 across microphones 202, 204, 206, and 208, and a lack of correlation between portions 914, 916, 918, 920, 922, 924, and 926 of the audio consultation information 106. In this example, the automated clinical documentation process 10 may identify speech activity within portions 900, 902, 904, 906, 908, 910, and 912 of the audio consultation information 106, at least in part, based on the threshold of correlation between portions 900, 902, 904, 906, 908, 910, and 912 of the audio consultation information 106. In some implementations, identifying speech activity within one or more portions of the 802 audio consultation information may include generating timestamps (e.g., start and end times for each portion) indicating the portions of the audio consultation information 106 that include speech activity. In some implementations, the automated clinical documentation process 10 may generate a vector of start and end times for each portion with detected speech activity. In some implementations, and as will be discussed in more detail below, the automated clinical documentation process 10 may label speech activity as a time-domain label (i.e., a set of samples of the signal including speech or speech) or a frequency-domain label set (i.e., a vector giving the probability that a specific frequency band in a specific time frame includes speech or speech).

[0096] In some implementations, the automated clinical documentation process 10 may determine location information for one or more portions of the audio patient information 804, at least in part, based on the correlation between signals received by each microphone of the microphone array. Location information may typically include the relative location or orientation of the source of the audio patient information within the acoustic environment. For example, the automated clinical documentation process 10 may determine location information associated with the source of the signals received by the microphone array 200. (See again...) Figure 4 Furthermore, in some implementations, the automated clinical documentation process 10 can receive audio patient information from various sources (e.g., patient participants 226, 228, 230; noise sources 244, 250; etc.). In some implementations, the automated clinical documentation process 10 can utilize information associated with the microphone array (e.g., the spacing between microphones in the microphone array; the positioning of microphones within the microphone array; etc.) to determine the location information associated with the audio patient information received by the microphone array at 804.

[0097] As described above, and in some implementations, the automated clinical documentation process 10 can determine a correlation between audio encounter information received from the microphone array 200. For example, the automated clinical documentation process 10 can compare one or more portions of the audio encounter information 106 received from individual microphones of the microphone array 200 (e.g., portions 900, 902, 904, 906, 908, 910, 912, 914, 916, 918, 920, 922, 924, 926) to determine a degree of correlation between the audio encounter information present in each portion across the plurality of microphones of the microphone array.

[0098] In some implementations, determining 804 the location information for the one or more portions of audio encounter information can include determining 808 a time difference of arrival between each pair of microphones of the microphone array for the one or more portions of audio encounter information. As known in the art, a time difference of arrival (TDOA) between a pair of microphones can include locating a source of a signal based at least in part on different times of arrival of the signal at individual receivers. In some implementations, the automated clinical documentation process 10 can determine 808 a time difference of arrival (TDOA) between each pair of microphones for the one or more portions of audio encounter information based at least in part on a correlation between the audio encounter information received by the microphone array.

[0099] For example, and in some implementations, highly correlated signals (e.g., signals having at least a threshold degree of correlation) can allow for accurate determination of a time difference between microphones. As described above, assume that the automated clinical documentation process 10 receives only ambient noise (i.e., no speech and no directional noise sources). The automated clinical documentation process 10 can determine that the spectrum observed in each microphone channel will be different (i.e., uncorrelated at each microphone). Thus, based on the lack of correlation between the microphone channels of the microphone array, a time difference between pairs of microphones can be difficult to determine and / or can not be accurate. However, assume that the automated clinical documentation process 10 receives speech or other “directive” signals within the audio encounter information. In this example, the automated clinical documentation process 10 can determine that one or more portions of the audio encounter information (e.g., portions of the audio encounter information having speech components) are highly correlated at each microphone of the microphone array. Thus, the automated clinical documentation process 10 can use the highly correlated audio encounter information across the microphone channels of the microphone array to more accurately determine a time difference of arrival between pairs of microphones.

[0100] In some implementations, identifying 802 the voice activity and determining 804 the location information can be jointly performed for one or more portions of the audio encounter information. For example and as described above, the correlation between portions of the audio encounter information can identify the presence of voice activity within one or more portions of the audio encounter information received by the microphone array. Further, and as described above, the correlation between portions of the audio encounter information can allow for more accurate determination of location information (e.g., time difference of arrival) between pairs of microphones of the microphone array. Accordingly, the automated clinical documentation process 10 can determine a correlation between portions of the audio encounter information across microphones or microphone channels of the microphone array and can utilize the determined correlation (e.g., by comparing the correlation to a threshold for voice activity detection and a threshold for time difference of arrival determination) to jointly identify 802 voice activity within one or more audio portions and determine 804 a time difference of arrival between pairs of microphones.

[0101] In some implementations, the joint identification 802 of voice activity and determination 804 of location information for one or more portions of the audio encounter information can be performed using a machine learning model. In some implementations, the automated clinical documentation process 10 can train a machine learning model to“learn” how to jointly identify voice activity and determine location information for one or more portions of the audio encounter information. As is known in the art, machine learning models generally can include an algorithm or combination of algorithms that are trained to identify certain types of patterns. For example, depending on the nature of the available signals, machine learning methods generally can be categorized into three types: supervised learning, unsupervised learning, and reinforcement learning. As is known in the art, supervised learning can include presenting an example input and its desired output to a computing device, which is given by a“teacher,” where the goal is to learn a general rule that maps inputs to outputs. In the case of unsupervised learning, the learning algorithm is not given labels, leaving it to find structure on its own in the inputs. Unsupervised learning can be an end in itself (discovering hidden patterns in the data) or a means to an end (feature learning) that can be used in supervised learning. As is known in the art, reinforcement learning generally can include a computing device that interacts with a dynamic environment, in which the computing device must perform a particular goal (e.g., driving a vehicle or playing a game against an opponent). As the program navigates its own problem space, the program is provided feedback similar to rewards that it tries to maximize. While three examples of machine learning methods have been provided, it will be appreciated that other machine learning methods are possible within the scope of the present disclosure.

[0102] Accordingly, the automated clinical documentation process 10 can utilize a machine learning model (e.g., the machine learning model 420) to identify speech activity and location information within one or more portions of the audio encounter information. For example, the automated clinical documentation process 10 can provide time-domain waveform data or frequency-domain features as input to the machine learning model 420. In some implementations, phase-spectrum features can provide spatial information for identifying speech activity and location information of one or more portions of the audio encounter information.

[0103] For example, the automated clinical documentation process 10 can provide training data to the machine learning model 420 associated with various portions of the audio encounter information pre-labeled with speech activity information and / or location information. The automated clinical documentation process 10 can train the machine learning model 420 to identify 802 speech activity and determine 804 location information of one or more portions of the audio encounter information based at least in part on a degree of correlation between the portions of the audio encounter information and the training data. In this way, the machine learning model 420 can be configured to learn how a degree of correlation between the portions of the audio encounter information maps to speech activity and accurate location information within the portions of the audio encounter information. Accordingly, and as shown, the automated clinical documentation process 10 can implement the VAD module 400 and the sound localization module 406 with the machine learning model 420. Figure 6

[0104] In some implementations, the automated clinical documentation process 10 can receive 810 information associated with an acoustic environment. As described above, the acoustic environment can represent a layout and acoustic properties of a room or other space in which a plurality of microphone systems can be deployed. For example, the information associated with the acoustic environment can describe a type / size of the room, activity zones in the room in which participants operate, acoustic properties of the room (e.g., expected noise types, reverberation range, etc.), a location of the microphone system(s) within the acoustic environment, etc. In some implementations, the automated clinical documentation process 10 can provide a user interface to receive the information associated with the acoustic environment. Accordingly, a user and / or the automated clinical documentation process 10 can provide the information associated with the acoustic environment via the user interface. However, it should be understood that the information associated with the acoustic environment can be received in various ways (e.g., default information for a default acoustic environment, automatically defined by the automated clinical documentation process 10, etc.).

[0105] In some implementations, the information associated with the acoustic environment can indicate one or more target speaker locations within the acoustic environment. Again referring to FIG. 4, the automated clinical documentation process 10 can receive 812 information associated with the acoustic environment. For example, the automated clinical documentation process 10 can receive 812 information associated with the acoustic environment from a user via a user interface, from a default acoustic environment, etc. Figure 10 ​and in some implementations, assume that the acoustic environment 600 includes one or more visit participants (e.g., visit participants 226, 228, 230) of the above-described patient visit. As described above, examples of visit participants (e.g., visit participants 226, 228, 230) can include, but are not limited to: medical professionals (e.g., doctors, nurses, physician assistants, laboratory technicians, physical therapists, scribes (e.g., transcriptionists), and / or staff participating in the patient visit), patients (e.g., persons visiting the above-described clinical environment for a patient visit), and third parties (e.g., friends of the patient participating in the patient visit, relatives of the patient, and / or acquaintances of the patient).

[0106] In some implementations, the automated clinical documentation process 10 can receive 810 information associated with the acoustic environment 600, which can indicate locations within the acoustic environment 600 where the visit participants 226, 228, 230 can be located. In some implementations, the information associated with the acoustic environment 600 can indicate locations within the acoustic environment where one or more target speakers are likely to be located when speaking. For example, assume that the automated clinical documentation process 10 determines that the exam table is located at a position of, for example, about 45° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that the patient is most likely to speak while sitting on or near the exam table at about 45° from the base of the microphone array 200. Further assume that the automated clinical documentation process 10 determines that the doctor’s desk is located at a position of, for example, about 90° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that the doctor is most likely to speak while sitting on or near the desk at about 90° from the base of the microphone array 200. Additionally, assume that the automated clinical documentation process 10 determines that the waiting area is located at a position of, for example, about 120° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that other patients or other third parties are most likely to speak from at about 120° from the base of the microphone array 200. While three examples of relative locations of target speakers within the acoustic environment have been provided, it should be understood that the information associated with the acoustic environment can include any number of target speaker locations or probability-based target speaker locations for any number of target speakers within the scope of the present disclosure.

[0107] In some implementations, identifying 802 the speech activity within the one or more portions of audio encounter information can be based at least in part on location information for the one or more portions of audio encounter information and information associated with the acoustic environment. For example, the automated clinical documentation process 10 can utilize the information associated with the acoustic environment to more accurately identify the speech activity within the one or more portions of audio encounter information. In one example, assume that the automated clinical documentation process 10 receives information associated with the acoustic environment that indicates a location where a speaker can be in the acoustic environment. In this example, the automated clinical documentation process 10 can determine whether the one or more portions of audio encounter information are within the potential speaker location defined by the acoustic environment information based on the location information for the one or more portions of audio encounter information. In response to determining that the one or more portions of audio encounter information originated from the potential speaker location within the acoustic environment, the automated clinical documentation process 10 can determine a higher probability that the one or more portions of audio encounter information include speech activity. Thus, the automated clinical documentation process 10 can utilize the location information and the acoustic environment information to identify 802 the speech activity within the one or more portions of audio encounter information.

[0108] In some implementations, the automated clinical documentation process 10 can tag 806 the one or more portions of audio encounter information with the speech activity and location information. Tagging the one or more portions of audio encounter information with the speech activity and location information can generally include generating metadata for the one or more portions of audio encounter information utilizing the speech activity and location information associated with each respective portion of audio encounter information. Also refer to Figure 4and in some implementations, the automated clinical documentation process 10 can tag 806 the portions 900, 902, 904, 906, 906, 908, 910, 912 with speech activity and location information associated with each respective portion by generating acoustic metadata (e.g., acoustic metadata 1000, 1002, 1004, 1006, 1008, 1010, 1012) for each portion of the audio encounter information 106. In some implementations, acoustic metadata associated with audio encounter information can include speech activity and location information associated with the audio encounter information. In some implementations, the automated clinical documentation process 10 can generate acoustic metadata 1000, 1002, 1004, 1006, 1008, 1010, 1012 that identifies portions (e.g., portions 900, 902, 904, 906, 908, 910, 912) of the audio encounter information 106 that include speech activity. In one example, the automated clinical documentation process 10 can generate acoustic metadata with timestamps that indicate portions (e.g., start and end times for each portion) of the audio encounter information 106 that include speech activity. In some implementations, the automated clinical documentation process 10 can tag 806 speech activity as a time domain tag (i.e., a set of samples of a signal include speech or are speech) or a set of frequency domain tags (i.e., a vector giving the likelihood that a particular frequency bin in a particular time frame includes speech or is speech).

[0109] Referring again to Figures 11-14 and as will be discussed in more detail below, the automated clinical documentation process 10 can provide acoustic metadata (e.g., acoustic metadata 1014) to a speaker identification and tracking module 404 to identify speakers from the audio encounter information and / or to track locations of speakers within the acoustic environment based at least in part on the acoustic metadata 1014.

[0110] Referring again to Figure 6 , the automated clinical documentation process 10 can receive 1100 information associated with an acoustic environment. Acoustic metadata associated with audio encounter information received by a first microphone system can be received 1102. One or more speaker representations can be defined 1104 based at least in part on the acoustic metadata associated with the audio encounter information and the information associated with the acoustic environment. One or more portions of the audio encounter information can be tagged 1106 with the one or more speaker representations and speaker locations within the acoustic environment.

[0111] In some implementations, the automated clinical documentation process 10 can receive 1100 information associated with an acoustic environment. As described above, the acoustic environment can represent a layout and acoustic properties of a room or other space in which a plurality of microphone systems can be deployed. For example, the information associated with the acoustic environment can describe a type / size of the room, activity zones in the room in which participants operate, acoustic properties of the room (e.g., expected noise types, reverberation ranges, etc.), locations of the microphone system(s) within the acoustic environment, etc. In some implementations, the automated clinical documentation process 10 can provide a user interface to receive the information associated with the acoustic environment. Thus, a user and / or the automated clinical documentation process 10 can provide the information associated with the acoustic environment via the user interface. However, it should be appreciated that the information associated with the acoustic environment can be received in a variety of ways (e.g., default information for a default acoustic environment, automatically defined by the automated clinical documentation process 10, etc.).

[0112] In some implementations, the information associated with the acoustic environment can indicate one or more target speaker locations within the acoustic environment. Again referring to Figure 4 and in some implementations, assume that the acoustic environment 600 includes one or more visit participants (e.g., visit participants 226, 228, 230) of the patient visit described above. As described above, examples of visit participants (e.g., visit participants 226, 228, 230) can include, but are not limited to: medical professionals (e.g., doctors, nurses, physician assistants, lab technicians, physical therapists, scribes (e.g., transcriptionists), and / or staff participating in the patient visit), patients (e.g., people visiting the clinical environment described above for a patient visit), and third parties (e.g., friends of the patient participating in the patient visit, relatives of the patient, and / or acquaintances of the patient).

[0113] In some implementations, the automated clinical documentation process 10 can receive 1100 information associated with the acoustic environment 600, which can indicate locations within the acoustic environment 600 where the visit participants 226, 228, 230 can be located. In some implementations, the information associated with the acoustic environment 600 can indicate locations within the acoustic environment where one or more target speakers can be located when speaking. For example, assume that the automated clinical documentation process 10 determines that the exam table is located at approximately, e.g., 45°, from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that the patient is most likely to speak while sitting on or near the exam table at approximately 45° from the base of the microphone array 200. Further assume that the automated clinical documentation process 10 determines that the doctor’s desk is located at approximately, e.g., 90°, from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that the doctor is most likely to speak while sitting on or near the desk at approximately 90° from the base of the microphone array 200.

[0114] Further, assume that the automated clinical documentation process 10 determines that the waiting area is located at approximately, e.g., 120°, from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can determine that other patients or other third parties are most likely to speak from approximately 120° from the base of the microphone array 200. While three examples of relative locations of target speakers within an acoustic environment have been provided, it should be understood that information associated with an acoustic environment can include any number of target speaker locations or probability-based target speaker locations for any number of target speakers within the scope of the present disclosure.

[0115] In some implementations, the automated clinical documentation process 10 can receive 1102 acoustic metadata associated with the audio visit information received by the first microphone system. Again referring to Figure 12 and in some implementations, the VAD module 400 and the sound localization module 406 can generate acoustic metadata associated with the audio visit information. For example, and in some implementations, the acoustic metadata associated with the audio visit information can include voice activity information and signal location information associated with the audio visit information. As described above, and in some implementations, the automated clinical documentation process 10 can identify portions of the audio visit information (e.g., the audio visit information 106) that have speech components. The automated clinical documentation process 10 can associate or label the portions of the audio visit information 106 as having voice activity. In some implementations, the automated clinical documentation process 10 can generate acoustic metadata that identifies the portions of the audio visit information 106 that include voice activity.

[0116] In one example, the automated clinical documentation process 10 can generate timestamps indicating portions of the audio encounter information 106 that include speech activity (e.g., a start and end time for each portion). Referring also to Figure 12 , the automated clinical documentation process 10 can receive acoustic metadata 1200 for one or more portions of the audio encounter information 106 (e.g., portions 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224, 1226, 1228). As described above, the automated clinical documentation process 10 can generate acoustic metadata having speech activity and signal location information associated with each portion of the audio encounter information 106 (e.g., in Figure 6 , respectively, as acoustic metadata 1230, 1232, 1234, 1236, 1238, 1240, 1242, 1244, 1246, 1248, 1250, 1252, 1254, 1256 for portions 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224, 1226, 1228 of the audio encounter information 106).

[0117] Further, the automated clinical documentation process 10 can determine location information for the audio encounter information 106. For example, the automated clinical documentation process 10 can determine spatial information from the microphone array (e.g., time difference of arrival (TDOA) between microphones). As known in the art, once a signal is received at two reference points, the time difference of arrival can be used to calculate the difference in distance between the target and the two reference points. In this example, the automated clinical documentation process 10 can determine the time difference of arrival (TDOA) between microphones of the microphone array. In some implementations, the automated clinical documentation process 10 can generate acoustic metadata having TDOA for a particular microphone array.

[0118] In some implementations, the automated clinical documentation process 10 can define one or more speaker representations based at least in part on acoustic metadata associated with the audio encounter information and information associated with the acoustic environment. A speaker representation can generally include a cluster of data associated with a unique speaker within the acoustic environment. For example, the automated clinical documentation process 10 can cluster spatial information and spectral information into separate speaker representations to account for a combination of spatial and spectral data referring to a unique speaker in the acoustic environment.

[0119] Referring again to Figure 13and in some implementations, assume that the automated clinical documentation process 10 receives 1100 information associated with the acoustic environment 600, which can indicate locations at which the visit participants 226, 228, 230 can be located within the acoustic environment 600. Further assume that the automated clinical documentation process 10 receives 1102 acoustic metadata (e.g., acoustic metadata 1200) associated with the audio visit information 106 received by the microphone array 200. In this example, assume that the acoustic metadata 1200 includes spatial information (e.g., location information of the audio visit information 106) and spectral information (e.g., voice activity information). In some implementations, the spectral information can include acoustic features (e.g., MEL frequency cepstral coefficients (MFCCs)) associated with particular speakers (e.g., the visit participants 226, 228, 230).

[0120] Also refer to Figure 2 and in some implementations, the automated clinical documentation process 10 can define 1104 speaker representations for the visit participants based at least in part on the acoustic metadata associated with the audio visit information and the information associated with the acoustic environment. For example, the automated clinical documentation process 10 can use a dynamic model of speaker movement and the information associated with the acoustic environment to cluster the spatial information (e.g., TDOAs) and the spectral information (e.g., acoustic features such as MFCCs) into separate speaker representations (e.g., a speaker representation 1300 for the visit participant 226; a speaker representation 1302 for the visit participant 228; and a speaker representation 1304 for the visit participant 230). While the above example includes defining, for example, three speaker representations, it can be appreciated that any number of speaker representations can be defined by the automated clinical documentation process 10 within the scope of the present disclosure.

[0121] In some implementations, the automated clinical documentation process 10 can receive 1108 visual metadata associated with one or more visit participants within the acoustic environment. For example, the automated clinical documentation process 10 can be configured to track movements and / or interactions of human-like shapes within a monitored space (e.g., the acoustic environment 600) during a patient visit (e.g., a visit to a doctor’s office). Thus, again refer to Figure 13 the automated clinical documentation process 10 can process machine vision visit information (e.g., the machine vision visit information 102) to identify one or more human-like shapes. As noted above, examples of the machine vision system 100 (and in particular, the ACD client electronic device 34) can generally include, but are not limited to, one or more of an RGB imaging system, an infrared imaging system, an ultraviolet imaging system, a laser imaging system, a sonar imaging system, a radar imaging system, and a thermal imaging system.

[0122] When the ACD client electronic device 34 includes a visible light imaging system (e.g., an RGB imaging system), the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by recording motion video in the visible light spectrum of the various objects. When the ACD client electronic device 34 includes a non-visible light imaging system (e.g., a laser imaging system, an infrared imaging system, and / or an ultraviolet imaging system), the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by recording motion video in the non-visible light spectrum of the various objects. When the ACD client electronic device 34 includes an X-ray imaging system, the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by recording energy in the X-ray spectrum of the various objects. When the ACD client electronic device 34 includes a sonar imaging system, the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by emitting sound waves that can reflect off of the various objects. When the ACD client electronic device 34 includes a radar imaging system, the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by emitting radio waves that can reflect off of the various objects. When the ACD client electronic device 34 includes a thermal imaging system, the ACD client electronic device 34 can be configured to monitor the various objects within the acoustic environment 600 by tracking thermal energy of the various objects.

[0123] As described above, the ACD computer system 12 can be configured to access one or more data sources 118 (e.g., a plurality of separate data sources 120, 122, 124, 126, 128), examples of which can include, but are not limited to, one or more of the following: a user profile data source, a voiceprint data source, a sound characteristic data source (e.g., for adapting automatic speech recognition models), a faceprint data source, an anthropomorphic shape data source, a discourse identifier data source, a wearable token identifier data source, an interaction identifier data source, a medical condition symptom data source, a prescription compatibility data source, a medical insurance coverage data source, and a home health data source.

[0124] Accordingly, and when processing the machine vision encounter information (e.g., the machine vision encounter information 102) to identify one or more anthropomorphic shapes, the automated clinical documentation process 10 can be configured to compare the anthropomorphic shapes defined within the one or more data sources 118 to the potential anthropomorphic shapes within the machine vision encounter information (e.g., the machine vision encounter information 102).

[0125] When processing machine vision encounter information (e.g., machine vision encounter information 102) to identify one or more anthropomorphic shapes, the automated clinical documentation process 10 can track movement of the one or more anthropomorphic shapes within a monitored space (e.g., acoustic environment 600). For example, and when tracking movement of the one or more anthropomorphic shapes within acoustic environment 600, the automated clinical documentation process 10 can add a new anthropomorphic shape to the one or more anthropomorphic shapes when the new anthropomorphic shape enters the monitored space (e.g., acoustic environment 600), and / or can remove an existing anthropomorphic shape from the one or more anthropomorphic shapes when the existing anthropomorphic shape exits the monitored space (e.g., acoustic environment 600).

[0126] Further, and when tracking movement of the one or more anthropomorphic shapes within acoustic environment 600, the automated clinical documentation process 10 can monitor trajectories of various anthropomorphic shapes within acoustic environment 600. Thus, assume that, upon exiting acoustic environment 600, encounter participant 242 walks in front of (or behind) encounter participant 226. Since the automated clinical documentation process 10 is monitoring trajectories of encounter participant 242 (e.g., moving from left to right) and encounter participant 226 (e.g., stationary), the identities of these two anthropomorphic shapes can not be confused by the automated clinical documentation process 10 when encounter participant 242 passes in front of (or behind) encounter participant 226.

[0127] The automated clinical documentation process 10 can be configured to obtain encounter information for a patient encounter (e.g., visiting a physician’s office), which can include machine vision encounter information 102 (in the manner described above) and / or audio encounter information 106. In some implementations, the automated clinical documentation process 10 can generate visual metadata associated with machine vision encounter information 102. For example, and as described above, the automated clinical documentation process 10 can generate visual metadata 1306 that indicates a direction or location of a speaker within acoustic environment 600, a number of speakers within acoustic environment 600, and / or an identity of a speaker within acoustic environment 600. As Figure 6 As shown in the example of FIG. 13, visual metadata 1306 can be defined for each portion of audio encounter information 106 (e.g., portions 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224, 1226, 1228 of audio encounter information 106) and represented as separate visual metadata for each portion (e.g., visual metadata 1308, 1310, 1312, 1314, 1316, 1318, 1320, 1322, 1324, 1326, 1328, 1330, 1332, 1334).

[0128] In some implementations, defining 1104 the one or more speaker representations can include defining 1110 the one or more speaker representations based at least in part on visual metadata associated with one or more appointment participants within the acoustic environment. For example and as described above, the machine vision system 100 can detect and track a location estimate for a particular speaker representation by tracking a human-like shape within the acoustic environment. The automated clinical documentation process 10 can “fuse” the location estimate from the visual metadata with the acoustic positioning information to cluster the spatial-spectral features of the visual metadata and the acoustic metadata into the plurality of speaker representations.

[0129] Referring again to the example of FIGS. 6-8, Figure 14 In this example, the visual metadata 1306 can indicate a relative location of each speaker within the acoustic environment 600; a number of speakers; and / or an identity of the speakers within the acoustic environment 600. For example and as described above, the machine vision system 100 can be configured to identify one or more human-like shapes and track movement of the one or more human-like shapes within the acoustic environment. The automated clinical documentation process 10 can receive visual metadata 1306 associated with an identity and / or a location of the appointment participants 226, 228, 230. Assume that the visual metadata 1306 includes an identity of each participant (e.g., based at least in part on comparing human-like shapes defined within one or more data sources 118 to potential human-like shapes within the machine vision appointment information (e.g., the machine vision appointment information 102)).

[0130] Further assume that the automated clinical documentation process 10 receives 1102 acoustic metadata 1200 associated with the audio appointment information 106. The automated clinical documentation process 10 can define 1110 one or more speaker representations (e.g., speaker representations 1300, 1302, 1304 for the appointment participants 226, 228, 230, respectively) based at least in part on the visual metadata 1306 associated with the appointment participants 226, 228, 230 within the acoustic environment 600 and the acoustic metadata 1200 associated with the audio appointment information 106. In this example, the automated clinical documentation process 10 can combine the location estimates from the visual metadata 1306 with the location information of the acoustic metadata 1200 to cluster the spatial-spectral features of the visual metadata 1306 and the acoustic metadata 1200 into the speaker representations 1300, 1302, 1304 for the appointment participants 226, 228, 230, respectively.

[0131] In some implementations, the automated clinical documentation process 10 can receive 1112 weighting metadata associated with the audio encounter information received by the second microphone system. As will be discussed in greater detail below, the device selection and weighting module 410 of the ACD computer system 12 can be configured to weight multiple audio streams (e.g., from different microphone systems) based at least in part on a signal-to-noise ratio of each audio stream, thereby defining a weight for each audio stream. In some implementations, a weight can be defined for each audio stream such that an estimated speech processing system performance is maximized. For example, the automated clinical documentation process 10 can train the device selection and weighting module based at least in part on a signal-to-noise ratio (SNR) of each portion or frame of each audio stream. However, it will be appreciated that other metrics or attributes associated with each audio stream can be utilized to select and weight each audio stream. For example, the automated clinical documentation process 10 can train the device selection and weighting module based at least in part on a reverberation level (e.g., a C50 ratio) of each portion or frame of each audio stream. While two examples of particular metrics that can be utilized to select and weight each audio stream have been provided, it will be appreciated that any metric or attribute can be utilized to select and / or weight the various audio streams within the scope of the present disclosure. In some implementations, the device selection and weighting module 410 can provide a previously processed or weighted portion of the audio encounter information to define a weighting for each audio stream from the multiple microphone systems. In this manner, the automated clinical documentation process 10 can utilize audio encounter information from multiple microphone systems when tracking and / or identifying a speaker.

[0132] In some implementations, defining 1104 the one or more speaker representations can include defining 1114 the one or more speaker representations based at least in part on weighting metadata associated with the audio encounter information received by the second microphone system. For example, the automated clinical documentation process 10 can utilize weighting metadata (e.g., weighting metadata 1336) associated with the audio encounter information received by the second microphone system (e.g., the second microphone system 416) to assist in identifying which speaker is speaking within an acoustic environment (e.g., the acoustic environment 600) at a given time.

[0133] For example, assume that the encounter participant 226 has a second microphone system (e.g., the mobile electronic device 416) nearby (e.g., in a pocket). In some implementations, the second microphone system (e.g., the mobile electronic device 416) can receive audio encounter information 106 from the encounter participants 226, 228, 230. As will be discussed in greater detail below, the automated clinical documentation process 10 can apply a weight to portions of the audio encounter information received by the first microphone system (e.g., the microphone array 200) and apply a weight to portions of the audio encounter information received by the second microphone system (e.g., the mobile electronic device 416). In some implementations, the automated clinical documentation process 10 can generate weighted metadata (e.g., the weighted metadata 1336) associated with the audio encounter information received by the second microphone system (e.g., the mobile electronic device 416). For example, the weighted metadata 1336 can indicate that the encounter participant 226 is speaking based on the audio encounter information received by the mobile electronic device 416. Accordingly, the automated clinical documentation process 10 can define 1114 a speaker representation for the encounter participant 226 based at least in part on the weighted metadata 1336 associated with the audio encounter information received by the second microphone system.

[0134] In some implementations, defining 1104 one or more speaker representations can include defining 1116 one or more of: at least one known speaker representation and at least one unknown speaker representation. For example and as described above, the ACD computer system 12 can be configured to access one or more data sources 118 (e.g., a plurality of separate data sources 120, 122, 124, 126, 128), examples of which can include, but are not limited to, one or more of: a user profile data source, a voiceprint data source, a voice characteristic data source (e.g., for adapting automated speech recognition models), a faceprint data source, an anthropomorphic shape data source, a speech identifier data source, a wearable token identifier data source, an interaction identifier data source, a medical condition symptom data source, a prescription compatibility data source, a medical insurance coverage data source, and a home health data source.

[0135] In some implementations, the process 10 can compare data included within a user profile (defined within a user profile data source) to at least a portion of the audio encounter information and / or the machine vision encounter information. For example, the data included within the user profile can include voice-related data (e.g., a voiceprint defined locally within the user profile or remotely within a voiceprint data source), language usage patterns, a user accent identifier, user-defined macros, and user-defined shortcuts. In particular, and when attempting to associate at least a portion of the audio encounter information to at least one known encounter participant, the automated clinical documentation process 10 can compare one or more voiceprints (defined in a voiceprint data source) to one or more voices defined in the audio encounter information.

[0136] As noted above, and for this example, assume that: the encounter participant 226 is a medical professional with a voiceprint / profile; the encounter participant 228 is a patient with a voiceprint / profile; and the encounter participant 230 is a third party (an acquaintance of the encounter participant 228) and thus has no voiceprint / profile. Thus, and for this example: assume that the automated clinical documentation process 10 will be successful and identify the encounter participant 226 when comparing the audio encounter information 106A to the various voiceprints / profiles included within the voiceprint data source; assume that the automated clinical documentation process 10 will be successful and identify the encounter participant 228 when comparing the audio encounter information 106B to the various voiceprints / profiles included within the voiceprint data source; and assume that the automated clinical documentation process 10 will be unsuccessful and not identify the encounter participant 230 when comparing the audio encounter information 106C to the various voiceprints / profiles included within the voiceprint data source.

[0137] Thus, and when processing encounter information (e.g., the machine vision encounter information 102 and / or the audio encounter information 106), the automated clinical documentation process 10 can associate the audio encounter information 106A with the voiceprint / profile of the physician, Susan Jones, and can identify the encounter participant 226 as "physician Susan Jones." The automated clinical documentation process 10 can also associate the audio encounter information 106B with the voiceprint / profile of the patient, Paul Smith, and can identify the encounter participant 228 as "patient Paul Smith." Further, the automated clinical documentation process 10 can not be able to associate the audio encounter information 106C with any voiceprint / profile and can identify the encounter participant 230 as an "unknown participant."

[0138] As noted above, the automated clinical documentation process 10 can define 1116 a known speaker representation for "physician Susan Jones" (e.g., the speaker representation 1300 for the encounter participant 226) and a known speaker representation for "patient Paul Smith" (e.g., the speaker representation 1302 for the encounter participant 228) based at least in part on the acoustic metadata 1200 associated with the audio encounter information 106 and the visual metadata 1306 associated with the machine vision encounter information 102. Similarly, the automated clinical documentation process 10 can define an unknown speaker representation for the participant 230 (e.g., the speaker representation 1304 for the encounter participant 230).

[0139] In some implementations, the automated clinical documentation process 10 can tag 1106 one or more portions of the audio encounter information with one or more speaker representations and speaker locations within the acoustic environment. For example, tagging 1106 one or more portions of the audio encounter information with one or more speaker representations and speaker locations within the acoustic environment can generally include associating speaker representation and speaker location information within the acoustic environment with a particular portion (e.g., segment or frame) of the audio encounter information. For example, the automated clinical documentation process 10 can generate a label for each portion of the audio encounter information that has a speaker representation and a speaker location within the acoustic environment associated with each portion.

[0140] Referring also to Figures 15-16 and in some implementations, the automated clinical documentation process 10 can tag one or more portions of the audio encounter information (e.g., portions 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224, 1226, 1228 of the audio encounter information 106) with one or more speaker representations and speaker locations within the acoustic environment. For example, for each portion of the audio encounter information 106, the automated clinical documentation process 10 can generate a“label” or speaker metadata 1400 (represented as speaker metadata 1402, 1404, 1406, 1408, 1410, 1412, 1414, 1416, 1418, 1420, 1422, 1424, 1426, 1428 defined for portions 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1214, 1216, 1218, 1220, 1222, 1214, 1216, 1218, 1220, 1222, 1224, 1226, 1228, respectively) utilizing the speaker representation and speaker location information associated with that portion of the audio encounter information 106.

[0141] Continuing the example above, assume that portions 1202, 1204, 1206, 1208, 1210 include speech from encounter participant 226; portions 1212, 1214, 1216, 1218, 1220 include speech from encounter participant 228; and portions 1222, 1224, 1226, 1228 include speech from encounter participant 230. In this example, the automated clinical documentation process 10 can label portions 1202, 1204, 1206, 1208, 1210 with speaker representation 1300 and a speaker location within the acoustic environment associated with speaker 226 during the associated portions of the audio encounter information; can label portions 1212, 1214, 1216, 1218, 1220 with speaker representation 1302 and a speaker location within the acoustic environment associated with encounter participant 228 during the associated portions of the audio encounter information; and can label portions 1222, 1224, 1226, 1228 with speaker representation 1304 and a speaker location within the acoustic environment associated with encounter participant 230 during the associated portions of the audio encounter information. As will be discussed in greater detail below, the speaker metadata 1400 can be provided to a beam and null selection module (e.g., beam and null selection module 406) of the ACD computing system 12 and can allow for selection and / or combination of particular beams and / or nulls based at least in part on speaker representations and speaker locations defined within the speaker metadata 1400.

[0142] With reference to at least Figure 6 , the automated clinical documentation process 10 can receive 1500 a plurality of predefined beams associated with a microphone array. A plurality of predefined nulls associated with the microphone array can be received 1502. One or more predefined beams from the plurality of predefined beams or one or more predefined nulls from the plurality of predefined nulls can be selected 1504. The microphone array can obtain 1506 audio encounter information via the microphone array using at least one of the one or more selected beams and the one or more selected nulls.

[0143] In some implementations, the automated clinical documentation process 10 can receive 1500 a plurality of predefined beams associated with the microphone array. As described above, the automated clinical documentation process 10 can predefine a plurality of predefined beams associated with the microphone array. For example, the automated clinical documentation process 10 can predefine the plurality of beams based at least in part on information associated with the acoustic environment. As described above, a beam can generally include a constructive interference pattern between the microphones of the microphone array produced by modifying the phase and / or amplitude of the signal at each microphone of the microphone array. The constructive interference pattern can improve the signal processing performance of the microphone array. In some implementations, the automated clinical documentation process 10 can predefine the plurality of beams by adjusting the phase and / or amplitude of the signal at each microphone of the microphone array. In some implementations, the automated clinical documentation process 10 can provide the predefined beams to the beam and null selection module 408. In one example, the automated clinical documentation process 10 can receive 1500 the plurality of predefined beams as a vector of the phase and / or amplitude of each signal for each microphone channel of the microphone array to achieve a desired sensitivity pattern. In another example, and as described above, the automated clinical documentation process 10 can receive 1500 the plurality of predefined beams as a plurality of predefined filters that produce the plurality of predefined beams.

[0144] In some implementations, the plurality of predefined beams can include one or more predefined beams configured to receive audio patient information from one or more target speaker locations within the acoustic environment. For example, and as described above, the automated clinical documentation process 10 can receive information associated with the acoustic environment. In some implementations, the acoustic environment information can indicate one or more target speaker locations from which a speaker can speak within the acoustic environment. Accordingly, the automated clinical documentation process 10 can receive 1500 predefined beams configured to receive audio patient information from the one or more target speaker locations within the acoustic environment.

[0145] For example, and again with reference to Figure 6The automated clinical documentation process 10 can receive 1500 a plurality of predefined beams configured to receive audio encounter information from one or more target speaker locations within the acoustic environment. Assume that the automated clinical documentation process 10 determines, based on information associated with the acoustic environment 600, that the patient is most likely to speak while seated at or near an exam table that is approximately 45° from the base of the microphone array 200; the physician is most likely to speak while seated at or near a desk that is approximately 90° from the base of the microphone array 200; and other patients or other third parties are most likely to speak from approximately 120° from the base of the microphone array 200. In this example, the automated clinical documentation process 10 can receive 1500 a predefined beam 220 configured to receive audio encounter information from the patient (e.g., encounter participant 228); a beam 222 configured to receive audio encounter information from the physician (e.g., encounter participant 226); and a beam 224 configured to receive audio encounter information from another patient / third party (e.g., encounter participant 230).

[0146] In some implementations, the automated clinical documentation process 10 can receive 1502 a plurality of predefined nulls associated with the microphone array. As described above, a null can generally include a pattern of destructive interference between the microphones of a microphone array that is produced by modifying the phase and / or amplitude of the signal at each microphone of the microphone array. The pattern of destructive interference can limit the reception of signals by the microphone array. In some implementations, the automated clinical documentation process 10 can predefine the plurality of nulls by adjusting the phase and / or amplitude of the signal at each microphone of the microphone array via a plurality of filters (e.g., a plurality of FIR filters). In one example, the automated clinical documentation process 10 can receive 1502 the plurality of predefined nulls as a vector of the phase and / or amplitude of each signal for each microphone channel of the microphone array to achieve a desired sensitivity pattern. In another example, and as described above, the automated clinical documentation process 10 can receive 1502 the plurality of predefined nulls as a plurality of predefined filters that produce the plurality of predefined nulls.

[0147] In some implementations, the plurality of predefined nulls can include one or more predefined nulls configured to limit the reception of audio encounter information from one or more target speaker locations within the acoustic environment. For example, and as described above, the automated clinical documentation process 10 can receive information associated with the acoustic environment. In some implementations, the acoustic environment information can indicate one or more target speaker locations from which a speaker can speak within the acoustic environment. Accordingly, the automated clinical documentation process 10 can receive 1502 a predefined beam configured to limit the reception of audio encounter information from the one or more target speaker locations within the acoustic environment.

[0148] For example, and again with reference to Figure 7 The automated clinical documentation process 10 can receive 1502 a plurality of predefined nulls configured to receive audio encounter information from one or more target speaker locations within the acoustic environment. Assume that the automated clinical documentation process 10 determines, from information associated with the acoustic environment 600, that the patient is most likely to speak while seated at or near the exam table that is approximately 45° from the base of the microphone array 200; that the physician is most likely to speak while seated at or near the desk that is approximately 90° from the base of the microphone array 200; and that the other patient or other third party is most likely to speak from approximately 120° from the base of the microphone array 200. In this example and also with reference to Figure 6 The automated clinical documentation process 10 can predefine null 252 to limit receiving audio encounter information from the patient (e.g., encounter participant 228); null 254 to limit receiving audio encounter information from the physician (e.g., encounter participant 226); and null 256 to limit receiving audio encounter information from the other patient / third party (e.g., encounter participant 230).

[0149] In some implementations, the automated clinical documentation process 10 can select at least one of: one or more predefined beams from the plurality of predefined beams, thereby defining one or more selected beams; and one or more predefined nulls from the plurality of predefined nulls, thereby defining one or more selected nulls. Selecting the predefined beams and / or predefined nulls can include selecting a pattern or combination of the predefined beams and / or predefined nulls to achieve a particular microphone array sensitivity. For example, and as will be described in greater detail below, assume that the automated clinical documentation process 10 determines that the physician (e.g., encounter participant 226) is speaking. In this example, the automated clinical documentation process 10 can select 1504 one or more beams and / or nulls to achieve a beamforming pattern that enables the microphone array 200 to receive audio encounter information from the physician (e.g., encounter participant 226). In some implementations, the one or more selected beams can enable receiving audio encounter information from the physician (e.g., encounter participant 226), and the one or more selected nulls can limit receiving audio encounter information from other speakers or noise sources.

[0150] In some implementations, the automated clinical documentation process 10 can receive 1508 speaker metadata associated with one or more portions of the audio encounter information. For example and as described above, for each portion of the audio encounter information 106, the automated clinical documentation process 10 can utilize a speaker representation and speaker location information associated with the portion of the audio encounter information to generate a“label” or speaker metadata. As described above, the speaker representation can include a cluster of data associated with a unique speaker within the acoustic environment, and the speaker location information can include information indicating a location of the speaker in the acoustic environment. In some implementations, the automated clinical documentation process 10 can utilize the speaker representation and speaker location information to select particular beams and / or nulls for beamforming via the microphone array.

[0151] In some implementations, selecting 1504 at least one of the one or more beams and the one or more nulls can include selecting 1510 at least one of the one or more beams and the one or more nulls based at least in part on speaker metadata associated with one or more portions of the audio encounter information. For example, the automated clinical documentation process 10 can determine a location of a speaker in the acoustic environment and a speaker identity from the speaker metadata. In one instance, assume that the speaker metadata 1400 indicates that the doctor (e.g., the encounter participant 226) is speaking (e.g., based at least in part on the speaker metadata, including a speaker identity associated with the doctor (e.g., the encounter participant 226) and speaker location information indicating that the audio encounter information is being received from near the doctor’s desk). In this example, the automated clinical documentation process 10 can select one or more beams and / or one or more nulls based at least in part on the speaker metadata indicating that the doctor (e.g., the encounter participant 226) is speaking.

[0152] In some implementations, selecting 1510 the one or more beams and the one or more nulls based at least in part on the speaker metadata associated with one or more portions of the audio encounter information can include selecting 1512 one or more predefined beams configured to receive the audio encounter information from one or more target speaker locations within the acoustic environment based at least in part on the speaker metadata associated with one or more portions of the audio encounter information. Continuing the example above, where the speaker metadata 1400 indicates that the doctor (e.g., the encounter participant 226) is speaking, the automated clinical documentation process 10 can determine which beam(s) provide microphone sensitivity to the speaker locations included in the speaker metadata 1400. For example, assume that the automated clinical documentation process 10 determines that beam 222 (as shown in FIG. 2) provides microphone sensitivity to the speaker locations included in the speaker metadata 1400. In this example, the automated clinical documentation process 10 can select beam 222 for beamforming via the microphone array. Figure 7The microphone sensitivity is provided at or near the speaker location included in the speaker metadata 1400. In this example, the automated clinical documentation process 10 can select 1512 the beam 222 for receiving audio encounter information from the physician (e.g., the encounter participant 226). While an example of selecting a single beam has been provided, it will be appreciated that any number of beams can be selected within the scope of the present disclosure.

[0153] In some implementations, selecting 1504 at least one of the one or more beams and the one or more nulls based at least in part on the speaker metadata associated with the one or more portions of the audio encounter information can include selecting 1514 one or more predefined nulls configured to limit receiving the audio encounter information from one or more target speaker locations within the acoustic environment based at least in part on the speaker metadata associated with the one or more portions of the audio encounter information. Continuing the example above, where the speaker metadata 1400 indicates that the physician (e.g., the encounter participant 226) is speaking, the automated clinical documentation process 10 can determine which null(s) limit the microphone sensitivity at or near the other speaker locations and / or noise sources included in the speaker metadata 1400. For example, assume that the automated clinical documentation process 10 determines that the null 252 (as Figure 7 illustrated) limits the microphone sensitivity at or near the speaker location associated with the encounter participant 228, and the null 256 (as Figure 16 illustrated) limits the microphone sensitivity at or near the speaker locations associated with the encounter participants 230, 242.

[0154] Further, assume that the automated clinical documentation process 10 determines that the null 248 limits the microphone sensitivity at or near the first noise source (e.g., the fan 244), and the null 250 limits the microphone sensitivity at or near the second noise source (e.g., the doorway 246). In this example, the automated clinical documentation process 10 can select 1514 the nulls 248, 250, 252, 254 to limit receiving the audio encounter information from the other encounter participants and noise sources. While an example of selecting four nulls has been provided, it will be appreciated that any number of nulls can be selected within the scope of the present disclosure.

[0155] In some implementations, the automatic clinical documentation process 10 can select at least one of the one or more beams and the one or more nulls based at least in part on information associated with the acoustic environment. As discussed above, and in some implementations, the automatic clinical documentation process 10 can receive information associated with the acoustic environment. In some implementations, the automatic clinical documentation process 10 can select particular beams and / or nulls to use when obtaining 1506 audio encounter information with the microphone array based at least in part on acoustic properties of the acoustic environment. For example, assume that the acoustic environment information indicates that the acoustic environment 600 includes a particular reverberation level. In this example, the automatic clinical documentation process 10 can select particular beams and / or nulls to account for and / or minimize signal degradation associated with the reverberation level of the acoustic environment 600. Thus, the automatic clinical documentation process 10 can dynamically select beams and / or nulls based at least in part on various acoustic properties of the acoustic environment.

[0156] In some implementations, the automatic clinical documentation process 10 can obtain 1506 audio encounter information via the microphone array using at least one of the one or more selected beams and the one or more selected nulls. In some implementations, the automatic clinical documentation process 10 can utilize a predefined plurality of beams and a plurality of nulls to obtain 1506 audio encounter information from a particular speaker and limit receiving audio encounter information from other speakers. Continuing with the example above and as shown in FIG. 2, the automatic clinical documentation process 10 can obtain 1506 audio encounter information 106A from the physician (e.g., encounter participant 226) through the microphone array 200. The audio encounter information 106A is obtained 1506 using one or more selected beams (e.g., beam 222) and one or more selected nulls (e.g., nulls 248, 250, 252, 254). Figure 4

[0157] In some implementations, and again referring to Figures 17-20B , the automatic clinical documentation process 10 can provide audio encounter information obtained 1506 using the selected beams and / or selected nulls to a device selection and weighting module (e.g., device selection and weighting module 410). As will be discussed in greater detail below, the automatic clinical documentation process 10 can provide audio encounter information received from a microphone array (e.g., a first microphone system) to the device selection and weighting module 410 to determine which audio stream (i.e., audio encounter information stream) to process with a speech processing system (e.g., speech processing system 418). As will be discussed in greater detail below, the device selection and weighting module can select from a plurality of audio streams (e.g., an audio stream from a first microphone system and an audio stream from a second microphone system). In this way, the automatic clinical documentation process 10 can select an audio stream (or from a portion of an audio stream) from a plurality of audio streams to provide to a speech processing system.

[0158] ​At least with reference to Figure 4 , the automated clinical documentation process 10 can receive 1700 audio encounter information from a first microphone system, thereby defining a first audio stream. Audio encounter information can be received 1702 from a second microphone system, thereby defining a second audio stream. Speech activity in one or more portions of the first audio stream can be detected 1704, thereby defining one or more speech portions of the first audio stream. Speech activity in one or more portions of the second audio stream can be detected 1706, thereby defining one or more speech portions of the second audio stream. The first audio stream and the second audio stream can be aligned 1708 based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

[0159] In some implementations, the automated clinical documentation process 10 can receive 1700 audio encounter information from a first microphone system, thereby defining a first audio stream. As noted above, and in some implementations, the first microphone system can be a microphone array. For example, and as shown in Figure 6 , the automated clinical documentation process 10 can receive audio encounter information (e.g., audio encounter information 106) from a first microphone system (e.g., microphone array 200 having audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218). Again with reference to Figure 4 , the audio encounter information 106 can be received 1700 with one or more beams and / or one or more nulls generated by the microphone array 200. As shown in Figure 18 , the solid lines between the beam and null selection module 408 and the device selection and weighting module 410 and between the beam and null selection module 408 and the alignment module 412 can represent the first audio stream received from the microphone array 200. Also with reference to Figure 4 , and in some implementations, the automated clinical documentation process 10 can receive audio encounter information having one or more portions (e.g., portions 1800, 1802, 1804, 1806, 1808, 1810, 1812, 1814, 1816, 1818, 1820, 1822, 1824, 1826) from a first microphone system (e.g., microphone array 200), thereby defining a first audio stream.

[0160] In some implementations, the automated clinical documentation process 10 can receive 1702 audio encounter information from a second microphone system, thereby defining a second audio stream. In some implementations, the second microphone system can be a mobile electronic device. For example, and as shown in Figure 4 , the automated clinical documentation process 10 can receive audio encounter information from a second microphone system (e.g., mobile electronic device 416). As shown in Figure 18The line with the dashed line and dots between the mobile electronic device 516 and the VAD module 414, between the VAD module 414 and the alignment module 412, and between the alignment module 412 and the device selection and weighting module 410 shown can represent the first audio stream received from the microphone array 200. Again referring to Figure 9 And in some implementations, the automated clinical documentation process 10 can receive audio encounter information having one or more portions (e.g., portions 1828, 1830, 1832, 1834, 1836, 1838, 1840, 1842, 1844, 1846, 1848, 1850, 1852, 1854) from a second microphone system (e.g., the mobile electronic device 416), thereby defining a second audio stream.

[0161] In some implementations, the automated clinical documentation process 10 can detect 1704 speech activity in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream. As described above, and in some implementations, the automated clinical documentation process 10 can identify speech activity within one or more portions of audio encounter information based at least in part on a correlation between audio encounter information received from a microphone array. Also refer to Figure 18 In the example of FIG. 18, and in some implementations, the audio encounter information 106 can include multiple portions or frames of audio information. In some implementations, the automated clinical documentation process 10 can determine a correlation between audio encounter information received from the microphone array 200. For example, the automated clinical documentation process 10 can compare one or more portions of the audio encounter information 106 to determine a degree of correlation between audio encounter information present in each portion across multiple microphones of the microphone array. In some implementations and known in the art, the automated clinical documentation process 10 can perform various cross-correlation processes to determine a degree of similarity between one or more portions of the audio encounter information 106 across multiple microphones of the microphone array 200.

[0162] For example, assume that the automated clinical documentation process 10 receives only ambient noise (i.e., no speech and no directional noise source). The automated clinical documentation process 10 can determine that the spectrum observed in each microphone channel will be different (i.e., uncorrelated at each microphone). However, assume that the automated clinical documentation process 10 receives speech or other “directional” signal within the audio encounter information. In this example, the automated clinical documentation process 10 can determine that one or more portions of the audio encounter information (e.g., portions of the audio encounter information having speech components) are highly correlated at each microphone in the microphone array.

[0163] In some implementations, the automated clinical documentation process 10 can identify speech activity within one or more portions of the audio encounter information based at least in part on determining a threshold amount or degree of correlation between portions of the audio encounter information received from the microphone array. For example, and as described above, various thresholds can be defined (e.g., user-defined, default thresholds, automatically defined via the automated clinical documentation process 10, etc.) to determine when portions of the audio encounter information are sufficiently correlated. Thus, in response to determining at least a threshold degree of correlation between portions of the audio encounter information across multiple microphones, the automated clinical documentation process 10 can determine or identify speech activity within one or more portions of the audio encounter information. In conjunction with determining a threshold correlation between portions of the audio encounter information, the automated clinical documentation process 10 can use other methods known in the art for voice activity detection (VAD) such as filtering, noise reduction, applying classification rules, etc. In this way, traditional VAD techniques can be used in conjunction with determining a threshold correlation between one or more portions of the audio encounter information to identify speech activity within one or more portions of the audio encounter information. Again with reference to Figure 18 , the automated clinical documentation process 10 can detect speech activity in portions 1800, 1802, 1804, 1806, 1808, 1810, 1812.

[0164] In some implementations, the automated clinical documentation process 10 can detect 1706 speech activity in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream. For example, and in some implementations, the automated clinical documentation process 10 can perform various known voice activity detection (VAD) processes to determine which portions of the audio encounter information received by the mobile electronic device 416 include speech activity. In this way, the automated clinical documentation process 10 can detect 1706 speech activity in one or more portions of the second audio stream. Again with reference to Figure 19 , the automated clinical documentation process 10 can detect speech activity in portions 1830, 1832, 1834, 1836, 1838, 1840, 1842 using various known VAD processes.

[0165] In some implementations, the automated clinical documentation process 10 can align 1708 the first audio stream and the second audio stream based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream. In some implementations, the automated clinical documentation process 10 can utilize an adaptive filtering method based on echo cancellation, where the delay between the first audio stream and the second audio stream is estimated from the overall delay of the filter, and the sparsity of the filter is used to ensure that the two audio streams can be aligned. Also with reference to Figure 19And in some implementations, the automated clinical documentation process 10 can align 1900 the first audio stream 1900 and the second audio stream 1902 based at least in part on one or more speech portions of the first audio stream (e.g., speech portions 1800, 1802, 1804, 1806, 1808, 1810, 1812 of the first audio stream 1900) and one or more speech portions of the second audio stream (e.g., speech portions 1830, 1832, 1834, 1836, 1838, 1840, 1842 of the second audio stream 1902). In Figures 20A-20B In examples, while each audio stream can have different amplitude values at a given point in time, the first and second audio streams can be aligned in time.

[0166] In some implementations, and in response to aligning the first audio stream and the second audio stream, the automated clinical documentation process 10 can process 1710 the first audio stream and the second audio stream with one or more speech processing systems. Also referring to Figure 6 , the automated clinical documentation process 10 can provide the first audio stream and the second audio stream from the device selection and weighting module (e.g., the device selection and weighting module 410) to the one or more speech processing systems (e.g., the speech processing systems 418, 2000, 2002). Examples of speech processing systems can generally include automatic speech recognition (ASR) systems, voice biometric systems, emotion detection systems, medical symptom detection systems, hearing enhancement systems, etc. As will be discussed in greater detail below, because different portions of the audio streams can be more accurately processed by the one or more speech processing systems based at least in part on a signal quality of each audio stream, the automated clinical documentation process 10 can selectively process 1701 particular portions from each audio stream with the one or more speech processing systems.

[0167] In some implementations, processing 1710 the first audio stream and the second audio stream with the one or more speech processing systems can include weighting 1712 the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio of the first audio stream and a signal-to-noise ratio of the second audio stream, thereby defining a first audio stream weight and a second audio stream weight. In some implementations, the device selection and weighting module 410 can be a machine learning system or model (e.g., a neural network-based model) configured to be trained to weight each audio stream such that an estimated speech processing system performance is maximized. For example, the automated clinical documentation process 10 can train the device selection and weighting model based at least in part on a signal-to-noise ratio (SNR) of each portion or frame of each audio stream. In some implementations, and as will be discussed in greater detail below, the machine learning model of the device selection and weighting module can be jointly trained in an end-to-end manner to weight and select particular portions of the audio streams for processing by the one or more speech processing systems.

[0168] Refer again Figure 20A For example, suppose a doctor (e.g., doctor 226) has a mobile electronic device 416 near them while they are talking to a patient (e.g., patient 228). In this example, a first microphone system (e.g., microphone array 200) and a second microphone system (e.g., mobile electronic device 416) can receive audio consultation information while the doctor (e.g., consultation participant 226) is speaking. As described above, the automated clinical documentation process 10 can detect speech activity within each audio stream and can align audio streams at least in part based on speech activity in each audio stream. The automated clinical documentation process 10 can weight 1712 each portion or frame of each audio stream at least in part based on the SNR ratio of each audio stream. In this example, suppose the doctor (e.g., consultation participant 226) is speaking while near the doctor's desk. Suppose the doctor (e.g., consultation participant 226) quickly moves to the examination table to assist a patient (e.g., patient 228). During this time, the doctor (e.g., consultation participant 226) may still be speaking. Therefore, even though the physician (e.g., patient participant 226) may move out of beam 222 and toward null 248, microphone array 200 may receive audio patient information of lower quality than that received by mobile electronic device 416 (i.e., the audio patient information received by microphone array 200 may have a lower SNR ratio). Therefore, when the physician is moving, the automated clinical documentation process 10 may weight portions of the audio patient information received from microphone array 200 with a lower weight than that received from mobile electronic device 416. Alternatively / instead, when the physician is moving, the automated clinical documentation process 10 may weight portions of the audio patient information received from mobile electronic device 416 with a higher weight than that received from microphone array 200. In this way, the automated clinical documentation process 10 can utilize the weighting of portions of each audio stream to process each audio stream using one or more speech processing systems.

[0169] In some implementations, processing the 1710 first audio stream and the second audio stream using one or more speech processing systems may include: processing the 1714 first audio stream and the second audio stream using a single speech processing system at least in part based on the weights of the first audio stream and the second audio stream. (See again) Figure 20BAnd in some implementations, the automated clinical documentation process 10 can utilize a single speech processing system (e.g., speech processing system 418) to process the first audio stream (e.g., represented as a solid line between the device selection weighting module 410 and the speech processing system 418) and the second audio stream (e.g., represented as a dotted line between the device selection weighting module 410 and the speech processing system 418) based at least in part on the first audio stream weight and the second audio stream weight (e.g., where both audio stream weights are represented as a dashed line between the device selection weighting module 410 and the speech processing system 418). In some implementations, the automated clinical documentation process 10 can select, via the speech processing system 418, particular portions of either audio stream to process based at least in part on the first audio stream weight and the second audio stream weight.

[0170] In some implementations, processing 1710 the first audio stream and the second audio stream with one or more speech processing systems can include processing 1716 the first audio stream with a first speech processing system, thereby defining a first speech processing output; processing 1718 the second audio stream with a second speech processing system, thereby defining a second speech processing output; and combining 1720 the first speech processing output with the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight. Again, with reference to ​ And in some implementations, the automated clinical documentation process 10 can process 1716 the first audio stream (e.g., represented as a solid line between the device selection weighting module 410 and the speech processing system 418) with a first speech processing system (e.g., speech processing system 418) to generate a first speech processing output (e.g., represented as a solid line between the speech processing system 418 and the speech processing system 2002). The automated clinical documentation process 10 can process 1718 the second audio stream (e.g., represented as a dotted line between the device selection weighting module 410 and the speech processing system 2000) with a second speech processing system (e.g., speech processing system 2000) to generate a second speech processing output (e.g., represented as a solid line between the speech processing system 2000 and the speech processing system 2002). In some implementations, the automated clinical documentation process 10 can combine 1720 the first speech processing output and the second speech processing output via a third speech processing system (e.g., speech processing system 2002) based at least in part on the first audio stream weight and the second audio stream weight (e.g., where both audio stream weights are represented as a dashed line between the device selection weighting module 410 and the speech processing system 2002). In some implementations, the automated clinical documentation process 10 can select, via the speech processing system 2002, particular portions of the audio streams to process or output based at least in part on the first audio stream weight and the second audio stream weight. While an example of two audio streams has been provided, it will be appreciated that any number of audio streams can be used within the scope of the present disclosure.

[0171] In some implementations, the automated clinical documentation process 10 can utilize output of one or more speech processing systems to generate a visit transcription (e.g., visit transcription 234), where at least a portion of the visit transcription (e.g., visit transcription 234) can be processed to populate at least a portion of a medical record (e.g., medical record 236) associated with a patient visit (e.g., visit to a physician's office). For example, when generating a visit transcription, the automated clinical documentation process 10 can utilize speaker representations to identify speakers in audio visit information. For example, the automated clinical documentation process 10 can generate a diarized visit transcription (e.g., visit transcription 234) that identifies spoken comments and utterances made by particular speakers based at least in part on speaker representations defined for each visit participant. In the example above, the automated clinical documentation process 10 can utilize spoken comments and utterances made by "Dr. Susan Jones" (e.g., visit participant 226), "Patient Paul Smith" (e.g., visit participant 228), and "Unknown Participant" (e.g., visit participant 230) to generate the diarized visit transcription 234.

[0172] General Information:

[0173] As those skilled in the art will appreciate, the present disclosure can be embodied as a method, system, or computer program product. Accordingly, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that can all generally be referred to herein as a "circuit," "module" or "system." Furthermore, the present disclosure can take the form of a computer program product on a computer-usable storage medium having computer-usable program code embodied in the medium.

[0174] Any suitable computer-usable or computer-readable medium can be utilized. The computer-usable or computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a transmission medium such as those supporting the Internet or intranet, or magnetic storage devices. The computer-usable or computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, by optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in the computer memory. In the context of this document, a computer-usable or computer-readable medium can be any medium that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-usable medium can include a propagated data signal with the computer-usable program code embodied in or on the propagated data signal, such as the above-described carrier wave or other transport mechanism. The computer-usable program code can be transmitted using any appropriate medium, including but not limited to the Internet, wireline, optical fiber cable, RF, etc.

[0175] Computer program code for carrying out operations of the present disclosure can be written in an object oriented programming language such as Java, Smalltalk, C++, etc. However, the computer program code for carrying out operations of the present disclosure can also be written in a conventional procedural programming language, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through a local area network / wide area network / Internet (e.g., network 14).

[0176] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions / acts specified in the one or more flow diagrams and / or block diagrams.

[0178] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions / acts specified in the one or more flow diagrams and / or block diagrams.

[0179] The flow diagrams and block diagrams in the drawings are meant to illustrate possible architectures, functions and operations for systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions (s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in reverse order, or not executed at all, depending on the functions involved and the implementation. It will also be noted that each block and combination of blocks in the block diagrams and / or flow diagrams can be implemented by a dedicated hardware-based system that performs the specified functions or combinations of functions, or a combination of dedicated hardware and computer instructions.

[0180] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0181] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims that follow, as well as any claim that is directly or indirectly recited in the summary or abstract as such, are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Many modifications and variations will be apparent to practitioners skilled in the art. Embodiments were chosen and described in order to best explain the principles of the present disclosure and its practical application, and to thereby enable others skilled in the art to best utilize the present disclosure in various embodiments and with various modifications as are suited to the particular use contemplated.

[0182] A number of implementations have been described. Nevertheless, it will be understood that various modifications and changes can be made to the implementations described without departing from the scope of the disclosure as defined in the claims that follow.

Claims

1. A computer-implemented method executed on a computing device, comprising: receiving audio consultation information from a first microphone system, thereby defining a first audio stream; receiving audio consultation information from a second microphone system, thereby defining a second audio stream; detecting speech activity in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream, wherein the speech activity of the one or more speech portions of the first audio stream is identified based on a threshold amount of correlation determined between the audio consultation information received from the first microphone system; detecting speech activity in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream; and aligning the first audio stream and the second audio stream based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

2. The computer-implemented method of claim 1, wherein the first microphone system comprises an array of microphones.

3. The computer-implemented method of claim 1, wherein the second microphone system comprises a mobile electronic device.

4. The computer-implemented method of claim 1, further comprising: in response to aligning the first audio stream and the second audio stream, processing the first audio stream and the second audio stream with one or more speech processing systems.

5. The computer-implemented method of claim 4, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: weighting the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio for the first audio stream and a signal-to-noise ratio for the second audio stream, thereby defining a first audio stream weight and a second audio stream weight.

6. The computer-implemented method of claim 5, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: processing the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight.

7. The computer-implemented method of claim 5, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: processing the first audio stream with a first speech processing system, thereby defining a first speech processing output; processing the second audio stream with a second speech processing system, thereby defining a second speech processing output; and combining the first speech processing output and the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight.

8. A computer program product, the computer program product residing on a non-transitory computer readable medium having stored thereon a plurality of instructions that, when executed by a processor, cause the processor to perform operations comprising: receiving audio consultation information from a first microphone system, thereby defining a first audio stream; receiving audio consultation information from a second microphone system, thereby defining a second audio stream; detecting speech activity in one or more portions of the first audio stream, thereby defining one or more speech portions of the first audio stream, wherein the speech activity of the one or more speech portions of the first audio stream is identified based on a threshold amount of correlation determined between the audio session information received from the first microphone system; detecting speech activity in one or more portions of the second audio stream, thereby defining one or more speech portions of the second audio stream; and aligning the first audio stream and the second audio stream based at least in part on the one or more speech portions of the first audio stream and the one or more speech portions of the second audio stream.

9. The computer program product of claim 8, wherein the first microphone system comprises an array of microphones.

10. The computer program product of claim 8, wherein the second microphone system comprises a mobile electronic device.

11. The computer program product of claim 8, wherein the operations further comprise: processing the first audio stream and the second audio stream with one or more speech processing systems in response to aligning the first audio stream and the second audio stream.

12. The computer program product of claim 11, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: weighting the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio for the first audio stream and a signal-to-noise ratio for the second audio stream, thereby defining a first audio stream weight and a second audio stream weight.

13. The computer program product of claim 12, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: processing the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight.

14. The computer program product of claim 12, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: processing the first audio stream with a first speech processing system, thereby defining a first speech processing output; processing the second audio stream with a second speech processing system, thereby defining a second speech processing output; and combining the first speech processing output and the second speech processing output based at least in part on the first audio stream weight and the second audio stream weight.

15. A computing system comprising: a memory; and ​ a processor configured to receive audio consultation information from a first microphone system, thereby defining a first audio stream, wherein the processor is further configured to receive audio consultation information from a second microphone system, thereby defining a second audio stream, wherein the processor is further configured to detect voice activity in one or more portions of the first audio stream, thereby defining one or more voice portions of the first audio stream, wherein the voice activity of the one or more voice portions of the first audio stream is identified based on a threshold amount of correlation determined between the audio consultation information received from the first microphone system, wherein the processor is further configured to detect voice activity in one or more portions of the second audio stream, thereby defining one or more voice portions of the second audio stream, and wherein processor is further configured to align the first audio stream and the second audio stream based at least in part on the one or more voice portions of the first audio stream and the one or more voice portions of the second audio stream.

16. The computing system of claim 15, wherein the first microphone system comprises an array of microphones.

17. The computing system of claim 15, wherein the second microphone system comprises a mobile electronic device.

18. The computing system of claim 15, wherein the processor is further configured to: process the first audio stream and the second audio stream with one or more speech processing systems in response to aligning the first audio stream and the second audio stream.

19. The computing system of claim 18, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: weight the first audio stream and the second audio stream based at least in part on a signal-to-noise ratio for the first audio stream and a signal-to-noise ratio for the second audio stream, thereby defining a first audio stream weight and a second audio stream weight.

20. The computing system of claim 19, wherein processing the first audio stream and the second audio stream with one or more speech processing systems comprises: process the first audio stream and the second audio stream with a single speech processing system based at least in part on the first audio stream weight and the second audio stream weight.

Citation Information

Patent Citations

  • Far-field audio processing

    CN109564762A

  • Methods and Systems for Dictation and Transcription

    US20130204618A1