Systems and methods for data augmentation for multi-microphone signal processing

By receiving and analyzing the harmonic distortion parameters of multiple microphone signals, an enhanced signal is generated, which solves the harmonic distortion matching problem in the microphone array, improves the accuracy and robustness of speech processing, and avoids the expensive calibration process.

CN115605953BActive Publication Date: 2026-03-17MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing multi-microphone systems, harmonic distortion mismatch of MEMS microphones leads to reduced speech processing performance, which is difficult to compensate for effectively using traditional methods or relies on expensive calibration processes.

Method used

By receiving signals from multiple microphones, the harmonic distortion associated with each microphone is determined, an enhanced signal is generated based on the harmonic distortion parameters, and multiple signals are enhanced using a harmonic distortion coefficient table to form a microphone array.

Benefits of technology

It improves the robustness of the speech processing system to microphone system defects, avoids expensive calibration processes, and enhances the accuracy and robustness of speech processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605953B_ABST
    Figure CN115605953B_ABST
Patent Text Reader

Abstract

A method, computer program product, and computing system are provided for receiving signals from each of a plurality of microphones to define a plurality of signals. Harmonic distortion associated with at least one microphone can be determined. One or more harmonic distortion-based enhancements can be performed on the plurality of signals, at least in part, based on the harmonic distortion associated with the at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the rights of the following U.S. Provisional Application No. 63 / 022,269, filed May 8, 2020, the entire contents of which are incorporated herein by reference. Background Technology

[0003] Automated Clinical Documentation (ACD) can be used, for example, to convert transcribed conversations (e.g., between doctors, patients, and / or other participants such as family members, nurses, physician assistants, etc.) into formatted (e.g., medical) reports. Such reports can be checked, for example, to ensure the accuracy of reports from doctors, scribes, etc.

[0004] To improve the accuracy of speech processing in ACD, data augmentation can allow the generation of new training data for machine learning systems by enhancing existing data to represent new conditions. For example, data augmentation has been used to improve robustness to noise and reverberation, as well as other unpredictable features of speech in real-world deployments (e.g., the problems and unpredictable characteristics when capturing speech signals in a real-world environment compared to a controlled environment).

[0005] Various physical characteristics of audio recording systems can lead to a reduction in speech processing performance. For example, microelectromechanical systems (MEMS) microphones typically include mechanical devices that sense acoustic air pressure and form the main sensor for acoustic signal acquisition in most popular consumer devices, such as mobile phones, video conferencing systems, and multi-microphone array systems.

[0006] MEMS microphones may have various defects. For example, known defects in these MEMS microphones typically include microphone sensitivity defects, microphone self-noise, microphone frequency response, and harmonic distortion.

[0007] When designing multi-microphone systems or arrays, it is often assumed that all microphones in the system or array are perfectly matched. However, this is often not accurate in real-world systems. Therefore, while traditional methods attempt to estimate and compensate for these deficiencies (e.g., often considering only microphone sensitivity), or to establish and compensate for these deficiencies by relying on expensive calibration processes (which is not feasible on a large scale), the underlying enhancement algorithms rely on perfectly matched microphones. Summary of the Invention

[0008] In one implementation, the computer-implemented method, executed by a computer, may include, but is not limited to: receiving signals from each of a plurality of microphones, thereby defining a plurality of signals; determining harmonic distortion associated with at least one microphone; and performing one or more harmonic distortion-based enhancements on the plurality of signals, at least in part, based on the harmonic distortion associated with at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.

[0009] This may include one or more of the following features. Determining the total harmonic distortion associated with at least one microphone may include: receiving harmonic distortion parameters associated with at least one microphone. The harmonic distortion parameters may indicate the order of the harmonics associated with at least one microphone. Performing one or more harmonic distortion-based enhancements on multiple signals, at least in part based on the harmonic distortion parameters, may include: generating a harmonic distortion-based enhanced signal, at least in part based on the harmonic distortion parameters and a harmonic distortion coefficient table. The harmonic distortion coefficient table may be generated at least in part based on the total harmonic distortion measured from at least one microphone. Multiple microphones may define a microphone array.

[0010] In another implementation, the computer program product resides on a computer-readable medium and has a plurality of instructions stored thereon. When executed by a processor, the instructions cause the processor to perform operations including, but not limited to: receiving signals from each of a plurality of microphones, thereby defining a plurality of signals; determining harmonic distortion associated with at least one microphone; and performing one or more harmonic distortion-based enhancements on the plurality of signals, at least in part, based on the harmonic distortion associated with at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.

[0011] This may include one or more of the following features. Determining the total harmonic distortion (THD) associated with at least one microphone may include: receiving harmonic distortion parameters associated with at least one microphone. The harmonic distortion parameters may indicate the order of the harmonics associated with at least one microphone. Performing the one or more harmonic distortion-based enhancements on the plurality of signals, at least in part based on the harmonic distortion parameters, may include: generating a harmonic distortion-based enhanced signal, at least in part based on the harmonic distortion parameters and a harmonic distortion coefficient table. Determining the THD associated with at least one microphone may include: measuring the THD from at least one microphone. The harmonic distortion coefficient table may be generated, at least in part, based on the THD measured from at least one microphone. The plurality of microphones may define a microphone array.

[0012] In another implementation, the computing system includes a processor, and the memory is configured to perform operations including, but not limited to, receiving signals from each of a plurality of microphones, thereby defining a plurality of signals. The processor may also be configured to determine the total harmonic distortion associated with at least one microphone. The processor may further be configured to perform one or more harmonic distortion-based enhancements on the plurality of signals, at least in part, based on the total harmonic distortion associated with said at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.

[0013] This may include one or more of the following features. Determining the total harmonic distortion (THD) associated with at least one microphone may include: receiving harmonic distortion parameters associated with at least one microphone. The harmonic distortion parameters may indicate the order of the harmonics associated with at least one microphone. Performing the one or more harmonic distortion-based enhancements on multiple signals, at least in part based on the harmonic distortion parameters, may include: generating a harmonic distortion-based enhanced signal, at least in part based on the harmonic distortion parameters and a harmonic distortion coefficient table. Determining the THD associated with at least one microphone may include: measuring the THD from at least one microphone. The harmonic distortion coefficient table may be generated at least in part based on the THD measured from at least one microphone. Multiple microphones may define a microphone array.

[0014] Details of one or more implementations are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of an automated clinical documentation computer system and data augmentation process coupled to a distributed computing network;

[0016] Figure 2 It is merged Figure 1 A schematic diagram of a modular ACD system for an automated clinical documentation computer system;

[0017] Figure 3 Is included Figure 2 A schematic diagram of a hybrid media ACD device within a modular ACD system;

[0018] Figure 4 yes Figure 1 A flowchart illustrating the implementation of a data augmentation process;

[0019] Figures 5 to 6 It is based on Figure 1 A schematic diagram of a modular ACD system for various implementations of the data augmentation process;

[0020] Figure 7 yes Figure 1A flowchart illustrating the implementation of a data augmentation process;

[0021] Figure 8 It is based on Figure 1 A schematic diagram of a modular ACD system implementing the data augmentation process;

[0022] Figure 9 yes Figure 1 A flowchart illustrating the implementation of a data augmentation process;

[0023] Figure 10 It is based on Figure 1 A schematic diagram of a modular ACD system implementing the data augmentation process;

[0024] Figure 11 It is based on Figure 1 A schematic diagram of the microphone frequency response for an implementation of the data augmentation process;

[0025] Figure 12 yes Figure 1 A flowchart of an implementation of the data augmentation process; and

[0026] Figure 13 It is based on Figure 1 A schematic diagram of a modular ACD system implementing the data augmentation process.

[0027] The same reference numerals in the various figures denote the same elements. Detailed Implementation

[0028] System Overview:

[0029] refer to Figure 1 The diagram illustrates data augmentation process 10. As will be discussed in more detail below, data augmentation process 10 can be configured to automate the collection and processing of clinical visit information to generate / store / distribute medical records.

[0030] The data augmentation process 10 can be implemented as a server-side process, a client-side process, or a hybrid server-side / client-side process. For example, the data augmentation process 10 can be implemented as a purely server-side process via data augmentation process 10s. Alternatively, the data augmentation process 10 can be implemented as a purely client-side process via one or more of data augmentation processes 10c1, 10c2, 10c3, and 10c4. Alternatively, the data augmentation process 10 can be implemented as a hybrid server-side / client-side process via a combination of data augmentation process 10s and one or more of data augmentation processes 10c1, 10c2, 10c3, and 10c4.

[0031] Therefore, the data augmentation process 10 used in this disclosure may include any combination of data augmentation process 10s, data augmentation process 10c1, data augmentation process 10c2, data augmentation process 10c3 and data augmentation process 10c4.

[0032] The data augmentation process 10S can be a server application and can reside on and be executed by an Automated Clinical Documentation (ACD) computer system 12, which can be connected to a network 14 (e.g., the Internet or a local area network). The ACD computer system 12 can include various components, examples of which may include, but are not limited to: personal computers, server computers, a series of server computers, minicomputers, mainframes, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, cloud-based computing systems, and cloud-based storage platforms.

[0033] As is known in the art, a SAN can include one or more of a personal computer, a server computer, a series of server computers, a minicomputer, a mainframe computer, a RAID device, and a NAS system. Various components of the ACD computer system 12 can execute one or more operating systems, examples of which may include, but are not limited to, Microsoft Windows Server. tm Red Hat Linux tm Unix or a custom operating system.

[0034] The instruction set and subroutines of the data enhancement process 10s, which can be stored on a storage device 16 coupled to the ACD computer system 12, can be executed by one or more processors (not shown) and one or more memory architectures (not shown) included within the ACD computer system 12. Examples of storage device 16 may include, but are not limited to: hard disk drives; RAID devices; random access memory (RAM); read-only memory (ROM); and all forms of flash memory storage devices.

[0035] Network 14 may be connected to one or more auxiliary networks (e.g., network 18), examples of which may include, but are not limited to, local area networks (LANs), wide area networks (WANs), or intranets.

[0036] Various I / O requests (e.g., I / O request 56) may be sent from data augmentation process 10s, data augmentation process 10c1, data augmentation process 10c2, data augmentation process 10c3, and / or data augmentation process 10c4 to ACD computer system 12. Examples of I / O request 56 may include, but are not limited to, data write requests (i.e., requests to write content to ACD computer system 12) and data read requests (i.e., requests to read content from ACD computer system 12).

[0037] The instruction sets and subroutines of data enhancement processes 10c1, 10c2, 10c3, and / or 10c4, which can be stored (respectively) on storage devices 20, 22, 24, and 26 coupled to ACD client electronics 28, 30, 32, and 34, can be executed by one or more processors (not shown) and one or more memory architectures (not shown) incorporated into ACD client electronics 28, 30, 32, and 34. Storage devices 20, 22, 24, and 26 can include, but are not limited to: hard disk drives; optical disk drives; RAID devices; random access memory (RAM); read-only memory (ROM); and all forms of flash memory. Examples of ACD client electronic devices 28, 30, 32, and 34 may include, but are not limited to, personal computing devices 28 (e.g., smartphones, personal digital assistants, laptops, notebook computers, and desktop computers), audio input devices 30 (e.g., handheld microphones, lapel microphones, embedded microphones (such as microphones embedded in glasses, smartphones, tablets, and / or watches), and audio recording devices), display devices 32 (e.g., tablets, computer monitors, and smart TVs), machine vision input devices 34 (e.g., RGB imaging systems, infrared imaging systems, ultraviolet imaging systems, laser imaging systems, sonar imaging systems, radar imaging systems, and thermal imaging systems), hybrid devices (e.g., a single device that includes the functionality of one or more of the aforementioned reference devices; not shown), audio presentation devices (e.g., speaker systems, headphone systems, or earphone systems; not shown), various medical devices (e.g., medical imaging devices, heart monitors, scales, thermometers, and blood pressure monitors; not shown), and dedicated network devices (not shown).

[0038] Users 36, 38, 40, and 42 can directly access ACD computer system 12 via network 14 or via auxiliary network 18. Furthermore, ACD computer system 12 can be connected to network 14 via auxiliary network 18, as shown by link line 44.

[0039] Various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) can be directly or indirectly coupled to network 14 (or network 18). For example, personal computing device 28 is shown as being directly coupled to network 14 via a hardwired network connection. Furthermore, machine vision input device 34 is shown as being directly coupled to network 18 via a hardwired network connection. Audio input device 30 is shown as being wirelessly coupled to network 14 via a wireless communication channel 46 established between audio input device 30 and wireless access point (i.e., WAP) 48, and WAP 48 is shown as being directly coupled to network 14. WAP 48 can be, for example, an IEEE 802.11a, 802.11b, 802.11g, 802.11n, Wi-Fi, and / or Bluetooth device capable of establishing a wireless communication channel 46 between audio input device 30 and WAP 48. Display device 32 is shown wirelessly coupled to network 14 via a wireless communication channel 50 established between display device 32 and WAP 52, which is shown directly coupled to network 14.

[0040] Various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) can each run an operating system, examples of which may include, but are not limited to, Microsoft Windows. tm Apple Macintosh tm Red Hat Linux tm Alternatively, a custom operating system can be used, in which the combination of various ACD client electronic devices (e.g., ACD client electronic devices 28, 30, 32, 34) and ACD computer system 12 can form a modular ACD system 54.

[0041] Also refer to Figure 2 A simplified example embodiment of a modular ACD system 54 is illustrated, configured to automate clinical documentation. The modular ACD system 54 may include: a machine vision system 100 configured to acquire machine vision visit information 102 about a patient's visit; an audio recording system 104 configured to acquire audio visit information 106 about a patient's visit; and a computer system (e.g., an ACD computer system 12) configured to receive the machine vision visit information 102 and the audio visit information 106 from the machine vision system 100 and the audio recording system 104, respectively. The modular ACD system 54 may also include: a display rendering system 108 configured to render visual information 110; and an audio rendering system 112 configured to render audio information 114, wherein the ACD computer system 12 may be configured to provide the visual information 110 and the audio information 114 to the display rendering system 108 and the audio rendering system 112, respectively.

[0042] Examples of machine vision system 100 may include, but are not limited to, one or more ACD client electronics (e.g., ACD client electronics 34, examples of which may include, but are not limited to, RGB imaging systems, infrared imaging systems, ultraviolet imaging systems, laser imaging systems, sonar imaging systems, radar imaging systems, and thermal imaging systems). Examples of audio recording system 104 may include, but are not limited to, one or more ACD client electronics (e.g., ACD client electronics 30, examples of which may include, but are not limited to, handheld microphones, lapel microphones, embedded microphones (such as microphones embedded in glasses, smartphones, tablet computers, and / or watches), and audio recording devices). Examples of display presentation system 108 may include, but are not limited to, one or more ACD client electronics (e.g., ACD client electronics 32, examples of which may include, but are not limited to, tablet computers, computer monitors, and smart TVs). Examples of audio presentation system 112 may include, but are not limited to, one or more ACD client electronics (e.g., audio presentation device 116, examples of which may include, but are not limited to, speaker systems, headphone systems, and earphone systems).

[0043] As will be discussed in more detail below, the ACD computer system 12 can be configured to access one or more data sources 118 (e.g., multiple individual data sources 120, 122, 124, 126, 128), examples of which may include, but are not limited to, one or more of the following: user profile data source, voiceprint data source, voice characteristic data source (e.g., for adapting to an automatic speech recognition model), facial recognition data source, human-like data source, speech identifier data source, wearable token identifier data source, interaction identifier data source, medical condition symptom data source, prescription compatibility data source, health insurance coverage data source, and home health care data source. While five different examples of data source 118 are shown in this particular example, this is for illustrative purposes only and is not intended to be a limitation of this disclosure, as other configurations are possible and are considered to be within the scope of this disclosure.

[0044] As will be discussed in more detail below, the modular ACD system 54 can be configured to monitor monitoring spaces (e.g., monitoring space 130) within a clinical environment, examples of which may include, but are not limited to, physician offices, medical facilities, medical practices, medical laboratories, emergency care facilities, medical clinics, emergency rooms, operating rooms, hospitals, long-term care facilities, rehabilitation facilities, nursing rooms, and hospice facilities. Therefore, examples of patient visits may include, but are not limited to, patients visiting one or more of the aforementioned clinical environments (e.g., physician offices, medical facilities, medical practices, medical laboratories, emergency care facilities, medical clinics, emergency rooms, operating rooms, hospitals, long-term care facilities, rehabilitation facilities, nursing rooms, and hospice facilities).

[0045] When the clinical environment described above is larger or requires a higher level of resolution, the machine vision system 100 may include multiple discrete machine vision systems. As described above, examples of the machine vision system 100 may include, but are not limited to, one or more ACD client electronics devices (e.g., ACD client electronics device 34, examples of which may include, but are not limited to, RGB imaging systems, infrared imaging systems, ultraviolet imaging systems, laser imaging systems, sonar imaging systems, radar imaging systems, and thermal imaging systems). Therefore, the machine vision system 100 may include one or more of each of the following: RGB imaging systems, infrared imaging systems, ultraviolet imaging systems, laser imaging systems, sonar imaging systems, radar imaging systems, and thermal imaging systems.

[0046] When the clinical environment described above is larger or requires a higher level of resolution, the audio recording system 104 may include multiple discrete audio recording systems. As described above, examples of the audio recording system 104 may include, but are not limited to, one or more ACD client electronics devices (e.g., ACD client electronics device 30, examples of which may include, but are not limited to, handheld microphones, lapel microphones, in-set microphones (such as microphones embedded in glasses, smartphones, tablets, and / or watches), and audio recording devices). Therefore, the audio recording system 104 may include one or more of each of the following: handheld microphones, lapel microphones, in-set microphones (such as microphones embedded in glasses, smartphones, tablets, and / or watches), and audio recording devices.

[0047] When the clinical environment described above is larger or requires a higher level of resolution, the display presentation system 108 may include multiple discrete display presentation systems. As described above, examples of the display presentation system 108 may include, but are not limited to, one or more ACD client electronic devices (e.g., ACD client electronic device 32, examples of which may include, but are not limited to, tablet computers, computer monitors, and smart TVs). Therefore, the display presentation system 108 may include one or more of each of tablet computers, computer monitors, and smart TVs.

[0048] When the clinical environment described above is larger or requires a higher level of resolution, the audio presentation system 112 may include multiple discrete audio presentation systems. As mentioned above, examples of the audio presentation system 112 may include, but are not limited to, one or more ACD client electronic devices (e.g., audio presentation device 116, examples of which may include, but are not limited to, speaker systems, headphone systems, or earphone systems). Therefore, the audio presentation system 112 may include one or more of each of a speaker system, headphone system, or earphone system.

[0049] ACD Computer System 12 may include multiple discrete computer systems. As described above, ACD Computer System 12 may include various components, examples of which may include, but are not limited to: personal computers, server computers, a series of server computers, minicomputers, mainframes, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, cloud-based computing systems, and cloud-based storage platforms. Therefore, ACD Computer System 12 may include one or more of each of the following: personal computers, server computers, a series of server computers, minicomputers, mainframes, one or more network-attached storage (NAS) systems, one or more storage area network (SAN) systems, one or more platform-as-a-service (PaaS) systems, one or more infrastructure-as-a-service (IaaS) systems, one or more software-as-a-service (SaaS) systems, cloud-based computing systems, and cloud-based storage platforms.

[0050] Also refer to Figure 3The audio recording system 104 may include a directional microphone array 200 with multiple discrete microphone accessories. For example, the audio recording system 104 may include multiple discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that can form the microphone array 200. As will be discussed in more detail below, the modular ACD system 54 may be configured to form one or more audio recording beams (e.g., audio recording beams 220, 222, 224) via the discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) included in the audio recording system 104.

[0051] For example, the modular ACD system 54 can also be configured to steer one or more audio recording beams (e.g., audio recording beams 220, 222, 224) to one or more participants in the patient visit (e.g., participants 226, 228, 230). Examples of participants (e.g., participants 226, 228, 230) may include, but are not limited to: medical professionals (e.g., doctors, nurses, physician assistants, laboratory technicians, physical therapists, scribes (e.g., transcriptists) and / or staff involved in the patient visit), patients (e.g., people visiting the clinical environment for the patient visit), and third parties (e.g., friends, relatives, and / or acquaintances of the patient involved in the patient visit).

[0052] Therefore, the modular ACD system 54 and / or audio recording system 104 can be configured to form an audio recording beam using one or more discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218). For example, the modular ACD system 54 and / or audio recording system 104 can be configured to form an audio recording beam 220 using audio acquisition device 210, thereby enabling the capture of audio (e.g., speech) generated by the patient 226 (because audio acquisition device 210 is pointed (i.e., oriented) towards the patient 226). Furthermore, the modular ACD system 54 and / or audio recording system 104 can be configured to form an audio recording beam 222 using audio acquisition devices 204, 206, thereby enabling the capture of audio (e.g., speech) generated by the patient 228 (because audio acquisition devices 204, 206 are pointed (i.e., oriented) towards the patient 228). Furthermore, the modular ACD system 54 and / or audio recording system 104 can be configured to form an audio recording beam 224 using audio acquisition devices 212, 214, thereby enabling the capture of audio (e.g., speech) generated by the patient 230 (because the audio acquisition devices 212, 214 are pointed (i.e., oriented) towards the patient 230). Additionally, the modular ACD system 54 and / or audio recording system 104 can be configured to utilize null-steering precoding to eliminate inter-speaker interference and / or noise.

[0053] As is well known in the art, zero-control precoding is a spatial signal processing method that enables multi-antenna transmitters to reduce multi-user interference signals in wireless communication to zero. Zero-control precoding can mitigate the effects of background noise and unknown user interference.

[0054] Specifically, zero-control precoding can be a beamforming method for narrowband signals that can compensate for delays in signals received from a specific source at different elements of an antenna array. Generally, to improve the performance of an antenna array, the incoming signals can be summed and averaged, where some signals can be weighted and signal delays can be compensated.

[0055] The machine vision system 100 and the audio recording system 104 can be standalone devices (e.g., Figure 2(As shown). Alternatively / alternatively, the machine vision system 100 and the audio recording system 104 can be combined into a package to form a hybrid media ACD device 232. For example, the hybrid media ACD device 232 can be configured to be installed into a structure (e.g., wall, ceiling, beam, column) within the aforementioned clinical environment (e.g., physician's office, medical facility, medical practice, medical laboratory, emergency care facility, medical clinic, emergency room, operating room, hospital, long-term care facility, rehabilitation facility, nursing room, and hospice facility), thereby allowing for easy installation. Furthermore, the modular ACD system 54 can be configured to include multiple hybrid media ACD devices (e.g., hybrid media ACD device 232) when the aforementioned clinical environment is larger or requires a higher level of resolution.

[0056] The modular ACD system 54 can also be configured to direct one or more audio recording beams (e.g., audio recording beams 220, 222, 224) to one or more participants in the patient's visit (e.g., participants 226, 228, 230) based at least in part on machine vision visit information 102. As described above, the hybrid media ACD device 232 (and the machine vision system 100 / audio recording system 104 included therein) can be configured to monitor one or more participants in the patient's visit (e.g., participants 226, 228, 230).

[0057] Specifically, the machine vision system 100 (as a standalone system or as a component of the hybrid media ACD device 232) can be configured to detect humanoid shapes within the aforementioned clinical environments (e.g., physician offices, medical facilities, medical practices, medical laboratories, emergency care facilities, medical clinics, emergency rooms, operating rooms, hospitals, long-term care facilities, rehabilitation facilities, nursing rooms, and hospice facilities). When the machine vision system 100 detects these humanoid shapes, the modular ACD system 54 and / or the audio recording system 104 can be configured to utilize one or more discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) to form audio recording beams (e.g., audio recording beams 220, 222, 224) directed at each detected humanoid shape (e.g., patient participants 226, 228, 230).

[0058] As described above, the ACD computer system 12 can be configured to receive machine vision consultation information 102 and audio consultation information 106 from the machine vision system 100 and the audio recording system 104, respectively; and can be configured to provide visual information 110 and audio information 114 to the display presentation system 108 and the audio presentation system 112, respectively. Depending on the configuration of the modular ACD system 54 (and / or the hybrid media ACD device 232), the ACD computer system 12 can be included within or outside the hybrid media ACD device 232.

[0059] As described above, the ACD computer system 12 can execute all or part of the data enhancement process 10, wherein the instruction set and subroutines of the data enhancement process 10 (which may be stored on one or more of, for example, storage devices 16, 20, 22, 24, 26) can be executed by the ACD computer system 12 and / or one or more ACD client electronic devices 28, 30, 32, 34.

[0060] Data augmentation process:

[0061] In some implementations consistent with this disclosure, systems and methods may be provided for data augmentation of training data for multi-channel speech processing systems (e.g., neural augmentation (e.g., beamforming), multi-channel, end-to-end automatic speech recognition (MCE2E) systems, etc.), having a series of corrupted profiles that allow the underlying speech processing algorithms to “learn” to become more robust to the deficiencies of the microphone system. For example, and as described above, data augmentation allows for the generation of new training data for machine learning systems by augmenting existing data to represent new conditions. For example, data augmentation has been used to improve robustness to noise and reverberation, as well as other unpredictable characteristics of speech in real-world deployments (e.g., problems and unpredictable characteristics when capturing speech signals in real-world environments compared to controlled environments).

[0062] In some implementations, various physical characteristics of the audio recording system may lead to a reduction in speech processing performance. For example, a microelectromechanical system (MEMS) microphone typically includes mechanical devices that sense acoustic air pressure and form the main sensor for acoustic signal acquisition in most popular consumer devices, such as mobile phones, video conferencing systems, and multi-microphone array systems. In some implementations, the microphone may typically include discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218), amplifiers, and / or analog-to-digital conversion systems.

[0063] In some implementations, MEMS microphones may be affected by a variety of defects. Known defects in these MEMS microphones typically include microphone sensitivity defects, microphone self-noise, microphone frequency response, and harmonic distortion. As discussed in more detail below, microphone sensitivity typically refers to the microphone's response to a given sound pressure level. This can vary from device to device (e.g., from microphone to microphone in a microphone array). Microphone self-noise typically refers to the amount of noise output by the microphone in a completely quiet environment. In some implementations, the spectral shape of this noise may make it more effective at certain frequencies than at others, and different microphones may have different self-noise levels / characteristics. In some implementations, the microphone may have non-flat amplitude and / or non-linear frequency response at different frequencies. In some implementations, the housing of the microphone or microphone array may introduce spectral shaping into the microphone frequency response. Harmonic distortion can be a measure of the amount of distortion in the microphone output for a given pure-tone input signal. While several examples of microphone defects have been provided, it will be understood that, within the scope of this disclosure, other defects may introduce problems when performing speech processing operations using multiple microphones (e.g., as in microphone array 200).

[0064] When designing neural beamforming or MCE2E systems, it is often assumed that all microphones in the system or array are perfectly matched. However, this is often not accurate in real-world systems, at least for the reasons mentioned above. Therefore, while traditional methods attempt to estimate and compensate for these deficiencies (e.g., often considering only microphone sensitivity), or to establish and compensate for these deficiencies by relying on expensive calibration processes (which are not feasible on a large scale), the underlying enhancement algorithms typically rely on perfectly matched microphones.

[0065] As will be discussed in more detail below, implementations of this disclosure can address microphone-to-microphone defects by enhancing training data for beamforming and MCE2E systems with a set of corrupted profiles that allow the underlying speech processing algorithms to 'learn' to become more robust to microphone system defects. In some implementations, the underlying speech processing system can learn to address a set of microphone system or array defects by incorporating system optimization criteria; rather than relying on external calibration data or auxiliary processing, which may not be ideal in conventional systems. Implementations of this disclosure also avoid any additional processing overhead from the underlying speech processing system and do not require expensive and time-consuming microphone system calibration data. Implementations of this disclosure can address the problem of microphone system performance degradation over time by learning microphone system defects during training.

[0066] As described above and at least refer to Figures 4 to 6The data augmentation process 10 can receive 400 signals from each of the multiple microphones, thereby defining multiple signals. One or more inter-microphone gain-based augmentations can be performed on the multiple signals, thereby defining one or more inter-microphone gain-enhanced signals.

[0067] Also refer to Figure 5 In some implementations, the audio recording system 104 may include a directional microphone array 200 with multiple discrete microphone accessories. For example, the audio recording system 104 may include multiple discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that may form the microphone array 200. In some implementations, each audio acquisition device or microphone may include a microphone assembly, an amplifier, and an analog-to-digital conversion system. As mentioned above, each microphone (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) may have defects and / or mismatches in the configuration or operation of each microphone. For example, each microphone in the microphone array 200 may include various physical characteristics that affect the ability of each microphone to process speech signals. In some implementations, the combination of microphone accessories, amplifiers, analog-to-digital conversion systems, and / or microphone housings can change the inter-microphone gain associated with the signals received by the microphone array 200.

[0068] For example, suppose microphone 202 introduces, for example, a two-dB gain relative to other microphones, while microphone 212 introduces, for example, a one-dB gain relative to other microphones. In this example, the inter-microphone gain mismatch may cause the speech processing system (e.g., speech processing system 300) to perform incorrect or inaccurate signal processing. Therefore, data augmentation process 10 may perform 402 augmentation on existing training data and / or signals received from various microphones to generate inter-microphone gain-enhanced signals. These inter-microphone gain-enhanced signals can be used to train speech processing system 300 to account for gain mismatches between microphones in microphone array 200.

[0069] In some implementations, the data augmentation process 10 can receive 400 signals from each of multiple microphones, thus defining multiple signals. (See again...) Figure 5Furthermore, in some implementations, microphone array 200 can process speech from various sources (e.g., audio medical information 106A-106C). Therefore, microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218 can generate signals representing the speech processed by microphone array 200 (e.g., multiple signals 500). In some implementations, data augmentation process 10 can receive 400 signals from some or each of microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218.

[0070] In some implementations, the data augmentation process 10 may perform one or more inter-microphone gain-based augmentations on multiple signals, thereby defining one or more inter-microphone gain-enhanced signals. Inter-microphone gain-based augmented signals typically include enhancements to the gain of a signal or training data representing variability or defects associated with the relative gain levels between microphones in a microphone array. As described above, inter-microphone gain-based augmented signals allow speech processing systems (e.g., speech processing system 300) to account for mismatches or variations between microphone gain levels without the need for expensive and complex signal compensation techniques used in conventional speech processing systems with microphone arrays.

[0071] In some implementations, performing one or more inter-microphone gain-based enhancements on multiple signals may include applying gain levels from multiple gain levels to the signal from each microphone. (See again) Figure 5 Furthermore, in some implementations, the data augmentation process 10 may apply multiple gain levels (e.g., multiple gain levels 502) to multiple signals (e.g., multiple signals 500). In some implementations, the multiple signals 500 received from multiple microphones (e.g., microphone array 200) may be received at any time before performing one or more inter-microphone gain-based augmentations. For example, the multiple signals 500 may include training data generated using microphone array 200. In some implementations, the multiple signals may include signals received during real-time processing of the speech signal. In this way, the multiple signals can be used to perform inter-microphone gain-based augmentations at any time relative to when the multiple signals are received.

[0072] In some implementations, multiple gain levels can be associated with a specific microphone or a specific microphone array. For example, suppose a speaker is speaking in a conference room where a microphone array for a teleconferencing system is deployed. In this example, the properties of the microphones in the microphone array can introduce variations in gain between microphones into the speech signal processed by the microphone array. Now suppose a speaker is speaking to a virtual assistant from a separate computing device. In this example, although the environmental characteristics remain constant (i.e., the conference room), the microphone array of the virtual assistant may have factors and characteristics that could affect signal processing differently than the microphone array of the teleconferencing system. In some implementations, the differences between microphone arrays can have various effects on the performance of the speech processing system. Therefore, the data augmentation process 10 can allow the speech signal received by one microphone array to be used to train a speech processing system with other microphone arrays and / or to adapt a speech processing system or model with new adaptation data.

[0073] In some implementations, the data augmentation process 10 may receive a selection of a target microphone array. The target microphone array may include a type of microphone or a microphone array. In some implementations, the data augmentation process 10 may receive the selection of a target microphone array by providing specific inter-microphone gain-based characteristics associated with the target microphone array. In some implementations, the data augmentation process 10 may utilize a graphical user interface to receive the selection of a target microphone array from a library of target microphone arrays. In one example, the data augmentation process 10 may receive selections of various characteristics of the microphone array (e.g., the type of microphone array, the arrangement of the microphones in the microphone array, etc.) (e.g., via a graphical user interface) to define the target microphone array. As will be discussed in more detail below, and in some implementations, the data augmentation process 10 may receive a range or distribution of characteristics of the target microphone array. While examples of graphical user interfaces have been described, it will be understood that the target microphone array may be selected in various ways within the scope of this disclosure (e.g., manually by the user, automatically by the data augmentation process 10, a predefined target microphone array, etc.).

[0074] In some implementations, the data augmentation process 10 may perform one or more inter-microphone gain-based augmentations on multiple signals, at least in part, based on the target microphone array. As will be discussed in more detail below, it may be desirable to augment multiple signals associated with a particular microphone array for various reasons. For example, and in some implementations, the data augmentation process 10 may perform one or more inter-microphone gain-based augmentations on multiple signals to train a speech processing system with a target microphone array using multiple signals. In this example, the data augmentation process 10 may use speech signals received by different microphone arrays to train the speech processing system, which allows the speech processing system to effectively utilize the augmented training signal set along with various microphone arrays.

[0075] In another example, data augmentation process 10 may perform one or more inter-microphone gain-based augmentations on multiple signals to generate additional training data for a speech processing system that has different levels of gain mismatch or variation between microphone arrays of the same or similar types. In this way, data augmentation process 10 can train the speech processing system to be more robust to variations in gain by augmenting the training data set with various gain levels representing defects in the microphones of the microphone array. While two examples of signals utilizing inter-microphone gain augmentation have been provided, it will be understood that within the scope of this disclosure, data augmentation process 10 may perform inter-microphone gain-based augmentations on multiple signals for a variety of other purposes. For example, and in some implementations, inter-microphone gain-based augmentation may be used to adapt the speech processing system with new adaptation data (e.g., inter-microphone gain-based augmentations).

[0076] In some implementations, one or more machine learning models can be used to simulate multiple gain levels to be applied to multiple signals. For example, one or more machine learning models can be used to simulate multiple gain levels of a microphone array, which is configured to "learn" how the characteristics of the microphone array or individual microphones affect the gain level of the signal received from the microphone array. As is known in the art, machine learning models can generally include algorithms or combinations of algorithms trained to recognize certain types of patterns. For example, machine learning methods can generally be categorized into three types based on the nature of the available signals: supervised learning, unsupervised learning, and reinforcement learning. As is known in the art, supervised learning can include presenting a computing device with example inputs and their expected outputs, given by a "teacher," where the goal is to learn general rules that map the inputs to the outputs. In the case of unsupervised learning, the learning algorithm is not labeled and is allowed to find the structure in the inputs on its own. Unsupervised learning can be the goal itself (discovering hidden patterns in data) or it can be a means to achieve the goal (feature learning). As is known in the art, reinforcement learning can generally include computing devices interacting in a dynamic environment where the computing device must perform a specific objective (e.g., driving a vehicle or playing a game with an opponent). As the program navigates its problem space, it is provided with feedback similar to rewards, and it attempts to maximize these rewards. While three examples of machine learning methods have been provided, it is understood that other machine learning methods are possible within the scope of this disclosure. Thus, the data augmentation process 10 can utilize a machine learning model (e.g., machine learning model 302) to simulate how the characteristics of a microphone array or a single microphone affect the gain level of the signal received from the microphone array.

[0077] In some implementations, multiple gain levels to be applied to multiple signals can be measured from one or more microphone arrays. For example, and as described above, data augmentation process 10 can receive multiple signals from a microphone array. In some implementations, data augmentation process 10 can determine the gain level of each microphone in the microphone array. For example, data augmentation process 10 can define a range of gain levels for the microphone array (e.g., typically for each microphone and / or microphone array). As will be discussed in more detail below, data augmentation process 10 can define a distribution of gain levels for the microphone array (e.g., typically for each microphone and / or microphone array). In some implementations, the distribution of gain levels can be a function of frequency, such that different gain levels are observed as a function of frequency, typically for a particular microphone and / or microphone array.

[0078] In some implementations, applying gain levels from multiple gain levels 404 to the signal from each microphone may include applying gain levels from a predefined range of gain levels 406 to the signal from each microphone. For example, the predefined range of gain levels may include a maximum gain level and a minimum gain level. In one example, the predefined range of gain levels may be a default gain level range. In another example, the predefined range of gain levels may be determined based on a training dataset for a specific microphone array. In yet another example, the predefined range of gain levels may be manually defined (e.g., by a user through a user interface). While several examples of how a gain level range can be defined have been described, it will be understood that a predefined gain level range can be defined in various ways within the scope of this disclosure.

[0079] Continuing the example above, suppose microphone 202 introduces, for example, a two-dB gain relative to the other microphones, and microphone 212 introduces, for example, a one-dB gain relative to the other microphones in microphone array 200, even though microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218 are identical. In this example, data augmentation process 10 can define a predefined gain level range, for example, from zero dB to, for example, two dB. Data augmentation process 10 can apply gain level 406 502, ranging from, for example, zero dB to, for example, two dB, to each signal from each microphone 202, 204, 206, 208, 210, 212, 214, 216, and 218. Therefore, data augmentation process 10 can perform one or more inter-microphone gain-based enhancements on each signal by applying multiple gain levels 406 from the predefined gain level range to the signal from each microphone to generate an inter-microphone gain-enhanced signal 504.

[0080] In some implementations, applying gain levels 404 from multiple gain levels to the signal from each microphone may include applying random gain levels 408 from a predefined range of gain levels to the signal from each microphone. For example, data augmentation process 10 may apply gain levels randomly selected from a predefined range of gain levels to the signal from each microphone to generate one or more inter-microphone gain-enhanced signals (e.g., inter-microphone gain-enhanced signal 504).

[0081] In some implementations, gain variation can be controlled by specifying parameters for maximum and minimum variation across microphones. For example, the data augmentation process 10 may receive selections of gain level variation parameters (e.g., from a user via a user interface) to define the maximum and / or minimum variation of gain levels across multiple microphones. For example, the gain level variation parameters may include a distribution of gain levels. In some implementations, the gain level variation parameters may include random variations in gain levels, Gaussian-distributed gain level variations, Poisson-distributed gain level variations, and / or gain level variations configured to be learned by a machine learning model. Therefore, it should be understood that the gain level variation parameters may include any type of gain level distribution from which gain levels can be applied to one or more signals. In some implementations, the gain level variation parameters may include default gain level variation parameters defined for a specific microphone array or microphone type. In this way, the data augmentation process 10 can limit the variation of gain levels in signals that are augmented across one or more microphones.

[0082] In some implementations, performing one or more inter-microphone gain-based enhancements on one or more signals may include: transforming one or more signals to the frequency domain. See also... Figure 6 Furthermore, in some implementations, the data augmentation process 10 can transform 410 multiple signals received from multiple microphones (e.g., multiple microphones 202, 204, 206, 208, 210, 212, 214, 216, 218) into a frequency domain representation of the signal (e.g., multiple frequency-domain based signals 600). In some implementations, transforming one or more signals 410 to a characteristic domain may include obtaining frequency components from the signal. In some implementations, the data augmentation process 10 can obtain frequency components from the signal by applying a short-time Fourier transform (STFT) to the signal. While STFT has been discussed as a means of obtaining frequency components from a signal, it is understood that other transforms can be used within the scope of this disclosure to derive frequency components from a signal and / or convert a time-domain representation of the signal to a frequency-domain representation of the signal.

[0083] In some implementations, performing one or more inter-microphone gain-based enhancements on one or more signals 402 may include applying multiple gain levels 412 to multiple frequency bands of the one or more signals that have been converted to the frequency domain. For example, the gain level variability of the microphone array may be frequency-dependent. In some implementations, the data enhancement process 10 may define multiple gain levels for multiple frequency bands. For example, the data enhancement process 10 may be directed to a vector (e.g., gain level vector 602) that defines gain levels for various frequency bands. In this example, each entry of the gain level vector 602 may correspond to a specific frequency or frequency band. In some implementations, the data enhancement process 10 may apply the same gain level 412 to each frequency band, or it may apply a different gain level 412 to each frequency band of each microphone signal among multiple signals received from the microphone array.

[0084] In some implementations, performing one or more inter-microphone gain-based enhancements on one or more signals (402) may include one or more of the following: amplifying (414) at least a portion of one or more signals and attenuating (416) at least a portion of one or more signals. For example, suppose gain level vector 602 specifies a gain level greater than 1 for a particular frequency band. In this example, data augmentation process 10 may amplify the frequency band of the signal from each microphone by 414 gain levels. In another example, suppose gain level vector 602 specifies a gain level less than 1 for another frequency band. In this example, data augmentation process 10 may attenuate the frequency band of the signal from each microphone by 416 gain levels. Therefore, data augmentation process 10 may perform one or more inter-microphone gain-based enhancements on one or more signals by amplifying (414) and / or attenuating (416) the signal from each microphone in the microphone array. In this way, data augmentation process 10 may augment training data to interpret or represent inter-microphone gain level mismatches between microphones in the microphone array.

[0085] As described above and at least refer to Figures 7 to 8 The data augmentation process 10 can receive 700 speech signals from each of a plurality of microphones, thereby defining a plurality of signals. It can receive 702 one or more noise signals associated with microphone self-noise. It can perform 704 one or more self-noise-based augmentations on the plurality of signals, at least in part, based on the one or more noise signals associated with microphone self-noise, thereby defining one or more self-noise-based augmented signals.

[0086] Also refer to Figure 8In some implementations, the audio recording system 104 may include a directional microphone array 200 with multiple discrete microphone accessories. For example, the audio recording system 104 may include multiple discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that may form the microphone array 200. In some implementations, each audio acquisition device or microphone may include a microphone accessory, an amplifier, and / or an analog-to-digital conversion system. As mentioned above, each microphone (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) may have defects and / or mismatches in the configuration or operation of each microphone. For example, each microphone in the microphone array 200 may include various physical characteristics that affect the ability of each microphone to process speech signals. In some implementations, the combination of microphone accessories, amplifiers, and / or analog-to-digital conversion systems may introduce microphone "self-noise" associated with the signals received by the microphone array 200. As mentioned above, "self-noise" can refer to the amount of noise output by a microphone when it is positioned in an environment without external noise. The spectral shape of this noise may have a greater impact on certain frequencies or bands than on others, and different microphones may have different self-noise levels or characteristics.

[0087] For example, suppose microphone 204 outputs a first noise signal, while microphone 214 outputs a second noise signal. In this example, the self-noise signal output by each microphone may cause the speech processing system (e.g., speech processing system 300) to perform incorrect or inaccurate signal processing. Therefore, data augmentation process 10 may perform augmentation 704 on existing training data and / or signals received from various microphones 700 to generate augmented signals based on microphone self-noise. These augmented signals based on microphone self-noise can be used to train speech processing system 300 to take into account the self-noise output by specific microphones of microphone array 200.

[0088] In some implementations, the data augmentation process 10 can receive 700 signals from each of multiple microphones, thereby defining multiple signals. (See again...) Figure 8 Furthermore, in some implementations, microphone array 200 can process speech from various sources (e.g., audio medical information 106A-106C). Therefore, microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218 can generate signals representing the speech processed by microphone array 200 (e.g., multiple signals 500). In some implementations, data augmentation process 10 can receive 700 signals from some or each of microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218.

[0089] In some implementations, the data augmentation process 10 may receive 702 one or more noise signals associated with the microphone's self-noise. As described above, and in some implementations, each microphone may output a noise signal without any external noise. The characteristics of the output noise signal or microphone self-noise may be based on the electromechanical properties of the microphone accessories, amplifier, and / or analog-to-digital conversion system. (See again...) Figure 8 The data augmentation process 10 can receive 702 one or more noise signals (e.g., one or more noise signals 800) associated with microphone self-noise from various sources (e.g., one or more machine learning models, measurements of microphones deployed in a noise-free environment, etc.).

[0090] In some implementations, receiving 702 of one or more noise signals associated with microphone self-noise may include simulating 706 a model representing the microphone self-noise. For example, one or more machine learning models may be used to simulate one or more noise signals, these models being configured to "learn" how the characteristics of a microphone array or a single microphone generate noise. As described above and as is known in the art, machine learning models typically include algorithms or combinations of algorithms trained to recognize certain types of patterns. In some implementations, the machine learning model (e.g., machine learning model 302) may be configured to simulate microphone operation to generate one or more noise signals (e.g., one or more noise signals 800) associated with microphone self-noise.

[0091] In some implementations, receiving 702 of one or more noise signals associated with microphone self-noise may include measuring 708 the self-noise from at least one microphone. For example, and as described above, the data augmentation process 10 may receive multiple signals from the microphone array. In some implementations, the data augmentation process 10 may determine the self-noise (e.g., one or more noise signals 800) of each microphone in the microphone array. For example, the data augmentation process 10 may define the distribution of the self-noise signals of the microphone array (e.g., typically for each microphone and / or the microphone array). In some implementations, the distribution of the self-noise signals may be a function of frequency, such that, typically for a particular microphone and / or microphone array, different noise responses are observed as a function of frequency.

[0092] In some implementations, the data augmentation process 10 may perform one or more self-noise-based augmentations on multiple signals, at least in part, based on one or more noise signals associated with microphone self-noise, thereby defining one or more self-noise-based augmented signals. The self-noise-based augmented signals may typically include augmentations of the signals or training data to include noise representing the self-noise generated by the microphone. As described above, the self-noise-based augmented signals allow speech processing systems (e.g., speech processing system 300) to account for the self-noise output by the microphone without requiring the expensive and complex signal compensation techniques used in conventional speech processing systems with microphone arrays.

[0093] In some implementations, one or more noise signals may be associated with a specific microphone or microphone array. For example, suppose a speaker is speaking in a clinical environment where a microphone array of a modular ACD system 54 is deployed. In this example, the properties of the microphones in the microphone array can output various noise signals or noise signal distributions in the speech signal processed by the microphone array. Now suppose a speaker is speaking to a virtual assistant located within a separate computing device in the clinical environment. In this example, although the environmental characteristics remain the same (i.e., the clinical environment), the microphone array of the virtual assistant may have factors and characteristics that may affect signal processing differently than the microphone array of the modular ACD system 54. In some implementations, differences between microphone arrays can have various effects on the performance of the speech processing system. Therefore, the data augmentation process 10 can allow speech signals received by one microphone array to be used to train a speech processing system with other microphone arrays and / or to adapt a speech processing system or model with new adaptation data.

[0094] In some implementations, the data augmentation process 10 may receive a selection of a target microphone array. The target microphone array may include a type of microphone or a microphone array. In some implementations, the data augmentation process 10 may receive a selection of a target microphone array by providing specific self-noise characteristics associated with the target microphone array. In some implementations, the data augmentation process 10 may utilize a graphical user interface to receive a selection of a target microphone array from a library of target microphone arrays. In one example, the data augmentation process 10 may receive selections of various characteristics of the microphone array (e.g., the type of microphone array, the arrangement of the microphones in the microphone array, etc.) (e.g., via a graphical user interface) to define the target microphone array. As will be discussed in more detail below, and in some implementations, the data augmentation process 10 may receive a range or distribution of characteristics of the target microphone array. While examples of graphical user interfaces have been described, it will be understood that the target microphone array may be selected in various ways within the scope of this disclosure (e.g., manually by the user, automatically by the data augmentation process 10, a predefined target microphone array, etc.).

[0095] In some implementations, the data augmentation process 10 may perform one or more self-noise-based augmentations on multiple signals, at least in part, based on the target microphone array. As will be discussed in more detail below, it may be desirable to augment multiple signals associated with a particular microphone array for various reasons. For example, and in some implementations, the data augmentation process 10 may perform one or more self-noise-based augmentations on multiple signals to train a speech processing system with a target microphone array using multiple signals. In this example, the data augmentation process 10 may use speech signals received by different microphone arrays to train the speech processing system, which allows the speech processing system to effectively utilize the augmented training signal set along with various microphone arrays.

[0096] In another example, data augmentation process 10 may perform one or more self-noise-based augmentations on multiple signals to generate additional training data for a speech processing system, which has varying self-noise in microphones or microphone arrays of the same or similar types. In this way, data augmentation process 10 can train the speech processing system to be more robust to microphone self-noise by augmenting the training dataset with self-noise signals representing defects in the microphones of the microphone array. While two examples of utilizing self-noise-based augmented signals have been provided, it will be understood that within the scope of this disclosure, data augmentation process 10 may perform self-noise-based augmentations on multiple signals for a variety of other purposes. For example, and in some implementations, self-noise-based augmentation may be used to adapt the speech processing system with new adaptation data (e.g., self-noise-based augmented signals).

[0097] In some implementations, performing one or more self-noise-based enhancements on multiple signals, at least in part, based on one or more noise signals associated with microphone self-noise, may include adding noise signals from the one or more noise signals to the signal from each microphone. For example, suppose data augmentation process 10 receives 702 a first noise signal associated with the self-noise of microphone 204 and a second noise signal associated with the self-noise of microphone 214. In this example, data augmentation process 10 may add the first noise signal associated with the self-noise of microphone 204 and the second noise signal associated with the self-noise of microphone 214 to the signal from each microphone in a plurality of signals (e.g., a plurality of signals 500). Thus, data augmentation process 10 may generate one or more self-noise-based enhanced signals (e.g., self-noise-based enhanced signal 802) for the signal from each microphone. In this way, data augmentation process 10 may allow the generation of training data using the self-noise of microphone 204 and the self-noise of microphone 214. Although an example of two self-noise signals from two microphones in microphone array 200 has been described, it will be understood that within the scope of this disclosure, any number of self-noise signals from any number of microphones can be added to the signal from each microphone to generate one or more self-noise-based enhancement signals.

[0098] In some implementations, adding noise signal 710 to the signal from each microphone may include adding noise signal 712 from one or more noise signals to the signal from each microphone, at least in part based on a predefined signal-to-noise ratio (SNR) of one or more self-noise-based augmented signals. For example, the data augmentation process 10 may receive a selection of the SNR of one or more self-noise-based augmented signals. In some implementations, the SNR ratio may be received as a selection of the SNR parameter (e.g., from the user via a user interface). In some implementations, the SNR parameter may include a default SNR parameter defined for a specific microphone array or microphone type.

[0099] In some implementations, adding noise signals 710 from one or more noise signals to the signal from each microphone may include adding random noise signals 714 from one or more noise signals to the signal from each microphone. For example, and as described above, data augmentation process 10 may receive one or more noise signals associated with the microphone self-noise of one or more microphones in a microphone array. Continuing the example above, suppose data augmentation process 10 receives 702 a first noise signal associated with the self-noise of microphone 204 and a second noise signal associated with the self-noise of microphone 214. In this example, data augmentation process 10 may add random noise signals (e.g., the first noise signal and / or the second noise signal) 714 to the signal from each microphone (e.g., signals from each microphone 202, 204, 206, 208, 210, 212, 214, 216, 218). In this way, data augmentation process 10 can generate more diverse training data for the speech processing system, which allows the speech processing system to be more robust to the microphone self-noise of the microphone accessories in the microphone array.

[0100] As described above and at least refer to Figures 9 to 11 The data augmentation process 10 can receive a signal 900 from each of a plurality of microphones, thereby defining a plurality of signals. It can receive 902 one or more microphone frequency responses associated with at least one microphone. It can perform one or more microphone frequency response-based augmentations 904 on the plurality of signals, at least in part, based on the one or more microphone frequency responses, thereby defining one or more microphone frequency response-based augmented signals.

[0101] Also refer to Figure 10 Furthermore, in some implementations, the audio recording system 104 may include a directional microphone array 200 with multiple discrete microphone accessories. For example, the audio recording system 104 may include multiple discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that can form the microphone array 200. In some implementations, each audio acquisition device or microphone may include microphone accessories, an amplifier, and an analog-to-digital conversion system. As mentioned above, each microphone (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) may have defects and / or mismatches in the configuration or operation of each microphone.

[0102] For example, each microphone in microphone array 200 may include various physical characteristics that affect each microphone's ability to process speech signals. In some implementations, a combination of microphone accessories, amplifiers, analog-to-digital conversion systems, and / or the housing of each microphone may introduce microphone frequency response. In some implementations, microphone frequency response may refer to a non-flat frequency response in terms of amplitude and a non-linear frequency response in terms of phase, indicating changes in microphone sensitivity at different frequencies. Typical MEMS microphones exhibit a non-flat frequency response shape. For example, the microphone housing may also introduce spectral shaping into the microphone frequency response. See also Figure 11 Furthermore, in some implementations, the microphone frequency response can vary as a function of various types of covers or pads applied to the microphone. Therefore, it should be understood that the microphone frequency response can include variations in signal amplitude and / or phase for different physical characteristics of the microphone. In some implementations, altering the microphone frequency response of one or more microphones in a microphone array can lead to erroneous processing of the speech signal by the speech processing system.

[0103] For example, suppose microphone 206 is characterized by a first microphone frequency response, while microphone 216 is characterized by a second frequency response. In this example, the microphone frequency response generated by each microphone may cause the speech processing system (e.g., speech processing system 300) to perform incorrect or inaccurate signal processing. Therefore, data augmentation process 10 may perform 904 augmentation on existing training data and / or signals 900 received from various microphones to generate augmented signals based on microphone frequency responses. These augmented signals based on microphone frequency responses can be used to train speech processing system 300 to take into account the frequency responses generated by specific microphones in microphone array 200.

[0104] In some implementations, the data augmentation process 10 can receive signals 900 from each of multiple microphones, thereby defining multiple signals. (See again...) Figure 10 Furthermore, in some implementations, microphone array 200 can process speech from various sources (e.g., audio medical information 106A-106C). Therefore, microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218 can generate signals representing the speech processed by microphone array 200 (e.g., multiple signals 500). In some implementations, data augmentation process 10 can receive 900 signals from some or each of microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218.

[0105] In some implementations, the data augmentation process 10 may receive 902 one or more microphone frequency responses associated with at least one microphone. As described above, and in some implementations, each microphone may generate a frequency response based on the microphone's physical characteristics. The shape of the microphone frequency response (e.g., in terms of amplitude and phase) may be based on the electromechanical properties of microphone fittings, amplifiers, analog-to-digital conversion systems, and / or microphone housings. (See again...) Figure 10 The data augmentation process 10 can receive 902 one or more microphone frequency responses (e.g., one or more microphone frequency responses 1000) associated with at least one microphone from various sources (e.g., one or more machine learning models, measurements of the frequency response of at least one microphone, etc.).

[0106] In some implementations, receiving 902 of one or more frequency responses associated with at least one microphone may include: simulating 906 one or more models representing the microphone frequency responses. For example, one or more machine learning models may be used to simulate the frequency responses of one or more microphones, and these machine learning models are configured to "learn" the frequency responses of individual microphones. As described above and as is known in the art, machine learning models typically include algorithms or combinations of algorithms trained to recognize certain types of patterns. In some implementations, the machine learning model (e.g., machine learning model 302) may be configured to simulate microphone operation to generate one or more frequency responses (e.g., one or more microphone frequency responses 1000).

[0107] In some implementations, receiving 902 of one or more frequency responses associated with at least one microphone may include measuring 908 the frequency response from at least one microphone. For example, and as described above, the data augmentation process 10 may receive multiple signals from a microphone array. In some implementations, the data augmentation process 10 may determine the frequency response of each microphone in the microphone array. For example, the data augmentation process 10 may define a distribution of the frequency response of the microphone array (e.g., typically for each microphone and / or the microphone array).

[0108] In some implementations, the data augmentation process 10 may perform one or more microphone frequency response-based augmentations on multiple signals, at least in part, based on the frequency responses of one or more microphones, thereby defining one or more microphone frequency response-based augmented signals. The microphone frequency response-based augmented signals typically include augmentation of the signal or training data to include enhancement of the phase and / or amplitude of the signal or training data as a function of frequency. As described above, the microphone frequency response-based augmented signals allow a speech processing system (e.g., speech processing system 300) to account for phase and / or amplitude variations as a function of microphone frequency without requiring the expensive and complex signal compensation techniques used in conventional speech processing systems with microphone arrays.

[0109] In some implementations, the frequency responses of one or more microphones may be associated with a specific microphone or microphone array. For example, suppose a speaker is speaking in a clinical environment where a microphone array of a modular ACD system 54 is deployed. In this example, the properties of the microphones in the microphone array can generate various frequency responses. Now suppose the speaker is speaking to a virtual assistant located within a separate computing device in the clinical environment. In this example, although the environmental characteristics remain the same (i.e., the clinical environment), the microphone array of the virtual assistant may have factors and characteristics that could affect signal processing differently from the microphone array of the modular ACD system 54. In some implementations, the differences between microphone arrays can have various effects on the performance of the speech processing system. Therefore, the data augmentation process 10 can allow speech signals received by one microphone array to be used to train a speech processing system with other microphone arrays and / or to adapt a speech processing system or model with new adaptation data.

[0110] In some implementations, the data augmentation process 10 may receive a selection of a target microphone or microphone array. The target microphone or microphone array may include a type of microphone or microphone array. In some implementations, the data augmentation process 10 may receive the selection of a target microphone or microphone array by providing a specific frequency response associated with the target microphone or microphone array. In some implementations, the data augmentation process 10 may utilize a graphical user interface to receive the selection of a target microphone array from a library of target microphone arrays. In one example, the data augmentation process 10 may receive selections of various characteristics of the microphone array (e.g., the type of microphone array, the arrangement of the microphones in the microphone array, etc.) (e.g., via a graphical user interface) to define the target microphone array. As will be discussed in more detail below, and in some implementations, the data augmentation process 10 may receive a range or distribution of characteristics of the target microphone array. While examples of graphical user interfaces have been described, it will be understood that the target microphone array may be selected in various ways within the scope of this disclosure (e.g., manually by the user, automatically by the data augmentation process 10, a predefined target microphone array, etc.).

[0111] In some implementations, the data augmentation process 10 may perform one or more microphone frequency response-based augmentations on multiple signals, at least in part, based on a target microphone or microphone array. As will be discussed in more detail below, it may be desirable to augment multiple signals associated with a particular microphone array for various reasons. For example, and in some implementations, the data augmentation process 10 may perform one or more microphone frequency response-based augmentations on multiple signals to train a speech processing system with a target microphone array using multiple signals. In this example, the data augmentation process 10 may use speech signals received by different microphone arrays to train the speech processing system, which allows the speech processing system to effectively utilize the augmented training signal set together with various microphone arrays.

[0112] In another example, data augmentation process 10 may perform one or more microphone frequency response-based augmentations on multiple signals to generate additional training data for a speech processing system, which has varying frequency responses across microphone arrays of the same or similar types. In this way, data augmentation process 10 can train the speech processing system to be more robust to variations in frequency response by augmenting the training dataset with various frequency responses or frequency response distributions. While two examples of utilizing microphone frequency response-based augmented signals have been provided, it will be understood that within the scope of this disclosure, data augmentation process 10 may perform microphone frequency response-based augmentations on multiple signals for a variety of other purposes. For example, and in some implementations, frequency response-based augmentation may be used to adapt a speech processing system with new adaptation data (e.g., microphone frequency response-based augmentation).

[0113] In some implementations, performing one or more microphone frequency response-based enhancements on multiple signals, at least in part based on one or more microphone frequency responses, may include: enhancing one or more amplitude and phase components of the multiple signals, at least in part based on one or more microphone frequency responses. As described above, and in some implementations, each signal may include an amplitude component and a phase component. Continuing the example above, assume that microphone 206 outputs a first microphone frequency response, while microphone 216 outputs a second frequency response. In this example, data enhancement process 10 may utilize the amplitude and / or phase components of the first microphone frequency response associated with microphone 206 and / or the amplitude and / or phase components of the second microphone frequency response associated with microphone 216 to enhance the signal from each microphone (e.g., microphones 202, 204, 206, 208, 210, 212, 214, 216, 218).

[0114] In some implementations, performing one or more microphone frequency response-based enhancements on multiple signals, at least in part, can include filtering the multiple signals using one or more microphone frequency responses. For example, data enhancement process 10 can filter each microphone signal from multiple signals (e.g., multiple signals 500) using one or more microphone frequency responses (e.g., one or more microphone frequency responses 1000). As is known in the art, filtering signals can include convolving the signals in the time domain and multiplying the signals in the frequency domain. For example, signal convolution is a mathematical method of combining two signals to form a third signal, and convolving signals in the time domain is equivalent to multiplying the spectra of the signals in the frequency domain. In some implementations, filtering multiple signals (e.g., multiple signals 500) using one or more microphone frequency responses (e.g., one or more microphone frequency responses 1000) can generate one or more microphone frequency response-based enhanced signals (e.g., one or more microphone frequency response-based enhanced signals 1002).

[0115] Continuing the example above, the data augmentation process 10 can be at least partially based on the frequency responses of one or more microphones. This is achieved by filtering (912) multiple signals 500 using the microphone frequency response associated with microphone 206 to generate one or more microphone frequency response-based augmented signals 1002. In this example, filtering multiple signals 500 using the microphone frequency response associated with microphone 206 can generate amplitude and / or phase enhancement or changes in the multiple signals. In this way, the data augmentation process 10 can generate augmented signals (e.g., one or more microphone frequency response-based augmented signals 1002) that allow the speech processing system to consider the frequency response of a particular microphone when processing speech signals.

[0116] In some implementations, performing one or more microphone frequency response-based enhancements on multiple signals, at least in part, based on one or more microphone frequency responses, may include filtering the multiple signals using a microphone frequency response randomly selected from one or more microphone frequency responses. For example, data enhancement process 10 may perform one or more microphone frequency response-based enhancements on multiple signals 500, at least in part, by filtering the multiple signals 500 using amplitude and / or phase components randomly selected from one or more microphone frequency responses 1000. In this example, data enhancement process 10 may randomly select amplitude and / or phase components from the microphone frequency responses associated with microphone 206 and / or microphone frequency responses associated with microphone 216 to filter the microphone signals from the multiple signals 500. Although examples of two frequency responses have been provided, it will be understood that within the scope of this disclosure, the data augmentation process 10 may perform one or more microphone frequency response-based augmentations on multiple signals 500 by filtering the amplitude and / or phase components randomly selected from any number of microphone frequency responses 914.

[0117] As described above and at least refer to Figures 12 to 13 The data enhancement process 10 can receive 1200 signals from each of a plurality of microphones, thereby defining a plurality of signals. 1202 Harmonic distortion associated with at least one microphone can be determined. 1204 One or more harmonic distortion-based enhancements can be performed on the plurality of signals, at least in part, based on the harmonic distortion associated with at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.

[0118] Also refer to Figure 13 Furthermore, in some implementations, the audio recording system may include a directional microphone array 200 having multiple discrete microphone components. For example, the audio recording system 104 may include multiple discrete audio acquisition devices (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) that can form the microphone array 200. In some implementations, each audio acquisition device or microphone may include microphone accessories, an amplifier, and an analog-to-digital conversion system. As mentioned above, each microphone (e.g., audio acquisition devices 202, 204, 206, 208, 210, 212, 214, 216, 218) may have defects and / or mismatches in the configuration or operation of each microphone.

[0119] For example, each microphone in microphone array 200 may include various physical characteristics that affect each microphone's ability to process speech signals. In some implementations, a combination of microphone accessories, amplifiers, and / or analog-to-digital conversion systems may introduce harmonic distortion. In some implementations, harmonic distortion may refer to a measurement of the amount of distortion at the microphone output for a given pure-tone input signal. In some implementations, changing the total harmonic distortion value associated with one or more microphones in the microphone array may cause the speech processing system to misprocess the speech signal.

[0120] For example, suppose microphone 208 outputs a first total harmonic distortion (THD), while microphone 218 outputs a second THD. In this example, the THD generated by each microphone could cause the speech processing system (e.g., speech processing system 300) to perform incorrect or inaccurate signal processing. Therefore, data augmentation process 10 can perform augmentation 1204 on existing training data and / or signals received from various microphones 1200 to generate harmonic distortion-based augmented signals. These harmonic distortion-based augmented signals can be used to train speech processing system 300 to account for harmonic distortion generated by specific microphones in microphone array 200.

[0121] In some implementations, the data augmentation process 10 can receive 1200 signals from each of multiple microphones, thereby defining multiple signals. (See again...) Figure 13 Furthermore, in some implementations, microphone array 200 can process speech from various sources (e.g., audio medical information 106A-106C). Therefore, microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218 can generate signals representing the speech processed by microphone array 200 (e.g., multiple signals 500). In some implementations, data augmentation process 10 can receive 1200 signals from some or each of microphones 202, 204, 206, 208, 210, 212, 214, 216, and 218.

[0122] In some implementations, the data augmentation process 10 may determine the total harmonic distortion (THD) associated with at least one microphone at 1202. For example, and as described above, THD may refer to a measurement of the amount of distortion on the microphone output for a given pure-tone input signal. The microphone output may include a fundamental signal and multiple harmonics added together. In some implementations, the data augmentation process 10 may receive the THD associated with at least one microphone (e.g., THD 1300). In some implementations, the data augmentation process 10 may determine the THD associated with at least one microphone at 1202 by measuring the THD from at least one microphone at 1206.

[0123] For example, the data augmentation process 10 can be performed by inputting a frequency of " " to each microphone (e.g., a combination of microphone accessories, amplifiers, and / or analog-to-digital conversion systems). A sinusoidal signal. In this example, it can be at the original frequency (i.e., " ")of Additional content is added by multiples of (harmonics). The data augmentation process 10 can determine the total harmonic distortion of each microphone by measuring additional signal content that is not present in the input signal from the output signal. In some implementations, the data augmentation process 10 can determine multiple harmonic orders (e.g., multiples of the input frequency) for each microphone. As will be discussed in more detail below, the harmonic order of the total harmonic distortion can be used as a parameter (e.g., a harmonic distortion parameter) to perform one or more harmonic distortion-based augmentations.

[0124] Refer again Figure 13 Furthermore, in some implementations, it is assumed that the data augmentation process 10 provides a pure sine wave input to at least one microphone (e.g., microphones 208, 218). In this example, the data augmentation process 10 may measure any additional content (i.e., content not included in the input signal) from the outputs of microphones 208, 218. Therefore, the data augmentation process 10 may determine, at least in part, the first total harmonic distortion (THD) of microphone 208 output and the second THD of microphone 218 output based on the additional content from the outputs of microphones 208, 218. While examples of determining two THDs, for example, associated with two microphones, have been described, it is understood that any number of THDs for any number of microphones can be determined within the scope of this disclosure.

[0125] In some implementations, determining the total harmonic distortion (THD) associated with at least one microphone 1202 may include simulating one or more models representing the THD. For example, one or more machine learning models may be used to simulate the THD, and these models may be configured to "learn" the THD of the individual microphones. As described above and as is known in the art, machine learning models typically include algorithms or combinations of algorithms trained to recognize certain types of patterns. In some implementations, the machine learning model (e.g., machine learning model 302) may be configured to simulate microphone operation to generate the THD associated with at least one microphone (e.g., THD 1300).

[0126] In some implementations, determining the total harmonic distortion associated with at least one microphone by 1202 may include receiving harmonic distortion parameters associated with at least one microphone by 1208. (See again...) Figure 13Furthermore, in some implementations, harmonic distortion parameters (e.g., harmonic distortion parameter 1302) may indicate the order of the harmonics associated with at least one microphone. For example, harmonic distortion parameter 1302 may be multiple harmonics associated with or generated at the microphone output (e.g., harmonic distortion parameter "1" may refer to a first-order harmonic; harmonic distortion parameter "2" may refer to both first- and second-order harmonics; and harmonic distortion parameter "n" may refer to an "n"-order harmonic). In some implementations, the data augmentation process 10 may utilize a graphical user interface to receive selections of harmonic distortion parameter 1302. In some implementations, harmonic distortion parameter 1302 may be a default value that can be updated or replaced by a user-defined or selected value.

[0127] In some implementations, the data augmentation process 10 may receive a harmonic distortion parameter associated with at least one microphone in response to determining the total harmonic distortion (THD) associated with at least one microphone in 1202. For example, and as described above, when measuring or simulating the THD, the data augmentation process 10 may determine the number of harmonics output by at least one microphone. Continuing the example above, suppose the data augmentation process 10 determines that the output of microphone 208 in 1202 has a THD with, for example, the fifth harmonic. In this way, the data augmentation process 10 may define the harmonic distortion parameter 1302 as, for example, 5, to represent the order of the harmonics associated with microphone 208.

[0128] In some implementations, the data augmentation process 10 may perform one or more harmonic distortion-based augmentations on multiple signals, at least in part, based on the total harmonic distortion associated with at least one microphone, thereby defining one or more harmonic distortion-based augmented signals. The harmonic distortion-based augmented signals (e.g., harmonic distortion-based augmented signal 1304) typically include augmentation of the signal or training data to include enhancements in the signal representing the sum of harmonic components output by the microphones. As described above, harmonic distortion-based augmented signals allow a speech processing system (e.g., speech processing system 300) to consider the harmonic components generated in the microphone's output signal without requiring the expensive and complex signal compensation techniques used in conventional speech processing systems with microphone arrays.

[0129] In some implementations, one or more total harmonic distortions (THDs) (e.g., THD 1300) may be associated with a specific microphone or microphone array. For example, suppose a speaker is speaking in a clinical environment where a microphone array of a modular ACD system 54 is deployed. In this example, the properties of the microphones in the microphone array can output various THDs. Now suppose the speaker is speaking to a virtual assistant located within a separate computing device in the clinical environment. In this example, although the environmental characteristics remain the same (i.e., the clinical environment), the microphone array of the virtual assistant may have factors and characteristics that could affect signal processing differently from those of the microphone array of the modular ACD system 54. In some implementations, the differences between microphone arrays can have various effects on the performance of the speech processing system. Therefore, the data augmentation process 10 can allow speech signals received by one microphone array to be used to train a speech processing system with other microphone arrays and / or to adapt a speech processing system or model with new adaptation data.

[0130] In some implementations, the data augmentation process 10 may receive a selection of a target microphone or microphone array. The target microphone or microphone array may include a type of microphone or microphone array. In some implementations, the data augmentation process 10 may receive the selection of a target microphone or microphone array by providing a specific total harmonic distortion associated with the target microphone or microphone array. In some implementations, the data augmentation process 10 may utilize a graphical user interface to receive a selection of a target microphone array from a library of target microphone arrays. In one example, the data augmentation process 10 may receive selections of various characteristics of the microphone array (e.g., the type of microphone array, the arrangement of the microphones in the microphone array, etc.) (e.g., via a graphical user interface) to define the target microphone array. In some implementations, the data augmentation process 10 may receive a range or distribution of characteristics of the target microphone array. While examples of graphical user interfaces have been described, it will be understood that the target microphone array can be selected in various ways within the scope of this disclosure (e.g., manually by the user, automatically by the data augmentation process 10, a predefined target microphone array, etc.).

[0131] In some implementations, the data augmentation process 10 may perform one or more harmonic distortion-based augmentations on multiple signals, at least in part, based on the total harmonic distortion associated with at least one microphone. As will be discussed in more detail below, it may be desirable to augment multiple signals associated with a particular microphone array for various reasons. For example, and in some implementations, the data augmentation process 10 may perform one or more harmonic distortion-based augmentations on multiple signals to train a speech processing system with a target microphone array using the multiple signals. In this example, the data augmentation process 10 may use speech signals received by different microphone arrays to train the speech processing system, which allows the speech processing system to effectively utilize the augmented training signal set with various microphone arrays.

[0132] In another example, data augmentation process 10 may perform one or more harmonic distortion-based augmentations 1204 on multiple signals to generate additional training data for a speech processing system, which has varying total harmonic distortion across microphone arrays of the same or similar types. In this way, data augmentation process 10 can train the speech processing system to be more robust to variations in harmonic distortion by augmenting the training dataset with various harmonic distortions or harmonic distortion distributions. While two examples of utilizing harmonic distortion-based augmented signals have been provided, it is understood that within the scope of this disclosure, data augmentation process 10 may perform harmonic distortion-based augmentations on multiple signals for a variety of other purposes. For example, and in some implementations, harmonic distortion-based augmentation may be used to adapt the speech processing system with new adaptation data (e.g., harmonic distortion-based augmented signal 1304).

[0133] In some implementations, performing one or more harmonic distortion-based enhancements on multiple signals, at least in part based on harmonic distortion parameters, may include generating a harmonic distortion-based enhanced signal, at least in part based on the harmonic distortion parameters and a harmonic distortion coefficient table. For example, data enhancement process 10 may utilize harmonic distortion parameters 1302 and a harmonic distortion coefficient table to generate one or more harmonic distortion-based enhanced signals (e.g., harmonic distortion-based enhanced signal 1304). In some implementations, data enhancement process 10 may reference a common total harmonic distortion coefficient table (e.g., harmonic distortion coefficient table 1306) of samples from the device under test (e.g., total harmonic distortion measured from multiple microphones). In some implementations, the data augmentation process 10 can generate a harmonic distortion-based augmented signal 1210 based at least in part on the harmonic distortion parameters and harmonic distortion coefficient table shown in Equation 1 below, where “N” is the order of the highest harmonic distortion (e.g., based on the harmonic distortion parameters), and “p[N]” represents the contribution of the Nth harmonic to the total harmonic distortion:

[0134] (1) Harmonic distortion signal =

[0135] In some implementations, the data augmentation process 10 may generate a harmonic distortion coefficient table 1212 based at least in part on the total harmonic distortion measured from at least one microphone. In one example, assume that the data augmentation process 10 measures 1206 the total harmonic distortion associated with microphone 208. As described above, assume that the data augmentation process 10 provides a pure sinusoidal input to microphone 208. In this example, assume that the data augmentation process 10 measures 1206 for any additional content from the output of microphone 208 (i.e., content not included in the input signal) and determines that the microphone 208 output has, for example, a total harmonic distortion of five harmonic components (e.g., the 1st to 5th harmonics in the output signal of microphone 208). The data augmentation process 10 may generate a harmonic distortion coefficient table 1212 with one or more harmonic distortion coefficients based at least in part on the output signal of microphone 208. While the example above describes generating a harmonic distortion coefficient table with harmonic distortion coefficients from the total harmonic distortion of a single microphone, it is understood that within the scope of this disclosure, the data augmentation process 10 can generate a harmonic distortion coefficient table 1212 with an arbitrary number of harmonic distortion coefficients 1306 for any number of microphones.

[0136] Continuing the example above, suppose data augmentation process 10 determines that the output of microphone 208 1202 has a first total harmonic distortion (THD) of, for example, the 5th harmonic. In this example, data augmentation process 10 can define the harmonic distortion parameter 1302 of microphone 208 as “5”, indicating 5 harmonic components or orders. In some implementations, data augmentation process 10 can look up or identify one or more harmonic distortion coefficients to utilize the harmonic distortion parameter “5” of microphone 208 applied to Equation 1. Therefore, data augmentation process 10 can generate one or more harmonic distortion-based augmented signals 1304 from multiple signals (e.g., multiple signals 500) based at least in part on the harmonic distortion parameter 1302 and the harmonic distortion coefficient table 1306. In this way, data augmentation process 10 can augment multiple signals 500 to be more robust to the harmonic distortion of microphone 208. While the examples above include generating one or more harmonic distortion-based augmented signals based at least in part on harmonic distortion parameters associated with a microphone, it will be understood that within the scope of this disclosure, any number of harmonic distortion-based augmented signals can be generated for any number of microphones and / or for any number of harmonic distortion parameters and / or total harmonic distortion.

[0137] General information:

[0138] As those skilled in the art will understand, this disclosure can be implemented as a method, system, or computer program product. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which are generally referred to herein collectively as “circuit,” “module,” or “system.” Furthermore, this disclosure can take the form of a computer program product on a computer-usable storage medium in which computer-usable program code is implemented.

[0139] Any suitable computer-usable or computer-readable medium may be used. A computer-usable or computer-readable medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, devices, or transmission media. More specific examples of computer-readable media (a non-exhaustive list) may include the following: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, transmission media such as those supporting the Internet or intranets, or magnetic storage devices. A computer-usable or computer-readable medium can also be paper or other suitable media on which a program is printed, because the program can be electronically captured, for example, by optical scanning of the paper or other medium, and then compiled, interpreted, or otherwise processed as necessary, and then stored in computer memory. In the context of this document, a computer-usable or computer-readable medium can be any medium that can contain, store, communicate, propagate, or transmit a program used by or in connection with an instruction execution system, apparatus, or device. The computer-usable medium may include propagated data signals having computer-usable program code implemented therein in baseband or as part of a carrier wave. The computer-usable program code may be transmitted using any suitable medium, including but not limited to the Internet, wired, fiber optic cable, RF, etc.

[0140] Computer program code used to perform the operations of this disclosure may be written in an object-oriented programming language, such as Java, Smalltalk, C++, etc. However, computer program code used to perform the operations of this disclosure may also be written in a conventional procedural programming language, such as the "C" programming language or a similar programming language. The program code may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via a local area network / wide area network / Internet (e.g., network 14).

[0141] This disclosure has been described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer / special-purpose computer / other programmable data processing apparatus, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more of the flowchart illustration and / or block diagram blocks.

[0142] These computer program instructions may also be stored in a computer-readable storage medium, which may instruct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of writing comprising instruction means that implement the functions / actions specified in one or more flowchart and / or block diagram frames.

[0143] Computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing one or more functions / actions specified in a flowchart and / or block diagram box.

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in the order indicated in the drawings. For example, in fact, two blocks shown consecutively may be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, not executed at all, or executed according to any combination of the functions involved with any other flowchart. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware, or a combination of dedicated hardware and computer instructions, that performs the specified function or action.

[0145] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0146] All means or steps in the appended claims, along with corresponding structures, materials, actions, and equivalents of the functional elements, are intended to protect according to the specific claims, including any structure, material, or action used to perform the function in conjunction with other claimed elements. The description in this disclosure is for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit the disclosure to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of this disclosure. The embodiments were chosen and described in order to best explain the principles and practical application of this disclosure and to enable those skilled in the art to understand the disclosure of various embodiments with various modifications suitable for the particular intended use.

[0147] Several implementations have been described. Having described this disclosure in such detail and with reference to its embodiments, it will be apparent that modifications and variations are possible without departing from the scope of this disclosure as defined in the appended claims.

Claims

1. A computer-implemented method executed on a computing device, comprising: receiving a signal from each of a plurality of microphones, thereby defining a plurality of signals; determining a total harmonic distortion associated with at least one microphone; and performing one or more harmonic distortion based augmentations on the plurality of signals based at least in part on the total harmonic distortion associated with the at least one microphone, thereby defining one or more harmonic distortion based augmented signals.

2. The computer-implemented method of claim 1, wherein determining the total harmonic distortion associated with the at least one microphone comprises: receiving a harmonic distortion parameter associated with the at least one microphone.

3. The computer-implemented method of claim 2, wherein the harmonic distortion parameter indicates an order of a harmonic associated with the at least one microphone.

4. The computer-implemented method of claim 2, wherein performing the one or more harmonic distortion-based augmentations on the plurality of signals based at least in part on the harmonic distortion parameter comprises: generating a harmonic distortion based augmented signal based at least in part on the harmonic distortion parameter and a harmonic distortion coefficient table.

5. The computer-implemented method of claim 4, wherein determining the total harmonic distortion associated with the at least one microphone comprises: measuring the total harmonic distortion from the at least one microphone.

6. The computer-implemented method of claim 5, further comprising: generating the harmonic distortion coefficient table based at least in part on the total harmonic distortion measured from the at least one microphone.

7. The computer-implemented method of claim 1, wherein the plurality of microphones define a microphone array.

8. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon that, when executed by a processor, cause the processor to perform operations comprising: receiving a signal from each of a plurality of microphones, thereby defining a plurality of signals; determining a total harmonic distortion associated with at least one microphone; and performing one or more harmonic distortion based augmentations on the plurality of signals based at least in part on the total harmonic distortion associated with the at least one microphone, thereby defining one or more harmonic distortion based augmented signals.

9. The computer program product of claim 8, wherein determining the total harmonic distortion associated with the at least one microphone comprises: receiving a harmonic distortion parameter associated with the at least one microphone.

10. The computer program product of claim 9, wherein the harmonic distortion parameter indicates an order of a harmonic associated with the at least one microphone.

11. The computer program product of claim 9, wherein performing the one or more harmonic distortion based enhancements on the plurality of signals based at least in part on the harmonic distortion parameter comprises: generating a harmonic distortion based augmented signal based at least in part on the harmonic distortion parameter and a harmonic distortion coefficient table.

12. The computer program product of claim 11, wherein determining the total harmonic distortion associated with the at least one microphone comprises: measuring the total harmonic distortion from the at least one microphone.

13. The computer program product of claim 12, wherein the operations further comprise: generating the harmonic distortion coefficient table based at least in part on the total harmonic distortion measured from the at least one microphone.

14. The computer program product of claim 8, wherein the plurality of microphones define a microphone array.

15. A computing system comprising: a memory; and a processor coupled to the memory. a processor configured to receive a signal from each microphone of a plurality of microphones, thereby defining a plurality of signals, wherein the processor is further configured to determine a total harmonic distortion associated with at least one microphone, and wherein the processor is further configured to perform one or more harmonic distortion based augmentations on the plurality of signals based at least in part on the total harmonic distortion associated with the at least one microphone, thereby defining one or more harmonic distortion based augmented signals.

16. The computing system of claim 15, wherein determining the total harmonic distortion associated with the at least one microphone comprises: receiving a harmonic distortion parameter associated with the at least one microphone.

17. The computing system of claim 16, wherein the harmonic distortion parameter indicates an order of a harmonic associated with the at least one microphone.

18. The computing system of claim 16, wherein performing the one or more harmonic distortion based augmentations on the plurality of signals based at least in part on the harmonic distortion parameter comprises: generating a harmonic distortion based augmented signal based at least in part on the harmonic distortion parameter and a table of harmonic distortion coefficients.

19. The computing system of claim 18, wherein determining the total harmonic distortion associated with the at least one microphone comprises: measuring the total harmonic distortion from the at least one microphone.

20. The computing system of claim 15, wherein the processor is further configured to generate the table of harmonic distortion coefficients based at least in part on the plurality of harmonic distortions measured from the at least one microphone.

Citation Information

Patent Citations

  • Digital Compensation for Non-Linearity in Displacement Sensors

    US20140336967A1

  • Linearity error compensator

    US6570514B1