Speaker-specific voice amplification
A speaker-specific acoustic model enhances user voice and suppresses background noise in real-time, addressing the challenge of clear communication in noisy environments by selectively amplifying the user's voice and adapting to environmental changes.
Patent Information
- Application Number
- JP2023536933
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-18
- Filing Date
- 2021-11-17
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Background noise in audio conversations, such as those occurring in public spaces or remote work environments, often interferes with clear communication, as existing technologies struggle to selectively amplify a user's voice while suppressing background noise.
A speaker-specific acoustic model is created using a user's voice sample to enhance their speech, allowing the system to selectively amplify the user's voice and suppress background noise in real-time, using pre-trained models and dynamic filters to adjust for varying environmental conditions.
The system effectively enhances the user's voice while reducing background noise, ensuring clear communication without requiring specific hardware, and can adapt to changing environments and user conditions, such as illness or language changes.
Smart Images

Figure 0007790842000001 
Figure 0007790842000002 
Figure 0007790842000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates to digital signal processing, and more particularly to speaker-specific systems and methods for amplifying audio. [Background technology]
[0002] The development of the EDVAC system in 1948 is often cited as the beginning of the computer age. Since then, computer systems have evolved into extremely complex devices. Modern computer systems typically contain a combination of sophisticated hardware and software components, application programs, operating systems, processors, buses, memory, and input / output devices. As advances in semiconductor processing and computer architectures continue to push performance, more advanced computer software has evolved to take advantage of the increased capabilities of these features, resulting in computer systems today that are far more powerful than they were just a few years ago.
[0003] One application of these new features is the mobile phone. Today, people routinely make calls in public spaces (e.g., cafes or trains) or work remotely. In these environments, background noise from children, spouses, pets, construction, and many other factors can interfere with conversations. Summary of the Invention
[0004] According to an embodiment of the present disclosure, there is provided a method for using a computing device to amplify a single voice during an audio conversation. One embodiment of the method may include receiving, by the computing device, an audio sample of speech from a user and generating, by the computing device, a user-specific acoustic model for enhancing the user's speech based on the audio sample. The method may further include receiving, during the audio conversation, a live audiovisual stream including live speech by the user, the live audiovisual stream including background noise, and using, by the computing device, the user-specific acoustic model to selectively amplify the live speech in the live audiovisual stream without amplifying the background noise.
[0005] According to an embodiment of the present disclosure, a computer program product for selectively amplifying a user's voice using a pre-trained acoustic model is provided. One embodiment of the computer program product may include a computer-readable storage medium having program instructions embodied therewith. The program instructions may be executable by a processor to cause the processor to extract voice data for the user from existing voice samples, create a pre-trained acoustic model for the user from the voice data, analyze an audio stream from the conference call, detect the presence of background noise in the audio stream, and apply the pre-trained acoustic model to the audio stream to selectively amplify the user's voice and not the background noise.
[0006] According to an embodiment of the present disclosure, a computer system for amplifying a single voice during an audio conversation may include a processor configured to execute program instructions that, when executed on the processor, cause the processor to: receive an audio sample of speech from a user; generate a user-specific acoustic model for enhancement of the user's speech based on the audio sample; receive a live audiovisual stream, the live audiovisual stream including the live speech by the user during the audio conversation and including background noise; and use the user-specific acoustic model to selectively amplify the live speech in the live audiovisual stream without amplifying the background noise.
[0007] The above summary is not intended to describe each illustrated embodiment or every implementation of the present disclosure.
[0008] The drawings included in this application are incorporated into and form a part of the specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. The drawings are merely illustrative of particular embodiments and are not intended to limit the disclosure. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an embodiment of a data processing system (DPS), consistent with some embodiments. [Figure 2] FIG. 1 is a diagram depicting a cloud computing environment consistent with some embodiments. [Figure 3] FIG. 1 illustrates an abstraction model layer consistent with some embodiments. [Figure 4] FIG. 1 is a system diagram for a computing environment consistent with some embodiments. [Figure 5]1 is a flowchart of an operational noise reduction service consistent with some embodiments. [Figure 6] 1 is a flowchart illustrating a method for training a machine learning model, consistent with some embodiments. [Figure 7] 1 is a flowchart of a conferencing system in operation consistent with some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0010] While the invention is susceptible to various modifications and alternative forms, specifics thereof have been shown by way of example in the drawings and will be described in detail. It should be understood, however, that the intention is not to limit the invention to the particular embodiments described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives within the scope of the invention.
[0011] Aspects of the present disclosure relate to digital signal processing, and more particularly to speaker-specific systems and methods for amplifying audio. While the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure will be understood through a discussion of various examples using this context.
[0012] Some embodiments of the present disclosure may include a system that builds a speaker-specific acoustic model of a user from a recording. Some embodiments may then use the speaker-specific acoustic model to isolate the user's voice in a future live audio stream and / or live audiovisual stream so that the embodiments can boost the volume of only those portions of the signal containing the user's spoken content and / or reduce the volume of noise from undesirable sources (e.g., background). That is, some embodiments may (a) reduce / remove undesirable background noise and (b) emphasize the speaker's own voice. Background noise, in turn, may include relatively static sounds generated by the conferencing system itself (e.g., static sounds recorded and / or generated by the user's microphone, artifacts created during digital compression or transmission of the microphone's output, etc.), relatively static sounds generated by environmental sources (e.g., local HVAC equipment, vehicle engine and tire sounds, aircraft engine sounds, etc.), dynamic sounds (e.g., other people talking nearby, sounds generated by nearby construction equipment and work, dogs barking, ambulance sirens, etc.), and other sounds that are not the speaker's own voice.
[0013] In operation, a user may begin by recording their own voice in a quiet environment (i.e., substantially free of background noise) to create a clear target recording or set of target recordings. Then, in some embodiments, the clear recording may be analyzed to generate one or more parameters for a voice profile. Then, in some embodiments, the one or more parameters may be used to build a speaker-specific acoustic model of the user's voice. This speaker-specific acoustic model may be specifically tuned to the voice characteristics of a particular user (i.e., each user has a different, or even unique, acoustic model associated with it). In this way, the speaker-specific acoustic model can be adapted to the voice of that particular user and its speech pattern characteristics, such as its pitch and frequency range.
[0014] When an acoustic model is constructed and / or trained, the same features used to create the model can be extracted and processed from future live audio or audiovisual streams. If no modifications to these streams are required, the original input can be sent to the network platform; otherwise, the disclosed system can improve the signal and send it to the platform (and / or other participants in the conference call) in near real time (e.g., in less than 100 milliseconds). This process can include measuring the level of background noise and attenuating it to an acceptable level (e.g., similar to that seen during initial training). The model in these embodiments can observe and analyze distinct speech patterns associated with the user's speaker-specific characteristics specified during initial training.
[0015] In some embodiments, the system may detect when a user joins a conference call by analyzing the user's interaction with an associated electronic device, which may be a computer or phone with a built-in microphone. Also, in some embodiments, the system may determine that a user is currently presenting by comparing the user's name to the agenda or by detecting the content of the current speech, and selectively amplify the user's voice while suppressing background noise and the voices of other participants.
[0016] In response to the detection, some embodiments may apply a dynamic band-pass or band-stop filter to selectively amplify and / or suppress certain frequencies when transmitting and / or retransmitting the audio stream. This suppression and / or amplification may occur in the time domain and / or the frequency domain, in some embodiments. In this way, the user's voice signal may be effectively strengthened while being transmitted and / or retransmitted, and other noise (other voices, non-speech sounds) may be suppressed and / or not transmitted / retransmitted.
[0017] Additionally, some embodiments may detect that there is some environmental noise (e.g., a vehicle horn) or that other people (e.g., children) are speaking simultaneously in the background. The system may then identify the detected background noise and / or the detected voices of other people as undesirable signals and remove them from the transmit / retransmit stream. More specifically, some embodiments may automatically activate and deactivate pre-trained acoustic models associated with presenters to enhance the presenter's voice as they speak. Furthermore, some embodiments may automatically and selectively modify the pre-trained acoustic models to compensate for detected noise and / or voices. In this way, other conference participants can hear the presenter more clearly and without interference. Some embodiments may be configured with configurable thresholds for what is detected as undesirable background noise, such as thresholds based on duration, loudness, amount of deviation from a pre-trained user's pitch range, etc. In this way, users and applications requiring acoustic fidelity can reduce the amount and / or degree of filtering performed by some embodiments.
[0018] In some embodiments, the acoustic model can be based on a user's acoustic voiceprint (profile). These features include, but are not limited to, pitch variations and perturbations (e.g., jitter), periodicity measures (e.g., harmonic noise ratio), linear predictive coding coefficients (LPC), spectral shape measures, speech onset time, Mel-Frequency Cepstral Coefficients (MFCC), i-vectors, etc. The model can be initially trained on the user's speaker-specific voiceprint, with the user's approval. In some embodiments, the acoustic model can include unsupervised speech alignment algorithms, hidden Markov models, unwanted noise removal algorithms, phoneme token DNNs, etc. Some of these features (e.g., formants F1 and F2) may be used primarily to characterize the uniqueness of the user's voice (e.g., obtaining a fingerprint), while others (e.g., harmonic noise ratio) may be used primarily to characterize background noise.
[0019] In some embodiments, acoustic features that can be analyzed to construct a speaker-specific voice profile include, without limitation, one or more features selected from the group consisting of fundamental frequency, spectral envelope, pitch characteristics (e.g., mean, maximum, minimum, etc.), voice onset times (VOTs) of various consonants, F1 and F2 for characterizing vowel pronunciation, and vowel duration. These acoustic features can be used to distinguish one user's voice from other voices (especially the consonant and vowel features mentioned above) and non-speech sounds (especially the high-level pitch and frequency features mentioned above).
[0020] Some embodiments may also create supplemental acoustic models for the user when their voice sounds slightly different, such as when they have a clogged throat, are in pain, etc. Similarly, some embodiments may also create customized speaker-specific acoustic models for each language the user speaks, since different languages may have different phonological profiles even for the same speaker.
[0021] As a first illustrative example, assume that a user (“A”) works remotely. As a result, User A is at home much of the time with her husband, children, and pets, all of which create background noise while making business calls. In particular, when User A calls into a conference, she often hears her dog barking in the background. In this example, the acoustic model has been previously trained on User A's voice and has a unique acoustic profile customized for her voice and speech patterns. Thus, the system in this example can detect User A's voice and then selectively amplify only her voice while simultaneously reducing the noise made by her dog that is broadcast to other conference participants.
[0022] As a second illustrative example, user (“B”) is running late for a meeting. User B is calling into the meeting from within his or her car. As a result, there is a lot of background noise from the road, the car engine, and even his or her crying baby in the backseat. Traditionally, User B could partially mitigate this noise by muting himself or herself, but doing so would prevent him or her from fully participating in the meeting without distracting his or her colleagues with all of this extraneous noise. However, using some embodiments of the present disclosure, a speaker-specific acoustic model can be pre-trained on User B's voice. Then, when he or she calls from within his or her car, some embodiments can identify and enhance various acoustic characteristics unique to User B's voice while reducing acoustic characteristics inconsistent with his or her own voice (e.g., from the road, his or her car, etc.). In this way, User B can speak in the meeting without transmitting loud noises from his or her environment.
[0023] As a third illustrative example, user (“C”) has a slight cold and, as a result, is working remotely out of consideration for his or her coworkers. While user C is participating in a conference call by phone, user C occasionally coughs and sneezes. Using some embodiments, an acoustic model may be pre-trained on user C's normal speech, e.g., without coughs or sneezes, and thus selectively enhance his or her speech signal while selectively filtering out coughs and sneezes as undesirable background noise. Further, knowing that he or she is prone to coughing and sneezing during conference calls, user C may, in some embodiments, configure a threshold for what is considered undesirable background noise based on either length of time, loudness, or amount of deviation from the user's pre-trained pitch range, etc., to ensure that these coughs are filtered out.
[0024] As a fourth illustrative example, a user (“D”) is ill with a sore throat and is working remotely to prevent the spread of the infection to his or her coworkers. However, as a result of the infection, User D's voice now sounds different from his or her normal voice. Some embodiments of the present disclosure enable User D to create additional customized acoustic models (or adjust his or her associated acoustic models) of the user's voice (e.g., his or her “normal” speech as well as his or her current “abnormal” voice, e.g., due to the illness). In this way, User D can choose to train a separate version of the system for his or her sore throat when he or she telephones into a conference, so that even though his or her current voice has acoustic characteristics that are somewhat different from his or her normal voice, the system can recognize and enhance it as opposed to suppressing it as background noise. Furthermore, in some embodiments, User D's current voice can selectively have those acoustic characteristics modified so that it sounds more like his or her normal voice when presented to other conference participants.
[0025] Thus, one feature and advantage of some embodiments is that they do not require the user to have and use specific hardware, such as a directional microphone, to amplify the user's voice and / or suppress noise. As a result, some embodiments can be integrated as a plug-in into existing videoconferencing systems and microphone processing software. Another feature and advantage of some embodiments is that they selectively amplify the user's voice using a pre-trained acoustic model. In this way, some embodiments can continue to broadcast sounds with a larger dynamic range, including desired but unexpected noise. Another feature and advantage of some embodiments is that they can learn the user's voice characteristics and selectively enhance the signal based thereon. In this way, some embodiments enable noise reduction even when the user is moving around in their local physical environment, allowing two or more speakers to use one microphone at a time, or allowing two or more speakers to speak at once, or a combination thereof.
[0026] Data Processing System FIG. 1 illustrates one embodiment of a data processing system (DPS) 100a consistent with some embodiments. In this embodiment, the DPS 100a may be implemented as a personal computer; a server computer; a portable computer, such as a laptop or notebook computer, a personal digital assistant (PDA), a tablet computer, or a smartphone; a processor embedded in a larger device, such as an automobile, an airplane, a teleconferencing system, or an appliance; a smart device; or any other suitable type of electronic device. Components other than or in addition to those shown in FIG. 1 may also be present, and the number, type, and configuration of such components may vary. Also, FIG. 1 depicts only representative major components of the DPS 100a, and individual components may be more complex than depicted in FIG. 1.
[0027] The data processing system 100a of FIG. 1 includes multiple central processing units 110a-110d (collectively referred to herein as processors 110 or CPUs 110) connected by a system bus 122 to memory 112, a mass storage interface 114, a terminal / display interface 116, a network interface 118, and an input / output ("I / O") interface 120. The mass storage interface 114 in this embodiment connects the system bus 122 to one or more mass storage devices, such as a direct access storage device 140, a universal serial bus ("USB") storage device 141, or a readable / writable optical disk drive 142. The network interface 118 enables the DPS 100a to communicate with another DPS 100b via a communications medium 106. The memory 112 also includes an operating system 124, multiple application programs 126, and program data 128.
[0028] The data processing system 100a in the embodiment of FIG. 1 is a general-purpose computing device. Accordingly, the processor 110 may be any device capable of executing program instructions stored in memory 112 and may itself consist of one or more microprocessors and / or integrated circuits. In this embodiment, the DPS 100a includes multiple processors and / or processing cores, as is typical of larger, more powerful computer systems; however, in other embodiments, the data processing system 100a may comprise a single processor system, a single processor designed to emulate a multiprocessor system, or both. Furthermore, the processor 110 may be implemented using multiple heterogeneous data processing systems 100a, in which a main processor resides on a single chip with secondary processors. As another illustrative example, the processor 110 may be a symmetric multiprocessor system including multiple processors of the same type.
[0029] When data processing system 100a boots, the associated processor 110 initially executes program instructions that make up operating system 124, which manages the physical and logical resources of DPS 100a. These resources include memory 112, mass storage interface 114, terminal / display interface 116, network interface 118, and system bus 122. As with processor 110, some DPS 100a embodiments may utilize multiple system interfaces 114, 116, 118, 120, and bus 122, each of which may include its own separate, fully programmed microprocessor.
[0030] Instructions for the operating system, applications, or programs, or combinations thereof (commonly referred to as "program code," "computer-usable program code," or "computer-readable program code") may initially be located on mass storage devices 140, 141, 142, which are in communication with processor 110 via system bus 122. Program code in different embodiments may be embodied on different physical or tangible computer-readable media, such as system memory 112 or mass storage devices 140, 141, 142. In the illustrative example of FIG. 1, the instructions are stored in a functional form of persistent storage on direct access storage device 140. These instructions are then loaded into memory 112 for execution by processor 110. However, program code may also be located in a functional form on selectively removable computer-readable media and loaded or transferred to DPS 100a for execution by processor 110.
[0031] The system bus 122 may be any device that facilitates communication between the processor 110, the memory 112, and the interfaces 114, 116, 118, 120. Additionally, although the system bus 122 in this embodiment is a relatively simple single bus structure that provides a direct communication path between the system buses 122, other bus structures are also consistent with the present disclosure, including, but not limited to, hierarchical configurations, point-to-point links in a star or web configuration, multiple hierarchical buses, parallel and redundant paths, etc.
[0032] Memory 112 and mass storage devices 140, 141, and 142 work together to store operating system 124, application program 126, and program data 128. In this embodiment, memory 112 is a random-access semiconductor device capable of storing data and programs. While FIG. 1 conceptually depicts the device as a single monolithic entity, memory 112 in some embodiments may be a more complex configuration, such as a hierarchy of caches and other memory devices. For example, memory 112 may exist in multiple levels of caches, and these caches may be further divided by function, such as one cache holding instructions and another cache holding non-instruction data used by the processor. Memory 112 may also be further distributed and associated with different processors 110 or sets of processors 110, as known in any of a variety of so-called non-uniform memory access (NUMA) computer architectures. Additionally, some embodiments may utilize a virtual addressing mechanism that allows DPS 100a to behave as if it has access to a large single storage entity rather than accessing multiple smaller storage entities such as memory 112 and mass storage devices 140, 141, 142.
[0033] Although operating system 124, application programs 126, and program data 128 are illustrated as being contained within memory 112, some or all of them may be located on different physical computer systems and, in some embodiments, may be accessed remotely, for example, via communications medium 106. Thus, while operating system 124, application programs 126, and program data 128 are illustrated as being contained within memory 112, these elements are not necessarily all contained entirely within the same physical device at the same time, and may reside in the virtual memory of another DPS, such as DPS 100b.
[0034] System interfaces 114, 116, 118, and 120 support communication with various storage and I / O devices. Mass storage interface 114 supports attachment of one or more mass storage devices 140, 141, and 142, which may be a rotating magnetic disk drive storage device, a solid-state storage device (SSD) that uses integrated circuit assemblies as memory to persistently store data, typically using flash memory, or a combination of the two. However, mass storage devices 140, 141, and 142 may also comprise other devices, including arrays of disk drives (commonly referred to as RAID arrays) and / or archival storage media, such as hard disk drives, tape (e.g., miniDV), writable compact discs (e.g., CD-R and CD-RW), digital versatile discs (e.g., DVD, DVD-R, DVD+R, DVD+RW, DVD-RAM), holographic storage systems, blue laser discs, IBM® Millipede devices, and the like, configured to appear as a single large storage device to the host.
[0035] Terminal / display interface 116 is used to directly connect one or more display units 180, including monitors and the like, to data processing system 100a. These display units 180 may be non-intelligent (i.e., dumb) terminals, such as LED monitors, or may themselves be fully programmable workstations used to allow IT administrators and customers to communicate with DPS 100a. However, it should be noted that while display interface 116 is provided to support communication with one or more display units 180, data processing system 100a does not necessarily require a display unit 180, as all of the necessary interaction with customers and other processes occurs through network interface 118.
[0036] The communication medium 106 may be any suitable network or combination of networks and may support any suitable protocol for communicating data and / or code between the multiple DPSs 100a, 100b. Accordingly, the network interface 118 may be any device that facilitates such communication, regardless of whether the network connection is made using current analog and / or digital technologies, or via any future networking mechanism. Suitable communication media 106 include, but are not limited to, networks implemented using one or more of the following: "InfiniBand" or IEEE (Institute of Electrical and Electronics Engineers) 802.3x "Ethernet" specifications; cellular transmission networks; wireless networks implementing any of the IEEE 802.11x, IEEE 802.16, General Packet Radio Service (GPRS), Family Radio Service (FRS), or Bluetooth® specifications; ultra-wideband (UWB) technology as described in FCC 02-48; and others. Those skilled in the art will appreciate that many different network and transport protocols can be used to implement the communications medium 106. The Transmission Control Protocol / Internet Protocol ("TCP / IP") suite includes suitable network and transport protocols.
[0037] Cloud Computing 2 illustrates a cloud environment that includes one or more DPSs 100a, 100b consistent with some embodiments. While this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present disclosure may be implemented in conjunction with any other type of computing environment now known or later developed.
[0038] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0039] Its features are as follows: On-demand self-service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without requiring human interaction with the provider of the service. Broad network access: Functionality is available over the network and accessed through standard mechanisms, facilitating use across heterogeneous thin- or thick-client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources being dynamically allocated and reallocated according to demand. Consumers generally have no control over or knowledge of the exact location of the resources provided, but there is a sense of location independence in that they may be able to specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: Capabilities can be delivered quickly and elastically, sometimes automatically, allowing for rapid scaling out and rapid release to rapidly scale in. To the consumer, the capabilities available for provisioning often appear unlimited, and can be purchased in any quantity at any time. Measured service: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at an abstraction level appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). The ability to monitor, control, and report resource usage provides transparency to both providers and consumers of utilized services.
[0040] The service model is as follows: Software as a Service (SaaS): The functionality offered to the consumer is the use of the provider's applications running on a cloud infrastructure. These applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). With the possible exception of limited user-specific application configuration settings, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application functions. Platform as a Service (PaaS): The capability offered to consumers is the deployment onto a cloud infrastructure of applications they create or acquire using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application hosting environment configuration. Infrastructure as a Service (IaaS): The capability offered to consumers is to provision processing, storage, network, and other basic computing resources on which the consumer can deploy and run any software, including operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating systems, storage, deployed applications, and possibly limited control over select network components (e.g., host firewalls).
[0041] The deployment model is as follows: Private Cloud: Cloud infrastructure is operated exclusively for an organization. It is managed by the organization or a third party and may reside on or off premises. Community Cloud: Cloud infrastructure is shared by multiple organizations to support a specific community of shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by the organization or a third party and may reside on or off premises. Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services. Hybrid Cloud: A cloud infrastructure is a composite of two or more clouds (private, community, or public) that remain distinct entities but are tied together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0042] Cloud computing environments are service-oriented with an emphasis on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure, which includes a network of interconnected nodes.
[0043] Referring now to FIG. 2, an illustrative cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof, may communicate. The nodes 10 may communicate with each other. They may be physically or virtually grouped in one or more networks, or combinations thereof, such as private, community, public, or hybrid clouds as described herein above (not shown). This enables the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service without the cloud consumer having to maintain resources on their local computing device. The types of computing devices 54A-54N shown in FIG. 2 are for illustrative purposes only, and it will be understood that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device (e.g., using a web browser) over any type of network and / or network-addressable connection.
[0044] Referring now to Figure 3, a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 2) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 3 are for illustrative purposes only and are not intended to limit the scope of the present invention. As shown in the figure, the following layers and corresponding functions are provided:
[0045] The hardware / software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, a RISC (reduced instruction set computer) architecture-based server 62, a server 63, a blade server 64, storage devices 65, and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0046] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73, including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.
[0047] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection of data and other resources. Customer portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management to meet required service levels. Service level agreement (SLA) planning and fulfillment 85 provides advance arrangement and procurement of cloud computing resources for anticipated future requirements according to SLAs.
[0048] The workload layer 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91; software development and lifecycle management 92; virtual classroom instruction delivery 93; data analytics processing 94; transaction processing 95; and noise reduction services 96.
[0049] Acoustic Platform 4 is a system diagram of a computing environment 400 consistent with some embodiments. The computing environment 400 includes a conferencing computer platform 402 connected to multiple user devices 403 over a network 406. The conferencing computer platform 402, in turn, includes a conferencing module 480 and a noise reduction service 496. The noise reduction service 496 includes a trained machine learning model 498 that generates multiple customized acoustic profiles 499, one or more for each user of the conferencing computer platform 402. The conferencing module 480 in some embodiments includes a database 482 that includes the multiple customized acoustic profiles.
[0050] The conferencing computer platform 402 may be a standalone computing device, an administrative server, a web server, a mobile computing device, or any other electronic device or computing system capable of receiving, transmitting, and processing data. In some embodiments, the conferencing computer platform 402 may be part of a cloud computing environment 50 and represent a pool of computing resources within that environment. The conferencing computer platform 402 includes internal and external hardware components, as depicted and described in further detail with reference to the DPS 100 in FIG. 1 .
[0051] User device 403 can represent one or more programmable electronic devices, or a combination of programmable electronic devices, capable of executing machine-readable program instructions and communicating with other computing devices (not shown) over network 406 within computing environment 400. Suitable user devices 403 include, but are not limited to, desktop computers, laptop computers, tablet computers, smartphones, smart watches, and voice-grade telephone handsets.
[0052] In some embodiments, user device 403 includes one or more Voice over Internet Protocol (VoIP) compatible devices (i.e., Voice over IP, IP telephony, broadband telephony, and broadband telephone service). VoIP generally refers to a set of methodologies and technologies for delivering voice communications and multimedia sessions over Internet Protocol (IP) networks, such as the Internet. VoIP can be integrated into any common computing device, such as a smartphone, a personal computer, or a user device 403 capable of communicating with network 406. In some embodiments, user device 403 includes one or more speakerphones adapted for use in an audio or video conferencing environment. A speakerphone generally refers to an audio device that includes at least a loudspeaker, a microphone, and one or more microprocessors.
[0053] In some embodiments, the user devices 403 include a user interface (not shown separately in FIG. 1 ). This user interface can provide an interface between the user devices 403 and the conferencing computer platform. In some embodiments, the user interface can include information (such as graphics, text, and sound) that a program presents to a user, as well as control sequences that the user employs to control the program, as a graphical user interface (GUI) or web user interface (WUI), which can display text, documents, web browser windows, user options, application interfaces, and instructions for operation. In another embodiment, the user interface includes mobile application software that provides an interface between each user device 403 and the conferencing computer platform 402. Mobile application software, or “apps,” is a class of computer programs that typically run on smartphones, tablet computers, smart watches, and other mobile devices.
[0054] 4 may comprise, for example, a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination thereof, and may include wired, wireless, or fiber optic connections. Network 406 may utilize a combination of connections and protocols that support communications between conferencing computer platform 402 and user devices 403, such as receiving and transmitting data, voice, and / or video signals, including multimedia signals including voice, data, and video information.
[0055] The conferencing computer platform, in some embodiments, can provide a noise reduction service 496. FIG. 5 is a flowchart of the noise reduction service 496 in operation, consistent with some embodiments. In operation 505, a user may create an audio sample (e.g., a recording). In some embodiments, the user is asked to read a predetermined script in a quiet environment using a high-quality microphone. The script may be selected to include a number of distinctive audio features, such as the most common phonemes or all phonemes of a given language. This audio sample is then used as part of a training phase 508.
[0056] During the training phase 508, the noise reduction service 496 receives audio samples at operation 510. In response, the noise reduction service 496 may extract audio features from the audio samples at operation 515 and then send those features to the trained machine learning model 498. The machine learning model 498 may then generate a first customized acoustic profile 499a for the user from the features at operation 520. This acoustic profile 499a may be optimized to selectively identify and / or isolate the full dynamic range of the user's voice from the recording stream. At operations 525-530, the user optionally listens to the recording (created in operation 505) as processed by the customized acoustic profile 499a and is then given the opportunity to approve or reject the model. If the user rejects the customized acoustic profile 499a (530: NO), the noise reduction service 496 may return to operations 505 and 510 to collect and process new samples. If the user accepts the customized acoustic profile 499a (530: YES), the system can output a model for use in future live meetings.
[0057] The user and / or the noise reduction service 496 can then initiate an additional training phase 550. This additional training allows for the creation of additional supplemental acoustic profiles 499b-499n for the user. These supplemental profiles may be created depending on the user's current physical condition, such as a cold or sore throat, or may optimize acoustic profiles for alternative languages. For example, a user may have one profile 499a for their normal voice when speaking English, one profile 499b for their coughing voice when speaking English, one profile 499c for their normal voice when speaking Spanish, and one profile 499d for use on specific days when they are sick. In this illustrative example, the user may have four separate acoustic profiles 499 stored in the system and select one at the start of a conference call.
[0058] Similar to the initial training phase 508, the tuning phase begins with receiving, at 555, new audio samples submitted by the user. In response, the noise reduction service 496 can extract audio features from the new audio samples and then send those features to the trained machine learning model 498 at operation 560. The voice improvement module 497 of the machine learning model 498 can then generate supplemental acoustic profiles 499b-499n at operation 565. In some embodiments, this involves receiving and tuning the original acoustic profile 499. At operations 570-575, the user is optionally given the opportunity to listen to the recording (created in operation 505) as processed by the supplemental acoustic profile 499a and then approve or reject the updated model. If the user rejects the supplemental acoustic profiles 499b-499n (575: NO), the noise reduction service 496 can return to operations 505, 555 to collect and process new samples. If the user accepts the supplemental acoustic profiles 499b-499n (575: YES), the system can output a model for use in future live meetings. In some embodiments, this includes allowing the user to select which acoustic profiles 499a-499n to use in the meeting.
[0059] Model Training In some embodiments, the machine learning model 498 may be any software system that recognizes patterns. In some embodiments, the machine learning model includes multiple artificial neurons interconnected through connection points called synapses. Each synapse may encode the strength of the connection between the output of one neuron and the input of another neuron. The output of each neuron is then determined by the aggregate inputs it receives from the other neurons connected to it, and therefore by the outputs of these "upstream" connected neurons, with the strength of the connection determined by the synaptic weights.
[0060] ML models can be trained to solve specific problems (e.g., to generate customized user-specific acoustic model configurations) by adjusting synaptic weights so that specific classes of inputs produce desired outputs. This weight adjustment procedure in these embodiments is known as "learning." Ideally, during the learning process, these adjustments lead to a pattern of synaptic weights that converge toward an optimal solution for the given problem based on some cost function.
[0061] In some embodiments, artificial neurons can be organized into layers. The layer that receives external data is the input layer. The layer that produces the final result is the output layer. Some embodiments include hidden layers between the input and output layers, typically including hundreds of such hidden layers.
[0062] 6 is a flowchart illustrating one method 600 of training a machine learning model 498, consistent with some embodiments. The system manager may begin by loading training vectors at operation 610. These vectors may include recordings from multiple different users taken in a quiet room reading specially prepared transcripts.
[0063] At operation 612, the system manager may select the desired output (e.g., optimal settings for the acoustic model 499). At operation 614, the training data may be prepared to reduce sources of bias, typically including de-duplication, normalization, and order randomization. At operation 616, initial gate weights for the machine learning model may be randomized. At operation 618, the ML model may be used to predict an output using a set of input data vectors, and the prediction is compared to labeled data. The error (e.g., the difference between the predicted value and the labeled data) is then used to update the gate weights at operation 620. This process is repeated, updating the weights at each iteration, until the training data is exhausted or the ML model reaches an acceptable level of accuracy and / or precision. At operation 622, the resulting model may optionally be compared to previously unevaluated data to validate and test its performance. In operation 624, the resulting model is loaded into the noise reduction service 496 in the cloud computing environment 50 and can be used to analyze user recordings.
[0064] Conference System FIG. 7 is a flowchart of a conferencing system 700 in operation consistent with some embodiments. At operation 705, multiple participants in a conference call register and / or log in to the conferencing system 700, e.g., using a username and password. At operation 710, the conferencing system 700 will query the database 482 for current acoustic models 499 associated with each of the participants. If a model does not exist for one or more participants (711: NO), the system may prompt the one or more participants to record raw audio data of their speech to begin building a speaker-specific acoustic model at operation 712. Optionally, the system may also give participants the option of using a universal model (i.e., a model created to separate a wide variety of voices and languages) for this conference call, which may be desirable if they do not have the time and / or equipment to create a customized model. If a participant has created multiple customized models, the system may prompt the participant to select a model to use for this call at operation 713.
[0065] One of the participants can then begin speaking. Their user devices 403 can record their sounds at operation 715, convert the recording into an original audio stream at operation 720, and transmit the original audio stream to the conferencing module 480 at operation 725. In response, the conferencing system 700 applies an acoustic model customized for that particular speaker (identified at operation 705) to the received audio stream to generate an optimized audio stream (e.g., one in which the user's voice is amplified and / or any background noise is suppressed) at operation 730. The conferencing system 700 can then transmit the optimized audio stream to the other participants in the active conference call at operation 735. These embodiments may be desirable because they can be used with any user devices 403, such as "plain old telephone system" handsets.
[0066] Alternatively, in some embodiments, some or all of the user devices 403 apply customized acoustic models locally (e.g., by a processor within the user devices 403) to the original audio stream and then transmit the optimized audio stream (rather than the original audio stream) to the conferencing system 700 in operation 725. The conferencing system 700 can then proceed directly to operation 735 and retransmit the optimized audio stream to the other participants. These embodiments may be desirable because they can be used with any conferencing system 700.
[0067] Operations 715-735 may be repeated by the conferencing system 700, by the user device 403, or by a combination of both, each time a participant speaks during the duration of the conference call.
[0068] Computer Program Products The present invention is a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail, including a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0069] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge-in-groove structures having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through a fiber optic cable), or an electrical signal transmitted over a wire.
[0070] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0071] Computer-readable program instructions for carrying out the operations of the present invention may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or object-oriented programming languages such as Smalltalk®, C++, and the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server, as a stand-alone software package. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection may be made to the external computer (e.g., through the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0072] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0073] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular way, such that the computer-readable storage medium having instructions stored therein constitutes an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0074] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0075] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing specified logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be accomplished as a single step, executed concurrently, substantially concurrently, in a partially or fully time-overlapping manner, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by special-purpose hardware-based systems that perform the specified functions or operations or execute a combination of special-purpose hardware instructions and computer instructions.
[0076] general Any particular program naming used in this description is for convenience only, and thus the present invention should not be limited to use in any particular application identified and / or implied by such naming. Thus, for example, routines executed to implement embodiments of the present invention may be referred to as "programs," "applications," "servers," or other meaningful naming, whether implemented as part of an operating system or as part of a particular application, component, program, module, object, or sequence of instructions. Indeed, other alternative hardware and / or software environments may be used without departing from the scope of the present invention.
[0077] It is therefore intended that the presently described embodiments be considered in all respects as illustrative and not restrictive, and that reference should be made to the appended claims to determine the scope of the invention.
Claims
1. 1. A method of using a computing device to amplify a single voice in an audio conversation, comprising: generating, for a user, a plurality of user-specific acoustic models each identifying or separating speech by said user from an audiovisual stream, wherein for each user-specific acoustic model: receiving, by a computing device, an audio sample of speech from the user; generating, by the computing device, the respective user-specific acoustic models based on the audio samples; generating a plurality of user-specific acoustic models, wherein a first user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a first language and a second user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a second language; receiving a live audiovisual stream including live speech by the user during the audio conversation, the live audiovisual stream including background noise; selectively amplifying, by the computing device, the live speech by the user in the identified or separated live audiovisual stream without amplifying the background noise using selected ones of the plurality of user-specific acoustic models; A method comprising:
2. 10. The method of claim 1, further comprising selectively suppressing, by the computing device, the background noise in the live audiovisual stream using the selected one of the plurality of user-specific acoustic models.
3. The method of claim 1 or 2, wherein the respective user-specific acoustic models are integrated into the teleconferencing software by a plug-in to the teleconferencing software.
4. The method of claim 3, further comprising generating the plurality of user-specific acoustic models for each one of a plurality of users of the teleconferencing software.
5. collecting the audio samples from the user in an environment substantially free of background noise; using the audio samples to generate the respective user-specific acoustic models; The method of any one of claims 1 to 4, further comprising:
6. The method of claim 5 , further comprising using a trained machine learning model to generate the respective user-specific acoustic model.
7. 1. A computer program for selectively amplifying a user's voice using a pre-trained acoustic model, the computer comprising: generating, for a user, a plurality of user-specific acoustic models each identifying or separating speech by said user from an audiovisual stream, wherein for each user-specific acoustic model: extracting voice data for the user from existing voice samples; creating a pre-trained user-specific acoustic model for each of the users from the speech data; generating a plurality of user-specific acoustic models, wherein a first user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a first language and a second user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a second language; analyzing an audio stream from the conference call; detecting the presence of background noise in the audio stream; applying selected ones of the plurality of user-specific acoustic models to the audio stream to selectively amplify the voice of the identified or isolated user while not amplifying the background noise; A computer program for performing the following.
8. 8. The computer program product of claim 7, further causing the computer to selectively suppress the background noise from the audio stream using selected ones of the plurality of user-specific acoustic models.
9. 9. The computer program product of claim 7, wherein the respective user-specific acoustic models are integrated into teleconferencing software by means of a plug-in to the teleconferencing software.
10. 10. The computer program product of claim 9, further causing the computer to generate the plurality of user-specific acoustic models for each one of a plurality of users of the teleconferencing software.
11. The computer, collecting audio samples from the user in a substantially background noise-free environment; using the audio samples to generate the respective user-specific acoustic models; The computer program according to any one of claims 7 to 10, further comprising:
12. A computer program described in any one of claims 7 to 11, wherein the first user-specific acoustic model of the plurality of user-specific acoustic models is adapted to the user's normal physical condition, and the third user-specific acoustic model of the plurality of user-specific acoustic models is adapted to the user's current physical condition.
13. The computer, receiving, by a computing device, an audio sample of speech from a user from a substantially background noise-free environment; using a trained machine learning model to generate the pre-trained each user-specific acoustic model from the audio samples; The computer program according to any one of claims 7 to 12, further comprising:
14. 1. A system for amplifying a single voice in an audio conversation, comprising: a processor configured to execute program instructions that, when executed on the processor, cause the processor to: generating, for a user, a plurality of user-specific acoustic models each identifying or separating speech by said user from an audiovisual stream, wherein for each user-specific acoustic model: receiving an audio sample of speech from the user; generating the respective user-specific acoustic models based on the audio samples; generating a plurality of user-specific acoustic models, wherein a first user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a first language and a second user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speech in a second language; receiving a live audiovisual stream, the live audiovisual stream including live speech by the user during an audio conversation and including background noise; selectively amplifying the live speech in the identified or separated live audiovisual stream without amplifying the background noise using selected ones of the plurality of user-specific acoustic models; A system that allows the following to be performed.
15. The system described in claim 14, further comprising program instructions for selectively suppressing the background noise in the live audiovisual stream using the selected one of the plurality of user-specific acoustic models.
16. 16. The system of claim 14 or 15, wherein the respective user-specific acoustic models are integrated into the teleconferencing software by a plug-in to the teleconferencing software.
17. The system of claim 16, further comprising program instructions for generating the plurality of user-specific acoustic models for each one of the plurality of users of the teleconferencing software.
18. program instructions for collecting audio samples from the user in a substantially background noise-free environment; program instructions for using the audio samples to generate the respective user-specific acoustic models; The system of any one of claims 14 to 16, further comprising:
19. 20. The system of claim 18, further comprising program instructions for using a trained machine learning model to generate the respective user-specific acoustic model.
Citation Information
Patent Citations
Target person voice enhancement method based on conditional variation auto-encoder
CN111653288A
Speech function equipment and its speaker voice extracting and transmitting method
JP1999154998A
Speech processing
JP2013531275A
Systems and methods for speech modeling based on speaker dictionaries
JP2017506767A
Method and apparatus for separating speech data from background data in audio communication
JP2017532601A