Speaker-specific voice amplification
By establishing a user-specific acoustic model, analyzing the audio stream, and using dynamic filter technology, the problem of background noise interference was solved, achieving clear amplification of the user's voice and effective suppression of background noise, ensuring clear transmission of the dialogue.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2021-11-17
- Publication Date
- 2026-04-21
AI Technical Summary
In public places or remote work environments, background noise interferes with the clarity of conversations, and existing technologies struggle to effectively isolate and amplify user speech while suppressing background noise.
By establishing a user-specific acoustic model, analyzing the audio stream using computing devices, selectively amplifying the user's speech and suppressing background noise, and employing dynamic bandpass or bandstop filters for frequency-selective processing in the time and frequency domains.
It effectively improves the clarity of user voice signals, reduces background noise interference, and ensures clear transmission of conversations in noisy environments.
Smart Images

Figure CN116648746B_ABST
Abstract
Description
Background Technology
[0001] This disclosure relates to digital signal processing, and more specifically to speaker-specific systems and methods for amplifying speech.
[0002] The development of the EDVAC system in 1948 is often cited as the beginning of the computer age. Since then, computer systems have evolved into extremely complex devices. Modern computer systems typically consist of a combination of complex hardware and software components, applications, operating systems, processors, buses, memory, input / output devices, and more. Advances in semiconductor processing and computer architecture have driven ever-increasing performance, and even more advanced computer software has evolved to take advantage of those capabilities, resulting in today's computer systems being far more powerful than those of just a few years ago.
[0003] One application of these new capabilities is mobile phones. Nowadays, people often make phone calls in public places (such as coffee shops or trains) or may be working remotely. In these environments, background noise from children, spouses, pets, buildings, and many other factors can interfere with the conversation. Summary of the Invention
[0004] According to embodiments of this disclosure, a method for amplifying a single speech item during an audio conversation using a computing device. Embodiments of the method may include receiving audio samples of speech from a user by the computing device, and generating an enhanced user-specific acoustic model of the user's speech based on the audio samples. The method may further include: receiving a live audio-visual stream comprising live speech of the user during an audio conversation, wherein the live audio-visual stream includes background noise; and selectively amplifying the live speech during the live audio-visual stream without amplifying the background noise using the user-specific acoustic model.
[0005] According to embodiments of this disclosure, a computer program product is used to selectively amplify user speech using a pre-trained acoustic model. Embodiments of the computer program product may include a computer-readable storage medium containing program instructions. The program instructions are executable by a processor to cause the processor to extract user speech data from existing speech samples, create a pre-trained acoustic model for the user from the speech data, analyze an audio stream from a teleconference, detect the presence of background noise in the audio stream, and apply the pre-trained acoustic model to the audio stream to selectively amplify the user's speech without amplifying the background noise.
[0006] According to embodiments of this disclosure, a computer system is provided for amplifying a single voice during an audio conversation. Embodiments of the system may include a processor configured to execute program instructions that, when executed on the processor, cause the processor to: receive audio samples of speech from a user; generate an enhanced user-specific acoustic model of the user's speech based on the audio samples; receive a live audio-visual stream including live speech from the user during an audio session, wherein the live audio-visual stream includes background noise; and selectively amplify the live speech during the live audio-visual stream without amplifying the background noise using the user-specific acoustic model.
[0007] The above description of the invention is not intended to depict every illustrated embodiment or implementation of this disclosure. Attached Figure Description
[0008] The accompanying drawings included in this application are incorporated in and form a part of this specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are merely illustrative of certain embodiments and do not limit the scope of the disclosure.
[0009] Figure 1 An embodiment of a data processing system (DPS) according to some embodiments is shown.
[0010] Figure 2 A cloud computing environment according to some embodiments is described.
[0011] Figure 3 An abstract model layer consistent with some implementations is described.
[0012] Figure 4 This is a system diagram for a computing environment based on some embodiments.
[0013] Figure 5 This is a flowchart of a noise reduction service in operation according to some embodiments.
[0014] Figure 6 This is a flowchart illustrating one method of training a machine learning model according to some embodiments.
[0015] Figure 7 This is a flowchart of a conference system in operation according to some embodiments.
[0016] While the invention is subject to various modifications and alternatives, its details have been illustrated by way of example in the accompanying drawings and will be described in detail. However, it should be understood that the invention is not limited to the specific embodiments described. Rather, the invention is intended to cover all modifications, equivalents, and substitutions that fall within its scope. Detailed Implementation
[0017] This disclosure relates to aspects of digital signal processing; more specifically, it relates to a speaker-specific system and method for amplifying speech. While this disclosure is not necessarily limited to such applications, its various aspects can be understood through the discussion of different examples using this context.
[0018] Some embodiments of this disclosure may include a system for building a speaker-specific acoustic model of a user from recordings. Some embodiments may then use the speaker-specific acoustic model to isolate the user's speech in future live audio and / or audiovisual streams, such that these embodiments may increase the volume of the portion of the signal containing only what the user has spoken and / or reduce the volume of noise from unwanted (e.g., background) sources. That is, some implementations may: (a) reduce / remove unwanted background noise; and (b) enhance the speaker's own speech. Background noise may include relatively static sounds generated by the conferencing system itself (e.g., static sounds recorded and / or generated by the user's microphone, artifacts generated during digital compression or microphone output transmission, etc.), relatively static sounds generated by environmental sources (e.g., local HVAC equipment, car engine and tire sounds, aircraft engine sounds, etc.), dynamic sounds (e.g., other people speaking nearby, sounds generated by nearby building equipment and processes, dog barking, ambulance sirens, etc.), and other sounds that are not the speaker's own speech.
[0019] In operation, users can begin by recording their speech in a quiet environment (i.e., with virtually no background noise), thus creating a clean target recording or set of recordings. Some embodiments may then analyze the clean recording to generate one or more parameters for a speech profile. Some embodiments may then use said one or more parameters to construct a speaker-specific acoustic model of the user's speech. This speaker-specific acoustic model may be specifically tuned to the vocal characteristics of said particular user (i.e., each user may have a different or even unique acoustic model associated with them). In this way, the speaker-specific acoustic model can be adapted to the characteristics of a particular user's speech and speaking patterns, such as their pitch and frequency range.
[0020] Once the acoustic model is built and / or trained, the same features used to create the model can be extracted and processed from future live audio and / or audiovisual streams. If no modification to those streams is required, the raw input can be sent to a network platform; otherwise, the disclosed system can improve the signal and send it to the platform (and / or other participants in a teleconference) in near real-time (e.g., less than 100 ms). This process may include measuring the level of background noise and attenuating it to an acceptable level (e.g., similar to those found during initial training). The model in these embodiments can view and analyze different speech patterns associated with speaker-specific characteristics of the user specified during initial training.
[0021] Some embodiments can detect when a user joins a conference call by analyzing the user's interactions with their associated electronic device (which may be a computer or a telephone with a microphone). Some embodiments can also determine when a user is currently presenting by comparing the user's name to the agenda or by detecting the content of their current speech, and can selectively amplify the user's voice while suppressing background noise and the voices of other participants.
[0022] In response to this detection, some embodiments may apply dynamic bandpass or bandstop filters to selectively amplify certain frequencies and / or selectively suppress certain frequencies during transmission and / or retransmission of the audio stream. In some embodiments, such suppression and / or amplification may occur in the time domain and / or frequency domain. In this way, the user's voice signal can be effectively boosted while being transmitted and / or retransmitted, and any other noise (other speech, non-spoken sounds) can be suppressed and / or not transmitted / retransmitted.
[0023] Furthermore, some embodiments can detect the presence of ambient noise (e.g., car horns) or other people speaking simultaneously in the background (e.g., children). The system can then identify the detected background noise and / or the detected speech of others as unwanted signals and remove them from the transmitted / retransmitted stream. More specifically, some embodiments can automatically activate and deactivate a pre-trained acoustic model associated with the presenter to enhance the presenter's speech while the presenter is speaking. Additionally, some embodiments automatically and selectively modify the pre-trained acoustic model to compensate for detected noise and / or speech. This allows other meeting participants to hear the presenter more clearly without interference. Some embodiments can also be configured with a configurable threshold for background noise detected as unwanted—such as a threshold based on duration, loudness, or the amount of deviation from the pitch range of a pre-trained user. In this way, users and applications requiring acoustic fidelity can reduce the amount and / or degree of filtering performed by some embodiments.
[0024] In some embodiments, the acoustic model may be based on a user's voiceprint (profile). These features may include, but are not limited to: pitch variations and perturbations (e.g., jitter), periodic measurements (e.g., harmonic noise ratio), linear predictive decoding coefficients (LPC), spectral shape measurements, speech onset time, Mel-frequency cepstral coefficients (MFCC), i-vectors, etc. The model may be trained on a user's voiceprint, which is speaker-specific and approved by the user at the outset. In some embodiments, the acoustic model may include unsupervised speech alignment algorithms, hidden Markov models, unwanted noise removal algorithms, phoneme token DNNs, etc. Some of these features (e.g., formants: F1 and F2) may primarily be used to characterize (e.g., to obtain the voiceprint) the uniqueness of the user's speech, while others (e.g., harmonic noise ratio) may primarily be used to characterize background noise.
[0025] In some embodiments, acoustic features that can be analyzed to construct a speaker-specific speech profile include, but are not limited to, one or more features selected from the group consisting of: fundamental frequency, spectral envelope, pitch characteristics (e.g., average, maximum, minimum, etc.), speech onset time (VOT) for different consonants, F1 and F2 for characterizing vowel articulation, and vowel duration. These acoustic features can be used to distinguish a user's speech from other human speech (particularly the consonant and vowel features mentioned above) and from non-speech (particularly the higher-level pitch and frequency features mentioned above).
[0026] Some embodiments may also create supplemental acoustic models for users when their speech differs to some extent, such as when they have a stuffy nose or a sore throat. Similarly, some embodiments may create custom, speaker-specific acoustic models for each language spoken by the user, since different languages may have different acoustic profiles, even for the same speaker.
[0027] As a first illustrative example, suppose user (“A”) is working remotely. As a result, user A is largely at home with her husband, children, and pets, who may create background noise when she attempts to make a business call. Specifically, her dog may frequently be heard barking in the background when user A calls into a meeting. In this example, the acoustic model has been previously trained on user A's voice and has a specific acoustic profile tailored to her voice and speech patterns. Thus, the system in this example can detect user A's speech and then selectively amplify only her voice while reducing any noise from her dog so that it is not broadcast to other meeting participants.
[0028] As a second illustrative example, User (“B”) arrives late for a meeting. User B takes the conference call in her car. As a result, there is considerable background noise from the road, the car engine, and possibly even a baby crying in the back seat. Traditionally, User B could partially mitigate this noise by muting herself, but then would not be able to fully engage in the meeting without all the external noise distracting her colleagues. However, using some embodiments of the invention, a speaker-specific acoustic model may have been pre-trained based on User B’s voice. Now, when she calls in from her car, some embodiments can identify and amplify different acoustic properties specific to User B’s voice while reducing acoustic properties inconsistent with her voice (e.g., from the road, her car, etc.). Thus, User B is able to speak in her meeting without transmitting loud noise from her environment.
[0029] As a third illustrative example, User (“C”) has a moderate cold and is therefore working remotely, away from her colleagues. While User C is in a meeting, she occasionally coughs and sneezes. Using some implementations, the acoustic model may have been pre-trained based on User C’s normal speech (e.g., without coughing and sneezing), thus allowing for selective amplification of her speech signal while selectively removing coughs and sneezes as unwanted background noise. Furthermore, knowing that she may cough or sneeze during the meeting, in some embodiments, User C may be configured with a threshold considered as unwanted background noise based on factors such as duration, loudness, or the amount of deviation from the user’s pre-trained pitch range to ensure that those coughs are filtered out.
[0030] As a fourth illustrative example, a user (“D”) is ill with a sore throat and is working remotely to prevent the infection from spreading to his colleagues. However, as a result of the infection, user D’s speech currently sounds different from his normal speech. Some embodiments of this disclosure may allow user D to create (or adjust) an additional, customized acoustic model of his / her speech (e.g., his / her “normal” speech and his / her current, for example, “abnormal” speech due to illness). Thus, user D may choose to train another version of the system for his / her sore throat speech, so that when he / she calls a meeting, the system is able to recognize and amplify his / her current speech, rather than suppressing it as background noise, even if it has some acoustic properties that differ from his / her normal speech. Furthermore, some embodiments may selectively modify those acoustic properties so that when presented to other meeting participants, user D’s current speech sounds more like his / her normal speech.
[0031] Accordingly, a feature and advantage of some embodiments is that they may not require users to have and use specific hardware (such as directional microphones) to amplify their speech and / or suppress noise. Therefore, some implementations can be integrated as plug-ins into existing video conferencing systems and telephone microphone processing software. Another feature and advantage of some embodiments is that they can use pre-trained acoustic models to selectively amplify user speech. In this way, some embodiments can continue to broadcast sound with a large dynamic range, including desired but unintended noise. Another feature and advantage of some implementations is that they can learn the user's voice characteristics and selectively amplify the signal based on these characteristics. In this way, some embodiments can allow noise reduction, allow more than one speaker to use the microphone at a time, and / or allow more than one speaker to speak at a time, even if the user is moving around within their local physical environment.
[0032] Data processing system
[0033] Figure 1 An embodiment of a data processing system (DPS) 100a according to some embodiments is shown. The DPS 100a in this embodiment can be implemented as a personal computer; a server computer; a portable computer, such as a laptop or notebook computer, PDA (personal digital assistant), tablet computer, or smartphone; or a processor embedded in a larger device (e.g., a car, airplane, teleconferencing system, appliance, smart device, or any other suitable type of electronic device). Furthermore, there may be processors different from or other than those described above. Figure 1 Components other than those shown, and the number, type, and configuration of such components can vary. Furthermore, Figure 1 Only representative major components of the DPS100a are depicted, and individual components can have more than Figure 1 The greater complexity is represented in the text.
[0034] Figure 1 The data processing system 100a includes multiple central processing units 110a-110d (collectively referred to herein as processor 110 or CPU 110) connected via a system bus 122 to a memory 112, a mass storage interface 114, a terminal / display interface 116, a network interface 118, and an input / output (“I / O”) interface 120. In this embodiment, the mass storage interface 114 connects the system bus 122 to one or more mass storage devices, such as a direct access storage device 140, a universal serial bus (“USB”) storage device 141, or a read / write optical disc drive 142. The network interface 118 allows DPS 100a to communicate with other DPS 100b via a communication medium 106. The memory 112 also contains an operating system 124, multiple application programs 126, and program data 128.
[0035] Figure 1 The data processing system 100a embodiment is a general-purpose computing device. Therefore, the processor 110 can be any device capable of executing program instructions stored in memory 112 and can itself be constructed from one or more microprocessors and / or integrated circuits. In this embodiment, the DPS 100a includes multiple processors and / or processing cores, as in a typical larger, more powerful computer system; however, in other embodiments, the data processing system 100a may include a single-processor system and / or a single processor designed to simulate a multiprocessor system. Furthermore, the processor 110 can be implemented using several heterogeneous data processing systems 100a, where a primary processor and secondary processors reside on a single chip. As another illustrative example, the processor 110 can be a symmetric multiprocessor system containing multiple processors of the same type.
[0036] When the data processing system 100a starts, the associated processor 110 initially executes program instructions that constitute the operating system 124, which manages the physical and logical resources of the DPS 100a. These resources include memory 112, mass storage interface 114, terminal / display interface 116, network interface 118, and system bus 122. Like the processor 110, some embodiments of the DPS 100a may utilize multiple system interfaces 114, 116, 118, 120, and bus 122, each of which may further include its own separate, fully programmable microprocessor.
[0037] Instructions for operating systems, applications, and / or programs (collectively, "program code," "computer-usable program code," or "computer-readable program code") may initially reside in mass storage devices 140, 141, and 142, which communicate with processor 110 via system bus 122. The program code in different embodiments may be implemented on different physical or tangible computer-readable media, such as system memory 112 or mass storage devices 140, 141, and 142. Figure 1 In the illustrative example, instructions are stored in functional form on direct access memory 140 as permanent memory. These instructions are then loaded into memory 112 for execution by processor 110. However, program code may also be located in functional form on selectively removable computer-readable media and may be loaded into or transferred to DPS 100a for execution by processor 110.
[0038] System bus 122 can be any means that facilitates communication between processor 110, memory 112, and interfaces 114, 116, 118, 120. Furthermore, although system bus 122 in this embodiment is a relatively simple, single-bus structure providing a direct communication path between system buses 122, other bus structures, including but not limited to point-to-point links in hierarchical, star, or network configurations, multiple hierarchical buses, parallel and redundant paths, etc., are also possible according to this disclosure.
[0039] Memory 112 and mass storage devices 140, 141, and 142 work together to store operating system 124, application program 126, and program data 128. In this embodiment, memory 112 is a random access semiconductor device capable of storing data and programs. Although Figure 1 Conceptually, the device is described as a single monolithic entity, but in some embodiments, memory 112 can be a more complex arrangement, such as a hierarchy of caches and other memory devices. For example, memory 112 may reside in multi-level caches, and these caches may be further functionally partitioned such that one cache holds instructions while another cache holds non-instruction data used by one or more processors. Memory 112 may be further distributed and associated with different processors 110 or sets of processors 110, as known in any of the various so-called Non-Unified Memory Access (NUMA) computer architectures. Furthermore, some embodiments may utilize virtual addressing mechanisms that allow DPS 100a to behave as if it were accessing a large single storage entity rather than multiple smaller storage entities, such as memory 112 and mass storage devices 140, 141, 142.
[0040] Although the operating system 124, application program 126, and program data 128 are shown as being contained within memory 112, in some embodiments, some or all of them may be physically located on different computer systems and may be remotely accessed, for example, via communication medium 106. Thus, although the operating system 124, application program 126, and program data 128 are illustrated as being contained within memory 112, these elements are not necessarily all completely contained in the same physical device at the same time, and may even reside in the virtual memory of other DPSs (e.g., DPS 100b).
[0041] System interfaces 114, 116, 118, and 120 support communication with various storage and I / O devices. Mass storage interface 114 supports the attachment of one or more mass storage devices 140, 141, and 142, which are typically spinning disk drive storage devices, solid-state storage devices (SSDs) that use integrated circuit components as memory to persistently store data (typically using flash memory, or a combination of both). However, mass storage devices 140, 141, and 142 may also include other devices, including disk drive arrays (often referred to as RAID arrays) and / or archival storage media configured to appear as a single large storage device to the host, such as hard disk drives, magnetic tapes (e.g., mini-DV), writable compact discs (e.g., CD-R and CD-RW), digital universal discs (e.g., DVD, DVD-R, DVD+R, DVD+RW, DVD-RAM), holographic storage systems, Blue LaserDiscs, IBM Millipede devices, etc.
[0042] Terminal / display interface 116 is used to directly connect one or more display units 180 (which may include monitors, etc.) to data processing system 100a. These display units 180 may be non-intelligent (i.e., dumb) terminals, such as LED monitors, or they may be fully programmable workstations that allow IT administrators and customers to communicate with DPS 100a. However, note that while display interface 116 is provided to support communication with one or more display units 180, data processing system 100a does not necessarily require display units 180, as all necessary interactions with customers and other processes can occur via network interface 118.
[0043] Communication medium 106 can be any suitable network or combination of networks and can support any suitable protocol suitable for transmitting data and / or code to / from multiple DPS100a, 100b. Therefore, network interface 118 can be any device facilitating such communication, regardless of whether the network connection uses current analog and / or digital technologies or is made via some future networking mechanism. Suitable communication medium 106 includes, but is not limited to, networks using “InfiniBand” or IEEE (Institute of Electrical and Electronics Engineers) 802.3x “Ethernet”, cellular transport networks; wireless networks implementing one of the IEEE 802.11x, IEEE 802.16, General Packet Radio Service (“GPRS”), FRS (Family Radio Service), or Bluetooth specifications; networks implementing one or more of the Ultra Wideband (“UWB”) technologies (such as those described in FCC 02-48), etc. Those skilled in the art will understand that many different network and transport protocols can be used to implement communication medium 106. The Transmission Control Protocol / Internet Protocol (“TCP / IP”) suite contains suitable network and transport protocols.
[0044] cloud computing
[0045] Figure 2 A cloud environment comprising one or more DPS100a, 100b is illustrated according to some embodiments. It should be understood that while this disclosure includes a detailed description of cloud computing, implementations of the teachings cited herein are not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0046] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing power, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0047] The characteristics are as follows:
[0048] • On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring human interaction with the service provider.
[0049] • Wide network access: Capabilities are available via the network and through standard mechanisms that facilitate access to heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0050] • Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. There is a sense of location independence because consumers typically do not have control or knowledge of the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0051] • Rapid and flexible: Capable of providing capacity quickly and flexibly, automatically scaling down and up rapidly in some situations to scale up quickly. For consumers, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.
[0052] • Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active customer accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0053] The service model is as follows:
[0054] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. The applications can be accessed from different client devices via thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities; possible exceptions include limited consumer-specific application configuration settings.
[0055] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created or acquired by the consumer using programming languages and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environments.
[0056] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, deployed applications, and potentially limited control over selected networking components (e.g., host firewalls).
[0057] The deployment model is as follows:
[0058] • Private cloud: A cloud infrastructure that operates solely for an organization. It can be managed by the organization or a third party and can exist on-site or off-site.
[0059] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0060] • Public cloud: Cloud infrastructure that is available to the general public or large industry groups and is owned by an organization that sells cloud services. • Hybrid cloud: Cloud infrastructure that is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported (e.g., cloud bursting for load balancing between clouds).
[0061] Cloud computing environments are service-oriented, focusing on statefulness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure that includes a network of interconnected nodes.
[0062] See now Figure 2 The diagram illustrates an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 to which local computing devices used by cloud consumers can communicate. These local computing devices include, for example, personal digital assistants (PDAs) or cellular phones 54A, desktop computers 54B, laptop computers 54C, and / or automotive computer systems 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platforms, and / or software as services that cloud consumers do not need to maintain on their local computing devices. It should be understood that... Figure 2 The types of computing devices 54A-N shown are intended to be illustrative only, and computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network-addressable connectivity (e.g., using a web browser).
[0063] See now Figure 3 This demonstrates a cloud computing environment of 50 ( Figure 2 This provides a set of functional abstractions. It should be understood beforehand. Figure 3 The components, layers, and functions shown are intended to be illustrative only, and embodiments of the invention are not limited thereto. As described, the following layers and corresponding functions are provided:
[0064] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: a mainframe 61; a RISC (Reduced Instruction Set Computer) based server 62; a server 63; a blade server 64; a storage device 65; and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0065] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71; virtual storage 72; virtual network 73, including virtual private network; virtual application and operating system 74; and virtual client 75.
[0066] In one example, management layer 80 may provide the following functionalities: Resource Provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. Customer Portal 83 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 84 provides cloud resource allocation and management to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 85 provides pre-scheduling and procurement of cloud resources based on anticipated future needs according to the SLA.
[0067] Workload layer 90 provides examples of functionalities that can leverage a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytics and processing 94; transaction processing 95; and noise reduction services 96.
[0068] Acoustic Platform
[0069] Figure 4 This is a system diagram of a computing environment 400 according to some embodiments. The computing environment 400 may include a conferencing computer platform 402 connected to multiple user devices 403 via a network 406. The conferencing computer platform 402 may further include a conferencing module 480 and a noise reduction service 496. The noise reduction service 496 may include a trained machine learning model 498 that generates multiple custom acoustic profiles 499. In some embodiments, the conferencing module 480 may include a database 482 containing the multiple custom acoustic profiles.
[0070] The conference computer platform 402 may be a standalone computing device, management server, network server, mobile computing device, or any other electronic device or computing system capable of receiving, sending, and processing data. In some embodiments, the conference computer platform 402 may be part of a cloud computing environment 50 and represent a pool of computing resources within that environment. The conference computer platform 402 may include internal and external hardware components, as referenced... Figure 1 The DPS 100 is described and depicted in more detail.
[0071] User equipment 403 may represent one or more programmable electronic devices or combinations of programmable electronic devices capable of executing machine-readable program instructions and communicating with other computing devices (not shown) within computing environment 400 via network 406. Suitable user equipment 403 includes, but is not limited to: desktop computers, laptop computers, tablet computers, smartphones, smartwatches, and voice-enabled handsets.
[0072] In some embodiments, user equipment 403 may include one or more Voice over Internet Protocol (VoIP) compatible devices (i.e., VoIP, IP telephony, broadband telephony, and broadband telephony services). VoIP generally refers to a set of methods and technologies for transmitting voice communications and multimedia sessions over an Internet Protocol (IP) network (such as the Internet). VoIP can be integrated into smartphones, personal computers, and any general-purpose computing device (such as user equipment 403) capable of communicating with network 406. In some embodiments, user equipment 403 may include one or more speakerphones suitable for use in audio or video conferencing environments. A speakerphone generally refers to an audio device that includes at least a speaker, a microphone, and one or more microprocessors.
[0073] In some embodiments, user device 403 may include a user interface ( Figure 1 (Not shown independently above). This user interface provides an interface between user device 403 and the conference computer platform. In some embodiments, the user interface may be a graphical user interface (GUI) or a web user interface (WUI), which may display text, documents, web browser windows, user options, application interfaces, and operation instructions, and includes information presented to the user by the program (such as graphics, text, and sound) and control sequences adopted by the user to control the program. In another embodiment, the user interface may include mobile application software that provides an interface between each user device 403 and the conference computer platform 402. Mobile application software or "applications" are a class of computer programs that typically run on smartphones, tablets, smartwatches, and other mobile devices.
[0074] Figure 4Network 406 may include, for example, a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN) (such as the Internet), or a combination thereof, and may include wired, wireless, or fiber optic connections. Network 406 may utilize a combination of connections and protocols supporting communication between the conference computer platform 402 and user equipment 403, such as receiving and transmitting data, voice, and / or video signals, including multimedia signals containing voice, data, and video information.
[0075] In some embodiments, the conference computer platform may provide noise reduction services 496. Figure 5 This is a flowchart of noise reduction service 496 in operation according to some embodiments. At operation 505, the user can create an audio sample (e.g., record). In some embodiments, the user may be asked to read a predetermined script using a high-quality microphone in a quiet environment. The script can be selected to contain a large number of unique audio features, such as the most common and / or all phonemes of a given language. The audio sample can then be used as part of training phase 508.
[0076] During training phase 508, noise reduction service 496 may receive audio samples at operation 510. In response, noise reduction service 496 may extract audio features from the audio samples and then feed those features into the trained machine learning model 498 at operation 515. Next, at operation 520, machine learning model 498 may generate a first customized acoustic profile 499a for the user based on the features. This acoustic profile 499a may be optimized to selectively identify and / or isolate the full dynamic range of the user's speech from the recording stream. At operations 525-530, the user may optionally be given the opportunity to hear the recording processed by the customized acoustic profile 499a (created at operation 505) and then approve or reject the model. If the user rejects the customized acoustic profile 499a (530: No), noise reduction service 496 may return to operations 505, 510 to collect and process new samples. If the user accepts the customized acoustic profile 499a (530: Yes), the system can output the model for future live conferences.
[0077] Later, the user and / or noise reduction service 496 can initiate an additional training phase 550. This additional training can allow the creation of additional, supplementary acoustic profiles 499b-499n for the user. These supplementary profiles can be created in response to the user's current physical condition (e.g., a cold or sore throat), or the acoustic profile can be optimized for an alternative language. For example, a user can have a profile 499a for speaking English with their normal voice, a profile 499b for speaking English with their coughing voice, a profile 499c for speaking Spanish with their normal voice, and a profile 499d for a specific day when they are sick. In this illustrative example, the user has four separate acoustic profiles 499 stored in the system and can select one at the start of a conference call.
[0078] As in the initial training phase 508, the adjustment phase begins by receiving new audio samples submitted by the user at 555. In response, the noise reduction service 496 can extract audio features from the new audio samples and then feed those features into the trained machine learning model 498 at operation 560. Next, at operation 565, the speech improvement module 497 of the machine learning model 498 can generate supplementary acoustic profiles 499b-499n. In some embodiments, this may include receiving and adjusting the original acoustic profile 499. At operations 570-575, the user is optionally given the opportunity to hear the recording processed by the supplementary acoustic profile 499a (created in operation 505) and then approve or reject the updated model. If the user rejects the supplementary acoustic profiles 499b-499n (575: No), the noise reduction service 496 can return to operations 505, 555 to collect and process new samples. If the user accepts the supplementary acoustic profiles 499b-499n (575: Yes), the system can output the model for future live conferences. In some implementations, this may include allowing the user to select which acoustic profile 499a-499n to use in the conference.
[0079] Model training
[0080] In some embodiments, the machine learning model 498 can be any software system that recognizes patterns. In some embodiments, the machine learning can include multiple artificial neurons interconnected by connection points called synapses. Each synapse can encode the strength of the connection between the output of one neuron and the input of another neuron. The output of each neuron can then be determined by aggregated inputs received from other neurons connected to it, and thus by the outputs of these “upstream” connected neurons and the connection strength determined by the synaptic weights.
[0081] ML models can be trained to solve specific problems (e.g., generating configuration settings for a customized, user-specific acoustic model) by adjusting the weights of synapses, thereby causing inputs of a specific class to produce the desired output. This weight adjustment procedure in these implementations is called "learning." Ideally, these adjustments result in a pattern of synaptic weights that converges towards the optimal solution for a given problem based on some cost function during the learning process.
[0082] In some embodiments, artificial neurons can be organized into layers. The layer that receives external data is the input layer. The layer that generates the final result is the output layer. Some embodiments include hidden layers between the input and output layers, and there are typically hundreds of such hidden layers.
[0083] Figure 6 This is a flowchart illustrating a method 600 for training a machine learning model 498 according to some embodiments. The system manager can begin at operation 610 by loading training vectors. These vectors may include audio recordings from multiple different users filmed in a quiet room and reading specially prepared scripts.
[0084] In operation 612, the system manager can select the desired output (e.g., optimal settings for acoustic model 499). In operation 614, training data can be prepared to reduce bias sources, which typically include deduplication, normalization, and sequential randomization. In operation 616, the initial weights of the gates of the machine learning model can be randomized. In operation 618, the ML model can be used to predict the output using the input data vector set and compare that prediction with the labeled data. The gate weights are then updated at operation 620 using the error (e.g., the difference between the predicted value and the labeled data). This process can be repeated with each iteration of weight updates until the training data is exhausted or the ML model reaches an acceptable level of accuracy and / or precision. In operation 622, the resulting model can optionally be compared with previously unevaluated data to validate and test its performance. In operation 624, the resulting model can be loaded onto the noise reduction service 496 in the cloud computing environment 50 and used for analyzing user records.
[0085] Conference System
[0086] Figure 7This is a flowchart of a conference system 700 in operation according to some embodiments. At operation 705, multiple participants in a telephone conference can register and / or log in to the conference system 700 using, for example, a username and password. At operation 710, the conference system 700 queries database 482 for the current acoustic model 499 associated with each participant. If one or more of the participants do not have a model (711: No), then at operation 712, the system may prompt one or more participants to record the raw audio data of their speech to begin building a speaker-specific audio model. Alternatively, the system may also, at operation 712, give participants the option to use a generic model (i.e., a model created to isolate a wide variety of voices and languages) for the conference, which may be desirable if participants do not have the time and / or equipment to create a custom model. If a participant has already created multiple custom models, then at operation 713, the system may prompt the participant to select that model for the current call.
[0087] One of the participants can then begin speaking. Their user device 403 can record those sounds at operation 715, convert the recording into a raw audio stream at operation 720, and transmit the raw audio stream to the conference module 480 at operation 725. In response, at operation 730, the conference system 700 applies a customized acoustic model (identified at operation 705) for that particular speaker to the received audio stream to generate an optimized audio stream (e.g., an audio stream in which the user's voice is amplified and / or any background noise is suppressed). Then, at operation 735, the conference system 700 can transmit the optimized audio stream to the other participants in the conference call. These embodiments may be desirable because they can be used with any user device 403 (such as a "plain telephone system" mobile phone).
[0088] Alternatively, in some embodiments, some or all of the user equipment 403 may apply a custom acoustic model to the original audio stream at operation 725 (e.g., via a processor in user equipment 403), and then transmit the optimized audio stream (instead of the original audio stream) to the conferencing system 700. The conferencing system 700 can then proceed directly to operation 735 and retransmit the optimized audio stream to other participants. These embodiments may be desirable because they can be used with any conferencing system 700.
[0089] Operations 715-735 can be repeated during the duration of a teleconference, by the conference system 700, by the user equipment 403, and / or a combination of both, whenever a participant speaks.
[0090] Computer program products
[0091] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.
[0092] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0093] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device.
[0094] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this invention.
[0095] This document describes various aspects of the invention with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0096] These computer-readable program instructions may be provided to a computer processor or other programmable data processing apparatus to generate a machine, such that these instructions, which execute via the computer processor or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, thereby comprising an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0097] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable apparatus, or other device to perform the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, two consecutively shown blocks may actually be completed as a single step, executed simultaneously, substantially simultaneously, or with partial or complete temporal overlap, or the blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0099] Overview
[0100] Any specific program nomenclature used in this specification is for convenience only, and therefore the invention should not be limited to use only in any specific application identified and / or implied by such nomenclature. Thus, for example, a routine executed to implement embodiments of the invention (whether implemented as part of an operating system or as a particular application, component, program, module, object, or sequence of instructions) may be referred to as a "program," "application," "server," or other meaningful term. In fact, other alternative hardware and / or software environments may be used without departing from the scope of the invention.
[0101] Therefore, it is intended that the embodiments described herein be considered illustrative rather than restrictive in all respects, and that reference be made to the appended claims to define the scope of the invention.
Claims
1. A method for amplifying a user's voice during an audio session using a computing device, the method comprising: Generate multiple customized acoustic profiles for the user, wherein the generation includes: The computing device receives an audio sample of a speech submitted by a user for a first customized acoustic profile among the plurality of customized acoustic profiles; and The computing device generates a user-specific acoustic model and a supplementary acoustic model based on the audio samples to enhance the user's speech, wherein the user-specific acoustic model is adapted to the user's normal physical condition, and wherein the supplementary acoustic model is adapted to the user's speech in cases of illness. In response to generating the plurality of customized acoustic profiles, a live audio-visual stream including the user’s live speech during an audio session is received, wherein the live audio-visual stream includes background noise. In response to receiving the live audio-visual stream, the user is prompted to select an acoustic profile for the live audio-visual stream; The computing device receives the selection of a first customized acoustic profile by the user in response to the prompt; and In response to the selection of a first custom acoustic profile, the computing device applies the supplementary acoustic model to selectively amplify the live speech during the live audiovisual stream without amplifying the background noise, wherein the computing device is configured to modify the acoustic characteristics of the user's disease speech when the supplementary acoustic model is selected, such that the user's disease speech sounds like normal speech, and wherein the amplification is based on the selected first custom acoustic profile.
2. The method according to claim 1, further comprising: The computing device uses the user-specific sound model to selectively suppress the background noise during the live audiovisual stream.
3. The method according to claim 1, wherein, The user-specific acoustic model is a plugin for the teleconferencing software.
4. The method of claim 3, further comprising generating a plurality of user-specific acoustic models, wherein, Each user-specific acoustic model is used for one of the multiple users in the teleconferencing software.
5. The method according to claim 1, further comprising: The audio samples were collected from the user in an environment with virtually no background noise. as well as The audio samples are used to generate the user-specific acoustic model.
6. The method of claim 5, further comprising using a trained machine learning model to generate the user-specific acoustic model.
7. A computer program product for selectively amplifying user speech using a pre-trained acoustic model, the computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a processor to cause the processor to: Generate multiple customized acoustic profiles for users, among which, The generation includes: Extract user voice data from existing audio samples submitted by the user for a first customized acoustic profile among the plurality of customized acoustic profiles; and A pre-trained acoustic model and a supplementary acoustic model are created for the user from the speech data, wherein the pre-trained acoustic model is adapted to the user's normal physical condition, and wherein the supplementary acoustic model is adapted to the user's speech in cases of illness. In response to the generation of the multiple customized acoustic profiles, a conference call is initiated among multiple users; The user is prompted to select an acoustic profile; Receive the user's selection of a first customized acoustic profile in response to the prompt; In response to receiving the selection, the audio stream from the conference call is analyzed; Based on the analysis, the presence of background noise in the audio stream was detected; and In response to the detection, the supplementary acoustic model is applied to the audio stream to selectively amplify the user's speech without amplifying the background noise, wherein, when the supplementary acoustic model is selected, the acoustic characteristics of the user's disease speech are modified so that the user's disease speech sounds like normal speech, and wherein the amplification is based on a selected first custom acoustic profile.
8. The computer program product of claim 7, further comprising program instructions for selectively suppressing the background noise from the audio stream using the pre-trained acoustic model.
9. The computer program product according to claim 7, wherein, The pre-trained acoustic model is a plugin for teleconferencing software.
10. The computer program product of claim 9, further comprising program instructions for generating a plurality of user-specific acoustic models, one of each of the plurality of users of the teleconference software.
11. The computer program product of claim 7, further comprising program instructions for performing the following operations: The audio samples were collected from the user in an environment with virtually no background noise; and The audio samples are used to generate user-specific acoustic models.
12. The computer program product of claim 7, further comprising program instructions for generating a plurality of user-specific acoustic models for the user, wherein a first user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speak in a first language, and wherein a second user-specific acoustic model of the plurality of user-specific acoustic models is adapted to speak in a second language.
13. The computer program product according to claim 7, further comprising: The computing device receives audio samples of the user's speech from an environment with virtually no background noise. as well as The pre-trained acoustic model is generated from the audio samples using a trained machine learning model.
14. A system for amplifying a user's voice during an audio session, wherein the system includes a processor configured to execute program instructions, the program instructions being executed on the processor, causing the processor to: Generate multiple customized acoustic profiles for users, among which, The generation includes: Receive audio samples of speech submitted by the user for a first customized acoustic profile among the plurality of customized acoustic profiles; and Based on the audio samples, a user-specific acoustic model and a supplementary acoustic model are generated to enhance the user's speech, wherein the user-specific acoustic model is adapted to the user's normal physical condition, and wherein the supplementary acoustic model is adapted to the user's speech in cases of illness. In response to generating the plurality of customized acoustic profiles, a live audio-visual stream including the user’s live speech during an audio session is received, wherein the live audio-visual stream includes background noise. In response to receiving the live audio-visual stream, the user is prompted to select an acoustic profile for the live audio-visual stream; Receives a selection of a first customized acoustic profile by the user in response to the prompt; and In response to the selection of a first custom acoustic profile, the supplementary acoustic model is used to selectively amplify the live speech during the live audiovisual stream without amplifying the background noise, wherein, when the supplementary acoustic model is selected, the acoustic characteristics of the user's disease speech are modified so that the user's disease speech sounds like normal speech, and wherein the amplification is based on the selected first custom acoustic profile.
15. The system of claim 14, further comprising program instructions for selectively suppressing the background noise during the live audiovisual stream using the user-specific acoustic model.
16. The system according to claim 14, wherein, The user-specific acoustic model is a plugin for the teleconferencing software.
17. The system of claim 16, further comprising program instructions for generating a plurality of user-specific acoustic models, wherein, Each user-specific acoustic model is used for one of the multiple users in the teleconferencing software.
18. The system of claim 14, further comprising program instructions for the following operations: Collect audio samples from the user in an environment with virtually no background noise; and The audio samples are used to generate the user-specific acoustic model.
19. The system of claim 18, further comprising program instructions for generating the user-specific acoustic model using a trained machine learning model.
Citation Information
Patent Citations
System and method for managing a mute button setting for a conference call
US20200110572A1