Full deep learning minimum variance distortionless response beamformer for speech separation and enhancement
By using a GRU network to estimate the speech and noise covariance matrix and combining it with a complex ratio filter and a frame-level weighting module, the problems of high residual noise and unstable matrix inversion in the prior art are solved, the effect of speech separation and enhancement is improved, and the performance of automatic speech recognition is enhanced.
Patent Information
- Application Number
- CN202180033782.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2021-06-23
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing deep learning-based speech enhancement and speech separation methods suffer from high residual noise levels in low signal-to-noise ratio or overlapping speech conditions, and traditional matrix inversion and principal component analysis are unstable, affecting the performance of automatic speech recognition systems.
A gated recurrent unit (GRU) network is used to estimate the covariance matrix of the target speech and noise. The predicted target waveform is generated by the minimum variance distortionless response function. Speech separation is performed using a complex ratio filter and a frame-level weighting module to achieve adaptive adjustment of frame-by-frame weights.
It reduces residual noise and improves the performance of automatic speech recognition systems, especially in low signal-to-noise ratio or overlapping speech conditions, reducing speech distortion and improving word error rate (WER) performance.
Smart Images

Figure CN115516554B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 038,498 (filed September 30, 2020), the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of data processing, and more particularly to speech recognition. Background Technology
[0004] Deep learning-based speech enhancement and speech separation methods have attracted widespread research attention. Mask-based minimum variance distortionless response (MVDR) beamformers can be used to reduce speech distortion, which is beneficial for automatic speech recognition. Multi-tap MVDR based on complex-valued masks can further improve the performance of automatic speech recognition in mask-based beamforming architectures. Summary of the Invention
[0005] The embodiments relate to methods, systems, and computer-readable media for speech recognition. According to one aspect, a method for speech recognition is provided. The method may include receiving audio data corresponding to one or more speakers; estimating a covariance matrix of target speech and noise associated with the received audio data using a gated recurrent unit-based network; and generating a predicted target waveform corresponding to a target speaker among the one or more speakers based on the estimated covariance matrix using a minimum variance distortion-free response function.
[0006] According to another aspect, a computer system for speech recognition is provided. The computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions. The program instructions are stored on at least one of the one or more storage devices and are executed by at least one of the one or more processors via at least one of the one or more memories. Thus, the computer system is capable of performing a method. The method may include receiving audio data corresponding to one or more speakers; estimating a covariance matrix of target speech and noise associated with the received audio data using a gated recurrent unit network; and generating a predicted target waveform corresponding to a target speaker among the one or more speakers using a minimum variance distortion-free response function based on the estimated covariance matrix.
[0007] According to another aspect, a computer-readable medium for speech recognition is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored on at least one of the one or more tangible storage devices. The program instructions are executable by a processor. The program instructions are executable by the processor to implement a method, which accordingly includes receiving audio data corresponding to one or more speakers. A network based on gated recurrent units is used to estimate a covariance matrix of target speech and noise associated with the received audio data. Based on the estimated covariance matrix, a predicted target waveform corresponding to a target speaker among the one or more speakers is generated using a minimum variance distortion-free response function. Attached Figure Description
[0008] These and other objects, features, and advantages will become apparent from the following detailed description of illustrative embodiments, which is taken in conjunction with the accompanying drawings. The various features in the drawings are not to scale, as they are illustrated to facilitate a clear understanding by those skilled in the art in conjunction with the detailed description. In the drawings:
[0009] Figure 1 A networked computer environment according to at least one embodiment is shown;
[0010] Figure 2 This is an exemplary speech recognition system according to at least one embodiment;
[0011] Figure 3 This is an operation flowchart of the steps performed by a procedure for separating the speech of a target speaker according to at least one embodiment;
[0012] Figure 4 According to at least one embodiment Figure 1 The diagram shows the internal and external components of the computer and server.
[0013] Figure 5 It includes, according to at least one embodiment Figure 1 The diagram illustrates a cloud computing environment for the computer system shown; and
[0014] Figure 6 According to at least one embodiment Figure 5 A block diagram illustrating the functional layers of an illustrative cloud computing environment. Detailed Implementation
[0015] This document discloses detailed embodiments of the claimed structures and methods. However, it is to be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which can be implemented in various forms. These structures and methods may be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope to those skilled in the art. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0016] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0017] The embodiments generally relate to the field of data processing, and more specifically to speech recognition. Furthermore, the exemplary embodiments described below provide a system, method, and computer program for separating the speech of a target speaker using a fully neural network approach. Therefore, some embodiments have the ability to improve the computational domain by allowing computer-enhanced speech enhancement, speech separation, and dereverberation tasks. Moreover, the disclosed methods, systems, and computer-readable media can be used to improve the performance of automatic speech recognition in fields such as hearing aids and communications.
[0018] As mentioned earlier, deep learning-based speech enhancement and speech separation methods have received extensive research attention. Mask-based minimum variance distortionless response (MVDR) beamformers can be used to reduce speech distortion, which is beneficial for automatic speech recognition. Multi-tap MVDR based on complex-valued masks can further improve the performance of automatic speech recognition in mask-based beamforming architectures. However, residual noise levels remain high, especially in cases of low signal-to-noise ratio or overlapping speech. Furthermore, the principal component analysis (PCA) of the inverse of the noise covariance matrix and the target speech covariance matrix involved in the jointly trained MVDR and neural network is unstable, leading to fewer optimal results. In addition, ambient noise and harmful indoor noise can significantly affect the quality of speech signals, thereby reducing the effectiveness of many speech communication systems, such as digital hearing aids and automatic speech recognition (ASR) systems.
[0019] To alleviate this problem, speech enhancement and speech separation algorithms have been proposed. With the resurgence of neural networks, deep learning methods can achieve better objective performance. However, significant nonlinear distortion often occurs in the separated target speech, thus impairing the performance of the ASR system. Minimum variance distortionless response (MVDR) filters aim to reduce noise while maintaining the target speech without distortion. In recent years, MVDR systems based on neural network (NN) time-frequency (TF) mask predictors have significantly reduced the word error rate (WER) of ASR systems with relatively small distortion, but residual noise remains because block-level or speech-level beamforming weights are not optimal for noise reduction. Several frame-level MVDR weight estimation methods have been proposed, and the authors estimate the covariance matrix recursively. However, when jointly trained with NN, the calculated frame-by-frame weights are not stable. Existing research shows that recurrent neural networks (RNNs) can effectively learn matrix inversion, and when RRNs are jointly trained with NNs, RRNs can better stabilize the matrix inversion and principal component analysis (PCA) processes.
[0020] Therefore, for mask-based MVDR beamforming architectures, using RNNs instead of traditional mathematical methods to predict the matrix inversion of noise covariance and the steering vector PCA of the target speech covariance matrix may be advantageous. This allows the entire architecture to be trained in a single, jointly trained deep learning module. Unlike traditional mask-based beamforming algorithms that can only compute block-level or utterance-level weights, the proposed ADL-MVDR adaptively acquires frame-by-frame weights, which helps reduce residual noise. Since RNNs are recursive models, the covariance matrices of noise and target speech can be updated automatically recursively without requiring manual parameter setting. Furthermore, complex-valued filters can be used instead of the commonly used per-TF-point masks to compute the noise and target speech covariance matrices. This can lead to more accurate estimation of the covariance matrix and stabilize the training of RNN-based matrix inversion and PCA. The jointly optimized complex-valued filters and ADL-MVDR can be used in an end-to-end manner.
[0021] This document describes aspects with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0022] Now for reference Figure 1 The functional block diagram of the networked computing environment illustrates a speech recognition system 100 (hereinafter referred to as the "System") for separating the speech of a target speaker using a fully neural network approach. It should be understood that... Figure 1This is merely a description of one implementation method and does not imply any limitation on the environments in which different implementation methods can be implemented. Many modifications can be made to the described environment based on design and implementation requirements.
[0023] System 100 may include computer 102 and server computer 114. Computer 102 may communicate with server computer 114 via communication network 110 (hereinafter referred to as the "network"). Computer 102 may include processor 104 and software program 108, which is stored on data storage device 106 and is capable of interfacing with a user and communicating with server computer 114. Reference will be made below. Figure 4 The computer 102 may include internal component 800A and external component 900A, and the server computer 114 may include internal component 800B and external component 900B. The computer 102 may be, for example, a mobile device, telephone, personal digital assistant, netbook, laptop, tablet, desktop computer, or any type of computing device capable of running programs, accessing networks, and accessing databases.
[0024] Server computer 114 can also operate in cloud computing service models, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), as described below relative to... Figure 5 and Figure 6 The server computer 114 can also be located in a cloud computing deployment model, such as a private cloud, community cloud, public cloud, or hybrid cloud.
[0025] A server computer 114, which can be used for speech recognition, is permitted to run a speech recognition program 116 (hereinafter referred to as the "program") that can interact with a database 112. The following is in contrast to... Figure 3 To illustrate the speech recognition program method in more detail, in one embodiment, computer 102 may operate as an input device including a user interface, while program 116 may run primarily on server computer 114. In an alternative embodiment, program 116 may run primarily on one or more computers 102, while server computer 114 may be used to process and store data used by program 116. It should be noted that program 116 may be a standalone program or may be integrated into a larger speech recognition program.
[0026] However, it should be noted that in some cases, the processing for program 116 can be shared between computer 102 and server computer 114 in any proportion. In another embodiment, program 116 can operate on more than one computer, server computer, or some combination of computers and server computers; for example, multiple computers 102 communicate with a single server computer 114 via network 110. In another embodiment, for example, program 116 can operate on multiple server computers 114 that communicate with multiple client computers via network 110. Alternatively, the program can run on a web server that communicates with servers and multiple client computers via a network.
[0027] Network 110 may include wired connections, wireless connections, fiber optic connections, or combinations thereof. Typically, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 may include various types of networks, such as a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as the Public Switched Telephone Network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., 5G, LTE, 3G, CDMA, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, a fiber optic-based network, and / or combinations of these or other types of networks.
[0028] by Figure 1 The number and arrangement of devices and networks shown are for illustrative purposes only. In reality, there may be other devices and / or networks, fewer devices and / or networks, different devices and / or networks, or networks similar to those shown. Figure 1 The devices and / or networks shown are arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown can be implemented as multiple distributed devices. Additionally or alternatively, a group of devices in system 100 (e.g., one or more devices) can perform one or more functions described as being performed by another group of devices in system 100.
[0029] Now for reference Figure 2The document describes an exemplary speech recognition system 200 according to one or more embodiments. The speech recognition system 200 may include an audio input 202, a camera 204, a complex ratio filter 206, gated recurrent unit (GRU) based networks (GRU-Nets) 208A and 208B, linear layers 210A and 210B, a frame-level weighting module 212, and a speech separation module 214.
[0030] The target speaker's direction of arrival (DOA) can be used to instruct a dilated convolutional neural network (CNN) to extract the target speech from a multi-speaker mix. Audio input 202 can receive speaker-independent features (e.g., logarithmic power spectrum (LPS) and interaural phase difference (IPD)) and speaker-related features (e.g., directional feature d(θ)). For example, audio input 202 could be a 15-element non-uniform linear microphone array located in the same position as camera 204 (which could be a wide-angle 180-degree camera). The position of the target speaker's face within the entire view of camera 204 can provide a coarse estimate of the target speaker's DOA. The position-guided direction feature (DF) d(θ) can be used to extract the target speech from a specific DOA. The cosine similarity between the target turning vector v(θ) and the IPDS can be calculated. The estimated mask or filter will help in calculating the covariance matrix Φ.
[0031] Consider a mixture of noisy speech y = [y1, y2, ..., y] recorded using a microphone array of size M. M ] T S can represent a clear speech signal, and n can represent interference noise with M channels. Y(t,f) = S(t,f) + N(t,f), where (t,f) can indicate the time and frequency indices of the sound signal in the TF domain, and Y, S, and N can represent the corresponding variables in the TF domain. Separate speech s MVDR (t, f) can be obtained as follows:
[0032]
[0033] in, The MVDR weight at frequency index f can be represented, where H represents the Hermitian operator. The goal of an MVDR beamformer might be to minimize noise power while maintaining undistorted target speech, which can be expressed as:
[0034]
[0035] Where, Φ NN The covariance matrix represents the noise power density spectrum (PSD), and This represents the direction vector of the target speech. Different approaches can be used to derive the MVDR beamforming weights. One solution is based on the direction vector and can be derived by applying principal component analysis (PCA) to the speech covariance matrix. Another solution is based on reference channel selection.
[0036]
[0037] Where Φ SS The covariance matrix of the speech PSD is represented. This is the one-hot vector for selecting the reference microphone channel. Note that matrix inversion and PCA can be unstable, especially when jointly trained with a neural network.
[0038] The complex ratio filter 206 can accurately estimate the target speech with less phase distortion using a complex ratio mask (denoted as cRM), which is advantageous for human listeners. In this case, the estimated speech... and speech covariance matrix Φ SS It can be calculated as follows:
[0039]
[0040] Where * denotes a complex multiplier, cRM S The estimated covariance matrix (cRM) represents the speech target. The noise covariance matrix Φ is also present. NN It can be obtained in a similar manner. However, the covariance matrix Φ derived here is at the discourse level, which is not optimal for each frame, resulting in a high level of residual noise.
[0041] GRU networks 208A and 208B can be used to replace matrix inversion and PCA for frame-level beamforming weight estimation. Using RNNs can utilize weighted information from all previous frames and eliminates the need for any heuristic update factors between consecutive frames required in recursive methods.
[0042] To better utilize nearby TF information and stabilize the estimated statistical variable (denoted as Φ). SS and Φ NN The complex ratio filter (cRF) 206 can be used to estimate speech and noise components. For each TF window, the cRF 206 can be applied to its K×L neighborhood windows, where K and L represent the number of neighborhood time and frequency windows.
[0043]
[0044] in This represents the estimated speech using a complex ratio filter. The cRF206 is equivalent to K×L cRMs, each applied to a corresponding shifted version of the noise spectrogram (i.e., along the time and frequency axes). The center mask of the cRF (i.e., the cRM) used for normalization is then utilized. S We use (t,f) to calculate the frame-level speech covariance matrix. It can be understood that, in order to preserve frame-level temporal information, in Φ... SS The time dimension of (t,f) may not contain a sum. The frame-level noise covariance matrix Φ NN (t,f) can be obtained in a similar way.
[0045] Two GRU networks, 208A and 208B, can be used to estimate the direction vector and the inverse of the noise covariance matrix. For h v2 The speech covariance matrix is also reweighted using another GRU network. Compared to the traditional frame-by-frame method based on heuristic update factors, GRU networks 208A and 208B can better utilize temporal information from previous frames for statistical term estimation. Furthermore, replacing matrix inversion with GRU networks 208A and 208B addresses the instability problem during joint training with neural networks. The MVDR coefficients can be obtained through the GRU network, as shown below:
[0046]
[0047] The real and imaginary parts of the complex covariance matrix Φ are concatenated and used as inputs to GRU networks 208A and 208B. It can be assumed that explicitly computed speech and noise covariance matrices may be important for RNN learning spatial filtering, which may differ from beamforming weights learned directly by the NN. Utilizing the temporal structure of the RNN, the model recursively accumulates and updates the covariance matrix for each frame. The output of each of GRU networks 208A and 208B can be injected into linear layers 210A and 210B to obtain the final real and imaginary parts of the complex covariance matrix or direction vector. Frame-level ADL-MVDR weights can be computed by frame-level weight module 212 as follows:
[0048]
[0049] Where h(t,f) is frame-by-frame and differs from the speech-level weights in traditional mask-based MVDR. Finally, the enhanced speech is obtained by the speech separation module 214, as shown below:
[0050]
[0051] Now for reference Figure 3 The diagram depicts the operational flowchart illustrating the steps of a method 300 for speech recognition. In some embodiments, Figure 3One or more processing boxes can be generated by computer 102 ( Figure 1 ) and server computer 114 ( Figure 1 ) Execution. In some implementations, Figure 3 One or more processing blocks may be executed by another device or a group of devices that are separate from or include computer 102 and server computer 114.
[0052] In operation 302, method 300 may include receiving audio data corresponding to one or more speakers.
[0053] In operation 304, method 300 includes a gated recurrent unit-based network to estimate the covariance matrix of the target speech and noise associated with the received audio data.
[0054] In operation 306, method 300 includes generating a predicted target waveform corresponding to the target speaker among the one or more speakers, based on the estimated covariance matrix and a minimum variance-free response function.
[0055] It should be understood that, Figure 3 This is merely an illustration of one implementation and does not imply any limitation on how different embodiments can be implemented. Many modifications can be made to the depicted environment based on design and implementation requirements.
[0056] Figure 4 According to the illustrative embodiments Figure 1 The diagram shown is a block diagram of the internal and external components of the computer. It should be understood that... Figure 4 This is merely a description of one implementation method and does not imply any limitation on the environments in which different implementation methods can be implemented. Many modifications can be made to the described environment based on design and implementation requirements.
[0057] Computer 102 ( Figure 1 ) and server computer 114 ( Figure 1 ) can include Figure 4 The corresponding sets of internal components 800A, 800B and external components 900A, 900B are shown. Each set of internal components 800 includes one or more processors 820 on one or more buses 826, one or more computer-readable RAMs 822 and one or more computer-readable ROMs 824, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0058] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. In some embodiments, processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that allow communication between internal components 800A, 800B.
[0059] Server computer 114 ( Figure 1 One or more operating systems 828 and software programs 108 on ) Figure 1 ), and speech recognition program 116 ( Figure 1 The data is stored on one or more corresponding computer-readable tangible storage devices 830 for execution by one or more corresponding processors 820 via one or more corresponding RAMs 822 (typically including cache memory). Figure 4 In the illustrated embodiment, each computer-readable tangible storage device 830 is a disk storage device of an internal hard disk drive. Alternatively, each computer-readable tangible storage device 830 is a semiconductor storage device, such as ROM 824, EPROM, flash memory, optical disc, magneto-optical disc, solid-state drive, optical disc (CD), digital universal disc (DVD), floppy disk, cassette tape, magnetic tape, and / or another type of non-volatile computer-readable tangible storage device capable of storing computer programs and digital information.
[0060] Each group of internal components 800A, 800B also includes an R / W drive or interface 832 for reading from or writing to one or more portable computer-readable tangible storage devices 936 (e.g., CD-ROM, DVD, Memory Stick, magnetic tape, disk, optical disc, or semiconductor storage device). Software programs, such as software program 108 ( Figure 1 ) and speech recognition program 116 ( Figure 1 The data can be stored on one or more of the corresponding portable computer-readable tangible storage devices 936, and read and loaded into the corresponding hard disk drive 830 via the corresponding R / W drive or interface 832.
[0061] Each group of internal components 800A and 800B also includes a network adapter or interface 836, such as a TCP / IP adapter card; a wireless Wi-Fi interface card; or a 3G, 4G, or 5G wireless interface card or other wired or wireless communication links. Server computer 114 ( Figure 1 Software program 108 on ) Figure 1 ), speech recognition program 116 ( Figure 1It can be downloaded from an external computer to computer 102 via a network (such as the Internet, a local area network, or other wide area networks) and a corresponding network adapter or interface 836. Figure 1 The network includes a network adapter or interface 836 and a server computer 114. Software programs 108 and 116 on the server computer 114 are loaded into the corresponding hard disk drives 830 from the network adapter or interface 836. The network may include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0062] Each group of external components 900A, 900B may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. External components 900A, B may also include a touchscreen, virtual keyboard, touchpad, pointing device, and other human-machine interface devices. Each group of internal components 800A, 800B also includes a device driver 840 that interfaces with the computer display monitor 920, keyboard 930, and computer mouse 934. Device driver 840, R / W driver or interface 832, and network adapter or interface 836 include hardware and software (stored in storage device 830 and / or ROM 824).
[0063] It should be understood in advance that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings herein is not limited to a cloud computing environment. Rather, some embodiments can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0064] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing power, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and deployed with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0065] The features are as follows:
[0066] On-demand self-service: Cloud users can unilaterally provide computing power, such as server time and network storage, which can be done automatically on demand without human interaction with the service provider.
[0067] Broad network access: Capabilities are available through the network and accessed via standard mechanisms to facilitate the use of a wide variety of thin or fat client platforms (e.g., mobile phones, laptops, and PDAs).
[0068] Resource pooling: Pooling vendor computing resources to dynamically allocate or pre-allocate different physical and virtual resources based on demand, thereby serving multiple users in a multi-tenant model. It feels location-agnostic because customers typically cannot control or know the precise location of the resources provided, but can specify the location in a highly abstract way (e.g., country, continent, or data center).
[0069] Rapid scalability: Capacity can be supplied quickly and elastically under certain conditions to rapidly expand outward and rapidly release inward. For users, the available capacity appears unlimited, allowing them to appropriately purchase any amount of capacity at any time.
[0070] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metrics that vary across certain levels of abstraction applicable to service types (e.g., storage, processors, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported to provide transparency to both the service providers and users.
[0071] The service model is as follows:
[0072] Software as a Service (SaaS): This provides users with the ability to use a vendor's applications running on cloud infrastructure. These applications can be accessed from various client-side devices through thin client interfaces such as web browsers (e.g., web-based email). Users do not manage or control the underlying cloud infrastructure (including networks, servers, operating systems, storage, or even the applications themselves), except for limited user-specific application settings.
[0073] Platform as a Service (PaaS): This provides users with the ability to deploy user-created or acquired applications onto cloud infrastructure. These applications are created using vendor-supported programming languages and tools. Users no longer manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, if any, the environment configuration for application hosting.
[0074] Infrastructure as a Service (IaaS): This provides users with the capability to provision processing, storage, networking, and other basic computing resources on which users can deploy and run any software, including operating systems and applications. Users no longer manage or control the underlying cloud infrastructure, but they can manage the operating system, storage, and deployed applications, and have limited control over selected network components (e.g., host firewalls).
[0075] The deployment modes are as follows:
[0076] Private cloud: The cloud infrastructure operates for a single organization. It may be managed by the organization itself or a third party, and can be deployed internally or externally.
[0077] Community cloud: A cloud infrastructure shared by multiple organizations that supports a specific community with common concerns (e.g., missions, security needs, policies, and compliance considerations). It may be managed by the organization itself or a third party, and can be deployed on-premises or externally.
[0078] Public cloud: Cloud infrastructure provided to the general public or large industry groups and owned by organizations that sell cloud services.
[0079] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that support the portability of data and applications (e.g., cloud bursting for load balancing between clouds).
[0080] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure comprised of a network of interconnected nodes.
[0081] refer to Figure 5 This describes an illustrative cloud computing environment 500. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 10. Local computing devices used by cloud users, such as personal digital assistants (PDAs) or cellular phones 54A, desktop computers 54B, laptops 54C, and / or automotive computer systems 54N, can communicate with the cloud computing nodes. The cloud computing nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks (e.g., private, community, public, or hybrid clouds, or combinations thereof, as described above). This allows the cloud computing environment 500 to provide infrastructure, platform, and / or software as services that cloud users do not need to maintain on their local computing devices. It should be understood that... Figure 5 The types of computing devices 54A-N shown are intended to illustrate only that cloud computing node 10 and cloud computing environment 500 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0082] refer to Figure 6 This demonstrates the 500 (cloud computing environment) Figure 5 This provides a set of functional abstraction layers, 600. It should be understood beforehand that... Figure 6 The components, layers, and functions shown are for illustrative purposes only, and the embodiments are not limited thereto. As shown, the following layers and corresponding functions are provided.
[0083] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: a mainframe computer 61; a server 62 based on a Reduced Instruction Set Computer (RISC) architecture; a server 63; a blade server 64; a storage device 65; and a network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0084] The virtualization layer 70 provides an abstraction layer from which examples of the following virtual entities can be provided: virtual server 71; virtual storage 72; virtual network 73, including virtual private network; virtual application and operating system 74; and virtual client 75.
[0085] In one example, management layer 80 can provide the following functionalities: Resource Provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 82 provides cost detection and billing or invoicing for the consumption of these resources as they are utilized in the cloud computing environment. In one example, these resources may include application software licenses. Security provides authentication for cloud users and tasks, as well as protection for data and other resources. User Portal 83 provides access to the cloud computing environment for users and system administrators. Service Level Management 84 provides cloud resource allocation and management to meet the required service level. Service Level Agreement (SLA) Planning and Implementation 85 provides pre-scheduling and procurement of cloud resources based on anticipated future needs according to the SLA.
[0086] Workload layer 90 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytics and processing 94; transaction processing 95; and speech recognition 96. Speech recognition 96 can use a fully neural network approach to separate the speech of the target speaker.
[0087] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of technical detail in the integration. A computer-readable medium may include a computer-readable non-volatile storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform operations.
[0088] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to: electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or those with raised structures in grooves on which instructions are recorded, and any suitable combination of the foregoing. The computer-readable storage media used herein should not be construed as transient signals, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0089] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. This network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.
[0090] Computer-readable program code / instructions for performing operations can be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++, and process programming languages such as "C" or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, as part of a standalone software package on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet provided by an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) can execute computer-readable program instructions by utilizing the status information of the computer-readable program instructions to customize the electronic circuitry for performance aspects or operations.
[0091] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or the other programmable data processing apparatus, create means for implementing the functions / actions specified in the blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other apparatus to operate in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture comprising instructions for implementing aspects of the functions / actions specified in the blocks of the flowchart and / or block diagram.
[0092] Computer-readable program instructions can be incorporated into a calculator, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device, thereby producing computer-achievable processing, such that the instructions executed on the computer, other programmable apparatus or other device perform the functions or actions specified in the boxes of the flowchart and / or block diagram.
[0093] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each box in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. The method, computer system, and computer-readable medium may include more boxes, fewer boxes, different boxes, or boxes arranged differently than those illustrated. In some alternative implementations, the functions recorded in the boxes may appear out of the order shown in the figures. For example, in fact, two boxes shown consecutively may be executed simultaneously or substantially simultaneously, or these boxes may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0094] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit these implementations. Therefore, since the operation and behavior of the systems and / or methods are described herein without reference to specific software code, it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0095] No element, action, or instruction used herein should be construed as critical or necessary unless explicitly described as such. Furthermore, as used herein, the article “a” is intended to include one or more items and may be used interchangeably with “one or more.” Additionally, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with “one or more.” If only one item is intended, the term “a” or similar language is used. Furthermore, as used herein, the terms “having,” “possessing,” etc., are intended to indicate open-ended terms. Furthermore, unless explicitly stated otherwise, “based on” is intended to mean “at least partially based on.”
[0096] For illustrative purposes, descriptions of various aspects and embodiments have been presented, but are not exhaustive or limited to the disclosed embodiments. Even though combinations of features are listed in the claims and / or disclosed in the specification, these combinations are not intended to limit the possible implementations disclosed. In fact, many of these features can be combined in ways not specifically stated in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the possible disclosure includes each dependent claim combined with every other claim in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, practical applications of techniques found in the market, or improvements to the technology, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A speech recognition method executed by a processor, comprising: Receive audio data corresponding to one or more speakers; A gated recurrent unit network (GRU-Net) is used to estimate the covariance matrix of the target speech and noise associated with the received audio data; as well as Based on the estimated covariance matrix, a predicted target waveform corresponding to the target speaker among the one or more speakers is generated using the minimum variance distortionless response function (MVDR). as well as The covariance matrix of one or more frames is recursively accumulated and updated by the GRU-Net.
2. The method according to claim 1, characterized in that, The covariance matrix corresponds to the noise power density spectrum and the speech power density spectrum.
3. The method according to claim 1, characterized in that, The predicted target waveform is generated using the MVDR coefficients corresponding to the covariance matrix.
4. The method according to claim 3, characterized in that, The MVDR coefficients are calculated by GRU-Net based on the real and imaginary parts of the covariance matrix connected by GRU-Net.
5. The method according to claim 1, characterized in that, Also includes: A linear layer is used to obtain the final real and imaginary parts of the covariance matrix.
6. The method according to claim 1, characterized in that, The target speaker is identified based on the direction of arrival corresponding to the received audio data.
7. A computer system for speech recognition, the computer system comprising: One or more computer-readable non-volatile storage media are configured to store computer program code; and One or more computer processors are configured to access and operate in accordance with the instructions of the computer program code, the computer program code comprising: A receiving code is configured to cause the one or more computer processors to receive audio data corresponding to one or more speakers; Estimation code is configured to enable the one or more computer processors to estimate the covariance matrix of target speech and noise associated with the received audio data based on a gated recurrent unit network (GRU-Net); and The code is configured to cause the one or more computer processors to generate a predicted target waveform corresponding to the target speaker among the one or more speakers, based on the estimated covariance matrix and using a minimum variance distortion-free response function (MVDR); and Accumulation and update codes are configured to cause the one or more computer processors to recursively accumulate and update the covariance matrix of one or more frames via the GRU-Net.
8. The computer system according to claim 7, characterized in that, The covariance matrix corresponds to the noise power density spectrum and the speech power density spectrum.
9. The computer system according to claim 7, characterized in that, The predicted target waveform is generated using the MVDR coefficients corresponding to the covariance matrix.
10. The computer system according to claim 9, characterized in that, The MVDR coefficients are calculated by GRU-Net based on the real and imaginary parts of the covariance matrix connected by GRU-Net.
11. The computer system according to claim 7, further comprising: The code is configured to cause the one or more computer processors to use a linear layer to obtain the final real and imaginary parts of the covariance matrix.
12. The computer system according to claim 7, characterized in that, The target speaker is identified based on the direction of arrival corresponding to the received audio data.
13. A non-volatile computer-readable medium having a computer program stored thereon for speech recognition, the computer program being configured to cause one or more computer processors to: Receive audio data corresponding to one or more speakers; A gated recurrent unit network (GRU-Net) is used to estimate the covariance matrix of the target speech and noise associated with the received audio data; as well as Based on the estimated covariance matrix, a predicted target waveform corresponding to the target speaker among the one or more speakers is generated using the minimum variance distortionless response function (MVDR). as well as The covariance matrix of one or more frames is recursively accumulated and updated by the GRU-Net.
14. The computer-readable medium according to claim 13, characterized in that, The covariance matrix corresponds to the noise power density spectrum and the speech power density spectrum.
15. The computer-readable medium according to claim 13, characterized in that, The predicted target waveform is generated using the MVDR coefficients corresponding to the covariance matrix.
16. The computer-readable medium according to claim 15, characterized in that, The MVDR coefficients are calculated by GRU-Net based on the real and imaginary parts of the covariance matrix connected by GRU-Net.
17. The computer-readable medium according to claim 13, characterized in that, The computer program is further configured to cause the one or more computer processors to use a linear layer to obtain the final real and imaginary parts of the covariance matrix.