Domain adaptive speech recognition using artificial intelligence
By introducing the domain adaptive method of artificial intelligence into speech recognition technology, processing phoneme sequences and generating candidate language data, the problem of resource-intensive and error-prone traditional speech recognition methods is solved, and more efficient and accurate speech recognition effects are achieved.
Patent Information
- Application Number
- CN202380071705.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional speech recognition methods are resource-intensive and error-prone, making it difficult to effectively process complex language data.
The domain adaptive speech recognition technology of artificial intelligence is adopted to process phoneme sequences through an artificial intelligence-based data transformation model, generate a collection of language data candidates, and generate the final speech recognition output using the biased language model and speech recognition model.
Improve the accuracy and efficiency of speech recognition, reduce resource consumption, and enable automatic execution of related actions.
Smart Images

Figure CN120019431A_ABST
Abstract
Description
Background Art
[0001] The present application relates generally to information technology, and more particularly to language data processing. More specifically, speech recognition presents many challenges. For example, conventional speech recognition methods are often resource-intensive and error-prone. Summary of the invention
[0002] In at least one embodiment, a domain adaptive speech recognition technique using artificial intelligence is provided. An example computer-implemented method includes processing a phoneme sequence associated with input speech data using an artificial intelligence-based data conversion model, generating a set of language data candidates, each language data candidate including one or more graphemes, and determining a grapheme subset from the language data candidate set for a target pair of one or more phonemes and one or more graphemes. The method also includes processing at least a portion of the subset of graphemes using at least one biased language model and an artificial intelligence-based speech recognition model to generate a first speech recognition output. In addition, the method includes replacing at least a portion of the grapheme subset in the first speech recognition output with at least one grapheme from the one or more graphemes in the target pair, generating a second speech recognition output, and performing one or more automated actions based at least in part on the second speech recognition output.
[0003] Another embodiment of the present invention or its elements may be implemented in the form of a computer program product that tangibly embodies computer-readable instructions that, when implemented, cause a computer to perform a plurality of method steps, as described herein. In addition, another embodiment of the present invention or its elements may be implemented in the form of a system comprising a memory and at least one processor coupled to the memory and configured to perform the method steps. In addition, another embodiment of the present invention or its elements may be implemented in the form of an apparatus or its elements for performing the method steps described herein; the apparatus may include a hardware module or a combination of hardware and software modules, wherein the software modules are stored in a tangible computer-readable storage medium (or multiple such media).
[0004] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 is a diagram showing an illustrative speech recognition workflow according to an example embodiment of the present invention;
[0006] Figure 2 is a diagram illustrating a neural network training process according to an exemplary embodiment of the present invention;
[0007] Figure 3 is a flow chart illustrating a technique according to an example embodiment of the present invention; and
[0008] Figure 4 is a diagram illustrating a computing environment in which at least one embodiment of the invention may be implemented. DETAILED DESCRIPTION
[0009] As described herein, at least one embodiment includes domain adaptive speech recognition using artificial intelligence. Such an embodiment includes training a first generator that receives a phoneme sequence and outputs a plurality of candidate words, each candidate word including one or more graphemes. As an illustration, in one or more embodiments, the first generator may use and / or include at least a portion of a trained end-to-end automatic speech recognition (ASR) model.
[0010] In one or more embodiments, the generator includes at least one neural network model. As described in further detail below, Figure 1 In the example embodiment depicted in , the artificial intelligence-based data conversion model 104 (also referred to herein as an enhanced sound-like generator) includes a neural network model, and the prediction portion of the generator reuses the prediction portion of a recurrent neural network transducer (RNN-T) model (e.g., the RNN-T decoder 110). That is, in one or more additional embodiments, reusing that portion of the RNN-T model is not mandatory, and the generator may include a non-neural network model, such as, for example, a rule-based generator that can convert phonemes into graphemes.
[0011] Additionally, for a target pair of a phoneme sequence and a word of one or more graphemes, at least one embodiment includes using the trained first generator to determine N best words for the one or more graphemes. For example, the N best words may be output in order of probability. In such an embodiment, the generator reads the sound-like elements of the speech input and outputs the N best enhanced sound-like words. Additionally, one or more embodiments include training at least one biased language model (LM) using the determined N best words for the one or more graphemes. In such an embodiment, the RNN-T model (e.g., a predictive network and a speech encoder, such as by Figure 1 The unit 110 in FIG. 1 has an inseparable LM, which can be trained with a pair of speech data and text data. In addition, the bias LM can be trained with only text data.
[0012] Further, in such an embodiment, the trained bias LM can be used with the trained end-to-end ASR model to generate one or more output transcripts and replace at least one of the N best words for one or more graphemes appearing in the one or more transcripts with a word for one or more graphemes from the target pair. As an illustration, in the case of English, a user may display a setting as: "IEEE" and sounds like: "AY TR IH P AX L IY", and the generator converts the sounds like to enhanced sounds like: ["I triple E," "eye triple iye"]. Further, in such an embodiment, the language model is trained with these enhanced sound like entries, and the output transcript includes enhanced sounds like, such as "I joined I triple E". Further, since the end user may expect "IEEE", a word replacement may be performed to convert "I triple E" to "IEEE". Thus, in such an example, the end user receives "I joined IEEE".
[0013] As described in further detail herein, in one or more embodiments, the trained end-to-end ASR model may include a trained RNN-T, and the first generator uses the prediction network of the trained RNN-T as its own prediction network. As described above, in such embodiments, the separate language model and acoustic model cannot be separated, but the prediction network is similar to the language model because, as Figure 2 As shown, for example, the prediction network 210 reads the previous character (y u-1 ) and output the next character (y u ). Thus, in such embodiments, the prediction network may be implemented as a language model.
[0014] Additionally, at least one embodiment includes obtaining at least one phoneme sequence from processing a word of graphemes as an input to the generator.
[0015] Figure 1 is a diagram showing an illustrative speech recognition workflow associated with the domain adaptive speech recognition system 126 according to an exemplary embodiment of the present invention. Specifically, Figure 1 An example embodiment of using the LM customization component 102 and the decoding component 114 to recognize the unknown word "IEEE" in English is described. Figure 1 The example embodiment depicted in includes generating a list of words (each including one or more graphemes) from a given sequence of phonemes using a translator component referred to herein as an artificial intelligence-based data conversion model 104 (or an enhanced sound-like generator). Figure 1As shown, the artificial intelligence-based data conversion model 104 includes a phoneme encoder and a prediction network shared with the RNN-T decoder 110 , wherein the prediction network allows the artificial intelligence-based data conversion model 104 to generate graphemes 106 that mimic the behavior of the RNN-T decoder 110 .
[0016] In addition, if Figure 1 As shown, the speech data 116 from the decoding component 114 is decoded by the RNN-T decoder 110 in combination with the bias LM 108 into a word list representing an initial speech recognition result 118. Moreover, one or more words in the initial speech recognition result 118 are replaced by at least a portion of the original input speech data 103 processed by the converter 112, thereby producing a final speech recognition result 120. In one or more embodiments, the converter 112 converts the enhanced sound-like output (e.g., graphemes 106) into a display in the original input speech data 103.
[0017] One or more embodiments include using phoneme and grapheme pairs to train the artificial intelligence-based data conversion model 104. In at least one embodiment, the artificial intelligence-based data conversion model 104 may include an encoder-decoder network. Figure 1 As depicted in conjunction with the LM customization component 102, a user adds raw input speech data 103 comprising one or more pairs of data that appears as (w) and data that sounds like (p). A trained model (i.e., an artificial intelligence-based data conversion model 104) generates up to N optimal graphemes 106 from a larger set of graphemes based at least in part on character sequence generation probabilities from the input phonemes (p). As an example, in the case of an English input of "AY TR IH P AX L IY," the output of the artificial intelligence-based data conversion model 104 may be "I triple E" with a probability of 0.03 and / or "eye triple iye" with a probability of 0.02. In addition, the graphemes may then be used 106 trains the biased LM 108 .
[0018] like Figure 1 As further described in , decoding the input speech data may include using the RNN-T decoder 110, through the biased LM 108, to read and process the original input speech data 103, and output an initial speech recognition result 118 (e.g., one or more transcripts). One or more embodiments also include replacing the data (w) in the initial speech recognition result 118 with the corresponding string of data (w) displayed as from the original input speech data 103. Thereby, a final speech recognition result 120 is generated.
[0019] Figure 2is a diagram illustrating a neural network training process according to an exemplary embodiment of the present invention. Specifically, Figure 2 Describes the use of one or more phonetic character pairs (e.g., one or more y u-1 and x i..T The RNN-T decoder 210 is trained with the trained prediction network of the RNN-T decoder 210, and the artificial intelligence-based data conversion model 204 is trained, as described in detail below. For example, in at least one embodiment, the training of the artificial intelligence-based data conversion model 204 includes initializing the prediction network of the artificial intelligence-based data conversion model 204 with the trained prediction network of the RNN-T decoder 210, freezing the prediction network of the artificial intelligence-based data conversion model 204 during the training of the artificial intelligence-based data conversion model 204, and initializing the prediction network of the artificial intelligence-based data conversion model 204 with one or more character (grapheme)-phoneme pairs (e.g., one or more y u-1 and p i…T The artificial intelligence-based data conversion model 204 is trained by using the neural network (for each pair). The training of the neural network is an iterative training, where each iteration includes a forward process and a reverse process (also called back propagation). The freezing step prohibits the reverse process in this training. For this reason, the weights of the frozen part of the neural network are unchanged during training. In addition, in the case of Figure 2 (and Figure 1 ), the RNN-T decoder 210 and the artificial intelligence-based data conversion model 204 use a common and / or shared prediction network.
[0020] In one or more embodiments, the speech recognition model described in detail herein is not limited to RNN-T, but may include, for example, one or more types of end-to-end ASR models. In addition, in combination with, for example, the biased LM described in detail herein, the input is not limited to one or more individual words, but may also include one or more sentences and / or more than one compound word.
[0021] Merely for illustration, consider an English use case example, where an end user provides raw input speech data 103 displayed as: "IEEE" and sounds like: "AY TR IH P AX L IY." Further, the AI-based data conversion model 104 converts the sounds like: "AY TR IH P AX L IY." to an enhanced sounds like: ["I triple E", "eyetriple iye"] (i.e., grapheme 106), and the biased LM 108 is implemented using the enhanced sounds like grapheme 106. Additionally, the converter 112 prepares the enhanced sounds like ["I triple E", "eye triple iye"] to display as: "IEEE." During the run-time (decoding) process, the RNN-T decoder 110 may output "I joined I triple E", and word replacement is performed with the converter 112 to convert "I triple E" to "IEEE", such that the end user receives "I joined IEEE". Thus, as seen by such examples, one or more embodiments include training the biased LM 108 with easy-to-learn graphemes, rather than training the biased LM 108 with the original user input 103 that appears and sounds like , and then restoring the output to the original word (displayed as ).
[0022] In addition, in at least one embodiment, the artificial intelligence-based data conversion model may include at least one RNN-T converter model having a prediction network of an RNN-T speech recognition network, but such a model type merely represents an example and may be implemented as other model types. Figure 1 The example embodiments depicted in the examples include the use of the English language, but one or more embodiments are not limited to English, but may include the use of various other languages (e.g., Japanese). Additionally, in such embodiments, the varying pronunciations of one or more words (within a given language) may be utilized in conjunction with the speech recognition models detailed herein.
[0023] Figure 3 302 includes processing a phoneme sequence associated with input speech data using an artificial intelligence-based data conversion model to generate a set of language data candidates, each of which includes one or more phonemes. In at least one embodiment, the artificial intelligence-based data conversion model includes at least one encoder-decoder neural network. In addition, in one or more embodiments, the artificial intelligence-based data conversion model implements a prediction network shared by an artificial intelligence-based speech recognition model.
[0024] One or more embodiments may also include training an artificial intelligence-based data conversion model using one or more grapheme-phoneme pairs. In addition, in at least one embodiment, the input speech data may include one or more of one or more individual words, one or more sentences, and one or more compound words.
[0025] Step 304 includes determining, for a target pair of one or more phonemes and one or more graphemes, a subset of graphemes from the candidate set of language data.
[0026] Step 306 includes generating a first speech recognition output by processing at least a portion of the grapheme subset using at least one biased language model and an artificial intelligence-based speech recognition model. In at least one embodiment, the artificial intelligence-based speech recognition model includes a recurrent neural network transducer. Additionally or alternatively, the artificial intelligence-based speech recognition model may include at least one end-to-end automatic speech recognition model.
[0027] One or more embodiments may also include training at least one biased language model using at least a portion of the grapheme subset.
[0028] Step 308 includes generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one grapheme from the one or more graphemes of the target pair.
[0029] Step 310 includes performing one or more automated actions based at least in part on the second speech recognition output. In at least one embodiment, performing the one or more automated actions includes automatically training at least one of an artificial intelligence-based data conversion model and an artificial intelligence-based speech recognition model using feedback data associated with the second speech recognition output. Additionally or alternatively, performing the one or more automated actions may include outputting the second speech recognition output to at least one user in response to the input speech data.
[0030] Furthermore, in one or more embodiments, implementing Figure 3 Software of the techniques described in can be provided as a service in a cloud environment.
[0031] It should be understood that a "model" as used herein refers to an electronically digitally stored set of executable instructions and data values associated with one another that is capable of receiving and responding to a programming or other digital call, invocation, or parsing request based on specified input values to produce one or more output values that can be used as the basis for computer-implemented recommendations, output data displays, machine controls, etc. Those skilled in the art find it convenient to express the model using mathematical equations, but the form of expression is not intended to limit the models disclosed herein to abstract concepts; rather, each model herein has practical application in a computer in the form of stored executable instructions and data that implement the model using a computer.
[0032] As stated in this article, Figure 3 The technology depicted in the can also include providing a system, wherein the system includes different software modules, each different software module is contained on a tangible computer-readable recordable storage medium. For example, all modules (or any subset thereof) can be on the same medium, or each can be on a different medium. The module can include any or all components shown in the figure and / or described herein. In an embodiment of the present invention, the module can be run on a hardware processor, for example. Then, the different software modules of the system executed on the hardware processor as described above can be used to perform the method steps. In addition, the computer program product can include a tangible computer-readable recordable storage medium having a code suitable for being executed to perform at least one method step described herein.
[0033] in addition, Figure 3 The techniques described in the can be implemented via a computer program product, which can include a computer usable program code stored in a computer readable storage medium in a data processing system, and wherein the computer usable program code is downloaded from a remote data processing system over a network. In addition, in an embodiment of the present invention, the computer program product can include a computer usable program code stored in a computer readable storage medium in a server data processing system, and wherein the computer usable program code is downloaded to the remote data processing system over a network for use in the computer readable storage medium using the remote system.
[0034] Embodiments of the present invention or elements thereof may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and configured to perform the exemplary method steps.
[0035] Various aspects of the present disclosure are described by narrative text, flow charts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flow chart, depending on the technology involved, the operations may be performed in an order different from the order shown in a given flow chart. For example, again depending on the technology involved, two operations shown in consecutive flow chart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0036] Computer program product embodiments ("CPP embodiments" or "CPP") are terms used in this disclosure to describe any collection of one or more storage media (also referred to as "media") collectively included in a collection of one or more storage devices that collectively include machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. Without limitation, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), static random access memories (SRAM), compact disk read-only memories (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / land formed in a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, should not be construed as storing in the form of transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, light pulses transmitted through fiber optic cables, electrical signals transmitted through wires and / or other transmission media. As will be appreciated by those skilled in the art, data is typically moved at certain occasional points in time during normal operation of the storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device transient because the data is not transient while it is stored.
[0037] The computing environment 400 includes an example of an environment for executing at least some of the computer codes involved in performing the methods of the present invention, such as domain adaptive speech recognition code 426. In addition to the code 426, the computing environment 400 includes, for example, a computer 401, a wide area network (WAN) 402, an end user device (EUD) 403, a remote server 404, a public cloud 405, and a private cloud 406. In this embodiment, the computer 401 includes a processor group 410 (including a processing circuit 420 and a cache 421), a communication structure 411, a volatile memory 412, a permanent storage device 413 (including an operating system 422 and code 426 as described above), a peripheral device group 414 (including a user interface (UI) device group 423, a storage device 424, and an Internet of Things (IoT) sensor group 425) and a network module 415. The remote server 404 includes a remote database 430. Public cloud 405 includes a gateway 440 , a cloud orchestration module 441 , a host physical machine group 442 , a virtual machine group 443 , and a container group 444 .
[0038] Computer 401 may take the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or developed in the future that is capable of running programs, accessing a network, or querying a database such as remote database 430. As is well known in the art of computer technology, and depending on the technology, the performance of computer-implemented methods may be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of computing environment 400, the detailed discussion focuses on a single computer, particularly computer 401, to keep the presentation as simple as possible. Computer 401 may be located in the cloud, even if it is not in the cloud. Figure 4 4. Although shown in the cloud, on the other hand, computer 401 need not be in the cloud unless to any extent that can be positively indicated.
[0039] Processor group 410 includes one or more computer processors of any type known now or to be developed in the future. Processing circuit 420 can be distributed over multiple packages, such as multiple coordinated integrated circuit chips. Processing circuit 420 can implement multiple processor threads and / or multiple processor cores. Cache 421 is a memory located in the processor chip package, and is generally used for data or code that should be available for fast access by threads or cores running on processor group 410. Cache memory is generally organized into multiple levels according to relative proximity to the processing circuit. Alternatively, some or all of the caches of the processor group may be located "off chip". In some computing environments, processor group 410 may be designed to work with qubits and perform quantum computing.
[0040] Computer readable program instructions are typically loaded onto computer 401 to cause processor group 410 of computer 401 to perform a series of operating steps to implement a computer-implemented method, such that the instructions so executed will instantiate the method specified in the flow chart and / or the narrative description of the computer-implemented method included in this document (collectively referred to as "the present method"). These computer readable program instructions are stored in various types of computer readable storage media, such as cache 421 and other storage media discussed below. The program instructions and related data are accessed by processor group 410 to control and direct the execution of the present method. In computing environment 400, at least some of the instructions for executing the present method may be stored in code 426 in permanent storage device 413.
[0041] Communications fabric 411 is the signaling pathways that allow the various components of computer 401 to communicate with each other. Typically, the fabric is comprised of switches and conductive pathways, such as those that make up a bus, a bridge, physical input / output ports, etc. Other types of signal communication pathways may be used, such as fiber optic communication pathways and / or wireless communication pathways.
[0042] Volatile memory 412 is any type of volatile memory now known or developed in the future. Examples include dynamic type RAM or static type RAM. Typically, volatile memory 412 is characterized by random access, but this is not required unless expressly stated. In computer 401, volatile memory 412 is located in a single package and is internal to computer 401, but, alternatively or additionally, volatile memory can be distributed over multiple packages and / or located externally relative to computer 401.
[0043] Permanent memory 413 is any form of non-volatile memory for computers known now or developed in the future. The non-volatility of the memory means that the stored data is maintained regardless of whether power is supplied to computer 401 and / or directly to permanent memory 413. Permanent memory 413 can be a ROM, but usually at least a portion of persistent memory allows the writing of data, the deletion of data, and the rewriting of data. Some common forms of permanent storage include disks and solid-state storage devices. Operating system 422 can take several forms, such as various known proprietary operating systems or operating systems of open source portable operating system interface types using a kernel. The code included in code 426 typically includes at least some of the computer codes involved in executing the method of the present invention.
[0044] The peripheral device group 414 includes the peripheral device group of the computer 401. The data communication connection between the peripheral device and other components of the computer 401 can be implemented in various ways, such as a Bluetooth connection, a near field communication (NFC) connection, a connection made by a cable (such as a universal serial bus (USB) type cable), a plug-in type connection (e.g., a secure digital (SD) card), a connection made through a local area communication network, and even a connection made through a wide area network such as the Internet. In various embodiments, the UI device group 423 may include components such as a display screen, a speaker, a microphone, a wearable device (such as goggles and a smart watch), a keyboard, a mouse, a printer, a touchpad, a game controller, and a tactile device. The memory 424 is an external memory, such as an external hard drive, or a pluggable memory, such as an SD card. The storage 424 can be persistent and / or volatile. In some embodiments, the storage 424 can take the form of a quantum computing storage device for storing data in the form of quantum bits. In embodiments where computer 401 needs to have a large amount of storage (e.g., where computer 401 locally stores and manages a large database), the storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. IoT sensor group 425 consists of sensors that can be used in IoT applications. For example, one sensor may be a thermometer, while another sensor may be a motion detector.
[0045] The network module 415 is a collection of computer software, hardware, and firmware that allow the computer 401 to communicate with other computers via the WAN 402. The network module 415 may include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 415 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software defined networks (SDN)), the control function and the forwarding function of the network module 415 are executed on physically separated devices so that the control function manages several different network hardware devices. Computer-readable program instructions for executing the method of the present invention can be downloaded to the computer 401 from an external computer or an external storage device typically via a network adapter card or a network interface included in the network module 415.
[0046] WAN 402 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any technology now known or developed in the future for transmitting computer data. In some embodiments, WAN 402 may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area such as a Wi-Fi network. WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0047] End-user device 403 is any computer system used and controlled by an end-user (e.g., a customer of a business operating computer 401), and may take any of the forms discussed above in conjunction with computer 401. EUD 403 typically receives useful and useful data from the operation of computer 401. For example, in the hypothetical case where computer 401 is designed to provide recommendations to an end-user, the recommendations would typically be transmitted to EUD 403 from network module 415 of computer 401 via WAN 402. In this manner, EUD 403 may display or otherwise present the recommendations to the end-user. In some embodiments, EUD 403 may be a client device, such as a thin client, a heavy client, a mainframe computer, a desktop computer, etc.
[0048] Remote server 404 is any computer system that provides at least some data and / or functionality to computer 401. Remote server 404 may be controlled and used by the same entity that operates computer 401. Remote server 404 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 401. For example, in the hypothetical case where computer 401 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 401 from remote database 430 of remote server 404.
[0049] Public cloud 405 is any computer system that can be used by multiple entities, which provides on-demand availability of computer system resources and / or other computer capabilities, especially data storage (cloud storage) and computing capabilities, without the need for direct active management by users. Cloud computing generally uses the sharing of resources to achieve consistency and economy of scale. The direct and active management of the computing resources of public cloud 405 is performed by computer hardware and / or software of cloud coordination module 441. The computing resources provided by public cloud 405 are generally implemented by virtual computing environments running on various computers that constitute the host physical machine group 442, which is a full domain of physical computers in public cloud 405 and / or available for the public cloud. Virtual computing environments (VCEs) are generally in the form of virtual machines from virtual machine group 443 and / or containers from container group 444. It should be understood that these VCEs can be stored as images and can be transmitted between various physical machine hosts as images or after instantiation of VCEs. Cloud coordination module 441 manages the transmission and storage of images, deploys new instantiations of VCEs, and manages active instantiations of VCE deployments. Gateway 440 is a collection of computer software, hardware, and firmware that allows public cloud 405 to communicate over WAN 402 .
[0050] Some further explanation of the VCE will now be provided. A VCE can be stored as an "image". A new active instance of the VCE can be instantiated from the image. Two common types of VCEs are virtual machines and containers. Containers are VCEs that use operating system-level virtualization. This refers to an operating system feature in which the kernel allows the existence of multiple isolated user space instances, called containers. From the perspective of the programs running in them, these isolated user space instances typically behave like actual computers. Computer programs running on a normal operating system can utilize all of the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices assigned to the container, a feature known as containerization.
[0051] Private cloud 406 is similar to public cloud 405, except that the computing resources are only available to a single enterprise. Although private cloud 406 is depicted as communicating with WAN 402, in other embodiments, the private cloud can be completely disconnected from the Internet and can only be accessed through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable coordination, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, public cloud 405 and private cloud 406 are both part of a larger hybrid cloud.
[0052] In computing environment 400, computer 401 is shown as being connected to the Internet (see WAN 402). However, in one or more embodiments of the present invention, computer 401 will be isolated from communications through a communications network and not connected to the Internet, operating as an independent computer. In these embodiments, network module 415 of computer 401 may not be necessary or even desirable, in order to ensure isolation and prevent external communications from entering computer 401. Independent computer embodiments are potentially advantageous at least in some applications of the present invention because they are generally safer. In other embodiments, computer 401 is connected to a secure WAN or secure LAN instead of WAN 402 and / or the Internet. In these network connection (i.e., non-independent) embodiments, system designers may want to take appropriate security measures known now or developed in the future to reduce the risk that incoming network communications will not cause security vulnerabilities.
[0053] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "include" and / or "comprise" when used in this specification indicate the presence of the features, steps, operations, elements and / or parts, but do not exclude the presence or addition of another feature, step, operation, element, part and / or combination thereof.
[0054] The description of various embodiments of the present invention has been given for the purpose of illustration, but it is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method comprising: Generate a set of language data candidates by processing a phoneme sequence associated with the input speech data using an artificial intelligence-based data conversion model, each language data candidate including one or more graphemes; for a target pair of one or more phonemes and one or more graphemes, determining a grapheme subset from the candidate set of language data; generating a first speech recognition output by processing at least a portion of the subset of graphemes using at least one biased language model and an artificial intelligence-based speech recognition model; generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one grapheme from the one or more graphemes in the target pair; as well as performing one or more automated actions based at least in part on the second speech recognition output; The method is executed by at least one computing device.
2. The computer-implemented method of claim 1 , wherein: The artificial intelligence-based speech recognition model includes a recurrent neural network transducer.
3. The computer-implemented method of claim 1 , wherein: The artificial intelligence-based data conversion model includes at least one encoder-decoder neural network.
4. The computer-implemented method of claim 1 , wherein: The artificial intelligence-based data conversion model implements a prediction network shared by the artificial intelligence-based speech recognition model.
5. The computer-implemented method of claim 1 , wherein: The artificial intelligence-based speech recognition model includes at least one end-to-end automatic speech recognition model.
6. The computer-implemented method of claim 1 , further comprising: The artificial intelligence based data conversion model is trained using one or more grapheme-phoneme pairs.
7. The computer-implemented method of claim 1 , further comprising: The at least one biased language model is trained using at least a portion of the subset of graphemes.
8. The computer-implemented method of claim 1 , wherein: Performing one or more automated actions includes automatically training at least one of the artificial intelligence-based data conversion model and the artificial intelligence-based speech recognition model using feedback data associated with the second speech recognition output.
9. The computer-implemented method of claim 1 , wherein: Performing one or more automated actions includes outputting the second speech recognition output to at least one user in response to the input speech data.
10. The computer-implemented method of claim 1, wherein: The input speech data includes one or more of the following: one or more individual words, one or more sentences, and one or more compound words.
11. The computer-implemented method of claim 1 , wherein: The software implementing the method is provided as a service in a cloud environment.
12. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions being executable by a computing device to cause the computing device to: Generate a set of language data candidates by processing a phoneme sequence associated with the input speech data using an artificial intelligence-based data conversion model, each language data candidate including one or more graphemes; for a target pair of one or more phonemes and one or more graphemes, determining a grapheme subset from the candidate set of language data; generating a first speech recognition output by processing at least a portion of the subset of graphemes using at least one biased language model and an artificial intelligence-based speech recognition model; generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one grapheme from the one or more graphemes in the target pair; as well as Based at least in part on the second speech recognition output, one or more automated actions are performed.
13. The computer program product of claim 12, wherein: The artificial intelligence-based speech recognition model includes a recurrent neural network transducer.
14. The computer program product of claim 12, wherein: The artificial intelligence-based data conversion model includes at least one encoder-decoder neural network.
15. The computer program product of claim 12, wherein: The artificial intelligence-based data conversion model implements a prediction network shared by the artificial intelligence-based speech recognition model.
16. The computer program product of claim 12, wherein: The artificial intelligence-based speech recognition model includes at least one end-to-end automatic speech recognition model.
17. A system comprising: a memory configured to store program instructions; as well as a processor operatively coupled to the memory to execute the program instructions to: Generate a set of language data candidates by processing a phoneme sequence associated with the input speech data using an artificial intelligence-based data conversion model, each language data candidate including one or more graphemes; for a target pair of one or more phonemes and one or more graphemes, determining a grapheme subset from the candidate set of language data; generating a first speech recognition output by processing at least a portion of the subset of graphemes using at least one biased language model and an artificial intelligence-based speech recognition model; generating a second speech recognition output by replacing at least a portion of the subset of graphemes in the first speech recognition output with at least one grapheme from the one or more graphemes in the target pair; as well as Based at least in part on the second speech recognition output, one or more automated actions are performed.
18. The system of claim 17, wherein: The artificial intelligence-based speech recognition model includes a recurrent neural network transducer.
19. The system of claim 17, wherein: The artificial intelligence-based data conversion model includes at least one encoder-decoder neural network.
20. The system of claim 17, wherein: The artificial intelligence-based data conversion model implements a prediction network shared by the artificial intelligence-based speech recognition model.
Citation Information
Cited By
Domain adaptive speech recognition using artificial intelligence
US12451124B2