Speech-to-text automatic scaling for live use cases
Patent Information
- Application Number
- JP2023514502
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-03
- Filing Date
- 2021-09-02
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2041-09-02
AI Technical Summary
Existing speech-to-text systems face challenges in maintaining low latency during live interactions due to varying CPU requirements, which are often addressed by over-allocating computational resources, leading to inefficiencies.
A method that dynamically adjusts computational resources based on latency thresholds by calculating word delivery differences, using metrics servers to determine resource allocation adjustments, and implementing scale-up or scale-down operations to maintain optimal performance.
This approach ensures efficient resource allocation, reducing latency and improving user experience in real-time speech-to-text applications by dynamically matching computational resources with user demands.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background Art
[0001] The present invention generally relates to the field of computing, and more particularly to a speech-to-text auto-scaling system.
[0002] Speech-to-text relates to the field of writing words in text form based on spoken language. Actual use cases include, but are not limited to, closed-captioning, automated agents, and dictation. In these types of cases, a natural and smooth interaction between the user and the cloud-based speech-to-text service is desired. The live interaction between the user and the cloud-based speech-to-text service requires the fast and efficient delivery of recognition transcripts by an audio recognition engine. The backend server operating in a remote data center desirably has sufficient headroom, i.e., available computational resources, to accommodate the dynamic central processing unit (CPU) requirements of the audio recognition engine. Therefore, it may be essential to have a system that dynamically generates metrics correlated with the user experience and the computational resources required at a given point in time to allocate computational resources and maintain latency.
Summary of the Invention
Means for Solving the Problems
[0003] According to one embodiment, a method, computer system, and computer program product for speech-to-text automatic scaling of computing resources are provided. The embodiment may include calculating a difference for each word in a transcript. The calculated difference may be for each word in the transcript between a wall clock time and the time the word is delivered to a client. The embodiment may also include sending a plurality of such differences to a group of metrics servers, which are configured to collect the plurality of such differences. The embodiment may further include requesting the current values of the plurality of such differences from the group of metrics servers. The embodiment may also include determining whether the current values of the plurality of such differences exceed a predefined maximum latency threshold. The embodiment may further include adjusting the allocated computing resources based on the frequency of the current values of the differences exceeding the predefined maximum latency threshold. The embodiment may also include creating a histogram from the current values of the plurality of such differences. The embodiment may further include scaling up the allocated computing resources based on the percentage of data points that exceed the predefined maximum latency threshold, in response to creating the histogram.
[0004] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of its exemplary embodiments, which should be read in conjunction with the accompanying drawings. Various features of the drawings are not to scale, as they are made clear to facilitate understanding of the invention by those skilled in the art in conjunction with the detailed description of the invention. [Brief explanation of the drawing]
[0005] [Figure 1] Figure 1 shows an exemplary networked computer environment according to at least one embodiment. [Figure 2]Figure 2 shows an operational flowchart for allocating computational resources in speech-to-text auto-scaling processing according to at least one embodiment. [Figure 3] Figure 3 is a functional block diagram of a software build quality evaluation process according to at least one embodiment. [Figure 4] Figure 4 illustrates a cloud computing environment according to one embodiment of the present invention. [Figure 5] Figure 5 illustrates an abstraction model layer according to one embodiment of the present invention. [Modes for carrying out the invention]
[0006] Detailed embodiments of the claimed structure and method are disclosed herein. However, it should be understood that the disclosed embodiments are merely illustrative of the claimed structure and method, which may be embodied in various forms. Nevertheless, the present invention may be embodied in many different forms and should not be construed as being limited to the exemplary embodiments described herein. In the description, well-known features and technical details may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0007] Embodiments of the present invention relate to the field of computing, and more particularly to speech-to-text auto-scaling systems. The exemplary embodiments described below, among other things, provide systems, methods, and program products for dynamically generating metrics correlated between user experience and the computing resources required at any given time, and, accordingly, allocating the computing resources necessary to keep latency within acceptable limits, utilizing speech-to-text and natural language processing (NLP). It should be understood that "computing resources" and "backend servers" are used interchangeably herein. Therefore, these embodiments have the ability to improve the technical field of speech-to-text auto-scaling by enabling natural and smooth interaction between users and cloud-based speech-to-text services.
[0008] As previously mentioned, speech-to-text concerns the field of transcribing spoken language into text. Real-time assurance, particularly user-perceived latency, is central to the user experience in all speech-to-text use cases involving raw interaction between the user and the cloud-based speech-to-text service. Actual use cases include, but are not limited to, closed captioning, automated agents, and dictation. In these types of cases, the recognized transcript needs to be delivered by the speech recognition engine with minimal delay. However, several factors contribute to latency degradation between the audio source and the speech recognition engine running in a remote data center, such as the lack of headroom in the computing resources available to the backend server to handle the speech recognition engine's dynamic central processing unit (CPU) requirements. This problem is typically addressed by over-allocating computing resources to cover the worst-case scenario of difficulty in processing the audio stream. The amount of computing resources, typically CPU, required by the speech recognition engine to generate a stream of transcripts in real time with low latency varies greatly depending on the audio stream being processed by the backend server. Factors related to the amount of computing resources required may include, but are not limited to, speaker characteristics, style, crosstalk, silent areas, background noise, traffic patterns, and co-located workloads, i.e., other applications or other instances of the speech recognition engine running on the same host machine that are competing for computing resources. Thus, it may be advantageous to proactively take measures to address these factors, among other things, by automatically scaling up or down the number of backend servers to match the available computing resources with the traffic coming from users. Computing resources can then be allocated on a need-by-need basis, and latency and cost can be kept within acceptable limits.
[0009] In at least one embodiment, a difference may be calculated for each word in the recognition transcript. The calculated difference may be the time difference between the wall clock time when the audio for the word is sent to the speech recognition engine and the time when the word is delivered to the client, i.e., the application layer. The difference may be sent to a group of metrics servers configured to collect the difference. A group of horizontal auto-scalers may periodically request the current value of the difference from the metrics servers. In at least one embodiment, if the current value of multiple differences exceeds a predefined max-latency threshold, the group of horizontal auto-scalers may trigger a scale-up operation and add more computing resources to the speech recognition engine running in a remote data center. The computing resources may be added incrementally until the current value of the difference falls within an acceptable latency range. In at least one other embodiment, if the current value of the difference falls below a predefined minimum latency threshold, the group of horizontal autoscalers may trigger a scale-down operation and reduce the computing resources allocated to the speech recognition engine running in the remote data center. Similarly, the computing resources may be gradually reduced until the current value of the difference falls within an acceptable latency range.
[0010] The present invention may be a system, method, or computer program product or a combination thereof at any possible level of technical detail that can be integrated. The computer program product may encompass one or more computer-readable storage media having computer-readable program instructions for causing a processor to execute the aspects of the present invention.
[0011] The computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of the computer-readable storage medium includes: portable computer diskettes®, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices, such as punch cards or grooved raised structures on which instructions are recorded, or any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted via wires.
[0012] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to individual computing devices / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may consist of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing device / processing device receives computer-readable program instructions from the network and transfers them for storage in a computer-readable storage medium within the individual computing device / processing device.
[0013] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, such as object-oriented programming languages, such as Smalltalk, C++, or conventional procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, partially as a standalone software package on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or such connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to carry out the aspects of the present invention.
[0014] The aspects of the present invention are described herein with reference to methods, apparatus (systems), and computer program products or computer program flowcharts or block diagrams or combinations thereof according to embodiments of the present invention. It will be understood that each block in the flowchart or block diagram or combination thereof, and combinations of blocks in the flowchart or block diagram or combination thereof, can be implemented by computer-readable program instructions.
[0015] These computer-readable program instructions may be provided to a computer processor or other programmable data processing device to create a machine, such that instructions executed via the processor of the computer or other programmable data processing device create means for implementing functions / operations specified in one or more blocks of the flowchart or block diagram or a combination thereof. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer-programmable data processing device or other device or a combination thereof to function in a particular manner, just as a product containing stored instructions includes instructions that implement the functional / operational aspects specified in one or more blocks of the flowchart or block diagram or a combination thereof.
[0016] The computer-readable program instructions may also be loaded onto the computer, other programmable data processing device, or other device such that instructions executed on the computer, other programmable data processing device, or other device implement the functions / operations specified in one or more blocks of the flowchart or block diagram or combination thereof, thereby causing a series of operations on the computer, other programmable device, or other device to generate a computer-implemented process.
[0017] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of systems, methods, and computer program products or possible implementations of computer programs according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or part thereof of instructions, which includes one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions shown in the block may occur in a different order than that shown in the drawings. For example, two consecutively shown blocks may actually be achieved as a single process executed simultaneously, substantially simultaneously, partially or entirely in a temporally overlapping manner, depending on the functions involved, or the blocks may be executed in reverse order. Note that each block in the block diagram or flowchart or a combination thereof, and any combination of multiple blocks in the block diagram or flowchart or a combination thereof, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or by a combination of special-purpose hardware and computer instructions.
[0018] The exemplary embodiments described below provide a system, method, and program product for determining in real time whether the allocated computing resources are sufficient to maintain acceptable latency or whether adjustments are necessary.
[0019] Referring to Figure 1, an exemplary networked computer environment 100 is depicted according to at least one embodiment. The networked computer environment 100 may comprise client computing devices 102 and servers 112 interconnected via a communication network 114. According to at least one implementation, the networked computer environment 100 may comprise multiple client computing devices 102 and servers 112, only one of which is shown for illustrative purposes.
[0020] The communication network 114 may comprise various types of communication networks, such as wide area networks (WANs), local area networks (LANs), communication networks, wireless networks, public switching networks, or satellite networks, or combinations thereof. The communication network 114 may also comprise connections, such as wired communication links, wireless communication links, or fiber optic cables. It should be understood that Figure 1 provides only an example of one implementation and does not imply any limitations regarding the environment in which different embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0021] The client computing device 102 may comprise a processor 104 and a data storage device 106 capable of hosting and executing a software program 108 and a speech-to-text auto-scaling program 110A, and communicating with a server 112 via a communication network 114, according to one embodiment of the present invention. The client computing device 102 may be, for example, a mobile device, a telephone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running a program and accessing a network. As will be illustrated with reference to Figure 3, the client computing device 102 may comprise internal components 302a and external components 304a, respectively.
[0022] The server computer 112 may be a laptop computer, a netbook computer, a personal computer (PC), a desktop computer, or any programmable electronic device or any network of programmable electronic devices that can host and run the speech-to-text auto-scaling program 110B and database 116, and communicate with client computing devices 102 via the communication network 114, according to embodiments of the present invention. As will be illustrated with reference to Figure 3, the server computer 112 may comprise internal components 302b and external components 304b, respectively. The server 112 may also operate in a cloud computing service model, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). The server 112 may also be deployed within a cloud computing deployment model, such as a private cloud, community cloud, public cloud, or hybrid cloud.
[0023] According to this embodiment, speech-to-text auto-scaling programs 110A and 110B may be programs that calculate, for each word in a transcript, the difference between the wall clock time and the time the word is delivered to the client, i.e., the application layer; send a plurality of such differences to one group of metrics servers configured to collect a plurality of such differences; request the current values of the plurality of such differences from the metrics servers; determine whether the current values of the plurality of such differences exceed a predefined threshold; and adjust the allocation of computational resources based on the frequency of the current values of the differences exceeding the predefined threshold. The speech-to-text auto-scaling method is described in more detail below in relation to Figure 2.
[0024] Referring now to Figure 2, an operation flowchart for allocating computing resources in a speech-to-text automatic scaling process 200 according to at least one embodiment is depicted. In step 202, speech-to-text automatic scaling programs 110A and 110B calculate a difference for each word in the recognition transcript. The calculated difference can be the time difference between the wall clock time when the audio of the word is sent to the speech recognition engine and the time when the word is delivered to the client. It may be considered that the audio is sent to the speech recognition engine at the moment the word is spoken by the user. Similarly, it may be considered that the audio is delivered to the client at the moment the word is transcribed. For example, if a word is spoken and timestamped at 1:00:00 PM, and the same word is transcribed and timestamped at 1:00:12 PM, the calculated difference can be determined to be 12 seconds. The calculated difference can be an end-to-end latency measurement value representing the quality of service from the perspective of the latency experienced by the user for a particular audio session.
[0025] In at least one other embodiment, a histogram can be created from the current values of a plurality of the differences to obtain a global estimate of the latency. Accordingly, the scaling up or down of the computing resources can be based on the percentage of data points that exceed the predefined maximum latency threshold. For example, if 5% of the calculated differences are greater than X seconds, speech-to-text automatic scaling programs 110A and 110B can scale up the number of backend servers by Y%. The predefined maximum latency threshold and the predefined minimum latency threshold can be changed according to user cases with different sensitivities to latency. For example, in an audio session with crosstalk and background noise, the predefined maximum latency threshold can be lowered to allocate more computing resources for that particular session.
[0026] Next, in step 204, the speech-to-text automatic scaling programs 110A and 110B transmit the current values of the plurality of differences to one group of the metrics servers configured to collect the plurality of differences. The current values of the plurality of differences may be transmitted via inter-process communication, where the speech-to-text automatic scaling programs 110A and 110B send a request, and the one group of the metrics servers responds to the request. For example, the speech-to-text automatic scaling programs 110A and 110B may transmit the current values of the plurality of differences to the metrics servers, and the metrics servers may accept the transmission of these current values.
[0027] Next, in step 206, the speech-to-text automatic scaling programs 110A and 110B cause the group of horizontal auto scalers to periodically request the current values of the plurality of differences from the metrics servers. For example, the request may be made every 10 seconds. According to at least one other embodiment, the timing of the periodic request may be changed. For example, the group of horizontal auto scalers may make a request every 8 seconds or every 12 seconds. The group of horizontal auto scalers may utilize a metrics application programming interface (API) to obtain the current values of the plurality of differences from the metrics servers. The metrics API may be a software intermediary that sends requests similar to the inter-process communication described above and delivers responses. For example, the group of horizontal auto scalers may request the current values of the plurality of differences from the metrics servers every 10 seconds, and the API may send the current values back to the group of horizontal auto scalers.
[0028] In at least one other embodiment, the metrics server may be configured to send the current values of the multiple differences to the group of horizontal autoscalers without a request from the group of horizontal autoscalers. The group of metrics servers may use the metrics API described above to send the current values of the multiple differences to the group of horizontal autoscalers. For example, the metrics server may send the current values of the multiple differences to the group of horizontal autoscalers, and the group of horizontal autoscalers may accept the transmission of the current values of the multiple differences.
[0029] Next, in step 208, the speech-to-text auto-scaling programs 110A and 110B determine whether the current value of the difference exceeds a predefined maximum latency (max-latency) threshold. This determination may be made based on an analysis of whether some differences exceed the predefined maximum latency threshold and some differences fall below the predefined minimum latency (min-latency) threshold. If the speech-to-text auto-scaling programs 110A and 110B determine that the current value of the difference exceeds the predefined maximum latency threshold, the speech-to-text auto-scaling process proceeds to step 210, where it adjusts the allocation of computing resources based on the frequency of the current value of the difference exceeding the predefined maximum latency threshold.
[0030] In at least one other embodiment, in step 209, the speech-to-text auto-scaling programs 110A and 110B determine whether the current value of the difference is below the predefined minimum latency threshold. If the speech-to-text auto-scaling programs 110A and 110B determine that the current value of the difference is below the predefined minimum latency threshold, the speech-to-text auto-scaling process proceeds to step 212, where it adjusts the allocation of computing resources based on the frequency of the current value of the difference being below the predefined minimum latency threshold.
[0031] Next, in step 210, the speech-to-text autoscaling programs 110A and 110B adjust the allocation of computing resources based on the frequency with which the current value of the difference exceeds the predefined maximum latency threshold. For example, if the predefined maximum latency threshold is exceeded more than Y times in at least Z seconds, the group of horizontal autoscalers may trigger a scale-up operation and add more computing resources to the speech recognition engine running in the remote data center. A default number of computing resources may be initially allocated to the speech recognition engine before the current value of the difference becomes available. For example, at least one computing resource may be allocated to the speech recognition engine at the start of an audio session. The group of horizontal autoscalers may trigger a scale-up operation by sending a signal to a monitoring circuit inside the backend server to start the backend server. The computing resources may be added incrementally until the current value of the difference falls within an acceptable latency range.
[0032] In at least one other embodiment, the maximum number of computing resources may be allocated to the speech recognition engine to prevent over-allocation of computing resources to any single audio session. The speech-to-text auto-scaling programs 110A and 110B may be configured to throttle new audio session requests from users in order to ensure the quality of service of already established audio sessions. For example, if 90% of the available computing resources are currently allocated to all combined audio sessions, the speech-to-text auto-scaling programs 110A and 110B may throttle new audio session requests.
[0033] In at least one other embodiment, in step 212, the speech-to-text autoscaling programs 110A and 110B adjust the allocation of computing resources based on the frequency with which the current value of the difference falls below a predefined min-latency threshold. For example, if the predefined min-latency threshold does not exceed Y times in at least Z seconds, the group of horizontal autoscalers may trigger a scale-down operation and reduce the computing resources allocated to the speech recognition engine running in the remote data center. The group of horizontal autoscalers may also trigger a scale-down operation by sending a signal to a monitoring circuit within the backend server to power down the backend server. Similarly, the computing resources may be gradually reduced until the current value of the difference falls within an acceptable range for latency.
[0034] In at least one other embodiment, a minimum number of computing resources may be allocated to the speech recognition engine so as not to significantly affect the quality of service of the audio session. For example, at least 10% of the available computing resources may be allocated to the speech recognition engine at any given time. Therefore, even if the current value of the difference is within acceptable latency limits, a large number of computing resources may be allocated to each audio session to ensure quality of service.
[0035] It should be understood that Figure 2 provides only an example of one implementation and does not imply any limitations on how different embodiments may be implemented. Many changes to the illustrated environment may be made based on design and implementation requirements.
[0036] Figure 3 is a block diagram 300 of the internal and external components of the client computing device 102 and server 112 depicted in Figure 1, according to an embodiment of the present invention. It will be understood that Figure 3 provides only an example of one implementation and does not imply any limitations with respect to the environment in which different embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0037] Data processing systems 302 and 304 represent any electronic device capable of executing machine-readable program instructions. Data processing systems 302 and 304 may represent smartphones, computer systems, PDAs, or other electronic devices. Examples of computing systems, environments, configurations, or combinations thereof that may be represented by data processing systems 302 and 304 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputer systems, and distributed cloud computing environments comprising any of the above systems or devices.
[0038] The client computing device 102 and the server 112 may comprise individual sets of internal components 302a and 302b and external components 304a and 304b, as illustrated in Figure 3. Each of the multiple sets of internal components 302 comprises one or more processors 320 on one or more buses 326, one or more computer-readable RAMs 322, one or more computer-readable ROMs 324, one or more operating systems 328, and one or more computer-readable tangible storage devices 330. One or more operating systems 328, software programs 108, and speech-to-text auto-scaling programs 110A in the client computing device 102, and speech-to-text auto-scaling programs 110B in the server 112, are stored in one or more individual computer-readable tangible storage devices 330 for execution by one or more of the individual processors 320 via one or more of the individual RAMs 322 (which typically have cache memory). In the embodiment shown in Figure 3, each of the computer-readable tangible memory devices 330 is a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible memory devices 330 is a semiconductor storage device, such as a ROM 324, EPROM, flash memory, or any other computer-readable tangible memory device capable of storing computer programs and digital information.
[0039] Each set of internal components 302a and 302b also includes an R / W drive or interface 332 for reading from and writing to one or more portable computer-readable tangible memory devices 338, such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor memory devices. Software programs, such as speech-to-text auto-scaling programs 110A and 110B, can be stored in one or more of the individual portable computer-readable tangible memory devices 338, can be read via the individual R / W drives or interfaces 332, and can be loaded into individual hard drives 330.
[0040] Each set of internal components 302a and 302b also includes a network adapter or interface 336, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G or 4G wireless interface card, or other wired or wireless communication link. The software program 108 and speech-to-text auto-scaling program 110A in the client computing device 102, and the speech-to-text auto-scaling program 110B in the server 112, can be downloaded from an external computer to the client computing device 102 and server 112 via a network (e.g., the Internet, a local area network, or other wide area network) and individual network adapters or interfaces 336. From the network adapter or interface 336, the software program 108 and speech-to-text auto-scaling program 110A in the client computing device 102, and the speech-to-text auto-scaling program 110B in the server 112, are loaded into individual hard drives 330. The network may include copper, optical fiber, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof.
[0041] Each of the multiple sets of external components 304a and 304b may include a computer display monitor 344, a keyboard 342, and a computer mouse 334. External components 304a and 304b may also include a touchscreen, a virtual keyboard, a touchpad, a pointing device, and other human interface devices. Each of the multiple sets of internal components 302a and 302b also includes a device driver 340 for interface with the computer display monitor 344, the keyboard 342, and the computer mouse 334. The device driver 340, the R / W drive or interface 332, and the network adapter or interface 336 include hardware and software (stored in the storage device 330 or ROM 324 or a combination thereof).
[0042] While this disclosure includes a detailed description of cloud computing, it should be understood in advance that the implementation of the teachings enumerated herein is not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in combination with any other type of computing environment that is currently known or may be developed in the future.
[0043] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0044] The features are as follows:
[0045] On-demand self-service: Cloud consumers can unilaterally provision computing functions, such as server time and network storage, as needed, without requiring human interaction with the service provider.
[0046] Broad network access: The functionality is available over a network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0047] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically allocated and reallocated according to demand. Consumers generally do not have control or knowledge of the exact location of the resources provided, but can identify the location at a higher level of abstraction (e.g., country, state, or data center), making it location-independent.
[0048] Rapid Adaptability: Features can be provisioned quickly and flexibly, and in some cases automatically, they can scale out quickly, be released quickly, and scale in quickly. For consumers, the features available for provisioning are often unlimited and can be purchased at any amount at any time.
[0049] Measured Services: Cloud systems automatically control and optimize resource usage by employing metric functions at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0050] The service model is as follows:
[0051] Software as a Service (SaaS): This refers to the functionality provided to consumers for using a provider's applications running on a cloud infrastructure. These applications are accessible from various client devices through a thin client interface, such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating system, storage, or even the underlying cloud infrastructure encompassing individual application functions, with the possible exception of limited, user-specific application configuration settings.
[0052] Platform as a Service (PaaS): A service provided to a consumer to deploy applications they have created or acquired, generated using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating system, or storage, but has control over the deployed applications and, in some cases, the application hosting environment configuration.
[0053] Infrastructure as a Service (IaaS): This is a service provided to a consumer to provision processing, storage, networking, and other basic computing resources, enabling the consumer to deploy and run any software, including operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has limited control over the operating system, storage, deployed applications, and, in some cases, network components (e.g., the host's firewall).
[0054] The deployment models are as follows:
[0055] Private Cloud: A cloud infrastructure is operated solely for a specific organization. This cloud infrastructure may be managed by that organization or a third party, and may reside on-premises or off-premises.
[0056] Community Cloud: Cloud infrastructure is shared by several organizations and supports a specific community that shares common interests (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organization or a third party and may reside on-premises or off-premises.
[0057] Public cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services.
[0058] Hybrid Cloud: Cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain separate entities but are brought together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application migration.
[0059] Cloud computing environments are oriented services that focus on statelessness, low coupling, modularity, and semantic interoperability. The heart of cloud computing is the infrastructure, which includes a network of interconnected nodes.
[0060] Referring now to Figure 4, an exemplary cloud computing environment 40 is illustrated. As shown, the cloud computing environment 40 comprises one or more cloud computing nodes 100 that can communicate with local computing devices used by cloud consumers, such as personal digital assistants (PDAs) or mobile phones 44A, desktop computers 44B, laptop computers 44C, or automotive computer systems 44N, or a combination thereof. The nodes 100 can communicate with each other. They may be physically or virtually grouped in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or a combination thereof, as described herein (not shown). This allows the cloud computing environment 40 to provide infrastructure, platforms, or software, or a combination thereof, as a service that does not require cloud consumers to maintain resources on their local computing devices. It is understood that the types of computing devices 44A-N shown are for illustrative purposes only, and that the computing node 100 and the cloud computing environment 40 may communicate with any type of computerized device via any type of network or network addressable connection or a combination thereof (e.g., using a web browser).
[0061] Referring here to Figure 5, a set of functional abstraction layers 500 provided by the cloud computing environment 40 is shown. It should be understood that the components, layers, and functions shown in Figure 5 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. As illustrated, several layers and corresponding functions are provided below.
[0062] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based servers 62; servers 63; blade servers 64; storage devices 65; and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0063] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities may be provided: namely, a virtual server 71; virtual storage 72; a virtual network 73; such a virtual network 73 encompassing, for example, a virtual private network; a virtual application and operating system 74; and a virtual client 75.
[0064] In one example, the management layer 80 may provide several functions as described below: Resource provisioning 81 provides the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 82 provides cost tracking when resources are used within the cloud computing environment and billing or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identification and verification for cloud consumers and tasks, and protection for data and other resources. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides the allocation and management of cloud computing resources to ensure that the required service levels are met. Service level agreement (SLA) planning and execution 85 provides the pre-placement and procurement of cloud computing resources for which future requirements are anticipated in accordance with the SLA.
[0065] The workload layer 90 provides examples of several functions that the cloud computing environment may utilize. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91; software development and lifecycle management 92; provision of virtual classroom education 93; data analysis processing 94; transaction processing 95; and speech-to-text autoscaling using natural language description 96. Speech-to-text autoscaling using natural language description 96 may involve dynamically generating metrics that correlate with the user experience and the computing resources required at any given time, and allocating the computing resources necessary to keep latency within acceptable limits.
[0066] The various embodiments described herein are presented for illustrative purposes only and are not intended to be exhaustive or to limit oneself to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described herein. The terms used herein have been selected to best describe the principles, practical applications, or technical improvements to the technologies available on the market of the embodiments, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-based method, comprising: For each word in the transcript, calculating a difference between the time when the audio corresponding to the word is sent to the speech recognition engine that generates the transcript (hereinafter referred to as the wall clock time) and the time when the transcript of the audio corresponding to the word is delivered to the client; Sending a plurality of said differences to one group of metrics servers, where said one group of metrics servers is configured to collect a plurality of said differences; Requesting the current values of a plurality of said differences from said one group of metrics servers; Determining whether the current values of a plurality of said differences exceed a predefined maximum latency threshold; Increasing the computing resources allocated to the speech recognition engine in response to a determination that the current value of said difference exceeds the predefined maximum latency threshold and the current value of said difference exceeds the maximum latency threshold Y times within Z seconds; The method as described above.
2. The method according to claim 1, wherein the predefined maximum latency threshold is configured to be changed according to user cases having various sensitivities to latency.
3. The method according to claim 1, wherein the allocated computing resources are incrementally added until the current value of said difference falls within the acceptable range for latency.
4. Determining whether the current values of a plurality of said differences are below a predefined minimum latency threshold; Reducing the computing resources allocated to the speech recognition engine in response to a determination that the current value of said difference is below the predefined minimum latency threshold and a determination that the current value of said difference does not exceed the predefined minimum latency threshold Y times within Z seconds; The method according to claim 1, further comprising the above.
5. The method according to claim 4, wherein the predefined minimum latency threshold is configured to be changed according to user cases having various sensitivities to latency.
6. The method according to claim 1, wherein the allocated computing resources are gradually reduced until the current value of said difference falls within the acceptable range for latency.
7. Calculating said difference is Creating a histogram from the current values of the plurality of the differences including wherein the method determining a percentage of data points in the histogram that exceed the predefined maximum latency threshold further including increasing the allocated computing resources being executed when the percentage of the data points exceeding the maximum latency threshold is greater than a predetermined value The method according to claim 1
8. A computer system comprising one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and a plurality of program instructions stored on at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories wherein the computer system for each word in the transcript, calculating a difference between the time (hereinafter referred to as wall clock time) when the audio corresponding to the word is sent to the speech recognition engine that generates the transcript and the time when the transcript of the audio corresponding to the word is delivered to the client sending the plurality of the differences to a group of one of the metrics servers, where the group of one of the metrics servers is configured to collect the plurality of the differences requesting current values of the plurality of the differences from the group of one of the metrics servers determining whether the current values of the plurality of the differences exceed a predefined maximum latency threshold increasing the computing resources allocated to the speech recognition engine in response to a determination that the current values of the differences exceed the predefined maximum latency threshold and the current values of the differences exceeding the maximum latency threshold Y times within Z seconds The computer system capable of executing a method including
9. The computer system according to claim 8, wherein the predefined maximum latency threshold is configured to be changed according to user cases having various sensitivities to latency
10. The computer system according to claim 8, wherein the allocated computing resources are gradually added until the current value of the difference falls within an acceptable range for latency.
11. Determining whether the current values of a plurality of the differences are less than a predefined minimum latency threshold; Reducing the computing resources allocated to the speech recognition engine in response to a determination that the current value of the difference is less than the predefined minimum latency threshold and a determination that the current value of the difference does not exceed the predefined minimum latency threshold more than Y times within Z seconds. The computer system according to claim 8, further comprising.
13. The computer system according to claim 8, wherein the allocated computing resources are gradually reduced until the current value of the difference falls within an acceptable range for latency.
14. Calculating the difference comprises Creating a histogram from the current values of a plurality of the differences including The method is Determining a percentage of data points in the histogram that exceed a predefined maximum latency threshold further including Increasing the allocated computing resources is performed when the percentage of the data points exceeding the maximum latency threshold is greater than a predetermined value. The computer system according to claim 8.
15. A computer program, For each word in the transcript, calculating the difference between the time when the audio corresponding to the word was sent to the speech recognition engine that generated the transcript (hereinafter referred to as the wall clock time) and the time when the transcript of the audio corresponding to the word was delivered to the client; Sending a plurality of the differences to one group of metrics servers, where the one group of metrics servers is configured to collect a plurality of the differences; Requesting the current values of a plurality of the differences from the one group of metrics servers; Determining whether the current values of a plurality of the differences exceed a predefined maximum latency threshold; Determining that the current value of the difference exceeds the pre-defined maximum latency threshold and increasing the computing resources allocated to the speech recognition engine in response to the current value of the difference exceeding the maximum latency threshold Y times within Z seconds The computer program that causes a processor to execute each step of the method including this **Claim 16** The computer program according to claim 15, wherein the pre-defined maximum latency threshold is configured to be changed according to user cases having various sensitivities to latency. **Claim 17** The computer program according to claim 15, wherein the allocated computing resources are gradually added until the current value of the difference falls within the allowable range for latency. **Claim 18** Determining whether the current values of a plurality of the differences are below a pre-defined minimum latency threshold; Determining that the current value of the difference is below the pre-defined minimum latency threshold and reducing the computing resources allocated to the speech recognition engine in response to the determination that the current value of the difference does not exceed the pre-defined minimum latency threshold more than Y times within Z seconds The computer program according to claim 15, further including this. **Claim 19** The computer program according to claim 18, wherein the pre-defined minimum latency threshold is configured to be changed according to user cases having various sensitivities to latency. **Claim 20** The computer program according to claim 15, wherein the allocated computing resources are gradually reduced until the current value of the difference falls within the allowable range for latency.