Real-time vs. non-real-time audio streaming

By determining the bitrate classification of audio streams, the method effectively addresses the challenge of differentiating between real-time and non-real-time audio streams, enhancing the accuracy and cost-effectiveness of speech-to-text conversion.

JP7695019B2Active Publication Date: 2025-06-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023515622
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-22
Filing Date
2021-07-22
Publication Date
2025-06-18
Estimated Expiration
2041-07-22

AI Technical Summary

Technical Problem

Existing audio streaming technologies struggle to differentiate between real-time and non-real-time audio streams, which affects the accuracy and cost-effectiveness of speech-to-text conversion.

Method used

A computer-implemented method that determines the expected and input bitrates of audio data, calculates an R value, and compares it to a threshold to classify the audio stream as real-time or non-real-time, thereby enabling more accurate and cost-effective speech-to-text processing.

Benefits of technology

This approach allows for more accurate speech-to-text conversion by applying suitable techniques based on the classification of audio streams as real-time or non-real-time, thereby reducing costs and improving service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695019000001
    Figure 0007695019000001
  • Figure 0007695019000002
    Figure 0007695019000002
  • Figure 0007695019000003
    Figure 0007695019000003
Patent Text Reader

Abstract

One or more audio data are received. An expected bit rate of the one or more audio data is determined. An input bit rate of the one or more audio data is determined. An R value is determined using the expected bit rate and the input bit rate. The R value is compared to an R threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of audio streaming, and more particularly to the stage of determining real-time versus non-real-time audio streaming for improving text conversion of audio.

[0002] Speech-to-text conversion transforms spoken language into text. Speech recognition develops methodologies and technologies that enable the recognition and conversion of spoken language by a computer into text. Speech recognition may also be known as automatic speech recognition (ASR), computer speech recognition, or speech-to-text (STT). Some speech recognition systems require "training" in which an individual speaker reads text or isolated vocabulary into the system. The system analyzes a particular voice of a person and uses it to fine-tune the recognition of that person's voice, resulting in increased accuracy. Some systems do not use training and are called "speaker-independent" systems.

Summary of the Invention

[0003] Embodiments of the present invention include a computer-implemented method, a computer program product, and a system for determining real-time versus non-real-time audio streaming. In one embodiment, one or more audio data are received. An expected bitrate of the one or more audio data is determined. An input bitrate of the one or more audio data is determined. An R value using the expected bitrate and the input bitrate is determined. The R value is compared with an R threshold.

Brief Description of the Drawings

[0004]

Figure 1

[0005]

Figure 2

[0006]

Figure 3

DETAILED DESCRIPTION OF THE INVENTION

[0007] The present invention provides a method, computer program product, and computer system for automatically distinguishing between real-time and non-real-time, such that the most suitable technique for voice text conversion can be applied in any case, resulting in a more accurate transcript and a lower overall cost. The detection is based on performing a comparison between the bitrate of the received incoming audio and the encoding bitrate of the audio.

[0008] Embodiments of the present invention recognize that not all audio is either real-time or non-real-time. Embodiments of the present invention recognize that it may be more costly to perform voice text conversion of real-time audio versus non-real-time audio.

[0009] Embodiments of the present invention recognize that there are benefits to knowing whether an audio stream must be processed in real time with low latency, from both the perspectives of service quality and cost. If it is known that an audio stream can be processed in batches (i.e., the recognition transcript does not need to be sent to the user in real time), it is possible to apply more sophisticated speech-to-text techniques that function from the start to the end of the utterance in audio, and vice versa (bidirectionality), which can result in higher accuracy. Also, since the end of the utterance is known in advance, it is possible to perform multi-passes over the same utterance in audio, which results in increased accuracy. Further, it is possible to utilize larger batches during inference, which increases the speed of the speech-to-text calculations since the low-latency constraint does not accommodate larger batch processing of audio. Additionally, if the audio stream can be processed in batches, a lower priority can be given to the audio stream first to align computing resources to latency-critical audio jobs.

[0010] Referring now more particularly to various embodiments of the invention, FIG. 1 is a functional block diagram of a network computing environment generally designated 100 that is suitable for operation of an audio program 112 according to at least one embodiment of the present invention. FIG. 1 provides only an illustration of one implementation and does not imply any limitation as to the environments in which different embodiments may be implemented. Many modifications to the illustrated environment may be made by those skilled in the art without departing from the scope of the present invention as recited in the claims.

[0011] The network computing environment 100 has computing devices 110 interconnected via a network 120. In an embodiment of the present invention, the network 120 may be a remote communication network, a local area network (LAN), a wide area network (WAN), such as the Internet, or a combination of the three, and may include wired, wireless, or fiber optic connections. The network 120 may have one or more wired or wireless networks, or both, capable of receiving and transmitting data, voice, or video signals, or a combination thereof, including multimedia signals including voice, data, and video forms. Generally, the network 120 may be any combination of connections and protocols that support communication between the computing devices 110 within the network computing environment 100 and other computing devices (not shown).

[0012] The computing device 110 is a computing device that can be a laptop computer, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a personal digital assistant (PDA), a smartphone, a smartwatch, or any programmable electronic device capable of receiving, transmitting, and processing data. Generally, the computing device 110 represents any programmable electronic device or combination of programmable electronic devices that can execute machine-readable program instructions to communicate with other computing devices (not shown) within the computing environment 100 via a network such as the network 120.

[0013] In various embodiments of the present invention, computing device 110 can be a stand-alone device, a management server, a web server, a media server, a mobile computing device, or any other programmable electronic device or computing system capable of receiving, transmitting, and processing data. In other embodiments, computing device 110 represents a server computing system that utilizes multiple computers as a server system, for example, in a cloud computing environment. In one embodiment, computing device 110 represents a computing system that utilizes clustered computers and components (e.g., database server computers, application server computers, web servers, and media servers) that act as a single pool of seamless resources when accessed within network computing environment 100.

[0014] In various embodiments of the present invention, computing device 110 includes audio program 112 and information repository 114.

[0015] In one embodiment, computing device 110 includes a user interface (not shown). The user interface is a program that provides an interface between the user and the application. The user interface refers to the information presented by the program to the user (e.g., graphics, text, and sound) and the control sequences used by the user to control the program. There are many types of user interfaces. In one embodiment, the user interface can be a graphical user interface (GUI). A GUI is a type of user interface that enables a user to interact with an electronic device such as a keyboard and mouse through graphical icons and visual indicators, e.g., as opposed to a text-based interface, typed command labels, or text navigation through secondary notation. In a computer, the GUI was introduced in response to the recognized steep learning curve of the command line interface, where commands had to be typed on the keyboard. Movement in a GUI is often performed through direct manipulation of graphical elements.

[0016] In one embodiment, computing device 110 includes an audio program 112. Embodiments of the present invention provide an audio program 112 that receives audio data. In embodiments of the present invention, audio program 112 determines an expected bitrate. In embodiments of the present invention, audio program 112 determines an input bitrate. In embodiments of the present invention, audio program 112 determines an R value. In embodiments of the present invention, audio program 112 determines whether the R value is greater than a threshold. In embodiments of the present invention, audio program 112 adds the audio to a real-time audio list. In embodiments of the present invention, audio program 112 adds the audio to a batch audio list. In embodiments of the present invention, audio program 112 waits for a time threshold.

[0017] In one embodiment, computing device 110 includes an information repository 114. In one embodiment, information repository 114 may be managed by audio program 112. In an alternative embodiment, information repository 114 may be managed by the operating system of computing device 110, by another program (not shown) alone, or in combination with audio program 112. Information repository 114 is a data repository that can store, collect, or analyze information, or a combination thereof. In some embodiments, information repository 114 is located external to computing device 110 and is accessed through a communication network such as network 120. In some embodiments, information repository 114 is stored on computing device 110. In some embodiments, information repository 114 may exist on another computing device (not shown) provided that information repository 114 can be accessed by computing device 110. Information repository 114 may include, but is not limited to, speech-to-text settings, R values, input bitrates, expected bitrates, lists of batch audio, lists of real-time audio, R thresholds, time thresholds, and the like.

[0018] As is known in the art, information repository 114 may be implemented using any volatile or non-volatile storage medium for storing information. For example, information repository 114 may be implemented in a tape library, an optical library, one or more independent hard disk drives, multiple hard disk drives in a redundant array of independent disks (RAID), a solid state drive (SSD), or random access memory (RAM). Similarly, information repository 114 may be implemented in any suitable storage architecture known in the art, such as a relational database, an object-oriented database, or one or more tables.

[0019] As mentioned in this specification, all data obtained, collected, and used is used in an opt-in manner, that is, the data provider has given permission for the data to be used. For example, the received data is received and used by an audio program 112 that determines real-time versus non-real-time audio streaming.

[0020] FIG. 2 is a flowchart diagram of a workflow 200 representing the operational steps related to an audio program 112 according to at least one embodiment of the present invention. In alternative embodiments, the steps of the workflow 200 may be performed by any other program in conjunction with the audio program 112. It should be understood that embodiments of the present invention provide steps for at least determining real-time versus non-real-time audio streaming. However, FIG. 2 provides merely an example of one implementation and does not suggest any limitation with respect to the environments in which different embodiments may be implemented. Many modifications to the illustrated environment may be made by those skilled in the art without departing from the scope of the present invention as recited by the claims. In a preferred embodiment, if a user desires to have an audio program 112 determine real-time versus non-real-time audio streaming, the user can invoke the workflow 200 via a user interface (not shown).

[0021] The audio program 112 receives audio data (step 202). In step 202, the audio program 112 receives audio data via the network 120. In one embodiment, the audio data may not be classified as real-time vs. non-real-time audio. In an alternative embodiment, the audio program 112 receives audio data stored in the information repository 114 via a display from the user via the user interface. In one embodiment, the audio program 112 may receive one piece of audio data. In an alternative embodiment, the audio program 112 may receive one or more pieces of audio data in a continuous flow. In one embodiment, the audio program 112 may receive audio associated with corresponding video data. In one embodiment, the audio program 112 may analyze a first chunk of audio data received within a threshold time. In other words, the audio program 112 starts receiving audio data, waits for a threshold time (i.e., 3 seconds), and can then proceed to the next step regardless of whether all the audio data has been received. In one embodiment, the received audio data may include a header indicating whether the audio data is either real-time or non-real-time audio. Here, the process can proceed directly to step 212 for real-time audio and to step 214 for non-real-time audio. In one embodiment, the received audio may be a request to convert speech (audio) to text. In one embodiment, the request may be an HTTP / websocket request or any other known network protocol known in the art.

[0022] The audio program 112 determines the expected bitrate (step 204). In step 204, the audio program 112 determines the expected bitrate of the received audio data. In one embodiment, the expected bitrate is determined from the header of the received audio data, more specifically from the metadata found within the header. In one embodiment, the expected bitrate is listed in the metadata of the header. Alternatively, in one embodiment, the sampling rate and sampling size are listed in the metadata of the header, and the expected bitrate is calculated by multiplying the sampling rate by the sample size. In an alternative embodiment, if the audio data does not include a header, the expected bitrate is calculated based on the sampling rate value and bytes per sample value found in the content type header within the received audio.

[0023] The audio program 112 determines the input bitrate (step 206). In step 206, the audio program 112 determines the input bitrate of the received audio data. In one embodiment, the input bitrate is determined based on how fast the bytes of the audio are received by the computing device 110 via the network 120. In one embodiment, the audio program 112 may receive this information from the operating system of the computing device 110 or any other program functioning on the computing device 110.

[0024] The audio program 112 determines the R value (decision step 208). In step 208, the audio program 112 determines the R value for the received audio data. In one embodiment, the R value is calculated by dividing the input bitrate by the expected bitrate. In one embodiment, the R value is stored in the information repository 114.

[0025] The audio program 112 determines whether R is greater than the R threshold (decision step 210). In decision step 210, the audio program 112 compares the R value with the R threshold. If the R value is not greater than the R threshold (decision step 210, no branch), the process proceeds to real-time audio (step 212). If R is greater than the R threshold (decision step 210, yes branch), the process proceeds to batch audio (step 214).

[0026] The audio program 112 determines real-time audio (step 212). In step 212, the audio program 112 determines that the received audio data is real-time audio. In one embodiment, the audio program 112 may process the received audio using techniques known in the art for processing real-time audio. In one embodiment, the audio program 112 may add a queue for processing to the received audio data by a real-time audio list or another program not shown. In one embodiment, the audio on the real-time audio list is processed with a higher priority than the audio on the batch audio list.

[0027] The audio program 112 determines batch audio (step 214). In step 214, the audio program 112 determines that the received audio data is batch audio. In one embodiment, the audio program 112 may process the received audio using techniques known in the art for processing batch audio. In one embodiment, the audio program 112 may add a queue for processing to the received audio data by a batch audio list or another program not shown. In one embodiment, the audio on the batch audio list is processed with a lower priority than the audio on the real-time audio list.

[0028] The audio program 112 waits for a time threshold (stage 216). In stage 216, after the audio program 112 waits for the time threshold, it proceeds to stage 206. In one embodiment, the audio program 112 may receive the time threshold when receiving audio data from the user via the user interface. In an alternative embodiment, the time threshold may be stored in the information repository 114 as a user preference. In yet another alternative embodiment, the time threshold may vary based on the processing time or date, or both, based on the preferences stored in the information repository 114. In one embodiment, the larger the time threshold, the better the reliability of the real-time versus batch audio detection. In an alternative embodiment, the smaller the time threshold, the faster the real-time versus batch audio detection. By waiting for the time threshold, it becomes possible to classify the audio file as real-time, and subsequently, if more time passes, the audio file can be changed to be classified as batch audio. This allows for a change in the speed of receiving the audio file due to network 120 issues such as bandwidth and latency. In one embodiment, the time threshold may be based on the time of day, the day of the week, or the calendar. For example, the time threshold may be larger on weekends (when less is being processed) and smaller during the week (when more is being processed).

[0029] Note that at any point when one audio data can end, for example, the audio data being processed has reached the end of the audio, and then a new piece of audio data can be inlined and processed by the audio program 112 next. In this embodiment, when a new piece of audio data is received, the processing starts at stage 202.

[0030] In one example, twelve MB audio files are received by computing device 110 for processing by audio program 112. Here, the files contain ten minutes of audio. In a first example, the files are received by audio program 112 in just a few seconds and can be limited only by the bandwidth over network 120. Thus, the audio files are determined to be batch audio. However, in a second example, if the audio files come in at a rate closer to real-time audio, the audio files are determined to be real-time audio.

[0031] To provide a more detailed example, the received audio stream is encoded at 16 kHz and 2 bytes per sample, which corresponds to 32 KB / second. If the audio comes in at an average bitrate of 40 KB / second, the inventors know that the stream is not coming in real-time and is coming in much faster, and thus, the inventors have a batch usage case. However, if the bitrate of the input stream is approximately 32 KB / second + / - a small delta, the inventors know that the stream is likely to be real-time.

[0032] FIG. 3 is a block diagram representing components of a computer 300 suitable for audio program 112 according to at least one embodiment of the present invention. FIG. 3 shows computer 300, one or more processors 304 (including one or more computer processors), communication fabric 302, memory 306 including RAM 316 and cache 318, persistent storage 308, communication unit 312, I / O interface 314, display 322, and external device 320. It should be understood that FIG. 3 provides only an illustration of one embodiment and does not imply any limitation regarding the environments in which different embodiments may be implemented. Many modifications may be made to the illustrated environment.

[0033] As shown, computer 300 operates across communication fabric 302, which provides communication between communication computer processor 304, memory 306, persistent storage 308, communication unit 312, and input / output (I / O) interface 314. Communication fabric 302 can be implemented in an architecture suitable for passing data or control information between processor 304 (e.g., a microprocessor, communication processor, and network processor), memory 306, external device 320, and any other hardware components within the system. For example, communication fabric 302 can be implemented with one or more buses.

[0034] Memory 306 and persistent storage 308 are computer-readable storage media. In the illustrated embodiment, memory 306 includes random access memory (RAM) 316 and cache 318. Generally, memory 306 can include any suitable one or more volatile or non-volatile computer-readable storage media.

[0035] Program instructions for audio program 112 can be stored in persistent storage 308, or more generally, in any computer-readable storage media, via one or more memories of memory 306, for execution by one or more of the respective computer processors 304. Persistent storage 408 can be a magnetic hard disk drive, solid state disk drive, semiconductor storage device, read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.

[0036] The medium used by the persistent storage 308 can also be removable. For example, a removable hard drive may be used for the persistent storage 308. Other examples include optical disks and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer onto another computer-readable storage medium that is also part of the persistent storage 308.

[0037] In these examples, the communication unit 312 provides communication with other data processing systems or devices. In these examples, the communication unit 312 can have one or more network interface cards. The communication unit 312 can provide communication using one or both of physical and wireless communication links. In the context of embodiments of the present invention, various sources of input data can be physically remote from the computer 300 such that input data is received and output is similarly transmitted via the communication unit 312.

[0038] The I / O interface 314 enables the input and output of data using other devices that can operate in conjunction with the computer 300. For example, the I / O interface 314 provides a connection to an external device 320, which can exist as a keyboard, keypad, touch screen, or other suitable input device. The external device 320 can also include a portable computer-readable storage medium, such as a thumb drive, portable optical disk or magnetic disk, and memory card. The software and data used to implement embodiments of the present invention may be stored on such a portable computer-readable storage medium and loaded onto the persistent storage 308 via the I / O interface 314. The I / O interface(s) 314 may similarly be connected to a display 322. The display 322 provides a mechanism for displaying data to the user and can be, for example, a computer monitor.

[0039] The present invention can be a system, method, or computer program product integrated at any possible level of technical detail, or a combination thereof. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0040] The computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, punch cards, or mechanically encoded devices such as raised structures in grooves recording instructions, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0041] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.

[0042] The computer-readable program instructions for carrying out the operations of the present invention may be written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages, in either source code or object code. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for implementing aspects of the present invention.

[0043] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0044] Computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, or other device to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions which implement the aspect of the function / act specified in one or more blocks of the flowchart and / or block diagram.

[0045] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0046] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of computer program instructions having one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions described in a block may occur in an order different from that depicted in the figures. For example, two blocks shown in succession may in fact be implemented as one step, executed simultaneously, substantially simultaneously, in a partially or wholly temporally overlapping manner, or, depending on the related functionality, the blocks may in some cases be executed in the reverse order. It should also be noted that each block of a block diagram or flowchart diagram, or combinations of blocks in a block diagram or flowchart diagram or both, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0047] The description of various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or to limit the invention to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application, or technical improvements made to the technology found in the marketplace, or to enable other practitioners in the art to understand the embodiments disclosed herein. [Item 1] A computer-implemented method for determining real-time versus non-real-time audio streaming, comprising: receiving, by one or more computer processors, one or more audio data; determining, by one or more computer processors, an expected bit rate of the one or more audio data; determining, by one or more computer processors, an input bit rate of the one or more audio data; calculating, by one or more computer processors, an R value using the expected bit rate and the input bit rate; comparing, by one or more computer processors, the R value with an R threshold A computer-implemented method comprising the above steps. [Item 2] In response to the R value being greater than the R threshold, adding, by one or more computer processors, the one or more audio data to a batch audio list The computer-implemented method according to Item 1, further comprising the above step. [Item 3] In response to the R value being less than the R threshold, adding, by one or more computer processors, the one or more audio data to a real-time audio list The computer-implemented method according to Item 1, further comprising the above step. [Item 4] The comparing step is performed after a threshold period. The computer-implemented method according to any one of Items 1 to 3. [Item 5] Processing, by one or more computer processors, the one or more audio data on the batch audio list using a lower priority than any one or more audio on the real-time audio list The computer-implemented method according to Item 2, further comprising the above step. [Item 6] Processing, by one or more computer processors, the one or more audio data on the real-time audio list using a higher priority than any one or more audio on the batch audio list The computer-implemented method according to Item 3, further comprising the above step. [Item 7] The R value is calculated by dividing the input bit rate by the expected bit rate. The computer-implemented method according to any one of Items 1 to 6. [Item 8] A computer program for determining real-time versus non-real-time audio streaming, the program causing one or more computer processors to receive one or more audio data; determine an expected bitrate of the one or more audio data; determine an input bitrate of the one or more audio data; calculate an R value using the expected bitrate and the input bitrate; compare the R value with an R threshold and execute the steps. [Item 9] The computer program according to item 8, further causing the one or more computer processors to add the one or more audio data to a batch audio list in response to the R value being greater than the R threshold. [Item 10] The computer program according to item 8, further causing the one or more computer processors to add the one or more audio data to a real-time audio list in response to the R value being less than the R threshold. [Item 11] The computer program according to item 8, wherein the comparing step is performed after a threshold period. [Item 12] The computer program according to item 9, causing the one or more computer processors to process the one or more audio data on the batch audio list using a lower priority than any one or more audio on the real-time audio list. The computer program according to item 9. [Item 13] The computer program according to item 10, causing the one or more computer processors to process the one or more audio data on the real-time audio list using a higher priority than any one or more audio on the batch audio list. [Item 14] The computer program according to item 8, wherein the R value is calculated by dividing the input bitrate by the expected bitrate. [Item 15] A computer system for determining real-time versus non-real-time audio streaming, the computer system comprising one or more computer processors; one or more computer-readable storage media; and Program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors comprising, the program instructions being program instructions for receiving one or more audio data; program instructions for determining an expected bitrate of the one or more audio data by one or more computer processors; program instructions for determining an input bitrate of the one or more audio data by one or more computer processors; program instructions for calculating an R value using the expected bitrate and the input bitrate by one or more computer processors; program instructions for comparing the R value with an R threshold by one or more computer processors of a computer system. [Item 16] The computer system according to item 15, further comprising program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions adding the one or more audio data to a batch audio list in response to the R value being greater than the R threshold. [Item 17] The computer system according to item 15, further comprising program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions adding the one or more audio data to a real-time audio list in response to the R value being less than the R threshold. [Item 18] The computer system according to item 15, wherein the comparison is performed after a threshold period. [Item 19] The computer system according to item 16, further comprising program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions processing the one or more audio data on the batch audio list using a lower priority than any one or more audio on the real-time audio list. [Item 20] The one or more computer-readable storage media further store program instructions for execution by at least one of the one or more computer processors, and the program instructions process the one or more audio data on the real-time audio list using a priority higher than any one or more audios on the batch audio list. The computer system according to item 17.

Claims

1. A computer-implemented method for determining real-time versus non-real-time audio streaming, comprising: receiving, by one or more computer processors, one or more audio data; determining, by one or more computer processors, an expected bitrate of the one or more audio data; determining, by one or more computer processors, an input bitrate of the one or more audio data; calculating, by one or more computer processors, an R value using the expected bitrate and the input bitrate; comparing, by one or more computer processors, the R value with an R threshold; in response to the R value being greater than the R threshold, adding, by one or more computer processors, the one or more audio data to a batch audio list A computer-implemented method comprising the above steps.

2. A computer-implemented method for determining real-time versus non-real-time audio streaming, comprising: receiving, by one or more computer processors, one or more audio data; determining, by one or more computer processors, an expected bitrate of the one or more audio data; determining, by one or more computer processors, an input bitrate of the one or more audio data; calculating, by one or more computer processors, an R value using the expected bitrate and the input bitrate; comparing, by one or more computer processors, the R value with an R threshold; In response to the R value being less than the R threshold value, adding, by one or more computer processors, the one or more audio data to a real-time audio list A computer-implemented method comprising. **Claim 3** The comparing step is performed after a threshold period, the computer-implemented method according to claim 1 or 2. **Claim 4** Processing, by one or more computer processors, the one or more audio data on the batch audio list using a lower priority than any one or more audio on the real-time audio list The computer-implemented method according to claim 1, further comprising. **Claim 5** Processing, by one or more computer processors, the one or more audio data on the real-time audio list using a higher priority than any one or more audio on the batch audio list The computer-implemented method according to claim 2, further comprising. **Claim 6** The R value is calculated by dividing the input bitrate by the expected bitrate, the computer-implemented method according to any one of claims 1 to 5. **Claim 7** A computer program for determining real-time versus non-real-time audio streaming, for one or more computer processors, A procedure for receiving one or more audio data; A procedure for determining the expected bitrate of the one or more audio data; A procedure for determining the input bitrate of the one or more audio data; A procedure for calculating an R value using the expected bitrate and the input bitrate; A procedure for comparing the R value with an R threshold; causing the one or more computer processors to add the one or more audio data to a batch audio list in response to the R value being greater than the R threshold A computer program.

8. A computer program for determining real-time versus non-real-time audio streaming, the method comprising, for one or more computer processors, receiving one or more audio data; determining an expected bitrate of the one or more audio data; determining an input bitrate of the one or more audio data; calculating an R value using the expected bitrate and the input bitrate; comparing the R value with an R threshold; causing the one or more computer processors to add the one or more audio data to a real-time audio list in response to the R value being less than the R threshold A computer program.

9. The computer program according to claim 7 or 8, wherein the comparing step is performed after a threshold period.

10. causing the one or more computer processors to process the one or more audio data on the batch audio list using a lower priority than any one or more audio on the real-time audio list The computer program according to claim 7.

11. The computer program according to claim 8, wherein the one or more computer processors are caused to execute a procedure for processing the one or more audio data on the real-time audio list using a higher priority than any one or more audio on the batch audio list.

12. The computer program according to any one of claims 7 to 11, wherein the R value is calculated by dividing the input bit rate by the expected bit rate.

13. A computer system for determining real-time versus non-real-time audio streaming, the computer system comprising: One or more computer processors; One or more computer-readable storage media; and Program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors comprising, the program instructions Program instructions for receiving one or more audio data; Program instructions for determining, by one or more computer processors, an expected bit rate of the one or more audio data; Program instructions for determining, by one or more computer processors, an input bit rate of the one or more audio data; Program instructions for calculating an R value using the expected bit rate and the input bit rate by one or more computer processors; Program instructions for comparing the R value with an R threshold by one or more computer processors; Program instructions for adding the one or more audio data to a batch audio list in response to the R value being greater than the R threshold A computer system having.

14. A computer system for determining real-time versus non-real-time audio streaming, the computer system comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors comprising, the program instructions comprising: program instructions for receiving one or more audio data; program instructions for determining, by one or more computer processors, an expected bit rate of the one or more audio data; program instructions for determining, by one or more computer processors, an input bit rate of the one or more audio data; program instructions for calculating an R value using the expected bit rate and the input bit rate by one or more computer processors; program instructions for comparing the R value with an R threshold by one or more computer processors; program instructions for adding the one or more audio data to a real-time audio list in response to the R value being less than the R threshold A computer system having.

15. The computer system according to claim 13 or 14, wherein the comparison is made after a threshold period.

16. The computer system according to claim 13, further comprising program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions using a lower priority than any one or more of the audios on the real-time audio list to process the one or more audio data on the batch audio list.

17. The computer system according to claim 14, further comprising program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions using a higher priority than any one or more of the audios on the batch audio list to process the one or more audio data on the real-time audio list. The computer system according to claim 14.

Citation Information

Patent Citations

  • Reproducing device and decoding control method

    JP2006186580A