Adjusting the playback speed of call audio
The system addresses frame loss in calls by playing missed audio frames at a faster speed, ensuring uninterrupted communication by quickly catching up with the conversation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2026-04-08
AI Technical Summary
Network problems during calls cause frame loss, leading to missed audio frames, which disrupt the user experience as the recipient must guess or ask the sender to repeat missed parts, wasting time and disrupting conversation flow.
A communication device or method that detects frame loss and initiates playback of missed audio frames at a faster speed than the normal playback speed to catch up with the conversation.
Enables seamless continuation of calls by quickly recovering missed audio frames through accelerated playback, minimizing user disruption and improving communication efficiency.
Smart Images

Figure 0007842767000001 
Figure 0007842767000002 
Figure 0007842767000003
Abstract
Description
Claim of Priority
[0001]
[0001] This application claims the benefit of priority from U.S. Non-Provisional Patent Application No. 17 / 142,022, owned by the applicant of this application, filed on January 5, 2021, the content of which is hereby expressly incorporated by reference in its entirety.
Technical Field
[0002]
[0002] This disclosure generally relates to adjusting the playback speed of call audio.
Background Art
[0003]
[0003] With the advancement of technology, computer devices have become smaller and more powerful. For example, there are now various portable personal computer devices, including wireless telephones such as mobile phones and smartphones that are small, lightweight, and easily carried by users, tablets, and laptop computers. These devices can transmit voice packets and data packets through wireless networks. Further, many of such devices incorporate additional functions such as digital still cameras, digital video cameras, digital recorders, and audio file players. Also, such devices can process executable instructions, including software applications that can be used to access the Internet, such as web browser applications. As such, these devices can include important computer functions.
[0004]
[0004] Such computer devices often incorporate functionality for receiving audio signals from one or more microphones. For example, the audio signals may represent the user's voice captured by the microphones, external sounds picked up by the microphones, or a combination thereof. Such devices may include communication devices for making calls, such as voice calls or video calls. Network problems during a call between a first user and a second user can cause frame loss, such that some audio frames transmitted from the first user's first device are not received by the second user's second device. In some cases, audio frames are received by the second device, but the second user misses part of the call because they are temporarily unable to attend (for example, they have to leave). The second user must guess what was missed or ask the first user to repeat the missed part, which negatively impacts the user experience. [Overview of the project]
[0005]
[0005] According to one embodiment of the present disclosure, a communication device includes one or more processors configured to receive a sequence of audio frames from a first device during a call. The one or more processors are also configured to initiate a frame loss indication to the first device in response to determining that no audio frames of the sequence have been received for a threshold duration since the last audio frame received in the sequence. The one or more processors are further configured to receive from the first device, in response to the frame loss indication, a set of audio frames of the sequence and an indication of a second playback speed. The one or more processors are also configured to initiate playback of the set of audio frames through a speaker based on the second playback speed, the second playback speed being greater than the first playback speed of the first set of audio frames of the sequence.
[0006]
[0006] According to another embodiment of the present disclosure, a method of communication includes, during a call, a device receiving a sequence of audio frames from a first device. The method also includes, in response to determining that no audio frames of the sequence have been received by the device during a threshold duration since the last audio frame received in the sequence, initiating the transmission of a frame loss indication from the device to the first device. The method further includes, in response to the frame loss indication, the device receiving from the first device an indication of a set of audio frames of the sequence and a second playback speed. The method also includes, initiating playback of the set of audio frames through a speaker based on a second playback speed, the second playback speed being greater than a first playback speed of the first set of audio frames of the sequence.
[0007]
[0007] According to another embodiment of the present disclosure, a communication device includes one or more processors configured to receive a sequence of audio frames from a first device during a call. The one or more processors are also configured to start playback of a set of audio frames based on a second playback speed, in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, the second playback speed being greater than a first playback speed of the first set of audio frames in the sequence.
[0008]
[0008] According to another embodiment of the present disclosure, a communication method includes a device receiving a sequence of audio frames from a first device. The method also includes, in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, starting playback of the set of audio frames on the device at at least a second playback speed. The second playback speed is greater than the first playback speed of the first set of audio frames in the sequence.
[0009]
[0009] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections: “Brief Description of the Drawings,” “Modes for Carrying Out the Invention,” and “Claims.” [Brief explanation of the drawing]
[0010] [Figure 1]
[0010] Block diagram of a particular exemplary embodiment of a system that can operate to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 2]
[0011] A diagram illustrating an exemplary embodiment of a system capable of adjusting the call audio playback speed, as shown in some examples of the present disclosure. [Figure 3]
[0012] A diagram illustrating an exemplary embodiment of a system capable of adjusting the call audio playback speed, as shown in some examples of the present disclosure. [Figure 4]
[0013] Ladder diagrams illustrating exemplary modes of operation of any component of the systems shown in Figures 1 to 3, according to some examples of the present disclosure. [Figure 5]
[0014] Figures illustrating exemplary embodiments of the operation of any component of the system shown in Figures 1 to 3, according to some examples of the present disclosure. [Figure 6]
[0015] A diagram illustrating an exemplary embodiment of a system capable of adjusting the call audio playback speed, as shown in some examples of the present disclosure. [Figure 7]
[0016] A diagram illustrating exemplary embodiments of the operation of the components of the system in Figure 6, based on several examples of the present disclosure. [Figure 8]
[0017] Figures illustrating specific embodiments of methods for adjusting the call audio playback speed, which may be implemented by the devices shown in Figures 1-3, as illustrated by some examples of the present disclosure. [Figure 9]
[0018] Figures illustrating specific embodiments of methods for adjusting the call audio playback speed, which may be implemented by the devices shown in Figures 1-3, as illustrated by some examples of the present disclosure. [Figure 10]
[0019] Figures illustrating specific embodiments of a method for adjusting the call audio playback speed, which may be implemented by the devices shown in Figures 1-3, as illustrated by some examples of the present disclosure. [Figure 11]
[0020] Figure 6 illustrates a specific embodiment of a method for adjusting the call audio playback speed, which may be performed by the device shown in some examples of the present disclosure. [Figure 12]
[0021] Figure 12 shows an example of an integrated circuit that can operate to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 13]
[0022] A diagram of a mobile device capable of adjusting the call audio playback speed, as illustrated in some examples of the present disclosure. [Figure 14]
[0023] A diagram of a headset capable of adjusting the call audio playback speed, as illustrated in some examples of the present disclosure. [Figure 15]
[0024] A diagram of a wearable electronic device capable of adjusting the playback speed of call audio, as illustrated by some examples of the present disclosure. [Figure 16]
[0025] A diagram of a voice-controlled speaker system that can operate to adjust the playback speed of call audio, as illustrated by some examples of the present disclosure. [Figure 17]
[0026] Diagram of a camera operable to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 18]
[0027] Diagram of a headset, such as a virtual reality or augmented reality headset, operable to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 19]
[0028] Diagram of a first example of a means of conveyance operable to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 20]
[0029] Diagram of a second example of a means of conveyance operable to adjust the call audio playback speed, according to some examples of the present disclosure. [Figure 21]
[0030] Block diagram of a specific illustrative example of a device operable to adjust the call audio playback speed, according to some examples of the present disclosure.
Mode for Carrying Out the Invention
[0011]
[0031] Missing a part of a call can have an adverse effect on the user experience. For example, during a call between a first user and a second user, if some audio frames transmitted from the first user's first device are not received by the second user's second device, the second user may miss a part of the first user's speech. As another example, the second user may miss a part of the call for other reasons, such as leaving the scene or being distracted. The second user has to either guess what the first user said or ask the first user to repeat what was missed. This can cause transmission errors, disrupt the flow of the conversation, and waste time.
[0012]
[0032] A system and method for adjusting the playback speed of call audio are disclosed. For example, a first call manager of a first device establishes a call with a second call manager of a second device. During the call, the first call manager sends a sequence of audio frames to the second device and buffers at least the most recently sent audio frames. The second call manager receives at least a portion of the sequence of audio frames and buffers the received audio frames for playback.
[0013]
[0033] In a specific example, a second frame loss manager in a second device sends a frame loss indication to the first device upon detection of frame loss. For example, the second frame loss manager detects frame loss when it determines that no audio frame was received within a specific period of time during which the last received audio frame (e.g., the most recently received one) was received. The frame loss indication shows the last received audio frame.
[0014]
[0034] The first frame loss manager of the first device, in response to receiving a frame loss indication, retransmits a set of audio frames to the second device. For example, this set of audio frames follows the last received audio frame in the sequence. The second call manager plays this set of audio frames at a second playback speed that is faster than the first playback speed of the previous audio frames. In a particular example, the first frame loss manager sends a second set of audio frames based on the set of audio frames such that playback of the second set of audio frames at the first playback speed corresponds to the effective second playback speed of that set of audio frames. For example, if the second set of audio frames contains every other frame from the set of audio frames, playback of the second set of audio frames at the first playback speed corresponds to the effective playback speed of the set of audio frames, which is twice the first playback speed. By playing the audio frames at a faster speed (or a faster effective speed), the first device can catch up with the call.
[0015]
[0035] In a specific example, a second user pauses audio playback during a call and then resumes it. For example, playback pauses after the last played audio frame (e.g., the one played immediately before). The second call manager on the second device, having decided to resume playback, plays a set of audio frames at a second playback speed that is faster than the first playback speed of the previous audio frames. This set of audio frames follows the last played audio frame in the sequence.
[0016]
[0036] The second call manager plays the subsequent audio frames at the first playback speed. In a specific example, the second call manager transitions between the second and first playback speeds in such a way that the change in playback speed is not very noticeable (e.g., not noticeable).
[0017]
[0037] Specific aspects of this disclosure are described below with reference to the drawings. In this specification, common functional parts are designated by common reference numerals. In this specification, various terms are used for the purpose of describing only specific embodiments and are not limited to embodiments. For example, the singular forms “a,” “an,” and “the” also include the plural unless otherwise indicated. Furthermore, some functional parts described herein are singular in some embodiments and plural in other embodiments. For example, Figure 1 shows a device 102 including one or more processors (“processor” 120 in Figure 1), which shows that in some embodiments the device 102 includes a single processor 120 and in other embodiments the device 102 includes multiple processors 120. For ease of reference in this specification, such functional parts are generally introduced as “one or more” functional parts and are referred to hereafter in the singular unless an embodiment relating to multiple functional parts is described.
[0018]
[0038] In this specification, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Furthermore, the term “wherein” may be used interchangeably with “where.” In this specification, “exemplary” should not be construed as indicating, limiting, or preferential, or preferential embodiment, an example, embodiment, and / or aspect. In this specification, ordering terms used to modify elements such as structure, components, and actions (e.g., “first,” “second,” “third,” etc.) do not in themselves indicate any priority or order of that element over another element, but merely distinguish that element from another element with the same name (with respect to the use of ordering terms). In this specification, the term “set” refers to one or more of a particular element, and the term “plural” refers to multiple (e.g., two or more) of a particular element.
[0019]
[0039] In this specification, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or a combination thereof). Two electrically coupled devices (or components) may be contained within the same device or within separate devices, and may be connected via electronic equipment, one or more connectors, or inductive coupling, as descriptive, non-limiting examples. In some embodiments, two devices (or components) that are communicatively coupled, such as by electrical communication, may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without any intervening components.
[0020]
[0040] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” may be used to describe how one or more actions are performed. Such terms should not be interpreted restrictively, and it should be noted that other techniques may be used to perform similar actions. Furthermore, as used herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” may be used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device, for example.
[0021]
[0041] Referring to Figure 1, a specific exemplary embodiment of a system configured to adjust the call audio playback speed is disclosed, shown as 100 in its entirety. System 100 includes device 102 coupled to device 104 via network 106.
[0022]
[0042] Device 102 is coupled to speaker 129. Device 102 includes memory 132 coupled to one or more processors 120. In a particular example, memory 132 includes a receive buffer 134 (e.g., a circular buffer) configured to store the most recently received audio frame for playback. Memory 132 is configured to store data indicating a first playback rate 105. In a particular embodiment, the first playback rate 105 is based on configuration settings, default data, user input, or a combination thereof. In a particular embodiment, the first playback rate 105 corresponds to the normal (e.g., expected) playback rate of a call. One or more processors 120 include a call manager 122 and a frame loss manager 124.
[0023]
[0043] Device 104 is coupled to a microphone 146. Device 104 includes memory 154 coupled to one or more processors 150. In a particular example, memory 154 includes a transmit buffer 110 (e.g., a circular buffer) configured to store the most recently transmitted audio frame. One or more processors 150 include a call manager 152 and a frame loss manager 156.
[0024]
[0044] Call managers 122 and 152 are configured to manage calls (e.g., audio calls, video calls, or both) on devices 102 and 104, respectively. In certain embodiments, call managers 122 and 152 correspond to clients of a communication application (e.g., an online meeting application). In certain embodiments, frame loss managers 124 and 156 are configured to manage frame loss recovery during calls on devices 102 and 104, respectively.
[0025]
[0045] In some embodiments, call managers 122 and 152 are blind to (e.g., unaware of) any frame losses managed by frame loss managers 124 and 156. In some embodiments, call managers 122 and 152 correspond to the upper layers (e.g., the application layer) of the network protocol stack (e.g., the Open System Integrity (OSI) model) of devices 102 and 104, respectively. In some embodiments, frame loss managers 124 and 156 correspond to the lower layers (e.g., the transport layer) of the network protocol stack of devices 102 and 104, respectively.
[0026]
[0046] In some embodiments, device 102, device 104, or both correspond to or are included in one or various types of devices. In a descriptive example, one or more processors 120, one or more processors 150, or a combination thereof are incorporated into a headset device, including a microphone, speaker, or both, as further described with reference to Figure 14. In another example, one or more processors 120, one or more processors 150, or a combination thereof are incorporated into at least one of the following: a mobile phone or tablet computer device, as described with reference to Figure 13; a wearable electronic device, as described with reference to Figure 15; a voice-controlled speaker system, as described with reference to Figure 16; a camera device, as described with reference to Figure 17; or a virtual reality headset or augmented reality headset, as described with reference to Figure 18. In yet another descriptive example, one or more processors 120, one or more processors 150, or a combination thereof are incorporated into a carrying device, also including a microphone, speaker, or both, as further described with reference to Figures 19 and 20.
[0027]
[0047] During operation, call managers 122 and 152 establish a call (e.g., an audio call, a video call, an online meeting, or a combination thereof) between device 102 and device 104. For example, a call is made between user 144 on device 104 and user 142 on device 102. Microphone 146 captures user 144's voice while user 144 is speaking and supplies device 104 with an audio input 141 representing that voice. Call manager 152 generates a sequence of audio frames 119 based on the audio input 141 and sends the sequence of audio frames 119 to device 102. For example, sequence 119 includes a set of audio frames 109, audio frame 111, a set of audio frames 113, the next audio frame 117, one or more additional audio frames, or a combination thereof. For example, the call manager 152 generates audio frames of sequence 119 when audio input 141 is received, and transmits sequence 119 of audio frames when the audio frames are generated (for example, by initiating the transmission of sequence 119).
[0028]
[0048] In a particular embodiment, the frame loss manager 156 buffers the audio frames of sequence 119 in the transmit buffer 110 when the audio frames are transmitted by the call manager 152. For example, the call manager 152 transmits each of the set of audio frames 109 to device 102 via the network 106. The call manager 152 transmits audio frame 111 to device 102 via the network 106 at a first transmission time. Similarly, the call manager 152 transmits each of the set of audio frames 113 to device 102 via the network 106. The frame loss manager 156 stores each of the set of audio frames 109, audio frame 111, and set of audio frames 113 in the transmit buffer 110.
[0029]
[0049] Device 102 receives a sequence of audio frames 119 from device 104 via network 106. In a particular embodiment, device 102 receives a set of audio frames (e.g., a burst) of sequence 119. In an alternative embodiment, device 102 receives one audio frame at a certain point in sequence 119.
[0030]
[0050] In a specific example, the frame loss manager 124 receives each of the sets 109 of audio frames in sequence 119 and stores each of the sets 109 in the receive buffer 134. The call manager 122 retrieves one or more of the sets 109 of audio frames from the receive buffer 134 and outputs the retrieved audio frames by the speaker 129 at a first playback speed 105. The sets 109 of audio frames precede audio frame 111 in sequence 119.
[0031]
[0051] The frame loss manager 124 receives the audio frame 111 at a first reception time. In a particular embodiment, the frame loss manager 124 stores the audio frame 111 in the receive buffer 134 for playback (for example, while one or more of the set of audio frames 109 are being played at the corresponding playback time). For example, the frame loss manager 124 adds a delay between the reception of the audio frame 111 and the playback of the audio frame 111 by the call manager 122 at the first playback time to increase the likelihood that subsequent frames will be available at the corresponding playback time in the receive buffer 134 (for example, at the second playback time).
[0032]
[0052] In a specific example, the frame loss manager 124 detects frame loss following the reception of audio frame 111. For example, the frame loss manager 124 detects frame loss in response to determining that no audio frames in sequence 119 have been received for a threshold duration since audio frame 111 (e.g., the last received audio frame in sequence 119) was received at a first reception point.
[0033]
[0053] In response to the detection of a frame loss, the frame loss manager 124 initiates the transmission of a frame loss indication 121 to the device 104 over the network 106. In a particular embodiment, the frame loss indication 121 indicates an audio frame 111 (e.g., the last received audio frame) (e.g., including its identifier). In another embodiment, the frame loss indication 121 includes a request to retransmit a previous audio frame corresponding to the estimated playback duration of the lost frame at a first playback speed 105. For example, at a particular point in time, the frame loss manager 124 determines the estimated playback duration based on a first reception time (e.g., of the last received audio frame) and a specific point in time (e.g., estimated playback duration = specific point in time - first reception time).
[0034]
[0054] The frame loss manager 156 receives a frame loss indication 121 from device 102 via network 106. In a particular embodiment, the frame loss manager 156 determines that the frame loss indication 121 indicates that audio frame 111 corresponds to the last audio frame received by device 102 (for example, including its identifier). In an alternative embodiment, the frame loss manager 156 identifies the last received audio frame based on its estimated playback duration, in response to determining that the frame loss indication 121 includes a request to retransmit a previous audio frame corresponding to the estimated playback duration of the lost frame at a first playback speed 105. For example, the frame loss manager 156 determines the last transmission time at a particular point in time based on the difference between that point in time and the estimated playback duration (for example, last transmission time = specific point in time - estimated playback duration), and identifies audio frame 111 as the last received audio frame, in response to determining that the last transmission time coincides with the transmission time of audio frame 111.
[0035]
[0055] The frame loss manager 156 determines which previous audio frames should be retransmitted based on the last received audio frame. For example, the frame loss manager 156 determines that audio frame 111 corresponds to the last received audio frame at device 102, and that set of audio frames 113 follows the last received audio frame in sequence 119 and has been previously transmitted, and therefore identifies set of audio frames 113 as the previous audio frames to be retransmitted.
[0036]
[0056] The frame loss manager 156, upon determining that the previous audio frames to be retransmitted include a set of audio frames 113 and that the set of audio frames 113 is available in the transmit buffer 110, generates a set of audio frames 123 based on the set of audio frames 113 and begins transmitting the set of audio frames 123, an indication of a second playback speed 115, or both, to the device 102 via the network 106. The second playback speed 115 is greater than the first playback speed 105. In a particular embodiment, the set of audio frames 123 is the same as the set of audio frames 113, and the frame loss manager 156 begins transmitting the set of audio frames 123 and an indication of a second playback speed 115.
[0037]
[0057] In an alternative embodiment, the audio frame set 123 includes a subset of the audio frame set 113 such that playback of the audio frame set 123 at a first playback speed 105 corresponds to effective playback of the audio frame set 113 at a second playback speed 115. For example, if the audio frame set 123 includes all other frames of the audio frame set 113, then playback of the audio frame set 123 at a first playback speed 105 corresponds to the effective playback speed of the audio frame set 113 (e.g., a second playback speed 115), which is twice as fast as the first playback speed 105. The frame loss manager 156 starts transmitting the audio frame set 123 even without an indication of the second playback speed 115 because playback of the audio frame set 123 at a first playback speed 105 (e.g., normal playback speed) corresponds to an effective playback speed adjustment of the second playback speed 115.
[0038]
[0058] In certain embodiments, the frame loss manager 156 determines a second playback speed 115 based on the number of audio frame sets 113, default values, configuration settings, user input, or a combination thereof. For example, the second playback speed 115 increases with the number of retransmitted audio frames.
[0039]
[0059] In certain embodiments, the frame loss manager 156 suppresses silence in a set of audio frames 113 within a set of audio frames 123. For example, the frame loss manager 156 sends an indication of a third playback speed for the corresponding silent frame in the set of audio frames 123, depending on whether it has determined that the set of audio frames 113 contains silent frames. In certain embodiments, the frame loss manager 156 selects fewer silent frames than audio frames in the set of audio frames 113 to generate the set of audio frames 123. For example, the frame loss manager 156 selects every other audio frame in the set of audio frames 113, as well as every four silent frames in the set of audio frames 113. The third playback speed is greater than the second playback speed 115. For example, by shortening the silent portions, the device 102 can catch up to the call faster (for example, by returning to playing subsequent audio frames at the first playback speed 105).
[0040]
[0060] In certain embodiments, the frame loss manager 156 selectively initiates the transmission of a set of audio frames 123 depending on whether it has determined that the number of previous audio frames to be retransmitted is greater than a first threshold, less than a second threshold, or both. For example, if too few audio frames are lost, the frame loss may not be significant enough to warrant implementing playback speed adjustment. As another example, if too many audio frames are lost, the playback speed may need to be significantly increased to catch up with the call, and playback speed adjustment may have a more negative impact on the user experience than missing audio frames. In certain embodiments, the frame loss manager 156 initiates the transmission of a set of audio frames 123 depending on whether it has determined that the network link with device 102 has been re-established.
[0041]
[0061] In a particular embodiment, the frame loss manager 156 sends a skip-ahead notification to device 104 in response to determining that the number of previous audio frames is less than or equal to a first threshold, or greater than or equal to a second threshold. In a particular embodiment, the skip-ahead notification indicates the subsequent audio frame to be sent by device 104. For example, the skip-ahead notification indicates the first audio frame of the next audio frame 117 of sequence 119. Following (or simultaneously with) sending the skip-ahead notification, the frame loss manager 156 begins sending the next audio frame 117. The next audio frame 117 follows the set of audio frames 113 of sequence 119. In a particular embodiment, the next audio frame 117 has not been previously sent during the call.
[0042]
[0062] In a particular embodiment, the frame loss manager 156 initiates the transmission of the next audio frame 117 simultaneously with the transmission of the set of audio frames 123, depending on whether it has determined that a previous audio frame should be retransmitted. For example, the transmission of the set of audio frames 123 (corresponding to the retransmission of at least a subset of the set of audio frames 113) does not delay the initial transmission of the next audio frame 117. In a particular embodiment, the next audio frame 117 is associated with a transition between the second playback speed 115 and the first playback speed 105 (for example, the normal playback speed) of the set of audio frames 123 (for example, the retransmitted audio frames). For example, the frame loss manager 156 transmits a first subset of the next audio frame 117 with an indication of a second playback speed 115, a second subset of the next audio frame 117 with an indication of an intermediate playback speed between the second playback speed 115 and the first playback speed 105, a third subset of the next audio frame 117 with an indication of the first playback speed 105, or a combination thereof. In an alternative embodiment, the frame loss manager 156 selects a first audio frame from the first subset based on the second playback speed 115 (for example, every other audio frame) and a second audio frame from the second subset based on the intermediate playback speed. The frame loss manager 156 transitions the next audio frame 117 from the second playback speed 115 to the first playback speed 105 by initiating the transmission of the selected first audio frame of the first subset, the selected second audio frame of the second subset, and the audio frame of the third subset.
[0043]
[0063] The frame loss manager 124 receives, in response to the frame loss indication 121, an indication of a set of audio frames 123, a second playback speed 115, or both. The frame loss manager 124 then initiates playback of the set of audio frames 123 by the speaker 129 based on the second playback speed 115, as further explained with reference to Figure 5.
[0044]
[0064] In a particular embodiment, set 123 of audio frames is the same as set 113 of audio frames. Frame loss manager 124 adds set 123 of audio frames and an indication of a second playback speed 115 to receive buffer 134. Call manager 122 retrieves one or more of the sets of audio frames 123 from receive buffer 134 and performs playback of set 123 of audio frames based on the second playback speed 115. For example, call manager 122 supplies each of the sets of audio frames 123 to speaker 129 for playback based on the second playback speed 115. In an alternative embodiment, frame loss manager 124 retrieves each audio frame of set 123 from receive buffer 134 based on the second playback speed 115 and supplies the retrieved audio frames to call manager 122. Call manager 122 supplies each received audio frame to speaker 129 for playback.
[0045]
[0065] In an alternative embodiment, the audio frame set 123 includes a subset of the audio frame set 113 such that playback of the audio frame set 123 at a first playback speed 105 corresponds to the effective playback speed of the audio frame set 113 (e.g., a second playback speed 115). The frame loss manager 124 adds the audio frame set 123 to the receive buffer 134. The call manager 122 retrieves one or more of the audio frame sets 123 from the receive buffer 134 and performs playback of the audio frame set 123 based on the first playback speed 105. For example, the call manager 122 supplies each of the audio frames 123 of the set to the speaker 129 for playback based on the first playback speed 105. In an alternative embodiment, the frame loss manager 124 retrieves each audio frame of the audio frame set 123 from the receive buffer 134 based on the first playback speed 105 and supplies the retrieved audio frames to the call manager 122. The call manager 122 supplies each received audio frame to the speaker 129 for playback.
[0046]
[0066] In a particular embodiment, the frame loss manager 124 suppresses silence at least partially in order to begin playback of the set of audio frames 123. For example, the frame loss manager 124 begins playback of the silent frames at a third playback speed greater than a second playback speed 115, in response to determining that the set of audio frames 123 contains silent frames. As another example, the frame loss manager 124 begins playback of the silent frames at a third playback speed in response to receiving an indication of a third playback speed for the silent frames of the set of audio frames 123.
[0047]
[0067] In a specific example, the frame loss manager 124 adds silent frames and an indication of a third playback speed to the receive buffer 134, and the call manager 122 performs playback of the silent frames based on the third playback speed. For example, the call manager 122 supplies the silent frames to the speaker 129 for playback based on the third playback speed. In an alternative embodiment, the frame loss manager 124 retrieves silent frames from the receive buffer 134 based on the third playback speed and supplies the retrieved silent frames to the call manager 122. The call manager 122 supplies the received silent frames to the speaker 129 for playback.
[0048]
[0068] In a particular embodiment, the frame loss manager 124 receives the next audio frame 117 simultaneously with receiving the set of audio frames 123, as further described with reference to Figure 4. In a particular embodiment, the frame loss manager 124 transitions the playback of the next audio frame 117 from a second playback speed 115 to a first playback speed 105, as further described with reference to Figure 5.
[0049]
[0069] In a particular embodiment, the frame loss manager 124 shifts the playback speed of the next audio frame 117 based on an indication of playback speed received for the next audio frame 117. For example, the frame loss manager 124, upon receiving an indication of a second playback speed 115 for each of the first subsets, starts playback of the first subset of the next audio frame 117 based on the second playback speed 115; upon receiving an indication of an intermediate playback speed for each of the second subsets, starts playback of the second subset of the next audio frame 117 based on the intermediate playback speed; upon receiving an indication of a first playback speed 105 for each of the third subsets, starts playback of the third subset of the next audio frame 117 at the first playback speed 105; or starts a combination of these.
[0050]
[0070] In a particular embodiment, the frame loss manager 124 shifts the playback speed of the next audio frame 117 based on the received audio frames of the next audio frame 117. For example, the frame loss manager 124 plays the received audio frames of the next audio frame 117 at a first playback speed 105. Since a small number of first audio frames from a first subset of the next audio frame 117 are received, playback of the first audio frames at the first playback speed 105 corresponds to the effective playback speed of the second playback speed 115 for the first subset. Similarly, playback of the second audio frames of the second subset at the first playback speed 105 corresponds to the effective playback speed of the intermediate playback speed for the second subset. Playback of all frames of the third subset of the next audio frame 117 at the first playback speed 105 corresponds to the effective playback speed of the first playback speed 105 for the third subset.
[0051]
[0071] In certain embodiments, the frame loss manager 124 shifts the playback speed of the next audio frame 117 based on the number of audio frames stored (e.g., remaining) in the receive buffer 134 for playback. For example, if the frame loss manager 124 determines that more audio frames than a first threshold number are available in the receive buffer 134 for playback, it starts playback of a first subset of the next audio frame 117 at a second playback speed 115; if it determines that less than or equal to the first threshold number but more than the second threshold number of audio frames are available in the receive buffer 134 for playback, it starts playback of a second subset of the next audio frame 117 at an intermediate playback speed between the second playback speed 115 and the first playback speed 105; and if it determines that less than or equal to the second threshold number of audio frames are available in the receive buffer 134 for playback, it starts playback of a third subset of the next audio frame 117 at the first playback speed 105, or a combination thereof. By shifting the playback speed of the next audio frame 117, the playback speed adjustment can be made less noticeable (for example, not noticeable), allowing device 102 to return to the normal playback speed after catching up with the call.
[0052]
[0072] In a particular embodiment, the frame loss manager 124 receives a skip-ahead notification in response to a frame loss indication 121. Upon receiving the skip-ahead notification and the next audio frame 117, the frame loss manager 124 starts playback of the next audio frame 117 at a first playback speed 105.
[0053]
[0073] In this way, system 100 is able to receive previously lost audio frames and catch up with the call by playing back the audio frames at a faster playback speed. In some embodiments, the transmission of lost audio frames is selective based on the number of lost audio frames in order to balance the impact of the gaps in the call caused by the lost audio frames with speeding up the call audio in order to catch up.
[0054]
[0074] Microphone 146 is shown coupled to device 104, but in other embodiments, microphone 146 may be integrated into device 104. Speaker 129 is shown coupled to device 102, but in other embodiments, speaker 129 may be integrated into device 102. One microphone is shown, but in other embodiments, one or more additional microphones configured to capture user speech may be included.
[0055]
[0075] For ease of explanation, please understand that device 104 is described as a transmitting device and device 102 is described as a receiving device. During a call, when user 142 begins speaking, the roles of devices 102 and 104 can switch. For example, device 102 can become a transmitting device and device 104 can become a receiving device. In certain embodiments, for example, when both user 142 and user 144 are speaking simultaneously or at overlapping times, device 102 and device 104 can each become a transmitting and receiving device, respectively.
[0056]
[0076] In certain embodiments, the call manager 122 is also configured to perform one or more operations described with respect to the call manager 152, and vice versa. In certain embodiments, the frame loss manager 124 is also configured to perform one or more operations described with respect to the frame loss manager 156, and vice versa. In certain embodiments, each of device 102 and device 104 includes a receive buffer, a transmit buffer, a speaker, and a microphone.
[0057]
[0077] Referring to Figure 2, a system configured to adjust the call audio playback speed is disclosed, the whole of which is shown as 200. In a particular embodiment, system 100 in Figure 1 includes one or more components of system 200. System 200 includes a server 204 coupled to devices 102 and 104 via a network 106.
[0058]
[0078] The server 204 includes memory 232 coupled to one or more processors 220. Memory 232 includes buffers 210 (e.g., circular buffers). One or more processors 220 include a call manager 222 and a frame loss manager 224. In a particular embodiment, the frame loss manager 224 corresponds to the frame loss manager 156 in Figure 1.
[0059]
[0079] The call manager 222 is configured to receive and forward audio frames during a call. For example, the call manager 222 establishes a call between the call manager 152 and the call manager 122. During the call, the call manager 152 captures the voice of the user 144 and sends a sequence of audio frames 119 to the server 204. The server 204 stores the audio frames of sequence 119 in the buffer 210 and forwards the audio frames to the device 102.
[0060]
[0080] The frame loss manager 124 sends the frame loss indication 121, as described with reference to Figure 1, to the server 204 (for example, not to device 104). The frame loss manager 224 on the server 204 manages frame loss recovery. For example, device 104 is blind to (e.g., unaware of) the frame loss detected by device 102. For example, the frame loss manager 224 on the server 204 performs one or more of the actions described with respect to the frame loss manager 156 in Figure 1.
[0061]
[0081] In this way, system 200 enables frame loss recovery for legacy devices (e.g., device 104). In certain embodiments, server 204 may also be closer to device 102 (e.g., fewer network hops), and retransmitting the missing audio frames from server 204 (e.g., rather than from device 104) may save overall network resources. In certain embodiments, server 204 may have access to network information that may be useful in successfully retransmitting a set of audio frames 113 (e.g., the corresponding set of audio frames 123) to device 102. As an example, server 204 first transmits the set of audio frames 113 over a first network link. Server 204 receives a frame loss indication 121 and, at least in part, determines that the first network link is unavailable (e.g., down), and transmits the set of audio frames 123 using a second network link that appears to be available.
[0062]
[0082] Referring to Figure 3, a system configured to adjust the call audio playback speed is disclosed, shown as 300 in its entirety. In a particular embodiment, system 100 in Figure 1 includes one or more components of system 300. System 300 includes a server 204 coupled via network 106 to device 102, device 104, device 302, one or more additional devices, or a combination thereof.
[0063]
[0083] The call manager 222 is configured to establish a call between multiple devices. For example, the call manager 222 establishes a call between user 142 on device 102, user 144 on device 104, and user 344 on device 302. During the call, device 104 captures user 144's voice and sends a sequence 331 of audio frames to server 204, and device 302 captures user 344's voice and sends a sequence 333 to server 204.
[0064]
[0084] In a particular embodiment, server 204 receives audio frames of sequence 331 interspersed with receiving audio frames of sequence 333. For example, user 144 and user 344 are speaking alternately, simultaneously, or overlapping in time. Sequence 331 is illustrated as containing multiple blocks. It should be understood that each block of a sequence represents one or more audio frames. One block of a sequence may represent the same or a different number (e.g., 100) of audio frames as another block of that sequence. Server 204 stores sequences 331 and 333 as sequence 119 in buffer 210 for device 102. Server 204 begins transmitting sequence 119, at least in part, based on its determination that the audio frames of sequence 119 are available in buffer 210 for transmission to device 102.
[0065]
[0085] Server 204 transmits the audio frames of sequence 119 to device 102, receives the frame loss indication 121, and transmits the set of audio frames 123 to device 102, as described with reference to Figure 2. In a particular embodiment, server 204 transfers sequence 333 to device 104, as described with reference to Figure 2. In a particular embodiment, server 204 transfers sequence 331 to device 302, as described with reference to Figure 2.
[0066]
[0086] In this way, system 300 enables server 204 to perform frame loss recovery for audio frames received from multiple devices during a call. For example, the playback speed of missed audio frames in set 113 from devices 104 and 302 is similarly adjusted in set 123 of audio frames to maintain the relative timing of the conversation between user 144 and user 344.
[0067]
[0087] Referring to Figure 4, a diagram is shown, with the entire system represented by 400. Figure 400 illustrates exemplary modes of operation of components of System 100 in Figure 1, System 200 in Figure 2, System 300 in Figure 3, or combinations thereof. The timing and operation shown in Figure 4 are for illustrative purposes only and are not limiting. In other embodiments, additional or fewer operations may be performed, and the timing may differ.
[0068]
[0088] Figure 400 illustrates the timing of the transmission of the audio frame of sequence 119 from device 402 to device 102. In certain embodiments, device 402 corresponds to device 104 in Figure 1 or server 204 in Figures 2-3.
[0069]
[0089] Device 402 sends audio frame 111 to device 102. Audio frame 111 is received by device 102 at time t0. Device 402 sends a set of audio frames 113 to device 102. For example, device 402 sends audio frame 411, audio frame 413, and audio frame 415 to device 102. Audio frames 411, 413, and 415 are expected to be received by device 102 around time t1, t2, and t3 (e.g., without network issues). The set of audio frames 113 is described as containing three audio frames for ease of explanation. In other embodiments, the set of audio frames 113 may contain fewer than three audio frames or more than three audio frames.
[0070]
[0090] Device 102 (for example, the frame loss manager 124 in Figure 1) detects frame loss in accordance with the determination that, following time t3, the audio frame 111 was not received within a threshold duration (for example, the threshold duration is less than or equal to the difference between time t3 and time t0) from time t0 of reception of the audio frame 111.
[0071]
[0091] In response to detecting a frame loss, device 102 (for example, frame loss manager 124) sends a frame loss indication 121 to device 402. In a particular embodiment, the frame loss indication 121 indicates an audio frame 111 (for example, an identifier of audio frame 111, or the playback duration of the missed audio frame) as the last audio frame received in sequence 119, as described with reference to Figure 1.
[0072]
[0092] Device 402 generates a set of audio frames 123 based on the set of audio frames 113 in response to receiving a frame loss indication 121 and determining that the set of audio frames 113 corresponds to previous audio frames that were not received by device 102, as described with reference to Figure 1. For example, the set of audio frames 123 includes audio frames 451, 453, 457, one or more additional frames, or a combination thereof, which are based on audio frames 411, 413, 415, one or more additional frames, or a combination thereof from the set of audio frames 113. In a particular embodiment, the set of audio frames 123 includes the same audio frames as the set of audio frames 113. In an alternative embodiment, the set of audio frames 123 includes a subset of the set of audio frames 113.
[0073]
[0093] Device 402 sends a set of audio frames 123 to device 102. In a particular embodiment, the set of audio frames 123 is sent as a single transmission to device 102. In an alternative embodiment, one or more audio frames from the set of audio frames 123 are sent to device 102 at shorter transmission intervals compared to the first transmission of the audio frames of sequence 119 to device 102. For example, in normal operation, audio frames are expected to be sent or received at expected intervals (e.g., the average of the difference between time t1 and time t0, the difference between time t2 and time t1, and the difference between time t3 and time t2). Device 102 sends the set of audio frames 123 at shorter intervals than expected.
[0074]
[0094] In a particular embodiment, the next audio frame 117 transitions from a second playback speed 115 to a first playback speed 105. For example, device 402 generates the next audio frame 495 based on a first subset of the next audio frame 117 (e.g., the next audio frame 491). The next audio frame 495 has a second playback speed 115. The second subset (e.g., the next audio frame 493) has a first playback speed 105.
[0075]
[0095] In a particular embodiment, the next audio frame 495 contains the same audio frames as the next audio frame 491. For example, audio frames 457, 459, and 461 of the next audio frame 495 are the same as audio frames 417, 419, and 421 of the next audio frame 491, respectively. Device 402 transmits the next audio frame 495 along with a second playback speed indication 115 for each of the next audio frames 495. Device 102 receives audio frame 457 at time t4 (for example, the expected time of reception of audio frame 417). Device 102 receives audio frame 459 at time t5 (for example, the expected time of reception of audio frame 419). Device 102 receives audio frame 461 at time t6 (for example, the expected time of reception of audio frame 421).
[0076]
[0096] In a particular embodiment, the next audio frame 495 includes a subset of the next audio frame 491 (e.g., every other audio frame). As an illustrative example, device 402 transmits audio frame 457 (e.g., based on audio frame 417) and audio frame 461 (e.g., based on audio frame 421) without transmitting any audio frame based on audio frame 419 (e.g., audio frame 459). In this example, device 102 receives audio frame 457 at time t4, does not receive audio frame 459, and receives audio frame 461 at time t5. Since the next audio frame 495 includes a subset of the next audio frame 491, playback of the next audio frame 495 at a first playback speed 105 corresponds to an effective playback speed of the next audio frame 491 that is greater than the first playback speed 105 (e.g., a second playback speed 115).
[0077]
[0097] In a specific example, device 402 transmits a set of audio frames 123 at the same time as transmitting the next audio frame 117. For example, device 402 transmits audio frame 457 at the expected transmission time of audio frame 417, based on the transmission time of audio frame 415 and the expected transmission interval (e.g., the default transmission interval).
[0078]
[0098] Device 402 transmits the next audio frame 493 (for example, audio frame 423, audio frame 425, one or more additional audio frames, or a combination thereof) at a first playback speed 105. For example, device 102 receives audio frame 423 at time t7, audio frame 425 at time t8, or both.
[0079]
[0099] Thus, Figure 400 illustrates that transmitting set 123 of audio frames does not delay the transmission of the next audio frame 117. In a particular embodiment, the next audio frame 117 transitions from a second playback speed 115 to a first playback speed 105. For example, the next audio frame 491 has a second playback speed 115, and the next audio frame 493 has a first playback speed 105.
[0080]
[0100] Referring to Figure 5, a diagram is shown, with the entire system represented by 500. Diagram 500 illustrates exemplary modes of operation of components of System 100 in Figure 1, System 200 in Figure 2, System 300 in Figure 3, or combinations thereof. The timings and operations shown in Figure 5 are for illustrative purposes only and are not limiting. In other embodiments, additional or fewer operations may be performed, and the timings may differ.
[0081]
[0101] Device 402 generates the audio frames 111, 411, 413, 415, 417, 419, 421, 423, 425, one or more additional audio frames, or a combination thereof, of sequence 119.
[0082]
[0102] Device 402 sends audio frame 411 to device 102. Device 102 plays audio frame 411 based on a first playback speed 105. Device 402 sends audio frame 413, audio frame 415, and audio frame 417, respectively. Device 102 does not receive any of audio frame 413, audio frame 415, or audio frame 417.
[0083]
[0103] Device 402 transmits a set of audio frames 123 (for example, audio frames 451, 453, and 455) at the same time as transmitting the next audio frame 117. For example, as described with reference to Figure 4, device 402 generates the next audio frame 495 (for example, audio frames 457, 459, and 461) corresponding to the next audio frame 491 based on a second playback speed 115. Device 402 begins transmitting the next audio frame 495 and the next audio frame 493 at the expected time when it will begin transmitting the next audio frame 117. For example, device 402 transmits audio frames 457, 459, 461, 423, and 425 at the expected time when it will transmit the corresponding audio frames 417, 419, 421, 423, and 425.
[0084]
[0104] Device 102 plays back the set of audio frames 123 and the next audio frame 491 based on a second playback speed 115 (for example, twice the speed of the first playback speed 105). For example, it takes the same amount of time for device 102 to play back audio frames 451 and 453 at the second playback speed 115 as it takes to play back a single audio frame at the first playback speed 105. Similarly, it takes the same amount of time for device 102 to play back audio frames 455 and 457 at the second playback speed 115 as it takes to play back a single audio frame at the first playback speed 105.
[0085]
[0105] In certain embodiments, there is a gap between the playback output of audio frame 111 and the playback output of audio frame 451 in device 102. The next audio frame 493 can be played back at a first playback speed 105 so that the call can be kept up by playing back set of audio frames 123 and the next audio frame 491. For example, device 102 plays back audio frames 423 and 425 at the first playback speed 105.
[0086]
[0106] The next audio frame 491 is played back based on the second playback speed 115 to compensate for the delay caused by playing back the set of audio frames 123 when starting the playback output of the next audio frame 117. For example, audio frames 451 and 453 are played back at the point in time when audio frame 417 would have been played back under normal conditions. To catch up with the call, if the second playback speed 115 is twice as fast as the first playback speed 105, then an equal number of additional audio frames (e.g., 3 additional audio frames) of the sequence 119 must be played back based on the second playback speed 115, equal to the number of lost audio frames (e.g., 3 audio frames). For example, if an audio frame corresponding to a playback duration (e.g., 30 seconds) at the first playback speed 105 is lost, then if the second playback speed 115 is twice as fast as the first playback speed 105, the audio frame corresponding to a playback duration (e.g., 30 seconds) at the second playback speed 115 will be played back to catch up with the call. For example, the first half of the playback duration (e.g., 15 seconds) is used to output the missed audio frames based on a second playback speed 115, and the second half of the playback duration (e.g., 15 seconds) is used to output the audio frames that would have been played at a first playback speed 105 under normal conditions during the playback duration (e.g., 30 seconds). In certain embodiments, partial silence suppression may allow for faster catching up or a reduction in the speed at which audio frames must be played in order to achieve a second playback speed 115 overall for a set of audio frames 123, or to achieve a second playback speed 115 as an effective playback speed overall for a set of audio frames 113.
[0087]
[0107] Referring to Figure 6, a system capable of adjusting the call audio playback speed is shown, the whole of which is shown as 600. In a particular embodiment, system 100 in Figure 1 includes one or more components of system 600.
[0088]
[0108] One or more processors 120 include a hiatus manager 624. Device 102 receives a sequence 119 of audio frames during a call and stores the audio frames of sequence 119 in a receive buffer 134. In a particular embodiment, sequence 119 corresponds to audio frames received during a call with a single additional device, as described with reference to Figures 1 and 2. For example, the call is between two or fewer devices, including device 102 and device 104. In an alternative embodiment, sequence 119 corresponds to audio frames received during a call with multiple second devices, as described with reference to Figure 3. The call manager 122 retrieves the audio frames from the receive buffer 134 and outputs them to the speaker 129. For example, the call manager 122 generates an audio output 143 based on the retrieved audio frames and supplies that audio output 143 to the speaker 129. For example, the call manager 122 outputs a set of audio frames 109, audio frame 111, or a combination thereof at a first playback speed 105.
[0089]
[0109] Device 102 receives a pause command 620 from user 142 (for example, via the user input). The pause command 620 indicates a user request to pause playback of the call. In response to receiving the pause command 620, the pause manager 624 stops the playback output of the audio frames of sequence 119. For example, the pause manager 624 receives the pause command 620 following the playback output of audio frame 111 by call manager 122 at a certain playback point. The pause manager 624 stops call manager 122 from playing and outputting subsequent audio frames of sequence 119. In certain embodiments, the pause manager 624 marks the audio frames following audio frame 111 (for example, set of audio frames 113) that are stored in the receive buffer 134 as unavailable for playback.
[0090]
[0110] The pause manager 624 receives a resume command 622 from the user 142 (for example, via the user input) at the resume time. The resume command 622 indicates a user request to resume playback of the call. In a particular embodiment, the pause manager 624 identifies a set of audio frames 113 as previous audio frames that were missed during playback. For example, the pause manager 624 determines that the set of audio frames 113 should have been played back during the pause duration 623 between the playback duration of the last played audio frame (for example, audio frame 111) and the time of resume.
[0091]
[0111] In certain embodiments, the pause manager 624 performs one or more operations described with reference to the frame loss manager 156, the frame loss manager 124, or both in Figure 1. For example, the pause manager 624 generates a set of audio frames 123 based on a set of audio frames 113, as described with reference to Figure 1. In certain embodiments, the set of audio frames 123 contains the same audio frames as the set of audio frames 113, and the pause manager 624 plays back the set of audio frames 123 (e.g., the set of audio frames 113) at a second playback speed 115. In another embodiment, the set of audio frames 123 contains a subset of the set of audio frames 113 such that the playback output of the set of audio frames 123 at a first playback speed 105 corresponds to the effective playback speed of the set of audio frames 113 (e.g., a second playback speed 115). The pause manager 624 transitions the next audio frame 117 from the second playback speed 115 to the first playback speed 105, as described with reference to Figure 1.
[0092]
[0112] In certain embodiments, the pause manager 624 performs at least partial silence suppression on a set of audio frames 123, on a subset of the next audio frames 117, or both, as described with reference to Figure 1. In certain embodiments, the pause manager 624 determines a second playback speed 115 based on the number of audio frames in the receive buffer 134 available for playback. In certain embodiments, the pause manager 624 determines the second playback speed 115 based on the pause duration 623. For example, a longer pause duration 623 corresponds to a higher second playback speed 115. In certain embodiments, the pause manager 624 selectively generates a set of audio frames 123 based on a set of audio frames 113, as described with reference to Figure 1. For example, the pause manager 624 generates a set of audio frames 123 for playback based on a second playback speed 115 (e.g., actual playback speed or effective playback speed), or jumps to playback of the next audio frame 117 at a first playback speed 105.
[0093]
[0113] In this way, system 600 allows user 142 to pause the call (for example, a real-time call). User 142 can resume the call and catch up on the conversation without having to ask other participants to repeat anything they missed.
[0094]
[0114] Referring to Figure 7, a diagram is shown, with the entire diagram shown as 700. Diagram 700 illustrates an exemplary mode of operation of the components of system 600 in Figure 6. The timing and operation shown in Figure 7 are for illustrative purposes only and are not limiting. In other modes, additional or fewer operations may be performed, and the timing may differ.
[0095]
[0115] During the call, device 102 receives from device 402 audio frames 111, 411, 413, 415, 417, 419, 421, 423, 425, one or more additional audio frames, or a combination thereof, from sequence 119.
[0096]
[0116] Device 102 (for example, call manager 122) plays back audio frame 111 at a first playback speed 105. Device 102 receives a pause command 620, as described with reference to Figure 6, and pauses playback of the call for a pause duration 623. For example, device 102 refrains from playing back audio frames 411, 413, and 415, which would have been played back at the first playback speed 105 during the pause duration 623 during normal operation.
[0097]
[0117] Device 102 receives a resume command 622, as described with reference to Figure 6, and resumes playback of the call. For example, device 102 (e.g., the pause manager 624) generates a set of audio frames 123 (e.g., audio frames 451, 453, and 455) from a set of audio frames 113, as described with reference to Figure 5, and plays back the set of audio frames 123 based on a second playback speed 115.
[0098]
[0118] The pause manager 624 transitions from the second playback speed 115 to the first playback speed 105 for the next audio frame 117. For example, the pause manager 624 generates the next audio frame 495, corresponding to the next audio frame 491, based on the second playback speed 115, as illustrated with reference to Figure 4. For example, the next audio frame 495 is the same as the next audio frame 491 and is played back at the second playback speed 115, or the next audio frame 495 is a subset of the next audio frame 491 such that playing back the next audio frame 495 at the first playback speed 105 corresponds to the effective playback speed of the next audio frame 491 (for example, the second playback speed 115), as illustrated with reference to Figure 5. The pause manager 624 transitions from the second playback speed 115 to the first playback speed 105 by playing back the next audio frame 493 based on the first playback speed 105. For example, the pause manager 624 outputs audio frames 423 and 425 at a first playback speed of 105.
[0099]
[0119] Referring to Figure 8, a specific embodiment of the method 800 for adjusting the call audio playback speed is shown. In a specific embodiment, one or more operations of the method 800 are performed by at least one of the following: the frame loss manager 124 in Figure 1, the call manager 122, one or more processors 120, the device 102, the system 100, the system 200 in Figure 2, the system 300 in Figure 3, or a combination thereof.
[0100]
[0120] Method 800 includes receiving a sequence of audio frames from a first device during a call, in 802. For example, device 102 receives a sequence of audio frames 119 from device 104 during a call, as described with reference to Figure 1.
[0101]
[0121] Method 800 also includes, in response to 804 determining that no audio frames of a sequence have been received during a threshold duration following the last received audio frame of the sequence, initiating the transmission of a frame loss indication to a first device. For example, the frame loss manager 124 in Figure 1, as described with reference to Figure 1, determines that no audio frames of sequence 119 have been received during a threshold duration following the last received frame of sequence 119 (e.g., audio frame 111), initiating the transmission of a frame loss indication 121 to device 104.
[0102]
[0122] Method 800 further includes, in 806, receiving from the first device a set of audio frames for a sequence and an indication of a second playback speed in response to a frame loss indication. For example, device 102 in Figure 1 receives from device 104 a set of audio frames for a sequence 119 and an indication of a second playback speed 115 in response to a frame loss indication 121, as described with reference to Figure 1.
[0103]
[0123] Method 800 also includes, in 808, initiating the playback of a set of audio frames through a speaker based on a second playback speed. For example, the frame loss manager 124, as illustrated with reference to Figure 1, initiates the playback of a set of audio frames 123 by speaker 129 based on a second playback speed 115. The second playback speed 115 is greater than the first playback speed 105 of the set of audio frames 109 of sequence 119. In this way, Method 800 enables the reception of previously lost audio frames and catching up on the call by outputting the audio frames at a faster playback speed.
[0104]
[0124] The method 800 in Figure 8 can be implemented by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, other hardware devices, a firmware device, or any combination thereof. For example, the method 800 in Figure 8 can be implemented by one or more processors that execute instructions, as described with reference to Figure 21.
[0105]
[0125] Referring to Figure 9, a specific embodiment of the method 900 for adjusting the call audio playback speed is shown. In a specific embodiment, one or more operations of the method 900 are performed by at least one of the following: the frame loss manager 156 in Figure 1, the call manager 152, one or more processors 150, the device 104, the system 100 in Figure 2, the system 200 in Figure 2, the system 300 in Figure 3, or a combination thereof.
[0106]
[0126] Method 900 includes, in 902, initiating the transmission of a sequence of audio frames to a second device during a call. For example, device 104 in Figure 1 initiates the transmission of a sequence of audio frames 119 to device 102 during a call, as described with reference to Figure 1.
[0107]
[0127] Method 900 also includes receiving a frame loss indication from a second device in 904. For example, the frame loss manager 156 in Figure 1 receives a frame loss indication 121 from device 102, as described with reference to Figure 1.
[0108]
[0128] Method 900 further includes determining the last received audio frame in the sequence received by the second device based on the frame loss indication in 906. For example, the frame loss manager 156 in Figure 1 determines the last received audio frame (e.g., audio frame 111) in the sequence 119 received by device 102 based on the frame loss indication 121, as described with reference to Figure 1.
[0109]
[0129] Method 900 also includes, at least in part, determining in 908 that a set of audio frames following the last received audio frame in the sequence is available, initiating the transmission of a set of audio frames and an indication of a second playback speed for the set of audio frames to a second device. For example, the frame loss manager 156, at least in part, determining that a set of audio frames 113 following audio frame 111 is available, as illustrated with reference to Figure 1, initiating the transmission of a set of audio frames 113 (for example, set of audio frames 123 is the same as set of audio frames 113) and an indication of a second playback speed 115. The second playback speed 115 is greater than the first playback speed 105 for set of audio frames 109 in sequence 119. In this way, Method 900 enables the retransmission of previously lost audio frames and allows device 102 to catch up on the call by playing the audio frames at a faster playback speed.
[0110]
[0130] The method 900 in Figure 9 can be implemented by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, other hardware devices, a firmware device, or any combination thereof. As an example, the method 900 in Figure 9 can be implemented by one or more processors that execute instructions, as described with reference to Figure 21.
[0111]
[0131] Referring to Figure 10, a specific embodiment of the method 1000 for adjusting the call audio playback speed is shown. In a specific embodiment, one or more operations of the method 1000 are performed by at least one of the following: the frame loss manager 156 in Figure 1, the call manager 152, one or more processors 150, the device 104, the system 100 in Figure 2, the system 200 in Figure 2, the system 300 in Figure 3, or a combination thereof.
[0112]
[0132] Method 1000 includes, in 1002, initiating the transmission of a sequence of audio frames to a second device during a call. For example, device 104 in Figure 1 initiates the transmission of a sequence of audio frames 119 to device 102 during a call, as described with reference to Figure 1.
[0113]
[0133] Method 1000 also includes receiving a frame loss indication from a second device in 1004. For example, the frame loss manager 156 in Figure 1 receives a frame loss indication 121 from device 102, as described with reference to Figure 1.
[0114]
[0134] Method 1000 further includes determining the last received audio frame in a sequence received by a second device based on a frame loss indication in 1006. For example, the frame loss manager 156 in Figure 1 determines the last received audio frame (e.g., audio frame 111) in a sequence 119 received by device 102 based on a frame loss indication 121, as described with reference to Figure 1.
[0115]
[0135] Method 1000 also includes generating an updated set of audio frames based on a set of audio frames such that the first playback speed of the updated set of audio frames corresponds to the effective second playback speed of the set of audio frames, based at least in part on the determination in 1008 that a set of audio frames following the last received audio frame in the sequence is available. For example, the frame loss manager 156 generates a set of audio frames 123 based on a subset of the set of audio frames 113 such that the first playback speed 105 of the set of audio frames 123 corresponds to the effective playback speed (e.g., second playback speed 115) of the set of audio frames 113, based at least in part on the determination that a set of audio frames 113 following audio frame 111 is available, as illustrated with reference to Figure 1. The effective second playback speed (e.g., second playback speed 115) is greater than the first playback speed 105. In this way, method 1000 enables the transmission of audio frames corresponding to lost audio frames, which have a faster effective playback speed to keep up with the call.
[0116]
[0136] Method 1000 in Figure 10 can be implemented by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, other hardware devices, a firmware device, or any combination thereof. As an example, Method 1000 in Figure 10 can be implemented by one or more processors that execute instructions, as described with reference to Figure 21.
[0117]
[0137] Referring to Figure 11, a specific embodiment of the call audio playback speed adjustment method 1100 is shown. In a specific embodiment, one or more operations of method 1100 are performed by at least one of the following: the frame loss manager 124, the call manager 122, one or more processors 120, the device 102, the system 100, the pause manager 624, the system 600, or a combination thereof, as shown in Figure 1.
[0118]
[0138] Method 1100 includes receiving a sequence of audio frames from a first device during a call, as described in 1102. For example, device 102 receives a sequence of audio frames 119 from device 104 during a call, as described with reference to Figure 1.
[0119]
[0139] Method 1100 also includes, in 1104, receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, and initiating playback of a set of audio frames based on a second playback speed. For example, the pause manager 624 in Figure 6, as described with reference to Figure 6, receives a resume command 622 indicating a user request to resume playback and determines that a set of audio frames 113 following the last played audio frame (e.g., audio frame 111) is available, and initiates playback of the set of audio frames 113 based on a second playback speed 115. The second playback speed 115 is greater than the first playback speed 105 of the set of audio frames 109 in sequence 119.
[0120]
[0140] In this way, method 1100 allows the user to pause a call (for example, a real-time call). The user can resume the call and catch up on the conversation without having to ask other participants to repeat anything they missed.
[0121]
[0141] The method 1100 in Figure 11 can be implemented by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, other hardware devices, a firmware device, or any combination thereof. As an example, the method 1100 in Figure 11 can be implemented by one or more processors that execute instructions, as described with reference to Figure 21.
[0122]
[0142] Figure 12 depicts an embodiment 1200 of device 102, device 104 in Figure 1, server 204 in Figure 2, or a combination thereof, as an integrated circuit 1202 including one or more processors 1220. The one or more processors 1220 include a call manager 1222, a frame loss manager 1224, a hiatus manager 624, or a combination thereof.
[0123]
[0143] In certain embodiments, call manager 1222 corresponds to call manager 122, call manager 152 in Figure 1, call manager 222 in Figure 2, or a combination thereof. In certain embodiments, frame loss manager 1224 corresponds to frame loss manager 124, frame loss manager 156 in Figure 1, frame loss manager 224 in Figure 2, or a combination thereof.
[0124]
[0144] The integrated circuit 1202 also includes an audio input section 1204, such as one or more bus interfaces, to enable audio data 1228 (for example, audio input 141) to be received for processing. The integrated circuit 1202 also includes an audio output section 1206, such as a bus interface, to enable the transmission of audio outputs 1243, such as audio output 143. The integrated circuit 1202 can enable the implementation of call audio playback speed adjustment as a component in a system such as a mobile phone or tablet as depicted in Figure 13, a headset as depicted in Figure 14, a wearable electronic device as depicted in Figure 15, a voice-controlled speaker system as depicted in Figure 16, a camera as depicted in Figure 17, a virtual reality headset or augmented reality headset as depicted in Figure 18, or a carrying device as depicted in Figure 19 or Figure 20.
[0125]
[0145] Figure 13 depicts an embodiment 1300 in which device 102, device 104, server 204, or a combination thereof includes a mobile device 1302, such as a telephone or tablet, as an explanatory, non-limiting example. The mobile device 1302 includes a speaker 129, a microphone 146, and a display screen 1304. Components of the processor 1220, including a call manager 1222, a frame loss manager 1224, a hiatus manager 624, or a combination thereof, are incorporated into the mobile device 1302 and are illustrated with dashed lines to show internal components that are not typically visible to the user of the mobile device 1302. In a specific example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of the user's speech, which is then processed by the mobile device 1302 to perform one or more actions, such as displaying information related to the user's speech on the display screen 1304 (for example, by the embedded "Smart Assistant" application).
[0126]
[0146] Figure 14 depicts an embodiment 1400 in which device 102, device 104, server 204, or a combination thereof includes a headset device 1402. The headset device 1402 includes a speaker 129, a microphone 146, or both. One or more components of processor 1220, including a frame loss manager 1224, a pause manager 624, a call manager 1222, or a combination thereof, are incorporated into the headset device 1402. In a particular example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of user voice, which may cause the headset device 1402 to perform one or more operations to send audio data corresponding to user voice activity to a second device (not shown) for further processing, or a combination thereof.
[0127]
[0147] Figure 15 depicts Embodiment 1500, in which device 102, device 104, server 204, or a combination thereof, includes a wearable electronic device 1502, illustrated as a “smartwatch.” A frame loss manager 1224, a pause manager 624, a call manager 1222, a microphone 146, a speaker 129, or a combination thereof are incorporated into the wearable electronic device 1502. In a particular example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of a user's speech, which is then processed by the wearable electronic device 1502 to perform one or more actions, such as launching a graphical user interface or otherwise displaying other information related to the user's speech on the display screen 1504 of the wearable electronic device 1502. For example, the wearable electronic device 1502 may include a display screen configured to display notifications based on the user's speech detected by the wearable electronic device 1502. In a specific example, the wearable electronic device 1502 includes a haptic device that provides haptic notifications (e.g., vibrates) in response to the detection of user voice activity. For example, the haptic notification allows the user to pay attention to the wearable electronic device 1502 and notice a visual notification indicating the detection of a keyword spoken by the user (e.g., to pause or resume playback of a call), or a warning indicating that call audio is being played in catch-up mode (e.g., at an increased speed). In this way, the wearable electronic device 1502 can warn a user with hearing impairment or a user wearing a headset that the user's voice activity has been detected.
[0128]
[0148] Figure 16 shows Embodiment 1600, in which device 102, device 104, server 204, or a combination thereof, includes a wireless speaker and voice-activated device 1602. The wireless speaker and voice-activated device 1602 may have wireless network connectivity and is configured to perform assistant operations. One or more processors 1220, including a frame loss manager 1224, a quiz manager 624, and a call manager 1222, a microphone 146, a speaker 129, or a combination thereof, are included in the wireless speaker and voice-activated device 1602. In a particular example, the frame loss manager 1224, the quiz manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of user speech. During activity, in response to receiving a verbal command identified as user speech, the wireless speaker and voice-activated device 1602 may perform assistant operations, such as by executing a voice-activated system (e.g., an integrated assistant application). Assistant actions may include pausing a call, resuming a paused call at an increased playback speed to catch up, adjusting the temperature, playing music, and turning on lights. For example, an assistant action may be performed in response to receiving a command following a keyword or key phrase (e.g., "Hello Assistant").
[0129]
[0149] Figure 17 depicts an embodiment 1700, which includes a portable electronic device corresponding to the camera device 1702, device 102, device 104, server 204, or a combination thereof. The camera device 1702 includes a frame loss manager 1224, a pause manager 624, a call manager 1222, a microphone 146, a speaker 129, or a combination thereof. In a particular example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of user speech. In operation, in response to receiving an oral command identified as user speech, the camera device 1702 may, as an illustrative example, adjust image or video capture settings, image or video playback settings, or perform actions that respond to the spoken user command, such as an image or video capture command.
[0130]
[0150] Figure 18 depicts an embodiment 1800 in which device 102, device 104, server 204, or a combination thereof, includes a portable electronic device that corresponds to a virtual reality, augmented reality, or mixed reality headset 1802. A frame loss manager 1224, a pause manager 624, a call manager 1222, a microphone 146, a speaker 129, or a combination thereof is incorporated into the headset 1802. In a particular example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof, operate to adjust the call audio playback speed of the user's speech. A visual interface device 1820 is positioned in front of the user's eyes to allow augmented reality or virtual reality images or scenes to be displayed to the user while the headset 1802 is being worn. In a particular example, the visual interface device is configured to display a notification indicating the detected user's speech.
[0131]
[0151] Figure 19 depicts an embodiment 1900 in which device 102, device 104, server 204, or a combination thereof corresponds to or is incorporated within a delivery means 1902 exemplified as a manned or unmanned aerial device (e.g., a parcel delivery drone). A frame loss manager 1224, a pause manager 624, a call manager 1222, a microphone 146, a speaker 129, or a combination thereof is incorporated into the delivery means 1902. In a particular example, the frame loss manager 1224, the pause manager 624, the call manager 1222, or a combination thereof operates to adjust the call audio playback speed of a user's voice in response to a delivery order from an authenticated user of the delivery means 1902, for example.
[0132]
[0152] Figure 20 depicts another embodiment 2000 in which device 102, device 104, server 204, or a combination thereof corresponds to, or is incorporated within, a transport system 2002 exemplified as an automobile. The transport system 2002 includes one or more processors 1220, including a frame loss manager 1224, a quiescent manager 624, a call manager 1222, or a combination thereof. The transport system 2002 also includes a microphone 146, a speaker 129, or both. In a particular example, the frame loss manager 1224, the quiescent manager 624, the call manager 1222, or a combination thereof, operates to adjust the call audio playback speed of a user's voice in response to voice commands from an authenticated passenger, for example. In certain embodiments, upon receiving an oral command identified as user speech, the voice-activated system initiates one or more operations of the carrier 2002 based on one or more keywords (e.g., "unlock", "start engine", "play music", "display weather forecast", or another voice command), such as by providing feedback or information via the display 2020 or one or more speakers (e.g., speaker 129).
[0133]
[0153] Referring to Figure 21, a block diagram of a particular exemplary embodiment of the device is depicted, with the whole being shown as 2100. In various embodiments, device 2100 may have more or fewer components than those shown in Figure 21. In exemplary embodiments, device 2100 may correspond to device 102, device 104, server 204, or a combination thereof. In exemplary embodiments, device 2100 may perform one or more operations described with reference to Figures 1 to 20.
[0134]
[0154] In certain embodiments, device 2100 includes a processor 2106 (e.g., a central processing unit (CPU)). Device 2100 may include one or more additional processors 2110 (e.g., one or more DSPs). In certain embodiments, one or more processors 1220 in Figure 12 correspond to processor 2106, processor 2110, or a combination thereof. Processor 2110 may include a speech and music coder decoder (CODEC) 2108, which includes a speech coder ("vocoder") encoder 2136, a vocoder decoder 2138, a call manager 1222, a frame loss manager 1224, a pause manager 624, or a combination thereof.
[0135]
[0155] Device 2100 may include memory 2186 and CODEC 2134. In certain embodiments, memory 2186 corresponds to memory 154, memory 132 in Figure 1, memory 232 in Figure 2, or a combination thereof. Memory 2186 may include instructions 2156 that can be executed by one or more additional processors 2110 (or processor 2106) to perform the functions described with reference to call manager 1222, frame loss manager 1224, hiatus manager 624, or a combination thereof. Device 2100 may include modem 2140 coupled to antenna 2152 via transceiver 2150.
[0136]
[0156] Device 2100 may include a display 2128 coupled to a display controller 2126. A speaker 129, a microphone 146, or both may be coupled to CODEC 2134. CODEC 2134 may include a digital-to-analog converter (DAC) 2102, an analog-to-digital converter (ADC) 2104, or both. In certain embodiments, CODEC 2134 may receive an analog signal from the microphone 146, convert the analog signal to a digital signal using the analog-to-digital converter 2104, and feed that digital signal to the speech and music codec 2108. The speech and music codec 2108 can process the digital signal, which may be further processed by the call manager 1222. In certain embodiments, the speech and music codec 2108 may feed the digital signal to CODEC 2134. The CODEC 2134 can use the digital-to-analog converter 2102 to convert a digital signal to an analog signal and supply that analog signal to the speaker 129.
[0137]
[0157] In certain embodiments, device 2100 may be included in a system-in-package device or a system-on-chip device 2122. In certain embodiments, memory 2186, processor 2106, processor 2110, display controller 2126, CODEC 2134, modem 2140, and transceiver 2150 are included in the system-in-package device or system-on-chip device 2122. In certain embodiments, input device 2130 and power supply 2144 are coupled to the system-on-chip device 2122. Furthermore, in certain embodiments, as shown in Figure 21, the display 2128, input device 2130, speaker 129, microphone 146, antenna 2152, and power supply 2144 are located outside the system-on-chip device 2122. In certain embodiments, the display 2128, input device 2130, speaker 129, microphone 146, antenna 2152, and power supply 2144 may each be coupled to a component of a system-on-chip device 2122, such as an interface or controller.
[0138]
[0158] Device 2100 may include virtual assistants, home appliances, smart devices, Internet of Things (IoT) devices, communication devices, computers, display devices, televisions, game consoles, music players, radios, video players, entertainment units, personal media players, digital video players, cameras, navigation devices, headsets, smart speakers, speaker covers, mobile communication devices, smartphones, cellular phones, laptop computers, tablets, personal digital assistants, digital video disc (DVD) players, tuners, transporters, augmented reality headsets, virtual reality headsets, mixed reality headsets, aerial transporters, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computer devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.
[0139]
[0159] In relation to the embodiments described, one device includes means for receiving a sequence of audio frames from a first device during a call. For example, the means for receiving may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to receive a sequence of audio frames during a call, or any combination thereof.
[0140]
[0160] The apparatus also includes means for initiating the transmission of frame loss indications to a first device, which is initiated when it is determined that no audio frames of a sequence have been received for a threshold duration following the last received audio frame of the sequence. For example, the means for initiating the transmission of frame loss indications may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to initiate the transmission of frame loss indications, or a combination thereof.
[0141]
[0161] The apparatus further includes means for receiving a set of audio frames in a sequence and an indication of a second playback speed from the first device, the set of audio frames and the indication being received in response to a frame loss indication. For example, the means for receiving the set of audio frames may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to receive the set of audio frames, or any combination thereof.
[0142]
[0162] The apparatus also includes means for initiating playback of a set of audio frames through a speaker based on a second playback speed, the second playback speed being greater than the first playback speed of a first set of audio frames in a sequence. For example, the means for initiating playback may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, processor 2106, one or more processors 2110, device 2100, one or more other circuits or components configured to initiate transmission of frame loss indications, or any combination thereof.
[0143]
[0163] In addition, in relation to the embodiments described, the apparatus includes means for receiving a sequence of audio frames from the first device. For example, the means for receiving may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, a quiz manager 624 in Figure 6, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to receive a sequence of audio frames during a call, or any combination thereof.
[0144]
[0164] The device also includes means for initiating playback of a set of audio frames at at least a second playback speed, in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. For example, the means for initiating playback may correspond to a frame loss manager 124, a call manager 122, one or more processors 120, device 102 in Figure 1, a pause manager 624 in Figure 6, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, device 2100, one or more other circuits or components configured to initiate playback of a set of audio frames, or any combination thereof.
[0145]
[0165] Furthermore, in relation to the embodiments described, the apparatus includes means for initiating the transmission of a sequence of audio frames to a second device during a call. For example, the means for initiating the transmission may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to initiate the transmission of a sequence of audio frames, or any combination thereof.
[0146]
[0166] The apparatus also includes means for receiving frame loss indications from a second device. For example, the means for receiving may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, a server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to receive frame loss indications, or any combination thereof.
[0147]
[0167] The apparatus further includes means for determining the last received audio frame in a sequence received by a second device, based on frame loss indication. For example, the means for determination may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, device 2100, one or more other circuits or components configured to determine the last received audio frame based on frame loss indication, or any combination thereof.
[0148]
[0168] The apparatus also includes means for initiating the transmission of a set of audio frames and an indication of a second playback speed for the set of audio frames to a second device, the transmission being initiated at least in part on the determination that a set of audio frames following the last received audio frame in the sequence is available. The second playback speed is greater than the first playback speed for the first set of audio frames in the sequence. For example, the means for initiating the transmission may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, a server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to initiate the transmission of a set of audio frames, or any combination thereof.
[0149]
[0169] In addition, in relation to the embodiments described, the apparatus includes means for initiating the transmission of a sequence of audio frames to a second device during a call. For example, the means for initiating the transmission may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to initiate the transmission of a sequence of audio frames, or any combination thereof.
[0150]
[0170] The apparatus also includes means for receiving frame loss indications from a second device. For example, the means for receiving may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, a server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to receive frame loss indications, or any combination thereof.
[0151]
[0171] The device further includes means for determining the last received audio frame in a sequence received by a second device, based on frame loss indication. For example, the means for determination may correspond to a frame loss manager 156, a call manager 152, one or more processors 150, device 104 in Figure 1, a call manager 222, a frame loss manager 224, one or more processors 220, server 204 in Figure 2, one or more processors 1220, a call manager 1222, a frame loss manager 1224 in Figure 12, device 2100, one or more other circuits or components configured to determine the last received audio frame based on frame loss indication, or any combination thereof.
[0152]
[0172] The apparatus also includes means for generating an updated set of audio frames based on a set of audio frames such that a first playback speed of the updated set of audio frames corresponds to an effective second playback speed of the set of audio frames, the updated set of audio frames is generated at least in part on the determination that a set of audio frames following the last received audio frame in the sequence is available. The effective second playback speed is greater than the first playback speed. For example, the means for generation may correspond to a frame loss manager 156, one or more processors 150, device 104 in Figure 1, a frame loss manager 224, one or more processors 220, a server 204 in Figure 2, one or more processors 1220, a frame loss manager 1224 in Figure 12, a transceiver 2150, a modem 2140, an antenna 2152, device 2100, one or more other circuits or components configured to generate an updated set of audio frames, or any combination thereof.
[0153]
[0173] In some embodiments, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 2186) includes an instruction (e.g., instruction 2156) which, when executed by one or more processors (e.g., one or more processors 2110 or processor 2106), causes one or more processors to receive a sequence of audio frames (e.g., sequence 119) from a first device (e.g., device 104) during a call. The instruction also, when executed by one or more processors, causes one or more processors to initiate sending a frame loss indication (e.g., frame loss indication 121) to the first device in response to determining that no audio frames of the sequence have been received during a threshold duration since the last received audio frame of the sequence (e.g., set of audio frames 111). The instruction, when executed by one or more processors, will receive from the first device, in response to a frame loss indication, an indication of a set of audio frames in the sequence (e.g., a set of audio frames 113) and an indication of a second playback speed (e.g., a second playback speed 115). The instruction, when executed by one or more processors, will also cause one or more processors to begin playback of the set of audio frames via a speaker (e.g., speaker 129) based on the second playback speed. The second playback speed is greater than the first playback speed (e.g., first playback speed 105) of the first set of audio frames in the sequence (e.g., a set of audio frames 109).
[0154]
[0174] In some embodiments, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 2186) includes an instruction (e.g., instruction 2156) which, when executed by one or more processors (e.g., one or more processors 2110 or processor 2106), causes one or more processors to receive a sequence of audio frames (e.g., sequence 119) from a first device (e.g., device 104) during a call. This instruction also, when executed by one or more processors, causes one or more processors to begin playing a set of audio frames based on a second playback rate (e.g., second playback rate 115) in response to receiving a user request to resume playback (e.g., resume command 622) and determining that a set of audio frames (e.g., set of audio frames 113) following the last played audio frame in the sequence (e.g., audio frame 111) is available. The second playback rate is greater than the first playback rate (e.g., first playback rate 105) of the first set of audio frames in the sequence (e.g., set of audio frames 109).
[0155]
[0175] In some embodiments, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 2186) includes an instruction (e.g., instruction 2156) which, when executed by one or more processors (e.g., one or more processors 2110 or processor 2106), causes one or more processors to begin sending a sequence of audio frames (e.g., sequence 119) to a second device (e.g., device 102) during a call. This instruction also, when executed by one or more processors, causes one or more processors to receive a frame loss indication (e.g., frame loss indication 121) from the second device (e.g., device 102). The instruction further, when executed by one or more processors, causes one or more processors to determine, based on the frame loss indication, the last audio frame (e.g., audio frame 111) of the sequence received by the second device. The instruction also, when executed by one or more processors, will initiate the transmission of a set of audio frames to a second device and an indication of a second playback speed for that set of audio frames (e.g., second playback speed 115), at least in part on the fact that one or more processors have determined that a set of audio frames following the last received audio frame in the sequence (e.g., set of audio frames 113) is available. The second playback speed is greater than the first playback speed (e.g., first playback speed 105) of the first set of audio frames in the sequence (e.g., set of audio frames 109).
[0156]
[0176] In some embodiments, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 2186) includes an instruction (e.g., instruction 2156) which, when executed by one or more processors (e.g., one or more processors 2110 or processor 2106), causes one or more processors to begin sending a sequence of audio frames (e.g., sequence 119) to a second device (e.g., device 102) during a call. This instruction also, when executed by one or more processors, causes one or more processors to receive a frame loss indication (e.g., frame loss indication 121) from the second device (e.g., device 102). The instruction further, when executed by one or more processors, causes one or more processors to determine, based on the frame loss indication, the last audio frame (e.g., audio frame 111) of the sequence received by the second device. The instruction also, when executed by one or more processors, will generate an updated set of audio frames (e.g., set of audio frames 123) based on the set of audio frames such that the first playback speed (e.g., first playback speed 105) of the updated set of audio frames corresponds to the effective second playback speed (e.g., second playback speed 115) of the set of audio frames, at least in part on the fact that one or more processors have determined that a set of audio frames following the last received audio frame in the sequence (e.g., set of audio frames 113) is available, such that the first playback speed (e.g., first playback speed 105) of the updated set of audio frames corresponds to the effective second playback speed (e.g., second playback speed 115) of the set of audio frames, where the effective second playback speed is greater than the first playback speed.
[0157]
[0177] Specific aspects of this disclosure are described below, in the first set of interrelated provisions.
[0158]
[0178] According to Clause 1, a communication device comprises one or more processors configured to receive a sequence of audio frames from a first device during a call; to initiate a transmission of a frame loss indication to the first device in response to determining that no audio frames of the sequence have been received during a threshold duration following the last received audio frame of the sequence; to receive a set of audio frames of the sequence and an indication of a second playback speed from the first device in response to the frame loss indication; and to initiate playback of the set of audio frames via a speaker based on the second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames of the sequence.
[0159]
[0179] Clause 2 includes the device described in Clause 1, and includes initiating playback of a set of audio frames at a second playback speed, thereby at least partially suppressing the silence represented by the set of audio frames.
[0160]
[0180] Clause 3 includes the device described in Clause 1 or Clause 2, wherein one or more processors are further configured to receive the next audio frame of a sequence from the first device, the next audio frame following a set of audio frames in the sequence, and one or more processors are further configured to transition the playback of the next audio frame from a second playback speed to a first playback speed.
[0161]
[0181] Clause 4 includes the device described in any of Clauses 1 to 3, and one or more processors are further configured to store a set of audio frames in a buffer; to receive the next audio frame of a sequence from the first device and store the next audio frame in a buffer, the next audio frame following a set of audio frames in the sequence, and one or more processors are further configured to transition playback from a second playback speed to a first playback speed when they determine that fewer than a threshold number of audio frames of the sequence are stored in the playback buffer.
[0162]
[0182] Clause 5 includes the device described in any of Clauses 1 to 4, wherein one or more processors are further configured to receive a set of audio frames from the first device and, at the same time, receive the next audio frame of the sequence from the first device, the next audio frame following the set of audio frames of the sequence.
[0163]
[0183] Clause 6 includes a device as described in any of Clauses 1 to 5, wherein one or more processors are configured to determine, at a given time, that no audio frame corresponding to a given playback duration at a first playback speed has been received since the last received audio frame, the given playback duration being based on a given time and the time of reception of the last received frame, and the frame loss indication includes a request to retransmit a previous audio frame corresponding to a given playback duration at a first playback speed.
[0164]
[0184] Clause 7 includes a device described in any of Clauses 1 through 6, where frame loss indication indicates the last received audio frame.
[0165]
[0185] Clause 8 includes the devices described in any of Clauses 1 through 6, and the second playback speed is based on the number of sets of audio frames.
[0166]
[0186] Clause 9 includes the devices described in any of Clauses 1 through 8, wherein one or more processors are configured to initiate playback of silent frames at a playback speed greater than a second playback speed, depending on whether the set of audio frames includes silent frames.
[0167]
[0187] Clause 10 includes devices described in any of Clauses 1 through 9, where one or more processors are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof.
[0168]
[0188] Specific aspects of this disclosure are described below in the second set of interrelated clauses.
[0169]
[0189] According to Clause 11, the method of communication comprises, during a call, receiving a sequence of audio frames from a first device, determining that no audio frames of the sequence have been received by the device during a threshold duration following the last received audio frame of the sequence, initiating the transmission of a frame loss indication from the device to the first device, receiving a set of audio frames of the sequence and an indication of a second playback speed from the first device in response to the frame loss indication, and initiating playback of the set of audio frames via a speaker based on the second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames of the sequence.
[0170]
[0190] Clause 12 includes the method described in Clause 11, wherein the first device includes a server, the call is with a plurality of second devices, and the sequence of audio frames is based on a plurality of sequences of audio frames received by the first device from the plurality of second devices.
[0171]
[0191] Clause 13 includes the method described in Clause 11, wherein a call is made between the device and a single additional device, and the single additional device includes the first device.
[0172]
[0192] Clause 14 includes a method according to any of Clauses 11-13, the method further comprising receiving a set of audio frames in one device from a first device and simultaneously receiving the next audio frame of a sequence in one device from a first device, the next audio frame following the set of audio frames of a sequence.
[0173]
[0193] Specific aspects of this disclosure are described below in a third set of interrelated clauses.
[0174] According to Clause 15, a communication device comprises one or more processors configured to receive a sequence of audio frames from a first device during a call; and to begin playback of a set of audio frames based on a second playback speed, in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence.
[0175]
[0194] Clause 16 includes the device described in Clause 15, and includes initiating playback of a set of audio frames at a second playback speed to at least partially suppress the silence represented by the set of audio frames.
[0176]
[0195] Clause 17 includes the device described in Clause 15 or Clause 16, wherein one or more processors are further configured to receive the next audio frame of a sequence from the first device, the next audio frame following a set of audio frames in the sequence, and one or more processors are further configured to transition the playback of the next audio frame from a second playback speed to a first playback speed.
[0177]
[0196] Clause 18 includes a device described in any of Clauses 15 to 17, wherein one or more processors are further configured to store a set of audio frames in a buffer; to receive the next audio frame of a sequence from the first device and store the next audio frame in a buffer, the next audio frame following a set of audio frames in the sequence, and one or more processors are further configured to transition playback from a second playback speed to a first playback speed when they determine that fewer than a threshold number of audio frames of the sequence are stored in the playback buffer.
[0178]
[0197] Clause 19 includes the devices described in any of Clauses 15 through 18, where the second playback speed is based on the number of sets of audio frames.
[0179]
[0198] Clause 20 includes a device described in any of Clauses 15 to 18, where a user request is received at the time of resumption, and the second playback speed is based on the pause duration between the time of playback of the last played audio frame and the time of resumption.
[0180]
[0199] Clause 21 includes a device described in any of Clauses 15 to 20, wherein one or more processors are configured to start playback of silent frames at a playback speed greater than a second playback speed, depending on whether the set of audio frames includes silent frames.
[0181]
[0200] Clause 22 includes devices described in any of Clauses 15 to 21, where one or more processors are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof.
[0182]
[0201] Specific aspects of this disclosure are described below, in the fourth set of interrelated clauses.
[0183]
[0202] According to Clause 23, the method of communication comprises, during a call, receiving a sequence of audio frames from a first device, receiving a user request to resume playback, and determining that a set of audio frames following the last played audio frame in the sequence is available, and initiating playback of the set of audio frames on the device at at least a second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence.
[0184]
[0203] Clause 24 includes the method described in Clause 23, which includes starting playback of a set of audio frames at a second playback speed, thereby suppressing at least partially any silence represented by the set of audio frames.
[0185]
[0204] Clause 25 includes the method described in Clause 23 or Clause 24, the method further comprising a device receiving the next audio frame of a sequence from a first device, the next audio frame following a set of audio frames in the sequence, and the method further comprising a device transitioning the playback of the next audio frame from a second playback speed to a first playback speed.
[0186]
[0205] Clause 26 includes a method of any of Clauses 23 to 25, the method further comprising storing a set of audio frames in a buffer, receiving the next audio frame of a sequence from a first device in a device, and storing the next audio frame in a buffer, the next audio frame following a set of audio frames in a sequence, and the method further comprises, in response to determining that fewer than a threshold number of audio frames of a sequence are stored in the buffer for playback, transitioning playback from a second playback speed to a first playback speed in the device.
[0187]
[0206] Clause 27 includes the method described in any of Clauses 23 to 26, wherein the second playback speed is based on the number of sets of audio frames.
[0188]
[0207] Clause 28 includes the method described in any of Clauses 23 to 26, wherein the user request is received at the time of resumption, and the second playback speed is based on the pause duration between the time of playback of the last played audio frame and the time of resumption.
[0189]
[0208] Clause 29 includes a method of any of Clauses 23 to 28, further comprising, depending on the determination that a set of audio frames includes silent frames, starting playback of silent frames at a playback speed greater than a second playback speed in the device.
[0190]
[0209] Clause 30 includes the method described in any of Clauses 23 to 29, wherein the first device includes a server, the call is to a plurality of second devices, and the sequence of audio frames is based on a plurality of sequences of audio frames received by the first device from the plurality of second devices.
[0191]
[0210] Specific aspects of this disclosure are described below in the fifth set of interrelated clauses.
[0192]
[0211] According to Clause 31, a non-temporary computer-readable medium, when executed by one or more processors, stores instructions that, during a call, one or more processors will receive a sequence of audio frames from a first device; determine that no audio frames of the sequence have been received for a threshold duration since the last audio frame received in the sequence, initiate sending a frame loss indication to the first device; receive from the first device a set of audio frames of the sequence and an indication of a second playback speed in response to the frame loss indication; and initiate playback of the set of audio frames through a speaker based on the second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames of the sequence.
[0193]
[0212] Clause 32 includes, with respect to the non-temporary computer-readable media described in Clause 31, initiating playback of a set of audio frames at a second playback speed, thereby suppressing at least partially any silence represented by the set of audio frames.
[0194]
[0213] Specific aspects of this disclosure are described below in the sixth set of interrelated provisions.
[0195]
[0214] According to Article 33, when a non-temporary computer-readable medium is executed by one or more processors, one or more processors will, during a call, receive a sequence of audio frames from a first device; and store instructions to begin playback of a set of audio frames based on a second playback speed, in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, the second playback speed being greater than the first playback speed of the first set of audio frames in the sequence.
[0196]
[0215] Clause 34 includes, with respect to the non-temporary computer-readable media described in Clause 33, initiating playback of a set of audio frames at a second playback speed, thereby suppressing at least partially any silence represented by the set of audio frames.
[0197]
[0216] Specific aspects of this disclosure are described below in the seventh set of interrelated clauses.
[0198]
[0217] According to Clause 35, the device comprises means for receiving a sequence of audio frames from a first device during a call, and means for initiating the transmission of a frame loss indication to the first device, which is initiated in response to the determination that no audio frames of the sequence have been received for a threshold duration following the last received audio frame of the sequence, and the device further comprises means for receiving a set of audio frames of the sequence and an indication of a second playback speed from the first device, which is received in response to the frame loss indication, and the device further comprises means for initiating playback of the set of audio frames via a speaker based on a second playback speed, the second playback speed being greater than the first playback speed of the first set of audio frames of the sequence.
[0199]
[0218] Clause 36 includes the devices described in Clause 35, wherein means for receiving a sequence, means for initiating the transmission of a frame loss indication, means for receiving a set of audio frames, and means for initiating playback are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof.
[0200]
[0219] Specific aspects of this disclosure are described below in the eighth set of interrelated provisions.
[0201]
[0220] According to Clause 37, the device includes means for receiving a sequence of audio frames from a first device during a call, and means for initiating playback of the set of audio frames based on a second playback speed, the playback being initiated in response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, the second playback speed being greater than the first playback speed of the first set of audio frames in the sequence.
[0202]
[0221] Clause 38 includes the devices described in Clause 37, where means for receiving and means for initiating playback are incorporated into at least one of the following: a virtual assistant, home appliance, smart device, Internet of Things (IoT) device, communication device, computer, display device, television, game console, music player, radio, video player, entertainment unit, personal media player, digital video player, camera, navigation device, headset, or any combination thereof.
[0203]
[0222] Specific aspects of this disclosure are described below in the ninth set of interrelated clauses.
[0204]
[0223] According to Clause 39, a communication device includes one or more processors configured to initiate transmission of a set of audio frames and an indication of a second playback speed for that set of audio frames to a second device, at least in part on the determination that a set of audio frames following the last received audio frame in the sequence is available, during a call, to initiate transmission of a set of audio frames and an indication of a second playback speed for that set of audio frames, where the second playback speed is greater than the first playback speed for the first set of audio frames in the sequence.
[0205]
[0224] Clause 40 includes the device of Clause 39, in which one or more processors are incorporated into the server.
[0206]
[0225] Clause 41 includes the device described in Clause 39 or Clause 40, wherein frame loss indication includes a request to retransmit a previous audio frame corresponding to a specific playback duration at a first playback speed, one or more processors are configured to determine the last received audio frame based on a specific playback duration, and one or more processors are configured to initiate the transmission of a set of audio frames, at least in part on having determined that the previous audio frame contains a set of audio frames.
[0207]
[0226] Clause 42 includes a device described in any of Clauses 39 to 41, which further comprises a buffer, and one or more processors are configured to receive a sequence of audio frames from the first device, and the transmission of a set of audio frames is initiated at least in part on the determination that a set of audio frames is available in the buffer.
[0208]
[0227] Clause 43 includes a device described in any of Clauses 39 to 42, wherein one or more processors are further configured to simultaneously begin transmitting the next audio frame in a sequence while transmitting a set of audio frames.
[0209]
[0228] Clause 44 includes the device described in Clause 43, and the next audio frame is associated with a transition to the first playback speed following the playback of a set of audio frames at the second playback speed.
[0210]
[0229] Clause 45 includes a device described in any of Clauses 39 through 44, wherein the first set of audio frames precedes the last received audio frame in the sequence.
[0211]
[0230] Clause 46 includes a device described in any of Clauses 39 to 45, wherein one or more processors are configured to selectively initiate transmission of a set of audio frames in response to determining that the number of audio frames in a set is below a threshold.
[0212]
[0231] Clause 47 includes the device described in any of Clauses 39 to 46, wherein one or more processors are configured to determine a second playback speed based on the number of sets of audio frames.
[0213]
[0232] Clause 48 includes a device described in any of Clauses 39 to 47, and one or more processors are configured to initiate transmission of a set of audio frames, at least in part, based on the determination that a network link with a second device has been re-established.
[0214]
[0233] Clause 49 includes a device as described in any of Clauses 39 to 48, wherein the set of audio frames includes silent frames, and one or more processors are further configured to transmit indications of a third playback speed for silent frames, the third playback speed being greater than a second playback speed.
[0215]
[0234] Specific aspects of this disclosure are described below in the tenth set of interrelated clauses.
[0216]
[0235] According to Clause 50, a communication device comprises one or more processors configured to, during a call, initiate transmission of a sequence of audio frames to a second device; receive a frame loss indication from the second device; determine, based on the frame loss indication, the last audio frame received in the sequence received by the second device; and, at least in part, determine that a set of audio frames following the last audio frame received in the sequence is available, and generate an updated set of audio frames based on the set of audio frames such that a first playback speed of the updated set of audio frames corresponds to an effective second playback speed of the set of audio frames, where the effective second playback speed is greater than the first playback speed.
[0217]
[0236] Clause 51 includes the device described in Clause 50, wherein one or more processors are configured to generate an updated set of audio frames by selecting a subset of the set of audio frames such that the subset has the same playback duration at a first playback speed as the set of audio frames at a second playback speed.
[0218]
[0237] Clause 52 includes the devices described in Clause 50 or Clause 51, and one or more processors are incorporated into the server.
[0219]
[0238] Clause 53 includes a device as described in any of Clauses 50 to 52, wherein frame loss indication includes a request to retransmit a previous audio frame corresponding to a specific playback duration at a first playback speed, one or more processors are configured to determine the last received audio frame based on a specific playback duration, and one or more processors are configured to initiate the transmission of an updated set of audio frames, at least in part on having determined that the previous audio frame contained a set of audio frames.
[0220]
[0239] Clause 54 includes a device described in any of Clauses 50 to 53, which further comprises a buffer, and one or more processors are configured to receive a sequence of audio frames from the first device, and the transmission of an updated set of audio frames is initiated at least in part on the determination that the set of audio frames is available in the buffer.
[0221]
[0240] Clause 55 includes the devices described in any of Clauses 50 to 54, wherein one or more processors are further configured to simultaneously begin transmitting the next audio frame in a sequence while transmitting an updated set of audio frames.
[0222]
[0241] Clause 56 includes a device described in any of Clauses 50 to 55, wherein one or more processors are configured to selectively initiate transmission of a set of audio frames in response to determining that the number of updated sets of audio frames is below a threshold.
[0223]
[0242] Clause 57 includes the device described in any of Clauses 50 to 56, wherein one or more processors are configured to determine an effective second playback speed based on the number of sets of audio frames.
[0224]
[0243] Clause 58 includes a device described in any of Clauses 50 to 57, and one or more processors are configured to initiate transmission of an updated set of audio frames, at least in part, based on the determination that a network link with a second device has been re-established.
[0225]
[0244] Clause 59 includes a device described in any of Clauses 50 to 58, and includes generating an updated set of audio frames, which suppresses at least a portion of the silence present in the set of audio frames.
[0226]
[0245] Specific aspects of this disclosure are described below in the eleventh set of interrelated provisions.
[0227]
[0246] According to Clause 61, the method of communication comprises, during a call, initiating the transmission of a sequence of audio frames from the first device to the second device; receiving a frame loss indication from the second device to the first device; determining, based on the frame loss indication, the last audio frame received by the second device in the sequence; and, at least in part, determining that a set of audio frames following the last audio frame received in the sequence is available, initiating the transmission of a set of audio frames from the first device to the second device, and transmitting an indication of a second playback speed for the set of audio frames, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence.
[0228]
[0247] Clause 62 includes the method of Clause 61, wherein the first device is incorporated into the server.
[0229]
[0248] Specific aspects of this disclosure are described below in 12 sets of interrelated clauses.
[0230]
[0249] According to Article 63, the method of communication comprises, during a call, initiating the transmission of a sequence of audio frames to a second device; receiving a frame loss indication from the second device; determining, based on the frame loss indication, the last audio frame received in the sequence received by the second device; and generating an updated set of audio frames based on the set of audio frames such that the first playback speed of the updated set of audio frames corresponds to an effective second playback speed of the set of audio frames, where the effective second playback speed is greater than the first playback speed.
[0231]
[0250] Clause 64 includes the method of Clause 63, which generates an updated set of audio frames, and includes selecting a subset of the set of audio frames such that the subset has the same playback duration at a first playback speed as the set of audio frames at a second playback speed.
[0232]
[0251] Specific aspects of this disclosure are described below in 13 sets of interrelated clauses.
[0233]
[0252] According to Article 65, when a non-temporary computer-readable medium is executed by one or more processors, the one or more processors store instructions to initiate transmission of a sequence of audio frames to a second device during a call; receive a frame loss indication from the second device; determine, based on the frame loss indication, the last audio frame received by the second device in the sequence; and, at least in part, determine that a set of audio frames following the last audio frame received in the sequence is available, to initiate transmission of the set of audio frames to the second device, along with an indication of a second playback speed for that set of audio frames, where the second playback speed is greater than the first playback speed for the first set of audio frames in the sequence.
[0234]
[0253] Clause 66 includes the non-temporary computer-readable storage medium described in Clause 65, and one or more processors are incorporated into the server.
[0235]
[0254] Specific aspects of this disclosure are described below in 14 sets of interrelated clauses.
[0236]
[0255] According to Article 67, a non-temporary computer-readable storage medium, when executed by one or more processors, stores instructions that, during a call, one or more processors will initiate transmission of a sequence of audio frames to a second device; receive a frame loss indication from the second device; determine, based on the frame loss indication, the last audio frame received in the sequence received by the second device; and, at least in part, determine that a set of audio frames following the last audio frame received in the sequence is available, to generate an updated set of audio frames based on the set of audio frames such that the first playback rate of the updated set of audio frames corresponds to the effective second playback rate of the set of audio frames, the effective second playback rate being greater than the first playback rate.
[0237]
[0256] Clause 68 includes generating an updated set of audio frames, which includes a non-temporary computer-readable storage medium as described in Clause 67, and includes selecting a subset of the set of audio frames such that the subset has the same playback duration at a first playback speed as the set of audio frames at a second playback speed.
[0238]
[0257] Specific aspects of this disclosure are described below in 15 sets of interrelated clauses.
[0239]
[0258] According to Clause 69, the device includes means for initiating the transmission of a sequence of audio frames to a second device during a call; means for receiving a frame loss indication from the second device; means for determining, based on the frame loss indication, the last audio frame received in the sequence received by the second device; and means for initiating the transmission of a set of audio frames and an indication of a second playback speed for the set of audio frames to the second device, the transmission being initiated at least in part on the determination that a set of audio frames following the last audio frame received in the sequence is available, and the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence.
[0240]
[0259] Clause 70 includes the devices described in Clause 69, where means for initiating transmission of a sequence of audio frames, means for determining, and means for initiating transmission of a set of audio frames and indications are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or any combination thereof.
[0241]
[0260] Specific aspects of this disclosure are described below in 16 sets of interrelated clauses.
[0242]
[0261] According to Clause 71, the device includes means for initiating the transmission of a sequence of audio frames to a second device during a call; means for receiving a frame loss indication from the second device; means for determining, based on the frame loss indication, the last audio frame received in the sequence received by the second device; and means for generating an updated set of audio frames based on a set of audio frames such that a first playback rate of the updated set of audio frames corresponds to an effective second playback rate of the set of audio frames, the updated set of audio frames being generated at least in part on the determination that a set of audio frames following the last audio frame received in the sequence is available, and the effective second playback rate is greater than the first playback rate.
[0243]
[0262] Clause 72 includes the devices described in Clause 71, and means for initiating the transmission of a sequence of audio frames, means for receiving a frame loss indication, means for determining, and means for generating an updated set of audio frames, which are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or any combination thereof.
[0244]
[0263] Those skilled in the art will further understand that various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by one or more processors, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the individual application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in various ways for individual applications, but such decisions should not be construed as resulting in a departure from the scope of this disclosure.
[0245]
[0264] Steps of methods or algorithms described in relation to embodiments disclosed herein may be embodied directly as hardware, as software modules executed by one or more processors, or as a combination of both. The software modules may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM®), registers, hard disks, removable disks, compact disk read-only memory (CD-ROM), or any other form of non-temporary recording medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information to and write information to the storage medium. In an alternative configuration, the storage medium may be integrated with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computer device or user terminal. In an alternative configuration, the processor and storage medium may reside as discrete components in a computer device or user terminal.
[0246]
[0265] The foregoing description of the disclosed embodiments is provided to enable those skilled in the art to manufacture or use the disclosed embodiments. Various modifications to these embodiments will be immediately apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of this disclosure. This disclosure is not limited to the embodiments shown herein and is provided to the broadest possible scope that is consistent with the principles and novel features defined by the appended claims. The invention described in the original claims of this application is listed below. [C1] A communication device, The system comprises one or more processors, and the one or more processors are During the call, Receiving a sequence of audio frames from the first device, In response to determining that no audio frames of the sequence have been received during the threshold duration following the last received audio frame of the sequence, the transmission of a frame loss indication to the first device is initiated. In response to the frame loss indication, the first device receives a set of audio frames for the sequence and an indication of a second playback speed. The playback of the set of audio frames via the speaker is initiated based on the second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. A communication device configured to perform the following actions. [C2] The device according to C1, wherein playback of the set of audio frames is initiated at the second playback speed, which includes at least partially suppressing the silence represented by the set of audio frames. [C3] The one or more processors further include: The next audio frame of the sequence is received from the first device, and the next audio frame follows the set of audio frames in the sequence. The playback of the next audio frame is transitioned from the second playback speed to the first playback speed, The device described in C1, configured to perform the following actions. [C4] The one or more processors further include: The set of audio frames is stored in a buffer, The next audio frame of the sequence is received from the first device, the next audio frame is stored in the buffer, and the next audio frame follows the set of audio frames in the sequence. In response to determining that fewer than a threshold number of audio frames of the sequence are stored in the buffer for playback, playback is shifted from the second playback speed to the first playback speed, The device described in C1, configured to perform the following actions. [C5] The one or more processors further include: It is configured to receive the set of audio frames from the first device and, at the same time, receive the next audio frame of the sequence from the first device. The next audio frame is the device C1, which follows the set of audio frames in the sequence. [C6] The one or more processors are configured to determine, at a specific point in time, that no audio frames corresponding to a specific playback duration at the first playback speed have been received since the last received audio frame. The aforementioned specific playback duration is determined based on the aforementioned specific time point and the time of reception of the last received frame. The frame loss indication includes a request to retransmit the previous audio frame, corresponding to the specific playback duration at the first playback speed. The device described in C1. [C7] The frame loss indication is the device described in C1, which indicates the last received audio frame. [C8] The second playback speed is based on the number of sets of audio frames, as described in C1. [C9] The device according to C1, wherein one or more processors are configured to start playback of the silent frames at a playback speed greater than the second playback speed in response to determining that the set of audio frames includes silent frames. [C10] The device described in C1, wherein one or more processors are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof. [C11] A method of communication, During the call, In the device, receiving a sequence of audio frames from the first device, In response to determining that no audio frames of the sequence have been received by the device during a threshold duration following the last received audio frame of the sequence, the device initiates transmission of a frame loss indication to the first device. In response to the frame loss indication, the device receives from the first device a set of audio frames for the sequence and an indication of a second playback speed. The playback of the set of audio frames via the speaker is initiated based on the second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. A method that includes [a certain feature]. [C12] The first device includes a server, and the call is a call with a plurality of second devices. The sequence of audio frames is based on a plurality of sequences of audio frames received by the first device from the plurality of second devices. Method described in C11. [C13] The method according to C11, wherein the call is made between the device and a single additional device, the single additional device including the first device. [C14] The method according to C11, further comprising receiving the set of audio frames from the first device in the device, and simultaneously receiving the next audio frame of the sequence from the first device in the device, wherein the next audio frame follows the set of audio frames of the sequence. [C15] A communication device, The system comprises one or more processors, and the one or more processors are During the call, Receiving a sequence of audio frames from the first device, In response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, playback of the set of audio frames is started based on a second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. A communication device configured to perform the following actions. [C16] The device according to C15, wherein playback of the set of audio frames is initiated at the second playback speed, which includes at least partially suppressing the silence represented by the set of audio frames. [C17] The one or more processors further include: The next audio frame of the sequence is received from the first device, and the next audio frame follows the set of audio frames in the sequence. The playback of the next audio frame is transitioned from the second playback speed to the first playback speed, A device as described in C15, configured to perform the following actions. [C18] The one or more processors further include: The set of audio frames is stored in a buffer, The next audio frame of the sequence is received from the first device, the next audio frame is stored in the buffer, and the next audio frame follows the set of audio frames in the sequence. In response to determining that fewer than a threshold number of audio frames of the sequence are stored in the buffer for playback, playback is shifted from the second playback speed to the first playback speed, A device as described in C15, configured to perform the following actions. [C19] The second playback speed is based on the number of sets of audio frames, as described in C15. [C20] The aforementioned user request is received at the time of resumption. The device according to C15, wherein the second playback speed is based on the pause duration between the playback time of the last played audio frame and the restart time. [C21] The device according to C15, wherein one or more processors are configured to initiate playback of the silent frames at a playback speed greater than the second playback speed in response to determining that the set of audio frames includes silent frames. [C22] The device described in C15, wherein one or more processors are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof. [C23] A method of communication, During the call, In the device, receiving a sequence of audio frames from the first device, In response to receiving a user request to resume playback and determining that a set of audio frames following the last played audio frame in the sequence is available, playback of the set of audio frames on the device is started at at least a second playback speed, wherein the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. A method that includes [a certain feature]. [C24] The method according to C23, wherein playback of the set of audio frames is initiated at the second playback speed, which includes at least partially suppressing the silence represented by the set of audio frames. [C25] The device receives the next audio frame of the sequence from the first device, and the next audio frame follows the set of audio frames in the sequence. In the device, the playback of the next audio frame is transitioned from the second playback speed to the first playback speed. A method using C23 that further incorporates these features. [C26] The set of audio frames is stored in a buffer, The next audio frame of the sequence is received by the device from the first device, The following means storing the next audio frame in the buffer, and the next audio frame following the set of audio frames in the sequence, In response to determining that fewer than a threshold number of audio frames of the sequence are stored in the buffer for playback, the device transitions playback from the second playback speed to the first playback speed, A method using C23 that further incorporates these features. [C27] The method according to C23, wherein the second playback speed is based on the number of sets of audio frames. [C28] The aforementioned user request is received at the time of resumption. The method according to C23, wherein the second playback speed is based on the pause duration between the playback time of the last played audio frame and the restart time. [C29] The method of C23, further comprising determining that the set of audio frames includes silent frames, and in response to determining that the set of silent frames includes silent frames, initiating playback of the silent frames in the device at a playback speed greater than the second playback speed. [C30] The first device includes a server, and the call is a call with a plurality of second devices. The sequence of audio frames is based on a plurality of sequences of audio frames received by the first device from the plurality of second devices. Methods used in C23.
Claims
1. A communication device, The system comprises one or more processors, and the one or more processors are During the call, Receiving a first set of audio frames of an audio frame sequence from a first device, In response to determining that no audio frames of the sequence have been received during the threshold duration following the last received audio frame of the sequence, the transmission of a frame loss indication to the first device is initiated. In response to the frame loss indication, the first device receives a second set of audio frames of the sequence and a second playback speed indication, wherein the second set of audio frames follows the last received audio frame of the sequence. The playback of the second set of audio frames is initiated via a speaker based on the second playback speed, and the second playback speed is greater than the first playback speed of the first set of audio frames. A communication device configured to perform the following actions.
2. The one or more processors described above further include: It is configured to receive the second set of audio frames from the first device and, at the same time, receive the next audio frame of the sequence from the first device. The device according to claim 1, wherein the next audio frame follows the second set of audio frames of the sequence.
3. The one or more processors are configured to determine, at a specific point in time, that no audio frames corresponding to a specific playback duration at the first playback speed have been received since the last received audio frame. The aforementioned specific playback duration is determined based on the aforementioned specific time point and the time of reception of the last received frame. The frame loss indication includes a request to retransmit the previous audio frame, corresponding to the specific playback duration at the first playback speed. The device according to claim 1.
4. The device according to claim 1, wherein the frame loss indication indicates the last received audio frame.
5. A method of communication, During the call, In the device, receiving a first set of audio frames of an audio frame sequence from the first device, In response to determining that no audio frames of the sequence have been received by the device during a threshold duration following the last received audio frame of the sequence, the device initiates transmission of a frame loss indication to the first device. In response to the frame loss indication, the device receives from the first device a second set of audio frames of the sequence and a second playback speed indication, wherein the second set of audio frames follows the last received audio frame of the sequence. Playback of the second set of audio frames is initiated via a speaker based on the second playback speed, and the second playback speed is greater than the first playback speed of the first set of audio frames in the sequence. A method that includes [a certain feature].
6. The first device includes a server, the call is a call with a plurality of second devices, and the sequence of audio frames is based on a plurality of sequences of audio frames received by the first device from the plurality of second devices, or The aforementioned call is made between the device and a single additional device, the single additional device including the first device. The method according to claim 5.
7. The device according to claim 1, wherein initiating playback of the second set of audio frames at the second playback speed includes at least partially suppressing the silence represented by the second set of audio frames.
8. The one or more processors described above further include: The next audio frame of the sequence is received from the first device, and the next audio frame follows the second set of audio frames in the sequence. The playback of the next audio frame is transitioned from the second playback speed to the first playback speed, The device according to claim 1, configured to perform the following:
9. The one or more processors described above further include: The second set of audio frames is stored in a buffer, The next audio frame of the sequence is received from the first device, and the next audio frame is stored in the buffer, and the next audio frame follows the second set of audio frames in the sequence. In response to determining that fewer than a threshold number of audio frames of the sequence are stored in the buffer for playback, playback is shifted from the second playback speed to the first playback speed, The device according to claim 1, configured to perform the following:
10. The device according to claim 1, wherein the second playback speed is based on the number of the second set of audio frames.
11. The device according to claim 1, wherein the one or more processors are configured to start playback of the silent frames at a playback speed greater than the second playback speed in response to determining that the second set of audio frames includes silent frames.
12. The device according to claim 1, wherein the one or more processors are incorporated into at least one of the following: a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a headset, or a combination thereof.
Citation Information
Patent Citations
Digital transmission method for sound
JP2001230761A
Digital audio with parameters for real-time time warping
JP2005512134A
Increasing video bit rates while maintaining video quality
US20200099972A1
Data processing device, data processing method, and program
WO2016088582A1