Real-time and file-based audio data processing method, electronic device and readable medium
By dynamically switching between real-time and batch data processing modes in the second electronic device, the delay problem in the voice activation function of low-cost devices is solved, and the processing efficiency of voice input and user experience are improved.
Patent Information
- Application Number
- CN202180067092.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-02
- Filing Date
- 2021-10-01
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-10-01
AI Technical Summary
Low-cost electronic devices experience delays when voice activation is activated due to insufficient communication, cache, or processing power.
By implementing a combination of real-time data processing and file-based batch data processing in the second electronic device, the processing mode is dynamically switched to process the audio data samples in real time or in batches according to the device capabilities.
This simplifies data processing and communication, improves performance and audio quality associated with voice input, and enhances the user experience.
Smart Images

Figure CN116325696B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 086,953, filed October 2, 2020, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present application relates generally to audio data transmission and processing, including but not limited to methods and systems for providing real-time and file-based audio data processing to facilitate data transmission and speech recognition in electronic devices. Background Art
[0004] Electronic devices with microphones are widely used as auxiliary devices to collect voice input from users and to start different voice activation functions according to the voice input. For example, many remote control devices coupled to digital television devices are configured to integrate microphones. The voice input of these remote control devices is streamed to the digital television device and is at least partially processed by the digital television device. The digital television device can submit the voice input (pre-processed or not) to a remote server system for additional audio processing. Due to the audio processing at the television device and / or remote server system, the user request is extracted from the voice input to start the voice activation function. Any defect in the communication, high-speed buffer memory and processing power of the television device may cause a delay in the startup of the voice activation function. This situation usually occurs on low-cost television devices with limited capabilities. It would be beneficial to have a more efficient data processing and transmission mechanism than current practice to compensate for the defects in the communication, high-speed buffer memory or processing power of these devices. Summary of the Invention
[0005] The present application relates to processing and transmitting audio data received from an electronic device having a microphone (e.g., a remote control device, an auxiliary device). The electronic device is coupled to another electronic device having audio processing capabilities (e.g., a television device) or to a server having audio processing capabilities. The two electronic devices are coupled via a communication channel. The audio data is transmitted in real time via the communication channel and processed by the receiving electronic device in real time or in batches, depending on whether the communication, computing, and storage of the receiving electronic device can support real-time processing of the audio data samples. Real-time audio data processing is thus supplemented by batch audio data processing, particularly in some electronic devices that do not always have sufficient resources to communicate, cache, or process audio data in real time.
[0006] In particular, in one aspect, a method is implemented to process audio data, e.g., to switch from a real-time data processing mode to a batch data processing mode. The method includes receiving, by a second electronic device (e.g., a television device), a first sequence of audio data samples and a second sequence of audio data samples from a first electronic device (e.g., a remote control device). The second sequence of audio data samples is after the first sequence of audio data samples in an audio signal captured by a microphone of the first electronic device. The method further includes processing, by the second electronic device, the first sequence of audio data samples according to the real-time data processing mode, and determining that the second electronic device is unable to support processing of the audio data samples in the real-time data processing mode. The method further includes, in accordance with the determination that the second electronic device is unable to support processing of the audio data samples in the real-time data processing mode, caching the second sequence of audio data samples in a buffer of the second electronic device, and generating a data file including the second sequence of audio data samples in a batch data processing mode.
[0007] Alternatively, in another aspect, a method is implemented to process audio data, e.g., to switch from a batch data processing mode to a real-time data processing mode. The method includes receiving, by a second electronic device (e.g., a television device), a first sequence of audio data samples and a second sequence of audio data samples from a first electronic device (e.g., a remote control device). The second sequence of audio data samples is after the first sequence of audio data samples in an audio signal captured by a microphone of the first electronic device. The method further includes processing, by the second electronic device, the first sequence of audio data samples according to the batch data processing mode, the method further including caching the first sequence of audio data samples in a buffer of the second electronic device and generating a data file including the first sequence of audio data samples. The method further includes determining that the second electronic device is able to support processing of the audio data samples in the real-time data processing mode. The method further includes, in accordance with the determination that the second electronic device is able to support processing of the audio data samples in the real-time data processing mode, processing, by the second electronic device, the second sequence of audio data samples according to the real-time data processing mode.
[0008] A non-transitory computer-readable medium has stored thereon instructions that, when executed by one or more processors, cause the processors to perform any of the methods described above. An electronic device includes one or more processors and memory having stored thereon instructions that, when executed by the one or more processors, cause the processors to perform any of the methods described above. BRIEF DESCRIPTION OF DRAWINGS
[0009] For a better understanding of the various implementations described herein, reference should be made to the following descriptions taken in connection with the accompanying drawings, in which like references refer to corresponding, but not necessarily identical, parts throughout the differential drawings.
[0010] Figure 1 is an example media environment according to some embodiments, in which a network-connected TV device, a remote control device, and a server system interact with each other via one or more communication networks.
[0011] Figure 2 is an example audio data transmission path between a first electronic device and a second electronic device according to some embodiments.
[0012] Figure 3 is a schematic diagram illustrating an example audio data processing process of switching from a real-time data processing mode to a batch data processing mode according to some embodiments.
[0013] Figure 4 is a schematic diagram illustrating an example audio data processing process of switching from a batch data processing mode to a real-time data processing mode according to some embodiments.
[0014] Figure 5 is a schematic diagram illustrating an example voice assistant process initiated by user action or voice input according to some embodiments.
[0015] Figure 6 Illustrated is an example remote control device configured to transmit audio data to a television device according to some embodiments.
[0016] Figure 7 is a flowchart of a method for dynamically processing audio data in two audio data processing modes according to some embodiments.
[0017] Figure 8 is a flowchart of another method for dynamically processing audio data in two audio data processing modes according to some embodiments.
[0018] Like reference numerals refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION
[0019] Many electronic devices (e.g., remote control devices, voice-activated displays, or speaker devices) include a microphone for collecting voice input from the environment in which the electronic device is placed. Such electronic devices can be configured to automatically collect voice input after detecting a hot word in an audio signal or in response to a user pressing a dedicated auxiliary button of the electronic device. After receiving the voice input, the electronic device transmits the voice input to a remote server system (e.g., an assistant server) via one or more communication networks, and the remote server system recognizes the user request in the voice input and responds to the user request. In an example, the electronic device includes a remote control device, and voice input is initiated to control a network-connected television (TV) device coupled to the remote control device. The remote control device sends the voice input to the remote server system via the TV device, and processes the voice input at the TV device before the TV device sends the voice input to the remote server system. During this process of processing the voice input, the TV device uses its communication, computing, and storage capabilities to bridge the remote control device and the remote server system.
[0020] The audio data delivered to the audio manager of TV equipment can be different from the audio data collected by the microphone of remote control device.This occurs due to various factors, such as loss and delay of data packets via the communication channel of coupling remote control and TV equipment, processor load of TV equipment.In various embodiments of the present application, the combination of real-time data processing and file-based batch data processing is realized at the second electronic device (e.g., TV equipment) to process the audio data collected by the microphone of the first electronic device (e.g., remote control device). In some embodiments, real-time data processing has the priority over file-based batch data processing. When determining that at least one of the communication, computing and storage capabilities of the second electronic device cannot support real-time processing audio data samples, subsequent data samples are cached and organized into data files at the second electronic device (e.g., processed by an audio data processing module different from the audio manager). Alternatively, when determining that the communication, computing and storage capabilities of the second electronic device can support real-time processing audio data samples, subsequent data samples are processed into data packets in real time by the second electronic device (e.g., processed by an audio manager, which is a part of the operating system of the second electronic device). This controlled audio data transfer process simplifies data processing and communication at the second electronic device and improves the performance, audio quality, and user experience associated with voice input that initiates user interaction with the electronic device.
[0021] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be understood by those skilled in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks are not described in detail to avoid unnecessarily obscuring aspects of the embodiments.
[0022] Figure 1 1 is an example media environment 100 according to some embodiments, wherein a network-connected TV device 102, a remote control device 104, and a server system 106 interact with each other via one or more communication networks 180. The media environment 100 corresponds to a virtual user domain created and hosted by the server system 106, and the virtual user domain includes multiple user accounts. For each user account, the server system 106 is coupled to a content source 110 and one or more media devices 102 and 116-126, and is configured to stream media content provided by the content source 110 for review by the user via the corresponding user account. Optionally, the content source 110 includes one or more of: an advertisement source, an electronic program guide (EPG) source, and a media content source.
[0023] Specifically, one or more media devices associated with a user and a user account are placed in the media environment 100 to provide the user with media content stored at and streamed from a content source 110. The content source 110 is optionally a third-party media content source or an internal media source hosted by a server system 106. In some embodiments, one or more media devices include a network-connected TV device 102 that directly streams media content from a remote content source or integrates an embedded delivery unit that is configured to stream media content for display to its audience. The network-connected TV device 102 is communicatively coupled to a dedicated remote control device 104 and / or an electronic device with a remote control application (e.g., a mobile phone 122, a tablet computer 124, a laptop computer 126, an auxiliary device 138). The dedicated remote control device 104 can be placed near the TV device 102 and is configured to communicate with the TV device 102 using digitally encoded pulses of infrared signals. Alternatively, in some cases, a dedicated remote control device 104 or an electronic device with a remote control application is configured to communicate with TV device 102 via communication network 180 (i.e., via a short-range communication link, a local area network, and / or a wide area network) and does not have to be physically close to TV device 102.
[0024] The network-connected TV device 102 includes one or more processors and a memory that stores instructions for execution by the one or more processors. The instructions stored on the network-connected TV device 102 include one or more of the following: a unified TV application, a local content delivery application, a remote control application, an assistant application, and one or more media playback applications associated with the content source 110. These applications are user applications that are distinct from the operating system of the TV device 102 and are optionally linked to a user account in the virtual user domain of the media environment 100. In addition, the network-connected TV device 102 includes an audio manager (e.g., Figure 2 234), the audio manager is integrated into its operating system to process audio data at the packet level.
[0025] Alternatively, in some embodiments, the media devices disposed in the media environment 100 include: a display device 116 that outputs media content directly to an audience; and a delivery device 118 that is coupled to the display device 116 and configured to stream the media content to the display device 116. Examples of the display device 116 include, but are not limited to, a television (TV) display device and a music player. Examples of the delivery device 118 include, but are not limited to, a set-top box (STB), a DVD player, and a TV box. Figure 1 In this example shown in , display device 116 comprises a TV display that is hardwired to a DVD player or set-top box 118. In contrast, in some embodiments, the media devices housed in media environment 100 include a computer screen 120A that outputs media content to an audience and a desktop computer 120B that streams the media content to computer screen 120A. In some embodiments, the media devices housed in media environment 100 include mobile devices, such as mobile phone 122, tablet computer 124, and laptop computer 126. Each of media devices 118 to 126 includes one or more media playback applications that are configured to receive and play media content items provided by content source 110 or an internal media source associated with server system 106.
[0026] The server system 106 includes a unified media platform (I-JWP) 128 that is configured to manage media content recommendations and streaming for one or more media devices in the media environment 100. The media content recommendations generated by the I-AAP 128 are presented on the network-connected TV device 102 via a server-side TV application 134, and the server-side TV application 134 is capable of displaying the media content on the unified TV application on the TV device 102 in response to a user selection from the media content recommendations. In addition, the UNIP 128 can also serve as a centralized media content management module that is configured to provide media content recommendations to other media devices 118 to 126 in addition to the TV device 102. In some embodiments, activity data associated with each user account is collected from the TV application 134 and the delivery service module 136, and the activity data is used to personalize the media content recommendations provided to the user of the user account.
[0027] In some embodiments, in addition to one or more of media devices 102, 104, and 116-126, a user account in the virtual user domain hosted by server system 106 is also associated with one or more other types of devices, such as network-connected auxiliary devices 138 installed in media environment 100. Examples of auxiliary devices 138 include speaker auxiliary device 142 and display auxiliary device 144. Speaker auxiliary device 142 can collect audio input, recognize user commands from the audio input, and implement operations in response to the user commands (e.g., play music, answer questions). Display auxiliary device 144 can collect audio and / or video input, recognize user commands from the audio and / or video input, and implement operations in response to the user commands (e.g., play music, present an image or video clip, answer questions). Each of auxiliary devices 138 is optionally managed by a dedicated device application or a general user application (e.g., a web browser) and is linked to a user account in the virtual domain in conjunction with the unified TV application of network-connected TV device 102.
[0028] In addition, in some embodiments, the server system 106 includes an auxiliary module 140, which is optionally powered by artificial intelligence. The auxiliary module 140 is configured to recognize user requests from voice input collected by the microphone and initiate operations such as searching the Internet, scheduling events and alarms, adjusting hardware settings, presenting public or private information, playing media content items, conducting a two-way conversation with the user, purchasing products, transferring money, etc. The microphone is integrated into any of the media devices 102, 104 and 116 to 126 and the auxiliary device 138 placed in the media environment 100. In some embodiments, the auxiliary module 140 is coupled to a speech recognition module 160, which is configured to process the voice input collected by the microphone and identify the user request from the voice input using, for example, a natural language processing (NLP) algorithm.
[0029] In some embodiments, the server system 106 includes a device and application registry 150 configured to store information for one or more user accounts managed by the server system 106 and information for user devices and applications associated with each of the one or more user accounts. For example, the device and application registry 150 stores information for the network-connected TV device 102, the remote control device 104, the media devices 116 to 126, the auxiliary device 138, and information for the corresponding unified TV application, remote control application, media playback application, and dedicated device application associated with the auxiliary device 138.
[0030] Optionally, these media devices and auxiliary devices associated with the same user account are distributed across different geographic regions. Optionally, these devices are located at the same physical location. Each media or auxiliary device communicates with another device or server system 106 using one or more communication networks 180. The communication network 180 used can be one or more networks with one or more types of topologies, including but not limited to the Internet, an intranet, a local area network (LAN), a cellular network, an Ethernet network, a storage area network (SAN), a telephone network, a Bluetooth personal area network (PAN), etc. In some embodiments, two or more devices in a subnet are coupled via a wired connection, while at least some devices in the same subnet are coupled via a local radio communication network (e.g., ZigBee, Z-Wave, Insteon, Bluetooth, Wi-Fi, and other radio communication networks).
[0031] In various embodiments, a first electronic device (e.g., remote control device 104, any of media devices 116 to 126, auxiliary device 138) having a microphone is coupled to a second electronic device (e.g., TV device 102 or any of media devices 116 to 126) via a communication channel, the second electronic device having one or more processors and a memory that stores instructions for execution by the one or more processors. The first electronic device captures an audio signal using its microphone. The audio signal is sampled to obtain a first sequence of audio data samples and a second sequence of audio data samples that follows the first sequence of audio data samples. Optionally, the first sequence and the second sequence of audio data samples are recorded in the same recording session or during two different recording sessions. Each recording session uses the first electronic device to be recorded by a corresponding user action (e.g., the user presses a button). Figure 6 The first electronic device transmits the first and second sequences of audio data samples to the second electronic device, which then processes each sequence of audio data samples via one of a real-time data processing mode and a batch data processing mode.
[0032] Simultaneously with or after transmitting the first sequence of audio data samples, the second electronic device determines whether the communication, computing, and storage capabilities of the second electronic device can support processing the audio data samples in the real-time data processing mode. If at least one of the communication, computing, and storage capabilities of the second electronic device cannot support processing the audio data samples in the real-time data processing mode, the second electronic device processes the second sequence of audio data samples in the batch data processing mode (e.g., by Figure 2 In contrast, if the second electronic device is capable of supporting the processing of audio data samples in a real-time data processing mode, the second electronic device processes the second sequence of audio data samples in the real-time data processing mode (e.g., by Figure 2 Therefore, based on determining that the second electronic device cannot support processing the audio data samples in the real-time data processing mode, for example, when the error rate or delay of the audio data exceeds the corresponding error or delay tolerance, the batch data processing mode is activated.
[0033] Figure 2is an example audio data transfer path 200 between a first electronic device 202 and a second electronic device 204 according to some embodiments. The first electronic device 202 has or is coupled to a microphone 206 configured to capture an audio signal 220, and the second electronic device 204 is coupled to the first electronic device 202 via a communication channel 208. The audio signal 220 is sampled at the first electronic device 202, and audio data samples are transferred from the first electronic device 202 to the second electronic device 204 via the communication channel 208. The communication channel 208 is enabled by one or more communication networks 180 including, but not limited to, local radio communication networks (e.g., ZigBee, Z-Wave, Insteon, Bluetooth, Wi-Fi, and other radio communication networks). In this example audio data transfer path 200, the communication channel 208 is formed via a Bluetooth communication link that is collectively enabled by a first Bluetooth stack 208A of the first electronic device 202 and a second Bluetooth stack 208B of the second electronic device 204.
[0034] The first electronic device 202 includes an audio streaming module 210 configured to obtain audio data samples of the audio signal 220 captured by the microphone 206 and to organize the audio data samples for transfer to the second electronic device over the communication channel 208. In some embodiments, the audio streaming module 210 groups a subset of the audio data samples into an ordered sequence of audio data packets. Each data packet includes one or more consecutive audio data samples, and optionally has a preamble, a message header, encoded packet data, dummy fields, and an integrity check field that conform to a predefined data format (e.g., MPEG-4 HE-ACC codec format, EVRC speech codec format). At the input of the first Bluetooth stack 208A, multiple ordered sequences of audio data packets are sequentially arranged into a stream of audio data 230 for transmission over the communication channel 208. In some embodiments, each data packet in the same data packet sequence corresponds to a predefined data format, while two different data packet sequences optionally correspond to the same data format or different data formats.
[0035] The second electronic device 204 includes three levels of programs, namely the kernel on the hardware abstraction layer (HAL) 216, device firmware 218, and applications and service programs 222. The kernel and device hardware 218 on the HAL 216 are part of the operating system of the second electronic device 204, while the applications and service programs 222 are outside the operating system and are installed by the manufacturer or user to implement specialized computer operations (e.g., games, web browsing, document editing, media playback). In some embodiments, after receiving operation data, optionally originating from the first electronic device 202 or any other electronic device, via the second Bluetooth stack 208B, the second electronic device 204 passes the received operation data to the input event identifier 226 of the kernel on the HAL 216. The input event identifier 226 identifies the operation data from the data received by the second Bluetooth stack 208B and provides the identified operation data to the input dispatcher 228 of the device firmware 218. The input dispatcher 228 assigns the operation data to the assistant application 232 installed on the second electronic device 204 to identify the user request in the operation data and initiate an operation in response to the user request.
[0036] In some embodiments, a stream of audio data 232 is collected by the microphone 206 of the first electronic device 202 and dispatched directly by the Kernal / HAL 216 to a remote control application 236 associated with the first electronic device 202. The remote control application 236 cooperates (260) with the audio manager 234 and assistant application 232 of the device firmware 218 installed on the second electronic device 204 to identify and respond to user requests in the stream of audio data 230. The audio manager 234 implements a real-time data processing mode and processes the audio data 230 at the packet level. That is, the audio manager 234 is configured to identify audio data samples from a plurality of ordered sequences of audio data packets received from the first electronic device 202; compensate for erroneous, lost, and out-of-order packets; organize the received audio data samples into another sequence of audio data packets; and pass (262) the sequence of audio data packets to the assistant application 232 for subsequent audio processing or transmission.
[0037] Alternatively, in some embodiments, the second electronic device 204 includes a file-based audio data processing module 238 that implements a batch data processing mode and processes the audio data 230 received from the first electronic device 202 into a data file. The file-based audio data processing module 238 is integrated (238A) into the device firmware 218 (i.e., the operating system) or installed (238B) as a user application among the applications and service programs 222. The remote control application 236 provides (270 or 280) the audio data 230 to the processing module 238. The processing module 238 is configured to identify audio data samples from a plurality of ordered sequences of audio data packets received from the first electronic device 202; cache the data packets in a data file; and provide (272 or 282) the data file to the assistant application 232 for subsequent audio processing or transmission. The second electronic device 204 further includes an audio data buffer 214 for storing data files.
[0038] The audio manager 234 and the processing module 238 process the audio data samples alternately in real-time data processing mode and in batch data processing mode, respectively. Depending on whether the communication, computing and storage capabilities of the second electronic device 204 can support the real-time transmission of the audio data samples, the second electronic device 204 switches between the two audio data processing modes. In some embodiments, when the audio manager 234 or the processing module 238 is processing an ordered sequence of audio data packets, this determination and mode switching are dynamically implemented. Alternatively, this determination and mode switching are implemented between two different recording sessions, i.e., between two different sequences of audio data packets, each of the two different sequences of audio data packets being independently processed by one of the real-time and batch data processing modes. Each recording session is optionally activated by a user action on the first electronic device 202 or by voice activation detected from the audio data 230 collected by the electronic device 202.
[0039] More specifically, after receiving the stream of audio data 230, the second electronic device 204 monitors the audio data 230 to determine whether the communication, computing, and storage capabilities of the second electronic device 204 can support processing the audio data samples in the real-time data processing mode. For example, the second electronic device 204 determines in real time whether a data sample latency of a data sample in the audio data 230 exceeds a latency tolerance, e.g., before or after the audio manager 234 processes the data sample in the audio data 230. In another example, the second electronic device 204 determines in real time whether an audio data sample loss rate of the data samples in the audio data 230 exceeds a loss rate tolerance, or whether an audio data sample out-of-order rate of the data samples in the audio data 230 exceeds an out-of-order rate tolerance, e.g., before or after the audio manager 234 processes the data sample in the audio data 230. In some embodiments, the second electronic device 204 monitors its central processing unit (CPU) utilization and determines that if the CPU utilization exceeds a predetermined utilization percentage (e.g., 85%), the second electronic device cannot support real-time processing of the audio data samples. According to the determination result, the second electronic device 204 selects the audio manager 234 or the file-based audio data processing module 238 to process the stream of the audio data 230 in real time or in batches, respectively.
[0040] In some embodiments, the second electronic device 204 is configured to locally process the output of the audio manager 234 or the processing module 238 to identify a user request from the output for the purpose of protecting user privacy, and optionally provide the processed output to the remote control application 236 for use in controlling the first electronic device 202. Alternatively, in some embodiments, the second electronic device 204 is configured to pre-process the output of the audio manager 234 or the processing module 238 before sending it to the remote server system 106 to identify a user request from the output. Alternatively, in some embodiments, the second electronic device 204 has limited speech recognition capabilities, such as when the second electronic device 204 is intended to be a low-cost device. The second electronic device 204 is configured to send the entire output of the audio manager 234 or the processing module 238 to the remote server system 106 and rely on the server system 106 to identify the user request from the output.
[0041] Figure 3FIG2 is a schematic diagram illustrating an example audio data processing process 300 for switching from a real-time data processing mode to a batch data processing mode according to some embodiments. An audio signal 220 is captured using a microphone 206 of a first electronic device 202 and sampled at an audio sampling rate to obtain a first sequence 302 of audio data samples and a second sequence 304 of audio data samples that follows the first sequence 302 of audio data samples. The first sequence 302 of audio data samples is transmitted to a second electronic device 204 via a communication channel 208. In some embodiments, the audio data samples in the first sequence 302 and the second sequence 304 are grouped into a plurality of audio data packets, each comprising one or more consecutive audio data samples. Optionally, the second sequence 304 of audio data samples immediately follows the first sequence 302. Optionally, the second sequence 304 of audio data samples is separated from the first sequence 302 by an interrupt. The plurality of audio data packets are streamed to the second electronic device 204. In some cases, one or more packets from the first sequence 302 and the second sequence 304 are reordered or dropped during transmission via the communication channel 208.
[0042] The audio data packets 308 are grouped from the first sequence 302 of audio data samples and processed by the second electronic device 204 in a real-time processing mode, wherein erroneous, lost, and out-of-order data packets in the audio data packets 308 may occur and be corrected. That is, one or more data packets 308 are discarded or reordered by the second electronic device 204 (specifically, by the audio manager 234). While or after processing the first sequence 302 of audio data samples, the second electronic device 204 determines that at least one of the communication, computing, and storage capabilities of the second electronic device 204 cannot support processing the audio data samples in the real-time data processing mode. In response to this determination (at time t F ), the second electronic device 204 suspends the real-time data processing mode and initiates the batch data processing mode to process the second sequence 304 of audio data samples. Specifically, the second electronic device 204 caches the second sequence 304 of audio data samples in the audio data buffer 214 and generates a data file 310 including the second sequence 304 of audio data samples in the batch data processing mode. In some implementations, the second electronic device 204 determines the corresponding capability of supporting the processing of audio data samples in the real-time data processing mode based on at least one of the following: data sample latency, audio data sample loss rate, audio data sample out-of-order rate, and CPU utilization associated with the second electronic device 204.
[0043] In some embodiments, upon determining that the communication, computing, and storage capabilities of the second electronic device 204 are capable of supporting the processing of audio data samples in real-time data processing mode, the second electronic device 204 stops processing the second sequence 304 of audio data samples. Alternatively, in some embodiments, the second electronic device 204 is configured to limit the second sequence 304 to include a predefined number of audio data samples. When the predefined number is reached, the second electronic device 204 processes the second sequence 304 of audio data samples to be cached in the first data file 310. The second electronic device 204 then organizes a third sequence 306 of audio data samples immediately following the second sequence 304 of audio data samples into a second data file 312. When the predefined number of audio data samples is included in the third sequence 306 of audio data samples, or when it is determined that the second electronic device 204 is capable of supporting the processing of audio data samples in real-time data processing mode, the second data file 312 is transmitted.
[0044] In some cases, when it is determined that the second electronic device 204 is capable of processing audio data samples in real-time data processing mode, the current number of audio data samples included in the data file 310 or 312 has not yet reached a predefined number. Based on this determination, the data file 310 or 312 can be immediately transmitted along with the current number of audio data samples. Optionally, the transmission of the data file 310 or 312 is aborted, and the current number of audio data samples are reorganized into data packets for real-time audio data processing by the audio manager 234 of the second electronic device 204.
[0045] In some embodiments, the first sequence 308 of processed audio data samples and the data file 310 including the second sequence of audio data samples are transmitted 320 to the server system 106. The first sequence 302 of processed audio data samples has a first data transmission rate corresponding to a real-time data processing mode, and the second sequence 304 of audio data samples has a second data transmission rate corresponding to a batch data processing mode. The second data transmission rate is greater than the first data transmission rate. In some embodiments, the first data transmission rate is slower than the audio sampling rate of the audio signal 220, and the second data transmission rate is greater than the audio sampling rate.
[0046] Figure 4is a schematic diagram illustrating an example audio data processing process 400 for switching from a batch data processing mode to a real-time data processing mode according to some embodiments. An audio signal 220 is captured using a microphone of a first electronic device 202 and sampled at an audio sampling rate to obtain a first sequence 402 of audio data samples and a second sequence 404 of audio data samples subsequent to the first sequence 402 of audio data samples. Both the first sequence 402 of audio data samples and the second sequence 404 of audio data samples are transmitted to a second electronic device 204 via a communication channel 208. The second electronic device 204 processes the first sequence 302 of audio data samples according to the batch data processing mode. The audio data samples in the first sequence 402 are cached in an audio data buffer 214 of the first electronic device 202 and organized into a data file 406.
[0047] While or after processing the first sequence 402 of audio data samples, the second electronic device 204 detects or determines that the second electronic device 204 is capable of supporting processing the audio data samples in a real-time data processing mode. Optionally, based on this determination (at the first time t A ), the second electronic device 204 continues to add more audio data samples to the first sequence 402 until the number of audio data samples in the first sequence 402 is greater than the number at the second time t B The second electronic device 204 completes caching the first sequence 402 of audio data samples in the data file 406 before it begins processing the second sequence 404 of audio data samples collected after the first sequence 402 of audio data samples in the real-time data processing mode. Alternatively, in some embodiments, based on the determination (at time t C ), the second electronic device 204 stops adding audio data samples to the first sequence 402, regardless of whether the number of audio data samples of the first sequence 402 has reached the predefined number. C The data file 406 is prepared, thereby terminating the batch data processing mode. In the real-time data processing mode, the second electronic device 204 immediately begins transmitting the second sequence 404 of audio data samples collected after the first sequence 402 of audio data samples. In addition, in some embodiments (not shown), in the determination (e.g., at time t C ), the second electronic device 204 ceases processing the first sequence 402 of audio data samples in the batch data processing mode and begins processing the first sequence 402 of audio data samples into data packets 408 in the real-time data processing mode immediately. After transmitting the first sequence 402 of audio data samples, the second electronic device 204 continues processing the second sequence 404 of audio data samples in the real-time data processing mode.
[0048] Each of the first sequences 302 and 402 of audio data samples optionally begins the stream of audio data 230 sent to the second electronic device 204, or is in the middle of the stream of audio data 230. Similarly, each of the second sequences 304 and 404 of audio data samples and the third sequence 306 of audio data samples optionally is the last sequence in the stream of audio data 230 sent to the second electronic device 204, or is in the middle of the stream of audio data 230. It should be noted that in some embodiments, the first sequence 302 or 402 of audio data samples does not immediately precede the second sequence of audio data samples 304 or 404. The first sequence and the second sequence are captured during two different recording sessions separated by an interruption. The second electronic device 204 determines whether it can support processing the audio data samples in real-time data processing mode during the interruption that separates the two recording sessions.
[0049] Figure 5 is a diagram illustrating an example voice assistant process 500 initiated by a user action or voice input according to some embodiments. The voice assistant process 500 is implemented by the first electronic device 202 and the second electronic device 204 in collaboration. In some embodiments, the first electronic device 202 includes a physical auxiliary button (e.g., Figure 6 6 in the figure), and allows the user to request to start or terminate the voice assistant function through a user action on the auxiliary button. For example, the user presses the physical auxiliary button to start the voice assistant process 500 (also known as a recording session). While the user keeps pressing the auxiliary button, the microphone 206 of the first electronic device 202 continuously collects audio signals for audio data sampling, transmission, and recognition via the voice assistant process 500. The microphone 206 does not stop collecting audio signals until the user releases the press of the auxiliary button to complete the corresponding recording session.
[0050] In some embodiments, in response to detecting the first user action, the first electronic device 202 sends an assistant invocation request 502 to the remote control application 236 and the assistant application 232 of the second electronic device 204. In response to the assistant invocation request 502, the assistant application 232 verifies that the first electronic device 202 is permitted to implement the voice assistant process 500 with the second electronic device 204, and sends an instruction to launch the assistant 504 to the remote control application 236. In response to the instruction to launch the assistant 504, the remote control application 236 sends an open microphone instruction 506 to the first electronic device 202. After the microphone 206 of the first electronic device 202 is turned on, audio data samples 230 are collected and transmitted to the second electronic device 204. After being transmitted to the second electronic device 204, the audio data samples 230 are processed (520 and 530) by the audio manager 234 associated with the real-time data processing mode or by the file-based audio data processing module 238 associated with the batch data processing mode. In some embodiments, the real-time and batch data processing modes are dynamically alternated at the second electronic device 204 based on, for example, a data sample latency, an audio data sample loss rate, an audio data sample out-of-order rate, or a CPU utilization rate associated with the second electronic device 204. In some embodiments, during each recording segment activated in response to detecting a first user action, only one of the real-time and batch data processing modes is activated.
[0051] In some embodiments, the assistant application 232 sends an instruction to start recording 508 to the remote control application 236 so that the remote control application 236 can control the first electronic device 202 to capture audio data collected by the first electronic device 202. The instruction to start recording 508 is optionally issued together with an instruction to launch the assistant 504 and is configured to trigger the open microphone instruction 506. In response to the instruction to start recording 508, the audio data samples 230 are recorded by the second electronic device 204 after they are transmitted from the first electronic device 202. Alternatively, in some embodiments, the assistant application issues an instruction to start recording 508' after the second electronic device 204 has received the subset of audio data samples 230. The instruction to start recording 508' can be issued based on the content of the subset of audio data samples (e.g., a user request in the content), and the second electronic device 204 does not record the subset of audio data samples 230.
[0052] In some embodiments, the assistant application 232 sends an instruction to stop recording 510 to both the audio manager 234 and the remote control application 236, so that the remote control application 236 can issue a close microphone instruction 512 to control the first electronic device 202 to close its microphone 206. Alternatively, in some embodiments, the user of the first electronic device 202 terminates the first user action that initiated the voice assistant process 500, or applies a second user action (e.g., releases the assist button) to terminate the voice assistant process 500. In response to the second user action, the microphone 206 of the first electronic device 202 is closed, and a request to end the assistant 514 is sent to the remote control application 236 and the assistant application 232 of the second electronic device 204.
[0053] With reference to Figure 3 and Figure 4 The microphone 206 of the first electronic device 202 is set to open and start capturing audio signals at different times relative to the first user pressing. In some embodiments, the first electronic device 202 receives the first user action requesting to record audio signals at a first time tl, and its microphone 206 opens at a second time t2 to capture audio signals immediately in response to the first user action, regardless of whether the open microphone instruction 506 is issued from the second electronic device 204. The second time t2 substantially coincides with the first time tl. Alternatively, in some embodiments, the first electronic device 202 receives the first user action requesting to record audio signals at a first time tl’, and the microphone captures audio signals at a second time t2, which is prior to the first time tl’. The audio signals corresponding to the duration between times t2 and tl’ are cached, but the audio data is still transmitted to the second electronic device 204 in response to the first user action. Additionally, in some embodiments, the first electronic device 202 receives the first user action requesting to record audio signals at a first time tl”, and starts capturing audio signals at a second time t2, which is after the first time tl”, e.g., delayed from the first time tl” by a predefined buffer time (such as 5 seconds).
[0054] It is noted that in some embodiments, the first electronic device 202 waits for receiving an audio data request including the open microphone instruction 506 from the second electronic device 204 in the duration between the time of receiving the first user action (tl, tl’, or tl”) and transmitting the captured audio data to the second electronic device 204 at t2. The second electronic device 204 obtains an approval to send the audio data request in response to the first user action, and this approval is granted by the assistant application 232 of the second electronic device 204 or by the remote server system 106.
[0055] In some cases, voice input initiates the voice assistant process 500. The microphone 206 of the first electronic device 202 is configured to continuously collect audio signals and provide corresponding audio data to the second electronic device 204, regardless of whether the first electronic device 202 is in sleep mode and active mode. In sleep mode, the audio data is not processed to identify any user request for the voice assistant function for controlling the media device or user application until one or more predefined hot words (e.g., "Hey, Google") are detected to enable active mode. The second electronic device 204 is configured to locally detect one or more predefined hot words in the audio data and initiate the voice assistant process 500 in response to detecting the one or more predefined hot words.
[0056] refer to Figure 5 In some embodiments, the assistant invocation request 502 includes one or more predefined hot words. After detecting the hot words in the audio data received from the first electronic device 202, the second electronic device 204 confirms the assistant invocation request 502. In response to the assistant invocation request 502, the assistant application 232 verifies that the first electronic device 202 is allowed to implement the voice assistant process 500 with the second electronic device 204, and sends an instruction to start the assistant 504 to the remote control application 236. The second electronic device 204 is controlled to operate in an activation mode to locally process the audio data samples 230 in a real-time data processing mode or a batch data processing mode to identify additional user requests. In addition, in some embodiments, the assistant application 232 identifies a stop recording (e.g., Figure 3 and Figure 4 The assistant application 232 generates an instruction to stop recording 510 in the audio data sample 230 and sends the instruction 510 to the audio manager 234 and the remote control application 236, so that the remote control application 236 can issue a mute microphone instruction 512 to control the first electronic device 202 to mute its microphone. By these means, the voice assistant process 500 is started and terminated based on the audio data captured by the first electronic device 202 without any physical user action.
[0057] refer to Figure 3 In the example, a request 314 to stop capturing the audio signal is identified in the audio signal 220 and is sent from the second electronic device 204 to the first electronic device 202. In response to the request 314, the second electronic device 204 completes processing the first data file 310 including the second sequence 304 of data samples, but aborts transmitting the second data file 312 including the third sequence 306 of data samples that immediately follows the second sequence 304. In contrast, referring to Figure 4In another example, after receiving the request 314 , the second electronic device 204 suspends processing subsequent data packets after the request 314 in the second sequence 404 of audio data samples.
[0058] Figure 6 An example remote control device 104 configured to transmit audio data to a television device according to some embodiments is illustrated. A plurality of user buttons of the remote control device 104 include one or more of a power button 602, a home button 604, an auxiliary button 606, a loop button 608 (also known as a play / loop button), a back button 610, a forward button 612, a preview / background button 614, and a volume control button 616. When a media device coupled to the remote control device 104 was off prior to user actuation, user actuation of the power button 602 powers on the media device, and when the media device was on prior to user actuation, user actuation of the power button 602 powers off the media device. User actuation of the home button 604 controls the media device coupled to the remote control device 104 to display a home screen. For example, the home screen displays a specific advertisement clip or a randomly selected media program provided by a predetermined internet content channel. In some embodiments, the power button 602 or the home button 604 functions as a quick play button configured to enable immediate playback of media content provided by a particular Internet content channel.
[0059] The user action on the auxiliary button 606 controls the microphone 206 integrated in the remote control device 104 to collect audio signals in the media environment 100 and extracts the user request from the audio signals to control one or more media playback devices (e.g., TV device 102) located in the media environment 100. In some embodiments, when a first short press is applied on the auxiliary button 606, the microphone 206 of the remote control device 104 starts collecting audio signals from the environment 100, and a second short press or user request is applied to stop collecting audio signals. Alternatively, in some embodiments, the microphone 206 of the remote control device 104 collects audio signals from the environment 100 only when the auxiliary button 606 is pressed, and stops collecting audio signals when the auxiliary button 606 is released. In addition, in some embodiments, the microphone 206 of the remote control device 104 continuously captures audio signals from the environment 100, and the audio signals include one or more predefined hot words and / or user requests. The user request can be used to control the remote control device 104 or one or more media devices or applications coupled to the remote control device 104.
[0060] In this application, audio data is processed at a second electronic device 204 (e.g., a network-connected television device 102) having two audio data processing modes, including a real-time data processing mode and a batch data processing mode. The second electronic device 204 determines whether its communication, caching, and processing capabilities can support processing audio data samples in the real-time data processing mode. Based on determining that the second electronic device 204 can support processing audio data samples in the real-time data processing mode, the second electronic device 204 processes subsequent audio data samples in the real-time data processing mode at the packet level. Based on determining that the second electronic device 204 cannot support processing audio data samples in the real-time data processing mode, the second electronic device 204 caches subsequent audio data samples in a buffer of the second electronic device 204 and generates a data file including these audio data samples in the batch data processing mode. In some embodiments, determination and mode activation are implemented during an interruption between two recording sessions. Alternatively, in some embodiments, determination and mode switching are dynamically implemented between the same recording sessions.
[0061] Figure 7 7 is a flow chart of a method 700 for dynamically processing audio data in two audio data processing modes (from real-time data processing mode to batch data processing mode) according to some embodiments. The method 700 is performed by the first electronic device 202 and the second electronic device 204 and is optionally governed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the respective electronic devices. Figure 7 Each operation shown in method 700 may correspond to an instruction stored in a computer memory or a non-transitory computer-readable storage medium. The computer-readable storage medium may include a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, or one or more other non-volatile memory devices. The instructions stored on the computer-readable storage medium may include one or more of the following: source code, assembly language code, object code, or other instruction formats interpreted by one or more processors. Some operations in method 700 may be combined and / or the order of some operations may be changed.
[0062] An audio signal is captured (702) using microphone 206 of first electronic device 202. First electronic device 202 obtains (704) a first sequence 302 of audio data samples and a second sequence 304 of audio data samples subsequent to first sequence 302 of audio data samples from the audio signal, and transmits the first sequence 302 of audio data samples to second electronic device 204 via communication channel 208 according to a real-time data processing mode. Second electronic device 204 receives (706) the first sequence 302 of audio data samples and the second sequence 304 of audio data samples from first electronic device 202 via communication channel 208. Simultaneously with or subsequent to the first sequence 302 of audio data samples, second electronic device 204 determines (708) that second electronic device 204 cannot support processing of the audio data samples in the real-time data processing mode. Based on determining that the second electronic device 204 cannot support processing the audio data samples in the real-time data processing mode, the second electronic device 204 caches (712) the second sequence 304 of audio data samples in a buffer of the second electronic device and generates (714) a data file 310 including the second sequence 304 of audio data samples in the batch data processing mode.
[0063] In some embodiments, the second electronic device 204 transmits the processed first sequence of audio data samples and a data file including the second sequence of audio data samples (eg, Figure 5 540) to the server system 106. The first sequence 302 of audio data samples has a first data transfer rate corresponding to a real-time data processing mode, and the second sequence 304 of audio data samples has a second data transfer rate corresponding to a batch data processing mode. The second data transfer rate is greater than the first data transfer rate. In addition, in the example, the audio signal captured by the first electronic device 202 is sampled at an audio sampling rate to obtain the first sequence 302 and the second sequence 304 of audio data samples. The first data transfer rate is slower than the audio sampling rate, and the second data transfer rate is greater than the audio sampling rate.
[0064] In some embodiments, the audio data samples in the first sequence 302 are grouped into a plurality of audio data packets in real-time data processing mode. Each audio data packet includes one or more consecutive audio data samples, optionally organized according to a consistent data format. The plurality of audio data packets are streamed to the server system 106.
[0065] In some embodiments, the second electronic device 204 is determined to not support transmitting audio data samples in the real-time data processing mode based on at least one of the following: a data sample latency, an audio data sample loss rate, and an audio data sample out-of-sequence rate associated with the second electronic device 204. Specifically, in an example, the data sample latency of a subset of the first sequence of processed audio data samples exceeds a latency tolerance. In another example, the audio data sample loss rate of the first sequence 302 of processed audio data samples exceeds a loss rate tolerance. In yet another example, the audio data sample out-of-sequence rate of the first sequence 302 of processed audio data samples exceeds an out-of-sequence rate tolerance.
[0066] In some embodiments, a first user action requesting the recording of an audio signal is received at the first electronic device 202. The audio signal is captured in response to the first user action. An example is the first user action pressing an auxiliary button of the first electronic device 202. Pressing the button initiates a process to obtain approval from the assistant application 232 or the server system 106 of the second electronic device 204. After receiving approval, the audio signal is captured, processed, and recorded. Specifically, in the example, the first electronic device 202 receives the first user action requesting the recording of an audio signal at a first time t1', wherein the capture of the audio signal is initiated at a second time t2 after the first time t1'. The second time t2 is delayed by a predefined buffer time from the first time. Specifically, in some cases, in response to the first user action, the first electronic device 202 receives an audio data request from the second electronic device 204. The second electronic device 204 is configured to obtain approval for sending the audio data request in response to the first user action. In response to the audio data request, a first sequence of transmitting audio data samples is initiated.
[0067] In some embodiments, the data files 310 include a first data file 310. After generating the first data file 310, the second electronic device 204 continues to generate a second data file 312 in batch data processing mode that includes a third sequence 306 of audio data samples. The third sequence 306 of audio data samples immediately follows the second sequence 304 of audio data samples in the audio data, and each of the second sequence 304 and the third sequence 306 of audio data samples has a predefined number of data samples.
[0068] In some embodiments, the second electronic device 204 is configured to transmit the processed first sequence 302 and second sequence 304 of audio data samples to the server system 106 for audio processing (e.g., speech recognition). The server system 106 hosts a virtual user domain that includes a user account. The first electronic device 202 and the second electronic device 204 are linked to the user account.
[0069] Alternatively, in some embodiments, the audio signal includes one or more predefined hot words or user requests. The second electronic device 204 is configured to locally process the first sequence and the second sequence of audio data samples to identify the one or more predefined hot words or user requests in the audio signal. Furthermore, in some embodiments, the user request includes a request to stop capturing the audio signal. The request to stop capturing the audio signal is recognized by the second electronic device 204 and provided to the first electronic device 202. In response to the request, the first electronic device 202 stops transmitting the sequence of audio data samples following the second sequence 304 of audio data samples.
[0070] In some embodiments, while transmitting the second sequence 304 of data samples, the second electronic device 204 receives a second user action for stopping capturing the audio signal. In response to the second user action, the second electronic device 204 suspends receiving the sequence of audio data samples that immediately follows the second sequence 304 of audio data samples.
[0071] Figure 8 8 is a flow chart of a method 800 for dynamically processing audio data in two audio data processing modes (from batch data processing mode to real-time data processing mode) according to some embodiments. The method 800 is performed by the first electronic device 202 and the second electronic device 204 and is optionally governed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the respective electronic devices. Figure 8 Each operation shown in can correspond to an instruction stored in a computer memory or a non-transitory computer-readable storage medium. The computer-readable storage medium may include a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, or one or more other non-volatile memory devices. The instructions stored on the computer-readable storage medium may include one or more of the following: source code, assembly language code, object code, or other instruction formats interpreted by one or more processors. Some operations in method 800 can be combined and / or the order of some operations can be changed.
[0072] An audio signal is captured (802) using a microphone of a first electronic device 202. The first electronic device 202 obtains (804) a first sequence 402 of audio data samples and a second sequence 404 of audio data samples, the second sequence of audio data samples following the second sequence 402 of audio data samples in the audio signal. The second electronic device 204 receives (806) the first sequence 402 of audio data samples and the second sequence 404 of audio data samples from the first electronic device 202. The second electronic device 204 processes the first sequence 402 of audio data samples according to a batch data processing mode, including caching (810) the first sequence 402 of audio data samples in a buffer of the second electronic device and generating (812) a data file including the first sequence 402 of audio data samples. While or after processing the first sequence of audio data samples, the second electronic device 204 determines (814) that the second electronic device 204 is capable of supporting processing the audio data samples in a real-time data processing mode. Based on determining that the second electronic device 204 is capable of supporting processing the audio data samples in the real-time data processing mode, the second electronic device 204 processes (816) the second sequence of audio data samples according to the real-time data processing mode.
[0073] In some embodiments, audio data samples 402 and 404 are transmitted to server system 106. First sequence 402 of audio data samples has a first data transmission rate corresponding to a real-time data processing mode, and second sequence 404 of audio data samples has a second data transmission rate corresponding to a batch data processing mode. The first data transmission rate is greater than the second data transmission rate. Furthermore, in an example, the second data transmission rate is slower than the audio sampling rate, and the first data transmission rate is greater than the audio sampling rate. In some embodiments, audio data samples in second sequence 404 are grouped into a plurality of audio data packets in the real-time data processing mode, and the plurality of audio data packets are optionally streamed to server system 106 in real time.
[0074] refer to Figure 7 and Figure 8 In some embodiments, the first electronic device 202 includes a remote control device 104, and the second electronic device 204 includes a network-connected TV device 102 configured to be controlled by the remote control device 104. In some embodiments, the second electronic device 204 includes one or more processors and a memory storing one or more programs configured to implement an Android operating system and one or more user applications on the second electronic device 204. In some embodiments, the first electronic device 202 is powered by a battery.
[0075] It should be understood that description Figure 7 and Figure 8The specific order of operations in each of the methods 700, 750, 800, and 850 is merely exemplary and is not intended to indicate that the described order is the only order in which the operations may be performed. One of ordinary skill in the art will recognize various ways to display information items and focused content in a unified user interface as described herein. Additionally, it should be noted that details described with respect to one of the methods 700, 750, 800, and 850 may also be applied in a similar manner to any other of the methods 700, 750, 800, and 850. For the sake of brevity, similar details are not repeated.
[0076] It will also be understood that although in some cases, the terms first, second, etc. are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various described embodiments, a first electronic device may be referred to as a second electronic device, and similarly, a second electronic device may be referred to as a first electronic device. The first electronic device and the second electronic device are both electronic devices, but they are not the same electronic device.
[0077] The terms used in the description of the various described embodiments herein are only for the purpose of describing specific embodiments and are not intended to be restrictive. As used in the description of the various described embodiments and the appended claims, the singular forms "one", "an" and "said" are intended to also include plural forms, unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" as used herein refer to and encompass any and all possible combinations of one or more items in the associated listed items. It will be further understood that when used in this specification, the terms "includes", "including", "comprises" and / or "comprising" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0078] As used herein, the term "if" is optionally interpreted to mean "when" or "after" or "in response to determining" or "in response to detecting" or "according to determining," depending on the context. Similarly, the phrases "if it is determined" or "if [the condition or event] is detected" are optionally interpreted to mean "after determining" or "in response to determining" or "after detecting [the condition or event]" or "in response to detecting [the condition or event]" or "according to detecting [the condition or event]," depending on the context.
[0079] Although the various figures illustrate multiple logical stages in a particular order, stages that are not order-dependent may be reordered and other stages may be combined or decomposed. While some reordering or other groupings are specifically mentioned, other groupings will be apparent to one of ordinary skill in the art, and thus the ordering and groupings presented herein are not an exhaustive list of alternatives. Furthermore, it should be appreciated that the stages may be implemented in hardware, firmware, software, or any combination thereof.
[0080] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussions are not intended to be exhaustive or to limit the scope of the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen to best explain the principles behind the claims and their practical application, thereby enabling others skilled in the art to best utilize the embodiments with various modifications as appropriate for the particular use contemplated.
Claims
1. A method for processing audio data, comprising: receiving, by a second electronic device different from the first electronic device, a first sequence of audio data samples and a second sequence of audio data samples from the first electronic device, wherein the second sequence of audio data samples follows the first sequence of audio data samples in an audio signal captured by a microphone of the first electronic device; processing, by the second electronic device, the first sequence of audio data samples according to a real-time data processing mode; determining that at least one of communication, computing, and storage capabilities of the second electronic device cannot support processing the audio data samples in the real-time data processing mode; as well as Based on determining that at least one of the communication, computing and storage capabilities of the second electronic device cannot support processing of audio data samples in the real-time data processing mode, the second sequence of audio data samples is cached in a buffer of the second electronic device, and a data file including the second sequence of audio data samples is generated in the batch data processing mode.
2. The method according to claim 1, further comprising: transmitting the processed first sequence of audio data samples and said data file comprising said second sequence of audio data samples to a server system; wherein the first sequence of audio data samples has a first data transfer rate corresponding to the real-time data processing mode, and the second sequence of audio data samples has a second data transfer rate corresponding to the batch data processing mode; and The second data transmission rate is greater than the first data transmission rate.
3. The method according to claim 2, wherein: sampling the audio data at an audio sampling rate to obtain the first and second sequences of audio data samples; and The first data transfer rate is slower than the audio sampling rate, and the second data transfer rate is greater than the audio sampling rate.
4. The method according to claim 1, further comprising: grouping the audio data samples in the first sequence into a plurality of audio data packets, each audio data packet comprising one or more consecutive audio data samples; as well as The plurality of audio data packets are streamed to a server system.
5. The method of claim 1 , wherein determining that the second electronic device cannot support processing audio data samples in the real-time data processing mode further comprises at least one of: determining that a data sample delay of a subset of the first sequence of processed audio data samples exceeds a delay tolerance; determining that an audio data sample loss rate of a first sequence of processed audio data samples exceeds a loss rate tolerance; and It is determined that an audio data sample out-of-sequence rate of a first sequence of processed audio data samples exceeds an out-of-sequence rate tolerance.
6. The method of claim 1 , wherein the second electronic device is determined to not support transmission of audio data samples in the real-time data processing mode based on at least one of: a data sample latency associated with the second electronic device, an audio data sample loss rate, and an audio data sample out-of-order rate.
7. The method according to claim 1, further comprising: A first user action requesting recording of the audio signal is received by the first electronic device, wherein the audio signal is captured in response to the first user action.
8. The method according to claim 1, further comprising: A first user action for requesting recording of the audio signal is received by the first electronic device at a first time, wherein capturing of the audio signal is initiated at a second time after the first time.
9. The method according to claim 8, further comprising: In response to the first user action, obtaining, by the second electronic device, approval to send the audio data request and generating an audio data request; as well as Wherein the first sequence of audio data samples is received in response to the audio data request.
10. The method of claim 1, wherein the data file comprises a first data file, the method further comprising: After generating the first data file, the second electronic device continues to generate a second data file including a third sequence of audio data samples in the batch data processing mode, wherein the third sequence of audio data samples immediately follows the second sequence of audio data samples in the audio signal, and each of the second and third sequences of audio data samples has a predefined number of data samples.
11. The method according to claim 1 , wherein: The second electronic device is configured to transmit the first and second sequences of processed audio data samples to a server system for audio processing; The server system hosts a virtual user domain including user accounts; as well as The first electronic device and the second electronic device are linked to the user account.
12. The method of claim 1 , wherein the audio signal comprises one or more predefined hot words or user requests, and the second electronic device is configured to locally process the first sequence and the second sequence of audio data samples to identify the one or more predefined hot words or user requests in the audio signal.
13. The method of claim 12, wherein the user request comprises a request to stop capturing the audio signal, the method further comprising: A request to stop capturing the audio signal is generated by the second electronic device, the first electronic device being configured to suspend transmitting a sequence of audio data samples subsequent to the second sequence of audio data samples in response to the request.
14. The method according to claim 1, further comprising: receiving a second user action to stop capturing the audio signal while transmitting the second sequence of data samples; as well as In response to the second user action, receiving a sequence of audio data samples immediately following the second sequence of audio data samples is discontinued.
15. The method of any one of claims 1-14, wherein the first electronic device comprises a remote control device and the second electronic device comprises a network-connected television device configured to be controlled by the remote control device.
16. The method according to any one of claims 1-14, wherein the second electronic device comprises one or more processors and a memory storing one or more programs, wherein the programs are configured to implement an Android operating system and one or more user applications on the second electronic device.
17. The method of claim 16, wherein the Android operating system includes an audio manager module having instructions for processing the first sequence of audio data samples according to the real-time data processing mode.
18. The method according to claim 16, wherein The one or more user applications of the second electronic device include a file-based audio data processing module for caching the second sequence of audio data samples in the buffer and generating the data file including the second sequence of audio data samples in the batch data processing mode.
19. The method according to any one of claims 1 to 14, wherein: The second electronic device is determined not to support processing audio data samples in the real-time data processing mode during a recording session activated by a user action and processes both the first sequence and the second sequence of data samples during the recording session.
20. The method according to any one of claims 1 to 14, wherein The second electronic device is determined not to support processing audio data samples in the real-time data processing mode in an interruption separating two different recording sessions activated by two different user actions, and captures the first sequence and the second sequence of data samples during the two different recording sessions.
21. A method for processing audio data, comprising: receiving, by a second electronic device different from the first electronic device, a first sequence of audio data samples and a second sequence of audio data samples from the first electronic device, wherein the second sequence of audio data samples follows the first sequence of audio data samples in an audio signal captured by a microphone of the first electronic device; processing, by the second electronic device, the first sequence of audio data samples according to a batch data processing mode, comprising: caching the first sequence of audio data samples in a buffer of the second electronic device, and generating a data file comprising the first sequence of audio data samples; determining that at least one of communication, computing, and storage capabilities of the second electronic device is capable of supporting processing of the audio data samples in a real-time data processing mode; and Based on determining that at least one of the communication, computing and storage capabilities of the second electronic device is capable of supporting processing audio data samples in the real-time data processing mode, the second electronic device processes the second sequence of audio data samples according to the real-time data processing mode.
22. An electronic device comprising: one or more processors; as well as A memory having stored thereon instructions which, when executed by the one or more processors, cause the processors to perform the method according to any one of claims 1 to 21.
23. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the processors to perform the method of any one of claims 1 to 21.
Citation Information
Patent Citations
Audio data transmission method and device
CN107481709A
Method and device for obtaining audio data, equipment and computer readable storage medium
CN109976696A