Method and system for seamless media synchronization and switching

By integrating sensors such as microphones and cameras into electronic devices, and utilizing acoustic and image signal analysis to achieve seamless synchronization and switching of audio content, the problem of switching between devices in different ecosystems is solved, providing a smoother user experience.

CN114257324BActive Publication Date: 2026-05-22APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2021-08-27
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies cannot achieve seamless synchronization and switching of audio content between electronic devices in different ecosystems, requiring users to manually switch devices to continue playing audio content.

Method used

By integrating sensors such as microphones and cameras into electronic devices, acoustic and image signal analysis is used to determine the recognition information of audio content. When the audio system stops outputting, the system automatically retrieves audio content from local or remote storage and drives the speaker to continue outputting audio content, achieving seamless switching.

Benefits of technology

It enables seamless continuation of audio playback after the audio system stops, providing a smoother user experience and avoiding the inconvenience of manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114257324B_ABST
    Figure CN114257324B_ABST
Patent Text Reader

Abstract

The invention is entitled "Method and system for seamless media synchronization and switching." A method performed by a portable media player device is disclosed. The method receives a microphone signal that includes audio content output by an audio playback device via a loudspeaker. The method determines identifying information about the audio content, where the identifying information is determined through acoustic signal analysis of the microphone signal. In response to determining that the audio playback device has ceased outputting the audio content, the method retrieves an audio signal corresponding to the audio content from a local memory of the portable media player device or a remote device communicatively coupled thereto, and uses the audio signal to drive a speaker that is not part of the audio playback device to continue outputting the audio content and any additional audio content related to the audio content.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit and priority of U.S. Provisional Patent Application Serial No. 63 / 082,896, filed on September 24, 2020, which is incorporated herein by reference in its entirety. Technical Field

[0003] One aspect of this disclosure relates to seamlessly synchronizing and switching audio playback from an audio system to another audio output device. Other aspects are also described. Background Technology

[0004] With the proliferation of wireless multimedia devices, people are able to stream multimedia content from virtually anywhere. For example, these devices provide access to music streaming platforms that allow users to stream music for free. As a result, people consume far more media content than ever before. For instance, on average, people spend about thirty hours a week listening to music. Summary of the Invention

[0005] One aspect of this disclosure is a method performed by an electronic device (e.g., a portable media player device, such as a smartphone) that receives a microphone signal from a microphone of the electronic device. The microphone signal includes audio content output by an audio system (or audio playback device) via a loudspeaker arranged to project sound into the surrounding environment of the electronic device. In one aspect, the audio system may be (or include) a (e.g., a standalone) wireless device playing a radio broadcast. For example, the audio system may be part of a vehicle audio system, which itself may include the wireless device. In some aspects, the electronic device may not be part of an ecosystem that includes the audio system. Specifically, the electronic device may not be configured to receive data (such as metadata) describing or identifying the audio content being output by the audio system. For example, the electronic device may not be communicatively coupled to the audio system. Instead, the electronic device determines identification information of the audio content, wherein this identification information is determined through acoustic signal analysis of the microphone signal. For example, this analysis may include an audio recognition algorithm configured to identify the audio content (e.g., when the audio content is a musical work, the recognition algorithm may identify the title and / or artist of the work). The electronic device determines that the audio system has stopped outputting audio content. For example, the device may determine that the sound output level of the loudspeaker is below a threshold (e.g., indicating that the audio system has been turned off). In response to determining that the audio system has stopped outputting audio content, the device retrieves an audio signal corresponding to the audio content from local memory or from a remote device communicatively coupled to the electronic device, and uses the audio signal to drive a speaker that is not part of the audio system to continue outputting the audio content.

[0006] In one aspect, audio playback switching can be synchronized so that the electronic device continues outputting audio content from when the audio system stops. Specifically, the audio content may have a playback duration (e.g., a three-minute musical piece). The audio system may stop outputting the (first) musical piece at some point within the playback duration (e.g., the playback stop time) (such as at the one-minute mark). In this case, the retrieved audio signal may be the remainder of the musical piece that began at or after the one-minute mark. In some aspects, a fade-in audio transition effect may be applied to the beginning of the audio signal to provide a smooth transition.

[0007] In one respect, once the speakers have output (or played) audio content, the electronic device can stream the playlist. Specifically, once the remainder of the first musical piece has been output via the speakers, the electronic device can sequentially stream each of the several musical pieces in the playlist (e.g., stream the musical pieces one after another). Each musical piece in the playlist can be related to the first musical piece. For example, each piece in the playlist can be of the same genre, the same artist, or belong to at least one of the same albums as the first musical piece.

[0008] As described herein, the audio system may be part of a vehicle audio system. In one aspect, determining that the audio system has stopped outputting audio content includes determining at least one of the vehicle having stopped or the vehicle's engine having been turned off. In another aspect, electronic devices receive image data from a camera. Determining that the vehicle has stopped or the engine has been turned off includes performing an object recognition algorithm on the image data to detect objects that indicate that the vehicle has stopped or the engine has been turned off. For example, the electronic devices may detect objects indicating that the vehicle has stopped, such as a tree that has remained in the camera's field of view for an extended period of time (e.g., one minute).

[0009] As described herein, the audio content can be a radio broadcast being played (or being played) by a wireless device. In one aspect, the electronic device can stream the radio broadcast when it determines that the audio system has stopped outputting the broadcast. For example, the electronic device can process microphone signals according to a speech recognition algorithm to detect the speech contained therein and determine, based on that speech, that the output audio content is a radio broadcast. For example, the device can determine that the audio content is a radio broadcast based on speech including a radio station call sign. Furthermore, the device can identify the specific radio station associated with that call sign. For example, the device can receive location information (e.g., Global Positioning Satellite (GPS) data) indicating the current location of the electronic device. The device can identify the specific radio station broadcasting the radio broadcast based on the radio station call sign and location information. The device can continue outputting audio content by streaming the radio broadcast being broadcast by the radio station via a computer network (e.g., using streaming data associated with the radio program).

[0010] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed embodiments below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description

[0011] Multiple aspects are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings indicate similar elements. It should be noted that references to "a" or "an" aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a single drawing may be used to illustrate features of more than one aspect, and for a particular aspect, not all elements in that drawing may be necessary.

[0012] Figure 1 It illustrates several stages of seamlessly synchronizing audio content and switching it to another device to continue the output of the audio content, based on one aspect.

[0013] Figure 2 A block diagram of an audio system that continuously outputs audio content based on one aspect is shown.

[0014] Figure 3 A block diagram of an audio system that continuously outputs radio broadcast programs according to another aspect is shown.

[0015] Figure 4 This is a flowchart of one aspect of the process executed by the controller and server to continue outputting audio content.

[0016] Figure 5 This is a flowchart of one aspect of the process executed by the controller and server to continue outputting radio broadcast programs. Detailed Implementation

[0017] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Unless the shape, relative position, and other aspects of the components described in any aspect are explicitly defined, the scope of this disclosure is not limited to the components shown, which are for illustrative purposes only. Furthermore, while numerous details have been set forth, it should be understood that some embodiments may be implemented without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description. Moreover, unless the meaning is explicitly contrary, all scopes shown herein are to be considered to include the endpoints of each scope.

[0018] People can experience media content, such as audio content (e.g., music or musical works), on different electronic devices throughout the day. For example, while working, a person might listen to music on a desktop computer via a music streaming platform. At the end of the day, the person can pack up and continue listening to music on a multimedia device (e.g., a smartphone) via the same music streaming platform. In some cases, a user might switch to a multimedia device in the middle of a song (e.g., at some point during the duration of the audio content's playback). Once the content is streamed to the multimedia device, the music streaming platform can start a different song, or the song playing on the desktop computer can continue playing on the multimedia device. In this case, the song continues because the streaming platform tracks the playback position, and both user devices are connected to the platform via the same user account. Therefore, when a user switches to a multimedia device, the platform can transfer the playback position to the multimedia device to allow the user to continue streaming the content.

[0019] While music streaming platforms can offer switching between two devices, this method is only possible when both devices are part of the same ecosystem. Specifically, both devices are associated with the same user account via the same streaming platform. However, in other cases, where the devices are not part of the same ecosystem, such switching may not be possible. For example, during a morning commute to work (and an evening commute), a person might listen to music via a car radio. The person can tune the car radio to a specific station playing (or broadcasting) a particular type of music, such as country music. Once the person reaches their destination, they park the car and turn off the engine, which may disable all (or most) of the vehicle's electronics, including the radio. Thus, the radio may abruptly turn off in the middle of a song the person originally wanted to listen to. To finish listening to the song, the person could search for the song on their multimedia device and manually play it. For example, the song could be stored in the device's memory, or the person could search for the song on the streaming platform. However, doing so would start the song from the beginning, which may not be preferred, as the person might want to listen to the song only from where the radio stopped. The person might have to manually fast-forward the song to reach that part. Therefore, unlike the previous example, multimedia devices cannot provide a seamless transition when the wireless device is off because the two devices are not part of the same ecosystem. Thus, a seamless switching of media playback from one device to another electronic device is required to continue audio output.

[0020] This disclosure describes an audio system including electronics for seamless media synchronization and switching to electronic devices (which are not part of a playback device that is outputting the media (e.g., communicatively coupled to the playback device)).

[0021] Figure 1 This diagram illustrates several stages for seamlessly synchronizing and switching audio content to another device to continue the output of the audio content, according to one aspect. Specifically, the diagram shows three stages 10-12, in which audio content is seamlessly switched to an electronic device to continue the output of that content. Each stage includes an audio system (or playback device) 1, an audio source device 2, and an audio output device 3. In one aspect, the audio source device and the audio output device may be part of (e.g., a second) audio system configured to perform audio signal processing operations to seamlessly synchronize and switch the playback of audio content, such as those described herein. Figure 2 and Figure 3 The audio system 20. This article describes further details about the audio system.

[0022] In one aspect, the audio system 1 is shown as (e.g., a standalone) wireless device. In another aspect, the audio system 1 can be any system (or electronic device) configured to output audio content (e.g., a part thereof). For example, the audio system can be part of a vehicle audio system. Other examples may include at least one of (e.g., a part thereof) a standalone speaker, a smart speaker, a home theater system, a tablet computer, a laptop computer, a desktop computer, etc.

[0023] Audio source device 2 is shown as a portable media player (or multimedia) device, more specifically a smartphone. In one aspect, the source device can be any electronic device capable of performing audio signal processing operations and / or networking operations. Examples of such devices can include any electronic device described herein, such as a laptop, desktop computer, etc. In one aspect, the source device can be a portable device, such as a smartphone as shown. In another aspect, the source device can be a head-mounted device such as smart glasses, or a wearable device such as a smartwatch.

[0024] Audio output device 3 is an in-ear headphone (earphone or earbud) arranged to direct sound into the wearer's ear. Although only the left earbud is shown, the audio output device may also include a right earbud, both of which are configured to output (e.g., stereo) audio content. In one aspect, the output device can be any electronic device including at least one speaker, arranged to output sound by driving the speaker with at least one audio signal. In another aspect, the output device can be any electronic (e.g., headband) device (e.g., headphones) arranged to be worn by a user (e.g., on the user's head), such as the earbud shown herein. Other examples of headband devices may include on-ear headphones and over-ear headphones.

[0025] In one aspect, audio source device 2 may be wirelessly coupled to audio output device 3. For example, the source device may be configured to establish a wireless connection with the output device via any wireless communication protocol (e.g., the BLUETOOTH protocol). During the established connection, the source device may exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the output device, which may include audio digital data. In another aspect, the source device may be coupled via a wired connection. In one aspect, the audio output device may be part of (or integrated into) the audio source device. For example, the two devices may be a single integrated electronic device. Thus, at least some of the components of the audio output device (e.g., at least one processor, memory, at least one speaker, etc.) may be part of the audio source device. As described herein, at least some (or all) of the synchronization and switching operations may be performed by audio source device 2 and / or the audio output device.

[0026] However, conversely, audio source device 2 (and / or audio output device 3) is not communicatively coupled to audio system 1 (shown as not connected to each other). Specifically, the audio system may be part of a “device ecosystem” where devices within (or belonging to) the ecosystem are configured to be communicatively coupled (or connected) to each other to exchange data. Audio source device 2 may be a non-ecosystem device relative to audio system 1 because the source device is not communicatively coupled to exchange data. In one aspect, the audio system may be the only system (or device) that is part of that ecosystem. For example, the wireless device shown in the figure may be a simple (or non-smart) device configured only to receive audio content as radio waves for output through one or more loudspeakers. Therefore, the source device may not be able to be coupled to the audio system to exchange data because the audio system does not have (e.g., wireless) connectivity capabilities. In another aspect, the audio source device may not be communicatively coupled due to lack of connection (e.g., not yet paired with the audio system), but would otherwise be able to connect. Therefore, the source device may remain a non-ecosystem device until it is paired with the audio system (or with any device within the system's ecosystem).

[0027] On the other hand, audio source device 2 can connect to audio system 1, but can still be a non-ecosystem device relative to a user account. As described herein, audio system 1 can be configured to stream audio content (e.g., music) via a multimedia streaming platform, to which a user may have an associated user account. For streaming content, the audio system can connect to the platform (e.g., one or more remote servers of the platform) and gain access to the media content via a user account. In one aspect, audio source device 2 can be a non-ecosystem device, such that the device is not communicatively coupled to the platform to exchange data (e.g., via the same user account connected to the audio system). In another aspect, the audio source device can be coupled to the system (e.g., via a Universal Serial Bus (USB) connector), but can still not be a non-ecosystem device if the two devices are not associated with the same user account. For example, when the audio system is a vehicle audio system, the audio source device can be coupled via a USB connector for charging.

[0028] return Figure 1 Phase 10 illustrates that the audio system 1 is outputting audio content (e.g., music) 5 to the surrounding environment, which includes the audio source device 2 and the user (e.g., a person wearing an audio output device 3). Specifically, the audio system is outputting the audio content via one or more loudspeakers (not shown), each loudspeaker being arranged to project sound into the surrounding environment where the source device and / or output device are located.

[0029] In one aspect, while audio system 1 is outputting audio content (before or after), audio source device 2 may be performing one or more audio signal processing operations to determine identification information of the audio content. For example, the source device may be facing one or more microphones (e.g., Figure 2 The audio source device performs acoustic signal analysis on one or more microphone signals from microphone 22 to determine identification information. Specifically, the audio source device may receive one or more microphone signals from microphone 22, wherein the signals include audio content output by the audio system via a loudspeaker, as described herein. The audio source device may determine identification information of the audio content by acoustic signal analysis of the microphone signals. For example, when the audio content is a musical work, the analysis may determine the title of the work, the duration of the work, the musical artist performing the work, the genre of the work, etc. Thus, the audio source device may determine information about the audio content without receiving such data from the audio system 1 that is playing the audio (e.g., as metadata received via wired and / or wireless connections). Furthermore, the source device may determine the current playback time of the audio content being output by the audio system. Further explanation of these operations is described herein.

[0030] At stage 11, audio system 1 has stopped (paused or stopped) outputting audio content 5. For example, the wireless device may have been turned off by the user (e.g., the user switched the power switch to the off position). Alternatively, the user may have adjusted the playback by manually pausing or stopping the audio content. Returning to the previous example, when the wireless device is part of a vehicle's audio system, the audio content may have stopped when the vehicle's engine is off. In one aspect, the source device can determine that the audio system has stopped outputting audio content. For example, the device can determine that the sound pressure level (SPL) of the microphone signal is below an SPL threshold. This document describes further explanation of how to determine that audio content has stopped. In another aspect, the source device can determine the moment (or playback stop time) when the musical piece stopped during the playback duration.

[0031] At stage 12, audio output device 3 continues to output audio content. Specifically, in response to the audio source device determining that audio system 1 has stopped outputting audio content, the audio source device can be configured to retrieve audio content (e.g., from local memory or from a remote device) as one or more audio signals using identification information, and use these audio signals to drive speakers to continue outputting audio content. In one aspect, the audio source device can use the audio signals to drive one or more speakers that are not part of the audio source device to continue outputting audio content. For example, the audio source device transmits the retrieved audio signals to the audio output device via a connection between two devices to drive the speakers of the output device. In one aspect, the sound produced by the audio output device (e.g., its speakers) can be synchronized with the output of audio system 1. Specifically, the retrieved audio content can be the remainder of the audio content output by audio system 1 (e.g., after the playback stop time). Therefore, the audio output device can continue outputting audio content as if switching audio outputs, without (or momentarily) pausing playback. Thus, even after audio system 1 has stopped playing, the user can continue to enjoy the audio content.

[0032] As described in this article, Figure 2 and Figure 3 A block diagram of audio system 20 is shown. In one aspect, the two figures show a block diagram in which identification information of audio content being output by an audio playback device (e.g., not communicatively coupled to audio system 20 by a wired and / or wireless (e.g., radio frequency (RF)) connection) is determined by acoustic signal analysis of the audio content, and in response to the (audio system) determining that the audio playback device has stopped outputting audio content, audio content continues to be output via one or more speakers, wherein, for example, Figure 2 It describes continuing to output the remaining portion of the audio content (e.g., the remainder of a musical piece), and Figure 3 The description continues the output of radio broadcasts. Further instructions regarding the operation of these accompanying figures are described below.

[0033] On one hand, in Figure 2 and Figure 3 The elements of the audio system 20 shown and described herein that perform audio digital signal processing operations can be implemented (e.g., as software) as one or more programmed digital processors (generally referred to herein as "processors" that execute instructions stored in memory). For example, at least some of the operations described herein can be performed by the processor of the audio source device 2 and / or the processor of the audio output device 3. Alternatively, all operations can be performed by at least one processor of either device.

[0034] Turn now Figure 2This figure illustrates a block diagram of an audio system 20 that continuously outputs audio content according to one aspect. Specifically, the system includes a camera 21, a microphone 22, a controller 25, a speaker 23, and a (remote) server 24. In one aspect, the system may include more or fewer elements (or components). For example, the system may include at least one (e.g., touch-sensitive) display screen. Alternatively, the system may not include server 24. In some aspects, the elements described herein may be part of either (or both of) an audio source device 2 and an audio output device 3, which may be part of the audio system 20. For example, the camera, microphone, and controller may be part of a source device (which may also include a display screen) (e.g., integrated into the source device), while the speaker may be part of the audio output device. In some aspects, the audio system 20 may include audio source device 2, audio output device 3, and / or server 24, as described herein.

[0035] Microphone 22 is an “external” microphone arranged to capture (or sense) sound from the surrounding environment as a microphone signal. Camera 21 is configured to generate image data (e.g., video and / or still images) of a scene containing the surrounding environment within the camera’s field of view. In one aspect, the camera may be part of an audio source device and / or audio output device (e.g., integrated with an audio source device and / or audio output device). In another aspect, the camera may be a separate electronic device. Speaker 23 may be, for example, an electrically driven driver specifically designed for sound output in a particular frequency band, such as a woofer, tweeter, or midrange driver. In one aspect, the speaker may be a “full-range” (or “full-frequency”) electrically driven driver that reproduces as much of the audible frequency range as possible. In another aspect, the speaker may be an “internal” speaker arranged to project sound into (or toward) the user’s ear. For example, as... Figure 1 As shown, the audio output device 3 may include one or more speakers 23, which are arranged to project sound directly into the ear canal of the user's ear when the output device is inserted into the ear canal of the user's ear.

[0036] In one aspect, the audio system 20 may include one or more “extra-ear” speakers (e.g., as speaker 23) arranged to project sound directly into the surrounding environment. Specifically, the audio output device may include an array of (two or more) extra-ear speakers configured to project directional beam patterns of sound at locations within the environment, such as directing the beam toward the user’s ears. For example, when the audio output device is a head-mounted device such as smart glasses, the extra-ear speakers can project sound into the user’s ears. Therefore, the audio system may include a sound output beamformer (e.g., where controller 25 is configured to perform beamformer operation), configured to receive one or more audio signals (e.g., audio signals containing audio content as described herein) and configured to generate speaker driver signals that, when used to drive the extra-ear speakers, produce a spatially selective sound output in the form of sound output beam patterns, each pattern containing at least a portion of the audio signal.

[0037] The (remote) server 24 can be any electronic device configured to communicatively couple to the controller 25 (e.g., via a computer network, such as the Internet) and configured to perform one or more signal processing operations as described herein. For example, the controller 25 can be coupled to the server via any network such as a wireless local area network (WLAN), a wide area network (WAN), a cellular network, etc., so that the controller can exchange data (e.g., IP packets) with the server 24. Further description of the server is provided herein.

[0038] Controller 25 may be a dedicated processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform seamless audio synchronization and switching to one or more electronic devices of the audio system 20, as described herein. The controller includes several operational blocks, such as an audio content recognizer 26, an audio stop detector 27, a playlist generator 28, and a content player 29. The operational blocks are discussed below.

[0039] Audio content recognizer 26 is configured to receive a microphone signal from microphone 22, the microphone signal including audio content being output by audio system 1. Recognizer 26 is configured to determine (or identify) identification information (or about the audio content) contained within the microphone signal. In one aspect, the recognizer determines this information through acoustic signal analysis of the microphone signal. Specifically, the recognizer may process the microphone signal according to an audio (or sound) recognition algorithm to identify one or more audio patterns (e.g., spectral content) known to be associated with a particular audio content. In one aspect, the algorithm compares the spectral content (e.g., at least a portion thereof) with known (e.g., stored) spectral content. For example, the algorithm may perform a lookup table in a data structure using the spectral content of the microphone signal, the data structure associating the spectral content with known audio content or more specifically with identification information about the audio content. In one aspect, the identification information may include characteristics of the audio content, such as a description of the audio content (e.g., title, type of audio content, etc.), the author of the audio content, and the playback duration of the audio content. For example, when the audio content is a musical work, the identification information may include characteristics such as the title of the musical work, the duration of the performance, the artist, and the genre. Upon identifying known spectral content that matches the spectral content of the microphone signal, the algorithm selects the identification information associated with that match. In one aspect, the audio content recognizer can use any known method for identifying audio content contained within the microphone signal and determining the associated identification information.

[0040] In one aspect, the audio content recognizer 26 can perform acoustic signal analysis by executing a speech recognition algorithm on the microphone signal to detect speech contained therein (e.g., at least one word). Similar to the previous process, the recognizer can identify the audio content by comparing this speech with (stored) speech associated with known audio content. Further explanation regarding the use of speech recognition algorithms is described herein.

[0041] In one aspect, the audio content recognizer 26 is configured to determine (or identify) the current playback time (or moment) within the playback duration of the identified audio content. Specifically, the recognizer determines (and stores) a timestamp of the current playback time. In one aspect, this timestamp may be a part of the determined identification information. For example, the spectral content of the microphone signal used to perform acoustic analysis may be associated with a specific moment during the playback duration. In another aspect, the timestamp may be the moment the recognizer determines the identification information (and / or receives the microphone signal) relative to the start time of the audio content. For example, the audio content recognizer may determine the start time of the audio content based on the microphone signal. Specifically, the recognizer analyzes the microphone signal to determine whether the signal's SPL is below a threshold for a period of time (e.g., two seconds). In the case of musical works, this period of time may correspond to the time between two tracks. The recognizer may store this as an initial timestamp, and when configured to determine identification information, may store the time increment from the initial timestamp to the current playback time (as a timestamp).

[0042] In one aspect, the audio content recognizer 26 can perform acoustic signal analysis periodically (e.g., every ten seconds). For example, once the analysis is performed to determine the recognition information of the audio content, the analysis can be performed periodically to update the timestamp and / or determine whether the audio content has changed. By performing this analysis periodically, the audio system can conserve system resources (e.g., battery power from which the controller can draw power). In another aspect, the audio content recognizer can perform acoustic signal analysis continuously.

[0043] An audio stop detector 27 is configured to detect whether the audio content being captured by microphone 22 has stopped. Specifically, the detector determines whether the audio system 1, which is outputting audio content, has stopped outputting audio content. In one aspect, the detector can determine that the audio system has stopped outputting content by determining that the sound output level of the audio system is below a (predetermined) threshold. Specifically, the detector determines whether the SPL of the microphone signal is below the SPL threshold. In some aspects, this determination may be based on whether the SPL has been below the threshold for a period of time (e.g., four seconds).

[0044] In one aspect, the audio stop detector 27 can detect whether audio output has stopped based on image data captured by the camera 21. Specifically, the detector can receive image data from the camera 21 and perform an object recognition algorithm on the image data to detect objects contained therein that indicate that the audio system has stopped (or will stop) outputting audio content. For example, the detector can detect objects that indicate that the radio device 1 is turned off, such as a dial or switch positioned in the off position. Alternatively, the detector can detect that light (e.g., the backlight used to illuminate the tuner of the radio device) is not lit, indicating that the radio device is not powered on.

[0045] On the other hand, the detector can determine that audio output has stopped (e.g., will be experienced by the user of the audio output device) based on objects not detected within the camera's field of view. For example, when the audio system 1 is detected to be outside the camera's field of view (e.g., for a period of time), the detector can determine that the user of the audio output device is not experiencing audio content. This might be the case, for example, when the camera is part of the audio output device and the user leaves the radio (e.g., and is in a different room).

[0046] As described herein, the audio system outputting audio content may be part of the vehicle's audio system. In one aspect, the detector may determine that the vehicle audio system has stopped outputting audio content based on the vehicle's state. For example, the detector may determine that the vehicle has stopped and / or the vehicle's engine has been turned off, each of which could indicate that the vehicle audio system has stopped outputting sound. The detector may determine these situations in various ways. For example, the detector may perform an object recognition algorithm on image data captured by camera 21 to detect one or more objects indicating whether the vehicle has stopped (e.g., is stationary and / or parked) and / or the engine is off. For example, the detector may determine that the vehicle has stopped by detecting that 1) a user is inside the vehicle (based on identification of the vehicle interior) and 2) an object outside the vehicle has remained stationary for a period of time. As another example, the detector may determine, based on image data, that a dashboard light that was illuminated when the vehicle engine was started is no longer illuminated, indicating that the vehicle engine is off. As yet another example, the detector may determine whether the vehicle has stopped and whether the engine is off based on the absence of engine and / or road noise. For example, the detector may obtain one or more microphone signals from microphone 22 configured to capture ambient sound. The detector can apply an audio classification model to the microphone signal to determine the presence of engine sound and / or road noise. If none are present, it can be determined that the vehicle has stopped and the engine is off.

[0047] In one respect, the detector can determine that the vehicle's engine is off by other methods. For example, the detector can determine that the source device (and / or output device) is no longer in motion. For example, the detector can obtain location information indicating the location of the device (e.g., from...). Figure 3The location identifier 33 shown (such as GPS data) determines that the device is no longer in motion based on the fact that the device's location has hardly changed (e.g., for at least a period of time). Alternatively, motion can be based on motion data (or IMU data) obtained from one or more inertial measurement units (IMUs) of the audio system (e.g., which may be integrated into any device in the system). Furthermore, the detector (or more specifically, the controller) can be communicatively coupled to an onboard computer that can notify the detector when the engine is started or stopped. Alternatively, the detector can determine that the vehicle is off based on whether a wireless connection (e.g., a BLUETOOTH connection) with the vehicle's audio system has been lost. In some aspects, the detector can use any method to determine whether the vehicle's audio system has stopped outputting audio content.

[0048] Content player 29 is configured to continue outputting audio content through speaker 23 (e.g., in response to controller 25 determining that the audio playback device has stopped outputting audio content). Specifically, the content player may be a media player algorithm (or software application) configured to retrieve an audio signal corresponding to (or containing) audio content and use that audio signal to drive speaker 23 to output the audio content. Specifically, the content player is configured to receive identification information of the audio content from audio content recognizer 26 and is configured to retrieve the audio content based on that information. For example, as described herein, the identification information may include characteristics of the audio content (e.g., title, author, etc.). The content player may perform a search (e.g., a search in local memory and / or on a remote device such as server 24) to retrieve an audio signal containing the audio content. Once retrieved, the player uses that audio signal to drive speaker 23.

[0049] As described herein, audio system 20 is configured to seamlessly synchronize and switch audio content to continue the output of audio content. The process of these operations performed by controller 25 will now be described. Specifically, audio stop detector 27 determines that audio system (or device) 1 has stopped outputting audio content, as described herein. The detector notifies audio content recognizer that the audio playback device is no longer outputting sound. In response, audio content recognizer determines the playback stop time as the moment when the audio content stops (e.g., will be output by the audio playback device) during the playback duration of the audio content. In one aspect, the recognizer can determine the playback stop time by performing acoustic signal analysis (e.g., identifying where a portion of the last received spectral content is located within the playback duration of the audio content). In another aspect, the audio content recognizer can determine the playback stop time as the time increment from the time when the timestamp of the last audio content was determined to the time when the detector notifies the recognizer that the audio content has stopped. The recognizer can transmit the identification information of the audio content and the playback stop time to content player 29. The content player uses the identification information and / or the playback stop time (e.g., from local or remote memory) to retrieve the audio content as an audio signal and uses the audio signal to drive the speaker. Specifically, the audio signal may include the remainder of the audio content that begins at or after the playback stop time (within the playback duration). On the other hand, the audio signal may include the entire playback duration of the audio content, but the content player may begin playback at or after the playback stop time. Therefore, playback of the audio content is switched to audio system 20 (e.g., an audio output device) and synchronized with audio system 1, so that the content continues to play, just as if playback were switched from audio system 1 to audio system 20.

[0050] In one aspect, the content player 29 performs one or more audio signal processing operations. In some aspects, the player may apply one or more audio effects to the audio signal. For example, the player may apply an audio transition effect to the audio signal while continuing to output audio content. For example, the player may apply a fade-in audio transition effect to the beginning of the audio signal by increasing the sound output level of speaker 23 to a preferred user level. Specifically, the player may adjust the direct-to-reverberation ratio of the audio signal over a period of time according to the playback duration. Fade-in audio content, which may be more preferable than abruptly outputting audio content at a preferred user level. In another aspect, the content player may apply other audio processing operations, such as equalization operations and spectrum shaping operations.

[0051] In one aspect, controller 25 may optionally include playlist generator 28 configured to generate playlists based on identified audio content. Specifically, the generator generates playlists containing audio content associated with the identified audio content for playback by audio system 20. For example, the generator receives identification information from an audio content recognizer and uses that information to generate the playlist. In one aspect, the generator may use that information to perform a lookup in a data structure that associates similar information with audio content. In some aspects, audio content having at least one similarity to the identified audio content (e.g., belonging to the same genre) may be selected by the generator as part of the playlist. In one aspect, the playlist may be a data structure that includes identification information of audio content selected in a specific order. For illustration, when the audio content is a (first) musical work, the playlist generator may use information from the first musical work to generate a list of several similar musical works (e.g., two or more musical works).

[0052] In one respect, the related audio content listed in a playlist may have one or more characteristics similar to (or identical to) the identified audio content. For example, when the audio content is country music, the playlist may include several country music works. Similarly, when a specific artist is identified as the author (performer or singer) of the country music work, the playlist may be a playlist for that specific artist, including several musical works by that artist. Therefore, a playlist may be a list of similar (or related) musical works, which may be at least one of the same genre, the same artist, or belong to the same album as the first musical work.

[0053] In one aspect, the playlist generator may be notified by audio content previously identified by the audio content recognizer 26. Specifically, the generator may generate playlists based on the playback history of the audio system 1. For example, the playlist generator may continuously (or periodically) retrieve identification information of audio content identified by the audio content recognizer 26. When the recognizer identifies different audio content, the playlist generator may receive this information and generate (or update) an existing playback model based on one or more characteristics of the identified audio content. In one aspect, the playback model may be a data structure indicating the characteristics of the most frequently identified audio content. In another aspect, the playlist generator may have models for different types of audio content. For example, the playlist generator may have a country music playback model indicating which country music works have been previously identified. Returning to the previous example, the country music playback model may indicate that a particular artist is the primary identified country music artist. Therefore, when the identification information of the current audio content is identified as country music genre, the playlist generator may use the country music model to generate a playlist that may have (e.g., primary) musical works by a particular artist. In one respect, each model can be updated over time.

[0054] In some aspects, the playlist generator can be notified via the playback history of the content player. For example, the content player 29 can be configured to retrieve audio content in response to user input, which can be received via an audio source device (e.g., via user touch input on a touch-sensitive display). Specifically, the audio source device can display a graphical user interface (GUI) of a media player software application and select songs for playback. In one aspect, such user selections can be recognized by the playlist generator to generate playlists based on playback history (e.g., playback models). In some aspects, the playlist generator can generate playlists based on user preferences. For example, the playlist generator can access the user account of a streaming platform associated with the user to identify characteristics of the audio content the user prefers to listen to.

[0055] In one aspect, playlist generator 28 can generate playlists for each identified audio content. In another aspect, once audio system 1 has stopped outputting audio content, the playlist generator can transmit the generated playlist to content player 29, which can use the playlist to continue outputting audio content after the remainder of the current audio content has been completed. Specifically, once the remainder has been output via the speaker, the player can stream the audio content listed in the generated playlist (e.g., in sequential order). For example, content player 29 uses the identification information of one (e.g., the first) audio content listed in the playlist to retrieve the associated audio signal (e.g., from local memory or from server 24). For example, when the audio content is a musical work, the playlist can be several (e.g., related) musical works, which are streamed by the player once the speaker outputs the musical work. In one aspect, content player 29 will continue playing the audio content in the playlist until a stop user input is received.

[0056] Figure 3 A block diagram of an audio system for continuously outputting radio broadcast programs is shown, according to another aspect. As described herein, a “radio broadcast program” can include audio content broadcast by a radio station. For example, the program can be a specific segment (e.g., a programming block) broadcast by a radio station. Examples of radio broadcast programs can include music, radio talk shows (or public shows), news programs, etc. In one aspect, radio broadcast programs can be broadcast via frequency modulation (FM) or amplitude modulation (AM) broadcasts or associated digital services (part of a broadcast channel).

[0057] As shown in the figure, this diagram illustrates operations performed by controller 25 for streaming radio broadcast programs over a computer network. In one aspect, controller 25 may perform at least some of the operations shown in the figure in response to determining that the audio content being played by audio system 1 is a radio broadcast program. Otherwise, the controller may perform operations such as... Figure 2 The aforementioned operations. This document provides further details on determining which operations to perform.

[0058] As shown in the figure, the controller includes several different (and similar) operation blocks, such as Figure 2 As shown. Specifically, controller 25 includes an audio stop detector 27, a radio station classifier 31, a radio station finder 32, a location identifier 33, and a content player 29.

[0059] In one aspect, the radio station classifier 31 is configured to determine whether the audio content captured by the microphone 22 is being broadcast by a radio station (and is being played by an audio playback device, e.g., Figure 1 The radio equipment 1) picks up radio broadcasts. In one aspect, a classifier can perform operations similar to an audio content recognizer to make this determination. Specifically, the classifier can determine identification information (of the audio content) that indicates the audio content is a radio broadcast (as associated with) and that the source of the audio content is the radio station broadcasting the program. For example, the classifier can perform acoustic signal analysis on the audio content of a microphone signal to determine identification information (of the audio content) that indicates the audio content is a radio broadcast (as associated with) by the radio station. For example, the classifier can process the signal according to an audio recognition algorithm (such as a speech recognition algorithm) to detect speech (as identification information) contained within the audio content that indicates the audio content is a radio broadcast. To determine whether the audio content is a radio broadcast, the classifier can determine whether the audio content contains speech (words or phrases) associated with radio broadcasting. For example, in the United States, the Federal Communications Commission (FCC) requires radio stations to broadcast their station call sign, which is a multi-letter (e.g., typically three to four letters) that uniquely identifies the radio station as close to the hour as possible. Therefore, a radio classifier can determine whether the speech contained within audio content includes a radio call sign. Specifically, the classifier can use the speech to perform a table lookup in a data structure that includes (for example, known) call signs. If a call sign is identified, the audio content is determined to be a radio broadcast program, and the source of the program is a radio station with the identified call sign.

[0060] On the other hand, the radio classifier 31 can determine whether the audio content is a radio broadcast program based on identification information (e.g., call sign) that can be displayed on the vehicle audio system's display. Most FM radio broadcasts include Radio Data System (RDS) information embedded in the FM signal. This identification information may include time, station information (e.g., station call sign), and program information (e.g., information related to the radio broadcast program being broadcast). When a station is tuned to, the vehicle audio system receives the station's RDS information and displays it on its display. In one aspect, the classifier can determine the identification information by receiving image data captured by the display's camera 21 and performing an object recognition algorithm on the image data to identify at least some of the displayed identification information to be used to determine whether the audio content is a radio broadcast program (e.g., whether the image data includes a station call sign, etc.).

[0061] On the other hand, RDS information can be obtained via other methods. For example, the vehicle audio system can transmit data to controller 25 (e.g., via any wireless protocol, such as the BLUETOOTH protocol).

[0062] On the other hand, the radio station classifier 31 can be configured to determine whether audio content is a radio broadcast program based on past acoustic analysis. For example, the classifier can compare the spectral content of the microphone signal with previously captured spectral content associated with a specific radio broadcast program. For example, a vehicle audio system may play the same radio broadcast program every day (e.g., during a morning commute). During previous acoustic analysis, the classifier may have already identified the radio stations associated with the radio program and stored the microphone's spectral content and information (e.g., as a data structure) for later comparison.

[0063] In one aspect, controller 25 can verify the validity of at least some certain identification information. A call sign may consist of several letters (e.g., two, three, four, etc.). In some cases, a radio station may share at least one common letter with another station. For example, a first station may have a four-letter call sign, while a different second station may have a three-letter call sign that may include all (or some) of the letters of the first station. To prevent misclassification (e.g., classifying the first station as the second station, which could occur if the last letter of the four-letter call sign is not recognized by the speech recognition algorithm), controller 25 is configured to confirm (e.g., with high certainty) the identified radio station based on (at least some) identification information (e.g., the identified station call sign) and additional information such as location information. Specifically, location identifier 33 is configured to identify the location of controller 25 (e.g., or more specifically, the location of the device when the controller is part of source device 2). For example, the location identifier may be an electronic component, such as a GPS component configured to determine the location information as coordinates. Radio station finder 32 is configured to receive location information and determined identification information (e.g., the identified radio station call sign), and is configured to identify (or confirm) a radio station broadcasting a radio program based on the determined identification information and the location information. For example, the radio station finder may include a data structure containing a list of broadcast radio stations, their call signs, and their locations. The radio station finder can perform a lookup in the data structure using the location information and the radio station call sign to identify radio stations with matching characteristics. Thus, returning to the previous example, the first radio station may be located in one city (e.g., Los Angeles, California), and the second radio station may be located in a different city (e.g., Denver, Colorado). When the location information indicates that the device is in Los Angeles (or nearby), radio station finder 32 may select the first radio station because three of the four letters of the call sign are identified and the station is located in Los Angeles. In one aspect, the radio station finder may (e.g., periodically) update the data structure via server 24.

[0064] In one respect, a radio station can be associated with one or more call signs. In this case, the station can broadcast from several different locations, each with a unique call sign. Therefore, the data structure used by the radio station finder 32 can associate each station with one or more call signs used by that station for broadcasting radio programs.

[0065] In one aspect, upon identifying (e.g., confirming) a radio station, the station finder 32 can determine the streaming data of the radio broadcast program. For example, most (if not all) radio stations stream radio broadcast programs (e.g., in real time). Therefore, upon identifying a station, the finder can determine streaming data, such as a streaming Uniform Resource Locator (URL) associated with the radio program. Specifically, the station finder's data structure may include streaming data. Thus, the station finder can perform a table lookup on the data structure using at least some identification information associated with the radio station (e.g., the station call sign) to determine the URL. In another aspect, the station finder can retrieve data from server 24. Upon identifying streaming data, the content player 29 can use that data (e.g., accessing the URL) to continue the output of audio content by streaming the radio broadcast program being broadcast by the radio station via a computer network.

[0066] In one aspect, the controller may be configured to stream similar (or related) audio content being broadcast by a radio station over a computer network, rather than streaming the radio broadcast program. In another aspect, similar audio content may include audio content having at least one similar characteristic, such as belonging to the same genre, having the same author, etc. In some aspects, the radio station finder 32 may be configured to identify radio broadcast programs similar to those identified by the classifier and stream those similar programs. This may be the case when the finder cannot find streaming data for the identified radio broadcast program. In this case, the finder may determine the characteristics of the identified radio program (e.g., the type of content, such as talk radio or music genre) and identify similar programs for streaming (e.g., by performing a table lookup, as described herein).

[0067] In one respect, the audio system (e.g., its controller 25) may perform based on confidence scores. Figure 2 or Figure 3 The operation described herein. Specifically, the controller may first execute... Figure 3 At least some of the operations in the process can then be performed when the confidence score falls below a threshold. Figure 2The operation can be performed as follows: In one aspect, the confidence score may be based on whether the audio system can determine (e.g., within deterministic limits) that the audio content played by the audio playback device is a radio broadcast. In this case, the radio station classifier 31 can monitor microphone signals to identify characteristics of the radio broadcast, such as call sign identification, as described herein. In one aspect, the confidence score may be based on multiple characteristics that the classifier can identify (e.g., call sign, type of audio content, etc.). In one aspect, the classifier may perform these operations for a period of time (e.g., one minute). However, if the classifier cannot identify (e.g., at least one) a characteristic or the confidence score is below a threshold, the controller 25 may begin execution. Figure 2 The operations described herein (e.g., audio content recognizer 26 can begin determining recognition information for the audio content). On the other hand, the audio system can perform operations based on user preferences.

[0068] The operations described so far are performed (at least in part) by the controller 25 of the audio system 20. In one aspect, at least some of the operations may be performed by one or more servers 24 communicatively coupled to the controller 25 (or more specifically, to the audio source device 2 and / or the audio output device 3). For example, the server 24 may perform the operations of the audio content recognizer 26 to determine recognition information. Once determined, the server 24 may (e.g., via a computer network) transmit the recognition information to the recognizer 26 (e.g., its controller 25). In another aspect, the server 24 may perform additional operations. The following flowchart includes at least some of the operations described herein performed by the server.

[0069] Figure 4 This is a flowchart of one aspect of a process 40 executed by controller 25 and server 24 for the continued output of audio content. Specifically, the diagram illustrates process 40, where, relative to... Figure 2 At least some of the operations described are performed by the controller and the server.

[0070] Process 50 begins with the controller receiving a microphone signal, which includes signals from an audio system (e.g., Figures 1 to 3The system 1 shown outputs audio content via a loudspeaker (at box 41). A controller transmits microphone signals to server 24 (e.g., via a computer network). In one aspect, the controller may transmit at least a portion of the audio data (e.g., spectral content) associated with the microphone signals. Server 24 determines identification information of the audio content by performing acoustic signal analysis of the microphone signals (at box 42). Specifically, the server may perform the operation of audio content recognizer 26 to determine identification information (e.g., description of the audio content, author of the audio content, etc.). Controller 25 determines that the audio system has stopped outputting audio content (at box 43). For example, an audio stop detector 27 may determine that the system has been shut down because the loudspeaker's sound output level (as measured by the microphone signals) is below a threshold. In response, the controller may transmit a message indicating that audio output has stopped. In one aspect, the message may include a timestamp indicating the time when playback of the audio content stopped (e.g., during the duration of playback of the audio content). The server retrieves the audio content (e.g., the audio signal corresponding to it) (at box 44). Specifically, the server can retrieve the audio signal using identification information and a timestamp to perform operations on the content player. Specifically, the audio signal can be the remainder of the audio signal based on the playback stop time indicated by the timestamp.

[0071] Server 24 transmits (or streams) the retrieved audio content (e.g., its remainder) to controller 25. Controller 25 receives the audio content (e.g., the audio signal corresponding to it) and continues to output the audio content (at box 45). Specifically, the controller synchronizes the continued output of the audio content with the paused output of audio system 1 by outputting the remainder of the audio content that begins at or after the playback stop time. In one aspect, for continued output, the controller uses the received audio signal to drive speaker 23 (which is not part of the audio playback device), as described herein. In another aspect, controller 25 may apply audio transition effects, as described herein.

[0072] Furthermore, server 24 can generate a playlist based on the identified audio content (at box 46). For example, the server can perform the operation of playlist generator 28 by generating a playlist of similar (related) audio content using identification information. For example, when the audio content is a musical work, the playlist may include several musical works (e.g., their identification information) that are each similar to the (original) musical work (e.g., of the same genre). For example, once the remainder of the audio content has been output by speaker 23, controller 25 streams the playlist (at box 47). For example, the playlist may be streamed sequentially for each audio item listed in the playlist (using its identification information).

[0073] Figure 5 This is a flowchart of one aspect of the process 50 for continuing to output radio broadcast programs, executed by controller 25 and server 24. Specifically, the diagram illustrates process 50, where, relative to... Figure 2 At least some of the operations described are performed by the controller and the server.

[0074] Process 50 begins with controller 25 receiving a microphone signal, which includes audio content output by audio system 1 via a loudspeaker (in box 51). The controller receives location information, such as GPS data (in box 52). For example, a location identifier can provide location information indicating the location of either audio source device 2 or audio output device 3 (or both). The controller then transmits the microphone signal and location information to a server.

[0075] Server 24 uses the microphone signal and the location information to determine that the audio content is a radio broadcast program being broadcast by a specific radio station (in box 53). At this stage, the server may perform operations performed by the radio station classifier 31 and / or the radio station finder 32, as described herein. For example, the server may determine identification information of the audio content being output by the audio system through acoustic signal analysis of the output audio content. Controller 25 determines that the audio system 1 has stopped outputting audio content (in box 54). For example, the audio stop detector 27 may determine that the SPL of the microphone signal is below the SPL threshold, as described herein. The controller transmits a message to the server indicating that audio output has stopped. The server determines the streaming data of the radio broadcast program, such as the streaming URL (in box 55). For example, the server may perform a table lookup in a data structure that associates the identification information with the streaming data using identification information associated with the radio broadcast program and / or the radio station. The server transmits the streaming data to the controller. The controller uses the streaming data to stream the radio broadcast program (in box 56).

[0076] Some aspects can be addressed Figure 4 and Figure 5The processes 40 and 50 described herein may be modified. For example, at least some specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects. For example, generating (and transmitting) a playlist at box 46 may be performed simultaneously with retrieving audio content at box 44. Thus, the generated playlist may be transmitted to controller 25 along with the retrieved audio content. On the other hand, the server may not retrieve audio content as described in box 44. Instead, the server may transmit identification information (along with timestamps) of the identified audio content to controller 25, which can use this information (e.g., from local or remote memory) to retrieve the audio content. As another example, instead of transmitting streaming data of a radio broadcast program in process 50, the server may retrieve the streaming data and transmit it to controller 25 for output.

[0077] As described herein, controller 25 continues the output of audio content (e.g., at boxes 45 and 56 in processes 40 and 50, respectively). In one aspect, these operations can be performed automatically. For example, the controller can continue outputting audio content in response to receiving audio content from server 24. In another aspect, controller 25 can continue outputting audio content in response to receiving authorization from a user. Specifically, controller 25 can provide a notification including a request to continue the output of audio content (e.g., via a speaker). For example, the controller can display a pop-up notification on the display of the audio source device, which includes a GUI item for authorizing (or denying) the continuation of audio playback. In one aspect, in response to receiving user authorization (e.g., in response to receiving a user selection of a GUI item via a touch-sensitive display), controller 25 can continue playing audio content (e.g., by driving speaker 23 with an audio signal).

[0078] In one aspect, at least some of the operations described herein (e.g., processes 40 and / or 50) may be performed by a machine learning algorithm configured to seamlessly synchronize and switch media to non-ecosystem devices. For example, to generate a playlist, a playlist generator may execute a machine learning algorithm to determine the optimal playlist that follows an audio content identified by audio content recognizer 26.

[0079] In one aspect, at least some of the operations described herein are operable or optional operations. Specifically, optional operations are shown as boxes with dashed lines or dashed borders. For example, by Figure 2The operation performed by the playlist generator 28 can be optional, allowing the audio system 20 to output only the remainder of the audio content. Once complete, the audio system can prompt the user (e.g., via an audible alert) asking if the user wishes to continue listening to (e.g., similar) audio content. Furthermore, Figure 2 and Figure 3 At least some of the operation blocks in the process are communicatively coupled to the server via dashed lines. In one aspect, each of these operation blocks may optionally communicate with the server to perform one or more operations. For example, Figure 3 The radio station finder 32 can (e.g., periodically) retrieve (e.g., an updated) list of broadcast radio stations, their call signs, and their locations. Alternatively, the radio station finder can retrieve this list from local memory.

[0080] Personal information used should comply with practices and privacy policies that are generally recognized as meeting (and / or exceeding) government and / or industry requirements for protecting user privacy. For example, any information should be managed to mitigate the risk of unauthorized or unintentional access or use, and users should be clearly informed of the nature of any authorized use.

[0081] As previously described, one aspect of this disclosure may be a non-transitory machine-readable medium (such as microelectronic memory) on which instructions are stored, programming one or more data processing units (generally referred to herein as a "processor") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing units and fixed hard-wired circuit components.

[0082] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and that this disclosure is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.

[0083] In some aspects, this disclosure may include languages ​​such as "[element A] and [element B]". This language can refer to one or more of these elements. For example, "at least one of A and B" can mean "A", "B", or "A and B". Specifically, "at least one of A and B" can mean "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include languages ​​such as "[element A], [element B], and / or [element C]". This language can refer to any of these elements or any combination thereof. For example, "A, B, and / or C" can mean "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".

Claims

1. A method performed by a portable media player device, the method comprising: The portable media player device receives a microphone signal from its microphone, the microphone signal including audio content output by an audio playback device via a loudspeaker, the loudspeaker being arranged to project sound into the surrounding environment where the portable media player device is located; Without receiving information about the audio content from another device, the identification information of the audio content is automatically determined by performing acoustic signal analysis on the microphone signal; as well as In response to determining that the audio playback device has stopped outputting the audio content, The identification information of the audio content is used to retrieve the audio signal corresponding to the audio content from the local memory of the portable media player device or a remote device communicatively coupled to the portable media player device. as well as The audio signal is used to drive a speaker that is not part of the audio playback device to continue outputting the audio content.

2. The method of claim 1, wherein the audio content is a musical work having a playback duration, wherein the audio playback device stops outputting the musical work at a certain moment during the playback duration, and wherein the audio signal includes the remainder of the musical work that begins at or after the moment during the playback duration.

3. The method according to claim 2, further comprising: Once the remainder of the musical work is output via the speaker, each of the multiple musical works in the playlist is streamed sequentially. Each of the multiple musical works in the playlist is associated with the musical work.

4. The method of claim 3, wherein each of the plurality of musical works in the playlist is related to the musical work in that it belongs to the same genre, the same artist, or is from at least one of the same album.

5. The method of claim 1, further comprising providing a notification including a request to continue the output of the audio content via the speaker, wherein the speaker is driven using the audio signal in response to receiving user authorization.

6. The method of claim 1, further comprising determining that the audio playback device has stopped outputting the audio content based on the sound output level of the loudspeaker being lower than a threshold.

7. The method of claim 1, wherein the audio playback device is part of a vehicle audio system, and wherein the method further comprises determining that the audio playback device has stopped outputting audio content by determining that at least one of the vehicle has stopped or the engine of the vehicle has been turned off.

8. The method of claim 1, wherein the portable media player device is not communicatively coupled to the audio playback device, such that the portable media player device cannot receive data describing or identifying the audio content from the audio playback device.

9. A method performed by an electronic device, the method comprising: Receive microphone signals; Based on the microphone signal, identification information of the audio content being output by the audio playback device is determined, wherein the identification information is determined by acoustic signal analysis of the audio content; Based on the fact that the sound level of the microphone signal is below a threshold, it is determined that the audio playback device has stopped outputting the audio content; as well as In response, the audio content continues to be output via a speaker based on the identification information.

10. The method of claim 9, wherein the microphone signal includes the audio content output by the audio playback device, and wherein the acoustic signal analysis is an acoustic signal analysis of the microphone signal.

11. The method of claim 9, wherein the identification information indicates that the audio content is a radio broadcast program being picked up by the audio playback device, and indicates the source of the audio content as a radio station broadcasting the radio broadcast program.

12. The method of claim 11, further comprising: Receive the location information of the electronic device; as well as The identification information and the location information are used to confirm that the source of the audio content is the radio station.

13. The method of claim 12, wherein continuing to output the audio content includes streaming the radio broadcast program via a computer network using the identification information.

14. The method of claim 9, wherein continuing to output the audio content includes streaming similar audio content being broadcast by a radio station via a computer network.

15. The method of claim 9, wherein continuing to output the audio content comprises: The identification information is used to retrieve the audio signal corresponding to the audio content from the local memory of the electronic device or a remote device; as well as The speaker is driven using the audio signal.

16. The method of claim 15, wherein the audio content has a playback duration, wherein the audio playback device stops outputting the audio content at a certain moment within the playback duration, wherein the retrieved audio signal includes the remainder of the audio content that begins at or after the moment within the playback duration.

17. The method of claim 9, wherein the audio content is a musical work, and wherein the method further comprises: Once the musical work has been completed and output by the audio playback device, each of the multiple musical works in the playlist is streamed sequentially. Each of the multiple musical works in the playlist is associated with the musical work.

18. The method of claim 17, wherein each of the plurality of musical works in the playlist is associated with the musical work in that it belongs to the same genre, the same artist, or is from at least one of the same album.

19. An audio system, comprising: processor; and The memory has instructions that, when executed by the processor, cause the audio system to: Receive microphone signals, the microphone signals including audio content output by an audio playback device via a loudspeaker; Without receiving additional information about the audio content from another device, acoustic signal analysis is performed on the microphone signal to automatically determine the identification information of the audio content; as well as In response to determining that the audio playback device has stopped outputting the audio content, the audio content is continued to be output via the speaker based on the identification information.

20. The audio system of claim 19, wherein performing the acoustic signal analysis comprises: The microphone signal is processed according to a speech recognition algorithm to detect the speech contained therein; as well as Based on the voice, it is determined that the audio content is a radio broadcast program.

21. The audio system of claim 20, wherein determining that the audio content is a radio broadcast program includes determining that the voice includes a radio call sign.

22. The audio system of claim 21, wherein the speaker is part of an electronic device, and wherein the memory further has instructions for performing the following operations: Receive the location information of the electronic device; and The radio station broadcasting the radio program is identified based on the radio call sign and the location information.

23. The audio system of claim 22, wherein continuing to output the audio content includes streaming the radio broadcast program being broadcast by the radio station via a computer network.

24. The audio system of claim 22, wherein continuing to output the audio content includes streaming audio content similar to the radio broadcast program via a computer network.

25. The audio system of claim 19, wherein the audio content is a musical work, and wherein the memory further includes instructions to: Generate a playlist of multiple musical works, each of which belongs to at least one of the same genre, the same artist, or the same album as the original musical work; and Once the musical works have been output, each of the multiple musical works is streamed sequentially.