Music service via video

The integration of video content with audio content is achieved through fingerprint data comparison and user profile-based synchronization, providing a dynamic and personalized user experience.

JP7830594B2Active Publication Date: 2026-03-16GRACENOTE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2026-03-16

AI Technical Summary

Technical Problem

The challenge of integrating video content with audio content, such as music, is complicated by the need to determine which video content to use and how to synchronize it effectively.

Method used

A method and system that identifies video content based on fingerprint data comparison and user profiles, synchronizing it with audio content using modules like reference determination and presentation modules to ensure seamless integration.

Benefits of technology

Enables dynamic and personalized presentation of video content with audio, enhancing user experience through synchronized and relevant video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007830594000001
    Figure 0007830594000001
  • Figure 0007830594000002
    Figure 0007830594000002
  • Figure 0007830594000003
    Figure 0007830594000003
Patent Text Reader

Abstract

To combine presentation of video content with presentation of audio content.SOLUTION: Techniques of providing motion video content along with audio content are disclosed. In some example embodiments, a computer-implemented system is configured to perform operations comprising: receiving primary audio content; determining that at least one reference audio content satisfies a predetermined similarity threshold based on comparison of the primary audio content with the at least one reference audio content; for each one of the at least one reference audio content, identifying motion video content based on the motion video content being stored in association with one of the at least one reference audio content and not stored in association with the primary audio content; and causing the identified motion video content to be displayed on a device concurrently with presentation of the primary audio content on the device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , [Figure 3B] , , , , , , ,

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Patent Application No. 15 / 475,488, filed Mar. 31, 2017, which is incorporated herein by reference in its entirety.

[0002] This application generally relates to the technical field of data processing and, in various embodiments, to methods and systems for providing video content along with audio content.

Background Art

[0003] The presentation of audio content often lacks corresponding video content. Combining the presentation of video content with such presentation of audio content raises many technical challenges, including but not limited to determining which video content to use and how to combine the video content with the audio content.

Summary of the Invention

[0004] Some embodiments of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate like elements.

Brief Description of the Drawings

[0005] [Figure 1] A block diagram representing a network environment suitable for providing video content along with audio content according to some exemplary embodiments. [Figure 2] Comparison of primary audio content with a plurality of reference audio contents according to some exemplary embodiments. [Figure 3A] A conceptual diagram representing synchronization of video content with primary audio content according to some exemplary embodiments. [Figure 3B] A conceptual diagram representing synchronization of video content with primary audio content according to some exemplary embodiments. [Figure 4A] This is a conceptual diagram illustrating the synchronization of different video content with primary audio content, following several embodiments as examples. [Figure 4B] This is a conceptual diagram illustrating the synchronization of different video content with primary audio content, following several embodiments as examples. [Figure 5] This flowchart illustrates a method for providing video content along with audio content, following several examples of embodiments. [Figure 6] This flowchart illustrates a method for simultaneously presenting audio content and displaying video content on a device, following several examples of embodiments. [Figure 7] This is a block diagram representing a mobile device according to some embodiments. [Figure 8] Here are some examples of computer system block diagrams, according to embodiments, that illustrate how the methodology described herein can be implemented. [Modes for carrying out the invention]

[0006] An example method and system for providing video content along with audio content is disclosed. In the following description, many specific details are given for illustrative purposes to provide a complete understanding of the exemplary embodiments. However, it will be apparent to those skilled in the art that these embodiments may be carried out without those specific details.

[0007] In some exemplary embodiments, a method performed by a computer includes receiving primary audio content, determining that at least one reference audio content satisfies a predetermined similarity threshold based on a comparison of the primary audio content with at least one reference audio content, identifying video content for each of the at least one reference audio content based on whether the video content is stored in association with one of the at least one reference audio content and not stored in association with the primary audio content, and displaying the identified video content on the device simultaneously with the presentation of the primary audio content on the device. In some exemplary embodiments, the primary audio content includes music.

[0008] In some exemplary embodiments, the comparison includes comparing the fingerprint data of the primary audio content with the fingerprint data of at least one reference audio content.

[0009] In some exemplary embodiments, the identification of video content is further based on the user profile associated with the device.

[0010] In some exemplary embodiments, displaying identified video content on a device simultaneously with the presentation of primary audio content on the device includes synchronizing data of at least one reference audio content with data of primary audio content, and synchronizing identified video content with primary audio content based on the synchronization of data of at least one reference audio content with data of primary audio content. In some exemplary embodiments, the synchronization of data of at least one reference audio content with data of primary audio content is based on a comparison of fingerprint data of at least one reference audio content with a fingerprint of primary audio content.

[0011] In some exemplary embodiments, at least one reference audio content comprises at least two reference audio contents, one of each of the at least two reference audio contents is stored in association with a different video content, and the identified video content comprises a portion of each of the different video contents.

[0012] The methods or embodiments disclosed herein can be implemented as a computer system having one or more modules (e.g., hardware modules or software modules). Such modules can be executed by one or more processors of the computer system. When the methods or embodiments disclosed herein are executed by one or more processors, they can be embodied as instructions stored in a machine-readable medium that causes one or more processors to execute the instructions.

[0013] Figure 1 is a block diagram representing a network environment 100 suitable for providing video content along with audio content, according to one example embodiment. The network environment 100 includes a content provider 110, one or more devices 130, and one or more data sources 140 (e.g., data source 140-1 to data source 140-N), all of which are interconnected and able to communicate with each other via a network 120. The content provider 110, device(s) 130, and data source(s) 140 may each be implemented in whole or in part in a computer system, as described below with respect to Figure 8.

[0014] Also shown in Figure 1 is User 132. User 132 may be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with device 130), or any appropriate combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). User 132 may not be part of the network environment 100, but may be associated with device 130 and be a user of device 130. For example, device 130 may be a desktop computer, vehicle computer, tablet computer, navigation device, portable media device, or smartphone belonging to User 132.

[0015] Any of the machines, providers, modules, databases, devices, or data sources shown in Figure 1 may be implemented on a computer that has been modified (e.g., configured or programmed) by software to become a special-purpose computer to perform one or more of the functions described herein for that machine, provider, module, database, device, or data source. For example, a computer system capable of implementing one or more of the methodologies described herein is discussed below with respect to Figure 8. As used herein, “database” is a data storage resource that may store data structured as text files, tables, spreadsheets, relational databases (e.g., object relational databases), triple stores, hierarchical data stores, or a suitable combination thereof. Furthermore, any two or more of the databases, devices, or data sources shown in Figure 1 may be combined on a single machine, and any of the functions described herein for any single database, device, or data source may be further divided among multiple databases, devices, or data sources.

[0016] Network 120 may be any network that enables communication between or within machines, databases, and devices. Therefore, Network 120 may be a wired network, a wireless network (e.g., a mobile or cellular network), or a suitable combination thereof. Network 120 may include one or more parts that constitute a private network, a public network (e.g., the Internet), or a suitable combination thereof. Therefore, Network 120 may include one or more parts that incorporate a local area network (LAN), a wide area network (WAN), the Internet, a mobile phone network (e.g., a cellular network), a wired telephone network (e.g., a pre-existing telephone system (POTS) network), a wireless data network (e.g., a WiFi network or WiMAX network), or a suitable combination thereof. One or more parts of Network 190 may communicate information via a transmission medium. As used herein, “transmission medium” should be understood to include any intangible medium having the ability to store, encode, or execute instructions for execution by a machine, and including digital or analog communication signals or other intangible media to facilitate the communication of such software.

[0017] The content provider 110 includes a computer system configured to provide audio and video content to a device such as device 130. In some exemplary embodiments, the content provider 110 includes a combination of one or more of the following: a criteria determination module 112, a video identification module 114, a presentation module 116, and one or more databases 118. In some exemplary embodiments, modules 112, 114, and 116, and database(s) 118 reside on a machine having memory and at least one processor. In some exemplary embodiments, modules 112, 114, and 116, and database(s) 118 reside on the same machine, but in other exemplary embodiments, one or more of modules 112, 114, and 116, and database(s) 118 reside on separate remote machines that communicate with each other over a network such as network 120.

[0018] In some exemplary embodiments, the reference determination module 112 is configured to receive primary audio content. The audio content may include music, such as a recording of a single song. However, it is considered that other types of audio content are also within the scope of this disclosure. In some exemplary embodiments, the reference determination module 112 is configured to identify, or otherwise determine, at least one reference audio content that satisfies a predetermined similarity threshold based on a comparison of the primary audio content with the reference audio content. For example, the reference determination module 112 may search a database 118 for reference audio content that satisfies a predetermined similarity threshold. In addition, or instead, the reference determination module 112 may search one or more external data sources 140 for reference audio content that satisfies a predetermined similarity threshold. The external data sources 140 are separate from the content provider 110 and may include independent data sources.

[0019] In some example embodiments, the comparison of the primary audio content with the reference audio content includes the comparison of the data of the primary audio content with the data of the reference audio content. The data being compared may include fingerprint data that uniquely identifies or characterizes the corresponding audio content. FIG. 2 represents the comparison of the primary audio content with multiple reference audio contents according to some example embodiments. In FIG. 2, the fingerprint data 212 of the primary audio content 210 is compared with the multiple fingerprint data 222 (e.g., fingerprint data 222-1, …, fingerprint data 222-N) of the multiple reference audio contents 220 (e.g., reference audio content 220-1, …, reference audio content 220-N). In some example embodiments, each comparison generates corresponding statistical data indicating the level of similarity between the primary audio content and the reference audio content. One example of such statistical data is the bit error rate. However, other statistical data are also contemplated within the scope of the present disclosure. In some example embodiments, the reference determination module 112 determines whether the statistical data corresponding to the reference audio content 220 meets a predetermined threshold.

[0020] In some example embodiments, the reference determination module 112 uses an exact fingerprint match between the fingerprint data 212 of the primary audio content 210 and the fingerprint data 222 of the reference audio content 220 as a predetermined threshold. For example, the reference determination module 112 searches for multiple reference audio contents 220 to match a version of an audio recording (e.g., compressed or noisy) to an identical version of the audio recording that has not been degraded.

[0021] In some example embodiments, the reference determination module 112 is configured to use a fuzzy fingerprint match between the fingerprint data 212 of the main audio content 210 and the fingerprint data 222 of the reference audio content 220 as a predetermined threshold. For example, the reference determination module 112 may search for a plurality of reference audio contents 220 and match a recorded song (or a live performance or narration of a play, etc.) with different performances or recordings of the same song (or a live performance or narration of a play, etc.).

[0022] In some example embodiments, the reference determination module 112 is configured to use a match between audio characteristics such as harmony, rhythm features, and start of instruments of the main audio content 210 and the reference audio content 220 as a predetermined threshold. For example, the reference determination module 112 may search for a plurality of reference audio contents 220 and match one audio recording with another based simply on a particular level of similarity between the audio characteristics of different audio recordings, such as determining a high level of similarity between the rhythm features of two different songs and matching the two different songs based thereon.

[0023] In some exemplary embodiments, for one or more reference audio contents 220 determined to meet a similarity threshold, the video identification module 114 identifies video content based on whether the video content is stored in association with the reference audio content and not in association with the primary audio content. In some exemplary embodiments, the video identification module 114 is also configured to identify video content based on a profile of user 132 associated with a device 130 on which the combination of primary audio content and identified video content is presented. In some exemplary embodiments, the user profile is stored in a database 118. The user 132 profile may include one or more combinations of the following: a history of audio content listened to by user 132, indications that user 132 prefers a particular type or category of audio content, a history of audio content purchases, a history of video content referenced by user 132, indications that user 132 prefers a particular type or category of video, and demographic information about user 132 (e.g., gender, age, geographical location). Other types of information indicating potential preferences for specific types of audio or video content may also be included in user 132's profile. In scenarios where several different video contents meet the similarity threshold, the video identification module 114 may use user 132's profile to select one or more video contents based on a determination of which video contents are most relevant to user 132.

[0024] In some exemplary embodiments, the presentation module 116 is configured to display video content identified by the video identification module 114 on device 130 simultaneously with the presentation of the main audio content on device 132. In some exemplary embodiments, where the main audio content includes a song, the content provider 110 thus dynamically generates a music video for a song for which the content provider 110 has stored a music video.

[0025] In some exemplary embodiments, the presentation module 116 is configured to synchronize data of a reference audio content with data of a main audio content, and then, based on the synchronization of the reference audio content data with the main audio content data, to synchronize identified video content with the main audio content. In some exemplary embodiments, the synchronization of the reference audio content data with the main audio content data is based on a comparison of the fingerprint data of the reference audio content with the fingerprint data of the main audio content.

[0026] Figures 3A-3B are conceptual diagrams illustrating the synchronization of video content with primary audio content according to embodiments as several examples. In Figure 3A, primary audio content 210 is shown as comprising audio segments 310-1, 310-2, 310-3, and 310-4, and reference audio content 220 is shown as comprising audio segments 320-1, 320-2, 320-3, and 320-4. Reference audio content 220 is also shown as being stored in relation to video content 320, which is shown as comprising video segments 322-1, 322-2, 322-3, and 322-4. Other segmented configurations are also considered to be within the scope of this disclosure. In Figure 3A, the audio segments 310 of primary audio content 210 and the audio segments 320 of reference audio content 220 are aligned along the time domain according to their respective timestamps as a result of the presentation module 116 synchronizing them. Similarly, the video segment 322 of the video content 320 is synchronized with the audio segment 320 of the reference audio content 220 to which it is associated.

[0027] In Figure 3B, the presentation module 116 synchronizes the video segment 322 of the video content 320 with the audio segment 310 of the main audio content 210, along with the synchronization of the audio segment 320 of the reference audio content 220 with the audio segment 320 of the reference audio content 220.

[0028] In some exemplary embodiments, portions from multiple different video content associated with multiple different reference audio content are combined with the main audio content. Figures 4A-4B are conceptual diagrams illustrating the synchronization of different video content with the main audio content according to some exemplary embodiments. In Figure 4A, as in Figure 3A, the main audio content 210 is shown as consisting of audio segments 310-1, 310-2, 310-3, and 310-4, and the reference audio content 220 is shown as consisting of audio segments 320-1, 320-2, 320-3, and 320-4. The reference audio content 220 is also shown as being stored in association with video content 320, which is shown as consisting of video segments 322-1, 322-2, 322-3, and 322-4. The audio segments 310 of the main audio content 210 and the audio segments 320 of the reference audio content 220 are adjusted along the time domain according to their respective timestamps as a result of the presentation module 116 synchronizing them. Similarly, the video segment 322 of the video content 320 is synchronized with the audio segment 320 of the reference audio content 220 to which it is associated.

[0029] In addition, Figure 4A shows that another reference audio content 420 consists of audio segments 420-1, 420-2, 420-3, and 420-4. The reference audio content 420 is also shown to be stored in association with video content 420, which consists of video segments 422-1, 422-2, 422-3, and 422-4. The audio segments 420-1, 420-2, 420-3, and 420-4 of the reference audio content 420, as well as the video segments 422-1, 422-2, 422-3, and 422-4, are coordinated with the audio segments 310-1, 310-2, 310-3, and 310-4 of the main audio content 210.

[0030] Using synchronization, the presentation module 116 generates video content 425 from portions of video content 320 and video content 420. As a result, video segment 322-1 is synchronized with audio segment 310-1, video segment 322-2 is synchronized with audio segment 310-2, video segment 422-3 is synchronized with audio segment 310-3, and video segment 422-4 is synchronized with audio segment 310-4.

[0031] In some exemplary embodiments, the presentation module 116 is configured to synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 based on an exact fingerprint match between the audio segment 310 of the main audio content 210 and the audio segment 320 of the reference audio content 220. For example, the presentation module 116 may synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 based on a match between one version of the audio recording (e.g., compressed or noisy) and an undegraded version of the same audio recording.

[0032] In some exemplary embodiments, the presentation module 116 is configured to synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 based on a fuzzy fingerprint match between the audio segment 310 of the main audio content 210 and the audio segment 320 of the reference audio content 220. For example, the presentation module 116 may synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 based on a match between a recording of a song (or a theatrical performance or narration, etc.) and a different performance or recording of the same song (or a theatrical performance or narration, etc.).

[0033] In some exemplary embodiments, the presentation module 116 is configured to synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 by using a match between the audio characteristics of the audio segment 310 of the main audio content 210 and such audio characteristics of the audio segment 320 of the reference audio content 220, such as chords, rhythmic features, and instrument initiation. For example, the presentation module 116 may synchronize the audio segment 310 of the main audio content 210 with the audio segment 320 of the reference audio content 220 based on a certain level of similarity between the audio characteristics of different audio recordings, such as synchronizing two different songs based on a high level of similarity determination between the rhythmic features of two different songs.

[0034] In some exemplary embodiments, the video identification module 114 and the presentation module 116 are configured to identify different video content to be synchronized and displayed simultaneously with the same main audio content, thereby changing the video experience for the same main audio content from one playback to the next. The change in the video experience from one presentation of the main audio content to the next may be partial, such as by swapping out one video segment or one scene for a different video segment or one scene, while retaining at least one video segment or one scene from one presentation to the next. Alternatively, the change in the video experience from one presentation of the main audio content to the next may be holistic, such as by replacing all the video segments used for the presentation of the main audio content with completely different video segments for subsequent presentations of the main audio content. For example, on a certain date, video content covering an entire live performance of a song may be synchronized and displayed simultaneously with the main audio content, and then, at a later date, video content covering an entire live performance of the same song in a studio (e.g., different from the live performance) may be synchronized and displayed simultaneously with the same main audio content instead of the video content covering the entire live performance of the song. Such changes in the video experience may be based on detected changes in the popularity of the video content (e.g., changes in the total number of daily YouTube® views of the video content), or on detected changes in the preferences or behaviors of users who will be presented with the video content alongside primary audio content (e.g., changes in users' browsing habits for video content on YouTube®), or may be random. It should be considered that other factors may be used to change the video content from one presentation to another.

[0035] Figure 5 is a flowchart illustrating a method 500 for providing video content along with audio content, according to several example embodiments. Method 500 can be implemented by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device), or a combination thereof. In one example embodiment, method 500 is implemented by the content provider 110 of Figure 1, or a combination of one or more of its components or modules.

[0036] In operation 510, the content provider 110 receives primary audio content. In some exemplary embodiments, the primary audio content includes music (e.g., a song). In operation 520, the content provider 110 determines, based on a comparison of the primary audio content with at least one reference audio content, that at least one reference audio content satisfies a predetermined similarity threshold. In some exemplary embodiments, the comparison includes a comparison of the fingerprint data of the primary audio content with the fingerprint data of the at least one reference audio content. In operation 530, for each of the at least one reference audio content, the content provider 110 identifies the video content based on the fact that the video content is stored in association with one of the at least one reference audio content and not stored in association with the primary audio content. In some exemplary embodiments, the identification of the video content is further based on a user profile associated with the device. In operation 540, the content provider 110 displays the identified video content on the device simultaneously with the presentation of the primary audio content on the device. It is considered that any of the other features described in this disclosure may be incorporated into method 500.

[0037] Figure 6 is a flowchart illustrating a method 600 for displaying video content on a device simultaneously with presenting audio content on the device, according to several exemplary embodiments. Method 600 can be implemented by processing logic that may include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device), or a combination thereof. In one exemplary embodiment, method 600 is implemented by the content provider 110 of Figure 1, or a combination of one or more of its components or modules.

[0038] In operation 610, the content provider synchronizes data of at least one reference audio content with data of the main audio content. In operation 620, the content provider 110 synchronizes identified video content with the main audio content based on the synchronization of data of at least one reference audio content with data of the main audio content. In some exemplary embodiments, the synchronization of data of at least one reference audio content with data of the main audio content is based on a comparison of fingerprint data of at least one reference audio content with fingerprint data of the main audio content. It is considered that any of the other features described in this disclosure may be incorporated into method 600.

[0039] Mobile devices as an example Figure 7 is a block diagram representing a mobile device 700 according to an exemplary embodiment. The mobile device 700 may include a processor 702. The processor 702 can be any of several different types of commercially available processors suitable for the mobile device 700 (e.g., an XScale architecture microprocessor, a microprocessor without an Interlocked Pipeline Stages (MIPS) architecture, or another type of processor). Memory 704, such as random access memory (RAM), flash memory, or other types of memory, is typically accessible to the processor 702. Memory 704 can be adapted to store application programs 708, such as mobile location-aware applications, which can provide LBS to the user, together with an operating system (OS) 706. The processor 702 can be coupled either directly or via appropriate intermediary hardware to a display 710, as well as one or more input / output (I / O) devices 712, such as a keypad, touch panel sensor, and microphone. Similarly, in some embodiments, the processor 702 can be coupled to a transceiver 714 that interfaces with an antenna 716. Depending on the nature of the mobile device 700, the transceiver 714 can be configured to both transmit and receive cellular network signals, radio data signals, or other types of signals via the antenna 716. Furthermore, in some configurations, the GPS receiver 718 can also utilize the antenna 716 to receive GPS signals.

[0040] Modules, components, and logic Certain embodiments of logic, or several components, modules, or mechanisms, are described herein. A module may consist of either a software module (e.g., code embodied on a machine-readable medium or in a transmitted signal) or a hardware module. A hardware module is a tangible unit capable of performing a particular operation and may be configured or arranged in a particular manner. In exemplary embodiments, one or more computer systems (e.g., standalone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or a group of processors), may be configured by software (e.g., an application or application portion) as hardware modules that operate to perform a particular operation as described herein.

[0041] In various embodiments, hardware modules can be implemented mechanically or electrically. For example, a hardware module may include dedicated circuitry or logic permanently configured to perform a specific operation (e.g., a special-purpose processor such as a field-programmable gate array (FPGA) or application-specific integrated circuit (ASIC)). A hardware module may also include programmable logic or circuitry temporarily configured by software to perform a specific operation (e.g., such as being incorporated within a general-purpose processor or other programmable processor). It is recognized that the decision to mechanically implement a hardware module in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software), may be derived from cost and time considerations.

[0042] Therefore, it should be understood that the term “hardware module” encompasses tangible entities that are physically constructed, permanently configured (e.g., built into hardware), or temporarily configured (e.g., programmed) to operate in a particular manner and / or perform specific operations as described herein. Considering embodiments in which a hardware module is temporarily configured (e.g., programmed), each hardware module does not need to be configured or instantiated in any one instance within a given time. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor can be configured as each of several different hardware modules at different times. Thus, software can configure the processor to, for example, build a particular hardware module in one instance of time and different hardware modules in different instances of time.

[0043] Hardware modules can provide information to other hardware modules and receive information from other hardware modules. Therefore, the hardware modules described can be considered to be communicatively coupled. When multiple such hardware modules exist simultaneously, communication can be achieved through signal transmissions connecting the hardware modules (e.g., through appropriate circuits and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, through the storage and retrieval of information in a memory structure that multiple hardware modules have access to. For example, one hardware module can perform an operation in a memory device to which it is communicatively coupled and store the output of that operation. Further hardware modules can then access the memory device to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and perform operations on resources (e.g., sets of information).

[0044] Various operations of the exemplary methods described herein can be performed, at least partially, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can be configured as processor implementation modules that operate to perform one or more operations or functions. The modules referred to herein may include processor implementation modules in some exemplary embodiments.

[0045] Similarly, the methods described herein can be implemented at least partially by processors. For example, at least some of the operations of a method can be performed by one or more processors or processor implementation modules. The execution of a particular operation can be distributed among one or more processors, deployed across several machines, as well as within a single machine. In some exemplary embodiments, the processor or processor(s) may be located in a single location (e.g., within a home environment, an office environment, or a server farm), while in other embodiments, the processors may be distributed across several locations.

[0046] One or more processors can also operate to support the execution of related operations in a “cloud computing” environment or as “software as a service” (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of a machine containing a processor), and those operations can be accessed over a network and through one or more appropriate interfaces (e.g., APIs).

[0047] An exemplary embodiment may be implemented in a digital electronic circuit, or in computer hardware, firmware, software, or a combination thereof. An exemplary embodiment may be implemented using a computer program product, such as an information carrier, such as a data processing device, such as a programmable processor, a computer, or a computer program tangibly embodied in a machine-readable medium for execution by or control of the operation of multiple computers.

[0048] Computer programs can be written in any form of programming language, including compiled or interpreted languages, and can be deployed as standalone programs or as modules, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed to run on one computer or multiple computers at one site, or they can be distributed across multiple sites and interconnected by communication networks.

[0049] In an exemplary embodiment, the operation can be performed by one or more programmable processors that execute a computer program to perform a function by operating with respect to input data and generating an output. The operation of the method can also be performed by a special-purpose logic circuit (e.g., FPGA or ASIC), and the device of the exemplary embodiment can be implemented as a special-purpose logic circuit.

[0050] A computing system can include clients and servers. Clients and servers are generally remote to each other and typically interact through a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with each other. In embodiments of deploying a programmable computing system, it is recognized that both hardware and software architectures are worth considering. In particular, it is recognized that the choice of whether to implement certain functionality in persistently configured hardware (e.g., ASICs), temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of persistent and temporarily configured hardware can be a design choice. The following shows hardware (e.g., machines) and software architectures that can be deployed in embodiments as various examples.

[0051] Figure 8 is a block diagram of a computer system 800 in exemplary form, on which instruction 824 can be executed, causing the machine to perform one or more of the methodologies discussed herein, according to an exemplary embodiment. In alternative embodiments, the machine may operate as a standalone device or be connected to other machines (e.g., networked). In a networked deployment, the machine may operate as a server or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, network router, switch or bridge, or any machine capable of executing (sequentially or otherwise) instructions that specify actions to be taken by that machine. Furthermore, although a single machine is represented, the term “machine” should be understood to also include any set of machines that individually or collectively execute a set (or more sets) of instructions to perform one or more of the methodologies discussed herein.

[0052] As an example, computer system 800 includes a processor 802 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), main memory 804, and static memory 806, which communicate with each other via bus 808. Computer system 800 may further include a video display unit 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)). Computer system 800 also includes an alphanumeric input device 812 (e.g., a keyboard), a user interface (UI) navigation (or cursor control) device 814 (e.g., a mouse), a disk drive unit 816, a signal generation device 818 (e.g., a speaker), and a network interface device 820.

[0053] The disk drive unit 816 includes a machine-readable medium 822 that stores one or more sets of data structures and instructions 824 (e.g., software) that embody or are utilized by any one or more methodologies or functions described herein. The instructions 824 may also reside entirely or at least partially in the main memory 804 and / or in the processor 802 during their execution by the computer system 800, which also constitutes the machine-readable medium, the main memory 804, and the processor 802. The instructions 824 may also reside entirely or at least partially in static memory 806.

[0054] While the machine-readable medium 822 is shown as a single medium in an exemplary embodiment, the term “machine-readable medium” can include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instructions 824 or data structures. The term “machine-readable medium” should also be understood to include any tangible medium having the ability to store, encode, or carry instructions for execution by a machine, and to cause a machine to execute one or more of the methodologies of this embodiment, or to store, encode, or carry data structures utilized by or associated with such instructions. Accordingly, the term “machine-readable medium” should be understood to include, but not be limited to, solid-state memory, as well as optical and magnetic media. Specific embodiments of machine-readable media include, as an example, semiconductor memory devices (e.g., erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), and flash memory devices), magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and non-volatile memory including compact disk read-only memory (CD-ROM) and digital multipurpose disk (or digital video disk) read-only memory (DVD-ROM) disks.

[0055] Instruction 824 may further be transmitted or received through a communication network 826 using a transmission medium. Instruction 824 may be transmitted using a network interface device 820 and one of several well-known transport protocols (e.g., HTTP). Embodiments of the communication network include LANs, WANs, the Internet, cellular networks, POTS networks, and wireless data networks (e.g., WiFi and WiMAX networks). The term “transmission medium” should be understood to include any intangible medium having the ability to store, encode, or carry instructions for execution by a machine, and includes digital or analog communication signals, or other intangible mediums that facilitate communication of such software.

[0056] While embodiments have been described with reference to specific examples, it is evident that various modifications and changes can be made to those embodiments without departing from the broader spirit and scope of the disclosure. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. The accompanying drawings, forming part thereof, illustrate, not restrictively, specific embodiments in which the subject matter can be carried out. The embodiments shown are described in sufficient detail to enable those skilled in the art to carry out the teachings disclosed herein. Other embodiments can be utilized and derived from them so that structural and logical substitutions and changes can be made without departing from the scope of the disclosure. Therefore, this detailed description is not to be considered restrictive, and the scope of the various embodiments, along with the same full scope as granted by the accompanying claims, is defined only by such claims.

[0057] While specific embodiments have been represented and described herein, it should be recognized that any arrangement calculated to achieve the same objective can be substituted for any particular embodiment shown. This disclosure is intended to cover any and all adaptations or variations of any of the various embodiments. Combinations of the embodiments described herein, and other embodiments not specifically described herein, will be apparent to those skilled in the art by referring to the above description.

Claims

1. A tangible computer-readable storage medium containing instructions, When the aforementioned instruction is executed, it will cause one or more processors to: Identifying at least one video content based on the fact that at least one video content is stored in association with a reference audio content that satisfies a predetermined similarity threshold with the primary audio content, and is not stored in association with the primary audio content; To detect changes in the popularity of at least one identified video content, In response to a change in the popularity of the at least one identified video content, the at least one identified video content is displayed on the device simultaneously with the presentation of the main audio content on the device. A tangible, computer-readable storage medium that enables the execution of a set of operations including [specific actions].

2. The main audio content includes music, as described in claim 1, in the tangible computer-readable storage medium.

3. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Simultaneously with the presentation of the first portion of the main audio content on the device, a portion of at least one identified video content is displayed on the device. A tangible computer-readable storage medium according to claim 1, further comprising:

4. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: To detect changes in the user behavior associated with the device, Based on the detected changes in the user's behavior associated with the device, the device will display the main audio content and, at the same time, the device will display the at least one identified video content. A tangible computer-readable storage medium according to claim 1, further comprising:

5. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: To detect changes in the user's preferences associated with the device, Based on the detected changes in the user's preferences associated with the device, the presentation of the primary audio content on the device is accompanied by the display of at least one identified video content on the device. A tangible computer-readable storage medium according to claim 1, further comprising:

6. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Synchronizing the data of the aforementioned reference audio content with the data of the aforementioned main audio content, Synchronizing the at least one identified video content with the main audio content based on the synchronization of the data of the reference audio content with the data of the main audio content, A tangible computer-readable storage medium according to claim 1, further comprising:

7. The aforementioned set of operations is Based on a comparison of the main audio content with the aforementioned reference audio content, it is determined that the reference audio content satisfies the predetermined similarity threshold. A tangible computer-readable storage medium according to claim 1, further comprising:

8. The tangible computer-readable storage medium according to claim 1, wherein the at least one identified video content includes a first identified video content and a second identified video content.

9. One or more processors, A tangible computer-readable storage medium containing instructions, A system equipped with, When the aforementioned instruction is executed, it will cause one or more processors to: Identifying at least one video content based on the fact that at least one video content is stored in association with a reference audio content that satisfies a predetermined similarity threshold with the primary audio content, and is not stored in association with the primary audio content; To detect changes in the popularity of at least one identified video content, In response to a change in the popularity of the at least one identified video content, the at least one identified video content is displayed on the device simultaneously with the presentation of the main audio content on the device. A system that executes a set of actions that include this.

10. The system according to claim 9, wherein the main audio content includes music.

11. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Simultaneously with the presentation of the first portion of the main audio content on the device, a portion of at least one identified video content is displayed on the device. The system according to claim 9, further comprising:

12. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: To detect changes in the user behavior associated with the device, Based on the detected changes in the user's behavior associated with the device, the device will display the main audio content and, at the same time, the device will display the at least one identified video content. The system according to claim 9, further comprising:

13. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: To detect changes in the user's preferences associated with the device, Based on the detected changes in the user's preferences associated with the device, the presentation of the primary audio content on the device is accompanied by the display of at least one identified video content on the device. The system according to claim 9, further comprising:

14. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Synchronizing the data of the aforementioned reference audio content with the data of the aforementioned main audio content, Synchronizing the at least one identified video content with the main audio content based on the synchronization of the data of the reference audio content with the data of the main audio content, The system according to claim 9, further comprising:

15. The aforementioned set of operations is Based on a comparison of the main audio content with the aforementioned reference audio content, it is determined that the reference audio content satisfies the predetermined similarity threshold. The system according to claim 9, further comprising:

16. The system according to claim 9, wherein the at least one identified video content includes a first identified video content and a second identified video content.

17. A method performed by at least one hardware processor, Identifying at least one video content based on the fact that at least one video content is stored in association with a reference audio content that satisfies a predetermined similarity threshold with the primary audio content, and is not stored in association with the primary audio content; To detect changes in the popularity of at least one identified video content, In response to a change in the popularity of the at least one identified video content, the at least one identified video content is displayed on the device simultaneously with the presentation of the main audio content on the device. Methods that include...

18. The method according to claim 17, wherein the main audio content includes music.

19. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Simultaneously with the presentation of the first portion of the main audio content on the device, a portion of at least one identified video content is displayed on the device. The method according to claim 17, further comprising:

20. Displaying the at least one identified video content on the device simultaneously with the presentation of the main audio content on the device is: Synchronizing the data of the aforementioned reference audio content with the data of the aforementioned main audio content, Synchronizing the at least one identified video content with the main audio content based on the synchronization of the data of the reference audio content with the data of the main audio content, The method according to claim 17, further comprising:

Citation Information

Patent Citations

  • Intelligent identification of multimedia content for synchronization

    US20060163358A1

  • Video synchronization based on an audio cue

    US20150373231A1

  • Audio broadcasting content synchronization system

    US20160337059A1