Generating a video presentation with audio
By synchronizing video transitions with musical beats, the method addresses the challenge of generating synchronized audio-video presentations, resulting in an enhanced user experience.
Patent Information
- Application Number
- JP2025139180
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-03-30
- Filing Date
- 2025-08-22
- Publication Date
- 2026-01-21
AI Technical Summary
Existing systems lack efficient methods for generating synchronized audio-video presentations that align transitions with musical beats, leading to suboptimal user experience.
The method involves selecting an audio track based on tempo, genre, or mood, and segmenting video sequences to align transitions with musical beats, using algorithms to synchronize video segments with audio tracks, and generating a synchronized audio-video presentation.
This approach ensures seamless synchronization of video transitions with audio beats, enhancing the user experience by providing a more engaging and coherent audio-visual experience.
Smart Images

Figure 2026009873000001_ABST
Abstract
Description
[Technical Field]
[0001] The subject matter disclosed herein generally relates to audio / video presentations. Specifically, the present disclosure is directed to systems and methods for generating video presentations accompanied by audio.
[0002] Certain embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. [Brief explanation of the drawings]
[0003] [Figure 1] 1 is a block diagram illustrating a network environment suitable for generating a video presentation with audio, according to some example embodiments. [Figure 2] 1 is a block diagram illustrating a database suitable for generating a video presentation with audio according to some example embodiments. [Figure 3] 1 is a block diagram illustrating segmented and non-segmented video data according to some example embodiments suitable for generating a video presentation with audio; [Figure 4] FIG. 1 is a block diagram illustrating alignment of an audio track with a video segment in a video presentation with audio, according to some example embodiments. [Figure 5] 1 is a flowchart illustrating a process in some example embodiments for generating a video presentation with audio. [Figure 6] 1 is a flowchart illustrating a process in some example embodiments for generating a video presentation with audio. [Figure 7] 1 is a flowchart illustrating a process in some example embodiments for generating a video presentation with audio. [Figure 8] 1 is a block diagram illustrating a user interface in some example embodiments for generating a video presentation with audio. [Figure 9]FIG. 1 is a block diagram illustrating components of a machine capable of reading instructions from a machine-readable medium and performing any one or more of the methodologies discussed herein, according to some example embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0004] An example method and system for generating a video presentation with audio is described. An audio track is selected explicitly or implicitly. An audio track may be explicitly selected by a user selecting an audio track from a set of available audio tracks. An audio track may also be implicitly selected by automatically selecting an audio track from a set of audio tracks based on the mood of the audio track, the genre of the audio track, the tempo of the audio track, or any suitable combination thereof.
[0005] The video presentation with the audio track is generated from one or more video sequences. The video sequences may be explicitly selected by a user or may be selected from a database of video sequences using search criteria. In some example embodiments, the video sequences are divided into video segments corresponding to the breaks between frames. The video segments are concatenated to form a video presentation to which the audio track is added.
[0006] In some example embodiments, only video segments having durations that correspond to an integral number of musical beats in the audio track are used to form the video presentation, and in these example embodiments, transitions between video segments in a video presentation with an audio track are aligned with musical beats.
[0007] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of example embodiments. However, it will be apparent to those skilled in the art that the present subject matter may be practiced without these specific details.
[0008] 1 is a network diagram illustrating a network environment 100 suitable for generating video presentations with audio, according to some example embodiments. Network environment 100 may include a server system 110 and a client device 150 or 160 connected by a network 140. Server system 110 has a video database 120 and an audio database 130.
[0009] Client device 150 or 160 is any device capable of receiving and presenting streams of media content (e.g., a television, a second set-top box, a laptop or other personal computer (PC), a tablet or other mobile device, a digital video recorder (DVR), or a gaming device). Client device 150 or 160 may also include a display or other user interface configured to display the generated video presentation. The display may be a flat-panel screen, a plasma screen, a light-emitting diode (LED) screen, a cathode ray tube (CRT), a liquid crystal display (LCD), a projector, or any suitable combination thereof. A user of client device 150 or 160 may interact with the client device via application interface 170 or browser interface 180.
[0010] Network 140 may be any network that enables communication between devices, such as a wired network, a wireless network (e.g., a mobile network), etc. Network 140 may include one or more portions that make up a private network (e.g., a cable or satellite television network), a public network (e.g., a wireless broadcast channel or the Internet), etc.
[0011] In some example implementations, client device 150 or 160 sends a request to server system 110 over network 140. The request identifies a search query for video content and a musical genre. Based on the musical genre, server system 110 identifies a particular audio track from audio database 130. Based on the search query for video content, server system 110 identifies one or more video sequences from video database 120. Using methods disclosed herein, server system 110 generates a video presentation having the identified audio track and video segments from the one or more identified video sequences. Server system 110 may transmit the generated video presentation to client device 150 or 160 for presentation on a display device associated with the client device.
[0012] 1, server system 110 has video database 120 and audio database 130. In some example embodiments, video database 120, audio database 130, or both, are implemented on separate computer systems accessible by server system 110 (e.g., via network 140 or another network).
[0013] Any of the machines, databases, or devices shown in FIG. 1 may be implemented on a general-purpose computer that has been modified (e.g., configured or programmed) with software to be a special-purpose computer that performs the functions described herein with respect to that machine. For example, a computer system capable of implementing any one or more of the methodologies described herein is discussed below with respect to FIG. 9. As used herein, a "database" is a data storage resource and may store data structured as text files, tables, spreadsheets, relational databases, document stores, key-value stores, triple stores, or any suitable combination thereof. Furthermore, any two or more of the machines shown in FIG. 1 may be combined into a single machine, and the functionality described herein with respect to any single machine may be further divided among multiple machines.
[0014] Furthermore, any of the modules, systems, and / or databases may be located on any of the machines, databases, or devices shown in Figure 1. For example, client device 150 may include video database 120 and audio database 130, among other configurations, and transmit identified video and audio data to server system 100.
[0015] 2 is a block diagram illustrating a database 200 suitable for generating video presentations with audio, according to some example embodiments. Database diagram 200 includes a video data table 210 and an audio data table 240. Video data table 210 uses fields 220 to provide title, keywords, creator, and data for each row in the table (e.g., rows 230A-230D). The video data may be in a variety of formats, such as Moving Picture Experts Group (MPEG)-4 Part 14 (MP4), Audio Video Interleaved (AVI), or QuickTime (QT).
[0016] Audio data table 240 uses fields 250 to provide title, genre, tempo, and data for each row in the table (e.g., rows 260A-260D). The audio data may be in a variety of formats, such as MP3, Windows Media Audio (WMA), Advanced Audio Coding (AAC), or Windows Wave (WAV).
[0017] FIG. 3 is a block diagram illustrating segmented and non-segmented video data suitable for generating a video presentation with audio, according to some example embodiments. Non-segmented video data 310 is shown to have a duration of 1 minute 24 seconds. Segmented video data 320 is divided into nine segments of varying individual durations, but still has the same video content with the same total duration of 1 minute 24 seconds. In some example embodiments, segments of video data are identified based on differences in a series of frames of non-segmented video data. For example, the distance measure in a series of frames may be compared to a predetermined threshold. If the distance measure exceeds the threshold, the series of frames may be identified as being part of a different segment. One example distance measure is the sum of the absolute values of the differences between corresponding pixels in RGB space. Illustratively, in a 1080x1920 high-definition frame, the difference in RGB values between each pair of corresponding pixels (out of 2,073,600 pixels) is determined, its absolute value is taken, and the 2,073,600 resulting values are summed. If the distance is 0, the two frames are exactly the same.
[0018] 4 is a block diagram 400 illustrating the alignment of an audio track and video segments in a video presentation with audio, according to some example embodiments. Block diagram 400 includes an audio track 410, beats 420, and video segments 430A, 430B, and 430C. Beats 420 refer to the instants at which beats occur in audio track 410. For example, if the music in audio track 410 has a tempo of 120 BPM, beats 420 are spaced at 0.5 second intervals. Video segments 430A-430C are aligned with beats 420. Thus, the transition between video segment 430A and video segment 430B occurs at one beat. Video segments 430A-430C may be taken from different video sequences (e.g., from video data table 210) or from a single video sequence. Furthermore, the video segments 430A-430C may be aligned with the audio track 410 in the same order as the video segments are presented in the video sequence from which they originate (eg, the video sequence of FIG. 3), or in a different order.
[0019] In some example embodiments, events other than scene transitions are aligned with beats 420 of the audio track 410. For example, in editing a boxing knockout, each of the video segments 430A-430C may be aligned with the audio track 410 so that the knockout blow lands on the beat.
[0020] The beats 420 may also refer to a subset of the beats in the audio track 410. For example, the beats 420 may be limited to the downbeats or downbeats of the music. The downbeats may be detected by detecting the strength or energy of the vocals in each beat and identifying the beats with the highest energy. For example, in music using a 4 / 4 time signature, one or two of each group of four beats may have higher energy than the other beats. Thus, the beats 420 used for alignment may be limited to one or two of each group of four beats.
[0021] In some example embodiments, transition points in the audio track 410 may be identified by audio signals other than beats 420. For example, an audio track that includes a recording of a running horse instead of music might have transition points identified by the beats of the horse's hooves slapping. As another example, an audio track that includes audio snippets from a movie or television program might have transition points identified by audio energy above a threshold, such as a person shouting, a gunshot, a vehicle approaching the microphone, or any suitable combination thereof.
[0022] 5 is a flow chart illustrating a process 500 in some example embodiments for generating a video presentation with audio. By way of example and not limitation, the steps of process 500 are described as being implemented by the system and device of FIG. 1 using database illustration 200.
[0023] In step 510, server system 110 accesses music tracks having a particular tempo. For example, the music tracks in column 260A may be accessed from audio data table 240. In some exemplary embodiments, client device 150 or 160 presents a user interface to the user via application interface 170 or browser interface 180. The presented user interface includes options that allow the user to select a particular tempo (e.g., a text field for entering a numeric tempo, a drop-down list of predefined tempos, a combo box with a text field and drop-down list, or any suitable combination thereof). Client device 150 or 160 transmits the received tempo to server system 110, which selects the accessed music track based on the tempo. For example, audio data table 240 of audio database 130 may be queried to identify columns having the selected tempo (or within a predetermined range of the selected tempo, e.g., within 5 BPM of the selected tempo).
[0024] In another example embodiment, the user interface includes an option that allows the user to select a genre. The client device transmits the received genre to the server system 110, which selects the accessed music track based on this genre. For example, a query may be made to the audio data table 240 of the audio database 130 to identify a column having the selected genre. Additionally or alternatively, the user may select a mood for selecting an audio track. For example, the audio data table 240 may be expanded to include one or more moods for each song and columns that match the user-selected moods used in step 510. In some example embodiments, the mood of an audio track is determined based on tempo (e.g., slow for sadness, fast for anger, medium for joy, etc.), key (e.g., music in a major key is joyful and music in a minor key is sad), instrument (e.g., bass is melancholic and piccolo is cheerful), keywords (e.g., joy, sadness, anger, or any suitable combination thereof), or any suitable combination thereof.
[0025] In step 520, the server system 110 accesses a video track having a plurality of video segments. For example, the video sequence in column 230A has video segments as shown in segmented video data 320 and may be accessed from video data table 210. The video sequence may be selected by a user (e.g., from a list of available video sequences) or automatically selected. For example, a video track having a mood matching the mood of the audio track may be automatically selected. In some example embodiments, the mood of the video track is determined based on facial recognition (e.g., smiling faces are happy, crying faces are sad, and serious faces are melancholy), color (e.g., bright colors are joyful, dull complexions are sad), recognized objects (e.g., rain is sad, weapons are aggressive, and toys are happy), or any suitable combination thereof.
[0026] In some example embodiments, the video tracks to be accessed are selected by the server system 110 based on the tempo and keywords associated with the video tracks in the video data table 210. For example, a video track associated with the keyword "hockey" may be likely to consist of many short video segments, while a video track associated with the keyword "soccer" may be likely to consist of longer video segments. Thus, a video track associated with the keyword "hockey" may be selected if the tempo is fast (e.g., above 110 BPM), and a video track associated with the keyword "soccer" may be selected if the tempo is slow (e.g., below 80 BPM).
[0027] In step 530, based on the tempo of the music track and the duration of the first video segment of the plurality of video segments, the server system 110 adds the first video segment to the set of video segments. For example, one or more video segments of the video sequence having a duration that is an integer multiple of the beat time of the music track may be identified and added to the set of video segments that can be synchronized to the music track. To illustrate, if the tempo of the music track is 120 BPM, the beat time of the music track is 0.5 seconds, and video segments whose durations are integer multiples of 0.5 seconds are identified as those that can be played in sync with the music track, with transitions between video segments synchronized with the beats of the music.
[0028] In some example embodiments, video segments that are within a predetermined number of frames that are an integer multiple of the beat duration are adjusted to align with the beat and added to a set of video segments in step 530. For example, if the video frame rate is 30 frames per second and the beat duration is 0.5 seconds, or 15 frames, a video segment that is 46 frames long is the only frame that is too long to align. Removing the first or last frame of the video segment generates an aligned video segment that can be used in step 540. Similarly, a video segment that is 44 frames long is the only frame that is too short to align. Repeating the first or last frame of the video segment generates an aligned video segment.
[0029] Server system 110 generates an audio / video sequence having a set of video segments and an audio track in step 540. For example, the audio / video sequence of Figure 4 includes three video segments 430A-430C that can be played while audio track 410 is played, with the transitions between video segments 430A-430C aligned with the beats of audio track 410. The generated audio / video sequence can be stored in video database 120 for later access, transmitted to client device 150 or 160 for playback to a user, or both.
[0030] In some example embodiments, one or more portions of an audio track are used in place of the entire audio track. For example, an audio track may be divided into a chorus and several verses. An audio / video sequence may be prepared using the chorus, a subset of the verses, or any suitable combination thereof. The selection of the portion may be based on the desired length of the audio / video sequence. For example, a three-minute song may be used to generate a one-minute audio / video sequence by selecting a one-minute portion of the song. The selected minute may be the first minute of the song, the last minute of the song, the opening minute at the beginning of the first chorus, one or more repeats of the chorus, one or more verses without the chorus, or another combination of verses and chorus.
[0031] In some example embodiments, multiple audio tracks are used in place of a single audio track. For example, a user may request a five-minute video containing punk music. Multiple songs in the punk genre may be accessed from audio data table 240, each of which is less than five minutes long. Two or more punk tracks of less than five minutes may be concatenated to create a five-minute audio track. The tracks to be concatenated may be selected based on matching tempos. For example, two 120 BPM songs may be selected instead of one 120 BPM song and another 116 BPM song. Alternatively, the tempos of one or more songs may be adjusted to match. For example, a 120 BPM song may be slowed down to 118 BPM, and a 116 BPM song may be sped up to 118 BPM. Either of these methods avoids the possibility of the tempo of an audio / video sequence changing midway.
[0032] 6 is a flow chart illustrating a process 600 in some example embodiments for generating a video presentation with audio. By way of example and not limitation, the steps of process 600 are described as being performed by the system and device of FIG. 1 using database illustration 200.
[0033] In step 610, the server system 110 accesses a music track having a particular tempo. For example, music track 260A may be accessed from the audio data table 240.
[0034] The server system 110 accesses a video track having multiple video segments in step 620. For example, the video sequence in column 230A may have video segments as shown in segmented video data 320 and may be accessed from video data table 210.
[0035] Based on the tempo of the music track and the duration of a particular video segment of the plurality of video segments, the server system 110 adds the video segment to a set of video segments in step 630. For example, video segments of a video sequence having durations that are integer multiples of the beat duration of the music track may be identified and added to a set of video segments that can be synchronized to the music track.
[0036] In step 640, the server system 110 determines whether the total duration of the set of video segments equals or exceeds the duration of the music track. For example, if the music track is one minute long, and only one video segment is added to the set of video segments, and that video segment is 30 seconds long, step 640 would determine that the total duration of 30 seconds is less than the duration of the music track. If the total duration does not equal or exceed the duration of the music track, process 600 repeats steps 620-640, adding another video segment to the set of video segments and repeating the duration check. If the total duration of the set of video segments matches or exceeds the duration of the music track, process 600 continues with step 650.
[0037] In an alternative embodiment, the comparison in step 640 is not with the duration of the music track, but with another duration. For example, a user may select a particular duration for the audio / video sequence. The duration may be shorter than the duration of the music track, in which case the music track may be truncated to fit the selected duration. The user may also select a duration longer than the duration of the music track, in which case the music track may be repeated to reach the selected duration, or additional music tracks of the same tempo may be retrieved from the audio data table 240 and added to the first music track.
[0038] In step 650, server system 110 generates an audio / video sequence having a set of music segments and a video track. For example, the audio / video sequence of FIG. 4 includes three video segments 430A-430C that can be played while audio track 410 is played, with the transitions between video segments 430A-430C aligned with the beats of audio track 410. The generated audio / video sequence may be stored in video database 120 for later access, transmitted to client device 150 or 160 for playback to a user, or both. In some example embodiments, if the total duration of the set of video segments exceeds the duration of the music track, one video segment (e.g., the last video segment) is truncated to fit the duration.
[0039] 7 is a flow chart illustrating a process 700 in some example embodiments for generating a video presentation with audio. By way of example and not limitation, the steps of process 700 are described as being performed by the system and device of FIG. 1 using database illustration 200.
[0040] In step 710, server system 110 accesses video sequences. For example, server system 110 may provide a web page rendered in browser interface 180 of client device 160. Using this web page, a user enters one or more keywords to identify a desired video sequence to be used for audio / video presentation. In this example, server system 110 accesses the video sequence in column 230A from video data table 210 based on a match between the user-provided keywords and the keywords stored in column 230A.
[0041] The server system 110 identifies video segments in the video sequence based on differences in the frames of the video sequence at step 720. For example, a distance measure may be calculated for each pair of frames of the sequence. If the distance measure exceeds a threshold, the pair of frames of the sequence may be determined to be in separate segments. One example distance measure is the sum of the absolute values of the differences in color values of corresponding pixels in the two frames. Thus, two identical frames would have a distance measure of zero.
[0042] In step 730, the multiple identified video segments are used in process 500 or process 600 (e.g., in step 520 or step 620) to generate an audio / video sequence having one or more of the identified video segments and a music track.
[0043] 8 is a block diagram illustrating a user interface 800 in some example embodiments for generating a video presentation with audio. User interface 800 includes a sporting event selector 810, a video style selector 820, and a video playback area 830. User interface 800 may be presented to a user by application interface 170 or browser interface 180.
[0044] A user may select a particular sport by operating a sporting event selector 810. For example, a drop-down menu may be presented that allows the user to select from a set of pre-defined options (e.g., football, hockey, or basketball). Similarly, a user may select a particular video style by operating a video style selector 820. A video style may correspond to a musical genre.
[0045] In response to receiving the selected sport and video style, client device 150 or 160 may transmit the selection to server system 110. Based on the selection, server system 110 identifies audio and video data from audio database 130 and video database 120 to use in performing one or more of processes 500, 600, and 700. After generating the video presentation with audio (e.g., via process 500 or 600), server system 110 may transmit the generated video presentation over network 140 to client device 150 or 160 for display in video playback area 830. Client device 150 or 160 plays the received video presentation for the user in video playback area 830.
[0046] According to various example embodiments, one or more of the methodologies described herein can facilitate generating a video presentation with audio. Accordingly, one or more of the methodologies described herein may obviate the need for certain efforts or resources that would otherwise be associated with generating a video presentation with audio. By utilizing one or more of the methodologies described herein, computing resources used by one or more machines, databases, or devices (e.g., in network environment 100) may be reduced. Examples of such computing resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, and cooling capacity.
[0047] 9 is a block diagram illustrating components of a machine 900, according to some example embodiments, capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium, a computer-readable storage medium, or any suitable combination thereof) and performing, in whole or in part, any one or more of the methodologies discussed herein. Specifically, FIG. 9 illustrates a schematic representation of the machine 900 in the form of an example computer system, within which instructions 924 (e.g., software, programs, applications, applets, apps, or other executable code) that cause the machine 900 to perform, in whole or in part, any one or more of the methodologies discussed herein. In alternative embodiments, the machine 900 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked arrangement, the machine 900 may operate in the capacity of a server machine or a client machine in a server-client network environment, or may operate as a peer machine in a distributed (e.g., peer-to-peer) network environment. The machine 900 may be a server computer, a client computer, a PC, a tablet computer, a laptop computer, a netbook, a set-top box (STB), a smart TV, a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 924 that specify actions to be taken by the machine. Additionally, while only a single machine is illustrated, the term "machine" should also be utilized to include a collection of machines that individually or collectively execute instructions 924 to perform all or a portion of any one or more of the methodologies discussed herein.
[0048] The machine 900 includes a processor 902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), or any combination thereof), a main memory 904, and a static memory 906, which are configured to communicate with each other via a bus 908. The machine 900 may further include a graphics display 910 (e.g., a plasma display panel (PDP), an LED display, an LCD, a projector, or a CRT). The machine 900 may also include an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing device), a storage unit 916, one or more GPUs 918, and a network interface device 920.
[0049] Storage unit 916 includes machine-readable medium 922 on which are stored instructions 924 that embody any one or more of the methodologies or functions described herein. The instructions 924 may also reside completely, at least partially, within main memory 904, within processor 902 (e.g., within a processor's cache memory), or both during execution by machine 900. Thus, main memory 904 and processor 902 may be considered machine-readable media. The instructions 924 may be transmitted or received over a network 926 (e.g., network 140 of FIG. 1 ) via network interface device 920.
[0050] As used herein, the term “memory” refers to a machine-readable medium capable of temporarily or permanently storing data and may be used to include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, and cache memory. While machine-readable medium 922 is shown in one example embodiment to be a single medium, the term “machine-readable medium” should be used to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) capable of storing instructions. The term “machine-readable medium” should also be used to include any medium, or combination of media, capable of storing instructions for execution by a machine (e.g., machine 900), such that, when executed by one or more processors (e.g., processor 902) of the machine, the instructions cause the machine to perform any one or more of the methodologies described herein. Thus, “machine-readable medium” refers to a single storage device or device, as well as a “cloud-based” storage system or storage network including multiple storage devices or devices. The term "machine-readable medium" should therefore be utilized to include, but is not limited to, one or more data storage locations in the form of solid-state memory, optical media, magnetic media, or any suitable combination thereof. The term "non-transitory machine-readable medium" refers to machine-readable media and excludes the signals themselves.
[0051] Throughout this specification, multiple examples may implement a component, operation, or structure that is described as a single example. While individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously, and nothing dictates that the operations be performed in the order shown. Structures and functionality presented as separate components in an example configuration may also be implemented as a combined structure or component. Similarly, structure and functionality presented as a single component may also be implemented as a separate component. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this specification.
[0052] Certain embodiments are described herein as including logic, or several components, modules, or mechanisms. A module may constitute a hardware module. A "hardware module" is a tangible unit capable of performing certain operations, and may be configured or arranged in a particular physical manner. In various exemplary embodiments, one or more computer systems (e.g., a standalone computer system, a client computer system, or a server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or portion of an application) as hardware modules that operate to perform certain operations as described herein.
[0053] In some embodiments, a hardware module may be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware module may include dedicated circuitry or logic that is permanently configured to perform specific operations. For example, a hardware module may be a special-purpose processor such as an FPGA or ASIC. A hardware module may also include programmable logic or circuitry that is temporarily configured by software to perform specific operations. For example, a hardware module may include software encapsulated within a general-purpose processor or other programmable processor. It should be understood that the decision to implement a hardware module mechanically, in dedicated permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.
[0054] Thus, the term "hardware module" should be understood to encompass tangible elements, i.e., permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) elements that are physically constructed to operate in a particular manner or to perform particular operations described herein. As used herein, "hardware-implemented module" refers to a hardware module. Considering embodiments in which the hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated together at any one instance of time. For example, if a hardware module includes a general-purpose processor configured by software to be a special-purpose processor, the general-purpose processor may be configured as different special-purpose processors (e.g., including different hardware modules) at different times. Software may thus configure the processor, for example, to configure a particular hardware module at one instance of time and a different hardware module at a different instance of time.
[0055] Hardware modules can provide information to and receive information from other hardware modules. Accordingly, the described hardware modules may be considered to be communicatively coupled. When multiple hardware modules are simultaneously present, communication may be achieved through signal transmission (e.g., over appropriate circuits and buses) between two or more of the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communication between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform a particular operation and store the output of that operation in a communicatively coupled memory device. Another hardware module may then, at a later time, access this memory device to retrieve and process the stored output. Hardware modules may also initiate communication with input or output devices and operate on particular resources (e.g., chunks of information).
[0056] Various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions described herein. As used herein, a "processor-implemented module" refers to a hardware module that is implemented using one or more processors.
[0057] Similarly, the methods described herein may be at least partially implemented on a processor, which is an example of hardware. For example, at least some of the operations of the methods may be performed by one or more processors or processor-implemented modules. Furthermore, one or more processors may also operate to support execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as an example of machines including processors), and these operations may be accessible via a network (e.g., the Internet) and one or more appropriate interfaces (e.g., application program interfaces (APIs)).
[0058] Execution of certain of the operations may reside within a single machine or may be distributed among one or more processors deployed across several machines. In some example embodiments, one or more processors or processor-implemented modules may be located in a single geographic location (e.g., in a home environment, in a work environment, or in a server farm). In other example embodiments, one or more processors or processor-implemented modules may be distributed across several geographic locations.
[0059] Portions of the subject matter discussed herein may be presented in terms of algorithms or symbolic representations of operations on data stored as bits or binary digital signals within a machine memory (e.g., computer memory). Such algorithms or symbolic representations are examples of techniques used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art. As used herein, an "algorithm" is a self-consistent sequence of operations or similar processes leading to a desired result. In this context, algorithms and operations involve physical manipulations of physical quantities. Typically, though not necessarily, such quantities take the form of electrical, magnetic, or optical signals that can be stored, accessed, transmitted, combined, compared, or otherwise manipulated by a machine. It is sometimes convenient, primarily for reasons of common usage, to refer to such signals using terms such as "data," "content," "bits," "values," "elements," "symbols," "characters," "terms," "numbers," "digits," and the like. However, these terms are merely convenient labels and should be associated with the appropriate physical quantities.
[0060] Unless specifically stated otherwise, discussions herein using words such as "processing," "operating," "calculating," "determining," "presenting," and "displaying" may refer to machine (e.g., computer) acts or processes that manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, nonvolatile memory, or any suitable combination thereof), registers, or other machine components that receive, store, transmit, or display information. Furthermore, unless specifically stated otherwise, the terms "a" or "an" are used herein to include one or more instances, as is common in patent documents. Finally, as used herein, the conjunction "or" refers to a non-exclusive "or" unless specifically stated otherwise.
Claims
1. a memory for storing instructions; one or more databases storing a plurality of music tracks and a plurality of video sequences; one or more processors configured by the instructions to perform operations; The operation is accessing a music track having a tempo from said one or more databases; accessing a video sequence from the one or more databases, the video sequence having a plurality of video segments; adding a first video segment of the plurality of video segments to a set of video segments based on the tempo of the music track and a duration of the first video segment; generating an audio / video sequence comprising said set of video segments and said music track; Including, the system.
2. The act of adding the first video segment to the set of video segments comprises: in an iterative process of identifying a plurality of video segments based on the tempo of the music track and the duration of each identified video segment; The system of claim 1 , wherein the first video segment is one of the plurality of video segments.
3. The operations further include receiving a tempo selection; accessing the music track from the database is based on the selected tempo and the tempo of the music track; The system of claim 1 .
4. The operation is accessing a second video sequence from the one or more databases, the second video sequence having a second plurality of video segments; adding a second video segment to the set of video segments based on the tempo of the music track and a duration of the second video segment of the second plurality of video segments; The system of claim 1 further comprising:
5. The operation is identifying transitions between the plurality of video segments in the video sequence based on a length of distance between successive frames in the video sequence; The system of claim 1 further comprising:
6. The operations further include accessing a search query; The system of claim 1 , wherein the accessing of the video sequences from the one or more databases is based on the search query.
7. 2. The system of claim 1, wherein adding the first video segment to the set of video segments is based on the duration of the first video segment being an integer multiple of a beat duration of the music track.
8. The system of claim 1 , wherein generating the audio / video sequence comprises generating the audio / video sequence having a predetermined duration.
9. The system of claim 1 , wherein generating the audio / video sequence comprises generating the audio / video sequence having a duration equal to a duration of the music track.
10. The system of claim 1 , wherein generating the audio / video sequence comprises generating the audio / video sequence having a user-selected duration.
11. accessing, by one or more processors, a music track having a tempo from an audio database; accessing, by the one or more processors, a video sequence from a video database, the video sequence having a plurality of video segments; adding, by the one or more processors, a first video segment of the plurality of video segments to a set of video segments based on the tempo of the music track and a duration of the first video segment; generating, by said one or more processors, an audio / video sequence having said set of video segments and said music track; A method comprising:
12. further comprising receiving a tempo selection; said accessing said music track from said audio database is based on said selected tempo and said tempo of said music track; The method of claim 11.
13. accessing a second video sequence from the video database, the second video sequence having a second plurality of video segments; adding a second video segment to the set of video segments based on the tempo of the music track and a duration of the second video segment of the second plurality of video segments; The method of claim 11 further comprising:
14. identifying transitions between the plurality of video segments in the video sequence based on a length of distance between successive frames in the video sequence; The method of claim 11 further comprising:
15. further comprising accessing the search query; The method of claim 11 , wherein the accessing the video sequences from the video database is based on the search query.
16. 12. The method of claim 11, wherein adding the first video segment to the set of video segments is based on the duration of the first video segment being an integer multiple of a beat duration of the music track.
17. The method of claim 11 , wherein generating the audio / video sequence comprises generating the audio / video sequence having a predetermined duration.
18. The method of claim 11 , wherein generating the audio / video sequence comprises generating the audio / video sequence having a duration equal to a duration of the music track.
19. The method of claim 11 , wherein generating the audio / video sequence includes generating the audio / video sequence having a user-selected duration.
20. A non-transitory machine-readable medium having instructions that, when executed by one or more processors of a machine, cause the machine to perform operations, the operations including: accessing a music track having a tempo from an audio database; accessing a video sequence from a video database, the video sequence having a plurality of video segments; adding a first video segment of the plurality of video segments to a set of video segments based on the tempo of the music track and a duration of the first video segment; generating an audio / video sequence comprising said set of video segments and said music track; The non-transitory machine-readable medium comprising: