Method, system and computer program for generating an audio output file
The method and system automate the selection of compatible instrumental content blocks in DAWs, addressing the time-consuming challenge of creating harmonious audio output files by using musical style-based templates and user customization.
Patent Information
- Application Number
- JP2025532954
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-11-20
- Publication Date
- 2025-12-16
AI Technical Summary
Creating harmonious and compatible audio output files in Digital Audio Workstations (DAWs) is time-consuming due to the challenge of selecting consistent and compatible pre-recorded audio content.
A method and system for generating audio output files by selecting instrumental content blocks based on musical style, using predetermined musical rules and templates, allowing for harmonious combinations and user customization.
Facilitates the creation of harmonious and customizable audio output files by automating the selection of compatible instrumental content blocks, reducing time and effort in the production process.
Smart Images

Figure 2025540804000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a method, system, and computer program for generating an audio output file. [Background technology]
[0002] Digital Audio Workstations (DAWs) have been developed to provide users with a production environment in which audio content can be created, recorded, edited, mixed, and optionally synchronized with target image or video content.
[0003] Such DAWs typically consist of an array of tools and a library of pre-recorded audio content that a user can select, edit, and combine to create audio output files that can be synchronized, if desired, with multimedia content such as image and / or video files.
[0004] However, in such a production environment, it can be very time-consuming for even the most skilled audio editor for a user to select a consistent and compatible pre-recorded audio content file for an audio output file.
[0005] It is therefore an object of the present disclosure to provide a system, method and computer program for generating audio output files that overcomes at least some of the above problems and / or provides a useful alternative for the public or industry.
[0006] Further aspects of the described embodiments will become apparent from the following description, which is given by way of example only. Summary of the Invention
[0007] According to one embodiment, there is provided a computer-implemented method for generating an audio output file, comprising the steps of: a. selecting a subset of instrumental content blocks from a group of instrumental content blocks by determining a musical style of the audio output file, the musical style being composed of a plurality of musical slots, each slot being associated with a predetermined musical rule; b. selecting a musical template from a plurality of musical templates for the slot using predetermined musical rules for each slot, the selected musical template defining a chord progression of a set of musical chords in a musical key and tempo; c. For each slot, selecting an instrumental content block that matches the chord progression defined by the selected music template; d. generating an audio output file by combining a subset of the instrumental content blocks.
[0008] Embodiments provide methods and systems for generating audio output files from blocks of instrumental content (also called stems) that, when combined, are harmonious and compatible, and therefore pleasant to listen to for the user.
[0009] The method provides for the configuration of musical slots, each specified by a musical rule that determines a musical template for the audio output file.
[0010] Instrumental content blocks that have matching chord progressions defined by the template are selected for use together in the audio output file.
[0011] The system provides the user with the option of generating an audio output file that includes a random selection of harmonically compatible instrumental content blocks or a selection of harmonically compatible instrumental content blocks based on a user's style selection (e.g., pop, synth, reggae, etc.) made via a user interface means. Once an initial selection of harmonically compatible instrumental content blocks is provided, the user may then apply editing and authoring tools to change, alter, adjust, shuffle, and / or delete the instrumental content blocks within the selection to tailor the acoustic audio output file to their preferences.
[0012] In one embodiment, each instrumental content block of a group of instrumental content blocks includes a plurality of tags, each tag associated with a musical parameter of the respective instrumental content block, such that the plurality of tags of the instrumental content block uniquely identifies the instrumental content block.
[0013] In one embodiment, each instrumental content block is created by a human performer according to a musical template, and each instrumental content block contains musical content from an instrument.
[0014] In one embodiment, the musical style is determined by one or more of the musical genre, musical style, artist name, song title, chord progression, tempo, and / or instruments involved in creating the audio output file.
[0015] In one embodiment, the method includes searching a database containing records of artist names and song titles to determine the musical style of the audio output file.
[0016] In one embodiment, a method includes receiving an audio input file including at least one vocal content block, each vocal content block including a vocal performance, and an audio output file generated by combining the vocal content block with a subset of the instrumental content blocks. Thus, the audio output file may include one or more vocal content blocks.
[0017] In one embodiment, the method comprises: a. receiving an audio input file including a vocal and / or instrumental performance; b. separating the audio input file into vocal content blocks and a subset of instrumental content blocks, the vocal content blocks comprising a vocal performance and each instrumental content block comprising audio content from one instrument involved in creating the instrumental performance; c. replacing the subset of instrumental content blocks with an alternative subset of instrumental content blocks; d. generating an audio output file by combining the vocal content block with one or more alternative instrumental content blocks.
[0018] Thus, the present invention may receive a song from a known artist, and a user may be provided with, or manually select, alternative instrumental content blocks in place of the original instrumental performance while retaining the vocal performance of the song. In this way, an audio output file is generated that retains the artist's vocal performance but has alternative acoustic accompaniment that can be adapted as desired by the user selecting the alternative instrumental content blocks until the user is satisfied with the sound of the final audio output file.
[0019] In one embodiment, the musical style of the audio input file is determined by analyzing vocal content blocks resulting from a vocal performance, and an alternative subset of instrumental content blocks is automatically selected according to the determined musical style.
[0020] In one embodiment, the musical style of the audio input file is determined by analyzing one or more subsets of instrumental content blocks resulting from the instrumental performance, and alternative subsets of instrumental content blocks are automatically selected according to the determined musical style.
[0021] In one embodiment, the alternative subset of instrumental content blocks is selected by a user operating a user interface means.
[0022] In one embodiment, the method comprises: a. receiving an audio input file containing an instrumental performance; b. Separating the audio input file into instrumental content blocks, each containing audio content from an instrument involved in creating an instrumental performance; c. determining the musical style of the audio input file by analyzing one or more instrumental content blocks; d. selecting a subset of alternative instrumental content blocks according to the determined musical style, such that the selected subset of alternative instrumental content blocks, when combined, produces a similar sound to the instrumental performance of the audio output file; e. replacing the instrumental content blocks with a subset of the alternative instrumental content blocks and generating an audio output file by combining the selected subset of the alternative instrumental content blocks.
[0023] In one embodiment, the method includes a step in which a user operates an audio recording means to record one or more audio input files.
[0024] Such audio recordings may be vocal and / or instrumental performances that may be incorporated into the audio output file as one or more vocal or instrumental content blocks. The audio recording means provides a selection of audio signal processors that enhance the acoustics of the recordings for the audio input file. Examples of these signal processors include reverb, delay, compressor, and pitch correction that manipulate and enhance the vocal and / or instrumental recordings. The audio recording(s) may be connected to specific sections or portions of the audio output file and may also be duplicated (copy / paste) and used in multiple portions of the audio output file.
[0025] In one embodiment, the method includes manipulating a user interface means provided by a back-end application programming interface (API) to create an audio output file. During operation, the native application calls the API that uses the back-end audio output file generator means to create the audio output file. If, during the creation process, the user modifies the vocal or instrumental content blocks of the audio output file or its audio parameters, this process is repeated.
[0026] In one embodiment, the present invention provides a web application that allows anyone to create music using a visual interface.
[0027] An application programming interface (API) provides a central communication point for all applications that connect to the audio output file generator means. Optionally, the API may be a partially open source interface so that third parties may use the API to create music for their own platform.
[0028] The audio output file generator means is the heart of the music creation. It communicates with the style database module and the template database module to create the audio output files and uses the logic entered into these modules. The audio output file generator means has built-in logic for creating the audio output files according to the style.
[0029] In one embodiment, the method includes operating a multimedia synchronization means to mix an audio output file with a composition, a photo, a video, or filtered multimedia.
[0030] In one embodiment, the method includes operating shuffling means configured to exchange instrumental content blocks in slots with different instrumental content blocks according to provided musical rules of the determined style.
[0031] For example, if a user is listening to an audio output file with five instrumental content blocks for various instruments (including one for guitar), and the user does not like the guitar instrumental content block, it may be shuffled or swiped to remove it, and an alternative instrumental content block that respects the determined style slot rules will be provided in place of the removed instrumental content block.
[0032] In one embodiment, for blocks of instrumental content that do not have a determined pitch or key, a special tag is applied to these blocks, allowing them to be used with any template in the same tempo range.
[0033] In one embodiment, the method includes reusing existing vocal content blocks in multiple compatible templates. To accommodate this feature, tables associated with the musical templates are provided for tempo (bpm), key, and chord progression, and such tables are used to locate the associated template with which the mono vocal content block is associated. Special tags are applied to these existing vocal content blocks, allowing them to be used in other associated templates in addition to the tables.
[0034] In one embodiment, a method includes importing an audio input file including a vocal performance, converting the vocal performance into one or more vocal content blocks, tagging the vocal content blocks, and using the vocal content blocks in one or more audio output files.
[0035] Such offerings facilitate remixes and new arrangements of existing songs, creating several different versions of well-known songs. A key component is capturing the parameters of the vocal performance, analyzing these parameters, and appropriately tagging the file so that it can function in an audio output file. Previously recorded vocal performances may be used as well.
[0036] In one embodiment, the method includes changing the key of the audio output file or a section thereof by replacing one or more instrumental content blocks within the audio output file or section with alternative instrumental content blocks in an alternative key.
[0037] In one embodiment, each template is divided into multiple template sections, each of which is 4 or 8 bars, so that each template section is tagged according to its position within the sections of the audio output file, and the template sections can be arranged in different orders. Such configuration gives the audio output file different song structures and can be done automatically using predefined instructions or by the user. Different musical genres can have different section arrangements. The use of manipulated sections can be used to lengthen and shorten the audio output file.
[0038] In one embodiment, the method further includes dividing each instrumental content block into sections, each section being a portion of the instrumental content block, and the method includes muting and unmuting the sections, such configuration being independent of the music slots and may be done automatically using predetermined instructions by the user.
[0039] In one embodiment, the method further includes selecting multiple musical templates for the musical slots using predetermined musical rules for each slot. Such an arrangement allows for the creation of an audio output file from multiple templates to provide a template "mashup." In this manner, sections from different associated templates are ordered to produce a musically pleasing result. A table associated with the musical templates may also be utilized to find associated templates.
[0040] In one embodiment, the method further includes providing a user interface means that allows a user to change the key and tempo of the audio output file, thus speeding up or slowing down the tempo (within set limits) and shifting the pitch of the audio output file until a pitch that best suits those voices recorded or imported as audio input files is found.
[0041] According to an embodiment, there is provided a computer-implemented system for generating an audio output file, comprising: a. means for selecting a subset of instrumental content blocks from a group of instrumental content blocks by means for determining a musical style of the audio output file, the musical style being comprised of a plurality of musical slots, each slot being associated with a predetermined musical rule; b. means for selecting a musical template from a plurality of musical templates for a slot using predetermined musical rules for each slot, the selected musical template defining a chord progression of a set of musical chords in a musical key and tempo; c. for each slot, a means for selecting a block of instrumental content that matches the chord progression defined by the selected music template; d. means for generating an audio output file by combining a subset of the instrumental content blocks.
[0042] In one embodiment, the tagging means provides for tagging each instrumental content block of the group of instrumental content blocks with a plurality of tags, each tag being associated with a musical parameter of the respective instrumental content block, whereby the plurality of tags of the instrumental content block uniquely identifies the instrumental content block.
[0043] In one embodiment, each instrumental content block is created by a human performer according to a musical template.
[0044] In one embodiment, the musical style is determined by one or more of the musical genre, musical style, artist name, song title, chord progression, tempo, and / or instruments involved in creating the audio output file.
[0045] In one embodiment, the system includes means for searching a database containing records of artist names and song titles to determine the musical style of the audio output file.
[0046] In one embodiment, the system includes means for receiving an audio input file including at least one vocal content block, the vocal content block including a vocal performance, and an audio output file generated by combining the vocal content block with a subset of the instrumental content blocks.
[0047] In one embodiment, the system comprises: a. means for receiving an audio input file comprising a vocal and / or instrumental performance; b. means for separating the audio input file into vocal content blocks and a subset of instrumental content blocks, the vocal content blocks comprising a vocal performance and each instrumental content block comprising audio content from an instrument involved in creating the instrumental performance; c. means for replacing a subset of instrumental content blocks with an alternative subset of instrumental content blocks; d. means for generating an audio output file by combining the vocal content block with one or more alternative instrumental content blocks.
[0048] In one embodiment, the system includes means for analyzing vocal content blocks resulting from a vocal performance to determine a musical style of the audio input file, and an alternative subset of instrumental content blocks is automatically selected according to the determined musical style.
[0049] In one embodiment, the system includes means for analyzing one or more of the subsets of instrumental content blocks resulting from the instrumental performance to determine a musical style of the audio input file, and an alternative subset of instrumental content blocks is automatically selected according to the determined musical style.
[0050] In one embodiment, the alternative subset of instrumental content blocks is selected by a user operating a user interface means.
[0051] In one embodiment, the system comprises: a. means for receiving an audio input file comprising an instrumental performance; b. means for separating the audio input file into instrumental content blocks, each of the instrumental content blocks containing audio content from an instrument involved in creating an instrumental performance; c. means for determining the musical style of the audio input file by analyzing one or more instrumental content blocks; d. means for selecting a subset of alternative instrumental content blocks according to the determined musical style such that the selected subset of alternative instrumental content blocks, when combined, produces a similar sound to the instrumental performance of the audio output file; and e. means for replacing the instrumental content blocks with a subset of the alternative instrumental content blocks and generating an audio output file by combining the selected subset of the alternative instrumental content blocks.
[0052] In one embodiment, the system includes an audio recording means whereby a user records one or more audio input files.
[0053] Such audio recordings may be vocal and / or instrumental performances that may be incorporated into the audio output file as one or more vocal or instrumental content blocks.
[0054] The audio recording means provides a selection of audio signal processors for the audio input file that enhance the sound of the recording. Examples of these signal processors include reverb, delay, compressor, and pitch correction to manipulate and enhance recordings of vocal and / or instrumental performances. The audio recording(s) can be connected to specific sections or portions of the audio output file and can also be duplicated (copy / paste) and used in multiple portions of the audio output file.
[0055] In one embodiment, the method includes a user interface means provided by a back-end application programming interface (API) for creating an audio output file. During operation, the native application calls the API that uses the back-end audio output file generator means to create the audio output file. If, during the creation process, the user modifies the vocal or instrumental content blocks of the audio output file or its audio parameters, this process is repeated.
[0056] In an embodiment, a web application is provided that allows anyone to create music using a visual interface.
[0057] An application programming interface (API) provides a central communication point for all applications that connect to the audio output file generator means. Optionally, the API may be a partially open source interface so that third parties may use the API to create music for their own platform.
[0058] The audio output file generator means is the heart of the music creation. It communicates with the style database module and the template database module to create the audio output files and uses the logic entered into these modules. The audio output file generator means has built-in logic for creating the audio output files according to the style.
[0059] In one embodiment, the system includes multimedia synchronization means for mixing audio output files with compositions, photos, videos, or filtered multimedia.
[0060] In one embodiment, the system includes shuffling means configured to exchange instrumental content blocks in slots with different instrumental content blocks according to provided musical rules of the determined style.
[0061] For example, if a user is listening to an audio output file with five instrumental content blocks for various instruments (including one for guitar), and the user does not like the guitar instrumental content block, it may be shuffled or swiped to remove it, and an alternative instrumental content block that respects the determined style slot rules will be provided in place of the removed instrumental content block.
[0062] In one embodiment, for blocks of instrumental content that do not have a determined pitch or key, a special tag is applied to these blocks, allowing them to be used with any template in the same tempo range.
[0063] In one embodiment, the system includes a means for reusing existing vocal content blocks across multiple compatible templates. To accommodate this feature, tables associated with musical templates are provided for tempo (bpm), key, and chord progression, and these tables are used to locate the associated template with which a mono vocal content block is associated. Special tags are applied to these existing vocal content blocks, allowing them to be used in addition to the tables in other associated templates.
[0064] In one embodiment, the system includes means for importing an audio input file including a vocal performance, converting the vocal performance into one or more vocal content blocks, tagging the vocal content blocks, and using the vocal content blocks in one or more audio output files.
[0065] Such offerings facilitate remixes and new arrangements of existing songs, creating several different versions of well-known songs. A key component is capturing the parameters of the vocal performance, analyzing these parameters, and appropriately tagging the file so that it can function in an audio output file. Previously recorded vocal performances may be used as well.
[0066] In one embodiment, the system includes means for changing the key of an audio output file or a section thereof by replacing one or more instrumental content blocks within the audio output file or section with alternative instrumental content blocks in an alternative key.
[0067] In one embodiment, the system includes means for dividing each template into multiple template sections, each of which is 4 or 8 bars long, so that each template section is tagged according to its position within the sections of the audio output file, and the template sections can be arranged in different orders. Such configuration gives the audio output file different song structures and can be done automatically using predefined instructions or by the user. Different musical genres can require different section arrangements. The use of manipulated sections can be used to lengthen and shorten the audio output file.
[0068] In one embodiment, the system includes means for dividing each instrumental content block into sections, each section being a portion of the instrumental content block, and the method includes muting and unmuting the sections, such configuration being independent of the music slots and may be done automatically using predetermined instructions by the user.
[0069] In one embodiment, the system includes means for selecting multiple musical templates for musical slots using predetermined musical rules for each slot. Such an arrangement allows audio output files to be created from multiple templates to provide a template "mashup." In this way, sections from different associated templates can be ordered to produce a musically pleasing result. Tables associated with musical templates may also be utilized to find associated templates.
[0070] In one embodiment, the system includes user interface means that allow a user to change the key and tempo of the audio output file, and thus may speed up or slow down the pitch of the audio output file (within set limits) until a pitch is found that best suits the voices recorded or imported as audio input files.
[0071] In a further embodiment of the present invention, there is provided a computer program comprising instructions which, when executed by one or more processors, cause the one or more processors to perform steps in accordance with the described method.
[0072] In yet another embodiment of the present invention, there is provided a computing device and / or arrangement of computing devices having one or more processors, memory, and display means operable to display an interactive user interface having the features described.
[0073] In a further embodiment, the present invention is configured to record an audio input file containing a vocal and / or instrumental performance and simultaneously output an audio output file over Bluetooth using a protocol such as A2DP.
[0074] Such an arrangement is particularly useful when the user is using a wireless headset or headphones that have a built-in speaker output for audio playback and a microphone input for recording.
[0075] This configuration introduces latency when an input stream is recorded by a microphone and an output stream is played back simultaneously by a speaker. By determining this latency, the output stream and the input stream can be synchronized. This configuration ensures that the recording made by the microphone is coordinated and synchronized with the audio playback when heard on the speaker, providing high fidelity audio input and high fidelity audio output simultaneously.
[0076] In a further embodiment, the present invention is configured to allow multiple audio tracks to be layered together to enable harmonies, overdubbing, and many other applications in music production.
[0077] However, smartphone devices have less processing power than digital audio workstations, limiting the number of vocal tracks they can process simultaneously. Processing more vocal tracks than the device can handle can result in serious audio artifacts that are unacceptable for music production.
[0078] To address this issue, the present invention allows a settings screen to allow the user to configure settings for a single actively selected vocal track in real time while pre-processing of the vocal track occurs in a background thread. After exiting the settings screen and after the vocal track has been pre-processed in the background thread, the settings allow a smooth transition from the real-time processed track to the pre-processed track. This transition and pre-processing of vocal tracks in the background thread allows the user to work with as many processed vocals as needed.
[0079] The embodiments will be more clearly understood from the following description of some embodiments thereof, given by way of example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0080] [Figure 1] 1 is a block diagram illustrating a system for generating an audio output file according to an embodiment of the present invention. [Figure 2] FIG. 2 is a detailed block diagram illustrating the use of styles and slots to select instrumental content blocks for an audio output file, according to an embodiment of the present invention. [Figure 3] FIG. 2 is a stylized diagram illustrating an embodiment of the present invention for use in generating an audio output file. DETAILED DESCRIPTION OF THE INVENTION
[0081] Embodiments of the present invention are implemented by one or more computer processors and memory containing computer software program instructions executable by the one or more processors, which may be provided by a computer server or network of connected and / or distributed computers.
[0082] Audio files of the present invention, including vocal content blocks, instrumental content blocks, audio input files, and audio output files, will be understood to be received, stored, or recorded files containing audio or MIDI data or content that, when processed by an audio or MIDI player, produces an acoustic output. Audio files may be recorded in any known audio file format, including, but not limited to, audio WAV format, MP3 format, Advanced Audio Coding (AAC) format, Ogg format, or in any other format, analog, digital, or otherwise, as desired. The desired audio or MIDI format may optionally be specified by the user.
[0083] A user may record one or more audio input files, which may be vocal and / or instrumental performances that may be incorporated into an audio output file as one or more vocal or instrumental content blocks.
[0084] An embodiment of the present invention provides a computer-implemented system 1 for generating an audio output file 20. The system includes generator means 10 that provides means for selecting a subset of instrumental content blocks (or stems) 70 from a group of instrumental content blocks (or stems) stored in a database 60. Each instrumental content block contains musical content from a musical instrument. Each instrumental content block 70 is created by a human performer according to a musical template 40 that defines a chord progression of a set of musical chords in a musical key and tempo that the performer must follow when creating the instrumental content block 70.
[0085] The generator 10 is configured to determine the musical style 30 according to one or more parameters including musical genre 90, musical style, artist name, song title, chord progression, tempo, and / or instruments involved in creating the audio output file.
[0086] The musical style may be determined based on an analysis of parameters of an audio input file, such as a recording of a vocal melody received from a user. Such vocal melody is converted for inclusion as a vocal content block in audio output file 20. The musical style of audio output file 20 may alternatively be provided by a user selecting from a range of styles provided by user interface means 120, or may be determined by system 1 analyzing musical parameters of an audio input file initially provided by the user, such as chord progressions and their tempos. Optionally, the user may search a database 100 containing records of artist names and song titles to determine the musical style of audio output file 20.
[0087] As shown in Figure 2, a musical style 30 consists of a number of musical slots 31, 32, 33, 34, 35, each associated with a predetermined musical rule. In the simplified example shown in Figure 2, five slots are provided for the style "disco / funk", each slot having a rule. For all five slots of the style "disco / funk", slot 1, indicated by reference number 31, has the rules "disco / funk pop" and "drums", slot 2, indicated by reference number 32, has the rules "disco / funk pop" and "bass", etc.
[0088] Using predetermined musical rules for each slot 31, 32, 33, 34, 35, the system 1 selects a musical template 50 from a database of musical templates for slots 40, where the selected musical template 50 defines a chord progression of a set of musical chords in a musical key and tempo. The system 1 then selects, for each slot 31, 32, 33, 34, 35, instrumental content blocks 70 from the database 60 that match the chord progression defined in the selected musical template 50 and satisfy other rules defined for the slot, and generates the audio output file 10 by combining a subset of the selected instrumental content blocks 70.
[0089] The tagging means 80 provides for tagging or labeling each instrumental content block 70 with an identifier, each tag being associated with a musical parameter of the instrumental content block 70, and the tags attached to an instrumental content block uniquely identify the instrumental content block 70.
[0090] For example, each instrumental content block or stem 70 in the system 1 is tagged by a human in a central tagging means to describe the characteristics of the instrumental content block or stem. Examples of tags for instrumental content blocks or stems are shown below: Instrumental content block provider (i.e., human performer name) = John Doe Key=C# Tempo=82bpm Instrumental music = guitar Style = Pop Strength = Large
[0091] A user may additionally and optionally provide an audio input file containing at least one vocal content block for the audio output file. Such vocal content block comprises a vocal performance, and the audio output file is generated by combining the user-provided vocal content block with a subset of instrumental content blocks 70 selected by system 1 for the vocal content block. As shown in Figure 1, vocal creation 140 is provided within the application and utilizes recording means to allow users to sing and record their own songs for use in the audio output file.
[0092] In one application of the present invention, an audio input file containing a vocal performance and / or an instrumental performance is received, the audio input file is separated into vocal content blocks and instrumental content blocks, the vocal content blocks containing the vocal performance and each instrumental content block containing audio content from an instrument involved in creating the instrumental performance.
[0093] A user may interact with the system to manually or automatically replace one or more of the instrumental content blocks with an alternative subset of instrumental content blocks.
[0094] In this embodiment, the musical style of the audio input file is determined by analyzing musical parameters of the vocal content blocks resulting from the vocal performance, and an alternative subset of the instrumental content blocks is automatically selected according to the musical style determined based on the parameters. Alternatively, the musical style of the audio input file is determined by analyzing one or more of the subsets of instrumental content blocks resulting from the instrumental performance, and an alternative subset of the instrumental content blocks is automatically selected according to the determined musical style. The alternative subset of the instrumental content blocks may also be selected by a user operating a user interface means.
[0095] An audio output file is generated by the system that combines vocal content blocks with one or more alternative instrumental content blocks to provide variations of the original audio input file.
[0096] In this way, the present invention may receive a song from a known artist, and while preserving the vocal performance of the song, the user may be provided with, or may manually select, alternative instrumental content blocks that harmoniously combine with the vocal performance in place of the original instrumental performance.
[0097] The user may also interact with the system to manually or automatically replace one or more of the instrumental content blocks with an alternative subset of instrumental content blocks and / or vocal content blocks.
[0098] In some applications, rules may be implemented to provide alternative audio output file offerings. Such rules may include: a. If the user chooses to retain the vocal content block in the audio output file, at least one instrumental content block in the audio output file must be modified. b. If the user chooses to remove a vocal content block, the user may add a new vocal content block and / or need to modify at least one instrumental content block in the audio output file. c. If no vocal content blocks are present in the audio output file, then at least one instrumental content block in the audio output file needs to be modified and / or a vocal content block is added.
[0099] Additionally, the pitch, tempo, and section layout of the audio output file may also be changed to provide alternative audio output file offerings.
[0100] 3, the system may receive an audio input file containing an instrumental performance via a microphone or receiver of a user electronic device 200, such as a mobile smartphone. For example, a user may switch the system into a listening mode in which the microphone receives a song or performance playing in the background as an audio input file.
[0101] The system uses a song analyzer 150 to separate the audio input file into instrumental content blocks 70, each containing audio content from an instrument involved in creating an instrumental performance, and together with a generator 10, determines the musical style 30 of the audio input file by analyzing parameters of one or more instrumental content blocks 70, such as chord progression, tempo, etc.
[0102] The generator 10 selects a subset of alternative instrumental content blocks 70 according to the determined musical style 30 such that, when combined, the selected subset of alternative instrumental content blocks 70 sounds similar to the instrumental performance of the audio input file.
[0103] An audio output file 20 is then created by combining a subset of the selected alternative instrumental content blocks 70 to create a "sound-alike."
[0104] An embodiment may be provided by a back-end application programming interface (API) 110 to create an audio output file. A software application or app 130 may be downloaded and installed on the electronic device to display a user interface that engages with the API 110. However, in some embodiments, the electronic device may execute a web browser application 120 that browses a website provided by a web server, where an embedded user interface is displayed.
[0105] In an embodiment, a web application may be provided that allows anyone to create music using a visual interface.
[0106] It will be understood that the invention is not limited to the specific details described herein, which are given by way of example only, and that various modifications and changes are possible without departing from the scope of the invention.
Claims
1. 1. A computer-implemented method for generating an audio output file, comprising: - selecting a subset of instrumental content blocks from a group of instrumental content blocks by determining a musical style of said audio output file, said musical style being composed of a plurality of musical slots, each slot being associated with a predefined musical rule; selecting a musical template from a plurality of musical templates for each slot using the predetermined musical rules for the slot, the selected musical template defining a chord progression of a set of musical chords in a musical key and tempo; selecting, for each slot, a block of instrumental content that matches the chord progression defined by the selected musical template; generating the audio output file by combining a subset of the instrumental content blocks.
2. The method of claim 1 , comprising providing a user with an audio output file having a random selection of harmonically compatible instrumental content blocks.
3. 10. The method of claim 1, comprising providing a user with an audio output file having a selection of harmonic and compatible instrumental content blocks based on a style selection made by the user.
4. receiving an audio input file including a vocal and / or instrumental performance; separating the audio input file into vocal content blocks and a subset of instrumental content blocks, the vocal content blocks comprising the vocal performance and each instrumental content block comprising audio content from an instrument involved in creating the instrumental performance; replacing the subset of instrumental content blocks with an alternative subset of instrumental content blocks; generating the audio output file by combining the vocal content block with the one or more alternative instrumental content blocks.
5. 5. The method of claim 4, wherein the musical style of the audio input file is determined by analyzing the vocal content blocks resulting from the vocal performance, and an alternative subset of the instrumental content blocks is automatically selected according to the determined musical style.
6. 5. The method of claim 4, wherein the musical style of the audio input file is determined by analyzing one or more of the subsets of the instrumental content blocks resulting from the instrumental performance, and an alternative subset of the instrumental content blocks is automatically selected according to the determined musical style.
7. The method of claim 4 , wherein the alternative subsets of instrumental content blocks are selected by a user operating a user interface means.
8. receiving an audio input file comprising an instrumental performance; separating the audio input file into instrumental content blocks, each of the instrumental content blocks containing audio content from an instrument involved in creating the instrumental performance; determining a musical style of the audio input file by analyzing the one or more instrumental content blocks; selecting a subset of alternative instrumental content blocks according to the determined musical style, such that the selected subset of alternative instrumental content blocks, when combined, sounds similar to the instrumental performance of the audio output file; and generating the audio output file by combining the selected subsets of alternative instrumental content blocks.
9. 2. The method of claim 1, wherein the musical style is determined by one or more of the musical genre, musical style, artist name, song title, chord progression, tempo, and / or instruments involved in creating the audio output file.
10. 2. The method of claim 1, wherein each instrumental content block of the subset of instrumental content blocks includes a plurality of tags, each tag associated with a musical parameter of the respective instrumental content block, whereby the plurality of tags of an instrumental content block uniquely identifies the instrumental content block.
11. 10. The method of claim 1, further comprising searching a database containing records of artist names and song titles to determine the musical style of the audio output file.
12. 2. The method of claim 1, wherein the method includes receiving an audio input file including at least one vocal content block, each vocal content block including a vocal performance, and the audio output file is generated by combining the vocal content block with a subset of the instrumental content blocks.
13. 10. The method of claim 1, further comprising operating a multimedia synchronization means to mix the audio output file with a composition, a photo, a video, or filtered multimedia.
14. 2. The method of claim 1, comprising operating shuffling means configured to exchange instrumental content blocks in slots with different instrumental content blocks according to the musical rules of the determined style.
15. 10. The method of claim 1, comprising changing the key of an audio output file or section by replacing one or more instrumental content blocks within the audio output file or section with alternative instrumental content blocks in an alternative key.
16. 10. The method of claim 1, comprising dividing each instrumental content block into sections, each section being a part of the instrumental content block, the method comprising allowing a user to mute and / or unmute sections.
17. 2. The method of claim 1, including the step of providing a user interface means that allows a user to modify audio parameters of the audio output file.
18. 1. A computer-implemented system for generating an audio output file, comprising: means for selecting a subset of instrumental content blocks from a group of instrumental content blocks by means for determining a musical style of said audio output file, said musical style being composed of a plurality of musical slots, each slot being associated with a predetermined musical rule; means for selecting a musical template from a plurality of musical templates for each slot using the predetermined musical rules for the slot, the selected musical template defining a chord progression of a set of musical chords in a musical key and tempo; means for selecting, for each slot, a block of instrumental content that matches the chord progression defined by the selected musical template; and means for generating the audio output file by combining a subset of the instrumental content blocks.
19. means for receiving an audio input file comprising a vocal and / or instrumental performance; means for separating the audio input file into vocal content blocks and a subset of instrumental content blocks, the vocal content blocks comprising the vocal performance and each instrumental content block comprising audio content from an instrument involved in creating the instrumental performance; means for replacing the subset of instrumental content blocks with an alternative subset of instrumental content blocks; and means for generating the audio output file by combining the vocal content block with the one or more alternative instrumental content blocks.
20. A computer program comprising instructions which, when executed by one or more processors, cause said one or more processors to perform steps in accordance with the method of claim 1.