Interactive Movement and Voice Engine
A system that translates user movements into musical sequences with music theory enforcement simplifies music composition, addressing the need for instrument proficiency and theory knowledge, enabling accessible music creation.
Patent Information
- Application Number
- JP2024537405
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-20
- Filing Date
- 2022-11-23
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Composing music requires extensive practice in instrument skills and music theory, making it difficult for individuals without formal training to create musically pleasing songs.
A system that captures a user's interactive movements using an image sensor, maps them to audio element identifiers, and enforces music theory rules to generate a musical sequence, reducing the need for instrument proficiency and music theory knowledge.
Enables users to create musically coherent compositions by simplifying the composition process, making it accessible to those without formal musical training.
Smart Images

Figure 0007736933000001 
Figure 0007736933000002 
Figure 0007736933000003
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent No. 17 / 556,178, filed December 20, 2021, entitled "INTERACTIVE MOTION AUDIO ENGINE," the disclosure of which is incorporated herein by reference in its entirety.
[0002] Composing and sharing songs and musical sequences is a common way for individuals to build social bonds. Some people learn to play instruments such as piano, guitar, or percussion instruments to compose for themselves and share with others. However, while mastering a single instrument requires years of practice and study, many songs and musical sequences may use several different instruments. Creating a short musical sequence that others will find pleasant to listen to (i.e., has a high level of musicality) may require several hours of writing out the notes, playing different instruments and recording them on different tracks, and editing the tracks together. Additionally, the rules of music theory that promote musicality are subtle and may be difficult to follow for users without a musical background.
[0003] The embodiments are described with reference to these and other general considerations, and while relatively specific problems are discussed, it should be understood that the embodiments should not be limited to solving the specific problems identified in the background. Summary of the Invention [Problem to be solved by the invention]
[0004] Aspects of the present disclosure relate to generating audio output. [Means for solving the problem]
[0005] In one aspect, a method for generating an audio output is provided. Image input of interactive movements by a user is received, captured by an image sensor. The interactive movements are mapped to a sequence of audio element identifiers. The sequence of audio element identifiers is processed to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers. An audio output representing the musical sequence is generated.
[0006] In another aspect, a system for generating an audio output is provided, the system comprising one or more hardware processors configured with machine-readable instructions to: receive image input of interactive movements by a user captured by an image sensor, map the interactive movements to a sequence of audio element identifiers, process the sequence of audio element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers, and generate an audio output representing the musical sequence.
[0007] In yet another aspect, a non-transitory computer-readable storage medium is provided, the medium including instructions executable by one or more processors that, when executed by the one or more processors, cause the one or more processors to receive image input of interactive movements by a user captured by an image sensor, map the interactive movements to a sequence of sound element identifiers, process the sequence of sound element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of sound element identifiers, and generate an audio output representing the musical sequence.
[0008] This Summary is provided to introduce in a simplified form a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. [Brief explanation of the drawings]
[0009] Non-limiting and non-exhaustive examples are described with reference to the following figures:
[0010] [Figure 1] FIG. 1 is a block diagram of an example system for generating audio output according to an example of the present disclosure.
[0011] [Figure 2] FIG. 2 is a block diagram of an example of a user input processor of a computing device according to an example of the present disclosure.
[0012] [Figure 3] FIG. 10 illustrates an exemplary output image for a graphical user interface according to an example of the present disclosure.
[0013] [Figure 4A] 1A-1C illustrate exemplary sequences of image inputs with facial expressions for generating audio outputs according to examples of the present disclosure.
[0014] [Figure 4B] FIG. 1 illustrates an exemplary image input with facial expression elements for generating an audio output according to an example of the present disclosure.
[0015] [Figure 5] 1 is a flowchart of an exemplary method for generating an audio output according to an example of the present disclosure.
[0016] [Figure 6]FIG. 1 is a block diagram illustrating exemplary physical components of a computing device that can be used to implement aspects of the present disclosure.
[0017] [Figure 7] FIG. 1 is a simplified block diagram of a mobile computing device that can be used to implement aspects of the present disclosure. [Figure 8] FIG. 1 is a simplified block diagram of a mobile computing device that can be used to implement aspects of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific embodiments or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the disclosure. The embodiments may be embodied as methods, systems, or apparatuses. Accordingly, the embodiments may take the form of a hardware implementation, an entirely software implementation, or an implementation combining software and hardware aspects. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the disclosure is limited only by the appended claims and their equivalents.
[0019] As mentioned above, composing and sharing music is common, but composing often requires extensive practice to master technical instrument skills and music theory rules to improve musicality. This disclosure provides methods and systems for music creation, such as generating audio output based on a user's interactive movements (e.g., moving hands, feet, arms, and / or legs). In some examples, a user may generate audio output by dancing in front of a camera on a computing device (e.g., a mobile phone), where the computing device maps the user's interactive movements to a sequence of audio element identifiers and then converts the sequence of audio element identifiers into a musical sequence and audio output. In other examples, a user may perform a facial expression (e.g., making a "happy" face) or a facial expression element (e.g., blinking, grinning, smiling) as an interactive movement. Using a user's interactive movements instead of keystrokes on a piano or fingering on a guitar significantly reduces the challenge of mastering instrument skills, making composition accessible to those who are not proficient on an instrument. In some examples, the computing device includes a music theory engine that generates musical sequences by performing music theory rule enforcement, e.g., by modifying sound element identifiers to improve their musicality. In this way, the music theory engine significantly reduces the demand for proficiency in music theory, further increasing accessibility to composition.
[0020] This and many other embodiments of computing devices are described herein. For example, FIG. 1 is a block diagram of an example system 100 for providing an interactive motion voice engine for generating audio output according to examples of the present disclosure. System 100 includes computing device 110 and computing device 120 communicatively coupled via network 150. Computing device 110 may be a smartphone, a mobile computer, or a mobile computing device (e.g., a Microsoft® Surface® device, a laptop computer, a notebook computer, an Apple® iPad®, etc.). TM Computing device 110 may be any type of computing device, including a tablet computer (e.g., a laptop, a netbook, etc.), or a fixed computing device such as a desktop computer or PC (personal computer). In some examples, computing device 110 is a user device or client device of user 102, and computing device 120 is a server device. In some examples, computing device 120 is a network server, a cloud server, or other suitable distributed computing system. Computing device 120 may be operated and / or maintained by a social media platform, a cloud processing provider, a software-as-a-service provider, or other suitable entity. Computing device 110 and / or computing device 120 may be configured to run one or more software applications (or “applications”) and / or services and / or manage hardware resources (e.g., processor, memory, etc.) that may be used by a user of computing device 110.
[0021] The computing device 110 includes an image sensor 112, a depth sensor 114, a user input processor 116, and a display 118. The image sensor 112 is configured to capture images and / or video of the user 102, for example, as the user 102 makes interactive movements. The image sensor 112 may be, for example, a front-facing "selfie" camera or a rear-facing camera of a smartphone. In various examples, images or still frames of the video may be used to identify facial expression elements, gestures, and user element movements performed by the user 102, as described below. The depth sensor 114 is configured to estimate the distance between the computing device 110 and the user 102, for example, to estimate the distance to the user's 102's hands, arms, legs, and / or head. The depth sensor 114 may provide depth information that enhances images captured by the image sensor 112, thereby enabling estimation of the user's 102's three-dimensional position.
[0022] The user input processor 116 is configured to identify interactive movements of the user 102 based on the images captured by the image sensor 112 and / or the depth information from the depth sensor 114. Furthermore, the user input processor 116 determines a user input identifier corresponding to the interactive movement. Advantageously, the user 102 does not need to use a touchscreen, mouse, keyboard, or other physical input device to provide user input.
[0023] Typically, the user input identifier is a discrete identifier, such as an integer or other suitable value, that uniquely identifies an interactive movement previously performed by the user 102. In various examples, the user input processor 116 may include one or more of a gesture processor that identifies gestures performed by the user 102, a facial expression processor that identifies facial expressions performed by the user 102, and / or a finger position processor that identifies finger positions of the user 102. Further details of the user input processor 116 are provided below with reference to FIG.
[0024] The display 118 is configured to present a user interface of the computing device 110. In various examples, the display 118 is a touchscreen display of a smartphone, a monitor of a desktop computer, etc. The display 118 displays the image captured by the image sensor 112. The image input may be configured to display an output image including a graphical user interface overlaid on the image input. For example, the output image may provide the user 102 with real-time feedback about their interactive movements.
[0025] The computing device 120 includes a sound element processor 122, a music theory engine 124, a synthesizer 126, an effects engine 128, and a beat quantizer 130. The sound element processor 122 is configured to map user-input identifiers representing user interactive movements to a sequence of sound element identifiers. The sound element identifiers are data structures representing musical notes, samples, loops, and / or timbres for musical instruments (e.g., piano, acoustic guitar, trumpet). In some examples, the sound element identifiers are in Musical Instrument Digital Interface (MIDI) format. For example, the sound element identifiers include information about pitch, velocity, vibrato, panning, timing clock signals, etc. For example, a phonetic element identifier for middle C on a piano (fundamental frequency approximately 261.63 Hz using the A440 pitch standard) may have an integer value pitch that is 60 with a value range from 0 to 127 (e.g., two octaves below middle C has an integer value of 36, and D above middle C has an integer value of 62). In other examples, different identifiers are used for pitch (e.g., absolute or relative pitch between notes), note duration, volume, etc.
[0026] In some examples, the audio element processor 122 maps a single user input identifier to an audio element identifier for a musical note. In other words, a single interactive movement (e.g., a nod) is mapped to a single musical note with a start time and an end time. In other examples, the audio element processor 122 maps a first user input identifier to an audio element identifier for a note start (i.e., a MIDI note-on event) and a second user input identifier to an audio element identifier for a note end (i.e., a MIDI note-off event). In some examples, the same user input identifier is alternately mapped to note-on and note-off events. In other examples, a subsequent different user input identifier is mapped to both a note-off of the previous note and a note-on of the current note.
[0027] Mapping a user input identifier to a sequence of audio element identifiers may include generating a single audio element identifier for a single user input identifier, or generating multiple audio element identifiers for a single user input identifier. For example, user 102 may hold one finger over an icon in a graphical user interface for a "one shot" of an audio sample, hold two fingers over an icon for a loop of the audio sample, etc.
[0028] In some examples, mapping the interactive movements includes selecting a predetermined set of instruments from a plurality of sets of instruments and mapping the interactive movements to instruments within the selected predetermined set of instruments. For example, the audio element processor 122 may generate the plurality of sets of instruments using a neural network engine that identifies a predetermined set of instruments from existing music samples (e.g., popular published music). In one example, each set of a predetermined group of instruments is selected so that the instruments sound pleasing together. The sample set may include percussion instruments such as a snare drum, kick drum, hi-hat, bass guitar, and overdriven electric guitar for a blues-style set of instruments; a violin, viola, and cello for a chamber quartet-style set of instruments; and a drum machine, keyboard, and drum audio samples for a bass-style set of instruments. The audio element processor 122 may include other sets of instruments for different musical styles, such as country music, electronic music, hip-hop music, jazz music, Latin music, pop music, rock music, and metal music. Conversely, a less ideal set of instruments might include a banjo, slide whistle, distorted electric guitar, and drum set.
[0029] The music theory engine 124 is configured to perform music theory rule enforcement on the sequence of audio element identifiers to generate a musical sequence. A musical sequence is also a sequence of audio element identifiers, but more likely to have a high level of musicality. Generally, the music theory engine 124 enforces rules that improve the musicality of the audio output 104 for audio output that is strictly based on the sequence of audio element identifiers. For example, the music theory engine 124 enforces rules to reduce dissonances, change “bad” or “incorrect” notes to “good” notes (i.e., change notes that are not in the current chord to be in the current chord, and change notes that are not in the current key signature to be in the current key signature), omit incorrect notes, or insert additional notes. The music theory engine 124 may enforce a rule by modifying at least one audio element identifier in the sequence of audio element identifiers that violates a music theory rule. Exemplary modifications include changing the pitch associated with at least one sound element identifier (e.g., matching a chord progression, scale, musical mode), omitting a sound element identifier (e.g., erasing a "bad" note), changing the duration of a sound element identifier, or changing other characteristics associated with the sound element identifier.
[0030] The music theory engine 124 may include one or more selectable rules that enforce various elements of music theory, such as maintaining the consistency of notes in a melody with the key signature, maintaining the harmony of notes in chords, preserving notes in chord progressions, ensuring that chords with dissonant intervals are followed by chords with consonant intervals, etc.
[0031] Synthesizer 126 is configured to convert musical sequences into audio output using synthesized or sampled audio and may be implemented, for example, by a library of MIDI output routines on computing device 120 as part of a sound card driver. In another example, synthesizer 126 is a keyboard or other suitable MIDI device that converts musical sequences into audio output. The audio output may be provided as a WAV file, an MP3 file, or other suitable audio output file.
[0032] The effects engine 128 is configured to add musical effects, such as reverb, chorus, delay, overdrive, distortion, filter cutoff, envelope filter, flange, tremolo, or other suitable effects, to the audio output. In some examples, the effects engine 128 adds a drum backing track, loop, or other sampled sound to the audio output.
[0033] The beat quantizer 130 is configured to modify the temporal characteristics (i.e., start and stop times) of the individual sound element identifiers of the musical sequence. Typically, the beat quantizer 130 is configured to align the individual sound element identifiers to a time signature, a predetermined rhythmic pattern, or other suitable temporal characteristics before the synthesizer 126 generates the audio output. For example, the beat quantizer 130 may align all sound element identifiers of the musical sequence to eighth notes in a 4 / 4 time signature. In some examples, the beat quantizer 130 randomly or pseudo-randomly selects a predetermined rhythmic pattern from a plurality of predetermined rhythmic patterns.
[0034] In some examples, one or more components of computing device 110 may be omitted or moved to other devices. In one example, depth sensor 114 is omitted from computing device 110. In another example, user input processor 116 is located in computing device 120 and omitted from computing device 110. In this example, computing device 120 receives image input from computing device 110. In yet another example, effects engine 128 is omitted from computing device 120. In another example, audio element processor 122, music theory engine 124, synthesizer 126, effects engine 128, and beat quantizer 130 are located in computing device 110 and omitted from computing device 120. In this example, computing device 120 and network 150 are omitted, and processing to generate audio output may be performed by computing device 110.
[0035] 2 is a block diagram of an example user input processor 200 of a computing device according to examples of the present disclosure. In some examples, the user input processor 200 generally corresponds to the user input processor 116 and generates user input identifiers as discrete representations of interactive movements from the user 102. The user input processor 200 includes a gesture processor 210, a facial expression processor 220, and a finger position processor 230.
[0036] The gesture processor 210 is configured to identify gestures performed by the user 102 and generate a corresponding user input identifier. Exemplary gestures may include a shrug, a pointing gesture, a nod, a wave, a clap, a wave of the arms, a foot lift, a jump, an arm raised to a predetermined position or angle, or other suitable gestures, and may be mapped to an integer identifier (e.g., 0 for shrug, 1 for pointing, 2 for nodding, ...). Similarly, the facial expression processor 220 is configured to identify facial expressions performed by the user 102 and generate a corresponding user input identifier. Exemplary facial expressions may include a smile, a frown, a surprised face, an angry face, or other suitable facial expressions. In some examples, the user input processor 200 identifies changes in the user 102 over time, such as when the user's 102 face moves in a circular motion or rotates (i.e., when the user 102 turns around). In some examples, the gesture processor 210 identifies only complete expressions for a single user input identifier. In other examples, the gesture processor 210 identifies partial expressions as facial expression elements corresponding to different user input identifiers, such as one eye (or both eyes) open, one eye (or both eyes) closed, winking with one eye, blinking, open or closed mouth, grinning with a partial mouth, etc.
[0037] The finger position processor 230 is configured to identify the current position of one or more fingers of the user 102. If the current position of a finger overlaps a portion of the graphical user interface, the finger position processor 230 may provide a user input identifier corresponding to the overlapping portion of the graphical user interface. In other examples, the user input processor 200 also identifies the movement of an object that overlaps the portion of the graphical user interface, such as waving a drumstick (or other implement that can mimic a drumstick, e.g., a pencil or pen). In some examples, the finger position processor 230 maps the finger position to a slider input or a pitch wheel. Ping and generate the corresponding user interface identifier.
[0038] In some examples, the user input processor 200 is configured to identify interactive movements that mimic the use of musical instruments such as "air guitar" or "air drumming." In the case of air guitar, the user input processor 200 may identify the user's 102 strumming hand to select note start and stop and another hand fretting notes or chords to select pitches on the virtual guitar. Additionally, the user input processor 200 may recognize the amount of mouth opening of the user 102 and map that amount to an effect, for example, a "wah" effect, where an open mouth corresponds to a forward pedal position of a wah pedal and a closed mouth corresponds to a backward pedal position of the wah pedal. In another example, the user input processor 200 maps an open mouth to a different effect, such as a volume, reverb, or delay level, that corresponds to an expression pedal. In the case of air drumming, the user input processor 200 may map the lowest point of the drumstick swing motion's travel to the start of a musical note, and map the lateral and / or depth position to a pitch (e.g., a xylophone key or cymbal pitch) or instrument (e.g., a hi-hat or snare drum).
[0039] Although the description herein is of a single user 102, the user input processor 200 may be configured to identify multiple users simultaneously. In some examples, the audio elements processor 122 may assign different users to different instruments within a given set of instruments (e.g., one user to acoustic guitar, one user to piano, and one user to percussion).
[0040] In some examples, one or two of processors 210, 220, and / or 230 may be omitted, for example, to simplify the operation of user input processor 200 and / or reduce power consumption of computing device 110. In other examples, two or more of processors 210, 220, and / or 230 may be combined with each other or with other elements of system 100. In one such example, processors 210, 220, 230 and sound element processor 122, music theory engine 124, synthesizer 126, effects engine 128, and beat quantizer 130 are implemented as a single processor.
[0041] FIG. 3 illustrates an exemplary output image 300 for a graphical user interface 302 according to an example of the present disclosure. In some examples, the computing device 110 displays the output image 300 on the display 118. The output image 300 includes a graphical user interface 302 that is overlaid with image input from the image sensor 112 and provides the user 102 with visual feedback about their interactive movements. The graphical user interface 302 includes multiple icons representing buttons, triggers, switches, sliders, or other user interface elements with which the user 102 can virtually interact. As an example, the user 102 may move various user elements (e.g., fingers, hands, arms, feet, legs, head, or objects held in hand) to overlap with the icons of the graphical user interface 302. In the example shown in FIG. 3 , the user 102 has a right hand 310 and a left hand 312, and the user 102 performs interactive movements using the right hand 310 and the left hand 312.
[0042] Graphical user interface 302 includes icons 320, 322, 324, 326, 330, 332, 334, and 336 as buttons, and icons 340 and 342 as slider elements. In other examples, icons such as dials, momentary switches, latching switches, or other suitable icons may be implemented within graphical user interface 302.
[0043] 3, user 102 "presses" icon 332 with left hand 312. The icons in graphical user interface 302 may be mapped to respective predefined sound element identifiers, for example, different notes on an instrument, parts of a drum kit, etc. The icons in graphical user interface 302 may also be mapped to discrete values, for example, volume levels, effect levels, frequency levels, etc. For example, icon 340 may correspond to a volume slider having a range of discrete values from 0 to 100, and the relative position of icon 342 within icon 340 may correspond to a discrete value (e.g., a volume level of 40).
[0044] 4A is a diagram illustrating an example sequence 400 of image inputs with facial expressions for generating audio output according to an example of the present disclosure. The user input processor 200 may receive the sequence 400, identify various facial expressions or portions of facial expressions (e.g., facial expression elements) in the sequence 400 (e.g., using the facial expression processor 220), and map the facial expressions and / or facial expression elements to different audio element identifiers.
[0045] 4A , user 102 executes facial expressions including a neutral expression 410 (e.g., neutral eyes, neutral brows, and closed mouth), a smiling expression 412, a neutral expression 414, an excited expression 416, and a neutral expression 418. In one example, user input processor 200 maps neutral expressions 410, 414, and 418 to a MIDI all notes off or "rest" identifier, maps smiling expression 412 to an audio element identifier corresponding to a first tuba sound, and maps excited expression 416 to a second tuba sound. In this example, beat quantizer 130 may time the various expressions to eighth notes within a 2 / 4 time signature, and synthesizer 126 may generate audio output that generally corresponds to a polka "oom-pah" rhythm played on the off-beats on a tuba.
[0046] 4B illustrates an exemplary image input 460 with facial expression elements for generating an audio output according to an example of the present disclosure. Facial expression elements 420, 422, and 424 correspond to the user 102 winking with the left eye, winking with the right eye, and blinking, respectively. Facial expression elements 430, 432, 434, and 436 correspond to a grinning upwards, a grinning upwards, a grinning downwards, and a grinning downwards, respectively. Facial expressions 440, 442, 444, 446, and 448 include various facial expression elements, such as raised eyebrows, frowning, grimacing, smiling, and an open mouth. Facial expression elements 450 and 452 correspond to the user 102 tilting his head to the left and right, respectively. Although this specification describes several facial expressions and facial expression elements that can be mapped to different audio element identifiers, in other examples, the user input processor 200 may be configured to recognize other combinations of facial expression elements, facial expressions, or other inputs.
[0047] FIG. 5 is a flowchart of an exemplary method 500 for generating audio output according to examples of the present disclosure. Unless otherwise indicated, the technical processes depicted in these figures are performed automatically. In any given embodiment, some steps of the processes may be repeated and operate using different parameters or data. Steps in an embodiment may also be performed in an order different from the top-to-bottom order presented in FIG. 5. Steps may be performed sequentially, partially overlapping, or entirely in parallel. Thus, the order in which steps of method 500 are performed may differ from one implementation of the process to another. Steps may be omitted, combined, renamed, regrouped, performed on one or more machines, or deviate from the illustrated flow, as long as the performed process is operable and consistent with at least one claim. The steps of FIG. 5 may be performed by computing device 110 (e.g., via user input processor 116, display 118), computing device 120 (e.g., via music theory engine 124, synthesizer 126, effects engine 128, or beat quantizer 130), or other suitable computing device.
[0048] Method 500 begins with step 502, in which image input of an interactive movement by a user is received as captured by an image sensor. In some examples, the image input corresponds to an image captured by image sensor 112 from user 102, such as the images shown in Figures 3, 4A, and / or 4B. For example, the image input may be received by user input processor 116 or user input processor 200.
[0049] In step 504, the interactive movements are mapped to a sequence of audio element identifiers. In some examples, the audio element processor 122 maps the interactive movements to a sequence of audio element identifiers, as described above. For example, the audio element processor 122 is configured to map a user-input identifier representing a user's interactive movements to a sequence of audio element identifiers. Audio element identifiers are data structures that represent musical notes, samples, loops, and / or timbres for musical instruments (e.g., piano, acoustic guitar, trumpet). In some examples, the audio element identifiers are in MIDI format. For example, the audio element identifiers include information about pitch, velocity, vibrato, panning, timing clock signals, etc.
[0050] In some examples, the audio element processor 122 maps a single user input identifier to an audio element identifier for a musical note. In other words, a single interactive movement (e.g., a nod) is mapped to a single musical note including a start time and an end time. In other examples, the audio element processor 122 maps a first user input identifier to an audio element identifier for a note onset (i.e., a MIDI note-on event) and maps a second user input identifier to an audio element identifier for a note end (i.e., a MIDI note-off event). In some examples, the same user input identifier is alternately mapped to note-on and note-off events. In other examples, a subsequent different user input identifier is mapped to both a note-off of the previous note and a note-on of the current note. In some examples, mapping the interactive movement includes selecting a predetermined set of instruments from a plurality of instrument sets and mapping the interactive movement to an instrument within the selected predetermined set of instruments.
[0051] In step 506, the sequence of voice element identifiers is processed to generate a musical sequence by performing music theory rule enforcement on the sequence of voice element identifiers. In some examples, as described above, the music theory engine 124 processes the sequence of voice element identifiers from the voice element processor 122 to generate a musical sequence. By enforcing the rules, the music theory engine 124 may reduce dissonances, change "bad" or "incorrect" notes to "good" notes (i.e., change notes that are not in the current chord to be in the current chord, and change notes that are not in the current key signature to be in the current key signature), omit incorrect notes, or insert additional notes. The music theory engine 124 may enforce the rules by modifying at least one voice element identifier in the sequence of voice element identifiers that violates a music theory rule. Exemplary modifications include changing the pitch associated with at least one sound element identifier (e.g., to match a chord progression, scale, musical mode), omitting a sound element identifier (e.g., to eliminate "bad" notes), changing the duration of a sound element identifier, or changing other characteristics associated with the sound element identifier. The music theory engine 124 may include one or more selectable rules that enforce various elements of music theory, such as maintaining the consistency of notes in a melody with a key signature, maintaining the harmony of notes in a chord, preserving notes in a chord progression, ensuring that chords with dissonant intervals are followed by chords with consonant intervals, etc.
[0052] At step 508, an audio output representing the musical sequence is generated. In some examples, the synthesizer 126 generates the audio output based on the musical sequence provided by the music theory engine 124. In some examples, generating the musical sequence further includes adding effects to the audio output (e.g., by the effects engine 128) and / or aligning the audio element identifiers to a time signature, rhythm, and / or chord progression (e.g., by the beat quantizer 130).
[0053] 6, 7, and 8 and the associated description provide descriptions of various operating environments in which various aspects of the present disclosure may be implemented. However, the devices and systems shown and described with reference to FIGS. 6, 7, and 8 are for purposes of example and illustration and are not intended to limit the numerous computing device configurations that may be used to implement various aspects of the present disclosure as described herein.
[0054] 6 is a block diagram illustrating the physical components (e.g., hardware) of a computing device 600 that can be used to implement aspects of the present disclosure. The computing device components described below may have computer-executable instructions for implementing an audio output generation application 620 on a computing device (e.g., computing device 110, computing device 120), including computer-executable instructions for an audio output generation application 620 that can be executed to implement methods disclosed herein. In a basic configuration, computing device 600 may include at least one processing unit 602 and system memory 604. Depending on the configuration and type of computing device, system memory 604 may include, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memory. The system memory 604 may include an operating system 605 and one or more program modules 606 suitable for executing an audio output generation application 620, such as one or more components related to Figures 1 and 2, in particular a user input processor 621 (e.g., corresponding to user input processor 116 or user input processor 200), an audio element processor 622 (e.g., corresponding to audio element processor 122), a music theory engine 623 (e.g., corresponding to music theory engine 124), a synthesizer 624 (e.g., corresponding to synthesizer 126), an effects processor 625 (e.g., corresponding to effects engine 128), and a beat quantizer 626 (e.g., corresponding to beat quantizer 130).
[0055] For example, operating system 605 may be adapted to control the operation of computing device 600. Additionally, embodiments of the present disclosure may be implemented in conjunction with a graphics library, other operating systems, or any other application program, but are not limited to any particular application or system. This basic configuration is illustrated in FIG. 6 by those components within dashed line 608. Computing device 600 may have additional features or functionality. For example, computing device 600 may further include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tape. Such additional storage is illustrated in FIG. 6 by removable storage device 609 and non-removable storage device 610.
[0056] As mentioned above, multiple program modules and data files may be stored in the system memory 604. The program modules 606 (e.g., audio output generation application 620), when executed on the processing unit 602, may perform processes including, but not limited to, aspects described herein. Other program modules usable in accordance with aspects of the present disclosure, particularly for generating audio output, may include a user input processor 621, an audio element processor 622, a music theory engine 623, a synthesizer 624, an effects processor 625, and a beat quantizer 626.
[0057] Furthermore, embodiments of the present disclosure may be implemented in electrical circuits including discrete electronic components, packaged or integrated electronic chips including logic gates, circuits utilizing a microprocessor, or a single chip including electronic components or a microprocessor. For example, embodiments of the present disclosure may be implemented via a system-on-chip (SOC) in which each or multiple components shown in FIG. 6 may be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communications units, system virtualization units, and various application functions, all integrated (or "baked") onto the chip substrate as a single integrated circuit. When operating via an SOC, the functionality described herein regarding the client's ability to switch protocols may operate via application-specific logic integrated with other components of the computing device 700 on a single integrated circuit (chip). Embodiments of the present disclosure may also be implemented using other technologies, including, but not limited to, mechanical, optical, fluidic, and quantum technologies, capable of performing logical operations such as AND, OR, and NOT. Additionally, embodiments of the present disclosure may be implemented within a general-purpose computer or any other circuit or system.
[0058] The computing device 600 may also have one or more input devices 612, such as a keyboard, mouse, pen, acoustic or voice input device, touch or swipe input device, etc. Output devices 614, such as a display, speakers, printer, etc., may also be included. The above devices are examples, and other devices may be used. The computing device 600 may include one or more communication connections 616 that enable communication with other computing devices 650. Examples of suitable communication connections 616 include, but are not limited to, radio frequency (RF) transmitter, receiver and / or transceiver circuitry, a universal serial bus (USB), a parallel port, and / or a serial port.
[0059] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, or program modules. System memory 604, removable storage device 609, and non-removable storage device 610 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article used to store information and accessible by computing device 600. Any such computer storage media may be part of computing device 600. Computer storage media do not include carrier waves or other propagated or modulated data signals.
[0060] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" may refer to a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0061] 7 and 8 illustrate a mobile computing device 700, such as a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., that can be used to implement embodiments of the present disclosure. In some aspects, a client may be a mobile computing device. Referring to FIG. 7, one aspect of a mobile computing device 700 for implementing these aspects is shown. In a basic configuration, the mobile computing device 700 is a handheld computer having both input and output elements. The mobile computing device 700 typically includes a display 705 and one or more input buttons 710 that allow a user to input information into the mobile computing device 700. The display 705 of the mobile computing device 700 can also function as an input device (e.g., a touchscreen display). An optional secondary input element 715 (if included) further enables additional user input. The secondary input element 715 may be a rotary switch, a button, or any other type of manual input element. In alternative aspects, the mobile computing device 700 may incorporate more or fewer input elements. For example, in some embodiments, the display 705 may not be a touchscreen. In yet another alternative embodiment, the mobile computing device 700 is a mobile telephone system, such as a cellular telephone. The mobile computing device 700 may include a front-facing camera 730. The mobile computing device 700 may further include an optional keypad 735. The optional keypad 735 may be a physical keypad or a "soft" keypad generated on the touchscreen display. In various embodiments, output elements include the display 705 for displaying a graphical user interface (GUI), a visual indicator 720 (e.g., a light-emitting diode), and / or an audio transducer 725 (e.g., a speaker).In some embodiments, the mobile computing device 700 incorporates vibration transducers to provide haptic feedback to the user. In yet other embodiments, the mobile computing device 700 incorporates input and / or output ports, such as an audio input (e.g., a microphone jack), an audio output (e.g., a headphone jack), and a video output (e.g., an HDMI port), to send and receive signals to and from external devices.
[0062] Figure 8 7 is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, mobile computing device 700 can implement several aspects by incorporating system (e.g., architecture) 802. In one embodiment, system 802 is implemented as a “smartphone” capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, games, and media client / player). In some aspects, system 802 is integrated as an integrated personal digital assistant (PDA), wireless telephone, or other computing device. System 802 may include a display 805 (similar to display 705), such as a touchscreen display, or other suitable user interface. System 802 may also include an optional keypad 835 (similar to keypad 735) and one or more peripheral ports 830, such as input and / or output ports for audio, video, control signals, or other suitable signals.
[0063] In some examples, system 802 may include a processor 860 coupled to memory 862. System 802 may also include a special-purpose processor 861, such as a neural network processor. One or more application programs 866 may be loaded into memory 862 and execute on or in association with operating system 864. Examples of application programs include a telephone dialing program, an email program, a personal information manager (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. System 802 also includes a non-volatile storage area 868 within memory 862. Non-volatile storage area 868 may be used to store persistent information that should not be lost when system 802 is powered off. Applications 866 may use and store information in non-volatile storage area 868, such as emails or other messages used by email applications. A synchronization application (not shown) also resides on system 802 and is programmed to interact with a corresponding synchronization application resident on the host computer to maintain synchronization between the information stored in non-volatile storage area 868 and corresponding information stored on the host computer.
[0064] The system 802 includes a power supply 870, which may be implemented as one or more batteries. The power supply 870 may also include an external power source, such as a powered storage base or AC adapter, that replenishes or recharges the batteries.
[0065] System 802 may further include a wireless interface layer 872 that performs the function of transmitting and receiving radio frequency communications. Wireless interface layer 872 facilitates wireless connectivity between system 802 and the "outside world" via a communications carrier or service provider. Transmissions to and from wireless interface layer 872 occur under the control of operating system 864. In other words, communications received by wireless interface layer 872 may be delivered to application programs 866 via operating system 864, and vice versa.
[0066] The visual indicator 820 can be used to provide a visual notification, and / or the audio interface 874 can be used to generate an audible notification via the audio transducer 725 (e.g., the audio transducer 725 shown in FIG. 7). In the illustrated embodiment, the visual indicator 820 is a light-emitting diode (LED), and the audio transducer 725 can be a speaker. These devices are directly coupled to the power source 870 so that when activated, they remain on for a duration specified by the notification mechanism, even if the processor 860 and other components are turned off to conserve battery power. The LED may be programmed to remain lit indefinitely until the user takes an action that indicates the device is powered on. The audio interface 874 is used to provide audible signals to and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 725, the audio interface 874 may be coupled to a microphone to receive audible input, for example, to facilitate a telephone conversation. According to embodiments of the present disclosure, the microphone can also function as an audio sensor to facilitate control of notifications, as described below. The system 802 may further include a video interface 876 that enables operation of peripheral devices 830 (eg, an on-board camera) to record still images, video streams, and the like.
[0067] Mobile computing device 700 implementing system 802 may have additional features or functionality. For example, mobile computing device 700 may further include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tape. Such additional storage is illustrated in FIG. 8 by non-volatile storage 868.
[0068] As described above, data / information generated or captured by mobile computing device 700 and stored via system 802 may be stored locally on mobile computing device 700, or the data may be stored in any number of storage media accessible to the device via wireless interface layer 872 or via a wired connection between mobile computing device 700 and another computing device associated with mobile computing device 700 (e.g., a server computer in a distributed computing network, such as the Internet). It should be understood that such data / information may be accessed via mobile computing device 700, via wireless interface layer 872, or via a distributed computing network. Similarly, such data / information may be readily transmitted between computing devices for storage and use in accordance with known data / information transmission and storage means, including email and collaborative data / information sharing systems.
[0069] It should be noted that Figures 7 and 8 are provided to illustrate the present method and system and are not intended to limit the present disclosure to any particular sequence of steps or combination of hardware or software components.
[0070] The terms "at least one," "one or more," "or," and "and / or" are open-ended expressions that are conjunctive and disjunctive in operation. For example, each of the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," "A, B, and / or C," and "A, B, or C" means A only, B only, C only, A and B, A and C, B and C, or A, B, and C.
[0071] The term "an" or "an" entity means one or more of that entity. Thus, the terms "one," "one or more," and "at least one" may be used interchangeably herein. It should also be noted that the terms "comprise," "include," and "have" may be used interchangeably.
[0072] As used herein, the term "automatic" and variations thereof refer to any process or operation that is performed without significant manual input, usually continuous or semi-continuous. However, performance of a process or operation can be automatic with significant or insignificant manual input, provided the input is received prior to performance of the process or operation. Manual input is considered significant input if it affects how the process or operation is performed. Manual input to consent to the performance of a process or operation is not considered "substantial."
[0073] Any of the steps, functions, and operations discussed herein may be performed sequentially and automatically.
[0074] The exemplary systems and methods of the present disclosure have been described in connection with computing devices. However, to avoid unnecessarily obscuring the present disclosure, the foregoing description omits some known structures and devices. This omission should not be construed as limiting. Specific details are set forth to provide an understanding of the present disclosure. However, it should be understood that the present disclosure may be practiced in a variety of ways in addition to the specific details set forth herein.
[0075] Furthermore, while the exemplary embodiments illustrated herein show various components of the system as being co-located, some components of the system may be located remotely in a distal portion of a distributed network, such as a LAN and / or the Internet, or may be located within a dedicated system. Accordingly, it should be understood that components of the system may be coupled to one or more devices, e.g., a server, a communications device, or may be co-located on a particular node of a distributed network, such as an analog and / or digital telecommunications network, a packet-switched network, or a circuit-switched network. As will be appreciated from the foregoing description, for reasons of computing efficiency, components of the system may be located anywhere within the distributed network of components without impacting the operation of the system.
[0076] Furthermore, it should be understood that the various links connecting the elements may be wired or wireless links, or any combination thereof, or any other known or future-developed elements capable of providing data to and / or communicating data from the connected elements. These wired or wireless links may also be secure links and may be capable of communicating encrypted information. For example, the transmission medium used as the link may be any suitable carrier of an electrical signal, including coaxial cable, copper wire, and optical fiber, or may take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0077] Although the flowcharts have been discussed and illustrated with reference to particular sequences of events, it should be understood that modifications, additions, and omissions may be made to the sequences without materially affecting the operation of the disclosed configurations and aspects.
[0078] Several variations and modifications of the present disclosure may be used, and it is possible to provide some features of the present disclosure and not other features.
[0079] In other settings, the systems and methods of the present disclosure may be implemented in combination with special purpose computers, programmed microprocessors or microcontrollers and peripheral integrated circuit elements, ASICs or other integrated circuits, digital signal processors, hardwired electronic or logic circuits (e.g., discrete element circuits), programmable logic devices or gate arrays (e.g., PLDs, PLAs, FPGAs, PALs), special purpose computers, any similar devices, etc. Overall, any device or means capable of implementing the methods set forth herein can be used to implement various aspects of the present disclosure. Exemplary hardware that can be used in the present disclosure is: Hardware includes computers, handheld devices, telephones (e.g., cellular telephones, Internet-enabled telephones, digital telephones, analog telephones, hybrid telephones, etc.), and other hardware known in the art. Some of these devices include processors (e.g., single or multiple microprocessors), memory, non-volatile storage, input devices, and output devices. Alternative software implementations may also be constructed to implement the methods described herein, including, but not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing.
[0080] In yet another configuration, the disclosed methods may be readily implemented in combination with software using object or object-oriented software development environments that provide portable source code usable on a variety of computer or workstation platforms. Alternatively, the disclosed systems may be implemented partially or fully in hardware using standard logic circuits or VLSI designs. Whether software or hardware is used to implement a system according to the present disclosure depends on the speed and / or efficiency requirements of the system, the particular functionality, and the particular software or hardware system or microprocessor or microcomputer system being used.
[0081] In yet another configuration, the disclosed methods may be partially implemented in software that can be stored on a storage medium and executed on a programmed general-purpose computer, special-purpose computer, microprocessor, etc., in cooperation with a controller and memory. In these instances, the disclosed systems and methods can be implemented as programs, e.g., applets, JAVA, or CGI scripts, installed on a personal computer, as resources resident on a server or computer workstation, as routines installed in dedicated measurement systems, system components, etc. A system can also be realized by physically incorporating the system and / or method into a software and / or hardware system.
[0082] This disclosure is not limited to the standards and protocols described. Other similar standards and protocols not described herein already exist and are included in this disclosure. Furthermore, the standards and protocols described herein, and other similar standards and protocols not described herein, are periodically replaced by faster or more efficient equivalents having substantially the same functionality. Such replacement standards and protocols having the same functionality are considered equivalents included in this disclosure.
[0083] In various configurations and aspects, the present disclosure includes components, methods, processes, systems, and / or devices as shown and described herein, including various combinations, subcombinations, and subsets thereof. Those skilled in the art, once they understand the present disclosure, will understand how to make and use the systems and methods of the present disclosure. In various configurations and aspects, the present disclosure includes providing devices and processes in the absence of items not shown and / or described herein or in its various configurations or aspects, including the absence of items that may have been used in previous devices or processes to, for example, improve performance, provide ease of use, and / or reduce implementation costs.
[0084] The present disclosure relates to systems and methods for generating audio output, at least according to the examples provided in the following sections.
[0085] (A1) In one aspect, some examples include a method for generating an audio output, the method including receiving image input of interactive movements by a user captured by an image sensor, mapping the interactive movements to a sequence of audio element identifiers, processing the sequence of audio element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of audio element identifiers, and generating an audio output representing the musical sequence.
[0086] (A2) In some examples of A1, processing the sound element identifiers includes modifying at least one sound element identifier in the sequence of sound element identifiers that violates a music theory rule, and generating the musical sequence based on the modified sound element identifier.
[0087] (A3) In some examples of A1 to A2, modifying the at least one sound element identifier includes changing a pitch associated with the sound element identifier.
[0088] (A4) In some examples of A1-A3, changing the pitch includes matching a chord progression that satisfies the music theory rules.
[0089] (A5) In some examples of A1 to A4, modifying the at least one sound element identifier includes omitting the at least one sound element identifier when generating the music sequence.
[0090] (A6) In some examples of A1 to A5, modifying the at least one sound element identifier includes changing a duration of the at least one sound element identifier.
[0091] (A7) In some examples of A1-A6, mapping the interactive movements includes selecting a predetermined set of instruments from a plurality of sets of instruments, and mapping the interactive movements to instruments within the selected predetermined set of instruments.
[0092] (A8) In some examples of A1-A7, the method further includes generating the plurality of sets of instruments using a neural network engine that identifies a set of predetermined instruments from a music sample.
[0093] (A9) In some examples of A1-A8, the method further includes displaying to the user an output image including a graphical user interface overlaid on the image input, and the interactive movement includes a user element of the user overlapping with the graphical user interface.
[0094] (A10) In some examples of A1 to A9, the user element of the user is the user's finger, hand, arm, foot, and / or leg.
[0095] (A11) In some examples of A1-A10, the graphical user interface includes a plurality of icons corresponding to a plurality of predetermined sound element identifiers, and mapping the interactive movements includes mapping an interactive movement having a user element overlapping an icon to a predetermined sound element identifier corresponding to the overlapped icon.
[0096] (A12) In some examples of A1 to A11, the plurality of predetermined voice element identifiers include a single element identifier and a multi-element identifier.
[0097] (A13) In some examples of A1 to A12, the interactive movement is a facial expression element performed by the user.
[0098] (A14) In some examples of A1 to A13, the interactive movement is a gesture performed by the user.
[0099] In yet another aspect, some examples include a computing system comprising one or more processors and a memory coupled to the one or more processors, the memory storing a plurality of instructions, the one or more instructions, when executed by the one or more processors, causing the one or more processors to perform any of the methods described herein (e.g., A1-A14 above).
[0100] In yet another aspect, some examples include a non-transitory computer-readable storage medium storing one or more programs for execution by one or more processors of a storage device, the one or more programs including instructions for performing any of the methods described herein (e.g., A1-A14 above).
[0101] In yet another aspect, some examples include a computing system comprising one or more processors and a memory coupled to the one or more processors, the memory storing a plurality of instructions, the one or more instructions, when executed by the one or more processors, causing the one or more processors to perform any of the methods described herein (e.g., method 500 above).
[0102] In yet another aspect, some examples include a non-transitory computer-readable storage medium storing one or more programs for execution by one or more processors of a storage device, the one or more programs including instructions for performing any of the methods described herein (e.g., method 500 above).
[0103] For example, aspects of the present disclosure are described above with reference to block diagrams and / or operational descriptions of methods, systems, and computer program products according to aspects of the present disclosure. The functions / acts noted in the blocks may occur out of the order shown in any flowchart. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functions / acts involved.
[0104] The description and illustration of one or more aspects provided herein are not intended to restrict or limit the scope of the claimed disclosure in any manner. The aspects, examples, and details described herein are deemed sufficient to convey proprietary rights and to enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to the aspects, examples, or details described herein. Various features (structural features and method features), whether illustrated or described in combination or individually, are intended to be selectively included or omitted to form embodiments having particular feature sets. The description and illustrations provided herein will enable those skilled in the art to envision changes, modifications, and alternative embodiments within the spirit of the broader aspects of the general inventive concept embodied herein, without departing from the broader scope of the claimed disclosure. [Explanation of symbols]
[0105] 104 Audio Output 110 Computing Devices 112 Image Sensor 114 Depth Sensor 116 User Input Processor 118 Display 120 Computing Devices 122 Audio Element Processor 124 Music Theory Engine 126 Synthesizer 128 Effects Engine 130 Beat Quantizer 150 Network 200 User Input Processor 210 Gesture Processor 220 Facial Expression Processor 230 Finger Position Processor 502 receives an image input of an interactive movement by a user captured by an image sensor. 504 Mapping interactive actions to sequences of audio element identifiers 506. Process a sequence of sound element identifiers to generate a musical sequence by performing music theory rule enforcement on the sequence of sound element identifiers. 508 Generate audio output representing a musical sequence 600 computing devices 602 Processing Unit 604 system memory 605 Operating Systems 606 Program Module 609 Removable Storage Devices 610 Non-removable storage devices 612 Input Device 614 Output Device 616 Communication Connections 620 Applications 621 User Input Processor 622 Audio Element Processor 623 Music Theory Engine 624 Synthesizer 625 Effects Processor 626 Beat Quantizer 650 Other Computing Devices 861 dedicated processor 860 processor 805 Display 830 Peripheral Port 835 keypad 862 memory 866 App 868 Storage device 870 Power supply 876 Video Interface 874 Voice Interface 872 Radio Interface Layer
Claims
1. 1. A method for generating an audio output, comprising: receiving an image input of an interactive movement by a user captured by an image sensor; mapping said interactive movements to a sequence of audio element identifiers; processing said sequence of sound element identifiers to generate a musical sequence by performing music theory rule enforcement on said sequence of sound element identifiers; generating an audio output representing the musical sequence; Including, Mapping the interactive movement includes: selecting a predetermined set of instruments from a plurality of sets of instruments; mapping said interactive movements to instruments within a selected set of predetermined instruments; Including, The method includes generating the plurality of sets of instruments using a neural network engine that identifies a set of predetermined instruments from a musical sample; further comprising: method.
2. Processing the speech element identifier includes: modifying at least one sound element identifier in said sequence of sound element identifiers that violates a music theory rule; generating the musical sequence based on the modified sound element identifiers; 10. The method of claim 1, comprising:
3. modifying the at least one sound element identifier includes changing a pitch associated with the sound element identifier; 3. The method of claim 2, comprising:
4. changing the pitches to match a chord progression that satisfies the music theory rules; 4. The method of claim 3, comprising:
5. modifying the at least one sound element identifier includes omitting the at least one sound element identifier when generating the musical sequence. The method of claim 2.
6. modifying the at least one sound element identifier includes changing a duration of the at least one sound element identifier; 3. The method of claim 2, comprising:
7. the method further comprising displaying to the user an output image including a graphical user interface overlaid on the image input, the interactive movement including a user element of the user overlapping with the graphical user interface; The method of claim 1.
8. the user element of the user is the user's finger, hand, arm, foot, and / or leg; The method of claim 7.
9. the graphical user interface includes a plurality of icons corresponding to a plurality of predetermined sound element identifiers; and mapping the interactive movement includes mapping an interactive movement having a user element overlapping an icon to a predetermined audio element identifier corresponding to the overlapped icon. The method of claim 7.
10. the plurality of predetermined voice element identifiers include single element identifiers and multi-element identifiers; 10. The method of claim 9.
11. the interactive movement is a facial expression element performed by the user; The method of claim 1.
12. The interactive movement is a gesture performed by the user. The method of claim 1.
13. A system for generating audio output, the system comprising one or more hardware processors configured with machine-readable instructions to carry out the method of any of claims 1-12.
14. A non-transitory computer-readable storage medium containing instructions executable by one or more processors, comprising: A non-transitory computer-readable storage medium, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method of any of claims 1-12.
15. A computer program product which, when executed by a computer, causes the computer to implement a method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Device and program for automatic music composition
JP2002311951A