Methods, systems, and instructions for dynamic music creation

By analyzing the mapping of music components and emotions, combining the characteristics of characters and scenes, dynamically generate game music, the problem that music in the existing technology cannot dynamically respond to player behavior, and a more immersive and personalized music experience is achieved.

CN120053986APending Publication Date: 2025-05-30SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411653221.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-11-15
Filing Date
2019-11-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The music in existing games is mainly composed of pre-written clips, which makes it difficult to dynamically respond to players' actions and emotions, resulting in the music experience not immersive and personalized enough.

Method used

Dynamic music is generated by analyzing music components (rhythm, time sign, melody structure, mode, harmonic structure, harmonic density, rhythm density, tone density) and mapping them with emotions expressed by human comments and social media. The system assigns emotions to musical thought groups based on character types, music composition components, or scenes, and generates music through vector mapping and fader control.

Benefits of technology

It realizes the dynamic creation of music in the game, which can better respond to players' movements and emotions, and enhances the immersion and personalization of the music experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120053986A_ABST
    Figure CN120053986A_ABST
Patent Text Reader

Abstract

Methods, systems, and instructions for dynamic music creation are disclosed. The method comprises the following steps: a) according to a role type, a music composition component or a scene, assigning an emotion to a group of one or more music ideas; b) associating a vector with the emotion; c) mapping a group of the one or more musical concepts to the vector based on the emotion; d) generating a music composition based on the vector and one or more desired emotions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention with the application date of November 7, 2019, application number 201980075286.0, and invention title "Dynamic Music Creation in Games". Technical Field

[0002] The present disclosure relates to the fields of music composition, music arrangement, machine learning, game design, and the psychological mapping of emotions. Background Art

[0003] Games have always been a dynamic pursuit, where the gameplay responds to the player's actions. As games become more cinematic and immersive, the importance of music continues to grow. Currently, most of the music in games is created from pre-written segments (usually pre-recorded), which are pieced together like puzzle pieces. They occasionally slow down, speed up, change key, and are often layered on top of each other. Composers can guess the possible paths throughout the gameplay, but since the gameplay is interactive, most of the gameplay is unpredictable, and of course, the timing of most parts is rarely predictable.

[0004] In parallel, machine learning and artificial intelligence have made it possible to generate content based on training sets of existing content marked by human reviewers. Additionally, there is a large corpus of review data and emotional mapping to various forms of artistic expression.

[0005] As another element, we are learning more and more about the players involved in games. Players who have opted in can be tracked on social media, and analyses of their personalities can be made based on their behavior. And as more users utilize biometric devices (skin conductance, pulse and respiration, body temperature, blood pressure, brain wave activity, genetic predisposition, etc.) that track them, environmental customization can be applied to the music environment. Summary of the Invention

[0006] The present disclosure describes a mechanism for analyzing music that separates musical components (rhythm, time signature, melodic structure, mode, harmonic structure, harmonic density, rhythmic density, and timbral density) and maps these components, individually and in combination, to emotional components based on published reviews and human opinions expressed on social media regarding concerts, recordings, etc. This is done at both the macro and micro levels within a musical composition. Based on this training set, faders (or virtual faders in software) are assigned emotional components such as tension, power, joy, surprise, softness, transcendence, peace, nostalgia, sadness, synesthesia, fear, etc. These musical components are mapped with reference to musical motifs that have been created for various elements / participants of a game, the elements / participants including but not limited to characters (leader, companion, main enemy, wizard, etc.), activity types (combat, rest, planning, hiding, etc.), regions (forest, city, desert, etc.), the personality of the person playing the game, etc. These motifs can be melodies, harmonies, rhythms, etc. Once the composer has created the motifs and assigned the desired emotions to the faders, the game simulation can be run, where the composer selects motif combinations and emotion mappings and applies them to various simulations. These simulations can be described a priori or generated using similar algorithms to map the game to a similar emotional environment. An actual physical fader (such as used in a computerized audio mixing console) may make the process of mapping emotions to situations more intuitive and instinctive, and the physicality will produce serendipitous results (e.g., even in a tense environment, raising the synesthesia fader may have a better and more interesting effect than raising the tension fader).

[0007] According to one embodiment of the present disclosure, a method for dynamic music creation is disclosed, comprising: a) assigning an emotion to a group of one or more musical motifs according to a character type, a musical composition component, or a scene; b) associating a vector with the emotion; c) mapping the group of one or more musical motifs to the vector based on the emotion; d) generating a musical composition based on the vector and one or more desired emotions.

[0008] According to one embodiment of the present disclosure, a system for dynamic music creation is disclosed, comprising: a processor; a memory coupled to the processor; and non-transitory instructions embedded in the memory, the non-transitory instructions, when executed, causing the processor to perform a method comprising: a) assigning an emotion to a group of one or more musical motifs according to a character type, a musical composition component, or a scene; b) associating a vector with the emotion; c) mapping the group of one or more musical motifs to the vector based on the emotion; d) generating a musical composition based on the vector and one or more desired emotions.

[0009] According to one embodiment of the present disclosure, a non - transitory instruction is disclosed, the non - transitory instruction being embedded in a computer - readable medium, and the non - transitory instruction, when executed, causes a computer to perform a method including the following: a) assign an emotion to a group of one or more musical ideas according to a role type, musical composition components, or a scene; b) associate a vector with the emotion; c) map the group of one or more musical ideas to the vector based on the emotion; d) generate a musical composition based on the vector and one or more desired emotions. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above - mentioned and other objects, features, and advantages of the present invention will become apparent from the following detailed description of some specific embodiments of the invention, particularly when considered in conjunction with the accompanying drawings, in which the same reference numerals in the various figures are used to designate the same components, and wherein:

[0011] Figure 1 is a block diagram depicting an overview of a dynamic music creation architecture according to aspects of the present disclosure.

[0012] Figure 2 is a block diagram showing the collection and analysis of music elements for generating a corpus of music elements according to aspects of the present disclosure.

[0013] Figure 3 is a block diagram depicting the collection and analysis of comments and reviews for generating a corpus of annotated performance data according to aspects of the present disclosure.

[0014] Figure 4 is a representative depiction of an embodiment of the emotional parameters of music according to aspects of the present disclosure.

[0015] Figure 5 is a systematic view of training a model using music and emotion comment elements according to aspects of the present disclosure.

[0016] Figure 6 is a block diagram showing the creation of an annotated set of game elements according to aspects of the present disclosure.

[0017] Figure 7 is a block diagram showing the creation and mapping of music elements with reference to game elements using virtual faders or real faders according to aspects of the present disclosure.

[0018] Figure 8 is a block diagram depicting the interaction of faders and switches with a game scenario playback engine according to aspects of the present disclosure.

[0019] Figure 9Is a block diagram showing how hint displays are used to preview upcoming events when mapping game and music elements, in accordance with aspects of the present disclosure.

[0020] Figure 10 Is a block diagram showing how hint displays are used to enable foreshadowing to manage upcoming music and game elements, in accordance with aspects of the present disclosure.

[0021] Figure 11 Depicts a schematic block diagram of a system for dynamic music creation in a game, in accordance with aspects of the present disclosure.

[0022] Figure 12A Is a simplified node diagram of a recurrent neural network used in dynamic music creation in a game, in accordance with aspects of the present disclosure.

[0023] Figure 12B Is a simplified node diagram of an unfolded recurrent neural network used in dynamic music creation in a game, in accordance with aspects of the present disclosure.

[0024] Figure 12C Is a simplified diagram of a conventional neural network used in dynamic music creation in a game, in accordance with aspects of the present disclosure.

[0025] Figure 12D Is a block diagram of a method for training a neural network for dynamic music creation in a game, in accordance with aspects of the present disclosure. Detailed Description

[0026] Although the following detailed description contains many specific details for illustrative purposes, any person of ordinary skill in the art should understand that many variations and changes to the following details are within the scope of the present invention. Accordingly, the exemplary embodiments of the present invention described below are set forth without loss of generality and without implying a limitation on the claimed invention.

[0027] Component Overview

[0028] From Figure 1 it can be seen that there are many components that make up the complete system. Details will be presented below, but it is useful to view the complete system from a high level. The components are broken down as follows:

[0029] First, music emotion data must be constructed: Music data 101 can be collected from a corpus of written and recorded music. This can be a score, a transcription created by a person, or a transcription created by an intelligent software system. Next, the system must break down the music data into its individual components – melody, harmony, rhythm, etc. 102, and these components must be stored as metadata associated with the work or a part of the work. Now, to determine the emotion components or tags associated with the individual music elements, we will rely on the wisdom of the crowd. Clearly, this is accurate as it is the emotion of the people related to the music we are trying to capture. We will use reviews, blog posts, album annotations, and other forms of commentary to generate emotion metadata for the entire work and individual parts 103. The next step is to map the music metadata to the emotion metadata 104. At this point, we will have a fairly large corpus of music data associated with emotion data, but it will never be complete. The next step is to use a Convolutional Neural Network (CNN) to compare the actual music metadata with the analysis of the emotion components 105. Then, this CNN can be used to suggest emotion associations for music that the model has not previously been trained on. Then the accuracy can be checked. Review and repeat – as the iterations continue, the accuracy will improve.

[0030] The next step is to analyze the game environment. The stages, component emotions, and characteristics of the game must first be collected and mapped. These scenarios and components include things such as locations, characters, environments, etc., actions such as battles, casting spells, etc., and so on 106. The next step is to map the game components to the emotion metadata components 107.

[0031] Once the game environment has been mapped, it is time to write musical ideas 108. Musical ideas can be written for all key characters, scenes, emotions, or any recurring elements of the game. Once the basic musical ideas have been written, faders (virtual or real) can be used to map the music emotion tags to the game tags 109. The personality or behavioral components of the user (game player) can also be mapped as additional emotion characteristics to be considered 110. Next, testing, review, and iteration continue, trying the model on different scenarios in the game under development 111.

[0032] Next is a more detailed description of the various processes.

[0033] Music Component Analysis

[0034] Some elements are needed to perform the analysis of the music components. These are in Figure 2Overview. First, there is a large music score corpus. The full scores are available in printed form 201. These printed full scores can be scanned and analyzed to convert them into electronic versions of the full scores. Any form of electronic score-making (MIDI, Sibelius, etc.) can be used. There are also full scores available in digital form 202, which can be added to the corpus. Finally, we can use machine learning to transcribe music into its musical elements (including full scores) 203. Machine learning methods can generate full scores for music that may not be in symbolic form. This will include multiple versions of the same song by the same artist. The advantage of analyzing live and recorded performances is that they can be mapped in more detail, including things like changes in tempo and dynamics; and for improvised material, many other variables can be considered. We must keep track of the analysis of different performances 204 so that we can know the differences between performances, for example, the performance of Beethoven's Symphony No. 5 by the New York Philharmonic conducted by Leonard Bernstein in 1964 will be different from the performance of the same work by Seiji Ozawa and the Boston Symphony Orchestra a few years later. These full scores must be analyzed for tempo 205, rhythmic components 206, melody 207, harmonic structure 208, dynamics 209, harmonic and improvisational elements 210. All of these are now stored in a detailed music performance corpus 211.

[0035] Once a thorough mechanical analysis of the large music corpus is done, then the descriptions from the reviews and the analysis of the literature about the works are used to map these analyses to the reviews of these works. This can be done in Figure 3seen in. The commentary and analysis can be in text, auditory, or visual form, and include, but are not limited to, articles 301, reviews 302, blogs 303, podcasts 304, album annotations 305, and social media comments 306. From this larger corpus of music analysis 307, descriptions 308 will be made of individual works and parts of works. In addition, these analyses must have a context within the period in which the works were composed and performed 309. For example, one might look at a Mozart piece and read the descriptions in the literature of his work (which could be a different period, such as a Mozart piece performed in the 20th century). Next, known structural music analysis 310 is applied to determine which works and / or parts of works or melodic phrases are beautiful, heroic, melancholy, jarring, sad, stately, powerful, expressive, etc. This may involve looking at music literature and applying general principles of compositional development to the analysis. For example, a melodic leap followed by a stepwise reversal of direction is considered beautiful. Melodic intervals can be associated with scales that are pleasing, unpleasing, or natural (e.g., a minor ninth leap is unpleasing, a fifth is heroic, and a single scale degree within a key signature is natural). All of this performance and compositional data can be used to map mood keywords to performances and segments 311.

[0036] The "mood" words may come from the process of numerically parsing the commentary and can include, and any manual analysis will almost necessarily, include the following words: tense, powerful, joy, surprise, soft, detached, peaceful, nostalgic, sad, and synesthesia.

[0037] Note that for these keywords to work, it is not necessarily important whether they are accurate, only that they have a consistent effect. When actually using the final system, it will not matter whether the faders associated with the word synesthesia actually create more sensually enjoyable music, but rather that it is predictable and emotionally understandable to the composer. There are two reasons for this: 1) the labels can always be changed to more intuitive names; and 2) humans are very adaptable when dealing with music and sound (in the history of music synthesizers, even though the labels were meaningless, knobs, buttons, and faders, such as on the Yamaha DX-7, did not in any way affect the intuitiveness of the sound based on the names, but rather had the effect of becoming part of the musician's sensory and muscle memory and thus were easy to use).

[0038] The detailed analysis of musical works and compositions is placed in a corpus of annotated performance data 312.

[0039] Note that mapping musical elements to create a corpus of musical performances need not be and will not be a one-to-one mapping, but rather is set on a scale of an emotional continuum, such as a set of musical vectors. For example, a piece might be eight-tenths on the synesthesia scale and four-tenths on the sadness scale. Once the musical elements are mapped to these emotional vectors, the model can be tested on titles not in the training dataset. One can listen to the results and fine-tune the model until it is more accurate. Eventually, the model will become very accurate at some point. Different models can be run in different time frames to create horror music in a 50s style or horror music in a 21st century style, keeping in mind that the actual descriptors are less important than the classification grouping (i.e., what might be soft for one composer might be boring for another).

[0040] Note that the model can map not only works and parts to emotional vectors, but also the constituent melodies, rhythms, and chord progressions to those same emotional vectors. At the same time, machine learn the melodic structures (e.g., a leap upward in melody followed by a step downward is generally considered nice, a leap of a minor ninth is generally considered dissonant, etc.).

[0041] Structural music analysis 310 can involve multiple components. For example, as Figure 4 shown, niceness, dissonance, and all subjective descriptive metrics can be used as compositional references in due course. For example, what was considered dissonant in Bach's time (e.g., the major seventh interval was considered too dissonant to play unless moving from a previously more consonant chord into the seventh) was considered nice in a later era (today, the major seventh chord is considered normal and even sometimes considered expressive).

[0042] In an existing method for classifying musical emotions 401, the emotions of songs are classified according to the traditional emotion model of psychologist Robert Thayer. The model divides songs along the lines of energy and stress into from happy to sad and from calm to energetic, respectively. The eight categories created by Thayer's model include each of the extremes of the two lines as well as the possible intersections of the lines (e.g., happy - energetic or sad - calm).

[0043] This historical analysis might have limited value, and the methods described herein might be more nuanced and flexible. Due to the size of the corpus and thus the size of the training dataset, a richer and more nuanced analysis can be applied.

[0044] Some 402 of the components to be analyzed can include, but are not limited to, harmonic groupings, modes and scales, time signatures, beats, harmonic density, rhythmic density, melodic structure, phrase length (and structure), dynamics, segmentation, and compositional techniques (modulation, inversion, retrograde, etc.), grooves (including rhythms such as funk, lounge, Latin, reggae, swing, tango, merengue, salsa, fado, 60s disco, 70s disco, heavy metal, etc., with hundreds of established rhythm styles).

[0045] Training model

[0046] Just as one would label pictures of peaches to train a convolutional neural network to recognize pictures of peaches when shown pictures of peaches it has never seen before, data on consumer and reviewer opinions can also be used to train a machine learning model, and then the model can be used to classify new material. Figure 5 An example of the process is shown. In the illustrated implementation, the wisdom of the crowd (e.g., reviews, posts, etc.) can be used to train our model to recognize the emotions associated with different musical segments.

[0047] In Figure 5 the illustrated implementation, a corpus of music data is divided into three buckets. The first bucket, Corpus A 501, is a set of music data 502 for which the collected emotion descriptions 503 have been mapped based on an analysis of listener review data. The second bucket, Corpus B 506, is a set of music data 507 for which its own set of collected emotion descriptions 508 have been mapped. The third bucket, Corpus C, is a set of music data 511 for which no collected emotion descriptions exist. The model is initially trained on Corpus A. The collected emotion descriptions 503 are mapped to the music data 502 using a convolutional neural network 504. The result of this mapping is a trained engine 505 that should be able to derive emotion descriptions from the music data. Then the music data 507 of Corpus B 506 is fed into the trained engine 505. The resulting output is the predicted emotion description 509 of the music of Corpus B. These predictions are then compared with the collected descriptions 510 to determine the accuracy of the predictions. The prediction, comparison, and iteration process (e.g., by moving musical elements from Corpus B to Corpus A and repeating) continues until the predicted emotion description 509 of the music of Corpus B closely matches the actual results. The process can continue indefinitely, with the system becoming increasingly intelligent over time. After the trained engine 505 has undergone sufficient iterations, it can be used as a prediction engine 512 on Corpus C 511 to predict the emotion description 513 of the music of Corpus C.

[0048] Mapping game elements

[0049] To match dynamic music elements with different elements in a game, a corpus of elements that may require music or music changes can be created. There are many different elements that can influence what the music should be at that moment. Of course, these elements are used by game developers and can be made available to composers and sound designers to match the music. As Figure 6 can be seen, there are many game elements to track and map. These include but are not limited to game characters 601, game environments or locations 602, game moods or tones 603, and perhaps most importantly, game events 604. Events in a game can be anything from the appearance of an enemy or friend to a change in location, to a battle, to casting a spell, to a race (and many more). All of these game elements 605, along with their triggers (inputs or outputs) and modifiers, are collected. Additionally, these can be collected as game vectors 606. They are called vectors because many of the components in the composition have multiple states. For example, a venue at night may be different from during the day, and a venue in the rain may be different from in the sun. The abilities of an enemy may be variable, increasing or decreasing based on time, location, or level. From this perspective, any element can have an array of vectors. All game elements and their vectors have an associated mood. The mood can change with the value of the vector, but nonetheless, the mood can be associated with the game element. Once the game element is mapped to its mood 607, we can store that relationship in a corpus or collection of annotated game elements 608.

[0050] Collecting ingredients required for the feed environment

[0051] Ultimately, the composition tool can be used to create music on its own from the gameplay, but this implementation focuses on creating music from composition primitives. The composer will create music elements or musical ideas and associate them with the characters or elements in the game.

[0052] The use of the tool here doesn't have to be used alone, but is often used to enhance traditional score-playing techniques, where pre-recorded music is combined with multiple components in sequence and in layers, thus being mixed together to create an entirety. That is to say, nowadays, different layers may start at different times, and some layers will be superimposed on other layers, and there are also some layers that may operate independently of other layers. The idea here is to develop additional tools that can 1) be used as a supplement to the prior art (e.g., on top of the prior art), 2) replace the prior art in some or all places, or 3) be used as a mechanism for notifying the use of the prior art. Additionally, these mechanisms can be used to create entirely new forms of interactive media (e.g., people trying to control blood pressure or brain wave states can use music feedback as a training tool, or even use biometric markers as a composition tool).

[0053] In traditional composition research, musical ideas usually refer to small melodic segments. However, in this context, musical ideas can be melodic segments, harmonic structures, rhythmic structures, and / or specific tones. Musical ideas can be created for as many elements / participants of the game as needed, the elements / participants including but not limited to characters (leader, partner, main enemy, wizard, etc.), activity types (combat, rest, planning, hiding, etc.), regions (forest, city, desert, etc.), the personalities of the people playing the game (young, old, male, female, introverted, extroverted, etc.). Musical ideas can be melodies, harmonies, rhythms, etc. Additionally, a single element can have multiple musical ideas, for example, rhythm patterns and melody patterns that can be used alone or together, or the same character may have both a sad musical idea and a happy musical idea for use in different situations.

[0054] Once the composer creates a musical idea, it can be assigned to an element / character. This can be done by the composer or the sound designer, and can be changed as the game is developed, or even dynamically changed within the game after release.

[0055] Mapping music components to game elements using tools

[0056] The foregoing aspects can be combined to map musical components to game elements. Throughout the document, it should be assumed that references to faders and buttons can be to real physical buttons or faders, or virtual buttons or faders. Intuitively, based on the experience of musicians and composers, it is expected that physical buttons and faders will be more intuitive to use and may produce better results. Actual physical faders (such as those used in computerized audio mixing consoles) may, even more likely, make the process of mapping emotions to situations more intuitive and instinctive, and the physicality will produce serendipitous results (for example, even in a tense environment, raising the synesthesia fader may have a better and more interesting effect than raising the tension fader). However, for the purposes of this application, either can work.

[0057] An example of a possible logical sequence of events is as Figure 7 shown, but any order can produce results, and the process will undoubtedly be iterative, recursive, and non-linear.

[0058] Figure 7 The logical sequence depicted in can start with a set 701 of annotated game elements, which can be transferred from the Figure 6 set 608 of annotated game elements in. Then, game triggers are assigned to switches at 702. These triggers can be the appearance of an enemy or obstacle or a level increase, etc. Next, a primary emotion vector 703 is selected, and a visual marker is assigned at 704 so that the element type can be shown to the composer. Then, game elements are assigned to emotion elements in a multi-dimensional array, as shown at 705. Then, the multi-dimensional array is assigned to faders and switches at 710, which are used to map musical emotion markers to game markers. Once a set of emotion markers has been established for the game elements, music can be applied to the game elements. The composer writes musical ideas 707. As described above, musical ideas can represent characters, events, regions, or emotions. Next, we select the baseline emotions 708 of the musical ideas. These are the default emotions. For example, whenever a wizard appears, there may be a default musical idea with its default emotion. These can vary based on the situation, but if no variables are applied, they will operate in their default mode. Now, we assign emotion vectors to faders 709. Now that the game vectors have been assigned and the music vectors have been assigned, at 710, the musical emotion markers can be mapped to game markers using the faders and switches. The faders do not have to correspond to a single emotion. Because it is like fine-tuning the tone of an instrument, the composer can be creative. Perhaps a fader that is 80% heroic and 20% sad will produce an unexpectedly pleasant result.

[0059] Next, various themes can be mapped to a set of buttons (possibly in a color matrix so that many can be seen at once). Some of these buttons may be the same as the game switches, but other buttons will be music switches. These buttons can be grouped by character type (hero, villain, wizard, etc.), musical composition elements (e.g., having rhythm on one side and melody on the other, with modes across the top), and scene (city, countryside, etc.).

[0060] Note that many of the game triggers that have been mapped to buttons or switches will ultimately be "driven" by the game-play process itself, but in the early stages, it will be useful to be able to simulate things like the arrival of an enemy, or a sunrise, or impending doom.

[0061] Now, the game simulation can be run in real time up to an early version of the game, even if it's just a plot summary, or even without a plot summary just to write music applicable to different situations. The composer can choose combinations of musical ideas and mood mappings and apply them to various simulations and test-drive them. As the game evolves, these situations can be fine-tuned.

[0062] In fact, this can also be used as generalized composition where the composer or performer can use a machine to create music based on primitives.

[0063] Programming music using faders and switches

[0064] Figure 8 An example of how game simulation and music programming can be actually used is shown. As Figure 7 shown in Figure 8As can be seen, musical ideas 801, game trigger factors 802, and game components 803 are mapped to faders and switches 804. These musical ideas 801, game trigger factors 802, and game components 803 mapped to faders and switches 804 are inserted into the game scenario playback engine 805. Now, while the composer is testing various game scenarios, music 806 is played and recorded by the recorded music scene / event mapping module 808. To see what elements are involved, the composer can view the cue display screen 808, which shows which elements are being used or are about to appear. If a fader is touched, the element it is mapped to may be highlighted on the screen to provide visual feedback. The cue display can show elements from the game scenario playback engine 805 as these elements are being played out from the real-time faders and switches. This is represented by the module 806 that plays music while playing the game and comes from the recorded music scene / event mapping module 807. Just as faders are updated when mixing with automated faders, all faders and switches can be updated and modified in real time at 809 and remembered for later playback. Additionally, musical ideas, components, and trigger factors can be updated at 810 and remembered by the system. Of course, there are an infinite number of undo operations, and different performances can be saved as different versions, which can be used at different times and combined with each other.

[0065] Figure 9 Additional details are shown regarding how a cue monitor 907 can be used in accordance with aspects of the present disclosure. Upcoming events 901, characters 902, and locations 903 can all be seen on the cue monitor (there may be multiple cue monitors). Based on the expected parameters of the gameplay process, there can be default weights 904. These and the fader (and switch) mapping matrix 905 are all shown on one or more cue monitors 907. The player history / prediction matrix 906 is another input to the fader mapping matrix 905 and is visible on the cue monitor. Based on the player's gaming style or other factors (time of day, minutes or hours in the current session, game state, etc.), this matrix can automatically change the music or change the music based on parameters set by the composer or sound designer.

[0066] Music foreshadowing changes

[0067] In movies, the music often changes before the visual changes. This process of foreshadowing or prefiguring is important both for connecting with the mood of the work and for preparing the viewer / listener for mood changes or other emotional preparations (even if it's a false prediction and it surprises the viewer). Now, when we're composing music on the fly, we'll want to be able to foreshadow changes. This could be associated with approaching the completion of a level or signaling the entry of a new character or new environment (or setting up one kind of change for the viewer but actually surprising them with another).

[0068] How will our composition engine effectively foreshadow? We can use triggers that are known in advance and use faders or dials to control the timing of the foreshadowing and the slope of the foreshadowing. For example, if a timer is running out on a level, the foreshadowing could be set to start 30 seconds before the time runs out and use an exponential curve (such as y = 2 x ) to increase the intensity. Many foreshadows are possible, from possible fear overtones to happy predictions. Again, because this is also designed to be instinctive, this feature will likely end up producing unexpected results, some of which will be useful for programming into the game. See Figure 10 , where a set of possible next events 1001 can be seen on the cue matrix 1003. There is also a next event timing matrix 1002, which can be used to set timing values and ramp characteristics. Another factor that might be involved in foreshadowing is the probability emotion weighting engine 1004. This engine weights the expected emotional state of upcoming events and is used to make the foreshadowing audio effective. It can be seen on the cue monitor, and its weighting effect can be changed. The likelihood of an upcoming event having any particular emotion and the predicted intensity of that emotion are mapped by the event-to-emotion likelihood mapping engine 1005. Foreshadowing is also affected by the player history and the player history / prediction matrix 1007, which can be seen on the cue monitor and is also fed into the foreshadowing matrix engine 1006. Although Figure 10 all of the elements in

[0069] Out-of-game use

[0070] The use of this composition tool is not limited to gaming. Even in non-gaming use, interactive VR environments can utilize these technologies. Additionally, this can be used as a composition tool for scoring traditional TV shows or movies. Moreover, one end use could be to use it in pure music creation to create the basis for albums or pop songs, etc.

[0071] System

[0072] Figure 11 FIG. depicts a system for dynamic music creation in a game according to aspects of the present disclosure. The system may include a computing device 1100 coupled to a user input device 1102. The user input device 1102 may be a controller, touch screen, microphone, keyboard, mouse, joystick, fader board, or other device that allows a user to input information including sound data into the system. The system may also be coupled to a biometric device 1123 configured to measure skin electrical activity, pulse and respiration, body temperature, blood pressure, or brain wave activity. The biometric device may be, for example but not limited to, a thermal or infrared camera pointed at the user and configured to determine the user's respiration and heart rate based on thermal signatures. For more information, see the co-owned patent No. 8,638,364, "USER INTERFACE SYSTEM AND METHOD USING THERMAL IMAGING" by Chen et al., the content of which is incorporated herein by reference. Alternatively, the biometric device may be, for example but not limited to, other devices such as a pulse oximeter, blood pressure cuff, electroencephalograph, electrocardiograph, wearable activity tracker, or smart watch with biometric sensing.

[0073] The computing device 1100 may include one or more processor units 1103, which may be configured according to well-known architectures (e.g., single-core, dual-core, quad-core, multi-core, processor-coprocessor, unit processor, etc.). The computing device may also include one or more memory units 1104 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), etc.).

[0074] The processor unit 1103 may execute one or more programs, portions of which may be stored in the memory 1104, and the processor 1103 is operatively coupled to the memory (e.g., by accessing the memory via the data bus 1105). The programs may be configured to generate or use the sonic musical ideas 1108 to create music based on the game vectors 1109 and emotion vectors 1110 of a video game. The sonic musical ideas may be short musical ideas composed by a musician, a user, or a machine. Additionally, the memory 1104 may contain programs implementing the training of the sound categorization and classification NN 1121. The memory 1104 may also contain one or more databases 1122 of annotated performance data and emotion descriptions. The neural network module 1121, such as a convolutional neural network for associating musical ideas with emotions, may also be stored in the memory 1104. The memory 1104 may store a report 1110 listing items not recognized by the neural network module 1121 as being in the database 1122. The sonic musical ideas, game vectors, emotion vectors, neural network module, and annotated performance data 1108, 1109, 1121, 1122 may also be stored as data 1118 in the mass storage 1118 or stored at a server coupled to a network 1120 accessed via the network interface 1114. Additionally, data for the video game may be stored in the memory 1104 as a database or elsewhere as data or as a program 1117 or data 1118 in the mass storage 1115.

[0075] The overall structure and probabilities of the NN may also be stored as data 1118 in the mass storage 1115. The processor unit 1103 is also configured to execute one or more programs 1117 stored in the mass storage 1115 or the memory 1104, the programs causing the processor to perform methods of dynamic music creation using the musical ideas 1108, game vectors 1109, and emotion vectors 1110 as described herein. The music generated from the musical ideas may be stored in the database 1122. Additionally, the processor may execute the methods described herein for the training of the NN 1121 and the classification of musical ideas to emotions. The system 1100 may generate neural networks 1122 as part of the NN training process and store them in the memory 1104. The complete NNs may be stored in the memory 1104 or stored as data 1118 in the mass storage 1115. Additionally, the NN 1121 may be trained using the actual responses from the user, where the biometric device 1123 is used to provide biofeedback from the user.

[0076] The computing device 1100 may also include well-known support circuits such as input / output (I / O) 1107, circuitry, power supply (P / S) 1111, clock (CLK) 1112, and cache 1113, which may communicate with other components of the system via, for example, bus 1105. The computing device may include a network interface 1114. The processor unit 1103 and the network interface 1114 may be configured to implement a local area network (LAN) or a personal area network (PAN) via a suitable network protocol for a personal area network (PAN) (e.g., Bluetooth). The computing device may optionally include a mass storage device 1115 (such as a disk drive, CD-ROM drive, tape drive, flash memory, etc.), and the mass storage device may store programs and / or data. The computing device may also include a user interface 1116 to facilitate interaction between the system and the user. The user interface may include a monitor, a television screen, speakers, headphones, or other devices that convey information to the user.

[0077] The computing device 1100 may include a network interface 1114 to facilitate communication via an electronic communication network 1120. The network interface 1114 may be configured to implement wired or wireless communication via local area networks and wide area networks (such as the Internet). The device 1100 may send and receive data and / or requests for files via the network 1120 in one or more packets. Packets sent via the network 1120 may be temporarily stored in a buffer 1109 in the memory 1104. Annotated performance data, audio music ideas, and annotated game elements may be obtained via the network 1120 and partially stored in the memory 1104 for use.

[0078] Neural network training

[0079] Generally, a neural network for dynamic music generation may include one or more of several different types of neural networks and may have many different layers. By way of example and not limitation, a classification neural network may be composed of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs).

[0080] Figure 12A Depicts a basic form of an RNN with a node layer 1220, each of the nodes being characterized by an activation function S, an input weight U, a recurrent hidden node transition weight W, and an output transition weight V. The activation function S may be any non-linear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S may be a Sigmoid or ReLu function. Different from other types of neural networks, an RNN has a set of activation functions and weights throughout the layer. As Figure 12BAs shown, an RNN can be considered a series of nodes 1220 with the same activation function moving in times T and T+1. Thus, the RNN maintains historical information by feeding the results from the previous time T to the current time T+1.

[0081] In some embodiments, a convolutional RNN can be used. Another type of RNN that can be used is a long short-term memory (LSTM) neural network, which adds storage blocks in the RNN nodes with input gate activation functions, output gate activation functions, and forget gate activation functions, thus forming a gated memory, which allows the network to retain some information for a longer time, as described in Hochreiter and Schmidhuber's "Long Short-Term Memory" (Neural Computation 9(8): 1735-1780 (1997)), which is incorporated herein by reference.

[0082] Figure 12C Depicts an exemplary layout of a convolutional neural network such as a CRNN according to aspects of the present disclosure. In this depiction, a convolutional neural network is generated to train data in the form of an array 1232 (e.g., having 4 rows and 4 columns, thus providing a total of 16 elements). The depicted convolutional neural network has filters 1233 of size 2 rows by 2 columns (with a stride value of 1) and channels 1236 of size 9. For clarity, only the connections 1234 between the first column of channels and their filter windows are depicted in Figure 12C However, aspects of the present disclosure are not limited to such implementations. According to aspects of the present disclosure, a convolutional neural network implementing classification 1229 can have any number of additional neural network node layers 1231 and can include any such layer types of any size, such as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, local contrast normalization layers, etc.

[0083] As seen in Figure 12D the training of a neural network (NN) begins with the initialization 1241 of the weights of the NN. Generally, the initial weights should be randomly assigned. For example, an NN with a tanh activation function should have random values distributed between and where n is the number of inputs to the node.

[0084] After initialization, the activation function and the optimizer are defined. Then, a feature vector or an input data set 1242 is provided to the NN. Each of the different feature vectors can be generated by the NN from inputs with known labels. Similarly, a feature vector can be provided to the NN that corresponds to an input with a known label or classification. The NN then predicts the label or classification 1243 of the feature or input. The predicted label or class is compared with the known label or class (also referred to as the ground truth), and the loss function measures the total error between the prediction and the ground truth over all training samples 1244. By way of example and not limitation, the loss function can be a cross-entropy loss function, quadratic cost, triplet contrast function, exponential cost, etc. Multiple different loss functions can be used depending on the purpose. By way of example and not limitation, for training a classifier, a cross-entropy loss function can be used, and for learning pre-trained embeddings, a triplet contrast function can be used. Then, the result of the loss function is used and the NN is optimized and trained using known training methods for neural networks such as backpropagation with adaptive gradient descent, etc. 1245. In each training epoch, the optimizer tries to pick the model parameters (i.e., weights) that minimize the training loss function (i.e., the total error). The data is divided into training samples, validation samples, and test samples.

[0085] During training, the optimizer minimizes the loss function on the training samples. After each training epoch, the model is evaluated on the validation samples by computing the validation loss and accuracy. If there is no significant change, training can stop, and the resulting trained model can be used to predict the labels of the test data.

[0086] Thus, a neural network can be trained from inputs with known labels or classifications to identify and classify those inputs.

[0087] While the foregoing is a complete description of the preferred embodiments of the present invention, the use of various alternatives, modifications, and equivalents is possible. Therefore, the scope of the present invention should not be determined with reference to the above description, but instead should be determined with reference to the appended claims and their equivalents in their entire scope. Any feature described herein (whether preferred or not) can be combined with any other feature described herein (whether preferred or not). In the appended claims, The indefinite article "a" or "an" refers to the quantity of one or more items after the article, unless otherwise expressly stated therein. The appended claims should not be construed as including a means-plus-function limitation, unless such a limitation is expressly recited in a given claim using the phrase "means for".

Claims

1. A method for dynamic music creation, which comprises: a) Assigning an emotion to a group of one or more musical ideas based on a character type, musical composition element, or scene; b) Associating a vector with the emotion; c) Mapping the group of one or more musical ideas to the vector based on the emotion; d) Generating a musical composition based on the vector and one or more desired emotions.

2. The method according to claim 1, wherein the emotion is assigned to the group of one or more musical ideas by a neural network trained to assign emotions to music based on a character, musical composition element, or scene.

3. The method according to claim 1, wherein a rank is assigned to the emotion, and the mapping of the group of one or more musical ideas to the vector changes with the rank assigned to the emotion.

4. The method according to claim 3, wherein an emotion rank fader is used to change the rank assigned to the emotion.

5. The method according to claim 1, further comprising associating user behavior with the emotion and mapping the group of one or more musical ideas to the user behavior based on the emotion.

6. The method according to claim 1, wherein the assignment of the group of one or more musical ideas to the emotion is based on the user's physical reaction to each idea in the group of one or more musical ideas.

7. The method according to claim 1, further comprising assigning one of the one or more musical ideas in the group of one or more musical ideas to a game vector as a default idea, and wherein the default idea is included in the musical composition whenever the vector is present within a video game.

8. The method according to claim 1, wherein the vector comprises game elements and trigger factors for the game elements within a video game.

9. The method according to claim 8, further comprising assigning one of the one or more musical ideas in the group of one or more musical ideas to a game element as a default idea for the game element, and including the default idea in the musical composition whenever the game element is present within the video game.

10. The method according to claim 8, further comprising assigning one of the one or more musical ideas in the group of one or more musical ideas to the trigger factor as a default idea for the trigger factor, and including the default idea in the musical composition whenever the trigger factor is activated within the video game.

11. The method according to claim 1, wherein generating the musical composition comprises: Selecting one musical idea from the one or more musical ideas in the group of one or more musical ideas to be played while playing a video game, wherein the one or more musical ideas in the group of one or more musical ideas are played simultaneously with the video game.

12. The method according to claim 1, further comprising generating an event cue for a video game, wherein the event cue includes an upcoming event within the video game, and wherein the event cue is used during the generation of the music composition.

13. The method according to claim 12, wherein the event cue includes an event probability, an event timing, and an event mood weight, wherein the event probability defines the likelihood that the event will occur within the event timing, and wherein the event mood weight provides the mood of the event.

14. The method according to claim 12, wherein the event mood weight includes a level of the mood.

15. A system for dynamic music creation, which comprises: a processor; a memory coupled to the processor; non-transitory instructions embedded in the memory, the non-transitory instructions, when executed, cause the processor to perform a method comprising the following: a) Assign an emotion to a group of one or more musical ideas based on a character type, a music composition component, or a scene; b) Associate a vector with the emotion; c) Map the group of one or more musical ideas to the vector based on the emotion; d) Generate a music composition based on the vector and one or more desired emotions.

16. The system according to claim 15, wherein the emotion is assigned to the group of one or more musical ideas by a neural network trained to assign emotions to music based on a character, a music composition component, or a scene.

17. The system according to claim 15, wherein a level is assigned to the emotion, and the mapping of the group of one or more musical ideas to the vector changes with the level assigned to the emotion.

18. The system according to claim 17, further comprising an emotion fader coupled to the processor, wherein the emotion fader is configured to control the level assigned to the emotion.

19. The system according to claim 15, wherein the assignment of the group of one or more musical ideas to the emotion is based on a user's physical reaction to each idea in the group of one or more musical ideas.

20. The system according to claim 19, further comprising a biometric device coupled to the processor and configured to track the user's physical reaction.

21. A non-transitory instruction embedded in a computer-readable medium, the non-transitory instruction, when executed, causes a computer to perform a method comprising the following: a) Assign an emotion to a group of one or more musical ideas based on a character type, a music composition component, or a scene; b) Associate a vector with the emotion; c) Map the group of one or more musical ideas to the vector based on the emotion; d) Generate a music composition based on the vector and one or more desired emotions.