An audio playing method, system, device and computer program product

By adjusting performance parameters and enriching performance content in real time through the audio playback system, the problem of the limited functionality of the intelligent drummer in the game was solved. This enabled real-time adjustment of music parameters and rich performance expression, improving the user experience and saving storage resources.

CN119680209BActive Publication Date: 2026-02-24NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411784615.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2026-02-24
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The existing intelligent drummers in games have limited functionality, cannot adjust music parameters in real time, and their performance content is too simple. It is difficult to customize the speed, style, and timbre of the drum part through simple operations, resulting in a poor user experience.

Method used

By adjusting performance parameters in real time through the audio playback system, and combining the user interface and audio engine, dynamic control of music clips and sound effect datasets can be achieved, including real-time adjustment of performance style, density and dynamics, and enriching the performance content by using combinations of preset music clips and sound effect datasets.

Benefits of technology

It enables real-time adjustment of music parameters and rich performance expression in the game, improving the user experience while saving storage resources and simplifying the interaction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119680209B_ABST
    Figure CN119680209B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio playing method, device, system and computer program product. The method comprises: obtaining a first input input through a user interface, the first input comprising music parameter information; obtaining a second input input through the user interface, the second input comprising sound effect data information; determining a target music segment based on the first input, the target music segment being described by target music parameters; determining a target sound effect data set based on the second input, the target sound effect data set comprising audio samples of different drum parts with the same sound effect style; and driving the target sound effect data set to play based on the target music segment. By selecting different music parameters and sound effect styles by the user, real-time adjustment of playing parameters and playing content can be realized, which can enrich the playing content and save storage resource space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of virtual musical instrument technology, and in particular to an audio playback method, system, device, and computer program product applied in games. Background Technology

[0002] In recent years, an increasing number of games have incorporated instrumental performances and ensembles into various game genres as a casual gameplay element to enrich the gaming experience. For example, an intelligent drummer is an automatic accompaniment system for drum parts in instrumental ensembles, which can significantly optimize the overall experience of instrumental performances and ensembles in games. Currently, intelligent drummers used in games have limited functionality, cannot achieve real-time adjustment of musical parameters, and their performance content is too simplistic. If a system is needed that can automatically play customizable drum parts with simple operations, and where the tempo, style, timbre, and other musical parameters and content of this drum part are controllable and variable to suit various performance styles, developing such a system using traditional methods is often too difficult and costly for game developers. This is the main reason why intelligent drummers have not yet appeared in games.

[0003] This application provides an audio playback system that enables real-time adjustment of music parameters and a wide range of performance options in games, while also taking into account sound quality, user-friendliness, and a reasonable development workload. Summary of the Invention

[0004] This specification provides one or more embodiments of an audio playback method, the method comprising: acquiring a first input input through a user interface, the first input including music parameter information; acquiring a second input input through the user interface, the second input including sound effect data information; determining a target music segment based on the first input, the target music segment being described by target music parameters; determining a target sound effect dataset based on the second input, the target sound effect dataset including audio samples of different drum parts having the same sound effect style; and driving the target sound effect dataset to play based on the target music segment.

[0005] According to one or more embodiments of this specification, the audio playback method includes music parameter information such as performance style information, performance density information, and performance dynamics information. Determining a target music segment based on a first input includes determining a target performance style, target performance density, and target performance dynamics based on the performance style information, performance density information, and performance dynamics information. The method also includes determining a target music segment from multiple preset music segments based on the target performance style, target performance density, and target performance dynamics. Each preset music segment is described by multiple music parameters, and the preset music segments are stored in a Digital Music Interface (MIDI) file format.

[0006] According to one or more embodiments of this specification, the frequency playback method provides that the same timbre in multiple preset music segments is at the same pitch.

[0007] According to one or more embodiments of this specification, the audio playback method includes sound effect data information, and the method for determining a target sound effect dataset based on a second input includes determining the target sound effect dataset from multiple preset sound effect datasets based on the sound effect style information. Each preset sound effect dataset includes a set of audio samples of each drum part in a drum part, and the audio samples of each drum part correspond to the same sound effect style.

[0008] The audio playback method provided according to one or more embodiments of this specification further includes: obtaining a third input input through a user interface, the third input including a target beat count (BPM) set by the user; determining a target music segment playback speed based on the mapping relationship between the beat count and the music segment playback speed and the target beat count; and playing a target sound effect dataset based on the target music segment playback speed.

[0009] According to one or more embodiments of this specification, the audio playback method includes a user interface including a style selector and an X / Y controller. Obtaining a first input from the user through the user interface includes: obtaining performance style information from the first input through the style selector; and obtaining performance density information and performance velocity information from the first input through the X / Y controller.

[0010] According to one or more embodiments of this specification, the audio playback method includes a user interface with a sound effect selector, and obtaining a second input through the user interface includes obtaining sound effect data information from the second input through the sound effect selector.

[0011] This specification provides one or more embodiments of an audio playback device, including a processor and a memory. The memory stores a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by the processor, they implement the audio playback method provided in one or more embodiments of this specification.

[0012] This specification provides one or more embodiments of a computer program product, characterized in that it includes a computer program, which, when at least a portion of the computer program is executed by a processor, is capable of implementing the audio playback method provided in one or more embodiments of the specification.

[0013] This specification provides one or more embodiments of an audio playback system, including: a game engine including a user interface, the user interface including a music selection module and a sound effect selection module, the music selection module being used to acquire a first input, the first input including music parameter information, and the sound effect selection module being used to acquire a second input, the second input including sound effect data information. The system also includes an audio engine, the audio engine including: a music module and a sound effect module, the music module being configured to determine a target music segment based on the first input, the target music segment being described by target music parameters; the sound effect module being configured to determine a target sound effect dataset based on the second input, the target sound effect dataset including audio samples of different drum parts having the same sound effect style; and to drive the target sound effect dataset to play based on the target music segment.

[0014] According to one or more embodiments of the audio playback system provided in this specification, the music selection module includes a style selector and an X / Y controller. The style selector is used to input performance style information from the music parameter information, and the X / Y controller is used to input performance density information and performance velocity information from the music parameter information.

[0015] According to one or more embodiments of the audio playback system provided in this specification, the user interface further includes a playback control module, which includes a beat count (BPM) selector for acquiring a third input, the third input including a target beat count set by the user, and the audio engine is further configured to: determine the playback speed of a target music clip based on the mapping relationship between the beat count and the playback speed of the music clip and the target beat count; and play the target sound effect dataset based on the playback speed of the target music clip.

[0016] The audio playback system provided according to one or more embodiments of this specification further includes a storage module for storing multiple preset sound effect datasets. The sound effect module is further configured to: determine a target sound effect dataset from the multiple preset sound effect datasets based on sound effect data information. Each preset sound effect dataset includes a set of audio samples of each drum part in a drum section, and the audio samples of each drum part correspond to the same sound effect style.

[0017] The audio playback system provided according to one or more embodiments of this specification further includes a storage module for storing a plurality of preset music segments. The music module is further configured to: determine a target music segment from the plurality of preset music segments based on a first input. Each preset music segment is described by a plurality of music parameters. The preset music segments are stored in a digital music interface (MIDI) file format. Attached Figure Description

[0018] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. The same numbers in the drawings denote the same structures or steps.

[0019] Figure 1 This is a schematic diagram of an audio playback system according to some embodiments of this specification.

[0020] Figure 2 These are specific embodiments of the audio playback system shown in some examples of this specification.

[0021] Figure 3 This is a schematic diagram of the music module configuration according to some embodiments of this specification.

[0022] Figure 4 This is a schematic diagram of the sound effect module configuration based on some embodiments described in this specification.

[0023] Figure 5 This is a schematic diagram of BPM control based on RTPC technology, as shown in some embodiments of this specification.

[0024] Figure 6 This is a schematic diagram illustrating the selection or switching of RTPC-based performance density according to some embodiments of this application.

[0025] Figure 7 This is a schematic diagram of a user interface according to some embodiments of this application.

[0026] Figure 8 This is a schematic diagram of an X / Y controller according to some embodiments of this application.

[0027] Figure 9 This is a schematic diagram of a BPM selector according to some embodiments of this application.

[0028] Figure 10 The flowchart of the audio playback method shown in some embodiments of this specification is illustrated.

[0029] Figure 11 A schematic diagram of an audio playback device is shown based on some embodiments of this specification.

[0030] Figure 12 An electronic device for audio playback is shown according to some embodiments of this specification. Detailed Implementation

[0031] To more clearly illustrate the technical solutions of the embodiments in this specification, the embodiments will be described in detail below with reference to the accompanying drawings. Obviously, the content described below are some examples or embodiments of this specification. For those skilled in the art, without creative effort, the technical solutions or means disclosed in this specification can be applied to other scenarios based on this technical content.

[0032] It should be understood that the terms "system," "device," "unit," and / or "module" used in this specification are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0033] Unless otherwise specified, the technical terms used to describe components, elements, etc. in this specification are not singular but may include plural. Generally speaking, terms such as "comprising" or "including" only indicate that explicitly identified steps, elements, or components are included, and these steps, elements, and components do not constitute an exclusive list, as the described method or apparatus may also include other steps or components.

[0034] This specification uses flowcharts to illustrate the operational steps performed by the apparatus or system of related embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of execution. Those skilled in the art can adjust the order of these steps based on the knowledge and information conveyed by the embodiments in this specification. Such adjustments include, but are not limited to, reversing the order of steps, merging multiple steps, and splitting a step.

[0035] In games, instrument sound triggering, including drums, typically doesn't include a "single-timbre real-time trigger" mode or a "looping of a specific instrument's audio" mode. A "single-timbre real-time trigger" mode plays a sound only once per button click. A "looping of a specific instrument's audio" mode plays a pre-defined audio clip each time a button is clicked. For example, in "single-timbre real-time trigger" mode, clicking the bass drum button and snare drum button on a drum kit would produce a "boom" and a "da" sound, respectively. In "looping of a specific instrument's audio" mode, when the user clicks the start button, the pre-prepared audio clip "boom~da boom~boom-da~" begins looping until the user clicks the stop button. Users can trigger other instrument sounds simultaneously during this looping drum part, creating an "ensemble" effect within the game. When playing expressive drums (especially complex, fast-paced drum kits) using "single-timbre real-time trigger," it demands a higher level of skill from the user, as they need to play each part like a professional drummer using a touchscreen or computer keyboard. To address this issue, games typically use a looping playback of a specific audio segment to mitigate the problem. However, in some embodiments, this looping mode may not allow for real-time BPM adjustment and offers a monotonous presentation with virtually no adjustable or variable performance parameters. Furthermore, in some embodiments, adding more audio segments requires additional storage space, unnecessarily increasing the game's file size. Moreover, adding numerous audio segments not only complicates storage acquisition but also complicates the user's search and selection process, thus deviating from the initial goal of "reducing difficulty."

[0036] Therefore, some embodiments of this specification propose an audio playback system that allows for real-time adjustment of performance parameters and content within a game, providing a simple and user-friendly interactive experience with flexible control and rich expression. The audio playback system includes a game engine and a user interface. The user interface includes a music selection module and a sound effect selection module. The music selection module acquires a first user input, which includes music parameter information. The sound effect selection module acquires a second user input, which includes sound effect data information. The audio playback system also includes an audio engine. The audio engine includes a music module and a sound effect module. The music module is configured to determine a target music segment based on the first input, the target music segment including target music parameters. The sound effect module is configured to determine a target sound effect dataset based on the second input, the target sound effect dataset including audio samples of different drum parts with the same sound effect style. The system also drives the playback of the target sound effect dataset based on the target music segment. By allowing the user to select different music parameters and sound effect styles, real-time adjustment of performance parameters can be achieved. By pre-setting multiple music segments and multiple pre-setting sound effect datasets, and through random combinations of music segments and sound effect datasets, the performance content can be enriched while saving storage space.

[0037] Figure 1 This is a schematic diagram of a music playback system according to some embodiments of this specification.

[0038] The audio playback system 100 can be used to implement the audio of various virtual musical instruments (e.g., drums, pianos, etc.) in games. In some embodiments, the audio playback system may also be referred to as a virtual musical instrument system, such as a virtual drummer system. This specification uses a virtual drummer system as an example, but it should not be limited to the scope of the embodiments described. The virtual drummer system can use computer software to simulate the sound (timbre and expression) of a real drum part. The drum part includes the drum and cymbals. The drum may include a snare drum, bass drum, floor drum, tom-tom, etc. The cymbals may include hi-hats, crash cymbals, saddle cymbals, stacked cymbals, Chinese cymbals, splash cymbals, etc. In some embodiments, the drum may also include a conga, bongo, tambourine, etc.

[0039] At least some components of the audio playback system 100 may be integrated into the terminal 110. In some embodiments, the terminal 110 may include a mobile device, tablet computer, ..., laptop computer, or any combination thereof. For example, a mobile device may include a mobile phone, personal digital assistant (PDA), gaming device, navigation device, laptop computer, tablet computer, desktop computer, or any combination thereof. In some embodiments, the terminal 110 may include input devices, output devices, etc. Input devices may include alphanumeric keys and other keys that can be input via a keyboard, touchscreen (e.g., with haptic or haptic feedback), voice input, eye-tracking input, brain monitoring system, or any other similar input mechanism. Input information received through the input devices may be transmitted, for example, via a bus, to the processing device 112 for further processing. Other types of input devices may include cursor control devices, such as a mouse, trackball, or arrow keys. Output devices may include a display, speakers, or similar devices, or combinations thereof.

[0040] Terminal 110 may provide a user interface. A user interface refers to the interface through which a user interacts with the audio playback system 100, the processing device 112, or the software. In some implementations, the user interface may include a graphical user interface, a touch user interface, a voice user interface, a natural user interface, etc. A graphical user interface can use visual elements such as windows, icons, buttons, and menus for interaction. A touch user interface allows users to interact by touching the screen. A voice user interface allows users to interact with the system by voice. A natural user interface utilizes natural user behavior for interaction, such as gesture recognition, facial recognition, etc. In some embodiments, terminal 110 may provide multiple types of user interface windows for the user to choose from.

[0041] In some embodiments, a user can interact with the processing device 112 through a user interface to achieve audio playback. For example, a user can input a first input and / or a second input through the user interface, and the processing device 112 can determine the target music segment (e.g., target MIDI file) and target sound effect dataset selected by the user based on the first and second inputs, respectively. As another example, a user can input music output parameters (e.g., target BPM) through the user interface, and the processing device 112 can adjust the playback parameters (e.g., BPM) of the target music segment (e.g., target MIDI file) according to the music output parameters input by the user. In some embodiments, the user interface may include input controls (e.g., buttons, text boxes, drop-down menus, tabs, X / Y controllers, etc.), navigation components (e.g., menu bars), information displays (e.g., labels, icons, message boxes, etc.). Input controls can be used for user input of operation instructions. For example, a user can trigger the audio engine to play the target MDI file selected by the user through a play button. As another example, a user can select a performance style through a tab or drop-down menu. Yet another example, a user can determine the performance density and / or performance dynamics through the X / Y controller. Information displays can provide prompts to the user, such as operation prompts for the X / Y controller. Navigation components can help users switch between different interfaces and functions. For example, navigation components can include game entry points and audio entry points to enable switching between the game interface and the audio interface.

[0042] In some embodiments, when the audio playback system 100 is used in a game, the user interface may be the user interface of the game engine.

[0043] Terminal 110 may include processing device 112 and memory (not shown in the figure).

[0044] Processing device 112 can process data acquired from the input device of terminal 110 and / or send processed data to the output device of terminal 110. For example, processing device 112 can acquire a first user input through a user interface, the first input including music parameter information; and acquire a second user input through the user interface, the second input including sound effect data. Processing device 112 can determine a target music segment (e.g., a target MIDI file) based on the first input; determine a target audio dataset based on the second input; and output the target audio dataset based on the target music segment (e.g., a target MIDI file). The target audio dataset can be output to the speaker of terminal 110 to play the target audio sample. The target audio dataset includes multiple target audio samples.

[0045] In some embodiments, processing device 112 may include a processor. The processor executes computer instructions (e.g., program code) as described herein and performs the functions of processing device 112. For example, the computer instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions that perform specific functions described herein. For example, the processor may process data acquired from memory, terminal 110, and / or any other component. In some embodiments, the processor may include one or more hardware processors, such as a microcontroller, microprocessor, reduced instruction set computer (RISC), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), high-order RISC system (ARM), programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof.

[0046] Processing device 112 can be used to run audio engine 130 and / or game engine 120. In some embodiments, audio engine 130 and game engine 120 can run on different processors. For example, audio engine 130 can run on a first processor, and game engine 120 can run on a second processor. The first and second processors can communicate to perform data calls and / or transmissions. For example, when a user runs game engine 120, the user's first and second inputs can be sent to audio engine 130 through game engine 120. Audio engine 130 can determine a target music segment and a target audio dataset based on the user's first and second inputs, respectively, and drive the target audio dataset to play using the target music segment.

[0047] The audio engine 130 includes a music module 134, a sound effects module 136, and an event triggering module 132.

[0048] Music module 134 is used to provide and switch descriptive information about musical performance. This descriptive information includes musical parameters related to the musical performance. These parameters may include performance content, dynamics, tempo, rhythm, and performance style.

[0049] Performance density describes the number of notes played within a given time period in a musical performance. The more notes played within a given time, the higher the performance density. In some implementations, performance density can be categorized into different levels. Figure 2 These are schematic diagrams illustrating specific embodiments of an audio playback system based on some of the embodiments described in this specification. For example... Figure 2As shown, performance density can be divided into three levels: "sparse," "medium," and "dense." Alternatively, performance density can be divided into five levels: "very sparse," "sparse," "medium," "dense," and "very dense." In some embodiments, performance density can be represented by the number of notes played within a given time period.

[0050] Dynamics refer to the variations in volume during musical performance; it can also be called performance intensity. In some implementations, dynamics can be categorized into different degrees. For example, ... Figure 2 As shown, playing dynamics can be divided into three levels: "heavy," "medium," and "light." Alternatively, playing dynamics can be divided into six levels: "very soft," "soft," "medium-soft," "medium-loud," "loud," and "very loud." In some embodiments, playing dynamics can be expressed by the volume level.

[0051] A beat is a repetitive, evenly distributed unit of time in musical performance. The speed of a beat is determined by tempo, which is measured in beats per minute (BPM). For example, 60 BPM means 60 beats per minute. In some performances, beats can be categorized into different levels. For instance, beats can be classified as "very slow," "slow," "medium," "fast," and "extremely fast" to indicate different tempos.

[0052] Performance styles can include classical, jazz, rock, pop, folk, neoclassical and experimental, disco, etc.

[0053] The performance content can include notes or pitches that need to be played at different times.

[0054] Pitch refers to the degree of highness or lowness of a sound. In some embodiments, pitch can be marked by octaves, such as C1, C2, C3, C4, ... B4. In some embodiments, pitch can be marked by the twelve-tone equal temperament, such as C, D, E, F, G, A, B.

[0055] In some embodiments, the music module 134 may provide and / or store multiple sets of descriptive information (i.e., music parameter information) corresponding to multiple preset music segments. Each preset music segment may include a temporal arrangement of different music parameters. In some embodiments, each preset music segment may be stored as a MIDI (Musical Instrument Digital Interface) file, also known as a MIDI music segment or MIDI segment. Since MIDI files only describe musical instructions rather than actual audio, their file size is relatively small, thus saving storage space.

[0056] A MIDI music clip is a set of descriptive information (i.e., different musical parameters) arranged in a temporal sequence to represent musical performance. The musical data instructs electronic instruments (e.g., synthesizers or computer software) how to generate or play specific notes. Furthermore, a MIDI music clip includes the temporal distribution of different musical parameters; for example, a MIDI music clip may include notes played at different times with varying playing styles, densities, and dynamics at different pitches.

[0057] In some embodiments, each MIDI file may include only one track. Since the embodiments in this specification only require the drum section, only one track is needed in a MIDI file. Additional tracks are redundant and meaningless; setting only one track avoids unexpected errors. In some embodiments, when the audio playback system is used to generate sounds from multiple types of instruments, such as drums, xylophones, and pianos, the number of tracks in the MIDI file can be set according to the number of instrument types.

[0058] Each pitch in different frequencies within multiple preset music segments corresponds to the timbre of a drum part (e.g., snare drum, bass drum, tom-tom, hi-hat, cymbal, spiking cymbal, Chinese cymbal, splash cymbal, sub-bass drum, etc.). For example, C1 corresponds to the bass drum, D1 corresponds to the snare drum, etc.

[0059] The creation of preset music clips can be designed according to the needs of the application scenario. For example, multiple different preset music clips can be created based on performance style, performance density, and performance dynamics. Each preset music clip can include one performance style, one performance density, and one performance dynamics. Assuming there are M performance styles, N performance densities, and P performance dynamics, a total of M*N*P preset music clips (i.e., MIDI files) can be created.

[0060] For example, assuming there are two performance styles, Rock 1 and Rock 2, three performance densities (sparse, medium, dense), and three performance dynamics (heavy, medium, light), then 9 MIDI files (3*3) are needed for each performance style. For example, for Rock 1, these would include MIDI file 1 (Rock 1, sparse, heavy), MIDI file 2 (Rock 1, sparse, medium), MIDI file 3 (Rock 1, sparse, light), MIDI file 4 (Rock 1, medium, heavy), MIDI file 5 (Rock 1, medium, medium), MIDI file 6 (Rock 1, medium, light), MIDI file 7 (Rock 1, dense, heavy), MIDI file 8 (Rock 1, dense, medium), and MIDI file 9 (Rock 1, dense, light). The MIDI files for Rock 2 are similar to those for Rock 1 and will not be described again here. Based on the two assumed different performance styles, a total of 2*3*3=18 MIDI clips, or 18 MIDI files, are needed for the Rock style. Following the same specifications, for each additional performance style, 9 more MIDI files are created according to the required style.

[0061] It's worth noting that the more musical parameters such as playing style, dynamics, and density are broken down into, the more preset music segments there are, resulting in more preset MIDI files. This finer granularity leads to a more linear adjustment process during use. Following the same specifications, for each additional playing style, nine more MIDI files are created according to the desired style. Of course, this number is just an example; in actual production, there can be more or fewer, but the content must maintain this design logic.

[0062] In a preset music clip, each pitch corresponds to a specific drum or cymbal timbre, and similar timbres in different preset music clips are played at the same pitch. For example, if D1 in MIDI file 1 is configured to correspond to the snare drum, then D1 in all other MIDI files must also be configured to correspond to the snare drum or a snare drum-like timbre (e.g., the clapping timbre in an electronic drum kit), and cannot be anything else such as the bass drum or hi-hat. This is to facilitate the "arbitrary combination" of timbres between different music clips and different drum kits. If pitch C1 in MIDI file 1 is configured to correspond to the bass drum, and the bass drum in drum kit A is also configured to correspond to pitch C1, then when MIDI file 1 is played and pitch C1 is reached, the bass drum in drum kit A will be triggered, and the bass drum sound will be heard. However, if MIDI file 2 incorrectly configures pitch C1 to correspond to the snare drum, then when drum kit A is played using MIDI file 2, the bass drum will actually be heard, which will not match the performance designed in MIDI file 2.

[0063] In some embodiments, all preset music segments have the same number of bars.

[0064] In some embodiments, the music module 134 can be built using the Wwise engine.

[0065] All preset music clips (e.g., preset MIDI files) can be imported into the music module. In some embodiments, these preset music clips (e.g., preset MIDI files) can be organized into a hierarchical structure based on the types of music parameters (e.g., playing density, playing velocity, and playing style) included in the music clips. The number of layers in the hierarchical structure can be related to the types of music parameters. For example, the hierarchical structure may include multiple layers, namely a first layer, a second layer, and a third layer. The first layer is the parent layer, the second layer is a child layer of the first layer, and the third layer is a child layer of the second layer.

[0066] The first, second, and third levels are all Switch containers, each used to switch a specific music parameter. The fourth level is the specific music segment, i.e., the content being played (e.g., notes or pitch). Figure 3 This is a schematic diagram illustrating the music module configuration according to some embodiments described in this specification. For example, as... Figure 3 As shown, the first level is used to switch playing styles, for example, Figure 2 Rock 1, Rock 2, Disco 1, Disco 2, etc., with different performance styles, corresponding to... Figure 3 The first level is for playing styles like Disco_A, Disco_B, and Rock_A; the second level is used to switch the playing density, for example, Figure 2 The sparse, medium, and dense in the middle correspond to respectively Figure 3 The playing density is Density_1, Density_2, and Density_3; the third level is used to switch the playing intensity, i.e. Figure 2 The light, medium, and heavy weights correspond to... Figure 3 Intensity_1, Intensity_2, Intensity_3 in .

[0067] In some embodiments, the content of each child level within a Switch container needs to be mapped to a Switch Group. For example, Switch Groups representing playing style, playing density, and playing intensity can be created separately and added to the first-level, second-level, and third-level Switch containers, respectively. A Switch Group for playing style contains Switches such as Rock 1 and Rock 2; a Switch Group for playing density contains Switches such as sparse, medium, and dense; and a Switch Group for playing intensity contains Switches such as light, medium, and heavy.

[0068] In some embodiments, the first-level switch container includes an auxiliary MIDI file containing no musical segments, used to stop the playing of musical segments in other MIDI files; this is also called a silent MIDI segment. For example, if the user-selected MIDI segment includes different pitches C1, C2, etc., when a drum kit receives a note with pitch C1 in the MIDI segment, the audio sample of the bass drum corresponding to pitch C1 in that drum kit will be output. If a stop-play command is received while the bass drum audio sample is being output, a traditional smart drum system would directly stop the bass drum sound instead of playing the bass drum sound completely before stopping the performance. However, in this application, by switching to a silent MIDI segment, after the audio sample of the bass drum corresponding to pitch C1 is played completely, pitches C2, etc., in the user-selected MIDI segment after C1 will not continue to be played. Therefore, when the user performs the stop-play operation, the unplayed part of the bass drum audio sample (e.g., the tail note) does not disappear immediately, nor is it necessary to disrupt the original natural decay in the audio sample, thus making the ending of the performance more realistic and natural, just like a real drummer stopping playing.

[0069] In some embodiments, the hierarchical structure further includes a fourth level, which is a sublevel of the third level. This fourth level may include playlists of specific performance content. Placing each MIDI clip in a playlist enables "looping" of the MIDI clips. For example, descriptive information about the musical performance in a MIDI file might include "Rock 1 – Sparse – Light (...)" Figure 3 The playlist for the MIDI file "SmartDrummer_Rock_A_Density_1_Intensity_1" can be configured to be placed alongside "Rock 1 - Sparse - Medium" and "Rock 1 - Sparse - Heavy". Figure 3 It is located under the path "Rock_A——Density_1". It should be noted that... Figure 3 This example only demonstrates the creation of different rhythmic patterns of Rock and Disco, and only showcases a subset of Rock 1 (Rock_A). Further variations can be applied when more playing styles, density, and intensity are required.

[0070] In some embodiments, the musical parameters corresponding to the first, second, and third levels can be performance style, performance dynamics, and performance density, respectively. In other words, the second level can be used to switch performance dynamics, and the third level can be used to switch performance density.

[0071] In some embodiments, the musical parameters corresponding to the first, second, and third levels can be performance dynamics, performance density, and performance style, respectively. That is, the first level can be used to switch performance dynamics, the second level to switch performance density, and the third level to switch performance style. The musical parameters corresponding to different levels can be adjusted according to actual needs, and will not be repeated here.

[0072] The configuration of music module 134 also includes configuring the logic for the transition phases of the first, second, and third levels within music module 134. This involves configuring the timing and entry point for ending the old content and entering the new content when switching from one switch to another. For example, in the first level (e.g., performance style switching), the timing configuration for switching from one switch to another can be set to Entry Cue. Entry Cue refers to the precise moment or signal point at which the audio content begins playing. In other words, the performance style transition from one style to another is based on a specific moment; for example, the precise moment or signal point for switching from one style to another is the moment when music module 134 receives the user's performance style switching command sent by the game engine. In the second and third levels, the timing configuration for switching from one switch to another (e.g., switching performance volume or density) can be set to "Same Time as Playing Segment," meaning a synchronized change occurs at the end of the current audio segment or at a specific point within it. In this way, the triggering of audio events is combined with the natural end or rhythmic point of the currently playing content. The purpose is to allow the intelligent drummer to start playing from the beginning only when it begins playing, and any switching during the process can take effect immediately while continuing the previous beat. For example, if the player has already played to the end of the second measure in MIDI file 1 before switching, then after switching to MIDI file 2, it will continue playing the third measure in MIDI file 2 instead of returning to the beginning of the first measure. This is crucial for a smooth ensemble and performance experience, especially when you want to change the rhythm pattern during a section of a song, without disrupting the original beat of the music due to the switch.

[0073] In some embodiments, the configuration of the music module 134 further includes establishing an association between the music module and the sound effects module, such that the audio samples of the sound effects module 136 are driven by music segments sent by the music module 134. For example, in the first level of the music module 134, the target of receiving the MIDI file is set to the first level of the sound effects module 136, thus establishing an association between the music module 134 and the sound effects module 136, so that the specific content of the sound effects module 136 is triggered by MIDI music segments.

[0074] The sound effects module 136 is used to provide multiple preset audio datasets and to provide switching functionality between different audio datasets. In some embodiments, the sound effects module 136 is configured to receive a target music segment output by the music module 134, and in response to receiving the target music segment, play the target audio dataset based on the target music segment.

[0075] Each preset audio dataset includes audio samples of different drum parts with the same sound effect style. Different drum parts with the same sound effect style can also be called a drum kit.

[0076] A drum kit consists of a combination of multiple different types of drums and / or cymbals, such as snare drum, bass drum, tom-tom, hi-hat, crash cymbal, stacked cymbals, Chinese cymbals, splash cymbals, bass drum, etc. Each drum kit corresponds to a specific sound style; that is, the timbre produced by the different drum parts within the kit will give the music played through those drum parts a corresponding sound style.

[0077] Different types of drums or cymbals can produce different timbres. For example, the snare drum produces a crisper sound with a more pronounced springy sound; the bass drum produces a deeper sound.

[0078] The same type of drum or cymbal can produce different timbres. Different playing techniques, striking positions, and equipment can also produce different timbres. For example, striking the center of the drumhead usually produces a deeper, more viscous sound; striking near the edge usually produces a brighter, higher-frequency sound. Similarly, using different drumsticks or brushes to strike the same drum will produce different timbres; further, using a brush instead of a drumstick will produce a softer, gentler sound, often used in jazz. Finally, different striking techniques and force control will significantly alter the timbre.

[0079] An audio sample can be defined as the sound of a single drum part (e.g., a drum or cymbal) producing a particular timbre. Drum parts of the same type (e.g., a drum or cymbal) can produce different timbres, thus generating multiple independent audio samples. For example, audio samples for a snare drum could include a sample produced by hitting the rim and a sample produced by hitting the center, each corresponding to a different timbre.

[0080] In some embodiments, based on the timbre characteristics of different types of drums or cymbals, the sounds produced by different drums or cymbals can be classified into different sound effect styles; that is, different drums or cymbals are suitable for different sound effect styles. Sound effect styles can include, for example, pop, jazz, heavy metal, rock, hip-hop, etc. In some embodiments, the sounds produced by different types of drums or cymbals can be classified into different drum groups according to sound effect styles. For example, a pop style may include drum kit 1 (snare drum 1, bass drum 1, tom-tom 1, hi-hat 1, cymbal 1, and stacked cymbals 1), where the sound timbre of all the drums or cymbals in drum kit 1 is suitable for pop style; a jazz style may include drum kit 2 (snare drum 2, bass drum 2, tom-tom 2, hi-hat 2, cymbal 2, and stacked cymbals 2), where the sound timbre of all the drums or cymbals in drum kit 2 is suitable for jazz style; a heavy metal style may include drum kit 3 (snare drum 3, bass drum 3, tom-tom 3, hi-hat 3, cymbal 3, and stacked cymbals 3), where the sound timbre of all the drums or cymbals in drum kit 3 is suitable for heavy metal style.

[0081] In some embodiments, different drum kits can be named using sound effects styles. For example, drum kit 1 is named pop style; drum kit 2 is named jazz style; and drum kit 3 is named heavy metal style.

[0082] In some embodiments, the sound effects module 136 can be built using the Wwise engine.

[0083] Figure 4 This is a schematic diagram illustrating the configuration of the sound effect module according to some embodiments described in this specification. For example... Figure 4 As shown, the audio samples of each drum kit (i.e., the preset audio dataset) are imported into the Actor-Mixer Hierarchy of the Wwise engine, which is the sound effects module. Each drum kit is named according to its musical style. For example, the drum kit name field includes Disco, Funk, and Rock, indicating that this is a drum kit tone suitable for Disco, Funk, and Rock styles, respectively. In actual application, the naming rules can be defined according to the actual situation.

[0084] Figure 4The first layer in the structure, Virtus1llnstrument-Drum (the parent layer), is a Switch container; a Switch container, also called a Switch Group, is used to organize and manage a group of related Switches. Each Switch corresponds to a drum kit. By selecting different Switches, different sound effects or music can be switched. The second layer consists of several mixing containers representing drum kits with different timbres (e.g., Virtus1llnstrument-Drum-Rock, Disco, Fun). The third layer contains audio samples of the individual drums in the different timbres within the mixing containers. These three layers correspond to... Figure 2 The text refers to a "Switch container containing multiple drum kits," "Acoustic Drum 1 (mixing container)," and "each triangle (audio sample) within Acoustic Drum 1." The maximum number of sounds produced by the mixing container can be limited to one. In some embodiments, the Hihat in the Rock drum kit produces two sounds, including the audio samples corresponding to HihatOpen (open cymbal) and HihatClose (close cymbal). Adding such a mixing container to the hihat is to ensure that whenever the closed cymbal is triggered, it interrupts the adjacent open cymbal, which is more consistent with the actual operation mechanism of a real hihat.

[0085] After importing the audio samples, three configurations are required for the first-layer Switch container: the specific Switch assignment, the mapping between MIDI pitch and audio samples, and the mapping between MIDI performance velocity and audio sample volume.

[0086] First, the specific Switch configuration includes creating a Switch Group for switching between different drum kits. Each switch within a Switch Group is assigned a corresponding drum kit tone. Each switch can be named according to the playing style associated with its drum kit. This way, when the game engine calls the names of the switches, it can switch to the corresponding drum kit tone (e.g., ...). Figure 2 The "Switch" for switching drum kit sounds is displayed in the middle right area.

[0087] The mapping of music clip pitches to audio datasets involves mapping different types of drums or cymbals in each drum kit to individual pitches, establishing a mapping relationship between the pitches in the music clip and the corresponding audio samples for each drum kit. Different types of drums or cymbals in each drum kit correspond to individual pitches, and drums or cymbals of the same type correspond to the same pitch. Similarly, different pitches in a music clip correspond to different timbres or drum types, and the same timbre (i.e., drum or cymbal) is at the same pitch. The mapping relationship between drums and pitches in each drum kit is configured to be the same as the mapping relationship between pitches and drums in the music clip, thus establishing a mapping relationship between music clip pitches (e.g., MDII pitches) and audio datasets (e.g., audio samples). For example, if D1 in the first MIDI file is configured to correspond to a snare drum, then D1 in all other MIDI files must also be configured to correspond to a snare drum or a snare-like function timbre (e.g., a clapping timbre in an electronic drum kit), and cannot be anything else, such as a bass drum, hi-hat, etc. Similarly, the snare drum in each drum kit corresponds to the pitch D1. Based on the mapping relationship between the pitch of a music clip and the audio dataset, different drums or cymbals in a drum kit can be driven by the music clip to play the corresponding audio samples. For example, when playing a D1 pitch or note in a MIDI file, the sound effects module will output the audio sample corresponding to the snare drum in the drum kit that has the same pitch D1.

[0088] The mapping between the performance intensity of a musical segment and the volume of an audio dataset refers to the correspondence between the performance intensity of the musical segment and the volume of the audio sample. In some embodiments, the mapping between the performance intensity of the musical segment and the volume of the audio dataset can be represented by functions, lookup tables, etc. The mapping between the performance intensity of the musical segment and the volume of the audio dataset can be linear or non-linear. For example, a Real-Time Parameter Control (RTPC) curve describing the relationship between the volume of the sample and the MIDI performance intensity can be configured in the first-level Switch container. The specific curvature and slope of the RTPC curve can be selected and configured as needed. The greater the performance intensity of the musical segment, the greater the volume of the played audio dataset, and vice versa. This is because the intensity of most instruments needs to change dynamically, and the notes in each musical segment must contain performance intensity information. When drums need to be played densely, one or more smaller taps are often added on the light taps before and after the heavy taps. Therefore, adding this configuration can better reproduce the original performance effect designed in the musical segment, especially for densely played sections.

[0089] The event triggering module 132 is used to control audio output. In some embodiments, after receiving control commands from the game engine 120, the event triggering module 132 can be triggered to execute related audio control actions, such as controlling audio playback to start or stop, beats per minute (BPM), MIDI playback rate, etc. In some embodiments, the event triggering module 132 includes an output parameter adjustment unit. The output parameter adjustment unit can be used to adjust audio output parameters. Audio output parameters include audio playback rate, audio playback speed, and music clip playback speed (i.e., MIDI playback rate). Beats per minute (BPM) refers to the number of beats played per minute (i.e., audio playback rate), used to describe the speed of music; the higher the value, the faster the music. Audio playback speed refers to the speed of an audio file or stream during playback. MIDI playback rate refers to the trigger interval between different audio samples; the smaller the trigger interval between two adjacent audio samples, the higher the MIDI playback rate. There is a direct relationship between MDI playback rate and BPM and audio playback speed. When the MIDI playback rate increases, the BPM and audio playback speed also increase.

[0090] In some embodiments, the output parameter adjustment unit can adjust the BPM based on the correspondence between MIDI playback rate and BPM, thereby adjusting the MIDI playback rate. By modifying the MIDI playback rate in real time, the effect of real-time BPM adjustment is achieved. This method of modifying the MIDI playback rate differs from directly modifying the audio playback rate. Modifying the MIDI playback rate actually modifies the trigger interval of audio materials such as the kick drum and snare drum, without changing the playback rate of the audio material itself, and therefore does not affect sound quality. In some embodiments, the correspondence between MIDI playback rate and BPM can be linear; for example, the ratio between MIDI playback rate and BPM is a constant. This constant is related to the settings of the audio playback system.

[0091] In some embodiments, the output parameter adjustment unit can obtain the target BPM input by the user through the user terminal, determine the target MIDI playback rate based on the target BPM and the correspondence between the MIDI playback rate and BPM, and thus adjust the BPM to the target BPM by adjusting the MIDI playback rate to the target MIDI playback rate.

[0092] In some embodiments, BPM can be adjusted using RTPC technology. RTPC (Real-Time Parameter Control) is a technology for dynamically adjusting audio parameters, allowing audio designers and developers to adjust audio parameters based on real-time game status or user interaction to achieve a more immersive and interactive audio experience. RTPC technology enables real-time adjustment of audio attributes based on in-game variables such as character speed, height, and health.

[0093] RTPC (Real-Time Parameter Control) is a type of control information that can exert a set influence on parameters (such as MIDI playback rate) in the Wwise engine through a control source. It consists of a control source name and its corresponding value. For example, the control source can be an in-game variable (such as character speed, height, health, etc.), and the parameters in the Wwise engine include MIDI playback rate.

[0094] In some implementations, the control source can be the BPM (i.e., target BPM) value set by the user within the game.

[0095] In some embodiments, the output parameter adjustment unit can be constructed using a hierarchical structure. This hierarchical structure includes a first level and a second level. The first level is used to set the mapping relationship between BPM and MIDI playback rate. The mapping relationship between BPM and MIDI playback rate can be linear. For example, the ratio between BPM and MIDI playback rate can be set to a constant such as 100 or 120. The second level is used to configure the RTPC based on the mapping relationship between BPM and MIDI playback rate, thereby allowing the MIDI music clip playback speed (i.e., MIDI playback rate) to be determined according to the target BPM value based on the mapping relationship between BPM and MIDI playback rate.

[0096] like Figure 5 As shown, Figure 5 This diagram illustrates BPM control based on RTPC technology according to some embodiments described in this specification. For example, at the first level, the ratio between BPM and MIDI playback rate is first set to 100; other values ​​are also possible, but 100 is recommended for ease of calculation. Then, RTPC is configured at the second level. RTPC represents the mapping relationship between BPM (horizontal axis) and MIDI playback rate (vertical axis). The value of RTPC can range from 30 to 240, depending on actual needs. The RTPC value will be sent by the game engine during game runtime based on user input or the game engine's default settings. For example, the RTPC value can be based on the user-set BPM value (i.e., the target BPM value, e.g., ...). Figure 2The "BPM slider value" is determined in the RTPC configuration. The RTPC's horizontal and vertical coordinate mapping can be set such that when the X-axis equals 30, Y equals 0.3, and when the X-axis equals 240, Y equals 2.4, etc., meaning the ratio between BPM (horizontal axis) and MIDI playback rate (vertical axis) is 100.

[0097] If the user sets the BPM value (the horizontal axis of the RTPC, i.e., the target BPM) to 40, the vertical axis of the RTPC will be 0.4, meaning the MIDI playback rate is 0.4. Since we previously set the overall BPM of the first level to 100, the actual playback speed will be 0.4 * 100 = 40 BPM, which is the BPM set by the user. The BPM control configuration is now complete.

[0098] In some embodiments, the Source of the MIDI Clip Tempo in the music module 134 must be set to Hierarchy. This ensures that the 100 BPM setting in the BPM control takes effect, guaranteeing that the actual BPM played matches the user-set BPM. MIDI Clip Tempo refers to the tempo value associated with a MIDI clip, defining the rhythm when playing MIDI information.

[0099] In some embodiments, the event triggering module 132 includes a music parameter conversion unit. The music parameter conversion unit is used to convert music parameters (e.g., playing density, playing dynamics, etc.). The music parameter conversion unit receives music parameter information set or input by the user in the game engine, determines target music parameters based on the music parameter information, and sends the target music parameters to the music module 134 to determine a target music segment (e.g., a target MIDI file), which includes the target music parameters. In some embodiments, the music parameter conversion unit can be integrated into the music module 134. In some embodiments, the music parameter conversion unit can be a separate module independent of the event triggering module 132.

[0100] In some embodiments, the music parameter conversion unit can acquire music parameter information set or input by the user. This music parameter information may include identifier information of the music parameter (e.g., the name of the playing style), the numerical range corresponding to the music parameter (e.g., RTPC value), or the numerical value corresponding to the music parameter. The music parameter conversion unit can determine the target music parameter based on the music parameter information input by the user and send the determined target music parameter to the music module 134. The music module 134 determines the target music segment (e.g., the target MIDI file) based on the received target music parameter. This target music segment (e.g., the target MIDI file) is the MIDI file that the user selects to play.

[0101] In some embodiments, for a performance style, the user can input a name for the performance style, and the music parameter conversion unit can determine the target performance style based on the name of the performance style input by the user.

[0102] In some embodiments, for musical parameters such as playing density or playing intensity, the user can input a numerical value corresponding to the playing density or playing intensity, and the musical parameter conversion unit can determine the target playing density and target playing intensity based on the user-input numerical value. In some embodiments, the musical parameter conversion unit can store the mapping relationship between numerical values ​​and playing density or playing intensity, and the musical parameter conversion unit can determine the target playing density and target playing intensity based on the mapping relationship and the user-input numerical value.

[0103] In some embodiments, for musical parameters such as playing density or playing intensity, the user can input the corresponding level of playing density or playing intensity (e.g., "light, medium, heavy" for playing intensity or "sparse, medium, dense" for playing density, etc.). The musical parameter conversion unit can determine the target playing density and target playing intensity based on the level input by the user.

[0104] In some embodiments, the selection or switching of music parameters can be controlled using RTPC technology. The RTPC value comes from the numerical value corresponding to the music parameters input by the user in the game engine. The event triggering module can determine the corresponding target music parameters based on the RTPC value. These determined target music parameters constitute the target music segment (e.g., the target MIDI file) selected by the user.

[0105] For example, an associated RTPC can be created for each of the two Switch Groups representing performance density and performance intensity. The selection or switching of the Switch within these two Switch Groups is not done by directly sending the Switch value, but rather by sending the RTPC value, which is then converted into a Switch value (corresponding to the specific performance density or performance intensity) in the Wwise engine. Avoiding direct Switch sending simplifies program development considerably, enabling the X / Y controller to function more easily, and also facilitating future expansion and flexibility.

[0106] For example, Figure 6 This is a schematic diagram illustrating the selection or switching of RTPC-based performance density according to some embodiments of this application. As shown in the figure, if the RTPC value range corresponding to the performance density is set to 0 to 100, the RTPC value can be obtained from the game engine through the user's X / Y controller (e.g., Figure 8 The value corresponding to the performance density input on the X-axis (as shown) is processed... Figure 6 After configuration, when the X-axis transmission value of the X / Y controller is in the range of 0 to 33.333, the switch representing the performance density will switch to "Density_1"; from 33.333 to 66.666, it will switch to "Density_2"; and from 66.666 to 100, it will switch to "Density_3". Similarly, the switch representing performance intensity will be associated with the RTPC representing the Y-axis value. For example, the RTPC value range corresponding to the performance density can be set to 0 to 100, and the RTPC value can be obtained from the game engine through the user's input via the X / Y controller (e.g.,...). Figure 8 The value corresponding to the performance density input on the Y-axis (as shown) is processed... Figure 6 After configuration, when the Y-axis transmission value of the X / Y controller is in the range of 0 to 33.333, the Switch representing the playing intensity will switch to "light (Intensity_1)", from 33.333 to 66.666 it will switch to "medium (Intensity_2)", and from 66.666 to 100 it will switch to "heavy (Intensity_3)".

[0107] In some embodiments, the event triggering module 132 further includes a control unit. The control unit is used to receive operation commands input by the user sent by the game engine, and to perform audio-related actions based on the operation commands, such as playing a certain audio, stopping a certain audio, resetting certain parameters, etc.

[0108] In some embodiments, the control unit may include events for starting and stopping playback. The target for starting playback includes the highest-level switch container in the music module. The target for stopping playback is the switch corresponding to a MIDI file in the music module that does not contain an audio segment. This method ensures that the instrument's final note plays out naturally rather than being forcibly interrupted when stopping.

[0109] Game engine 120 is responsible for implementing the terminal's game functions, such as rendering 3D or 2D graphics, playing sound effects and music, processing user input, and executing game logic. The game engine includes a user interface through which users access game functions and / or audio playback. For example, the game engine allows users to operate, navigate, and perform actions within the game environment.

[0110] The user interface includes an interactive control module. This interactive control module may include a playback control module 122, a music selection module 124, and a sound effect selection module 126.

[0111] The playback control module 122 receives user operation commands and sends them to the audio engine's event triggering module, music module, and / or audio module to execute the corresponding operations. For example, the playback control module 122 may include a play button, a stop button, a BPM selector, etc. Figure 7 These are schematic diagrams of user interfaces based on some embodiments described in this specification. For example... Figure 7 As shown, the user interface may include a play button 708, a stop button 710, and a BPM selector 712.

[0112] When the event triggering module 132 in the audio engine 130 receives the operation instruction from the game engine 120 that the user clicks the play button 708, it triggers a playback start event. In some embodiments, after the play button is clicked to enter the playback state, the play button will become a "stop button".

[0113] When the event triggering module 132 in the audio engine 130 receives the operation command from the game engine 120 that the user clicks the stop button 710, a stop playback event is triggered. In some embodiments, after the stop button is clicked and the system enters a stop state, the stop button will become a "play button".

[0114] The BPM selector 712 is used by the user to specify the desired target BPM. The game engine can send the user-specified target BPM to an event-triggered module (e.g., an output parameter adjustment unit), which adjusts the MIDI playback rate based on the user-specified target BPM to adjust the BPM to the target BPM. For a more detailed description of how the event-triggered module adjusts the MIDI playback rate based on the user-input target BPM, see other sections of the specification.

[0115] For example, Figure 9 This is a schematic diagram of a BPM selector according to some embodiments of this application. The BPM value range is consistent with the range set in the audio engine, for example, 30 to 240. The user can change the value by moving the cursor or clicking the plus or minus buttons, and simultaneously send the value to the audio engine. The audio engine will use the user input to control the playback rate multiplier (e.g., ...). Figure 6 As shown in the image, this allows for real-time control of the smart drummer's BPM. The rectangle to the left of the BPM selector represents the ratio between the BPM and the MIDI playback rate, which can be set to a default value, such as any value between 80 and 120, like 100.

[0116] The music selection module 124 receives music parameter information input by the user, such as performance style information, performance density information, and performance velocity information, and sends the user-input music parameter information to the music module 134 or event triggering module 132 (e.g., music parameter conversion unit) in the audio engine 130 to determine the target music segment (e.g., target MIDI file). In some embodiments, the music selection module 124 may include a voice input button, allowing the user to input music parameter information via voice input.

[0117] In some embodiments, the music selection module 124 includes a style selector and an X / Y controller. For example, Figure 7 The style selector 702 and the X / Y controller 704 are shown.

[0118] The style selector 702 is an input device for selecting a performance style. In some embodiments, the style selector can select various main performance styles such as rock and disco via tabs. Each main performance style contains sub-performance styles; for example, rock might include rock 1 and rock 2. The performance style names can be customized as needed, such as punk rock, metal rock, funk rock, etc. When a specific sub-style (i.e., the target performance style) is selected, the corresponding switch name for that sub-style is sent to the audio engine. The music module in the audio engine, for example, a "Switch container with multiple MIDI segments," will switch to that sub-performance style. In some embodiments, the style selector can also select the target performance style via voice.

[0119] In some embodiments, a default sub-style can be set to ensure that a style is selected by default even if the user does not make any selection.

[0120] The X / Y controller 704 is an input device for controlling two different parameters, such as playing density and playing velocity. For example, the X-axis of the X / Y controller can be used to control playing density, and the Y-axis can be used to control playing velocity. As the X-axis value increases, the playing density increases; as the Y-axis value increases, the playing velocity also increases.

[0121] In some embodiments, the X / Y controller can be implemented via a touchscreen. Users can control it by directly dragging their fingers across the screen of devices such as smartphones, tablets, and touchscreen displays. Figure 8 As shown, Figure 8 This is a schematic diagram of an X / Y controller according to some embodiments of this application. The X / Y controller may include a control point P. The control point P can be moved on the X / Y controller interface by the user, and corresponding coordinates, i.e., X-axis values ​​and Y-axis values, will be present at each position. The X-axis values ​​and Y-axis values ​​can represent the values ​​corresponding to the playing density and playing intensity, respectively. The value range of both the X-axis and Y-axis must be the same as the value range set by the event triggering module in the audio engine, for example, 0 to 100. Whenever the position of the control point P is moved, the X-axis and Y-axis values ​​are sent to the audio engine in real time in RTPC format to control the switches for switching playing density and playing intensity.

[0122] In some embodiments, the default position of control point P can be set at the center of the X / Y controller interface. In some embodiments, whenever the user adjusts the position of control point P, control point P can return to the default position in real time, facilitating the user's next adjustment of the position of control point P.

[0123] In some embodiments, the user can be prompted by text, voice, or other means to “move the control point horizontally to adjust the playing density (i.e., playing complexity), the further to the right the playing becomes more complex; move the control point vertically to adjust the playing intensity, the further up the playing becomes stronger.”

[0124] The sound effect selector is an input device used to select the sound effect style. For example, Figure 7 The sound effect selector 706 is shown. In some embodiments, the sound effect selector allows selection of various sound effect styles such as rock and disco via tabs. When a specific sound effect style (i.e., the target sound effect style) is selected, the corresponding Switch for that sound effect style is sent to the sound effect module in the audio engine. For example, a "Switch container with multiple drum kits" will switch to the drum kit corresponding to that sound effect style. In some embodiments, the sound effect selector allows selection of the target sound effect style via voice. It is basically the same as the previous style selector, the only difference being that it usually doesn't need two levels, but only one. Refer to the style selector for details; further explanation is unnecessary.

[0125] The memory can be used to store system data and / or instructions, such as preset music clips, preset sound effect datasets, mappings between audio playback rates and MIDI playback rates, mappings between music clip pitches and audio datasets, etc. The memory may include a storage module. In some embodiments, the storage module may be integrated into the music module, sound effect module, etc.

[0126] Figure 2 These are schematic diagrams illustrating specific embodiments of an audio playback system according to some examples of this specification. For example... Figure 2 As shown, the game engine includes a play button, a stop button, a style selector, an X / Y controller, a tone selector, and a BPM slider. The style selector can be displayed as tabs in the user interface. Users can select their desired playing style based on the name of the style in the tab. For example, the rock style includes subsets such as Rock 1, Rock 2, etc., and the disco style includes subsets such as Disco 1, Disco 2, etc. When the user clicks on the rock style, the subsets under the rock style will expand in the user interface for selection. The X / Y controller includes X-axis and Y-axis values, which users can input for playing density and dynamics, i.e., RTPC values, respectively. The tone selector (also called the sound effect selector) can be displayed as tabs in the user interface, such as Acoustic Drum 1, Acoustic Drum 2, Electronic Drum 1, etc. Users can select the drum kit for a sound effect style by clicking on the corresponding location in the user interface.

[0127] When the user inputs performance style information through the style selector, the game engine can call up the name of each performance style Switch (i.e., Figure 2 When the user inputs performance velocity and density information via the X / Y controller, the audio engine can convert the RTPC value, including the performance density information, into the name of the switch corresponding to the performance density (i.e., the name of the switch). Figure 2 If the RTPC is converted to a switch for selecting MIDI segments with different note densities, the music module can switch to the corresponding MIDI segment with the appropriate performance density. The audio engine can convert the RTPC value, which includes performance velocity information, input by the user into the name of the switch corresponding to the performance velocity (i.e., the name of the switch). Figure 2 If the RTPC is converted to a switch that selects MIDI clips with different playing strengths, the music module can switch to the MIDI clip corresponding to that playing strength. Using this method, a target MIDI music clip can be selected from a switch container containing multiple MIDI music clips; this target MIDI music clip is the one selected by the user through the style selector and X / Y controller.

[0128] When a user inputs sound effect style information via the sound selector, the game engine can call up the Switch's name for each sound effect style (i.e., Figure 2 If you select a switch in the sound effects module (e.g., a switch for changing drum kit sounds), the sound effects module can switch to the corresponding sound effect style. The sound effects module contains multiple drum kit switches, each containing multiple audio samples. Switching to the corresponding sound effect style in the sound effects module determines the audio sample set for the user-selected drum kit.

[0129] When the user inputs the desired BPM value (i.e., RTPC) via the BPM slider in the game engine, the music module determines the MIDI playback rate and adds this determined MIDI playback rate to the target MIDI segment before sending it to the sound effects module. The sound effects module can switch to the corresponding switch based on the user-selected sound effect style. When the sound effects module receives the target MIDI segment sent by the music module, it can output an audio sample set of the drum kit selected by the user based on the target MIDI segment.

[0130] Figure 10 The audio playback method is illustrated in the schematic diagram of some embodiments shown in this specification. The audio playback method can be executed by processing device 112 or electronic device 1200. The audio generation method 1000 includes the following steps.

[0131] Step 1002: Obtain the first input via the user interface. In some embodiments, step 1002 may be performed by the music selection module 124 of the game engine 120 in the processing device 112 or by the acquisition module 1102 in the audio playback device 1100.

[0132] The first input includes user-defined music parameter information. This music parameter information may include performance style information, performance density information, performance dynamics information, etc. For example, the music parameter information may include an identifier for the performance style (e.g., name), an identifier for the performance density (e.g., value or level), and an identifier for the performance dynamics (e.g., value or level). In some embodiments, the user can select the desired performance style from multiple preset performance styles using a style selector in the user interface to generate performance style information. In some embodiments, the user can input the values ​​corresponding to the performance density and performance dynamics using the X / Y controller in the user interface to generate performance density and performance dynamics information. In some embodiments, the user can input music parameter information via voice input through the user interface. Further descriptions of music parameter information and the user interface can be found in other parts of the specification, for example... Figure 1 , 2 as well as Figure 7-9 A detailed description.

[0133] Step 1004: Obtain a second input via a user interface. In some embodiments, step 1004 may be performed by the sound effect selection module 126 of the game engine in the processing device 112 or by the acquisition module 1102 in the audio playback device 1100.

[0134] The second input includes user-defined sound effect data. This sound effect data may include music style information. For example, it may include an identifier for the sound effect style (e.g., a name). In some embodiments, the user can select a desired sound effect style from multiple preset styles using a sound effect selector in the user interface to generate the sound effect data. In some embodiments, the user can input the sound effect data (e.g., a sound effect style name) via voice input through the user interface. Further descriptions of sound effect styles and the user interface can be found in other parts of the specification, for example... Figure 1 , 2 as well as Figure 7-9 A detailed description.

[0135] Step 1006: Determine the target music segment from multiple preset music segments based on the first input. In some embodiments, step 1006 may be performed by the playback control module 112 of the game engine in the processing device 112 or by the music segment determination module 1104 in the audio playback device 1100.

[0136] In some embodiments, each of the multiple preset music segments can have multiple different music parameters, such as descriptions of playing style, playing density, and playing dynamics. Different combinations of music parameters can produce different musical expressions. In some embodiments, the multiple music parameters can be saved in MIDI file format, that is, each music segment is stored in a MIDI file, also known as a MIDI music segment or MIDI segment.

[0137] In some embodiments, the target performance style corresponding to the target music segment can be determined based on the performance style information in the first input. The target performance density corresponding to the target music segment can be determined based on the performance density information in the first input. The target performance dynamics corresponding to the target music segment can be determined based on the performance dynamics information in the first input. Further, a target music segment matching the target performance style, target performance density, and target performance dynamics can be selected from multiple preset music segments based on the target performance style, target performance density, and target performance dynamics.

[0138] In some embodiments, for example, a first candidate music segment with the target performance style can be selected based on the performance style information (or target performance style) in the first input, a second candidate music segment can be selected from the first candidate music segment based on the performance density information (or target performance density) in the first input, and finally the target music segment can be selected from the second candidate music segment based on the performance dynamics information (or target performance dynamics) in the first input.

[0139] In some embodiments, for example, a first candidate music segment with the target performance style can be selected based on the performance style information (or target performance style) in the first input. Then, a second candidate music segment can be selected from the first candidate music segments based on the performance dynamics information (or target performance dynamics) in the first input. Finally, the target music segment can be selected from the second candidate music segments based on the performance density information (or target performance density) in the first input. The specific method for determining the target music segment may be related to the hierarchical construction of the music module 134 in the audio engine 130. For example, the music module 134 includes a first level, a second level, and a third level. If the first level includes a switch group corresponding to the performance density, the second level includes a switch group corresponding to the performance intensity, and the third level includes a switch group corresponding to the performance style, then the music module 134 can first filter out a first candidate music segment with the target performance density based on the performance density information (or target performance density) in the first input, then filter out a second candidate music segment with the target performance intensity from the first candidate music segment based on the performance intensity information (or target performance intensity) in the first input, and finally filter out a target music segment with the target performance style from the second candidate music segment based on the performance style information (or target performance style) in the first input.

[0140] In some embodiments, the playing density or playing intensity can be divided into multiple levels. For example, playing density can be divided into "sparse," "medium," and "dense," and playing intensity can be divided into "light," "medium," and "heavy." The music parameter information in the first input may include the values ​​of playing density and playing intensity. The event triggering module 132 can determine the level of playing density desired by the user (i.e., the target playing density) based on the mapping relationship between the value of playing density and the level of playing density. The event triggering module 132 can determine the level of playing intensity desired by the user (i.e., the target playing intensity) based on the mapping relationship between the value of playing intensity and the level of playing intensity. For more details on determining the target playing density and target playing intensity based on the values ​​of playing density and playing intensity, please refer to the descriptions in other parts of the specification.

[0141] In some embodiments, the music parameter information in the first input may include the level of performance density and the level of performance intensity. The event triggering module 132 may specify the level of performance density and the level of performance intensity in the first input as the target performance density and the target performance intensity, respectively.

[0142] In some embodiments, the event triggering module 132 or the music segment determination module 1104 can send the determined target music segment to the music module 134 or the output module 1108 respectively.

[0143] Step 1008: Determine the target sound effect dataset from a plurality of preset sound effect datasets based on the second input. In some embodiments, step 1008 may be performed by the sound effect module 136 of the game engine in the processing device 112 or by the sound effect data determination module 1106 in the audio playback device.

[0144] In some embodiments, a target sound effect dataset matching the sound effect style can be determined from multiple preset sound effect datasets based on the sound effect style information in the second input.

[0145] Each of the multiple preset sound effect datasets can include audio samples of different drum parts with the same sound effect style. Different drum parts with the same sound effect style can also form a drum kit. Each drum part in the same drum kit corresponds to one audio sample. For more detailed descriptions of the audio samples, please refer to the detailed descriptions in other parts of this manual, such as... Figure 1 and 2 And its detailed description.

[0146] In some embodiments, each preset sound effect dataset may have an identifier to identify the sound effect style corresponding to each preset sound effect dataset. The identifier for each preset sound effect dataset may be the name of the sound effect style corresponding to that preset sound effect dataset, such as rock, disco, etc., or named after the category of drums with that sound effect style (e.g., acoustic drums, electronic drums, etc.). The target sound effect dataset can be determined by comparing the name of the sound effect style in the sound effect style information in the second input with the sound effect style names corresponding to multiple preset sound effect datasets.

[0147] Step 1010: Drive the playback of the target sound effect dataset based on the target music segment. In some embodiments, step 1010 may be performed by the sound effect module 136 in the processing device 112 or by the output module 1108 in the audio playback device.

[0148] In some embodiments, the target music segment includes a temporal arrangement of different notes. Each note corresponds to a pitch. The target sound effect dataset includes multiple different audio samples, each corresponding to a drum part and a pitch. When the sound effect module receives a specific note from the target music segment, an audio sample of the drum part with the same pitch as that specific note will be output and played. A detailed description of how the target music segment drives the playback of sound effect data can be found elsewhere in the specification.

[0149] In some embodiments, a third user input may be obtained, which may include a user-set target playback rate (i.e., a target BPM value). The target music segment playback rate (i.e., the target MIDI playback rate) is determined based on the mapping relationship between audio playback rate and music segment playback speed (e.g., MIDI playback rate) and the target beat count; and the target sound effect dataset is played based on the target music segment playback speed. For a detailed description of the music segment playback speed determination, please refer to other parts of the specification.

[0150] Figure 11 The schematic diagram of an audio playback device shown in some embodiments of this specification is illustrated. The audio playback device 1100 is used to execute the audio playback method 1000. The audio playback device 1100 may include an acquisition module 1102, a music segment determination module 1104, a sound effect data determination module 1106, and an output module 1108.

[0151] The acquisition module 1102 is used to acquire a first input from the user through the user interface and a second input from the user through the user interface. The music segment determination module 1104 is used to determine a target music segment from multiple preset music segments based on the first input. The sound effect data determination module 1106 is used to determine a target sound effect dataset from multiple preset sound effect datasets based on the second input. The output module 1108 is used to drive the target sound effect dataset to play based on the target music segment.

[0152] This application also provides an electronic device for audio playback. For example... Figure 12 As shown, the electronic device 1200 includes a processor 1202 and a memory 1204 for storing a program for playing music. After the device is powered on and the program for playing music is run by the processor, the following steps are performed: obtaining a first input from the user through a user interface, the first input including music parameter information; obtaining a second input from the user through a user interface, the second input including sound effect data information; determining a target music segment from multiple preset music segments based on the first input; determining a target sound effect dataset from multiple preset sound effect datasets based on the second input; and outputting the target sound effect dataset based on the target music segment.

[0153] This application provides a computer-readable storage medium storing a program for playing music. The program is executed by a processor and performs the following steps: obtaining a first user input through a user interface, the first input including music parameter information; obtaining a second user input through a user interface, the second input including sound effect data information; determining a target music segment from a plurality of preset music segments based on the first input; determining a target sound effect dataset from a plurality of preset sound effect datasets based on the second input; and outputting the target sound effect dataset based on the target music segment.

[0154] Some embodiments of this specification also provide a computer program product, including computer instructions that, when at least a portion of the computer instructions are executed by a processor, can implement this specification. Figure 10 The method is illustrated. In some embodiments, the computer program product may relate only to computer instructions, which may be carried on a storage medium or processing device. In other embodiments, the computer program product may also be a storage medium or processing device containing the aforementioned computer instructions. The processing device may include one or more processors, and the storage medium.

[0155] In some embodiments, the processor may be a combination of one or more of the following processors: central processing unit (CPU), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), graphics processing unit (GPU), physical processing unit (PPU), digital signal processor (DSP), field-programmable gate array (FPGA), programmable logic device (PLD), programmable logic controller (PLC), reduced instruction set computer (RISC), and microprocessor.

[0156] In some embodiments, the storage medium may include one or more combinations of the following: mass storage, removable storage, volatile read-write memory, and read-only memory (ROM). Exemplary mass storage may include disks, optical disks, solid-state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, compressed hard disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary RAM may include dynamic random access memory (DRAM), dual data rate synchronous dynamic random access memory (DDRSDRAM), static random access memory (SRAM), silicon controlled retrieval memory (T-RAM), and zero-capacitance memory (Z-RAM), etc. Exemplary read-only memory may include masked read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compressed hard disk read-only memory (CD-ROM), and digital multifunction hard disk read-only memory, etc.

[0157] While this application discloses preferred embodiments as described above, it is not intended to limit the scope of this application. Any person skilled in the art can make possible variations and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims. For example, the electronic device 1200 also includes one or more input / output interfaces, network interfaces, and memory. Memory may include non-permanent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media. Computer-readable media includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage media, or any other non-transferable medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0158] For more details on each module, please refer to the relevant descriptions in other parts of the specification, which will not be repeated here. It should be understood that the systems and modules shown in this specification can be implemented in various ways. For example, in some embodiments, the systems and modules can be implemented using hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated hardware design. Those skilled in the art will understand that the methods and systems described above can be implemented using computer-executable instructions and / or included in the control code of a processor, such as on a media such as a disk, CD, or DVD-ROM, or in the memory of a programmable device. The systems and modules in this specification can be implemented not only with hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or programmable hardware devices such as field-programmable gate arrays and programmable logic devices, but also with software, for example, executed by various types of processors, or with a combination of the aforementioned hardware circuits and software (e.g., firmware).

[0159] It should be noted that the above description of the system and its modules is for convenience only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules without departing from these principles to form subsystems connected to other modules. Alternatively, some modules may be split to obtain more modules or multiple units under a single module. Such modifications are all within the scope of this specification.

[0160] The beneficial effects that the embodiments of this specification may bring include, but are not limited to: (1) Each drum set is divided into independent audio materials (i.e., audio samples) according to the specific timbre of the bass drum, snare drum, hi-hat, tom-tom, and slam cymbals contained therein, forming multiple preset audio material sets (i.e., sound effect datasets), combining MIDI files with sound effect datasets, and then using MIDI files to drive and organize playback. In this way, the separated drums are like musical instruments, which determine the timbre. The MIDI file is like a musical score, which determines when to trigger which specific timbre. The user is like the "conductor" or "DJ" here, who decides which drum kit and which score to choose. This allows the smart drummer's playing style, playing density, playing intensity, BPM and other musical parameters and instrument timbre to be conveniently adjusted, effective, and seamlessly switched in real time. Furthermore, the musical parameters and playing content can be freely combined with drum kits of different timbres to enrich the audio and facilitate user operation. (2) By modifying the MIDI playback rate, the BPM can be adjusted to any value without affecting the sound quality. (3) Many different styles, rhythms, note densities and playing intensities of MIDI can be pre-made to enrich the smart drummer's performance and provide more expression and choices. (4) Audio data is stored in the form of MIDI files, which are very small in size and can significantly save on the consumption of additional storage resources. (5) The audio playback system can be interacted with in a simple and friendly way in the game, with flexible control, rich expression, and low system development cost. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.

[0161] The basic concepts have been described above. It is obvious that the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, various modifications, improvements, and corrections may be made to this specification by those skilled in the art. Such modifications, improvements, and corrections are taught in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

Claims

1. An audio playback method, characterized in that, include: Obtain a first input through the user interface, the first input including music parameter information, the music parameter information including performance style information, performance density information, and performance dynamics information; Obtain a second input input through the user interface, the second input including sound effect data information, the sound effect data information including sound effect style information; Based on the performance style information, the performance density information, and the performance dynamics information, the target performance style, target performance density, and target performance dynamics are determined. A target music segment is determined from multiple preset music segments based on the target performance style, the target performance density, and the target performance dynamics. Each preset music segment is described by multiple music parameters. The preset music segments are stored in a digital music interface file format. The target music segment is described by target music parameters. Based on the sound effect style information, a target sound effect dataset is determined from multiple preset sound effect datasets. The target sound effect dataset includes audio samples of different drum parts with the same sound effect style. Each preset sound effect dataset includes audio samples of each drum part in a set of drum parts, and the audio samples of each drum part correspond to the same sound effect style. as well as The target music segment drives the playback of the target sound effect dataset.

2. The audio playback method according to claim 1, characterized in that, The same timbre in the multiple preset music segments is at the same pitch.

3. The audio playback method according to claim 1, characterized in that, The method further includes: Obtain a third input through the user interface, the third input including the target number of beats; Based on the mapping relationship between the number of beats and the playback speed of the music segment, and the target number of beats, the playback speed of the target music segment is determined; and The target sound effect dataset is played based on the playback speed of the target music segment.

4. The audio playback method according to claim 1, characterized in that, The user interface includes a style selector and an X / Y controller, and the first input obtained through the user interface includes: The performance style information from the first input is obtained through the style selector; and The X / Y controller acquires the performance density information and performance intensity information from the first input.

5. The audio playback method according to claim 1, characterized in that, The user interface includes a sound effect selector, and the acquisition of the second input through the user interface includes: The sound effect data information in the second input is obtained through the sound effect selector.

6. An audio playback device, characterized in that, It includes a processor and a memory, the memory storing a computer program or computer-executable instructions, which, when executed by the processor, implement the audio playback method according to any one of claims 1 to 5.

7. A computer program product, characterized in that, It includes a computer program that, when at least a portion of the computer program is executed by a processor, enables the implementation of the method as described in any one of claims 1 to 5.

8. An audio playback system, characterized in that, include: A game engine includes a user interface, which includes a music selection module and a sound effect selection module. The music selection module is used to acquire a first input, which includes music parameter information, including performance style information, performance density information, and performance velocity information. The sound effect selection module is used to acquire a second input, which includes sound effect data information, including sound effect style information. as well as The audio engine includes a music module and a sound effects module, wherein the music module is configured as follows: The target performance style, target performance density, and target performance intensity are determined based on the performance style information, the performance density information, and the performance intensity information. A target music segment is determined from multiple preset music segments based on the target performance style, the target performance density, and the target performance dynamics. The target music segment is described by target music parameters. Each preset music segment is described by multiple music parameters. The preset music segments are stored through a digital music interface file format. The sound effects module is configured as follows: Based on the sound effect style information, a target sound effect dataset is determined from multiple preset sound effect datasets. The target sound effect dataset includes audio samples of different drum parts with the same sound effect style. Each preset sound effect dataset includes audio samples of each drum part in a set of drum parts, and the audio samples of each drum part correspond to the same sound effect style. as well as The target music segment drives the playback of the target sound effect dataset.

9. The audio playback system as described in claim 8, characterized in that, The music selection module includes a style selector and an X / Y controller. The style selector is used to input the performance style information in the music parameter information, and the X / Y controller is used to input the performance density information and performance dynamics information in the music parameter information.

10. The audio playback system as described in claim 8, characterized in that, The user interface also includes a playback control module, which includes a beat count selector for acquiring a third input, the third input including a target beat count. The audio engine is further configured to: Based on the mapping relationship between the number of beats and the playback speed of the music segment, and the target number of beats, the playback speed of the target music segment is determined; and The target sound effect dataset is played based on the playback speed of the target music segment.

11. The audio playback system as described in claim 8, characterized in that, It further includes a storage module for storing the plurality of preset sound effect datasets.

12. The audio playback system as described in claim 8, characterized in that, It further includes a storage module for storing the plurality of preset music segments.

Citation Information

Patent Citations

  • Accompaniment music generation method and device and computer readable storage medium

    CN111933098A

  • Method of and system for automated musical arrangement and musical instrument performance style transformation supported within an automated music performance system

    US20210110802A1