Virtual reality music game system based on audio separation technology and use method

By employing audio separation technology and real-time signal processing, it is possible to automatically extract specific instrument tracks from any song and generate VR interactive prompts. This solves the problem that existing music rhythm games cannot process and automatically generate such prompts in real time, thus enhancing the player's personalized gaming experience and the system's flexibility.

CN121731744APending Publication Date: 2026-03-27FOSHAN CHUANGSHIJIA SCI&TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing music rhythm games cannot achieve real-time processing and automated generation of any song, and lack the function of automatically separating specific instrument tracks from the song and converting them into interactive prompts, thus limiting the player experience to preset content.

Method used

By employing audio separation technology combined with real-time signal processing and VR interactive design, the system automatically extracts specific instrument tracks from any song through a real-time processing engine and generates VR interactive prompts, allowing players to operate virtual instruments in a virtual reality environment to recreate the song accompaniment.

Benefits of technology

It enables personalized game content generation for any song, enhancing player immersion and enjoyment. It supports separate and collaborative performance of multiple instruments, reduces system setup costs, and offers real-time performance and high efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121731744A_ABST
    Figure CN121731744A_ABST
Patent Text Reader

Abstract

The invention provides a virtual reality music game system based on an audio separation technology, which comprises a VR interaction module and a real-time processing engine, and the real-time processing engine comprises an audio input module, an audio processing module, a rhythm prompt generation module, a control management module and a virtual scene generation module. The audio input module receives a song audio file; the audio processing module generates a first audio track file according to an audio track of a target musical instrument in the song audio file and removes the audio track of the target musical instrument to generate a second audio track file; the rhythm prompt generation module analyzes the rhythm of the first audio track file and the time sequence of the sound of the target musical instrument, and generates an interaction prompt information file; the control management module can control the system and receive, store and manage data files, the virtual scene generation module generates a virtual reality environment, and the VR interaction module generates a third audio track file. By automatically extracting the specific musical instrument audio track in the song and converting the specific musical instrument audio track into the VR interaction prompt, a player is supported to import any song and restore the accompaniment through the virtual musical instrument.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of training devices that incorporate virtual reality, and in particular to a virtual reality music game system and its usage method based on audio separation technology. Background Technology

[0002] Currently, music rhythm games rely on manually designed game levels, making it impossible to process any imported song in real-time and automatically generate game music. User experience is limited by preset content. Furthermore, while traditional audio editing tools can separate instrument tracks from songs, they haven't been deeply integrated with VR interaction, failing to meet players' needs for importing any song and instantly generating personalized game content.

[0003] For example, music rhythm games such as Guitar Hxxx and Beat Sxxxx mainly rely on manually designed game levels and cannot achieve real-time processing and automated generation of any song. The player experience is limited by preset content. At the same time, the above games do not belong to the category of song accompaniment content that requires players to use props to supplement instrument tracks. Existing music game systems have the following deficiencies: (1) They cannot perform real-time analysis and game content generation on any song provided by the player, and rely heavily on preset levels (preset levels refer to levels with core content such as layout, rules, objectives, and elements designed in advance by the developer); (2) They lack the function of automatically separating specific instrument tracks from songs and converting them into interactive prompts; (3) Players cannot flexibly choose sound sources (such as original song tracks or general sound source libraries) for personalized supplementation and restoration.

[0004] This invention proposes an innovative solution that combines audio separation technology, real-time signal processing, and VR interactive design to automatically extract specific instrument tracks from any song and convert them into game prompts in real time. This allows players to recreate the song accompaniment by operating virtual instruments in a VR environment, thereby enhancing the fun and engagement of music games. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a virtual reality music game system based on audio separation technology. By automatically extracting specific instrument tracks from songs, converting and matching them into corresponding VR interactive prompts, the system eliminates the need for manual pre-design of levels and allows players to import any song into the game system and automatically generate game content, making the game system personalized and flexible.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] Virtual reality music game systems based on audio separation technology include:

[0008] At least one VR interaction module is used to display virtual reality to players and allow players to interact with the virtual reality environment.

[0009] Real-time processing engine, including:

[0010] The audio input module is used to receive song audio files input by the player.

[0011] The audio processing module is used to extract the track of the target instrument from the song audio file, generate the first track file, and remove the track of the target instrument from the song audio file and generate the second track file of the corresponding accompaniment version of the song.

[0012] The rhythm cue generation module is used to analyze the rhythm and timing of the target instrument sounds in the extracted first audio track file and generate an interactive cue information file suitable for the VR environment.

[0013] The control and management module is used to control the system's operation, receive and store data files, and manage the data files.

[0014] The virtual scene generation module generates virtual reality environments;

[0015] During system operation, the control and management module receives the song audio file, the second audio track file, and the interactive prompt information file, and generates game content data based on the song audio file and virtual reality data. The VR interaction module receives the game content data, interactive prompt information file, and second audio track file output by the real-time processing engine, plays the second audio track file to the player, and displays the virtual reality environment. The displayed virtual reality environment will display rhythm or melody prompts. When the player interacts in the virtual reality environment, the VR interaction module generates a third audio track file. The VR interaction module overlays the third audio track file onto the second audio track file in real time and plays it. The third audio track file can be a first audio track file based on the target instrument audio in the input song, or it can be an audio track file based on the target instrument sound from an external digital audio library.

[0016] Compared with existing technologies, the virtual reality music game system based on audio separation technology of the present invention has the following beneficial effects:

[0017] (1) In the prior art (such as patent CN202010927502.4), drum beats or rhythms or melodies of various other instruments can be extracted from the audio of a song, and the original sounds of various instruments in a song can be extracted. Then, prompts are matched for the extracted sounds and presented in the VR scene according to the time sequence, so that players can operate the instruments according to the prompts and restore the original instrument accompaniment of the song. This invention is a specific implementation method of existing technology. Based on artificial intelligence technology, this invention achieves "automatic" and "real-time" processing of imported audio data. This solution can automatically extract specific instruments in the audio of a song, match corresponding prompts, and provide players with VR game interaction. It realizes real-time processing of audio data and prompt output, making it convenient for players to directly import any song and automatically generate game content. Specifically, in the VR environment, players can select any song. This system can automatically extract the original sound of various instruments in a song and use an audio separation tool to remove the rhythm points (such as drum beats), vocals, or melodies of specific instruments in the original song to form an accompaniment version. Then, when players play or strum virtual instruments in VR, the performance audio is superimposed on the accompaniment audio in real time, thereby restoring the rhythm points and melodies of the separated instruments in the accompaniment audio and achieving the purpose of restoring the original song.

[0018] (2) This system can extract the target instrument track from any input song file and generate interactive prompts as VR interactive content, which eliminates the need for manually selecting specific instrument tracks and designing interactive prompts in the song file, reducing the system's construction cost. Moreover, the above steps can be completed immediately after the player inputs the song file, which is real-time and efficient.

[0019] (3) This system is intelligent and personalized. No manual level design is required. Players can import any song and the system will automatically generate game content.

[0020] (4) This system can bring players a good sense of immersion. Through VR interaction and real-time audio track overlay, it provides a highly realistic experience of various musical instrument accompaniment.

[0021] (5) This system is scalable, and the algorithm supports the separation of multiple instruments, and is suitable for various musical elements such as drum beats, melodies, and chords;

[0022] (6) This invention combines audio separation technology with VR interaction, supports players to import any song and generate game content in real time, and provides sound source selection (original sound source or external sound source library) and accompaniment restoration function.

[0023] (7) The separated first audio track file is the target instrument sound extracted from the input song audio file; when the player interacts and plays in virtual reality, if the player chooses to overlay the first audio track file onto the second audio track, that is, the playing sound emitted by the player during interaction is the original sound of the target instrument of the input song, at this time the third audio track file is equivalent to the first audio track file; if the player chooses the target instrument sound from the external digital audio library for interactive playing, it is equivalent to generating a new audio track file that is different from the first audio track file and overlaying it onto the second audio track, that is, the playing sound emitted by the player during interaction is the sound of the target instrument of the input song with different timbre and loudness. For example, if the target instrument sound is the flute sound in the input song, then the target instrument sound in the third audio track file can be the flute sound with different timbre, loudness and pitch.

[0024] Preferably, the VR interaction module is equipped with a sound source selection module, which allows players to select the instrument sound source to be used when generating the third audio track file.

[0025] Preferably, the VR interaction module has at least two components, allowing at least two players to play the game simultaneously. This enables at least two players to collaboratively operate at least two virtual musical instruments in the virtual reality environment to input at least two third audio track files. The real-time processing engine receives at least two third audio track files and overlays them with a second audio track file. The real-time processing engine sends the interactive prompt information file corresponding to the virtual musical instrument selected by the player to the corresponding VR interaction module.

[0026] The setup of multiple VR interactive modules allows different people to play the same song simultaneously and overlay multiple third-track files input by multiple players onto the accompaniment (second-track file), thereby enabling multi-instrument collaborative performance, expanding the applicable scenarios of this system, and adding collaborative elements to the game, which can promote teamwork among player groups.

[0027] Preferably, the real-time processing engine is connected to an artificial intelligence system, and the audio processing module uses the artificial intelligence system to extract the audio tracks from the song audio file in real time.

[0028] When combined with an artificial intelligence system, the real-time processing engine can use lightweight deep learning algorithms to perform audio track separation (such as audio source separation technology based on the LiteDemucs model neural network, where the lightweight LiteDemucs model has a hidden layer dimension of 32 and a convolution kernel size of 8), extracting various specific instrument tracks (such as drum beats, melodies, etc.) from the song. At the same time, it combines Fast Fourier Transform (FFT) and beat detection algorithms to extract rhythm points in the audio tracks.

[0029] Preferably, the real-time processing engine employs a real-time processing mode and a pre-management mode. In real-time processing mode, after the player inputs a song audio file, the player can immediately operate the VR interaction module to play the game. In pre-management mode, after the player inputs a song audio file, the real-time processing engine completes the analysis and processing of the song audio file and outputs game content data, interactive prompt information files, and a second audio track file. The game content data, interactive prompt information files, and the second audio track file are stored in the database. When the player subsequently operates the VR interaction module to play the game, the game content data, interactive prompt information files, and the second audio track file are called.

[0030] This invention provides two data management modes: a real-time processing mode and a pre-management mode. The real-time processing mode utilizes high-performance computing and deep learning models (such as convolutional neural networks or Transformer architectures) to perform instant separation and rhythm analysis on the input song audio file, enabling "instant play." This mode allows players to quickly start the game, but the system has higher operational complexity and requires algorithm optimization to reduce latency. The pre-management mode, on the other hand, performs offline analysis and processing of the song audio, generating relevant data which is then stored in a database and retrieved when the game loads. This mode is suitable for scenarios with limited computing resources but cannot achieve "instant play." By providing two data management modes, this invention allows the system to select the appropriate mode based on the available computing resources for different scenarios, thereby reducing the deployment limitations of the system.

[0031] Preferably, the real-time processing engine enters real-time processing mode. After the player inputs the song audio file, the following steps are performed:

[0032] Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds.

[0033] Several audio segments are processed using a data model to generate a first audio track file and a second audio track file; the first audio track file is analyzed using the data model to extract rhythm points from the first audio track file and generate a rhythm point data file.

[0034] Convert the rhythm point data file into an interactive prompt information file that can be mapped by VR devices in the VR environment;

[0035] And / or, the real-time processing engine enters pre-management mode, and after the player inputs the song audio file, the following steps are performed:

[0036] Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds.

[0037] Several audio segments are processed using an offline data model to generate a first audio track file and a second audio track file. The second audio track file is then stored.

[0038] The first audio track file is analyzed using an offline data model to extract rhythm points from the first audio track file, generate a rhythm point data file, and store the rhythm point data file.

[0039] After the player operates the VR interaction module to load the game and selects a song and sound source, the real-time processing engine calls the second audio track file and rhythm point data file. The real-time processing engine converts the rhythm point data file into an interactive prompt information file that can be mapped by the VR device in the VR environment.

[0040] Preferably, in the process of processing several audio segments using a data model, hierarchical parallel computing is employed, including:

[0041] Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue;

[0042] Execute the separation thread: Run the INT8 quantization model to extract the track of the selected instrument from the audio segment, obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and integrate the several audio segments containing the selected instrument to generate the first track file;

[0043] The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file.

[0044] Use a data model to extract the rhythm points from the entire first audio track file and generate a rhythm point data file.

[0045] Alternatively, when processing several audio segments using a data model, hierarchical parallel computing can be employed, including:

[0046] Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue;

[0047] Execute the separation thread, run the INT8 quantization model, extract the track of the selected instrument from the audio segment, and obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and the several audio segments containing the selected instrument are used as a subset of the first track file.

[0048] The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file.

[0049] The rhythm analysis thread uses Fast Fourier Transform and dynamic beat detection algorithms to extract the rhythm points of each track segment within a subset of the first audio track file, and integrates the rhythm points of several track segments to generate a rhythm point data file.

[0050] In the process of processing several audio segments using a data model, the hierarchical parallel computing architecture can improve the processing efficiency of audio files, enabling rapid processing and generation of data after the song audio is input into the real-time processing engine, thereby enhancing the player's gaming experience in real-time processing mode.

[0051] Preferably, during the execution of the separate thread, redundant parameters are removed by model pruning and 8-bit integer quantization are used to reduce the processing latency of each audio segment to no more than 25 milliseconds.

[0052] By optimizing latency through separate threads, the impact of large data processing volumes on game screen latency can be reduced.

[0053] Preferably, during the execution of the rhythm analysis thread, predictive caching is performed based on the periodic detection of music beats to pre-generate rhythm cues for the next 5-20 seconds.

[0054] By implementing predictive caching in the rhythm analysis thread, real-time computation can be reduced by 40%, thereby improving data processing efficiency.

[0055] Another object of the present invention is to provide a method of using the above-mentioned virtual reality music game system based on audio separation technology, comprising:

[0056] Players can upload song audio files and select the target instrument to play through the system interface;

[0057] The real-time processing engine receives the song audio file, analyzes and processes it, and generates game content data, interactive prompt files, and a second audio track file. The interactive prompt file contains a sequence of operation prompts generated by the real-time processing engine based on the rhythm and timing of the target instrument's audio track in the song audio file.

[0058] Players interact with the VR interaction module, and the real-time processing engine sends interactive feedback to the VR interaction module.

[0059] Players select their audio source configuration;

[0060] Players play the selected virtual instrument in a virtual reality environment according to the interactive prompts to generate a third audio track file. The VR interaction module overlays the third audio track file onto the second audio track file in real time and plays it, allowing players to reproduce the accompaniment of the input song in the virtual reality environment.

[0061] Compared with existing technologies, the present invention provides a method for using a virtual reality music game system based on audio separation technology. This overcomes the problem that existing VR music games only provide songs, game prompts, and game levels pre-set by game developers, resulting in relatively fixed VR music game content. The invention provides a highly flexible and personalized method for using a virtual reality music game system based on audio separation technology, allowing players to flexibly choose the accompaniment of the songs they want to appear in the VR scene according to their own interests or needs, and use virtual instruments to restore the accompaniment, thereby improving the player's entertainment experience. Attached Figure Description

[0062] Figure 1 A schematic diagram illustrating one possible setup method for the overall architecture of a virtual reality music game system based on audio separation technology;

[0063] Figure 2 This is the virtual reality scene interface for this game system;

[0064] Figure 3 This is a schematic diagram illustrating another possible configuration of the overall architecture of a virtual reality music game system based on audio separation technology. Detailed Implementation

[0065] The embodiments of the present invention are described below with reference to the accompanying drawings:

[0066] Example 1

[0067] See Figure 1 This embodiment of a virtual reality music game system based on audio separation technology includes: at least one VR interaction module and a real-time processing engine; the VR interaction module is used to display virtual reality to players and allow players to interact in the virtual reality environment; the real-time processing engine includes: an audio input module, an audio processing module, a rhythm prompt generation module, and a control management module; the audio input module is used to receive song audio files input by players; the audio processing module is used to extract the track of the target instrument from the song audio file, generate a first track file, and remove the track of the target instrument from the song audio file, and generate a second track file for the accompaniment version of the corresponding song; the rhythm prompt generation module is used to analyze the rhythm of the first track file and the timing of the target instrument sound based on the extracted first track file, and generate an interactive prompt information file suitable for the VR environment; the control management module is used to control the system operation, receive and store data files, and manage the data files; the virtual scene generation module generates a virtual reality environment (generating a virtual reality environment is an application of prior art).

[0068] When this system is running, the control and management module receives the song audio file, the second audio track file, and the interactive prompt information file, and generates game content data based on the song audio file and virtual reality data. The VR interaction module receives the game content data, interactive prompt information file, and second audio track file output by the real-time processing engine, plays the second audio track file to the player, and displays the virtual reality environment. The displayed virtual reality environment will display rhythm or melody prompts. When the player interacts in the virtual reality environment, the VR interaction module generates a third audio track file. The VR interaction module overlays the third audio track file onto the second audio track file in real time and plays it (the real-time processing engine also receives the third audio track file). The third audio track file can be a first audio track file based on the target instrument audio in the input song, or it can be an audio track file based on the target instrument sound from an external digital audio library.

[0069] The real-time processing engine features a song scoring mode. During gameplay, it records the accuracy of the player's operation of the virtual instrument and provides corresponding operation prompts and feedback (such as real-time operation scores, sound effect adjustments, etc.). Then, it scores the player's operation and provides feedback on the player's performance in this game, so that the player can challenge themselves or choose the difficulty of the song later.

[0070] Introduction to VR Interaction Module:

[0071] The VR interaction module includes a control processing module, a display module, an operation sensing module, and a scene physics simulation module. However, the modules within the VR interaction module can be added according to the actual needs of the game, and are not limited to the above-mentioned functional modules. The above-mentioned functional modules are applications of existing technology and are not the main inventive point of this invention, therefore they will not be described in detail. In this embodiment, the system supports Quest2, Index, and Vive. If PSVR or Pico needs to be adapted, OpenXRFeature can be extended, meaning that the system is compatible with multiple types of VR devices.

[0072] During gameplay, the VR interaction module receives player input via virtual musical instruments (such as drumsticks, piano keys, guitar strings, etc.) and sends it to the real-time processing engine. The real-time processing engine then feeds back corresponding data to alter the VR interaction module's presentation of the virtual reality.

[0073] The process involves analyzing the first audio track file and then generating an interactive prompt information file. The main technical principle is to generate interactive prompts in a dynamic three-dimensional space in the VR environment based on the timing and characteristics of the audio track (such as rhythm points, pitch, and dynamics). The interactive prompt information displayed by the VR interaction module can be the intensity of light spot brightness, the size of the light spot, the position of the light spot, the change (trajectory) of the light spot position, or the striking position and hitting position prompts of virtual musical instruments, but is not limited to the above prompt forms.

[0074] The VR interaction module also includes a sound source selection module, allowing players to choose the instrument sound sources used when generating the third audio track file. Instrument sound sources include those from the player-input song audio file and those from external digital audio libraries. Players can select the instrument sound source before the game starts, or select it before the game starts and switch it during gameplay.

[0075] In this solution, the sources of instrument sound sources are diversified. They can be the sound sources of various instruments in the original song, or sound sources from other sound source libraries. For example, when players hit virtual instruments, they can use the sound source of the original song (such as the drum sound extracted from the original song), or they can choose to use the sound source of other sound source libraries (such as drum sounds of different timbres and types). Supporting multiple sound source selections makes this system highly flexible. The sound source selection function can meet the creative needs of different players, thereby giving players a personalized and highly immersive music game experience. In addition, players can switch instrument sound sources during the game, thereby improving the flexibility and fun of this system.

[0076] The VR interaction module also includes a haptic feedback module that provides haptic feedback to players. The haptic feedback module can generate corresponding vibrations according to different virtual musical instruments and adjust the vibration intensity and duration according to the intensity of the player's operation.

[0077] The haptic feedback module is designed to provide players with haptic feedback, outputting more realistic interactive feedback, thereby enhancing the player's sense of immersion.

[0078] See Figure 2 The interactive prompts in this invention include:

[0079] (1) Dynamic Spatialization Cue (3DRhythmCue)

[0080] Technical principle: The rhythm points of the instrument track are converted into three-dimensional spatial cues in the VR environment, and the position and movement trajectory of the cues are dynamically associated with the characteristics of the track (such as pitch and intensity).

[0081] Implementation steps: (1) Track feature extraction: Analyze the rhythm points, pitch and intensity of the target track. For example, the strength of the drum beat corresponds to the size of the cue, and the bass drum and treble cymbal correspond to different spatial heights; (2) Cue generation: Generate dynamic light spots or trajectories in the VR space, such as the representation of the bass drum: red light spot, rising from the ground, and the player needs to hit it downwards with a virtual drumstick; and the representation of the treble cymbal: blue light spot, falling from above, and the player needs to hit it upwards; and intensity changes: the size of the light spot changes with the volume (for example, a strong drum beat is a large light spot).

[0082] (2) Spatial mapping hint:

[0083] Technical principle: The prompts are distributed in a 360° circle around the player, with increased density of light spots in fast-paced segments and sparse distribution in slow-paced segments.

[0084] Traditional VR rhythm games mostly use a one-way trajectory (such as blocks flying in front) for prompts, while this solution uses the characteristics of the sound track to generate multi-dimensional and dynamic prompts, thereby enhancing the player's immersion (the dynamic prompts in this solution generate three-dimensional light spots based on pitch (height) and intensity (magnitude), which are randomly distributed around the player).

[0085] Audio processing module introduction:

[0086] The audio processing module includes an audio separation module and an accompaniment generation module. The audio separation module is used to extract the track of the target instrument from the song audio file and generate a first track file. The accompaniment generation module is used to remove the track of the target instrument from the song audio file and generate a second track file for the corresponding accompaniment version of the song.

[0087] The real-time processing engine is connected to an artificial intelligence system, and the audio processing module uses the artificial intelligence system to extract the audio tracks from the song's audio file in real time.

[0088] When combined with an artificial intelligence system, the real-time processing engine can use lightweight deep learning algorithms to perform audio track separation (such as audio source separation technology based on the LiteDemucs model neural network, where the lightweight LiteDemucs model has a hidden layer dimension of 32 and a convolution kernel size of 8), extracting various specific instrument tracks (such as drum beats, melodies, etc.) from the song. At the same time, it combines Fast Fourier Transform (FFT) and beat detection algorithms to extract rhythm points in the audio tracks.

[0089] The real-time processing engine employs both a real-time processing mode and a pre-management mode. In real-time processing mode, after the player inputs a song audio file, the player can immediately operate the VR interaction module to play the game. In pre-management mode, after the player inputs a song audio file, the real-time processing engine analyzes and processes the song audio file and outputs game content data, interactive prompt information files, and a second audio track file. The game content data, interactive prompt information files, and the second audio track file are stored in the database. When the player subsequently operates the VR interaction module to play the game, they can call upon the game content data, interactive prompt information files, and the second audio track file ("audio stream", "audio track stream", "rhythm data", and "virtual reality scene data").

[0090] This invention provides two data management modes: a real-time processing mode and a pre-management mode. The real-time processing mode utilizes high-performance computing and deep learning models (such as convolutional neural networks or Transformer architectures) to perform instant separation and rhythm analysis on the input song audio file, enabling "instant play." This mode allows players to quickly start the game, but the system has higher operational complexity and requires algorithm optimization to reduce latency. The pre-management mode, on the other hand, performs offline analysis and processing of the song audio, generating relevant data which is then stored in a database and retrieved when the game loads. This mode is suitable for scenarios with limited computing resources but cannot achieve "instant play." By providing two data management modes, this invention allows the system to select the appropriate mode based on the available computing resources for different scenarios, thereby reducing the deployment limitations of the system.

[0091] In this embodiment, the real-time processing engine is in real-time processing mode, using high-performance computing and deep learning models (such as convolutional neural networks or Transformer architecture) to perform real-time separation and rhythm analysis of the input audio.

[0092] In this embodiment, the real-time processing engine enters real-time processing mode. After the player inputs the song audio file, the following steps are performed:

[0093] Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds.

[0094] Several audio segments are processed using a data model to generate a first audio track file and a second audio track file; the first audio track file is analyzed using the data model to extract rhythm points from the first audio track file and generate a rhythm point data file.

[0095] Convert the rhythm point data file into an interactive cue information file that can be mapped by VR devices in the VR environment.

[0096] In this embodiment, the real-time processing engine enters pre-management mode. After the player inputs the song audio file, the following steps are performed:

[0097] Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds.

[0098] Several audio segments are processed using an offline data model to generate a first audio track file and a second audio track file. The second audio track file is then stored.

[0099] The first audio track file is analyzed using an offline data model to extract rhythm points from the first audio track file, generate a rhythm point data file, and store the rhythm point data file (stored in JSON format, which includes timestamps and pitches).

[0100] After the player operates the VR interaction module to load the game and selects a song and sound source, the real-time processing engine calls the second audio track file and rhythm point data file. The real-time processing engine converts the rhythm point data file into an interactive prompt information file that can be mapped by the VR device in the VR environment.

[0101] The selected audio source can be the original sound of the target instrument in the original song (original version), or it can be the target instrument sound from an external digital audio library, such as various virtual instrument sounds in Cubase and Battery.

[0102] This invention provides the following two processes for processing several audio segments using a data model:

[0103] The first method: Using a data model to process several audio segments, a hierarchical parallel computing approach is employed, including:

[0104] Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue;

[0105] Execute the separation thread: Run the INT8 quantization model (LiteDemucs model), extract the track of the selected instrument from the audio segment, obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and integrate several audio segments containing the selected instrument to generate the first track file;

[0106] The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file.

[0107] The data model is used to extract the rhythm points from the entire first audio track file, generating a rhythm point data file.

[0108] The second approach involves using a data model to process several audio segments, employing hierarchical parallel computing, including:

[0109] Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue;

[0110] Execute the separation thread, run the INT8 quantization model, extract the track of the selected instrument from the audio segment, and obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and the several audio segments containing the selected instrument are used as a subset of the first track file.

[0111] The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file.

[0112] The rhythm analysis thread uses Fast Fourier Transform (FFT) and dynamic beat detection algorithms (such as the improved LibrosaOnsetDetection) to extract the rhythm points of each track segment within a subset of the first track file, and integrates the rhythm points of several track segments to generate a rhythm point data file.

[0113] The difference between the first and second processes lies in the rhythm analysis process. The first process analyzes the entire first audio track file directly, while the second process analyzes a subset (segmented audio) of the first audio track file.

[0114] The rhythm analysis thread and the separation thread run synchronously, thereby shortening the rhythm analysis efficiency; the data synchronization of each thread adopts a timestamp alignment mechanism to ensure that the separation tracks, rhythm points and accompaniment are seamlessly connected on the timeline.

[0115] In the process of processing several audio segments using a data model, the audio data is processed based on a hierarchical parallel computing architecture (decomposing the audio processing flow into multiple sub-tasks and reducing serial processing through parallel computing), which can improve the processing efficiency of audio files and enable the rapid processing and generation of data after the song audio is input to the real-time processing engine, thereby improving the player's gaming experience in real-time processing mode.

[0116] Real-time instrument track separation with latency below 25 milliseconds is achieved through hierarchical parallel computing and a lightweight LiteDemucs model (32 hidden layer dimensions, quantized to INT8).

[0117] Preferably, during the execution of the separate thread, redundant parameters are removed by model pruning and 8-bit integer quantization is performed. Combined with GPU parallel computing, the processing latency of each audio segment is reduced to no more than 25 milliseconds.

[0118] By optimizing latency through separate threads, the impact of large data processing volumes on game screen latency can be reduced.

[0119] In this embodiment, 30% of redundant parameters can be removed through model pruning, thereby achieving latency optimization.

[0120] Preferably, during the execution of the rhythm analysis thread, predictive caching is performed based on music beat periodicity detection (BPM detection) to pre-generate rhythm cues for the next 5-20 seconds.

[0121] By implementing predictive caching in the rhythm analysis thread, real-time computation can be reduced by 40%, thereby improving data processing efficiency. The BPM detection method is an application of existing technology.

[0122] In this embodiment, the audio separation algorithm uses the Spleeter or Demucs model based on deep learning, which supports multi-channel separation.

[0123] In this embodiment, timing analysis involves extracting rhythm points through spectrum analysis and beat detection algorithms (such as the Librosa library).

[0124] In this embodiment, the VR is presented as being developed using Unity or Unreal Engine and is compatible with mainstream VR devices (such as Oculus Quest).

[0125] In this embodiment, the relevant explanations of model compression and hardware acceleration technologies are as follows:

[0126] Technical principle: Reduce computation by using model pruning and quantization, and accelerate real-time performance by using dedicated hardware.

[0127] Separation accuracy optimization algorithm: adopts multi-channel separation annotation and extracts spatial information from stereo to improve separation accuracy (by utilizing the difference between the left and right channels of stereo, the spatial characteristics of the specified instrument track are enhanced, making the real-time processing engine more sensitive to the rhythm point position of the instrument in the audio and improving separation accuracy).

[0128] Implementation details: (a) Pruning the pre-trained audio separation model (e.g., Demucs) to remove redundant parameters and retain feature extraction layers sensitive to target audio tracks such as drum beats and melodies; (b) Quantizing the model weights from 32-bit floating-point to 8-bit integers (INT8) to reduce memory usage and computational overhead; (c) Accelerating inference on VR devices (e.g., Oculus Quest) using built-in DSPs (Digital Signal Processors) or external TPUs (e.g., Google Coral). Inference time was reduced from approximately 150 milliseconds / segment to 30 milliseconds / segment, and memory usage was reduced by approximately 60% (optimization effect).

[0129] Regarding the song separation and accompaniment generation in the real-time processing engine, this invention provides two solutions. One solution is as follows: After the player inputs the song audio, they select the instrument they want to play through the VR interaction module, thereby determining the instrument audio that needs to be extracted from the original song and the instrument audio that needs to be removed from the original song. The real-time processing engine then processes the audio file according to the player's needs. The other solution is as follows: After the player inputs the song audio, the real-time processing engine automatically detects the audio tracks of all instruments in the song audio. Assuming there are n types of instruments in the song audio, the engine extracts the audio tracks of each of the n instruments from the original song to form n first audio track files. It also removes the audio tracks of the corresponding instruments from the original song to form n second audio track files (for example, the xth first audio track file and the xth third audio track file can be superimposed to restore the original song). When the player subsequently selects the instrument they want to play through the VR interaction module, the real-time processing engine calls the prompt information file generated by the first audio track file corresponding to that instrument and the third audio track file corresponding to that instrument.

[0130] Explanation of predictive caching mechanisms:

[0131] Technical principle: By utilizing the periodicity of music, subsequent rhythm points can be predicted in advance, reducing the burden of real-time calculation.

[0132] Implementation steps: (1) When analyzing the first few seconds of audio, identify the beat period (e.g., 120 beats per minute, BPM=120); (2) Based on the periodicity assumption, pre-generate the rhythm cue sequence for the next few seconds (e.g., the drum position for the next 5-20 seconds); (3) If a beat change is detected (e.g., speed change or complex rhythm), dynamically adjust the prediction model to ensure accuracy.

[0133] Optimization results: Reduces the amount of dynamic computation for real-time processing by about 40%, and further reduces latency to less than 20 milliseconds.

[0134] The following examples illustrate the real-time processing mode and the pre-management mode:

[0135] Example 1: The time processing engine is in real-time processing mode.

[0136] Players import the song "Bad." The real-time processing engine extracts the drum beat track (first audio file) from the song's audio using the audio separation module and generates a drum-free accompaniment audio (second audio file). The rhythm cue generation module analyzes the timing of the drum beats in the drum beat track. The real-time processing engine outputs game content data, interactive cues, and the second audio file. The VR interaction module receives the game content data, interactive cues, and the second audio file, and presents the hit cues as dynamic light dots in the VR environment. Players select the original drum beat source or a source from the sound library and use virtual drumsticks to hit the cued positions. The system then overlays the drum beat audio onto the accompaniment audio in real time (mixing), completing the song's restoration (restoring the drum accompaniment within the song's accompaniment).

[0137] The workflow steps of a real-time processing engine:

[0138] Step 1: Audio Input: The player uploads "Bad" (MP3 format, 44.1kHz sampling rate), and the system segments it into 0.5-second segments (22050 samples).

[0139] Step 2: Instrument Separation: Using the LiteDemucs model (trained on the MUSDB18 dataset), processed via parallel threads:

[0140] Execute the preprocessing thread: segment the audio and send it to the queue.

[0141] Execute the separation thread: Run the INT8 quantization model on the GPU to extract the drum track, with a separation accuracy SNR>15dB, and compress the model parameters to 1 / 4 of the original Demucus.

[0142] The accompaniment generation thread removes the drum beats from the original track through reverse mixing, generating a drum-free accompaniment audio.

[0143] Step 3: Rhythm Analysis: Using the onset_strength and beat_track functions of the Librosa library, drum beat rhythm points are detected with an error of <10ms, and cue sequences for the next 5-20 seconds are predicted and generated in advance (based on the average beat interval).

[0144] Step 4: VR cue generation: Map the rhythm points to red light spots in VR space (height 0m, size determined by intensity, i.e., position based on pitch (height) and intensity (magnitude)), and need to surround the player in 360°.

[0145] Step 5: Interaction and Recreation: Players use virtual drumsticks to hit the light spots, and the system overlays the original drum sound source onto the accompaniment, triggering haptic feedback with an amplitude of 1.0 and a duration of 0.1 seconds.

[0146] The algorithm related to instrument separation includes:

[0147] Input: Audio clip (mix, 2×22050)

[0148] Encoding: conv1d(mix, hidden=32, kernel=8, stride=4) → ReLU

[0149] Decoding: conv_transpose1d(enc, hidden=32→2, kernel=8, stride=4)

[0150] Output: Drum track (drums, 2×22050)

[0151] (1) LiteDemucs model:

[0152] structure:

[0153] Encoder: 2-layer Conv1d (2-channel input, 64-channel output, 4-step stride).

[0154] Decoder: 2-layer ConvTranspose1d (outputs 2 channels).

[0155] Number of parameters: approximately 500,000 (1 / 10 of the original Demucs).

[0156] train:

[0157] Dataset: MUSDB18 (150 songs, preprocessed into 0.5-second segments).

[0158] Loss function: Mean Squared Error (MSE).

[0159] Optimizer: Adam (learning rate 0.001).

[0160] Iteration: 10 rounds, batch size 4.

[0161] Optimize pruning:

[0162] Model pruning removes 30% of redundant parameters, and weights are quantized to INT8, reducing inference time from 150 milliseconds / segment to 30 milliseconds / segment; quantization processing converts weights from FP32 to INT8, improving inference speed by 2 times.

[0163] Hierarchical parallel computing:

[0164] The Python threading module is used to implement three parallel threads, and queue synchronization ensures timestamp alignment.

[0165] (2) Rhythm Analysis and Prediction:

[0166] onset_strength calculates time-frequency energy, beat_track detects BPM, and the prediction formula is: next beat point = last beat point + (60 / BPM); predicts the next 5 seconds, and error correction is achieved through dynamic adjustment.

[0167] (3) VR prompt generation:

[0168] Mapping rules:

[0169] Position: x = r * cos(θ), z = r * sin(θ), y = pitch * 2 (r = 2m, θ is random).

[0170] Size: scale = intensity * 1.5.

[0171] synchronous:

[0172] Aligned with song timestamps, with an error of <10ms.

[0173] Experimental results data:

[0174] Input: A 3-minute clip from "Bad".

[0175] Processing latency: Average end-to-end latency is 25 milliseconds (approximately 300 milliseconds before optimization).

[0176] Separation accuracy: The signal-to-noise ratio (SNR) of the drum track extraction is 15.2dB, which is better than Spleeter2stems' 13.8dB.

[0177] Frame rate: Stable 90 FPS on Oculus Quest 2.

[0178] Synchronization error warning: The positional error between the rhythm point and the light point is less than 10 milliseconds.

[0179] Player feedback: Of the 10 testers, 9 felt the experience was "highly immersive".

[0180] Example 2: The real-time processing engine is in pre-managed mode.

[0181] The song "Billie Jean" is imported, and the real-time processing engine preprocesses the song, separating the melody track and generating an accompaniment track and prompts. This data is then stored. Subsequently, the player uses the VR interaction module to enter the game. The game loads, the player selects the song and chooses either the original drum beat source or a source from a sound library. The player then uses virtual drumsticks to hit the indicated positions, and the system overlays the drum beat audio onto the accompaniment audio in real time, completing the song's recreation.

[0182] The workflow steps of a real-time processing engine:

[0183] Step 1: Audio preprocessing: Perform offline analysis on "Billie Jean", divide it into 0.5-second segments, and generate melody and accompaniment tracks.

[0184] Step 2: Data generation: Analyze the melody track, extract rhythm points, and store them in JSON format (timestamp + pitch).

[0185] Step 3: Game loading: Players select a song and sound source (original melody or piano sample), and the system loads the pre-generated data.

[0186] The algorithm related to instrument separation includes:

[0187] Audio separation algorithm:

[0188] (1) LiteDemucs training:

[0189] Dataset: MUSDB18 (150 songs), segmented into 0.5-second segments, trained for 10 epochs, batch size 4, Adam optimizer (learning rate 0.001).

[0190] Data augmentation: random volume adjustment (0.8-1.2x), time stretching (0.9-1.1x), noise injection (Gaussian noise, standard deviation 0.005).

[0191] Real-time optimization: Layered parallelism reduces computation by 40%, and predictive caching covers timing points 5-20 seconds in the future.

[0192] (2) Time series analysis:

[0193] Algorithm: Librosaonset_strength calculates the strength envelope, beat_track detects the beat, and predictive caching is based on BPM periodicity.

[0194] Parameters: Window size 1024, jump 256, BPM range 60-180.

[0195] (3) VR presentation:

[0196] Development tools: Unity 2021.3, OpenXR plugin supports multiple devices (Oculus Quest 2, Valve Index, HTC Vive).

[0197] Hint design: The position of the light spot is determined by the pitch (0-2 meters high) and intensity (0.5-1.0 times scaling), and it is distributed 360° around the player.

[0198] Haptic feedback: Drum beats (amplitude 1.0, 0.1 seconds), melody (amplitude 0.5, 0.05 seconds).

[0199] Experimental results data:

[0200] Processing time: Offline analysis and preprocessing of a 3-minute song took approximately 20 seconds (i7-9700K CPU).

[0201] Storage requirements: Approximately 500KB of data is required.

[0202] Separation accuracy: Melody track SNR 16.5dB.

[0203] Reproduction accuracy: Rhythm point matching rate 95%.

[0204] Loading delay: The database read prompt data took 0.2 seconds.

[0205] Compared with existing technologies, the virtual reality music game system based on audio separation technology of the present invention has the following beneficial effects:

[0206] (1) In the prior art (such as patent CN202010927502.4), drum beats or rhythms or melodies of various other instruments can be extracted from the audio of a song, and the original sounds of various instruments in a song can be extracted. Then, prompts are matched for the extracted sounds and presented in the VR scene according to the time sequence, so that players can operate the instruments according to the prompts and restore the original instrument accompaniment of the song. This invention is a specific implementation method of existing technology. Based on artificial intelligence technology, this invention achieves "automatic" and "real-time" processing of imported audio data. This solution can automatically extract specific instruments in the audio of a song, match corresponding prompts, and provide players with VR game interaction. It realizes real-time processing of audio data and prompt output, making it convenient for players to directly import any song and automatically generate game content. Specifically, in the VR environment, players can select any song. This system can automatically extract the original sound of various instruments in a song and use an audio separation tool to remove the rhythm points (such as drum beats), vocals, or melodies of specific instruments in the original song to form an accompaniment version. Then, when players play or strum virtual instruments in VR, the performance audio is superimposed on the accompaniment audio in real time, thereby restoring the rhythm points and melodies of the separated instruments in the accompaniment audio and achieving the purpose of restoring the original song.

[0207] (2) This system can extract the target instrument track from any input song file and generate interactive prompts as VR interactive content, which eliminates the need for manually selecting specific instrument tracks and designing interactive prompts in the song file, reducing the system's construction cost. Moreover, the above steps can be completed immediately after the player inputs the song file, which is real-time and efficient.

[0208] (3) This system is intelligent and personalized. No manual level design is required. Players can import any song and the system will automatically generate game content.

[0209] (4) This system can bring players a good sense of immersion. Through VR interaction and real-time audio track overlay, it provides a highly realistic experience of various musical instrument accompaniment.

[0210] (5) This system is scalable, and the algorithm supports the separation of multiple instruments, and is suitable for various musical elements such as drum beats, melodies, and chords;

[0211] (6) This invention combines audio separation technology with VR interaction, supports players to import any song and generate game content in real time, and provides sound source selection (original sound source or external sound source library) and accompaniment restoration function.

[0212] (7) The separated first audio track file is the target instrument sound extracted from the input song audio file; when the player interacts and plays in virtual reality, if the player chooses to overlay the first audio track file onto the second audio track, that is, the playing sound emitted by the player during interaction is the original sound of the target instrument of the input song, at this time the third audio track file is equivalent to the first audio track file; if the player chooses the target instrument sound from the external digital audio library for interactive playing, it is equivalent to generating a new audio track file that is different from the first audio track file and overlaying it onto the second audio track, that is, the playing sound emitted by the player during interaction is the sound of the target instrument of the input song with different timbre and loudness. For example, if the target instrument sound is the flute sound in the input song, then the target instrument sound in the third audio track file can be the flute sound with different timbre, loudness and pitch.

[0213] (8) This invention uses a lightweight LiteDemucs model and hierarchical parallel computing to extract specific instrument tracks from songs in real time and convert them into dynamic 3D VR interactive prompts; the system supports players to import any song and restore the accompaniment through virtual instruments, providing two modes: real-time processing (latency 25 milliseconds, SNR 15.2dB) and pre-management, combined with cross-device adaptation and haptic feedback, to achieve a personalized and highly immersive music game experience;

[0214] (9) This invention optimizes the real-time performance of processing by using layered parallelism, model compression, and hardware acceleration. It also reduces the real-time computing burden by using predictive caching (using the periodicity of music for prediction) and reduces the latency to an acceptable range, which is an improvement over existing rhythm analysis algorithms.

[0215] Example 2

[0216] See Figure 3 This embodiment is an improvement on embodiment one. In this embodiment, there are at least two VR interaction modules, allowing at least two players to play the game simultaneously. At least two players can collaboratively operate at least two virtual musical instruments in the virtual reality environment to input at least two third audio track files. The real-time processing engine receives at least two third audio track files and superimposes the at least two third audio track files onto a second audio track file. The real-time processing engine sends the interactive prompt information file corresponding to the virtual musical instrument selected by the player to the corresponding VR interaction module.

[0217] The setup of multiple VR interactive modules allows different people to play the same song simultaneously and overlay multiple third-track files input by multiple players onto the accompaniment (second-track file), thereby enabling multi-instrument collaborative performance, expanding the applicable scenarios of this system, and adding collaborative elements to the game, which can promote teamwork among player groups.

[0218] In this embodiment, the workflow of the virtual reality music game system is as follows: Two players simultaneously (either in different locations or in the same venue) play drums and flutes in VR. The real-time processing engine simultaneously extracts the drum and flute sounds from the imported audio files, generating a first audio track file corresponding to the two instruments, and a second audio track file with the two instruments removed. Two sets of interactive prompt information files are generated based on the two first audio track files. The real-time processing engine sends the corresponding game content data, the corresponding interactive prompt information files, and the second audio track files to the VR interaction modules of the two players respectively. One player supplements the drum sound using a virtual drum (third audio track file A), and the other player supplements the flute sound using a virtual flute (third audio track file B). The real-time processing engine receives the two third audio track files and superimposes them onto the second audio track file for playback within the VR interaction module.

[0219] Example 3

[0220] Another object of the present invention is to provide a method of using the virtual reality music game system based on audio separation technology according to Embodiment 1 or Embodiment 2 above, including:

[0221] Players can upload song audio files and select the target instrument to play through the system interface;

[0222] The real-time processing engine receives the song audio file, analyzes and processes it, and generates game content data, interactive prompt files, and a second audio track file. The interactive prompt file contains a sequence of operation prompts generated by the real-time processing engine based on the rhythm and timing of the target instrument's audio track in the song audio file.

[0223] Players interact with the VR interaction module, and the real-time processing engine sends interactive feedback to the VR interaction module.

[0224] Players select their audio source configuration;

[0225] Players play the selected virtual instrument in a virtual reality environment according to the interactive prompts to generate a third audio track file. The VR interactive module overlays the third audio track file onto the second audio track file in real time and plays it, allowing players to reproduce the accompaniment of the input song in a virtual reality environment.

[0226] The VR interaction module outputs real-time operation prompts and feedback from the real-time processing engine.

[0227] After the game ends, the VR interactive module displays the player's score for the game.

[0228] Compared with existing technologies, the present invention provides a method for using a virtual reality music game system based on audio separation technology. This overcomes the problem that existing VR music games only provide songs, game prompts, and game levels pre-set by game developers, resulting in relatively fixed VR music game content. The invention provides a highly flexible and personalized method for using a virtual reality music game system based on audio separation technology, allowing players to flexibly choose the accompaniment of the songs they want to appear in the VR scene according to their own interests or needs, and use virtual instruments to restore the accompaniment, thereby improving the player's entertainment experience.

[0229] Example 4

[0230] Another objective of this invention is to provide a collaborative performance method for a virtual reality music game system based on audio separation technology, which utilizes at least two VR interaction modules.

[0231] See Figure 3 In this embodiment, the VR interaction module of the virtual reality music game system based on audio separation technology is provided with at least two, allowing at least two players to play the game simultaneously. This enables at least two players to collaboratively operate at least two virtual instruments in the virtual reality environment to input at least two third audio track files. The real-time processing engine receives at least two third audio track files and superimposes the at least two third audio track files onto a second audio track file. The real-time processing engine sends the interactive prompt information file corresponding to the virtual instrument selected by the player to the corresponding VR interaction module.

[0232] The collaborative performance method includes: the execution flow of the first VR interaction module and the execution flow of other VR interaction modules (hereinafter referred to as the second VR interaction module execution flow).

[0233] The execution process of the first VR interaction module (inviting party) includes:

[0234] Importing a song: Player A imports the song audio file into the real-time processing engine through the first VR interaction module;

[0235] Sending invitation information: Player A sends a collaborative performance invitation to other VR interaction modules through the first VR interaction module, which includes at least the performance piece and performance scene information (performance scene information includes time and virtual scene theme);

[0236] Loading the virtual scene: If player B sends back an acceptance message through the second VR interaction module, the first VR interaction module and the second VR interaction module load the target virtual scene together.

[0237] Displaying the scene and performing the performance: The first VR interaction module displays the target virtual scene to player A. Player A selects the target instrument to be played in advance. Player A is immersed in the virtual scene through interactive devices such as head-mounted displays and headphones, and operates the virtual instrument to play according to the interactive prompts in the virtual scene (the virtual instrument / simple device and operation prompts (such as the timing and position of striking and pressing) are displayed when the instrument is played).

[0238] Feedback: Player A performs a performance action, and the system obtains the response data and outputs feedback information.

[0239] The execution process of the second VR interaction module (for the invited party) includes:

[0240] Receiving and responding to the invitation: Player B receives the invitation information through the second VR interaction module. After player B confirms, the second VR interaction module sends an acceptance message to the first VR interaction module.

[0241] Load the target scene: Load the corresponding target virtual scene from the server along with the first VR interaction module.

[0242] Participating in collaborative performance: The second VR interaction module displays the target virtual scene to player B. Player B selects the target instrument to be played in advance. Player B is immersed in the virtual scene through interactive devices such as head-mounted displays and headphones, and operates the virtual instrument according to the interactive prompts in the virtual scene, and performs collaborative performance with the player using the first VR interaction module in the same virtual scene.

[0243] Feedback: Player B performs a performance action, obtains its response data, and outputs feedback information.

[0244] The real-time processing engine acquires the response data generated by the playing actions of both players, outputs auditory and visual (including force perception in instrument playing) feedback information, and overlays the third audio track file generated by the VR interaction module onto the second audio track file. This is then played through the VR interaction module to achieve collaborative playing. After the performance, the engine outputs evaluation data (the matching degree of the players' playing, the accuracy of instrument operation, etc.) to the VR interaction module.

[0245] The aforementioned collaborative performance method utilizes virtual reality technology to achieve remote collaborative performance, bringing an immersive experience, breaking geographical limitations, and allowing non-professional users to participate through precise prompts and virtual performance devices, thus lowering the barrier to entry for performance. In addition, the combination of auditory, visual, and tactile feedback enhances the sense of immersion and realism, providing rich and intuitive feedback.

[0246] This embodiment can be combined with the solution in Embodiment 3.

[0247] Based on the disclosure and teachings of the foregoing specification, those skilled in the art can make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. Furthermore, although some specific terms are used in this specification, these terms are only for convenience of explanation and do not constitute any limitation on the present invention.

Claims

1. A virtual reality music game system based on audio separation technology, characterized in that, include: At least one VR interaction module is used to display virtual reality to players and allow players to interact with the virtual reality environment. Real-time processing engine, including: The audio input module is used to receive song audio files input by the player. The audio processing module is used to extract the track of the target instrument from the song audio file, generate the first track file, and remove the track of the target instrument from the song audio file and generate the second track file of the corresponding accompaniment version of the song. The rhythm cue generation module is used to analyze the rhythm and timing of the target instrument sounds in the extracted first audio track file and generate an interactive cue information file suitable for the VR environment. The control and management module is used to control the system's operation, receive and store data files, and manage the data files. The virtual scene generation module generates virtual reality environments; During system operation, the control and management module receives the song audio file, the second audio track file, and the interactive prompt information file, and generates game content data based on the song audio file and virtual reality data. The VR interaction module receives the game content data, interactive prompt information file, and second audio track file output by the real-time processing engine, plays the second audio track file to the player, and displays the virtual reality environment. The displayed virtual reality environment will display rhythm or melody prompts. When the player interacts in the virtual reality environment, the VR interaction module generates a third audio track file. The VR interaction module overlays the third audio track file onto the second audio track file in real time and plays it. The third audio track file can be a first audio track file based on the target instrument audio in the input song, or it can be an audio track file based on the target instrument sound from an external digital audio library.

2. The virtual reality music game system based on audio separation technology according to claim 1, characterized in that, The VR interaction module includes a sound source selection module, which allows players to choose the instrument sound source used when generating a third audio track file.

3. The virtual reality music game system based on audio separation technology according to claim 1, characterized in that, The VR interaction module is equipped with at least two, allowing at least two players to play the game simultaneously, enabling at least two players to collaboratively operate at least two virtual musical instruments in the virtual reality environment to input at least two third audio track files; The real-time processing engine receives at least two third audio track files and integrates the at least two third audio track files into the second audio track file; The real-time processing engine sends the interactive prompt information file corresponding to the virtual instrument selected by the player to the corresponding VR interaction module.

4. The virtual reality music game system based on audio separation technology according to claim 3, characterized in that, The virtual reality music game system features a collaborative performance mode: When the virtual reality music game system is in collaborative performance mode, it includes the execution flow of the first VR interaction module and the execution flow of other VR interaction modules; The execution flow of the first VR interaction module includes: Importing Songs: Players import song audio files into the real-time processing engine through the first VR interaction module; Sending invitation messages: Players send collaborative performance invitation messages to other VR interaction modules through the first VR interaction module, which must include at least the performance piece and performance scene information; Loading virtual scene: If another player sends back an invitation acceptance message through other VR interaction modules, the first VR interaction module loads the target virtual scene together with other VR interaction modules; Scene display and performance execution: The first VR interaction module displays the target virtual scene to the player. The player selects the target instrument to be played in advance and operates the virtual instrument to play according to the interactive prompts in the virtual scene. Feedback: When a player performs a performance action, the system acquires their response data and outputs feedback information. The execution flow of other VR interaction modules includes: Receiving and responding to invitations: Another player receives the invitation information through another VR interaction module. After the player confirms, the other VR interaction module sends an acceptance message to the first VR interaction module. Loading the target scene: The corresponding target virtual scene is loaded from the server along with the first VR interaction module; Participate in collaborative performance: Other VR interaction modules show players the target virtual scene. Players select the target instrument to be played in advance. Players operate the virtual instrument according to the interactive prompts in the virtual scene and perform collaborative performance with the player using the first VR interaction module in the same virtual scene. Feedback: Another player performs a performance action, and the system obtains the response data and outputs feedback information.

5. The virtual reality music game system based on audio separation technology according to claim 1, characterized in that, The real-time processing engine is connected to an artificial intelligence system, and the audio processing module uses the artificial intelligence system to extract the audio tracks from the song's audio file in real time.

6. The virtual reality music game system based on audio separation technology according to claim 1, characterized in that, The real-time processing engine employs both a real-time processing mode and a pre-management mode. In real-time processing mode, after the player inputs the song audio file, the player can immediately operate the VR interaction module to play the game; In the pre-management mode, after the player inputs the song audio file, the real-time processing engine analyzes and processes the song audio file and outputs game content data, interactive prompt information file, and second audio track file. The game content data, interactive prompt information file, and second audio track file are stored in the database. When the player subsequently operates the VR interaction module to play the game, the game content data, interactive prompt information file, and second audio track file are called.

7. The virtual reality music game system based on audio separation technology according to claim 6, characterized in that, The real-time processing engine enters real-time processing mode. After the player inputs the song audio file, the following steps are executed: Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds. Several audio segments are processed using a data model to generate a first audio track file and a second audio track file; the first audio track file is analyzed using the data model to extract rhythm points from the first audio track file and generate a rhythm point data file. Convert the rhythm point data file into an interactive prompt information file that can be mapped by VR devices in the VR environment; And / or, the real-time processing engine enters pre-management mode, and after the player inputs the song audio file, the following steps are performed: Divide the song audio file into several audio segments, with each segment lasting 0.5-1 seconds. Several audio segments are processed using an offline data model to generate a first audio track file and a second audio track file. The second audio track file is then stored. The first audio track file is analyzed using an offline data model to extract rhythm points from the first audio track file, generate a rhythm point data file, and store the rhythm point data file. After the player operates the VR interaction module to load the game and selects a song and sound source, the real-time processing engine calls the second audio track file and rhythm point data file. The real-time processing engine converts the rhythm point data file into an interactive prompt information file that can be mapped by the VR device in the VR environment.

8. The virtual reality music game system based on audio separation technology according to claim 7, characterized in that, In processing several audio segments using a data model, hierarchical parallel computing is employed, including: Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue; Execute the separation thread: Run the INT8 quantization model to extract the track of the selected instrument from the audio segment, obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and integrate the several audio segments containing the selected instrument to generate the first track file; The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file. Use a data model to extract the rhythm points from the entire first audio track file and generate a rhythm point data file. Alternatively, when processing several audio segments using a data model, hierarchical parallel computing can be employed, including: Execute the preprocessing thread: The song audio file is divided into several audio segments and sent to the analysis queue; Execute the separation thread: Run the INT8 quantization model to extract the track of the selected instrument from the audio segment, and obtain several audio segments containing the selected instrument, where the separation accuracy SNR>15dB, and the several audio segments containing the selected instrument are used as a subset of the first track file. The accompaniment generation thread performs reverse mixing of the selected instrument tracks from the removed audio segments, obtaining several audio segments with the selected instruments removed, and then integrating these audio segments to generate a second audio track file. The rhythm analysis thread uses Fast Fourier Transform and dynamic beat detection algorithms to extract the rhythm points of each track segment within a subset of the first audio track file, and integrates the rhythm points of several track segments to generate a rhythm point data file.

9. The virtual reality music game system based on audio separation technology according to claim 8, characterized in that, During the execution of the separate thread, redundant parameters are removed by model pruning and 8-bit integer quantization is used to reduce the processing latency of each audio segment to no more than 25 milliseconds. And / or, during the execution of the rhythm analysis thread, predictive caching is performed based on the periodic detection of music beats to pre-generate rhythm cues for the next 5-20 seconds.

10. A method of using the virtual reality music game system based on audio separation technology according to any one of claims 1 to 9, comprising: Players can upload song audio files and select the target instrument to play through the system interface; The real-time processing engine receives the song audio file, analyzes and processes it, and generates game content data, interactive prompt files, and a second audio track file. The interactive prompt file contains a sequence of operation prompts generated by the real-time processing engine based on the rhythm and timing of the target instrument's audio track in the song audio file. Players interact with the VR interaction module, and the real-time processing engine sends interactive feedback to the VR interaction module. Players select their audio source configuration; Players play the selected virtual instrument in a virtual reality environment according to interactive prompts to generate a third audio track file. The VR interaction module overlays the third audio track file onto the second audio track file in real time and plays it, allowing players to recreate the accompaniment of the input song in the virtual reality environment.

Citation Information

Patent Citations

  • Collaborative performance methods, systems, terminal devices, and storage media

    CN112203114B