Cognitive interaction method based on virtual reality rhythm game, electronic device and storage medium
By loading music and score data into a virtual reality environment, controlling dynamic note sequences, and generating real-time feedback, the problem of loose coupling between limb movement and brain cognition in virtual reality cognitive training systems is solved. This achieves highly precise cognitive intervention and adaptive adjustment, enhancing user immersion and training effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE HONG KONG POLYTECHNIC UNIV SHENZHEN RES INST
- Filing Date
- 2026-03-24
- Publication Date
- 2026-06-12
AI Technical Summary
Existing virtual reality cognitive training systems fail to effectively combine rigorous cognitive psychology paradigms with users' sensory-physical interactions, resulting in a loose coupling between users' physical movements and cognitive processing when performing tasks. They also fail to use music rhythms to guide time attention allocation, and the matching feedback lacks specificity, thus affecting training effectiveness.
By loading target music data and score data into a virtual reality environment, controlling the dynamic note sequence to move according to a preset rhythm, and presenting cognitive stimulation content at the trigger position, combined with user interaction to generate real-time feedback, a rigorous N-order retrospective cognitive psychology paradigm and multi-sensory physical interaction in three-dimensional space are deeply integrated.
It enhances user immersion and intrinsic motivation, closely couples physical movement with brain cognitive processing, improves the accuracy of cognitive intervention and the ability to adaptively adjust real-time cognitive load, and significantly improves the effectiveness of working memory training.
Smart Images

Figure CN122183169A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer graphics technology, and in particular to cognitive interaction methods, electronic devices, and storage media based on virtual reality rhythm games. Background Technology
[0002] In the fields of modern medical health and cognitive psychology, the training and intervention of core executive functions such as working memory is a crucial technical aspect, as its effectiveness directly impacts a user's reasoning ability and the quality of their daily activities. Traditional cognitive training is typically based on the classic N-back working memory task paradigm, requiring users to mechanically match and judge sequentially presented visual or auditory stimuli in front of a two-dimensional display device using a keyboard or mouse. While this traditional intervention method has a scientific theoretical basis, the highly repetitive and monotonous task process leads to extremely low intrinsic motivation and long-term training compliance, making it difficult to maintain training effects in the long term. With the development of computer graphics and human-computer interaction technologies, Virtual Reality (VR) technology has been gradually introduced into the cognitive rehabilitation and training industry to enhance user engagement and training experience through immersive three-dimensional environments and gamified elements.
[0003] In related technologies, to achieve more engaging cognitive training, simple memory test cases are often embedded into virtual reality environments, and basic visual cues are used to guide user interaction. However, existing virtual reality cognitive systems often fail to deeply integrate rigorous cognitive psychology paradigms with users' sensory-physical interactions, resulting in a loose coupling between the user's physical movements and the brain's cognitive processing during task execution. This approach not only fails to effectively utilize rhythmic cues such as musical rhythms to dynamically guide the user's time attention allocation, but also struggles to provide accurate matching feedback when faced with complex dynamic musical sequences, failing to combine historical stimuli with real-time interaction timing. Consequently, while existing virtual reality cognitive training methods enhance user immersion, their accuracy in cognitive intervention and their adaptive adjustment capabilities to real-time cognitive load remain relatively low. Summary of the Invention
[0004] This application provides a cognitive interaction method, electronic device, and storage medium based on virtual reality rhythm games, which can enhance user immersion while effectively improving the accuracy of cognitive intervention and the adaptive adjustment capability for real-time cognitive load.
[0005] To achieve the above objectives, a first aspect of this application proposes a cognitive interaction method based on a virtual reality rhythm game, the method comprising: Load the target music data and the corresponding score data according to the user's current interaction configuration parameters; When playing the target music data, a dynamic note sequence is controlled to move along the note track of the virtual environment to the interactive area according to a preset rhythm. The dynamic note sequence is generated based on the score data. When at least one target note in the dynamic note sequence moves to the trigger position, the cognitive stimulus content corresponding to the target note is presented. In response to a user's interactive operation on the target note based on the interactive area, the historical stimulus content of the historical note is obtained, and based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interactive operation, a cognitive matching result for the target note is obtained. Based on the cognitive matching results and the timing of the interactive operation, real-time feedback information is generated.
[0006] In some embodiments, the current interaction configuration parameters include a preset number of backtracking steps. The step of acquiring historical stimulus content for historical notes and obtaining a cognitive matching result for the target note based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interaction operation includes: Identify the historical notes in the dynamic note sequence that are separated from the target note by a preset number of backtracking steps, and obtain the historical stimulus content of the historical notes; The content consistency determination result is obtained by comparing the cognitive stimulus content with the historical stimulus content; The cognitive matching result is obtained based on the stimulus type corresponding to the target note, the content consistency judgment result, and the correspondence between the interaction position and the preset hitting sub-region.
[0007] In some embodiments, the stimulus type includes visual and auditory stimuli, and the preset striking sub-region includes a first striking sub-region, a second striking sub-region, and a third striking sub-region. The cognitive matching result is obtained based on the stimulus type corresponding to the target note, the content consistency determination result, and the correspondence between the interaction position and the preset striking sub-region, including: When the stimulus type is the visual stimulus, and the content consistency determination result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the first hitting sub-region, then the cognitive matching result indicating correctness is generated. When the stimulus type is the auditory stimulus, and the content consistency determination result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the second hitting sub-region, then the cognitive matching result indicating correctness is generated. When the content consistency determination result indicates that the cognitive stimulus content is inconsistent with the historical stimulus content, or when the target note is an ordinary note that does not contain cognitive stimulus content, and the interaction position is located in the third striking sub-region, then the cognitive matching result that indicates correctness is generated.
[0008] In some embodiments, the step of generating the dynamic note sequence includes: Based on the time node information and note type information in the spectrum data, an initial note sequence is generated; In the initial note sequence, an interference position is determined that is a target number of steps away from the stimulus note containing the target stimulus content, and an interference stimulus note is generated at the interference position according to a preset probability. The dynamic note sequence is obtained based on the updated initial note sequence. Wherein, the cognitive stimulus content of the interference stimulus note is the same as the target stimulus content, and the target number of steps is obtained based on the preset backtracking number of steps.
[0009] In some embodiments, generating real-time feedback information based on the cognitive matching result and the timing of the interaction operation includes: When the cognitive matching result is correctly represented, the actual trigger time of the interactive operation is obtained, and the time difference is obtained based on the difference between the actual trigger time and the theoretical arrival time of the target note to the center of the interactive area. The time difference is matched with multiple preset performance evaluation time windows to obtain a matching time window, and real-time feedback information corresponding to the evaluation level of the matching time window is generated based on the matching time window.
[0010] In some embodiments, the current interaction configuration parameters include the music difficulty level and the preset backtracking steps, and the method further includes: Within the current interaction task cycle, determine the accuracy of the cognitive matching results corresponding to multiple interaction operations; When the accuracy rate reaches a first preset threshold, the music difficulty level is increased; When the music difficulty level reaches the highest level and the accuracy reaches the second preset threshold, the music difficulty level is maintained and the preset backtracking steps are increased.
[0011] In some embodiments, before loading the target music data according to the user's current interaction configuration parameters, the method further includes: Obtain the original audio file, and obtain the periodic beat characteristics and beat speed characteristics of the original audio file; Based on the aforementioned periodic beat characteristics, the timeline of the original audio file is divided into multiple beat cycles; Within each beat cycle, based on the beat speed characteristics, empty modules, stimulus modules, and non-stimulation modules are arranged according to a preset repetition pattern to generate the spectral data.
[0012] In some embodiments, arranging empty modules, stimulation modules, and non-stimulation modules according to a preset repetition pattern based on the beat speed characteristics includes: Establish a mapping model between the beat velocity characteristics and the stimulus interval time; The time interval between the stimulation modules is determined based on the mapping model, and the empty module or the non-stimulation module is inserted within the time interval.
[0013] To achieve the above objectives, a second aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the cognitive interaction method based on a virtual reality rhythm game as described in the first aspect.
[0014] To achieve the above objectives, a third aspect of this application provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the cognitive interaction method based on virtual reality rhythm games described in the first aspect.
[0015] The cognitive interaction method, electronic device, and storage medium based on virtual reality rhythm games proposed in this application include: First, loading target music data and corresponding score data according to the user's current interaction configuration parameters; then, while playing the target music data, controlling a dynamic note sequence to move along the note track of the virtual environment to the interaction area according to a preset rhythm, the dynamic note sequence being generated based on the score data; subsequently, when at least one target note in the dynamic note sequence moves to the trigger position, presenting the cognitive stimulus content corresponding to the target note; next, in response to the user's interaction operation on the target note based on the interaction area, acquiring the historical stimulus content of historical notes, and obtaining a cognitive matching result for the target note based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interaction operation; finally, generating real-time feedback information based on the cognitive matching result and the interaction timing of the interaction operation. This application's embodiments load and play target music data and corresponding score data according to the user's current interaction configuration parameters, control the dynamic note sequence to move towards the interactive area on the note track of the virtual environment according to the preset music rhythm, and present cognitive stimulation content at the trigger position. It can effectively guide the user's time attention allocation by utilizing the rhythmic cues of music, breaking the limitations of the monotony of traditional working memory training, and greatly improving the user's immersion, intrinsic motivation, and long-term training compliance. At the same time, the solution of this application compares the extracted historical stimulation content with the current cognitive stimulation content in response to the user's tapping operation in the interactive area, and obtains a cognitive matching result by combining the specific interaction position. Then, based on the matching result and the timing of the tapping interaction, it generates millisecond-level real-time feedback information, thereby realizing the deep integration of the rigorous N-order back cognitive psychology paradigm and multi-sensory physical interaction in three-dimensional space. It closely couples the user's physical movement and brain cognitive processing, effectively solving the technical problems of low intervention accuracy and lack of targeted matching feedback in existing virtual reality cognitive systems, and significantly improving the training efficiency and intervention quality of core executive functions such as working memory.
[0016] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0017] Figure 1 This is a flowchart of a cognitive interaction method based on a virtual reality rhythm game provided in an embodiment of this application.
[0018] Figure 2 This is a flowchart of the generation process of spectral data provided in another embodiment of this application.
[0019] Figure 3 yes Figure 2 The flowchart for step 203.
[0020] Figure 4 This is a schematic diagram of the three-dimensional spatial layout and core interaction mechanism of an immersive virtual environment for cognitive interaction provided in another embodiment of this application.
[0021] Figure 5 This is a flowchart of the generation of dynamic note sequences provided in another embodiment of this application.
[0022] Figure 6 This is another embodiment of the present application, which provides a spatial arrangement of a dynamic note sequence on a note track and a detailed schematic diagram of the cognitive stimulus content presented by different types of target notes at the trigger position.
[0023] Figure 7 yes Figure 1 The flowchart for step 104.
[0024] Figure 8 yes Figure 7 The flowchart for step 703.
[0025] Figure 9 This is a schematic diagram of a virtual interactive scene in which a user performs interactive operations under a visual channel working memory task, as provided in another embodiment of this application.
[0026] Figure 10 This is another embodiment of the present application that provides a virtual interactive scenario in which a user performs interactive operations under an auditory channel working memory task.
[0027] Figure 11 This is another embodiment of the present application that provides a virtual scenario in which a user performs interactive operations when faced with interfering stimuli that do not conform to memory matching rules.
[0028] Figure 12 This is another embodiment of the present application, which provides a virtual interactive scenario in which a user performs interactive operations when faced with ordinary musical notes that do not require cognitive processing.
[0029] Figure 13 yes Figure 1 The flowchart for step 105.
[0030] Figure 14 This is a flowchart of an update feedback process provided in another embodiment of this application.
[0031] Figure 15 This is a schematic diagram of the overall execution flow of a cognitive interaction method based on a virtual reality rhythm game, provided in another embodiment of this application.
[0032] Figure 16This is a schematic diagram of the overall architecture and module data flow of a cognitive interaction system based on a virtual reality rhythm game, provided in another embodiment of this application.
[0033] Figure 17 This is a schematic diagram of the hardware structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0035] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0037] In the fields of modern medical health and cognitive psychology, the training and intervention of core executive functions such as working memory is a crucial technical aspect, as its effectiveness directly impacts a user's reasoning ability and the quality of their daily activities. Traditional cognitive training is typically based on the classic N-back working memory task paradigm, requiring users to mechanically match and judge sequentially presented visual or auditory stimuli in front of a two-dimensional display device using a keyboard or mouse. While this traditional intervention method has a scientific theoretical basis, the highly repetitive and monotonous task process leads to extremely low intrinsic motivation and long-term training compliance, making it difficult to maintain training effects in the long term. With the development of computer graphics and human-computer interaction technologies, Virtual Reality (VR) technology has been gradually introduced into the cognitive rehabilitation and training industry to enhance user engagement and training experience through immersive three-dimensional environments and gamified elements.
[0038] In related technologies, to achieve more engaging cognitive training, simple memory test cases are often embedded into virtual reality environments, and basic visual cues are used to guide user interaction. However, existing virtual reality cognitive systems often fail to deeply integrate rigorous cognitive psychology paradigms with users' sensory-physical interactions, resulting in a loose coupling between the user's physical movements and the brain's cognitive processing during task execution. This approach not only fails to effectively utilize rhythmic cues such as musical rhythms to dynamically guide the user's time attention allocation, but also struggles to provide accurate matching feedback when faced with complex dynamic musical sequences, failing to combine historical stimuli with real-time interaction timing. Consequently, while existing virtual reality cognitive training methods enhance user immersion, their accuracy in cognitive intervention and their adaptive adjustment capabilities to real-time cognitive load remain relatively low.
[0039] To enhance user immersion while effectively improving the accuracy of cognitive intervention and its adaptive adjustment capability to real-time cognitive load, this application's embodiments load and play target music data and corresponding score data based on the user's current interaction configuration parameters. It controls a dynamic note sequence to move along a preset musical rhythm on a note track within the virtual environment towards the interactive area, presenting cognitive stimuli at trigger locations. This effectively guides the user's time attention allocation using the rhythmic cues of music, breaking through the monotonous limitations of traditional working memory training and significantly improving user immersion, intrinsic motivation, and long-term training adherence. Simultaneously, this application's solution utilizes the rhythmic cues of music to effectively guide the user's time attention allocation, overcoming the limitations of traditional working memory training's monotony and greatly enhancing user immersion, intrinsic motivation, and long-term training compliance. In response to the user's tapping action in the interactive area, the extracted historical stimulus content is compared with the current cognitive stimulus content, and a cognitive matching result is obtained by combining the specific interaction location. Then, based on the matching result and the timing of the tapping interaction, millisecond-level real-time feedback information is generated. This achieves a deep integration of the rigorous N-back cognitive psychology paradigm with multi-sensory physical interaction in three-dimensional space, closely coupling the user's physical movement and brain cognitive processing. It effectively solves the technical problems of low intervention accuracy and lack of targeted matching feedback in existing virtual reality cognitive systems, and significantly improves the training efficiency and intervention quality of core executive functions such as working memory.
[0040] The cognitive interaction method, electronic device, and storage medium based on virtual reality rhythm games provided in this application will be further described below. First, the cognitive interaction method based on virtual reality rhythm games in the embodiments of this application will be described in detail. (Refer to...) Figure 1 This is an optional flowchart of a cognitive interaction method based on virtual reality rhythm games provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps 101 to 105. It is also understood that this embodiment... Figure 1The order of steps 101 to 105 is not specifically limited; the order of steps can be adjusted or certain steps can be added or removed according to actual needs. The cognitive interaction method based on virtual reality rhythm games provided in this application can be applied to processing systems with virtual reality computing capabilities (such as smart terminals, servers, computing processors, etc.).
[0041] Step 101: Load the target music data and the corresponding score data according to the user's current interaction configuration parameters.
[0042] Step 101 will be described in detail below.
[0043] In step 101 of some embodiments, the system loads target music data and corresponding score data based on the user's current interaction configuration parameters. Specifically, the user's current interaction configuration parameters may include the user's identity (ID), historical training records, and current cognitive load level or difficulty level. The data module of the processing system parses and matches the above parameters and retrieves target music data (e.g., audio files in mp3, wav, etc. formats) that meet the current difficulty requirements from a preset music library.
[0044] Simultaneously, the processing system loads the score data that strictly corresponds to the target music data. This score data refers to a dataset pre-generated using acoustic feature detection algorithms (such as drum beat detection and beats per minute, BPM), containing information on musical rhythm timing, note triggering times, and cognitive stimulus associations. Through this step, the system prepares the foundational data that balances musical rhythm and the underlying logic of psychological cognitive tasks for the user before they officially enter the interactive environment.
[0045] The following section will describe in more detail how to generate spectral data.
[0046] Reference Figure 2 Before loading the target music data based on the user's current interaction configuration parameters, the method also includes the following steps 201 to 203.
[0047] Step 201: Obtain the original audio file and its periodic beat characteristics and beat speed characteristics.
[0048] Step 202: Divide the timeline of the original audio file into multiple beat cycles based on the periodic beat characteristics.
[0049] Step 203: Within each beat cycle, based on the beat speed characteristics, empty modules, stimulus modules and non-stimulation modules are arranged according to a preset repetition pattern to generate spectral data.
[0050] Steps 201 to 203 are described in detail below.
[0051] In step 201 of some embodiments, the original audio file is acquired, and its periodic beat characteristics and tempo characteristics are obtained. Specifically, the system's score generation module imports the original audio file from a preset music library. The format of this audio file can include, but is not limited to, mainstream audio formats such as mp3, wav, flac, and ogg. Subsequently, the system uses a sound detection algorithm to extract features from the input original audio file to identify and obtain the periodic beat characteristics (i.e., the timestamps of drum beats in the music) and tempo characteristics (i.e., BPM, BeatsPerMinute, representing the number of beats or drum beats per minute) within the music file.
[0052] In practical applications, the extraction of such audio features can be achieved using existing audio processing techniques, such as using Python's Librosa module or the deep learning-based Madmom model for automated, high-precision detection, thereby providing accurate basic data support for the subsequent synchronization of game beats with cognitive stimuli.
[0053] In step 202 of some embodiments, the timeline of the original audio file is divided into multiple beat cycles based on periodic beat characteristics. After obtaining all key drum beat information in the original audio file, the system needs to establish a structured time segmentation logic. The system first defines a starting point on the timeline, for example, taking the detected 5th drum beat or other specific drum beats as the starting point for the generation of the entire score.
[0054] After determining the starting point, the system divides all drum beat sequences from that starting point to the end of the audio file into multiple consecutive cycles according to a fixed number of beats. For example, depending on the difficulty, each complete beat cycle can be set to consist of 2, 4, or 6 drum beats. This method of discretizing and periodizing the continuous audio timeline lays a strict time framework foundation for the subsequent regular insertion of cognitive interactive tasks.
[0055] In step 203 of some embodiments, within each beat cycle, based on the beat speed characteristics, empty modules, stimulus modules, and non-stimulus modules are arranged according to a preset repetition pattern to generate spectrogram data. To achieve a deep integration of rigorous cognitive psychology tasks and music games, the system defines three distinctly different underlying module units. An "empty module" indicates that no visual or auditory object appears at the beat point, primarily used to control task density and provide the user's brain with buffer time for information processing. A "stimulus module" represents an interactive object (i.e., a stimulus note) that requires substantial user participation and requires judgment based on historical stimulus content to determine whether it conforms to the N-back matching rule. A "non-stimulus module" indicates that an object appears at the beat point, but this object only serves to maintain the continuity of the game's beat; the user does not need to perform complex N-back cognitive matching judgments (corresponding to the ordinary notes mentioned earlier). It should be noted that, to facilitate rapid visual or auditory recognition within millisecond-level reaction time, stimulus modules and non-stimulus modules are designed to have clearly distinct performance characteristics.
[0056] In step 203 of some embodiments, the system systematically arranges these three modules according to the preset repetition pattern within each predefined beat cycle. The specific arrangement strategy can be flexibly configured according to actual cognitive load requirements: for example, if a cycle consists of 4 drumbeats, the system can generate the corresponding modules sequentially at these 4 drumbeats according to a fixed sequence of "empty-stimulus-empty-non-stimulus"; if a cycle consists of 2 drumbeats, the modules can be arranged in a compact pattern of "stimulus-non-stimulus" or "empty-stimulus"; if a cycle consists of 6 drumbeats, a loose pattern of "empty-empty-stimulus-empty-empty-non-stimulus" can be used. The system packages these regularly arranged module data with the time nodes, BPM, and associated cognitive stimulus content of the original audio file, ultimately generating spectrogram data that can be directly called by the cognitive interaction stage, as described below.
[0057] Reference Figure 3 Based on the beat speed characteristics, empty modules, stimulation modules and non-stimulation modules are arranged according to a preset repetition pattern, including the following steps 301 to 302.
[0058] Step 301: Establish a mapping model between beat speed characteristics and stimulus interval time.
[0059] Step 302: Determine the time interval between stimulus modules based on the mapping model, and insert empty modules or non-stimulation modules within the time interval.
[0060] Steps 301 to 302 are described in detail below.
[0061] In step 301 of some embodiments, the system establishes a mapping model between beat speed characteristics and stimulus interval time. Specifically, in the field of cognitive psychology, the stimulus interval time (ISI) is a core parameter determining the cognitive load of working memory tasks (such as the N-back paradigm), while the beat speed characteristic (i.e., beats per minute, BPM) is a fundamental attribute characterizing the physical rhythm of music itself. To scientifically couple these two, the system internally constructs a mathematical mapping relationship that automatically converts musical beats into cognitive load time windows. For example, the system can set the target note requiring cognitive matching judgment to appear once every two beats, and the system calculates the corresponding time interval accordingly. And ISI is set to In this mapping model, the system strictly sets the stimulus interval (ISI) of the training task to be equal to this time interval. By establishing this deterministic mapping model, the system can ensure that the frequency of cognitive training is directly controlled by the inherent rhythm of the music. For example, when the system detects that the BPM of the input music is in the range of 80 to 160, the mapping model can stably constrain the stimulus interval ISI and keep it between 1500 and 3000 milliseconds. This quantification range not only ensures the synchronization of the beat, but more importantly, it strictly meets the effective difficulty requirements of the N-back cognitive training task, avoiding the adverse effects of stimuli appearing too quickly or too slowly on the intervention effect.
[0062] In step 302 of some embodiments, the time interval between stimulus modules is determined based on a mapping model, and empty or non-stimulus modules are inserted within the time interval. After locking the absolute time span between two adjacent "stimulus modules" (i.e., notes containing cognitive memory tasks) through the mapping model in step 301, the system needs to ensure the continuity of interaction within this span to maintain the immersion of the virtual reality rhythm game. Since the user's cognitive processing channel requires a certain buffer time when processing N-back comparisons, if a high-load cognitive task is continued to be arranged between two stimulus modules, it will lead to overload of the user's working memory and frustration. Therefore, the system will intelligently insert modules that do not require working memory matching within the determined time interval, according to the preset periodic pattern of the spectrum data.
[0063] Specifically, the system can insert "empty modules" (i.e., no virtual objects are presented at that time point, giving the user a pure cognitive and visual buffer space) or "non-stimulation modules" (i.e., ordinary musical notes that only require the user to mechanically strike without the brain having to make N-back rule judgments). By interspersing these low-load or zero-load module units within fixed stimulation intervals, the system not only fills the gaps in the musical rhythm, ensuring that the density of the generated dynamic note sequence is perfectly synchronized with the auditory rhythm of the original audio, but also cleverly balances the coherence of game operation with the scientific nature of the brain's cognitive load.
[0064] In one example workflow, the original audio file is first acquired, and advanced computer sound detection technology is used to automatically identify the underlying features of the music file. Specifically, the system can call the Librosa module in the Python environment or use tools such as the Madmom model based on the deep learning framework to perform high-precision automatic detection on the original audio file, thereby accurately obtaining the periodic beat features (such as the timestamp position of each drum beat in the audio) and beat speed features (i.e., BPM, Beats Per Minute, representing the number of drum beats per minute or per second) of the original audio file.
[0065] After acquiring the aforementioned underlying features, the system divides the timeline of the original audio file into multiple beat cycles based on periodic beat characteristics. In the specific division process, the system first defines a starting point, for example, using the 5th drumbeat in the detected music file (or other suitable drumbeats after evaluation) as the starting point for the first appearance of the score. Then, the entire drumbeat sequence from this starting point to the end of the audio is divided into multiple cycles according to a fixed number. During this process, combined with step 301, the system establishes a mapping model between beat velocity (BPM) and stimulus interval time. The core function of this mapping model is to transform the physical beat rhythm of the music itself into the time span between two adjacent stimuli in cognitive training, ensuring that the final generated task rhythm meets the difficulty requirements of cognitive psychology.
[0066] After defining the periodic divisions and time intervals, the system determines the time intervals between stimulus modules based on a mapping model, and inserts empty or non-stimulus modules within these intervals. An "empty module" refers to a module where nothing appears at that beat point; a "stimulus module" is a module that presents information requiring the user to determine whether it conforms to the N-back working memory rule (this can be auditory or non-auditory visual stimuli); and a "non-stimulus module" is a module where a specific object appears, but the user does not need to perform an N-back judgment. To ensure the accuracy of the interaction, stimulus and non-stimulus modules are clearly distinguishable to facilitate quick user judgment.
[0067] Within each beat cycle, the system arranges the three modules according to a preset repetition pattern based on beat speed characteristics, generating spectrogram data. For example, the system specifies that only one stimulus module must appear in each cycle, with the remaining positions filled by empty and non-stimulus modules. If the system uses four drum beats as a cycle, these three modules can be repeated at the drum beats according to the pattern of "empty-stimulus-empty-non-stimulus"; if, to increase task density, a two-drum-beat cycle is used, modules can appear in a compact pattern of "stimulus-non-stimulus" or "empty-stimulus"; if, to reduce task density, a six-drum-beat cycle is used, modules can appear in a looser pattern of "empty-empty-stimulus-empty-empty-non-stimulus". Through this arrangement and insertion method based on a mapping model and cycle patterns, the system can automatically generate spectrogram data that balances musical immersion with scientific cognitive load.
[0068] Through steps 301 and 302 above, the underlying algorithm logic effectively separates and reintegrates the entertainment and serious medical cognitive attributes of music. This application's solution precisely maps and converts highly non-standardized external music file parameters (BPM) into standardized, quantifiable psychological experimental parameters (ISI), completely breaking through the technical bottleneck of conventional virtual reality music games' difficulty in controlling instantaneous cognitive load. Simultaneously, through the flexible insertion of "empty modules" and "non-stimulation modules," the system maintains the rhythmic continuity of the limb striking movements while ensuring the brain receives sufficient information processing space, thus providing users with a training program that combines scientific cognitive intervention validity with a highly immersive game-like flow experience.
[0069] Through the offline preprocessing steps 201 to 203 above, the technical effect of automatically converting ordinary commercial audio files into professional cognitive training materials is achieved. This application's solution not only eliminates the tedious work of manually creating scoreboards, which is heavily reliant on traditional music games, significantly improving the efficiency and versatility of scoreboard generation, but more importantly, it creatively quantifies and aligns the rigorous N-back cognitive psychology paradigm with the physical beat cycle of music itself. This periodic module arrangement ensures that the intervals between stimuli are precisely constrained by the music's beat rate (BPM), enabling the final generated dynamic note sequence to perfectly match the music's inherent rhythm while also meeting the scientific training requirements of specific cognitive loads. This lays a solid and scientific data foundation for subsequent highly immersive interaction in a virtual environment.
[0070] Step 102: When playing the target music data, control the dynamic note sequence to move along the note track of the virtual environment to the interactive area according to the preset rhythm. The dynamic note sequence is generated based on the score data.
[0071] Step 102 is described in detail below.
[0072] In step 102 of some embodiments, when playing target music data, the system controls a dynamic note sequence to move along the note track of the virtual environment towards the interactive area according to a preset rhythm. The dynamic note sequence is generated based on the score data. During this process, the virtual environment construction module renders a training scene with spatial depth, and the note generation module instantiates interactive note objects according to the loaded score data at preset time points.
[0073] To ensure visual consistency with underlying parameters, target notes are typically generated at a starting point P_0 at a preset distance (e.g., 25 meters) from the interactive area (e.g., the virtual drumhead position P_2), and move towards the user's interactive area at a certain speed. ,in, The distance between P_1 (stimulus occurrence point) and P_2 (drumhead position, i.e., interaction position) can be set to 7.5 meters. The movement process is strictly synchronized with the rhythm of the target music data (i.e., the preset rhythm). For example, the system will establish a linear coupling relationship between the stimulus interval (ISI) and the music's BPM. When the BPM is in the range of 80–160, the stimulus interval between adjacent notes is precisely controlled between 1500–3000 ms to ensure that the falling density of the note sequence can truly reflect and meet the set cognitive load requirements.
[0074] Reference Figure 4 This is a schematic diagram of the three-dimensional spatial layout and core interaction mechanism of an immersive virtual environment for cognitive interaction provided in an embodiment of this application. Figure 4 As shown in the diagram, in this virtual environment, the system constructs a central musical note track with a strong sense of spatial depth and motion visual feedback (such as the particle acceleration lines around the perimeter). A dynamic musical note sequence moves at a constant speed towards the user from a distance along this track, clearly distinguishing between square-shaped stimulus notes and circular ordinary notes. A preset trigger position P_1 (i.e., the stimulus appearance point) is set at the beginning of this track. When the square target note in the dynamic musical note sequence moves along the track and crosses the dotted line marking the P_1 position, the system will trigger and present the specific cognitive stimulus content corresponding to that target note (e.g., ...). Figure 4 The specific pattern displayed on the surface of the square block at the front, while the circular ordinary musical notes remain unchanged when passing through this area; at the end of the track, adjacent to the user's standing side, the hitting area P_2 (i.e., the virtual drumhead position, which is also the interaction position) is set as the core interactive interface. This P_2 area is clearly divided into three independent hitting sub-areas in physical space: left, middle, and right. Figure 4The text further describes the moment when the user interacts with the system using two virtual drumsticks mapped by the handheld device. Specifically, when the user's working memory determines that the visual stimulus pattern presented by the current square target note is consistent with the historical visual stimulus pattern N steps ago (i.e., it conforms to the single-channel visual N-back matching rule), the user controls the left virtual drumstick to strike the left striking sub-region (i.e., the first striking sub-region) in the P_2 area. This clearly reflects the complete execution link of the system transforming the abstract cognitive matching result into the corresponding physical space regional striking action.
[0075] The process of generating dynamic note sequences will be described in further detail below.
[0076] Reference Figure 5 The steps for generating a dynamic note sequence include steps 501 to 502.
[0077] Step 501: Generate an initial note sequence based on the time node information and note type information in the score data.
[0078] Step 502: In the initial note sequence, determine the interference position that is a target number of steps away from the stimulus note containing the target stimulus content, and generate interference stimulus notes at the interference position according to a preset probability, and obtain a dynamic note sequence based on the updated initial note sequence; wherein, the cognitive stimulus content of the interference stimulus note is the same as the target stimulus content, and the target number of steps is obtained based on the preset backtracking number of steps.
[0079] Steps 501 to 502 are described in detail below.
[0080] In step 501 of some embodiments, the system generates an initial note sequence based on the time node information and note type information in the score data. Specifically, the note generation module inside the system reads the pre-processed score data and extracts the trigger timestamp (i.e., time node information) corresponding to each beat and the classification label (i.e., note type information) bound to that node. The note type information here typically covers three basic types: ordinary notes that do not require additional cognitive stimulation tasks, visual stimulation notes that trigger visual cognitive stimulation elements (such as specific colors, graphics, etc.) at specific time points, and auditory stimulation notes that are associated with stimulation audio and trigger auditory stimulation elements (such as voice prompts). Based on this extracted basic information, the system sequentially instantiates and generates corresponding 3D models or audio objects in the virtual environment, thereby constructing a basic initial note sequence that is absolutely synchronized with the rhythm of the music track. This initial sequence constitutes the basic temporal and spatial framework for subsequent cognitive interaction, ensuring the smooth operation of the basic flow of the rhythm game.
[0081] In step 502 of some embodiments, the system determines, within the initial note sequence, interference positions that are a target number of steps away from the stimulus notes containing the target stimulus content. Interference stimulus notes are generated at these interference positions according to a preset probability, and a dynamic note sequence is obtained based on the updated initial note sequence. The cognitive stimulus content of the interference stimulus notes is the same as the target stimulus content, and the target number of steps is obtained based on a preset backtracking number. This step is the core mechanism for improving the depth of cognitive training, namely, introducing "tempting stimuli" to increase the interference level of the task within the classic N-back (N-order backtracking) paradigm. Specifically, assuming the current user's preset backtracking number of steps (i.e., the N value in N-order backtracking) is N, the system calculates the target number of steps as N±1. The system finds interference positions in the initial note sequence that are N+1 or N-1 positions away from the real target stimulus notes, and at these interference positions, randomly inserts interference stimulus notes that are exactly the same as the real target stimulus content (e.g., presenting highly recognizable geometric shapes or sound effects similar to the real target) according to a certain preset probability (e.g., approximately 33% in actual use). Because the cognitive stimuli of these interfering notes are highly familiar and deceptive, but their sequential positions do not conform to the strict N-step matching rule, users cannot rely solely on short-term memory familiarity with the surface features of the stimuli for conditioned reflex-like mechanical striking. Instead, they must rely on higher-order executive functions for cognitive inhibition control. The system seamlessly integrates these randomly generated interfering notes into the initial note sequence, updates the sequence data, and ultimately outputs a dynamic note sequence that directly drives virtual scene rendering and interaction decisions.
[0082] Through steps 501 and 502 above, a highly intelligent dynamic note sequence generation scheme with an "anti-suspicion" mechanism is realized. This scheme not only automatically transforms static musical score data into interactive object sequences with precise time and type attributes, but more importantly, by introducing an enticing stimulus mechanism based on N±1 steps, it significantly increases the anti-interference dimension and complexity of the cognitive matching task. This design effectively avoids the experience effect that users develop during long-term cognitive training, preventing users from guessing answers based on simple familiarity. This forces users to continuously utilize deep cognitive processing and inhibition abilities during fast-paced music interaction, significantly improving the scientific rigor and actual rehabilitation intervention effect of working memory training while ensuring an immersive gaming experience.
[0083] Step 103: When at least one target note in the dynamic note sequence moves to the trigger position, the cognitive stimulus content corresponding to the target note is presented.
[0084] Step 103 will be described in detail below.
[0085] In step 103 of some embodiments, when at least one target note in the dynamic note sequence moves to the trigger position, the cognitive stimulus content corresponding to the target note is presented. This trigger position (P_1) is typically located at a spatial node on the note's movement trajectory at a specific distance (e.g., 7.5 meters) from the interaction area P_2. When the target note reaches the trigger position P_1, the system activates and presents the corresponding cognitive stimulus content based on the note type information. If the target note is defined as a visual stimulus note (e.g., represented by a square block), specific patterns, colors, or other visual elements that need to be memorized are presented to the user at this moment; if it is an auditory stimulus note, specific sound effects or semantic cues are played synchronously through the audio channel; if it is an ordinary note (e.g., represented by a circular block), it does not contain any cognitive stimulus content that needs to be matched and continues to move towards the user as a non-stimulus element. This step realizes the transformation of the cognitive task from implicit data to explicit multi-sensory cues, using changes in spatial location to explicitly prompt the user to begin cognitive processing of working memory.
[0086] Reference Figure 6 This is a schematic diagram illustrating the spatial arrangement of a dynamic note sequence on a note track and the details of the cognitive stimulus content presented by different types of target notes at the trigger position, as provided in the embodiments of this application. Figure 6 As shown, in the virtual environment, different types of musical notes move from far to near towards the user's area along a note track according to the score data. The interactive area at the end of this track is clearly divided into three independent physical interactive sub-interfaces: a visual striking area, a normal striking area, and an auditory striking area. In the moving dynamic note sequence, the notes are visually distinguished into circular normal notes and square stimulus notes. When a square stimulus note falls and crosses the trigger position, the system will stimulate the corresponding cognitive stimulus content according to the set channel mode. Specifically, combined with... Figure 6 The "Examples of Stimulus Notes" module on the right side shows that if the target note is defined as a single-channel visual stimulus note, its square surface will display specific visual geometry (such as...). Figure 6 The five-pointed star pattern shown is used for visual memory processing by the user; if it is a single-channel auditory stimulus (such as...) Figure 6 (The system displays squares with speaker icons). Instead of directly visually displaying the memorized characters, it synchronously broadcasts specific auditory stimuli (such as...) through the virtual environment's audio channel. Figure 6The voice-announced number "1" indicated by the bubble in the middle provides users with auditory memory processing. If the system is in dual-channel mode, the target note is instantiated as an audiovisual stimulus note. The system will present a visual graphic (such as a triangle) on the target square while simultaneously announcing the corresponding auditory stimulus sound effect (such as the number "1"). Through the above-mentioned clearly categorized multi-channel stimulus presentation mechanism, the system can accurately input targeted cognitive processing tasks to users and naturally guide users to complete subsequent working memory interaction matching based on different sensory channels by utilizing the clearly defined striking areas at the end of the track.
[0087] Step 104: In response to the user's interactive operation on the target note based on the interactive area, obtain the historical stimulus content of the historical note, and obtain the cognitive matching result for the target note based on the historical stimulus content, cognitive stimulus content and the interactive position of the interactive operation.
[0088] Step 104 is described in detail below.
[0089] In step 104 of some embodiments, the system responds to the user's interactive operation on the target note based on the interactive area, acquires the historical stimulus content of the historical note, and obtains the cognitive matching result for the target note based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interactive operation. This step fully internalizes the logic of the classic N-back (N-order backtracking) working memory task paradigm. Specifically, when the target note reaches the interactive area P_2, the user needs to hold a virtual controller (such as a virtual drumstick) to perform a striking operation. At this time, the system will backtrack and extract the stimulus content of the historical note before a preset backtracking step number (i.e., N value), and compare it with the cognitive stimulus content of the current target note. At the same time, the system detects the specific interaction position where the user's interactive operation occurs (e.g., dividing the interactive area into the left drumhead, the right drumhead, and the middle drumhead).
[0090] For example, if the current stimulus is visual and consistent with the visual stimulus N steps ago, the correct interaction position for the user should be the left drumhead; if it is auditory and consistent with the stimulus N steps ago, the correct interaction position should be the right drumhead; if the stimuli are inconsistent or are ordinary musical notes, the correct interaction position is the middle drumhead. The system combines the comparison results of the stimulus content with the actual interaction position to finally output the cognitive matching result of whether the user's judgment was correct.
[0091] The following section will describe in more detail how to determine the cognitive matching results for the target note.
[0092] Reference Figure 7The current interaction configuration parameters include a preset number of backtracking steps, obtaining the historical stimulus content of historical notes, and obtaining the cognitive matching result for the target note based on the historical stimulus content, cognitive stimulus content, and the interaction position of the interaction operation, including the following steps 701 to 703.
[0093] Step 701: Identify the historical notes in the dynamic note sequence that are separated from the target note by a preset number of backtracking steps, and obtain the historical stimulus content of the historical notes.
[0094] Step 702: Compare the cognitive stimulus content with the historical stimulus content to obtain the content consistency judgment result.
[0095] Step 703: Based on the stimulus type corresponding to the target note, the content consistency judgment result, and the correspondence between the interaction position and the preset striking sub-region, the cognitive matching result is obtained.
[0096] Steps 701 to 703 are described in detail below.
[0097] In step 701 of some embodiments, historical notes that are separated from the target note by a preset number of backtracking steps in the dynamic note sequence are first determined, and the historical stimulus content of the historical notes is obtained. In this step, the "preset number of backtracking steps" corresponds to the core parameter "N" in the classic cognitive psychology N-back (N-order backtracking) working memory task paradigm in the underlying logic of the system, and its value directly represents the memory load level currently assigned to the user.
[0098] When the target note in the dynamic note sequence arrives at the interactive area along the note track and triggers the system's judgment mechanism, the system's data processing module will trace back the data on the generated sequence timeline to accurately locate the historical note that is exactly N steps away from the current target note (i.e., N valid positions apart in the order of appearance). After location, the system extracts the historical stimulus content carried by the historical note. For example, in a single-channel visual paradigm with N=3, the system will extract the specific visual pattern (such as a triangle, a pentagram, or a moon) of the third stimulus note before the current note, thus providing accurate reference benchmark data for subsequent memory comparison.
[0099] In step 702 of some embodiments, a content consistency determination result is obtained by comparing the cognitive stimulus content with the historical stimulus content. After acquiring the cognitive stimulus content presented in real time for the current target note and the extracted historical stimulus content, the system executes a feature matching and comparison algorithm on the corresponding perceptual channel dimension.
[0100] Specifically, if the current task involves visual stimuli, the system compares whether the visual models, geometric patterns, or color features of the two stimuli are completely identical. If it involves auditory stimuli, the system compares whether the specific sound effects or semantic content of the spoken words (such as the numbers "1" or "3") are identical. After the above rigorous channel feature comparison, the system finally outputs a quantitative content consistency judgment result, that is, it explicitly identifies in the system logic whether the two stimuli, separated by N steps, are in a "consistent" or "inconsistent" state.
[0101] In step 703 of some embodiments, a cognitive matching result is obtained based on the stimulus type corresponding to the target note, the content consistency judgment result, and the correspondence between the interaction position and the preset striking sub-region. This step is the core judgment hub for realizing multi-channel sensory collaborative processing and deep mapping of interactive actions in three-dimensional physical space. The system pre-divides the overall interaction area (such as a virtual drumhead) in the virtual reality environment into different preset striking sub-regions with clear functional definitions, including a first striking sub-region (the left drumhead corresponding to the interaction position), a second striking sub-region (the right drumhead corresponding to the interaction position), and a third striking sub-region (the middle drumhead corresponding to the interaction position).
[0102] When generating the final matching result, the system needs to comprehensively evaluate three dimensions: first, identify the stimulus type of the target note (visual or auditory stimulus); second, retrieve the content consistency judgment result; and finally, detect the spatial interaction position where the user's virtual controller actually collided. Following the preset neuroscience interaction rules, the system generates a cognitive matching result for the interaction operation based on strict multi-condition coupling logic, as described below.
[0103] Reference Figure 8 The stimulus types include visual and auditory stimuli. The preset striking sub-regions include a first striking sub-region, a second striking sub-region, and a third striking sub-region. Based on the stimulus type corresponding to the target note, the content consistency judgment result, and the correspondence between the interaction position and the preset striking sub-regions, a cognitive matching result is obtained, including the following steps 801 to 803.
[0104] Step 801: When the stimulus type is visual stimulus, and the content consistency judgment result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the first hitting sub-region, then a cognitive matching result indicating correctness is generated.
[0105] Step 802: When the stimulus type is auditory, and the content consistency judgment result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the second striking sub-region, then a cognitive matching result with correct representation is generated.
[0106] Step 803: When the content consistency determination result indicates that the cognitive stimulus content is inconsistent with the historical stimulus content, or the target note is an ordinary note that does not contain cognitive stimulus content, and the interaction position is located in the third striking sub-region, then a cognitive matching result indicating correctness is generated.
[0107] Steps 801 to 803 are described in detail below.
[0108] In step 801 of some embodiments, when the stimulus type is a visual stimulus, and the content consistency determination result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the first striking sub-region, a cognitive matching result indicating correctness is generated. Specifically, when the system recognizes that the target note currently presented at the trigger position contains visual dimension information input (e.g., the target note is a square block presenting a specific geometric pattern or color), the system will initiate the N-order backtracking (N-back) matching logic of the visual channel. If the underlying system, based on the comparison of historical data, concludes that the current visual stimulus content is completely identical to the historical visual stimulus content separated by N steps (i.e., the content consistency determination result indicates "consistency"), the system requires the user to perform a specific spatial mapping interaction action. According to the preset striking sub-region spatial layout rules, the user must swing the virtual reality controller (such as a handle configured as a virtual drumstick) to accurately strike the first striking sub-region (corresponding to the left drumhead in the virtual environment). Only when the system detects that the controller's interaction position falls precisely within the first striking sub-area will the system determine that the user's cognitive decision and physical execution fully conform to the visual matching rules, thereby generating and outputting a cognitive matching result representing "correctness" to trigger subsequent positive feedback.
[0109] In step 802 of some embodiments, when the stimulus type is auditory, and the content consistency judgment result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the second striking sub-region, a cognitive matching result representing correctness is generated. In a multi-channel sensory coordination interaction mode, the target note may not be prompted by visual patterns, but by sound through the audio channel. When the system recognizes that the current target note carries a specific sound effect or voice broadcast (e.g., an auditory stimulus note broadcasting a specific number or tone), the system activates the working memory matching logic of the auditory channel. If the comparison result shows that the current auditory stimulus is the same as the historical auditory stimulus content N steps away, after the user's brain completes the auditory working memory retrieval, it needs to convert the "consistency" judgment result into a physical spatial action on the other side. At this time, the correct interaction rule set by the system is to strike the second striking sub-region (corresponding to the right drumhead in the virtual environment). When the user's actual interaction position is captured by the system sensor and confirmed to be located in the second striking sub-region, the system will also generate a cognitive matching result representing "correctness". This spatial decoupling design, which maps visual matching to the left and auditory matching to the right, effectively forces users to separate information from multiple sensory channels and output directional movement within a short period of time.
[0110] In step 803 of some embodiments, when the content consistency determination result indicates that the cognitive stimulus content is inconsistent with the historical stimulus content, or the target note is an ordinary note that does not contain cognitive stimulus content, and the interaction position is located in the third striking sub-region, a cognitive matching result indicating correctness is generated. This step covers all "non-target" or "maintaining the beat baseline" arrangements in cognitive interaction except for "channel target matching". First, if the current stimulus content is different from the historical stimulus content in the N-order retrospective comparison of vision or hearing (i.e., there is a distractor or a new memory item, and the content consistency determination result is "inconsistent"), the user needs to perform cognitive inhibition and restrain the conditioned reflex impulse to strike to the left or right. Second, ordinary notes (e.g., circular squares without any specific patterns or sound effects) that are interspersed in a large number of dynamic note sequences to maintain the continuity of the musical beat do not require the user to perform cognitive load comparison. In the face of the above two situations, the system's preset interaction rules point to a default baseline operation area, namely the third striking sub-region (corresponding to the middle drum surface in the virtual environment). When the system detects that the user has correctly responded by hitting the third strike sub-region when these "inconsistent" stimulus notes or "no judgment required" ordinary notes are given, it generates a cognitive matching result that represents "correctness".
[0111] Reference Figure 9 This is a schematic diagram of a virtual interaction scenario provided in this application embodiment, illustrating a user's interactive operation under a visual channel working memory task. Figure 9As shown, in the presented 3D immersive virtual environment, when specific visual patterns (such as...) are included... Figure 9 When the square stimulus note (the apple-shaped pattern shown) moves along the note track to the user's interactive area, the system has already completed the comparison between the current visual stimulus content and the historical visual stimulus content at a preset backtracking step (N steps) in the background. In this scenario, since the visual features of the two are completely consistent (i.e., they meet the visual N-back matching rule), after the user's brain completes the visual memory retrieval and consistency judgment, it controls the virtual controller (i.e., the virtual drumstick) held in the left hand to swing downwards, accurately striking the corresponding first striking sub-area (i.e., the left drumhead) within the interactive area. Figure 9 This illustrates how the system transforms the abstract visual cognitive matching "consistency" decision into a specific interactive action of directional striking in the physical space to the left.
[0112] Reference Figure 10 This is a virtual interactive scenario provided in this application embodiment, where a user performs interactive operations under an auditory channel working memory task. For example... Figure 10 As shown, in a virtual scene, a target note moving towards the user along a note track triggers specific auditory stimuli (such as playing specific speech or sound effects). When the target note reaches the trigger position and subsequent interaction area, if the system's underlying logic determines that the currently played auditory stimulus is exactly the same as the historical auditory stimulus N steps away (i.e., it conforms to the auditory N-back matching rule), the user needs to convert the auditory memory "consistency" judgment result into spatial motion output on the other side. At this time, the user controls the virtual controller held in their right (or left) hand to accurately strike the corresponding second striking sub-area (i.e., the right drumhead) within the interaction area. Figure 10 This reflects the system's spatial decoupling design for multi-channel sensory information, meaning that the correct response in the auditory matching dimension is strictly limited to the physical interaction interface on the right.
[0113] Reference Figure 11 This is a virtual scenario provided in this application embodiment where a user performs interactive operations when faced with interfering stimuli that do not conform to memory matching rules. For example... Figure 11 As shown, during the movement of the dynamic note sequence, when the presented square stimulus notes (such as...) Figure 11When the user enters the interaction phase with the square containing the banana pattern, if, after N-order backtesting in their working memory, the user finds that the current visual or auditory stimulus is different from the historical content N steps ago (i.e., it does not conform to either the visual or auditory N-back rule, falling under the category of "inconsistent" interference in content consistency), the user must actively suppress their cognitive responses and restrain the conditioned reflex of striking to the left or right. At this point, the correct interaction rule requires the user to swing the virtual controller to strike the third striking sub-area (i.e., the central drumhead) located in the very center of the interaction area. Figure 11 The study demonstrates the physical action feedback path after the system introduces a cognitive inhibition mechanism, emphasizing the rules for directional spatial handling of "mismatched" information.
[0114] Reference Figure 12 This is a virtual interaction scenario provided in this application embodiment, where a user performs interactive operations when faced with ordinary musical notes that do not require cognitive processing. For example... Figure 12 As shown, on the immersive note track, to maintain the continuity of the musical rhythm and the flow experience of the game, the system intersperses circular ordinary notes (i.e., non-stimulus modules) between the intervals of the stimulating notes. When these circular ordinary notes, which do not contain any specific cognitive stimulus content (no specific memory pattern or auditory effect), move to the user's interaction area, the user does not need to invoke the working memory network for N-back comparison judgment, but directly treats the note as a simple rhythm beat point for mechanical interaction. At this time, the illustration clearly shows that the user controls the virtual controller to uniformly strike the third striking sub-area (i.e., the middle drumhead) to complete the correct response to the ordinary note. Figure 12 It supplements the cognitive training system with a fallback interaction logic in the "non-stimulation" state, ensuring that the user's physical striking actions can be seamlessly connected with the rhythm of the music, regardless of whether there is cognitive load or not.
[0115] Through steps 801 to 803 above, a set of logically rigorous and action-oriented multi-channel cognitive interaction response rules is constructed in a three-dimensional virtual space. This application's solution deeply maps and binds the classic cognitive psychology N-back experimental paradigm with the physical region-based striking action in virtual reality. It precisely externalizes the complex executive function processing processes that originally occurred implicitly within the brain, such as "visual memory retrieval," "auditory memory retrieval," and "cognitive inhibition decision-making," into explicit spatial limb movements of "hit left," "hit right," and "hit center." This design not only greatly enriches the interactive dimensions and playability of rhythm games, but more importantly, through independent spatial mapping of multi-sensory channels and anti-interference settings for default baseline areas, it forces users to complete high-intensity information diversion and directional movement control within millisecond-level beat intervals, thereby significantly improving the systematic intervention effect on working memory, attention allocation, and cognitive inhibition abilities.
[0116] Through steps 701 to 703 above, the abstract N-back working memory experimental paradigm is successfully transformed into a three-dimensional physical interaction logic that can be accurately read and judged by a computer. This application's solution not only achieves rigorous retrospective comparison and alignment of time-series stimuli across multiple channels, including visual and auditory, but also innovatively establishes a strong binding mapping between the brain's cognitive decision-making state (consistent or inconsistent stimulus content) and the body's differentiated action execution in three-dimensional virtual space (hitting different areas on the left, middle, or right). This design mechanism completely eliminates mechanical, blind hitting, forcing users to first complete deep information retrieval and memory matching within a very short musical beat response time, and then transform higher-order cognitive decisions into precise directional limb motor control. This maximizes the mobilization of the user's working memory network and motor execution function, providing a solid technical guarantee for significantly improving the effectiveness and scientific rigor of cognitive intervention.
[0117] Step 105: Generate real-time feedback information based on the cognitive matching results and the timing of the interactive operation.
[0118] Step 105 is described in detail below.
[0119] In step 105 of some embodiments, the system generates real-time feedback information based on the cognitive matching result and the timing of the interaction. This timing refers to the time difference between the actual trigger time of the collision between the user's virtual controller and the interaction area, and the theoretically optimal time for the target note to reach the center of the interaction area. Assuming the cognitive matching result is correct, the system further evaluates the performance evaluation time window into which this time difference falls. According to preset rules, if the hit timing is within a very short window of ±100ms, the system generates "perfect" feedback; within ±100-250ms, "very good" feedback; and within ±250-500ms, "good" feedback. Conversely, if the hit is incorrect or the time exceeds 500ms, "incorrect" or "missed" feedback is generated. These millisecond-level evaluation results are displayed to the user as real-time feedback information through visual UI effects or auditory sound effects, thus completing the closed loop of a single interaction.
[0120] The following section will describe in more detail how to generate real-time feedback information.
[0121] Reference Figure 13 Based on the cognitive matching results and the timing of the interactive operation, real-time feedback information is generated, including the following steps 1301 to 1302.
[0122] Step 1301: When the cognitive matching result is correctly represented, obtain the actual trigger time of the interactive operation, and obtain the time difference based on the difference between the actual trigger time and the theoretical arrival time of the target note to the center of the interactive area.
[0123] Step 1302: Match the time difference with multiple preset performance evaluation time windows to obtain a matching time window, and generate real-time feedback information on the evaluation level of the corresponding matching time window based on the matching time window.
[0124] Steps 1301 to 1302 are described in detail below.
[0125] In step 1301 of some embodiments, when the cognitive matching result is correctly represented, the actual trigger time of the interactive operation is obtained, and a time difference is obtained based on the difference between the actual trigger time and the theoretical arrival time of the target note reaching the center of the interactive area. Specifically, after the system completes the N-level backtracking (N-back) logical judgment of the multi-sensory channels, once it confirms that the user's memory comparison is not only correct, but also that the spatial position of the strike fully conforms to the preset rules (i.e., the cognitive matching result is "correct"), it will immediately start the accuracy assessment based on the time dimension. At this time, the underlying interactive control module of the system will accurately record the absolute timestamp of the physical collision between the virtual controller (such as a virtual drumstick) and the virtual interactive area (such as a drumhead), and use this as the "actual trigger time".
[0126] Simultaneously, based on the score data and musical rhythm, the system extracts a preset timestamp indicating when the target note falls at a uniform speed and perfectly aligns with the geometric center of the interactive area, using this as the "theoretical arrival time." The system then calculates the difference between these two timestamps, resulting in a "time difference" accurate to the millisecond level. This difference directly and objectively quantifies the accuracy of the user's beat-following during motion output, reflecting the degree of harmony between the user's physical movements and the musical rhythm.
[0127] In step 1302 of some embodiments, the time difference is matched with multiple preset performance evaluation time windows to obtain a matching time window, and real-time feedback information of the evaluation level corresponding to the matching time window is generated based on the matching time window. In this step, the "performance evaluation time window" is a series of gradient time tolerance intervals preset by the system to measure the user's agility and rhythm accuracy. The system compares the calculated millisecond-level time difference value with these intervals one by one to determine the specific range it falls into, and then locks the corresponding "matching time window".
[0128] In a specific setting example, if the time difference falls within a very short tolerance range of ±100 ms, the system determines that it matches the highest-level evaluation window and generates and displays a "Perfect" evaluation level prompt; if the time difference falls within the range of ±100 ms to 250 ms, the system generates a "Very Good" evaluation level prompt; and if the time difference falls within the range of ±250 ms to 500 ms, the system generates a "Good" evaluation level prompt. These evaluation levels generated based on matching time windows are ultimately transformed into real-time feedback information such as visual UI effects or auditory sound effects, presented to the user within millisecond-level latency, thus completing the closed loop of a single interaction.
[0129] Through steps 1301 to 1302 above, while ensuring the effectiveness of N-back working memory cognitive intervention, a fine-grained movement assessment mechanism based on the time dimension is introduced. This application's solution not only performs a simple "right or wrong" binary judgment of the user's cognitive assessment results, but also further quantifies and provides feedback on the "rhythmic accuracy" of the user's movements in multiple gradients. This design, which transforms rigid cognitive test results into gamified, hierarchical positive incentives (such as millisecond-level evaluations like "perfect" and "very good"), can greatly stimulate the user's desire for challenge and sense of accomplishment, prompting the user to continuously adjust their motor neural feedback with each strike to approach the optimal beat point. This not only significantly enhances the immersion and flow experience of the training process, but also effectively couples auditory rhythm perception, brain working memory retrieval, and fine motor control of the limbs, thereby comprehensively improving the long-term compliance and actual intervention quality of cognitive and motor dual rehabilitation training.
[0130] Furthermore, the cognitive interaction method based on virtual reality rhythm games provided in this application also includes an update feedback mechanism. (See reference...) Figure 14 The cognitive interaction method based on virtual reality rhythm games provided in this application also includes the following steps 1401 to 1403.
[0131] Step 1401: Within the current interaction task cycle, determine the accuracy of the cognitive matching results corresponding to multiple interaction operations.
[0132] Step 1402: When the accuracy reaches the first preset threshold, increase the music difficulty level parameter.
[0133] Step 1403: When the music difficulty level parameter reaches the highest level and the accuracy reaches the second preset threshold, maintain the music difficulty level and increase the preset backtracking steps.
[0134] Steps 1401 to 1403 are described in detail below.
[0135] In step 1401 of some embodiments, the accuracy of cognitive matching results corresponding to multiple interactive operations is determined within the current interactive task cycle. Specifically, the "current interactive task cycle" typically refers to the user completing a complete training track or experiencing a continuous training level phase of a preset duration. During the operation of this cycle, the system's internal data module monitors and collects in real time the user's striking response data for each falling stimulus note. Each response generates a clear "correct" or "incorrect" quantitative cognitive matching result based on the system's underlying working memory N-order backtracking paradigm rules. When the task cycle (such as a song) ends, the system summarizes these discrete single matching judgment results, calculates the percentage of correct striking matches out of the total number of stimuli to be matched in the cycle, and thus obtains an objective accuracy index. This accuracy not only provides long-term feedback to the user as part of the performance summary report, but also constitutes the core objective basis for the system to subsequently trigger an adaptive difficulty adjustment mechanism.
[0136] In step 1402 of some embodiments, when the accuracy reaches a first preset threshold, the music difficulty level parameter is increased. The system internally sets a first preset threshold to measure whether the user's current ability has fully adapted to the current training load. When the user's accuracy within the interactive task cycle reaches or exceeds this threshold, it indicates that the user has achieved a high level of proficiency under the current cognitive and operational rhythm. The system then triggers the first dimension adjustment in the two-dimensional adaptive adjustment mechanism, i.e., increasing the music difficulty level parameter. In specific implementations, the system divides songs into different progressive difficulty levels according to the music's tempo (BPM, i.e., the number of drumbeats per minute). For example, the system can divide them into three levels: Level 1 (corresponding to 80–95 BPM), Level 2 (corresponding to 95–110 BPM), and Level 3 (corresponding to 110–120 BPM). Increasing this parameter means that, while maintaining the preset backtracking steps (N value) unchanged, the system will automatically call up sheet music data with a faster tempo and shorter stimulus intervals for the user in the next training task cycle. This adjustment method, which increases the density of instantaneous information processing by accelerating the pace of physical operations, can effectively prevent users from experiencing a sense of experiential effect or boredom due to overly simple tasks, thus keeping them in a highly focused training state.
[0137] In step 1403 of some embodiments, when the music difficulty level parameter reaches the highest level and the accuracy reaches the second preset threshold, the music difficulty level is maintained and the preset backtracking steps are increased. As the user's cognitive and physical coordination training deepens, when the music difficulty level parameter is gradually increased to the preset highest level (e.g., level 3 corresponding to 110–120 BPM), the system determines that the user's reaction speed and beat-following ability at the current backtracking depth have reached a plateau. At this time, if the user's performance at the highest tempo level remains stable and the calculated accuracy meets the set second preset threshold, the system will trigger the second-dimensional leap adjustment of the adaptive mechanism. To prevent excessive increase in physical beat speed from causing the task to exceed the limits of normal human motor neural response, the system will lock and maintain the current highest music difficulty level (i.e., maintain a fixed highest BPM) and instead increase the preset backtracking steps (i.e., the core memory load parameter N value). For example, the system will jump the current task directly from level 1 backtracking (memorizing the previous stimulus) to level 2 backtracking (memorizing the stimulus before that), requiring the user to temporarily store and compare the historical stimulus content in their brain. This adjustment strategy essentially completes a smooth transition and upgrade from external "physical speed pressure" to internal "core cognitive capacity pressure".
[0138] Through steps 1401 to 1403 above, a sophisticated two-dimensional adaptive difficulty adjustment mechanism conforming to the principles of cognitive psychology was constructed. This application's solution completely abandons the rigid, linear difficulty progression model of traditional cognitive interventions, instead utilizing the user's real-time interaction accuracy as dynamic feedback to intelligently match the two core dimensions of "music tempo" and "working memory load capacity." This personalized adaptive adjustment based on the user's actual performance not only ensures that the difficulty of each interactive task accurately falls within the user's zone of proximal development, effectively inducing and maintaining the user's immersive flow state, but also fundamentally eliminates the frustration caused by excessive difficulty or the training fatigue caused by insufficient difficulty, thereby greatly enhancing the user's intrinsic motivation and long-term adherence to rehabilitation and cognitive training.
[0139] In summary, steps 101 to 105 have deeply and organically integrated the classic N-back working memory training paradigm with a rhythmic action game in a virtual reality environment. On the one hand, by utilizing the physical rhythm (BPM) of music and the spatial movement trajectory of notes, the user's time attention allocation is naturally guided, greatly enhancing the immersion and fun of the task and overcoming the shortcomings of traditional training, which is often tedious and leads to low compliance. On the other hand, by linking the cognitive memory comparison of the brain with the precise spatial segmentation of the body and supplementing it with millisecond-level real-time feedback evaluation, a highly close coupling between the brain's executive function processing and physical movement is achieved. This ensures the scientific effectiveness of cognitive training while significantly improving the accuracy of intervention and the long-term benefits for users.
[0140] Reference Figure 15 This is a schematic diagram illustrating the overall execution flow of a cognitive interaction method based on a virtual reality rhythm game provided in an embodiment of this application. Figure 15 As shown, the system is initially in the startup state. After entering the cognitive interaction system, the user inputs their unique user identifier (i.e., ...) through the user interface (UI) module. Figure 15 The input ID is then entered into the system's data module. The system's data module responds to this input instruction by performing a matching query on the underlying database to determine if a historical record exists (i.e., the user ID entered). Figure 15 (Regarding the "whether there is a record" option). If the system does not find the historical training record corresponding to the ID, it determines that the user is new and creates a brand new data profile for them (i.e., "create a new record"). If a matching historical record is found, the system directly retrieves and obtains the historical interaction performance and current cognitive ability assessment status corresponding to the ID (i.e., "obtain the record").
[0141] After completing the initial matching and acquisition of user data, the system enters the task preparation and core execution phase. Based on the acquired user's current data records (i.e., extracting the user's current interaction configuration parameters, such as the desired difficulty level), the system automatically loads a music library with a matching difficulty level and its corresponding score data (i.e.,...). Figure 15 The system loads the matching music library. Once the user selects a specific training track from the matching music library (i.e., "select track"), the system officially starts the training level. In the core step of "completing the training level," the system plays the target music data, controls the dynamic note sequence to move towards the interactive area, and acquires the user's tapping interaction based on the N-back paradigm in real time, while providing millisecond-level real-time feedback. This level execution process fully covers the aforementioned complete set of technical solutions for dynamic note generation, multi-channel stimulus presentation, and cognitive matching judgment.
[0142] When the set training track finishes playing and the system stops generating new dynamic notes, the training level is considered complete, and the system enters the data settlement and closed-loop update phase. The system's data module automatically calculates and obtains the user's overall performance results within the current interaction task cycle (i.e., "Obtain Settlement Data" in the diagram). This settlement data specifically includes objective indicators such as the user's hitting statistics, the accuracy of cognitive matching results, and the overall score. Finally, the system synchronously overwrites and saves this latest settlement data to the user ID's dedicated database (i.e., "Update Record"), thus completing the dynamic update of the user's current interaction configuration parameters. The updated data will serve as the direct basis for the next training session to call target music data and trigger adaptive difficulty adjustments (such as increasing the music difficulty level or preset backtracking steps). At this point, a complete cognitive interaction closed-loop process officially ends.
[0143] Reference Figure 16 This is a schematic diagram illustrating the overall architecture and module data flow of a cognitive interaction system based on a virtual reality rhythm game, provided in an embodiment of this application. Figure 16 As shown, on the system's data preparation and input side ( Figure 16 (As shown on the left), the system constructs independent "music library" and "score database". The "score generation module", as the core of offline data processing, is responsible for retrieving raw audio files from the "music library", extracting their periodic beat and tempo features, and generating score data adapted to the cognitive training task according to preset repetition patterns. This data is then centrally stored in the "score database". This architectural design effectively decouples the underlying task generation from the front-end real-time interactive computation, providing solid data support for a high frame rate immersive experience.
[0144] Within the core "virtual reality cognitive training system" ( Figure 16 (See the dashed box). It integrates eight key functional modules that support the entire interactive loop. Among them, the "Audio Playback Module" directly reads the target music data from the music library for playback and is responsible for synchronously and independently outputting cognitive stimulus audio to the auditory channel; the "Note Generation Module" reads the matching score data from the score database and dynamically instantiates and generates interactive note sequences according to strict time nodes; the "Virtual Environment Construction Module" is responsible for building a three-dimensional scene with spatial depth, including note tracks, user standing areas, and multi-hit sub-areas; and the "Data Module" runs throughout the training process, responsible for calling the user's current interaction configuration parameters (such as preset backtracking steps), managing records, and storing and updating settlement data after the task is completed.
[0145] At the level of physical interaction between the system and external hardware, this architecture Figure 16The bidirectional data communication relationship between the core system and the "virtual reality kit" (including the head-mounted display and spatial positioning controller) on the right side of the diagram is clearly defined. Specifically, the "UI module" is responsible for building the user interaction interface for the entire process; the "interaction control module" establishes a physical connection and collects real-time data on the user's head posture and the spatial position and swing trajectory of the controllers from the virtual reality kit; the "real-time feedback module" completes millisecond-level matching and judgment based on the collected interaction position and timing, combined with working memory paradigm rules, and generates a performance evaluation level; finally, the "rendering module" serves as the core hub of the system's visual output, rendering the constructed 3D scene, dynamic musical notes, UI interface, and real-time feedback effects across platforms and seamlessly projecting them into the user's virtual reality kit, thus constructing a highly closed-loop technical architecture of "data loading - multi-sensory presentation - physical action interaction - cognitive result feedback".
[0146] This application also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in a memory, and the processor executes the at least one program to implement the cognitive interaction method based on virtual reality rhythm games described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0147] Please see Figure 17 , Figure 17 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1702 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1702 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1702 and is called and executed by the processor 1701 to execute the cognitive interaction method based on virtual reality rhythm games in the embodiments of this application. The input / output interface 1703 is used to implement information input and output; The communication interface 1704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1705 transmits information between various components of the device (e.g., processor 1701, memory 1702, input / output interface 1703, and communication interface 1704); The processor 1701, memory 1702, input / output interface 1703 and communication interface 1704 are connected to each other within the device via bus 1705.
[0148] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described cognitive interaction method based on a virtual reality rhythm game.
[0149] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0150] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0151] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0153] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0154] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0155] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.
[0157] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0160] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A cognitive interaction method based on virtual reality rhythm games, characterized in that, The method includes: Load the target music data and the corresponding score data according to the user's current interaction configuration parameters; When playing the target music data, a dynamic note sequence is controlled to move along the note track of the virtual environment to the interactive area according to a preset rhythm. The dynamic note sequence is generated based on the score data. When at least one target note in the dynamic note sequence moves to the trigger position, the cognitive stimulus content corresponding to the target note is presented. In response to a user's interactive operation on the target note based on the interactive area, the historical stimulus content of the historical note is obtained, and based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interactive operation, a cognitive matching result for the target note is obtained. Based on the cognitive matching results and the timing of the interactive operation, real-time feedback information is generated.
2. The cognitive interaction method based on virtual reality rhythm games according to claim 1, characterized in that, The current interaction configuration parameters include a preset backtracking step count. The process of acquiring historical stimulus content for historical notes and obtaining a cognitive matching result for the target note based on the historical stimulus content, the cognitive stimulus content, and the interaction position of the interaction operation includes: Identify the historical notes in the dynamic note sequence that are separated from the target note by a preset number of backtracking steps, and obtain the historical stimulus content of the historical notes; The content consistency determination result is obtained by comparing the cognitive stimulus content with the historical stimulus content; The cognitive matching result is obtained based on the stimulus type corresponding to the target note, the content consistency judgment result, and the correspondence between the interaction position and the preset hitting sub-region.
3. The cognitive interaction method based on virtual reality rhythm games according to claim 2, characterized in that, The stimulus types include visual and auditory stimuli, and the preset striking sub-regions include a first striking sub-region, a second striking sub-region, and a third striking sub-region. The cognitive matching result, obtained based on the stimulus type corresponding to the target note, the content consistency determination result, and the correspondence between the interaction position and the preset striking sub-regions, includes: When the stimulus type is the visual stimulus, and the content consistency determination result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the first hitting sub-region, then the cognitive matching result indicating correctness is generated. When the stimulus type is the auditory stimulus, and the content consistency determination result indicates that the cognitive stimulus content is consistent with the historical stimulus content, and the interaction position is located in the second hitting sub-region, then the cognitive matching result indicating correctness is generated. When the content consistency determination result indicates that the cognitive stimulus content is inconsistent with the historical stimulus content, or when the target note is an ordinary note that does not contain cognitive stimulus content, and the interaction position is located in the third striking sub-region, then the cognitive matching result that indicates correctness is generated.
4. The cognitive interaction method based on virtual reality rhythm games according to claim 2, characterized in that, The steps for generating the dynamic note sequence include: Based on the time node information and note type information in the spectrum data, an initial note sequence is generated; In the initial note sequence, an interference position is determined that is a target number of steps away from the stimulus note containing the target stimulus content, and an interference stimulus note is generated at the interference position according to a preset probability. The dynamic note sequence is obtained based on the updated initial note sequence. Wherein, the cognitive stimulus content of the interference stimulus note is the same as the target stimulus content, and the target number of steps is obtained based on the preset backtracking number of steps.
5. The cognitive interaction method based on virtual reality rhythm games according to claim 1, characterized in that, The generation of real-time feedback information based on the cognitive matching result and the timing of the interaction operation includes: When the cognitive matching result is correctly represented, the actual trigger time of the interactive operation is obtained, and the time difference is obtained based on the difference between the actual trigger time and the theoretical arrival time of the target note to the center of the interactive area. The time difference is matched with multiple preset performance evaluation time windows to obtain a matching time window, and real-time feedback information corresponding to the evaluation level of the matching time window is generated based on the matching time window.
6. The cognitive interaction method based on virtual reality rhythm games according to claim 1, characterized in that, The current interaction configuration parameters include the music difficulty level and the preset number of backtracking steps, and the method further includes: Within the current interaction task cycle, determine the accuracy of the cognitive matching results corresponding to multiple interaction operations; When the accuracy rate reaches a first preset threshold, the music difficulty level is increased; When the music difficulty level reaches the highest level and the accuracy reaches the second preset threshold, the music difficulty level is maintained and the preset backtracking steps are increased.
7. The cognitive interaction method based on virtual reality rhythm games according to claim 1, characterized in that, Before loading the target music data based on the user's current interaction configuration parameters, the method further includes: Obtain the original audio file, and obtain the periodic beat characteristics and beat speed characteristics of the original audio file; Based on the aforementioned periodic beat characteristics, the timeline of the original audio file is divided into multiple beat cycles; Within each beat cycle, based on the beat speed characteristics, empty modules, stimulus modules, and non-stimulation modules are arranged according to a preset repetition pattern to generate the spectral data.
8. The cognitive interaction method based on virtual reality rhythm games according to claim 7, characterized in that, The arrangement of empty modules, stimulation modules, and non-stimulation modules based on the beat speed characteristics and according to a preset repetition pattern includes: Establish a mapping model between the beat velocity characteristics and the stimulus interval time; The time interval between the stimulation modules is determined based on the mapping model, and the empty module or the non-stimulation module is inserted within the time interval.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the cognitive interaction method based on virtual reality rhythm games as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the cognitive interaction method based on virtual reality rhythm game as described in any one of claims 1 to 8.