Multi-modal interaction control system and method based on virtual reality and sensing technology

Through multimodal data acquisition and fusion technology, combined with nostalgic scene generation and dynamic difficulty control, the problem of unfriendly interactive experience for elderly users was solved, a personalized VR rehabilitation training platform was realized, and the cognitive and physical rehabilitation effects of the elderly were improved.

CN120686973APending Publication Date: 2025-09-23王海梁
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760404.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing VR rehabilitation system is not user-friendly for elderly users, lacks multimodal data fusion, and the preset scenes cannot adapt to individual memory characteristics, resulting in rigid interaction methods and reduced personalized adaptation effects.

Method used

A multimodal data acquisition module is used, including eye movement, gesture, voice and skeletal movement data collection. User behavior feature vectors are generated through data fusion. Combined with the nostalgic scene generation module, the task difficulty is dynamically adjusted, and visual, auditory and movement feedback is provided to construct personalized nostalgic scenes.

Benefits of technology

It improves the interactive experience of elderly users, enhances their cognitive and physical rehabilitation effects through multimodal data fusion and personalized scene adaptation, and provides a personalized, efficient and easy-to-operate training platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686973A_ABST
    Figure CN120686973A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal interaction control system and method based on a virtual reality and sensing technology, and belongs to the field of multi-modal interaction control, and the method comprises the steps: obtaining the eye movement, gesture, voice and skeleton motion data of a user in real time through a multi-modal data collection module, and generating a user behavior feature vector through a data fusion algorithm; the nostalgic scene generation module dynamically generates personalized immersive nostalgic scenes, such as families, convenience stores and streets, based on the user feature vectors in combination with a 70-year article library and a three-dimensional reconstruction algorithm. The dynamic difficulty control module dynamically adjusts task difficulty according to user performance, the multi-mode feedback output module provides visual, auditory and action feedback, and user experience is enhanced. The system aims at improving the memory, the decision-making ability and the body coordination ability of the elderly through the reminiscence therapy and cognitive training, and a scientific and personalized rehabilitation scheme is provided for the elderly patients with mild cognitive impairment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multimodal interactive control, and in particular relates to a multimodal interactive control system and method based on virtual reality and sensing technology. Background Art

[0002] In recent years, the application of VR and sensor technologies in healthcare has become increasingly widespread. VR technology provides users with a highly simulated environment through immersive experiences, while sensor technologies can capture users' physiological and behavioral data in real time. The combination of these two technologies offers new possibilities for the management of mild cognitive impairment.

[0003] Existing VR rehabilitation systems mostly rely on a single sensor (such as a handle) to achieve interaction, which has the following limitations: (1) Hand-based operation relies on high-precision motion capture, resulting in a rigid interaction method that is not user-friendly for elderly users. (2) Only basic behavioral data (such as task completion time) is collected, the data dimension is single, and there is a lack of multimodal data fusion (physiological signals, operation paths). (3) Preset scenes cannot adapt to individual memory characteristics, resulting in static scene generation and reduced personalized adaptation effects. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a multimodal interactive control system based on virtual reality and sensing technology, comprising:

[0005] Multimodal data acquisition module, used to collect user's eye movement, gesture, voice and skeletal movement data to generate multimodal data;

[0006] The data fusion module is used to fuse the collected multimodal data and generate user behavior feature vectors;

[0007] A nostalgic scene generation module, configured to generate a personalized nostalgic scene based on the user behavior feature vector;

[0008] A dynamic difficulty control module, configured to dynamically adjust the difficulty of a task based on the user's performance in the personalized nostalgic scene;

[0009] A multimodal feedback output module, which provides visual, auditory, and motion feedback to users based on task difficulty;

[0010] Summary report module, used to generate user task performance summary report;

[0011] The data collection and analysis module is used to collect and store user operation data and task performance data.

[0012] Preferably, the multimodal data acquisition module includes an eye tracking unit, a gesture recognition unit, a speech recognition unit and a skeleton tracking unit;

[0013] The eye tracking unit is used to capture the user's line of sight focus coordinates;

[0014] The gesture recognition unit is used to capture the user's hand movements;

[0015] The speech recognition unit is used to receive and analyze the user's voice instructions;

[0016] The skeleton tracking unit is used to capture the coordinates of the user's entire body skeleton joints in real time.

[0017] Preferably, the data fusion module performs time stamp synchronization and spatial coordinate alignment on the eye movement, gesture, voice and skeletal movement data through a Kalman filter;

[0018] The data fusion module generates a user behavior feature vector based on the synchronized data.

[0019] Preferably, the nostalgic scene generation module includes a 1970s item library unit, a dynamic scene rendering engine unit, and a scene difficulty and task design unit;

[0020] The 1970s item library unit stores three-dimensional models and animation data of items in life scenes of the 1970s;

[0021] The dynamic scene rendering engine unit generates a personalized scene in real time based on the old photos provided by the user and the selection made during the virtual environment setting process;

[0022] The scenario difficulty and task design unit sets three difficulty modes according to the scenario, and the task design stimulates autobiographical memory through nostalgia therapy.

[0023] Preferably, the dynamic difficulty control module includes a user-selected difficulty unit and a system-customized difficulty unit;

[0024] The user-selectable difficulty unit allows the user to select from three difficulty levels.

[0025] The system personalized difficulty unit evaluates the user's exercise level based on the user's exercise performance and recommends a corresponding learning mode to the user.

[0026] Preferably, the multimodal feedback output module includes a visual feedback unit, an auditory feedback unit and a motion feedback unit;

[0027] The visual feedback unit uses Unity Shader programming to achieve target object highlighting, uses the Unity UI system to update task progress, countdown and prompt information in real time, and dynamically adjusts the user avatar skeleton color according to the action matching degree;

[0028] The auditory feedback unit plays pre-recorded voice commands through the Unity Audio Source class and plays scene sound effects in conjunction with animation events;

[0029] The motion feedback unit realizes the magnetic attraction effect of objects through the PhysX engine and guides the user to adjust the posture based on the motion capture data.

[0030] Preferably, the summary report module provides a task performance summary report to help users review their performance in cognitive and physical exercise games and track training progress; the report content includes the user's task completion status and performance level of task goal achievement.

[0031] On the other hand, the present invention also provides a multimodal interactive control method based on virtual reality and sensing technology, comprising:

[0032] Collect user's eye movement, gesture, voice and skeletal movement data;

[0033] Fuse the collected multimodal data to generate user behavior feature vectors;

[0034] Generate personalized nostalgic scenes based on user behavior feature vectors;

[0035] Dynamically adjust task difficulty based on user performance in nostalgic scenarios;

[0036] Provide users with visual, auditory, and motor feedback based on task difficulty;

[0037] Collect and store user operation data and task performance data.

[0038] On the other hand, the present invention further provides an electronic device, comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the method is implemented when the processor executes the computing program.

[0039] On the other hand, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the method when executed by a processor.

[0040] Compared with the prior art, the present invention has the following advantages and technical effects:

[0041] (1) Through age-friendly interactive design, based on eye tracking, gesture control, voice commands and body movements, the mixed input and output logic reduces operational complexity, improves user experience, and provides a personalized, efficient and easy-to-use training platform for the elderly;

[0042] (2) an algorithmic framework integrating cognitive training (spatial memory, executive function) and physical rehabilitation (coordination training);

[0043] (3) Based on real-time user data, evaluate their performance level and dynamically generate tasks with appropriate difficulty;

[0044] (4) Combining the historical object library with the 3D reconstruction algorithm, a personalized immersive environment of scenes such as homes, streets, and parks in the 1970s is constructed. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0046] Figure 1 This is a structural diagram of the home scene operation mechanism of the elderly cognitive and physical rehabilitation system based on virtual reality and sensing technology and its multimodal interactive control method according to an embodiment of the present invention;

[0047] Figure 2 This is a structural diagram of the convenience store scenario operation mechanism in the elderly cognitive and physical rehabilitation system based on virtual reality and sensing technology and its multimodal interactive control method in an embodiment of the present invention.

[0048] Figure 3 This is a structural schematic diagram of the street scene operation mechanism in the elderly cognitive and physical rehabilitation system based on virtual reality and sensing technology and its multimodal interactive control method in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0050] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0051] Example 1

[0052] like Figure 1-2 As shown, this embodiment provides a multimodal interactive control system based on virtual reality and sensing technology, including:

[0053] This embodiment aims to provide a cognitive and physical rehabilitation system for the elderly based on virtual reality and sensing technology, along with its multimodal interactive control method. This system addresses technical issues with existing self-management systems, including unfriendly user interactions for the elderly, insufficient personalized adaptation, insufficient functionality, and a lack of scientific psychotherapy. By designing a set of virtual scenarios and tasks tailored to the specific needs of the elderly and incorporating nostalgic therapy, the system provides cognitive rehabilitation training and personalized physical exercise guidance based on nostalgic 1970s home, convenience store, and street scenes. The system aims to improve cognitive and motor abilities, including memory, decision-making, eye-hand coordination, and executive function, among others.

[0054] This embodiment's technical solution, based on virtual reality and sensor technologies, builds a comprehensive cognitive and motor training platform. This technology utilizes the Unity platform (a real-time 3D interactive content creation and operation platform), the Neon XRCore Package (an eye-tracking plugin), and the Azure Kinect SDK (a motion capture camera). The system includes five modules: a multimodal data acquisition module, a nostalgic scene generation module, a dynamic difficulty control module, a multimodal feedback output module, a summary reporting module, and a data acquisition and analysis module.

[0055] 1) Multimodal data acquisition module, which includes a hardware layer unit and a data fusion algorithm unit.

[0056] 1.1) Hardware layer unit:

[0057] 1.1.1) Skeletal tracking unit: Based on the Azure Kinect sensor, it uses a depth camera (TOF technology) and an RGB camera to capture the coordinates of the user's entire body skeleton joints in real time (accuracy of ±2cm).

[0058] 1.1.2) Eye Tracking Unit: Neon XR eye tracking is integrated into the Meta Quest 3 system for gaze estimation and real-time streaming. It captures the user's gaze focus coordinates at a 120Hz sampling rate (with a visual angle error of less than 0.5°). Staring at a target for more than one second triggers a highlight in the interface. If the user does not actively trigger a gesture or voice command, the system defaults to eye lock as the starting point for interaction.

[0059] 1.1.3) Voice Interaction Unit: This unit utilizes the Oculus Quest 3's built-in microphone array and the Google Speech-to-Text API for voice recognition. Voice commands serve as a fallback when gestures fail. Users can select scene difficulty, view prompts, or complete simple commands (such as "Confirm" or "Exit") by voice.

[0060] 1.1.4) Gesture Interaction Unit: This unit uses the Oculus Quest 3 gesture recognition module to capture hand movements (including grasping, rotating, and pointing). If a hand movement deviates from the target area by more than 30%, it automatically switches to voice command assistance mode.

[0061] 1.2) Data fusion algorithm unit:

[0062] 1.2.1) Spatiotemporal Alignment Technology: Using a Kalman filter, we synchronize the timestamps of multi-source data (skeleton, eye movement, and voice) with the spatial coordinate system to generate a user behavior feature vector. This user behavior vector is then fed into the Unity event system to trigger corresponding interaction events (such as capture, animation playback, and scene switching).

[0063] The spatial coordinate unification process includes: coordinate system transformation matrix;

[0064] (1) Global coordinate system: Azure Kinect is used as the origin (world coordinate system), and other devices are converted through the calibration matrix:

[0065] 1) Eye movement coordinate system → world coordinate system;

[0066] 2) Gesture coordinate system: Directly use the local coordinate system of the Quest 3 hand model and convert it to world coordinates through Unity's XR InputSubsystem.

[0067] (2) Dynamic alignment strategy: Recalibration trigger condition: When the distance between the skeletal joint point and the eye focus space is continuously greater than 30 cm (for 5 seconds), the calibration process is automatically started, requiring the user to look at the three preset calibration points (corner, center of the desktop, and virtual cursor).

[0068] 1.2.2) Feature extraction formula:

[0069] F=[Jt,Gt,Vt]∈R 18×3 ;

[0070] Among them, Jt is the joint coordinate, Gt is the eye movement trajectory, and Vt is the voice command feature.

[0071] 2) Nostalgic Scene Generation Module, which includes a 1970s item library unit, a dynamic scene rendering engine unit, and a scene difficulty and task design unit. This module generates interactive and immersive nostalgic scenes based on user-defined needs, supporting user-defined scene content and old object layout. The specific implementation is as follows:

[0072] 2.1) 70s Object Library Unit: This unit stores 3D models (OBJ format) and texture maps (4K PBR textures) of three life scenes from the 1970s (convenience store scene, home scene, and street scene). It contains typical old objects including but not limited to (as shown in Table 1):

[0073] Table 1

[0074]

[0075]

[0076] 2.2) Dynamic scene rendering engine unit:

[0077] 2.2.1) User-defined scene construction:

[0078] 2.2.1.1) Vintage Object Selection System: Before training, users use a gesture-based selection interface to select objects from a pre-set library of 1970s-era objects to add to the scene. Based on the user's selections, the system renders the scene in real time using the Unity HDRP pipeline, ensuring that the lighting and materials adhere to the nostalgic style (e.g., warm lighting and desaturated textures).

[0079] 2.2.2) 3D Photo Puzzle Generation Technology:

[0080] 2.2.2.1) Old photo upload and processing: Users upload old 2D photos (JPEG / PNG format), and the system generates interactive 3D puzzles through the following process:

[0081] Step 1: Use Maya modeling software to pre-cut the standard 3D picture frame model into 9 regular puzzle pieces.

[0082] Step 2: Extract photo feature points through the OpenCV image processing library, map the two-dimensional photo uploaded by the user to the surface of the three-dimensional photo frame, and generate a PBR material map.

[0083] Step 3: Instantiate the cut puzzle model in Unity, assign independent rigid body properties to each sub-puzzle to support physical interaction.

[0084] 2.2.2.2) Jigsaw Puzzle Task Design: During cognitive training, puzzle pieces are randomly scattered throughout the scene, and users are required to grasp and place them within the frame using gestures. When the distance between the puzzle piece and the target position is less than 2 cm, a magnetic attraction effect (implemented through joint constraints in the PhysX engine) is triggered.

[0085] 2.3) Scenario Difficulty and Task Design Unit:

[0086] 2.3.1) 1970s Family Scene Task;

[0087] 2.3.1.1) Difficulty 1: Find the Item Mission (Find the Difference)

[0088] a) Task Description: Three target objects are displayed for 30 seconds and then moved to a different location (randomly shifted). After locking onto the target by gazing, the user must confirm the target object with gestures / voice commands. Specifically, a 30-second countdown is displayed at the top of the screen. When the countdown ends, the system automatically switches to the new, shuffled scene, and a prompt is displayed in the upper right corner of the screen. The user can click to view the prompt, which is the original correct scene. Each click displays for 10 seconds and then automatically closes. The user can click multiple times to view the prompt content until all target objects are found.

[0089] 2.3.1.2) Difficulty 2: Object Return Task (Reset);

[0090] a) Task description: The three target items are displayed for 30 seconds and then moved to other locations (random displacement). The user locks the target by gazing, and then uses gestures to grab and restore the items to their original positions. Specifically, a 30-second countdown is displayed at the top of the screen. When the countdown ends, the system automatically switches to the new scrambled scene, and a prompt is displayed in the upper right corner of the screen. The user can click to view the prompt, which is the original correct scene. It will be displayed for 10 seconds after each click and will automatically turn off. The user can click multiple times to view the prompt content. Until the item is returned to the correct position, when the distance between the selected item and the target position is less than 2cm, the magnetic effect is triggered (achieved through the joint constraints of the PhysX engine).

[0091] 2.3.1.3) Difficulty 3: Object Interaction Task (Interaction);

[0092] a) Task description: 3 target objects are displayed for 30 seconds and then moved to other locations (random displacement). The user locks the target by gazing, and then uses gestures to grab the objects and restore them to their original positions. Finally, the user completes a simple interactive task (such as playing TV programs, using sewing machines, playing records, etc.). Specifically, a 30-second countdown is displayed at the top of the screen. When the countdown ends, the system automatically switches to the new scrambled scene, and a prompt is displayed in the upper right corner of the screen. The user can click to view the prompt, which is the original correct scene. It will be displayed for 10 seconds after each click and will automatically turn off. The user can click multiple times to view the prompt content. Until the object interaction task is completed.

[0093] 2.3.1.4) The animation triggering mechanism for item interaction tasks is as follows:

[0094] Prefabricated animations (such as TV program animations, sewing machine thread animations, and clock hand movement animations) are triggered through collision detection. Animation playback is controlled by Unity Animator, and the event system is bound to sound effects and visual feedback. Specifically, Maya / 3D Max is used to fine-tune the modeling of old objects and create component motion animations. Taking TV interaction as an example, the user triggers a highlight prompt by staring at the TV knob for more than 1 second. The knob is rotated with a gesture, and the knob angle is mapped to the volume / channel parameters. Play TV programs from the 1970s (the video stream is rendered to the screen material through the Unity Video Player component).

[0095] 2.3.2) 1970s convenience store scenario task;

[0096] 2.3.2.1) Difficulty 1: Basic Shopping Task;

[0097] a) Task Description: There are 18 items on the shelf, with 6 items of each, for a total of 108 items. Three items will appear randomly during each task. The shopping list will display three items, with one item of each. The user must add all three items on the shopping list to the shopping cart. Total number of items: 108. Shopping List: Three items will appear randomly during each task. Task Objective: Collect a total of three items.

[0098] b) Data Recording Process: The system presets a shopping list (e.g., "Green Treasure Orange Juice Soda_Small x 1"). Each time a user grabs an item, the backend verifies that the item ID matches. At the end of the task, the number of correct items grabbed, any excess or insufficient items, and a score are calculated.

[0099] 2.3.2.2) Difficulty 2: Quantity-increasing tasks;

[0100] a) Task Description: There are 18 items on the shelf, with 6 items of each, for a total of 108 items. Three items will appear randomly in each task, and the shopping list will display three items, with 1-3 items of each. The user must place the items on the shopping list into the shopping cart, for a total of 6 items. Total number of items: 108. Shopping List: Three items will appear randomly in each task. Task Objective: Collect a total of 6 items.

[0101] b) If the shopping list contains dynamic quantities (e.g., "Vitasoy_Small×2"), the system records the actual quantity grabbed and compares the difference.

[0102] 2.3.2.3) Difficulty level 3: Size change task;

[0103] a) Task Description: There are 18 items on the shelf. Each item comes in two sizes, large and small, with 3 items of each size, for a total of 108 items. Three items will appear randomly in each task. The shopping list will display the quantity and size of the items. The user must place the items in the correct size and quantity according to the shopping list, for a total of 6 items. Total number of items: 108. Shopping List: Three items will appear randomly in each task. Task Objective: Collect a total of 6 items (including different sizes).

[0104] b) The system interprets the size requirements in the shopping list (e.g., "Green Treasure Orange Juice_Large x 2") and verifies the metaData.size field during fetching. If the user fetches a small size, the size is marked as incorrect.

[0105] 2.3.2.4) The shopping interaction logic implementation mechanism is as follows:

[0106] Through collision detection and metadata matching, the backend records user operation data and classifies error types. Specifically, first, an ItemMetaData script is attached to each product model to store the following attributes:

[0107] public stringitemID; / / Unique identifier of the product (such as "Green Treasure Orange Juice_Large");

[0108] public string category; / / Product category (such as "beverage");

[0109] public string size; / / size ("large" / "small");

[0110] Next, add a Box Collider to the shopping basket model and label it as the Basket tag, checking the Is Trigger property. Next, the system compares the ID of the product captured by the user with the expected value in the shopping list in real time, generating the following error types: (1) Category error (the captured product category does not match the list); (2) Size error (the product size does not match the list); (3) Quantity error (the total number of captured items exceeds or falls short of the list requirement).

[0111] 2.3.2.5) Convenience Store Scenario Scoring System and Report Generation:

[0112] a) Score calculation formula:

[0113]

[0114] Constraint: The final score must be lower than 0 points.

[0115] b) Generate a summary report page through the Unity UI system, and use the Unity Text component (TextMeshPro) to render the score text to display the final score graph.

[0116] 2.3.3) Difficulty setting for 1970s street scenes;

[0117] 2.3.3.1) Difficulty 1: Novice Mode (i.e., the user's movement accuracy is 0%-70%);

[0118] Task Description: The user performs a 30-minute physical exercise with a virtual trainer and their avatar. This difficulty level provides full virtual trainer and avatar functionality and comprehensive feedback, including guidance through voice commands, video demonstrations, and bone color cues, enabling the user to gradually fine-tune their performance and improve their movement precision. This guidance is implemented using various functional systems within the Unity Editor. Task Objective: Gradually fine-tune movements and improve movement precision.

[0119] 2.3.3.2) Difficulty Level 2: Experienced Mode (i.e., user's movement accuracy is 70%-75%);

[0120] Task Description: The user will perform a 30-minute physical exercise with a virtual trainer and avatar. This difficulty level introduces a virtual trainer and avatar, providing only voice commands and video demonstrations. Task Objective: Further improve athletic performance.

[0121] 2.3.3.3) Difficulty Level 3: Expert Mode (i.e., the user's movement accuracy is 75%-100%);

[0122] Task Description: The user follows a virtual trainer through a 30-minute workout. This difficulty level only provides basic instructions from the virtual trainer, minimizing visual complexity. Task Objective: Maintain a flow experience and complete a high-intensity workout.

[0123] 2.3.3.4) The implementation mechanism of virtual physical exercise interaction logic is as follows:

[0124] a) Motion Capture: The system monitors 32 skeletal joints throughout the user's body (including key areas such as shoulders, elbows, wrists, and knees) during warm-up exercises to obtain the user's warm-up motion data. Using a motion matching algorithm, the system calculates the difference between the user's warm-up motion data and the virtual coach's standard motion data. This difference is used to establish a correction threshold for the user's formal training session, deriving the user's motion accuracy and supporting the recommendation of learning modes during the warm-up training session. The motion matching accuracy formula is as follows:

[0125]

[0126] In the above formula, Represents the joint coordinates of the user or trainer, and n represents the number of selected joints used for calculation. k ) normalized Pose is a normalized vector calculated from two joint coordinates. match Calculate the average cosine similarity between the user's motion vector sequence and the standard motion vector sequence. Note that the present invention sets the maximum allowable angle between matching vectors to 90°. If the angle exceeds 90°, the matching accuracy is determined to be 0.

[0127] b) Movement Feedback: The system monitors and records the user's movement accuracy and compares the user's formal training movement data with the warm-up movement data. If the error between the two sets of data exceeds the correction threshold set by the warm-up training unit, the movement coaching unit provides the user with corresponding error correction guidance. In a further embodiment, the error correction guidance includes voice commands, video demonstrations, and bone color cues.

[0128] (1) The steps for calling language instructions are as follows: the system references the animation event in the coach animation. When the coach animation plays a specific action, it calls the play function of the Audio Source to play the corresponding language. In addition, the system calculates the matching degree between the user and the coach's movements (the calculation formula is the same as the movement matching accuracy formula mentioned above) and outputs the user's joint information. The language instruction code determines whether the user's movement matching accuracy meets the conditions for triggering the language instruction. If so, it plays the specified feedback audio.

[0129] (2) The steps of video demonstration are described as follows: the system calculates the matching degree between the user and the coach (calculated using the same formula as above for matching accuracy) and outputs the user's joint information. The video playback code determines whether the user's matching accuracy meets the conditions for the demonstration video. If so, it activates the object containing the Video Player component and demonstrates the correct movement of the move in a video.

[0130] (3) The steps for improving bone color are as follows: The system calculates the degree of motion matching between the user and the coach (calculated using the same formula as above for motion matching accuracy) and outputs the user's joint information. The video playback code determines whether the user's motion matching accuracy meets the conditions for displaying bone color. If so, the object containing the Image component in the parent Canvas object is replaced and the color of the corresponding bone position of the user is changed.

[0131] 3) Dynamic difficulty control module, which includes user-selected difficulty units and system-customized difficulty units.

[0132] 3.1) User-selected difficulty unit: In the 1970s family scene (schematic diagram of the operating mechanism is as follows Figure 1 ) and the convenience store scene in the 1970s (the operating mechanism structure diagram is shown in Figure 2 As shown in the figure, users can choose from three levels of difficulty independently, and the system does not make any recommendations.

[0133] 3.2) The system personalizes the difficulty unit: In the street scene of the 1970s, the user follows the virtual coach to perform a 1-minute warm-up activity. The system will evaluate the user's exercise level based on the user's exercise performance and recommend the corresponding learning mode to the user. Specifically, in the warm-up training unit of the system, the user follows the coach to perform warm-up training. Among them, the Eight-Section Brocade moves instructed by the coach are warm-up moves, that is, each move is simplified, and each move is only performed once (a complete move is a move movement that is repeated three times). Therefore, the system recommends a formal training learning mode suitable for the user based on the user's movement proficiency in the warm-up stage. There are three types of learning modes, and users can choose different modes for formal training. The learning modes include novice mode (that is, the user's movement accuracy is 0%-70%), experienced mode (that is, the user's movement accuracy is 70%-75%), and master mode (that is, the user's movement accuracy is 75%-100%). The specific operating mechanism structure diagram is shown as follows: Figure 3 shown.

[0134] 4) Multimodal feedback output module, which includes a visual feedback unit, an auditory feedback unit, and a motion feedback unit.

[0135] 4.1) Visual feedback unit:

[0136] 4.1.1) Highlighting: Use Unity Shader programming to make the target object flash red (such as staring at the locked object).

[0137] 4.1.2) Dynamic interface display: Use the Unity UI system (Canvas+TextMeshPro) to update task progress, countdown and prompt information in real time.

[0138] 4.1.3) Skeleton color feedback: Dynamically adjust the user avatar’s skeleton color based on the action matching degree (e.g. green indicates correct action, red indicates deviation).

[0139] 4.2) Auditory feedback unit:

[0140] 4.2.1) Voice Guidance: Play pre-recorded voice instructions (such as "Please put the items back in place") through the Unity Audio Source class.

[0141] 4.2.2) Sound Effect Trigger: Combined with animation events (Animation Event) to play scene sound effects (such as TV program sound, puzzle piece return sound effect).

[0142] 4.3) Action feedback unit:

[0143] 4.3.1) Physics simulation: The PhysX engine is used to achieve the magnetic attraction effect of objects (the puzzle piece will automatically adsorb when the distance between it and the target position is less than 2cm).

[0144] 4.3.2) Virtual Coach Demonstration: Based on motion capture data, the user's formal training movement data is compared with the warm-up movement data. If the error between the two sets of data exceeds the correction threshold set by the warm-up training unit, the user will be guided to adjust his posture.

[0145] Summary Report Module: This module provides task performance summary reports to help users review their performance in cognitive and physical training games and track their training progress. The summary report for the home scenario specifically includes the user's task completion status (success / failure) and the achievement of task goals (i.e., item finding accuracy %, item placement accuracy %, and interactive action completion %). The summary report for the convenience store scenario specifically includes the user's completion status (success / failure) and the achievement of task goals (i.e., shopping list completion accuracy %). The summary report for the community street scenario specifically includes the user's performance level and action accuracy for each action.

[0146] Data collection and analysis module: This module collects and stores user operation data. This data includes:

[0147] 5.1) User basic information: user ID, age, gender, and experiment start date.

[0148] 5.2) Task performance data: task ID, scenario type, difficulty level, task start time, task end time, task duration, score breakdown (i.e., number of correct items, number of incorrect items, number of size errors, number of quantity errors, number of category errors, other error types, number of correct actions, number of incorrect actions, etc.).

[0149] 5.3) Operational behavior data: number of grabs, number of prompt views, skeletal coordinate positions, etc.

[0150] Example 2

[0151] This embodiment provides a multimodal interactive control system based on virtual reality and sensing technology, including:

[0152] Step 1: System initialization and introduction;

[0153] This implementation uses the Oculus Quest2 VR headset. Users wear a head-mounted display and can experience VR without controllers. Users enter the virtual environment in a 6.5m x 2.8m x 2.8m high practice area. The module includes user ID input, scene introduction, and background introduction. The background introduction uses immersive graphics and voice guidance to guide users through system functions and operations, ensuring that elderly users can quickly get started.

[0154] Step 2: Enter the scene;

[0155] The user will briefly learn the story and press the OK button to enter the task execution phase. Task execution order: After the user enters the number to log in, the system automatically enters the home, convenience store and street scenes in the preset order and completes a task. Each scene is divided into three difficulty levels (difficulty 1, difficulty 2, difficulty 3), and each difficulty level contains 16 tasks. The specific execution order is:

[0156] Family scenario: Complete 16 tasks of difficulty 1, 16 tasks of difficulty 2, and 16 tasks of difficulty 3 in sequence;

[0157] Convenience store scenario: Complete 16 tasks of difficulty 1, 16 tasks of difficulty 2, and 16 tasks of difficulty 3 in sequence;

[0158] Street Scene: Complete 16 missions of Difficulty 1, 16 missions of Difficulty 2, and 16 missions of Difficulty 3 in sequence.

[0159] The user completes a total of 48 tasks (16 times at home × 3 difficulty levels + 16 times at convenience stores × 3 difficulty levels + 16 times on the street × 3 difficulty levels). The system automatically records and generates a task performance report.

[0160] Step 3: Task execution and summary report generation

[0161] The experiencer will perform tasks related to the scenario, including the task process for the home scenario, the task process for the convenience store scenario, and the task process for the street scenario.

[0162] Step 4: End the session;

[0163] Take off the VR headset to automatically exit and end all training.

[0164] Through the above implementation examples, the present invention provides a more efficient, safe, scientific and user-friendly self-management system for the elderly with mild cognitive impairment, which significantly improves the health intervention effect and user experience.

[0165] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A multimodal interactive control system based on virtual reality and sensing technology, characterized in that: include: Multimodal data acquisition module, used to collect user's eye movement, gesture, voice and skeletal movement data to generate multimodal data; The data fusion module is used to fuse the collected multimodal data and generate user behavior feature vectors; A nostalgic scene generation module, configured to generate a personalized nostalgic scene based on the user behavior feature vector; A dynamic difficulty control module, configured to dynamically adjust the difficulty of a task based on the user's performance in the personalized nostalgic scene; A multimodal feedback output module, which provides visual, auditory, and motion feedback to users based on task difficulty; Summary report module, used to generate user task performance summary report; The data collection and analysis module is used to collect and store user operation data and task performance data.

2. The system according to claim 1, wherein: The multimodal data acquisition module includes an eye tracking unit, a gesture recognition unit, a speech recognition unit and a skeleton tracking unit; The eye tracking unit is used to capture the user's line of sight focus coordinates; The gesture recognition unit is used to capture the user's hand movements; The speech recognition unit is used to receive and analyze the user's voice instructions; The skeleton tracking unit is used to capture the coordinates of the user's entire body skeleton joints in real time.

3. The system according to claim 1, wherein: The data fusion module synchronizes the timestamps and aligns the spatial coordinates of the eye movement, gesture, voice and skeletal movement data through a Kalman filter; The data fusion module generates a user behavior feature vector based on the synchronized data.

4. The system according to claim 1, wherein: The nostalgic scene generation module includes a 1970s item library unit, a dynamic scene rendering engine unit, and a scene difficulty and task design unit; The 1970s item library unit stores three-dimensional models and animation data of items in life scenes of the 1970s; The dynamic scene rendering engine unit generates a personalized scene in real time based on the old photos provided by the user and the selection made during the virtual environment setting process; The scenario difficulty and task design unit sets three difficulty modes according to the scenario, and the task design stimulates autobiographical memory through nostalgia therapy.

5. The system according to claim 4, characterized in that The dynamic difficulty control module includes user-selected difficulty units and system-customized difficulty units; The user-selectable difficulty unit allows the user to select from three difficulty levels. The system personalized difficulty unit evaluates the user's exercise level based on the user's exercise performance and recommends a corresponding learning mode to the user.

6. The system according to claim 1, wherein: The multimodal feedback output module includes a visual feedback unit, an auditory feedback unit and a motion feedback unit; The visual feedback unit uses Unity Shader programming to achieve target object highlighting, uses the UnityUI system to update task progress, countdown and prompt information in real time, and dynamically adjusts the user avatar skeleton color according to the action matching degree; The auditory feedback unit plays pre-recorded voice commands through the UnityAudio Source class and plays scene sound effects in combination with animation events; The motion feedback unit realizes the magnetic attraction effect of objects through the PhysX engine and guides the user to adjust the posture based on the motion capture data.

7. The system according to claim 1, wherein: The summary report module provides a task performance summary report to help users review their performance in cognitive and physical exercise games and track training progress; the report content includes the user's task completion status and performance level of task goal achievement.

8. A multimodal interactive control method based on virtual reality and sensing technology, characterized in that: include: Collect user's eye movement, gesture, voice and skeletal movement data; Fuse the collected multimodal data to generate user behavior feature vectors; Generate personalized nostalgic scenes based on user behavior feature vectors; Dynamically adjust task difficulty based on user performance in nostalgic scenarios; Provide users with visual, auditory, and motor feedback based on task difficulty; Collect and store user operation data and task performance data.

9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein: When the processor executes the computing program, the method according to claim 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to claim 8 is implemented.