Guidance method and device based on multi-modal perception, equipment and medium

CN120611149BActive Publication Date: 2026-09-15PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510701696.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2026-09-15
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

[0006]本发明的主要目的在于提供一种基于多模态感知的指导方法、装置、设备及存储介质,旨在解决现有技术缺乏对个体身心状态的精准感知与个性化交互优化,难以提供适配性的沉浸式指导的技术问题

Benefits of technology

[0025]Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health finance, financial technology, and mental health care. It discloses a guidance method based on multimodal perception, including: collecting users' physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body dataset; inputting a pre-trained mind-body association model to generate a mind-body state mapping relationship; generating an immersive interactive scene containing environmental simulation elements based on a three-dimensional scene model library and dynamic attention parameters in the mind-body state mapping relationship; parsing scene interaction commands in the immersive interactive scene according to entity association relationships in a domain knowledge graph; fusing historical behavioral feature vectors from user profiles with real-time state vectors in the mind-body state mapping relationship to generate comprehensive decision parameters; and selecting personalized guidance content from a guidance content library based on the feature similarity between the comprehensive decision parameters and scene interaction commands, and outputting the guidance content. This invention achieves accurate perception of users' physiological states and behavioral characteristics through multimodal data collection and mind-body state modeling, and improves the accuracy of personalized guidance by combining immersive scene interaction and knowledge graph parsing. The generation and matching optimization of comprehensive decision parameters enable interactive content to adapt to different users' behavioral patterns, improving the personalization and immersiveness of intelligent guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611149B_ABST
    Figure CN120611149B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, can be applied to business scenes such as medical health finance, financial technology and psychological health care, and discloses a guiding method, device, equipment and medium based on multi-modal perception, which comprises the following steps: collecting physiological index data and motion trajectory data, generating a multi-modal body data set, inputting a pre-trained mind-body correlation model, and generating a mind-body state mapping relationship; in combination with a three-dimensional scene model library and dynamic attention parameters, an immersive interactive scene is generated, and a scene interaction instruction is analyzed based on a domain knowledge graph; a comprehensive decision parameter is generated by fusing a user portrait and a real-time state vector, and personalized guidance content is selected from a guidance content library and output according to the feature similarity between the comprehensive decision parameter and the scene interaction instruction. Through multi-modal data acquisition and mind-body state modeling, the application realizes accurate perception of the physiological state and behavior characteristics of a user, and improves the accuracy of personalized guidance in combination with immersive scene interaction and knowledge graph analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a guidance method, apparatus, device, and storage medium based on multimodal perception. Background Technology

[0002] In the process of developing and applying intelligent interactive systems, although multimodal perception technology, immersive interactive systems and artificial intelligence have made significant progress in many industries, existing technologies still have many shortcomings. In particular, in the fields of medical and health care, fintech business and mental health care, intelligent systems still have significant limitations in simulating complex scenarios, accurately identifying individual behavioral patterns and providing personalized guidance, making it difficult to meet the in-depth needs of different fields.

[0003] In the healthcare sector, existing technologies have shortcomings in intelligent health monitoring and behavioral intervention. Current health management systems primarily rely on user-inputted data or static health data analysis, lacking a holistic understanding of users' multimodal physiological data (such as heart rate, skin conductivity, and movement patterns) and psychological states, resulting in an incomplete assessment of users' health status. Furthermore, existing technologies also have limitations in rehabilitation training and psychological intervention. For example, during rehabilitation, patients' movement behavior and psychological state are closely related, but current technologies struggle to accurately capture individual movement trajectories and emotional fluctuations, making it difficult to develop personalized intervention plans. In addition, in mental health interventions, existing intelligent systems struggle to integrate the dynamic relationship between physiological and psychological changes, failing to accurately identify patients' emotional states and behavioral patterns, leading to low accuracy and real-time performance of personalized treatment recommendations. While existing immersive therapies provide a partially immersive environment, they lack deep interaction with patients' physiological and psychological feedback, resulting in limited therapeutic effects.

[0004] In the field of mental health care, existing technologies primarily focus on the cognitive processing of information, lacking a deep understanding of the integration of mind and body. For example, during mental health care, an individual's breathing rhythm, body posture, and mental state are closely related; this mind-body interaction directly affects relaxation and mood regulation. However, current intelligent technologies struggle to capture and model this mind-body interaction, remaining at a superficial level of guidance, such as playing relaxing music or providing fixed meditation guidance, rather than offering precise mental health care assistance based on the individual's real-time state. Furthermore, current intelligent interactive systems cannot effectively simulate real-world care scenarios, such as a tranquil forest walk, soothing lakeside meditation, or a quiet resting space, resulting in a fragmented user experience in virtual environments, lacking immersion and realism.

[0005] In the fintech sector, existing intelligent analytics systems primarily rely on static data for financial behavior analysis, making it difficult to accurately identify users' true decision-making preferences and risk tolerance. Current credit assessment and financial recommendation systems mainly rely on users' historical transaction records, financial data, and market trends, neglecting multimodal information such as users' physiological states and behavioral patterns. For example, during financial decision-making, a user's emotional state, focus, and stress level can directly influence investment behavior and risk preferences, but current technologies struggle to perceive these factors in real time and make corresponding adjustments, resulting in insufficient accuracy and personalization in financial decision support systems. Furthermore, in intelligent financial services, systems often fail to accurately understand users' intentions and needs; financial advice still relies on fixed logic for recommendations, lacking the ability to dynamically adjust based on the user's current cognitive state. Meanwhile, in immersive trading and financial education scenarios, existing virtual systems still present static rules, failing to dynamically optimize the learning experience based on user behavioral feedback and psychological states, resulting in a lack of effective immersion and interactive experience for users during financial decision-making and learning. Summary of the Invention

[0006] The main objective of this invention is to provide a guidance method, device, equipment, and storage medium based on multimodal perception, aiming to solve the technical problem that the existing technology lacks accurate perception of individual physical and mental state and personalized interaction optimization, making it difficult to provide adaptive immersive guidance.

[0007] To achieve the above objectives, the present invention provides a guidance method based on multimodal perception, comprising:

[0008] Collect users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset;

[0009] The multimodal body dataset is input into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0010] Based on a 3D scene model library, and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated.

[0011] Based on the entity relationships in the domain knowledge graph, the scene interaction commands in the immersive interactive scene are parsed.

[0012] By integrating the historical behavioral feature vectors in the user profile with the real-time state vectors in the mapping relationship between the physical and mental states, comprehensive decision parameters are generated.

[0013] Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, personalized guidance content is selected from the guidance content library;

[0014] Output the personalized guidance content.

[0015] Furthermore, to achieve the above objectives, the present invention provides a guidance device based on multimodal perception, comprising:

[0016] The physiological and motion data acquisition module collects users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset.

[0017] The mind-body state modeling module inputs the multimodal body dataset into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0018] The immersive scene generation module, based on a 3D scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, generates an immersive interactive scene containing environmental simulation elements.

[0019] The interaction instruction parsing module parses the scene interaction instructions in the immersive interaction scene based on the entity association relationships in the domain knowledge graph;

[0020] The personalized behavior analysis module integrates historical behavior feature vectors from user profiles with real-time state vectors from the mapping relationship between physical and mental states to generate comprehensive decision parameters.

[0021] The guidance content matching module selects personalized guidance content from the guidance content library based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions.

[0022] The personalized content output module outputs the personalized guidance content.

[0023] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal perception-based guidance program stored in the memory and executable on the processor, wherein the multimodal perception-based guidance program, when executed by the processor, implements the steps of the multimodal perception-based guidance method as described above.

[0024] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multimodal perception-based guidance program, wherein the multimodal perception-based guidance program, when executed by a processor, implements the steps of the multimodal perception-based guidance method as described above.

[0025] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health finance, financial technology, and mental health care. It discloses a guidance method based on multimodal perception, including: collecting users' physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body dataset; inputting a pre-trained mind-body association model to generate a mind-body state mapping relationship; generating an immersive interactive scene containing environmental simulation elements based on a three-dimensional scene model library and dynamic attention parameters in the mind-body state mapping relationship; parsing scene interaction commands in the immersive interactive scene according to entity association relationships in a domain knowledge graph; fusing historical behavioral feature vectors from user profiles with real-time state vectors in the mind-body state mapping relationship to generate comprehensive decision parameters; and selecting personalized guidance content from a guidance content library based on the feature similarity between the comprehensive decision parameters and scene interaction commands, and outputting the guidance content. This invention achieves accurate perception of users' physiological states and behavioral characteristics through multimodal data collection and mind-body state modeling, and improves the accuracy of personalized guidance by combining immersive scene interaction and knowledge graph parsing. The generation and matching optimization of comprehensive decision parameters enable interactive content to adapt to different users' behavioral patterns, improving the personalization and immersiveness of intelligent guidance. Attached Figure Description

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0027] Figure 1 This is a schematic diagram of an application environment for a guidance method based on multimodal perception in one embodiment of the present invention;

[0028] Figure 2 This is a flowchart illustrating an embodiment of the guidance method based on multimodal perception of the present invention;

[0029] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the guidance device based on multimodal perception of the present invention;

[0030] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0031] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0032] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0033] The guidance method based on multimodal perception provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can collect users' physiological index data and 3D motion trajectory data from the user terminal to generate a multimodal body dataset; input a pre-trained mind-body association model to generate a mind-body state mapping relationship; based on the 3D scene model library and the dynamic attention parameters in the mind-body state mapping relationship, generate an immersive interactive scene containing environmental simulation elements; parse the scene interaction commands in the immersive interactive scene according to the entity association relationship in the domain knowledge graph; fuse the historical behavioral feature vector in the user profile with the real-time state vector in the mind-body state mapping relationship to generate comprehensive decision parameters; based on the feature similarity between the comprehensive decision parameters and the scene interaction commands, select personalized guidance content from the guidance content library and output the guidance content. This invention achieves accurate perception of users' physiological state and behavioral characteristics through multimodal data collection and mind-body state modeling, and improves the accuracy of personalized guidance by combining immersive scene interaction and knowledge graph parsing. The generation and matching optimization of comprehensive decision parameters enable the interactive content to adapt to different users' behavioral patterns, improving the personalization and immersion of intelligent guidance. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0034] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal perception-based guidance method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0035] like Figure 2 As shown, the guidance method based on multimodal perception proposed in this invention includes the following steps:

[0036] S10 collects users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset;

[0037] In this embodiment, the core objective of collecting users' physiological index data and 3D motion trajectory data to generate a multimodal body dataset is to comprehensively perceive the user's physiological state and movement patterns, constructing complete physical and mental state information to support subsequent behavioral modeling and immersive interaction. The physiological index data is collected from various wearable smart devices, including smart bracelets, smart clothing, and ear-worn monitoring devices. These devices can monitor key physiological parameters such as heart rate, skin conductivity, and blood oxygen saturation in real time. Heart rate monitoring reflects the user's autonomic nervous system activity, skin conductivity is used to perceive the user's emotional fluctuations, and blood oxygen saturation can be used to assess the user's respiratory status and oxygen supply. In addition to basic physiological parameters, brainwave sensors can be used to obtain brainwave activity data to more accurately characterize the user's psychological responses in different states.

[0038] The acquisition of 3D motion trajectory data involves inertial measurement units (IMUs), optical motion capture devices, and computer vision analysis techniques. IMUs, including accelerometers and gyroscopes, acquire user posture information, direction of motion, and acceleration changes. Optical motion capture devices use multiple high-precision cameras to capture user movement, identifying detailed features of limb movements such as joint angles, amplitude of motion, and gait characteristics. Furthermore, computer vision analysis techniques use deep learning models to segment, track, and identify user movements, making the analysis of complex movements more accurate. The construction of 3D motion trajectories not only includes the trajectory of body movement but can also incorporate data from force feedback sensors to obtain the force experienced by the user when performing specific actions, thereby further optimizing the accuracy of motion analysis.

[0039] Based on the collected physiological index data and 3D motion trajectory data, data preprocessing and fusion are required to ensure data consistency and accuracy. First, timestamp synchronization of data from different sources is necessary to enable analysis of physiological and motion data on the same timeline. Second, signal filtering is required for both physiological and motion data to remove noise interference and improve data reliability. Filtering algorithms such as Kalman filtering and low-pass filtering can be used to ensure high stability of sensor data even under high-frequency jitter. Furthermore, data format standardization is necessary during the fusion of physiological and motion data to support subsequent deep learning modeling and analysis. Standardization methods can be based on maximum / minimum value normalization or z-score normalization to eliminate the influence of different measurement units on data analysis.

[0040] Wearable devices can collect physiological data using photoplethysmography (PPG) technology, which detects minute changes in blood flow to obtain high-precision heart rate and blood oxygenation data. Additionally, electrophysiological signals can be acquired via skin electrodes to analyze changes in the user's sympathetic nervous system activity. In motion data acquisition, high-precision inertial measurement units (IMUs) combined with computer vision technology can improve motion capture accuracy. For example, when a user performs specific yoga poses, the system can detect the user's body posture using an IMU and correct detection errors using an optical camera to provide more accurate posture data.

[0041] During data fusion, different processing strategies can be adopted depending on the application scenario. In high-precision scenarios (such as rehabilitation training), a Bayesian-optimized multimodal fusion algorithm can be used to weight data from different sources to improve the accuracy of the fused data. In real-time interactive scenarios (such as immersive virtual experiences), recurrent neural networks can be used to perform time-series modeling of the data to ensure its timeliness and continuity. Regarding data transmission, edge computing technology can be combined to complete some data preprocessing tasks on wearable devices, reducing data transmission latency and improving the system's real-time response capabilities.

[0042] Example: In the healthcare field, remote patient monitoring and rehabilitation training guidance can be implemented. For instance, for postoperative rehabilitation patients, by collecting gait data and physiological parameters, the patient's recovery status can be assessed, and the rehabilitation training plan can be adjusted based on data changes. If gait instability is detected, along with an abnormally high heart rate, it can be inferred that the patient may be fatigued or uncomfortable, and the system can provide rest suggestions or adjust the training intensity.

[0043] In mental health monitoring, physiological and motor data can be combined to analyze a user's emotional state. For example, by detecting a user's heart rate variability, skin conductivity, and micro-expression features, their anxiety level can be assessed, and relaxation training or psychological intervention guidance can be provided at appropriate times. Furthermore, in anxiety treatment, three-dimensional motion data can be used to assess a user's breathing rhythm, guiding them to adjust their breathing patterns and improve their mental state.

[0044] In the financial sector, systems can be optimized to support user financial decisions. For example, when a user makes a large transaction decision, the system can assess their decision-making state by combining physiological indicators (such as stress levels) with operational behaviors (such as mouse movements and keyboard input patterns). If the system detects that the user is under high stress, it can appropriately postpone transaction confirmation or provide a cooling-off period suggestion to reduce the risk of impulsive decisions. Furthermore, in financial education, combining user attention data and cognitive load analysis can dynamically adjust the pace and difficulty of teaching content, improving learning efficiency.

[0045] By jointly collecting and fusing physiological and motion data, a comprehensive understanding of a user's physical and mental state can be achieved, enabling precise analysis of user behavior patterns. Compared to single data collection methods, this multimodal perception technology provides more complete and granular user state information, laying a solid data foundation for subsequent physical and mental state modeling, immersive scene generation, and personalized interaction.

[0046] S20, input the multimodal body dataset into the pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0047] In this embodiment, the core objective of inputting a multimodal body dataset into a pre-trained mind-body association model to generate a mind-body state mapping relationship is to model an individual's mind-body state using multimodal data, extract correlation features between physiological indicators, movement patterns, and psychological states, and construct a precise mind-body interaction relationship. The mind-body association model is a time-series analysis model built using deep learning methods, capable of feature extraction, state prediction, and pattern analysis of collected physiological and movement data. This model employs neural network structures such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or gated recurrent units (GRUs) to ensure the temporal dependence of the data, while also introducing an attention mechanism to enable the model to focus on key physiological and movement features.

[0048] Before inputting a multimodal body dataset, preprocessing is required, including timestamp alignment, data normalization, outlier detection, and noise removal. Timestamp alignment involves synchronizing data from different data sources to ensure all data can be modeled within the same time window. Data normalization can employ min-max normalization or standardization to reduce the impact between different data dimensions. Outlier detection addresses potential short-term distortions or data loss in sensor data, using sliding window statistical analysis or autoencoder methods for anomaly detection and data completion. Noise removal can be achieved through filtering methods such as low-pass filtering or wavelet denoising techniques to improve data stability.

[0049] After preprocessing, the input data is divided into time-series data segments and fed into the mind-body connection model. This model employs a multi-layer neural network structure, receiving subsets of physiological indicators, 3D motion trajectories, and action features at the input layer. Through a cross-modal feature fusion layer, convolutional neural networks (CNNs) or self-attention mechanisms are used to extract correlation features between different modalities, such as the synchronicity between breathing rhythm and gait, and the coordination between heart rate changes and limb movements. During deep learning training, the model can be optimized using supervised or self-supervised learning methods. Training data can come from large-scale health databases, motion behavior datasets, or data collected from specific experiments.

[0050] At the model's output, a mind-body state mapping relationship is generated, including dynamic attention parameters and mind-body state assessment parameters. Dynamic attention parameters measure an individual's dependence on different physiological and motor characteristics at specific points in time. For example, during meditation, the system might focus more on the user's breathing depth, while during running training, the system might focus more on heart rate changes. Mind-body state assessment parameters include stress index, focus weight, emotional stability, and behavioral normativity assessment coefficients. These parameters can be used for individual state analysis and subsequent interactive decision-making.

[0051] By deeply fusing multimodal data and modeling the connection between mind and body, we can accurately depict a user's physical and mental state, making up for the shortcomings of existing technologies in perceiving individual physical and mental states. It can not only analyze an individual's physiological state but also combine motor behavior for comprehensive assessment, improving our understanding of an individual's condition.

[0052] S30, based on the three-dimensional scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated.

[0053] In this embodiment, based on a 3D scene model library and combined with dynamic attention parameters in the mind-body state mapping relationship, the core objective of generating an immersive interactive scene containing environmental simulation elements is to enhance the individual's behavioral guidance effect through an immersive virtual environment, make the interactive experience more realistic, and adaptively adjust environmental features according to the user's mind-body state, thereby improving the user's focus, comfort, and interaction matching degree.

[0054] The 3D scene model library stores 3D models of multiple virtual environments, including buildings, natural scenes, interactive props, virtual characters, lighting systems, and sound elements. These models can be predefined scenes, such as psychological treatment rooms, gyms, and medical rehabilitation rooms, or personalized environments can be dynamically constructed based on user behavior needs using procedural generation methods.

[0055] The dynamic attention parameters in the mind-body state mapping relationship are derived from the user's multimodal physiological data (heart rate, stress index, focus weight, movement state, etc.) and are used to adjust the dynamic interactive elements of the virtual environment. For example, when a user's stress index is detected to be high, the brightness of the ambient light can be reduced and visual distractions reduced to provide a more relaxing interactive atmosphere; when a user's focus weight is detected to be low, the user's attention can be improved by enhancing audio guidance and adding visual cues.

[0056] The core elements of 3D environment simulation include: lighting systems, weather simulation, scene element interaction, and virtual character synchronization. The lighting system employs global illumination, adjusting the intensity of ambient light based on time, space, and user status. For example, in a meditation scene, the lighting can present a warm and soft tone, while in a sports training scene, the lighting can be more vibrant. Weather simulation enhances the realism of the environment; for example, during rehabilitation training, weather changes can increase the user's immersion and influence the training rhythm. Scene element interaction refers to the ability of objects in the virtual environment to provide real-time feedback to the user's actions. For instance, a virtual instructor character can synchronize with the user's movement rhythm, or the system can automatically trigger relevant prompts when the user approaches an interaction point.

[0057] By constructing environments based on a 3D scene model library, a highly immersive interactive experience can be provided to users, making behavioral guidance more realistic and intuitive. Combined with the mapping relationship between mind-body states, personalized adjustments to the scene can be achieved, allowing the virtual environment to dynamically adapt to changes in the user's state, thereby improving user focus and interactive experience.

[0058] S40, based on the entity association relationships in the domain knowledge graph, parse the scene interaction instructions in the immersive interactive scene;

[0059] In this embodiment, the core objective of parsing scene interaction instructions in an immersive interactive scenario based on entity relationships in a domain knowledge graph is to enable the system to understand and execute interaction logic that conforms to a specific context, ensuring that the interactive experience in the virtual scene can intelligently adapt to the user's state and behavior patterns. Traditional interactive systems rely on predefined rules for scene feedback and lack intelligent dynamic adaptation capabilities. In contrast, the domain knowledge graph-based method can make the interactive content more accurate and intelligent through semantic understanding, entity relationship reasoning, and context matching.

[0060] Domain knowledge graphs are structured systems for storing and representing knowledge within a specific domain. Their core elements include entities, entity attributes, and entity relationships. In immersive interactive scenarios, the construction of a knowledge graph involves elements such as scene elements, user behavior patterns, environmental variables, and interaction commands. For example, in an intelligent rehabilitation training scenario, scene elements might include "rehabilitation equipment," "exercise instructor," and "real-time feedback mechanism"; user behavior patterns might involve "gait training," "strength training," and "balance exercises"; environmental variables might include "light intensity," "background noise," and "time factor"; and interaction commands might involve "triggering a specific training program," "adjusting training difficulty," and "providing real-time voice feedback," etc.

[0061] Key technical aspects of parsing scene interaction commands include knowledge graph loading, scene element extraction, interaction pattern matching, command logic generation, and spatiotemporal alignment processing. Knowledge graph loading refers to retrieving stored knowledge structures for subsequent querying and reasoning. Scene element extraction involves computer vision, object recognition, or user input analysis to determine the interactive elements in the current virtual environment. For example, in a virtual yoga training scene, the system might detect that "the user is currently in a tree pose," "the virtual instructor is providing movement guidance," and "background music is playing meditation music." These elements need to be structured and stored in the knowledge graph for subsequent interaction logic reasoning.

[0062] Matching interaction patterns involves calculating the semantic similarity between the user's current behavior pattern and standard behavior patterns in the knowledge graph to find the interaction method that best suits the current context. For example, in a smart fitness scenario, if the user's gait data has a high match with the "jogging mode" in the knowledge graph, the system can automatically recommend the best movement adjustment suggestions for jogging training. The generation of instruction logic relies on the entity relationships in the knowledge graph. For example, if a user selects the "strength training" mode in a rehabilitation training scenario, and the knowledge graph shows that "strength training" is associated with "heart rate monitoring," the system can automatically enable the heart rate monitoring function.

[0063] Ultimately, spatiotemporal alignment ensures that scene interaction commands are triggered at the appropriate time and location. For example, in a virtual gym environment, if a user approaches the "dumbbell training area," the system can use knowledge graph reasoning to trigger the command "suggest upper limb muscle strengthening training," and play a demonstration video of the exercise after the user enters the specific location.

[0064] By leveraging domain knowledge graph-based interaction parsing, intelligent and personalized immersive interactive experiences can be achieved, enabling virtual environments to adapt to user behavior patterns and provide precise real-time feedback. It can dynamically adjust interaction logic and optimize scene responses based on user status, thereby enhancing the naturalness and intelligence of the interaction.

[0065] S50, integrate the historical behavioral feature vectors in the user profile with the real-time state vectors in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0066] In this embodiment, the core objective of generating comprehensive decision parameters by integrating historical behavioral feature vectors from user profiles with real-time state vectors from the mapping relationship between physical and mental states is to accurately model the user's physical and mental state and behavioral preferences. By integrating multimodal data, personalized decisions are generated to improve the adaptability and intelligence of interactive scenarios.

[0067] User profiles are personalized feature sets built upon historical data, containing information such as users' long-term behavioral patterns, preferences, interaction habits, and physiological trends. This data can be obtained from multiple sources, including past training records, psychological state monitoring data, and scenario-based interaction behavior data, and is stored in a structured manner using high-dimensional feature vectors. For example, in a health management scenario, a user profile might include information such as exercise frequency, sleep quality, mood fluctuation trends, and heart rate variability curves.

[0068] The real-time state vector in the mind-body state mapping relationship is a dynamic feature set obtained by modeling data such as physiological indicators, movement state, and psychological state at the current moment. Compared with long-term behavioral patterns in user profiles, the real-time state vector can reflect the user's immediate state changes, such as the user's current stress level, fatigue level, and concentration level.

[0069] During data fusion, the first step is to standardize and impute missing values ​​in the historical behavioral feature vectors of user profiles to ensure data consistency and integrity. For example, if some features in a user profile have missing values, interpolation methods or collaborative imputation strategies based on similar users can be used to complete them. Furthermore, to eliminate scale differences between different data sources, Z-score standardization or Min-Max normalization methods can be used to ensure that all feature values ​​are distributed within the same numerical range.

[0070] Secondly, the real-time state vector needs to be matched with historical behavioral feature vectors to ensure data alignment. During processing, a time window sliding method can be used to dynamically smooth the real-time state vector, generating a time-weighted feature vector to enhance the temporal consistency of the data. For example, in sports training scenarios, the moving average of the real-time state vector can be calculated based on physiological data from the past 30 seconds to reduce noise interference caused by short-term data fluctuations.

[0071] Then, a multimodal fusion network based on an attention mechanism is employed to deeply fuse the user profile's historical behavioral feature vectors with its real-time state vectors. This multimodal fusion network can learn the weight relationships between different feature modalities; for example, it pays more attention to physiological data when under high stress, and more to historical behavioral patterns when the state is stable. For instance, in meditation training, if the system detects that the user's historical training record shows "long-term high stress," while the current real-time state vector shows "rapid breathing rate and elevated heart rate," the system can increase the duration of the meditation guidance content or decrease the training intensity to adapt to the user's current state.

[0072] The integrated decision parameters after fusion include:

[0073] Match score: Used to measure the similarity between a user profile and the current state.

[0074] Priority tags: Used to sort different training schemes and guidance content, ensuring that high-priority content is recommended first.

[0075] Output confidence score: Used to evaluate the credibility of the current recommendation. If the confidence score is low, multiple interaction options can be provided to enhance the system's adaptability.

[0076] By integrating user profiles and real-time state vectors, more precise personalized decision-making can be achieved, enabling the system to dynamically adjust interactive content based on users' long-term behavioral patterns and immediate states. This allows for the optimization of interaction strategies across various application scenarios. For example, in sports training, training intensity can be adjusted based on the user's physical fatigue level; in mental health management, targeted relaxation training can be provided based on the user's emotional state; and in intelligent investment advisory systems, investment recommendation strategies can be optimized based on the user's cognitive load, reducing the risk of impulsive decisions under high pressure.

[0077] S60, Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, select personalized guidance content from the guidance content library;

[0078] In this embodiment, the core objective of selecting personalized guidance content from the guidance content library based on the feature similarity between comprehensive decision parameters and scene interaction commands is to provide content recommendations that highly match the user's state and interaction needs, ensuring that users receive guidance information most suitable for their current state and behavioral patterns in an immersive interactive environment. Traditional content recommendation methods typically rely on predefined rules or static matching based on historical behavior, making it difficult to dynamically adapt to the user's real-time physical and mental state. This method, however, achieves real-time adaptation of personalized recommendations through the calculation of feature similarity between comprehensive decision parameters and scene interaction commands.

[0079] Comprehensive decision parameters are decision data generated by fusing historical behavioral feature vectors from user profiles with real-time state vectors from the mapping relationship between physical and mental states. These parameters include metrics such as matching score, priority label, and output confidence. They quantify a user's long-term preferences and immediate state, enabling the recommendation system to adjust recommended content based on the user's current physiological and psychological state. For example, in a mental health training scenario, comprehensive decision parameters could include information such as the user's emotional stability index, anxiety level, and meditation training preferences to support personalized recommendations.

[0080] Scene interaction instructions are dynamic interaction data parsed based on domain knowledge graphs, reflecting the interaction rules, triggering conditions, and environmental feedback mechanisms in the current scene. For example, in a virtual fitness training environment, scene interaction instructions might include "provide relaxation training suggestions when the user reaches a specific heart rate threshold" or "provide correction prompts when the user performs an incorrect action."

[0081] Feature similarity calculation measures the degree of matching between comprehensive decision parameters and scene interaction commands. It typically employs cosine similarity, vector distance calculation, or deep learning-based matching models. For example, if the matching score in the comprehensive decision parameters is highly similar to the action triggering condition in the scene interaction command, the system can prioritize recommending guidance content related to that interaction rule.

[0082] During the calculation process, the feature vectors of the comprehensive decision parameters and scene interaction commands first need to be vectorized to ensure that they are in the same feature space. For example, word vector encoding, feature embedding, and normalization can be used to convert data from different modalities into numerical representations of the same dimension.

[0083] Secondly, a feature similarity-based matching strategy is adopted to calculate the similarity score between the comprehensive decision parameter vector and the scene interaction command vector.

[0084] Finally, based on the similarity calculation results, the content with the highest matching degree in the guidance content library is selected, and the priority of the recommended content is adjusted in conjunction with the output confidence score. For example, if a training scheme has a high matching degree and a high confidence score for user historical preferences, then that scheme should be given priority for recommendation.

[0085] By calculating the feature similarity between comprehensive decision parameters and scene interaction commands, accurate and highly personalized content recommendations can be achieved. This allows the system to dynamically adapt to the user's real-time state, enhancing the intelligence level of the interactive experience. It can adaptively adjust recommendation strategies to provide the most suitable guidance content in different contexts.

[0086] S70, output the personalized guidance content.

[0087] In this embodiment, the core objective of outputting personalized guidance content is to provide highly tailored guidance information based on the user's individual characteristics and real-time status, thereby enhancing the accuracy and effectiveness of the immersive interactive experience. Personalized guidance content includes not only multimodal output formats such as text messages, voice commands, video demonstrations, and real-time feedback, but also a dynamic adjustment mechanism to ensure that the output guidance content remains consistent with the user's physical and mental state, behavioral patterns, and interactive environment.

[0088] Personalized guidance content is the optimal guidance solution selected based on the aforementioned comprehensive decision parameters and the feature similarity of scene interaction commands. Its core features include:

[0089] Adaptive content generation: The guidance content should be dynamically adjusted according to the user's current state. For example, in virtual rehabilitation training, the system can adjust the pace and intensity of the training guidance according to the user's fatigue level.

[0090] Multimodal information fusion: Personalized guidance can utilize text, voice, video, and interactive animation to enhance user understanding. For example, in a smart fitness system, voice prompts combined with real-time video demonstrations can be provided to ensure users correctly execute training movements.

[0091] Real-time interactive optimization: After the guidance content is output, the system can adjust based on user feedback and physiological data. For example, in meditation training, if the system detects that the user's heart rate is decreasing slowly, it can increase the duration of deep breathing guidance to enhance the relaxation effect.

[0092] During the content output process, it is necessary to ensure intelligent adaptation of output format, feedback method, and interaction logic:

[0093] Intelligent output format adaptation: Different users may have different ways of receiving information, and the system can adjust the output method based on the user's historical preferences. For example, some users prefer text descriptions, while others prefer video demonstrations, and the system can automatically match the best output mode.

[0094] Intelligent optimization of feedback methods: When users perform guided content, the system can adjust the feedback strategy in real time based on physiological sensors and behavioral analysis. For example, in a virtual fitness environment, when a user does not complete a certain movement correctly, the system can issue an immediate voice prompt and provide real-time corrective video to guide the user to make adjustments.

[0095] Personalized interaction logic: The system can adjust the interaction process according to the user's real-time status. For example, in emotion management training, if the system detects that the user has entered a high-stress state, it can automatically reduce complex training instructions and instead provide simpler and more intuitive relaxation exercises.

[0096] Example Description: In the healthcare field, to address the personalized rehabilitation guidance needs of patients undergoing rehabilitation training, a multimodal perception-based immersive behavioral guidance method is employed to provide precise recommendations for rehabilitation training content. The system first collects patients' physiological data, including heart rate, blood oxygen saturation, and skin conductivity, using wearable smart devices (such as smart bracelets and smart insoles). Simultaneously, an inertial measurement unit (IMU) monitors the patient's three-dimensional motion trajectory, such as gait rhythm, joint angular velocity, and posture angles. Furthermore, optical motion capture devices record the patient's limb range of motion and posture changes, forming a multimodal body dataset. This data undergoes timestamp alignment and signal filtering to eliminate time errors between acquisition devices and remove anomalous signals. After acquiring the patient's multimodal body data, the system inputs it into a pre-trained mind-body association model to extract the mapping relationship between the patient's mind-body state. This model, based on historical training data, learns the correlation features between physiological indicators, movement patterns, and psychological states through a gated recurrent unit (GRU). The model generates a state mapping matrix at the output layer, including stress index and focus weights, to assess the patient's current physiological and psychological state. At the same time, the accuracy of the training movements is assessed by calculating the movement standardization evaluation coefficient based on the patient's movement trajectory deviation.

[0097] To enhance the immersive experience of rehabilitation training, the system uses a 3D scene model library to create virtual interactive environments tailored to patients' rehabilitation training needs. The system adjusts the ambient lighting intensity and spatial layout of the training scene based on the patient's physical and mental state. For example, if the patient's stress level is high, the system will reduce stimulating factors in the environment, such as lowering lighting intensity and reducing background noise; if focus is a high priority, the system will generate clearer path markers in the scene to guide the patient through training tasks. Furthermore, the virtual instructor synchronizes with the patient's movement rhythm, providing real-time motion demonstrations and voice guidance to help patients improve the effectiveness of their rehabilitation training. A physics engine constraint ensures that virtual objects in the training scene have realistic motion feedback, such as simulating the patient's interaction with assistive training equipment in the virtual environment.

[0098] During training, the system utilizes a domain knowledge graph to parse scene interaction commands and adjusts the interaction logic based on the patient's current behavioral patterns. The domain knowledge graph predefines training steps, rules for using rehabilitation equipment, and the correlation information between the patient's state and movement patterns. For example, if the patient's knee flexion angle does not reach the expected range, the system prompts the patient to adjust their movements via voice; if the patient performs a training movement for a duration exceeding a set threshold, the system dynamically adjusts the training pace and provides appropriate rest prompts.

[0099] Personalized training guidance is generated based on the patient's historical training data and current physical and mental state. The system analyzes the patient's performance in different training tasks and calculates the stress index and focus weight of the real-time state vector to adjust the recommendation weight of personalized training suggestions. Through a multimodal fusion model based on attention mechanisms, the system matches the patient's training behavior characteristics with the rehabilitation guidance content library, calculates a matching score, and selects training content with higher priority. For example, if the patient has a high matching score and a high focus weight, the system may recommend more challenging training content; conversely, it may recommend a gentler training mode.

[0100] Ultimately, based on the feature similarity between comprehensive decision parameters and scene interaction commands, the system selects personalized guidance content suitable for the patient's current state from the training content library. The matching score and scene interaction feature similarity score are calculated using linear weighting to determine the priority of personalized training content, and the recommendation threshold is dynamically adjusted to adapt to the training needs of patients in different states. If certain training tasks conflict with scene interaction logic (e.g., the patient cannot currently use certain rehabilitation equipment), the system will automatically remove conflicting content and output the final personalized training guidance, ensuring the safety and efficiency of the rehabilitation training process.

[0101] In the fintech field, this technology can be used for intelligent investment advisory, risk management, and personalized financial decision support systems. The system first uses multimodal perception technology to collect users' physiological indicators (such as heart rate, pupil dilation, and skin conductivity) to assess their emotional fluctuations. Simultaneously, it combines this data with the user's financial behavior data (such as transaction records, portfolio adjustment frequency, and risk preference test results) to generate a multimodal financial behavior dataset.

[0102] These data are processed by a pre-trained financial behavior analysis model to generate a risk tolerance assessment result for the user. This model calculates a stress index and focus weight based on historical trading data and market volatility, combined with the user's physiological feedback (such as emotional reactions caused by market fluctuations). If a user exhibits a high stress index during periods of high market volatility, the system can reduce the recommendation weight of high-risk investments to prevent impulsive investments when the user is emotionally unstable. If the user's focus weight is high, more complex investment strategies can be recommended, such as multi-factor portfolio optimization or quantitative trading strategies.

[0103] To enhance the investment experience, the system constructs a virtual financial analysis environment based on a 3D market scenario model library. For example, in the virtual market lobby, users can view the historical return curves of different investment portfolios in real time and adjust investment ratios through interactive data visualization tools. If a user's stress level is high, the system can adjust the visual elements of the scene, such as reducing the display frequency of market volatility information or using a softer color scheme to reduce visual burden. If a user's focus level is high, the system can generate key analysis areas in the scene, such as highlighting key economic indicators that affect investment decisions.

[0104] Personalized financial decisions are generated based on users' historical investment behavior and the current market environment. The system uses domain knowledge graphs to analyze market interaction instructions and adjusts recommendation strategies according to users' current investment behavior patterns. For example, if a user's trading frequency increases and their emotions fluctuate significantly, the system may suggest that the user rebalance their portfolio to reduce risk exposure; if a user increases their investment in high-risk assets, the system may trigger risk warnings and provide alternative investment suggestions based on user behavior data.

[0105] Ultimately, based on comprehensive decision-making parameters and market interaction logic, the system selects personalized investment recommendations from the financial product database that match the user's current situation. The matching score and market interaction feature similarity score are calculated using linear weighting to ensure that the recommended investment plan is both in line with the user's risk tolerance and adaptable to the current market environment.

[0106] In the field of mental health care, it can be applied to personalized meditation, relaxation guidance, and emotion regulation support systems. The system first collects users' physiological indicators (such as heart rate variability, respiratory rate, and electroencephalogram data) through wearable devices, and combines them with limb movement data (such as sitting posture, walking rhythm, and body movement amplitude) to form a multimodal mental health care behavior dataset.

[0107] These data are processed by a pre-trained mental state model to generate a mapping relationship between the user's mind-body state. This model learns the physiological and psychological characteristics of different mental states through long-term accumulation of conditioning data. For example, the system can identify the user's relaxation depth based on EEG data and calculate the user's focus weight by combining heart rate variability. If the user's focus is low, the system may suggest simpler guided breathing exercises, such as diaphragmatic breathing or rhythmic breathing; if the user's stress index is high, the system may guide the user into a more soothing conditioning mode, such as mindful breathing, progressive muscle relaxation, or mind-body scan exercises.

[0108] Based on a 3D psychological wellness scene model library, the system can construct virtual relaxation spaces, such as tranquil lakesides, forest trails, or quiet rest rooms, to enhance the immersive psychological wellness experience. The system can dynamically adjust the scene atmosphere according to the user's physical and mental state; for example, playing soft background music and reducing visual stimulation when the user is anxious; and enhancing visual focus when concentration is high, such as guiding the user to gaze at a virtual candle flame or a slowly moving point of light. A virtual facilitator can synchronize with the user's breathing rhythm, providing real-time voice guidance, such as "Take a deep breath and exhale slowly" or "Focus your attention on each breath," ensuring the user can deeply experience the psychological wellness process.

[0109] Ultimately, based on comprehensive decision-making parameters and the logic of the psychological well-being scenario, the system selects personalized guidance content suitable for the user's current state from the psychological well-being guidance library. For example, when the system detects that the user has entered a deep relaxation state, it may reduce the interference of voice prompts and adjust the rhythm of the guidance content. Through intelligent psychological well-being guidance, users can more efficiently enter a deep relaxation state and experience diverse psychological well-being methods in a virtual environment.

[0110] By providing intelligent, personalized guidance, the system ensures the accuracy of the interactive experience, enabling users to receive feedback that best suits their current state. It can dynamically adjust the guidance content based on the user's real-time status, achieving a more interactive and personalized user experience.

[0111] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health finance, financial technology, and mental health care. It discloses a guidance method based on multimodal perception, comprising: collecting users' physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body dataset; inputting a pre-trained mind-body association model to generate a mind-body state mapping relationship; generating an immersive interactive scene containing environmental simulation elements based on a three-dimensional scene model library and dynamic attention parameters in the mind-body state mapping relationship; parsing scene interaction commands in the immersive interactive scene according to entity association relationships in a domain knowledge graph; fusing historical behavioral feature vectors from user profiles with real-time state vectors in the mind-body state mapping relationship to generate comprehensive decision parameters; and selecting personalized guidance content from a guidance content library based on the feature similarity between the comprehensive decision parameters and scene interaction commands, and outputting the guidance content. This invention achieves accurate perception of users' physiological states and behavioral characteristics through multimodal data collection and mind-body state modeling, and improves the accuracy of personalized guidance by combining immersive scene interaction and knowledge graph parsing. The generation and matching optimization of comprehensive decision parameters enable interactive content to adapt to different users' behavioral patterns, improving the personalization and immersiveness of intelligent guidance.

[0112] In one embodiment, S10 includes:

[0113] S101 collects the user's skin conductivity data and blood oxygen saturation data through wearable smart devices to generate a first physiological index subset;

[0114] S102, through the inertial measurement unit, collects the acceleration data of the user's torso and the angular velocity data of the limb joints, and generates a subset of three-dimensional motion trajectories;

[0115] S103: Collects facial expression change data and limb movement amplitude data of the user through optical motion capture equipment, and generates a subset of motion features;

[0116] S104, perform timestamp alignment processing on the data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset;

[0117] S105 performs signal filtering on the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset after timestamp alignment processing to generate a standardized multimodal body dataset.

[0118] In this embodiment, the purpose of collecting users' physiological index data and three-dimensional motion trajectory data is to construct a multimodal body dataset to accurately obtain users' physiological state, movement patterns, and behavioral characteristics. The fusion of different types of physiological signals and motion features in multimodal data helps improve the accuracy of individual state identification and provides data support for subsequent physical and mental state modeling and personalized guidance.

[0119] Electrodermal activity (EDA) and blood oxygen saturation (SpO2) are important indicators for measuring the state of the human autonomic nervous system. EDA reflects the activity level of the sympathetic nervous system and is closely related to mood fluctuations, stress levels, and anxiety levels. Blood oxygen saturation monitors blood oxygen levels and reflects a user's respiratory status, fatigue level, etc. Wearable smart devices (such as smart bracelets and smart rings) are typically equipped with EDA sensors and optical blood oxygen sensors, enabling them to collect these physiological data in real time.

[0120] During data acquisition, a strategy combining continuous and intermittent measurements can be employed. For example, data can be collected every 30 seconds in a static state; during exercise or periods of intense emotional fluctuation, the sampling frequency can be increased to once per second. Furthermore, to ensure data stability, an adaptive data fusion algorithm can be used to correct for environmental factors such as temperature and humidity, preventing external interference from affecting skin conductivity.

[0121] An inertial measurement unit (IMU) typically contains sensors such as accelerometers, gyroscopes, and magnetometers, which can be used to detect a user's movement patterns. Trunk acceleration data reflects the user's overall movement trends, such as gait, walking speed, and exercise intensity. Angular velocity data from the limb joints can be used to analyze the user's hand gestures, limb coordination, and movement habits.

[0122] During data acquisition, IMU sensors need to be placed at multiple locations (such as the wrist, ankle, and waist) to ensure the integrity of the 3D trajectory. After data acquisition, Kalman filtering or complementary filtering can be used for data fusion to reduce noise and improve accuracy. Furthermore, six-degree-of-freedom (6-DoF) or nine-degree-of-freedom (9-DoF) attitude estimation algorithms can be employed to calculate the user's attitude angles, rotation angles, and other information to achieve more accurate motion trajectory modeling.

[0123] Optical motion capture devices (such as RGB cameras, depth cameras, and infrared cameras) can recognize a user's facial expressions, body movements, and postures. Facial expression data includes micro-expression features of the eyebrows, eyes, and mouth, which can be used to infer the user's emotional state, such as anxiety, pleasure, or fatigue. Body movement amplitude data can be used to analyze information such as the user's gesture interactions, posture adjustments, and motor coordination.

[0124] During data acquisition, a multi-view camera scheme is required to ensure accurate recognition of both facial and limb movements. For facial expression recognition, deep learning algorithms such as OpenFace and MediaPipe Face Mesh can be used to extract facial key points, and user emotional states can be classified based on Support Vector Machine (SVM) or Convolutional Neural Network (CNN). For limb movement recognition, skeletal tracking algorithms such as OpenPose and MediaPipe Pose can be used to perform 3D reconstruction of the user's limb key points and analyze features such as movement amplitude and smoothness.

[0125] Because different sensors have different sampling frequencies and data formats, the raw data often has time synchronization discrepancies, requiring timestamp alignment. Time alignment methods include:

[0126] Interpolation-based methods: For low sampling rate data (such as EDA, SpO2), linear interpolation or spline interpolation can be used to align it with high-frequency data (such as IMU data).

[0127] The time window-based method sets a fixed time window (e.g., 100ms) and maps all sensor data to the same time dimension to ensure synchronized data processing.

[0128] The method based on dynamic time warping (DTW) is used to align different modal data in the case of nonlinear time bias in time series data, so that they are consistent in the time dimension.

[0129] Raw sensor data may contain noise and outliers, therefore signal filtering is necessary to improve data quality. Signal filtering can include the following methods:

[0130] Low-pass filtering (Butterworth filter, mean filter): Removes high-frequency noise and improves signal stability. For example, in IMU data processing, it removes rapidly fluctuating high-frequency signals and improves the continuity of motion trajectories.

[0131] Median filtering: Used to remove short-term outlier data points. For example, in EDA signal processing, it can filter out transient spike noise to ensure the stability of skin conductivity data.

[0132] Autoregressive moving average (ARMA) filtering: In the process of multimodal data fusion, it can predict the current value based on the trend of historical data to improve the robustness of the data.

[0133] Standardized multimodal body datasets serve as the foundational data source for subsequent mind-body state mapping and personalized behavior guidance, and can be used in scenarios such as user state analysis and personalized training recommendations.

[0134] This embodiment, by integrating physiological signals, motion data, and movement characteristics, can more accurately depict the user's physical and mental state, making personalized behavioral guidance more scientific and efficient. It can play a role in areas such as sports training, health management, and psychological intervention, improving the accuracy of user status monitoring and supporting the intelligent optimization of personalized guidance programs.

[0135] In one embodiment, S20 above includes:

[0136] S201, the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset in the multimodal body dataset are divided into temporal segments according to a preset time window;

[0137] S202, the time sequence segment is input into the mind-body association model constructed based on the gated loop unit, and the cross-modal association features between the first physiological index subset, the three-dimensional motion trajectory subset and the action feature subset are extracted through the mind-body association model;

[0138] S203, Based on the cross-modal association features, generate a mind-body state mapping matrix containing dynamic attention parameters and joint angle thresholds in the output layer of the mind-body association model;

[0139] S204, Based on the stress index and focus weight in the dynamic attention parameters, determine the user's physical and mental state level;

[0140] S205, Based on the deviation between the joint angle threshold and the motion trajectory data in the subset of the three-dimensional motion trajectory, generate a motion standardization evaluation coefficient;

[0141] S206, integrate the physical and mental state level and the action standardization evaluation coefficient into a physical and mental state mapping relationship that includes real-time state vectors.

[0142] In this embodiment, the core objective of inputting a multimodal body dataset into a pre-trained mind-body association model and generating mind-body state mapping relationships is to establish the correlation between a user's physiological state, movement patterns, and behavioral characteristics to achieve personalized mind-body state assessment. Traditional human state analysis typically relies on data from only one modality, such as physiological signals or movement trajectories, while ignoring the potential relationships between multimodal information. This method extracts cross-modal correlation features through a deep learning model and combines it with time series analysis to generate mind-body state mapping relationships, thereby improving the accuracy and real-time performance of state assessment.

[0143] Multimodal data acquisition typically involves continuous time series, but different modalities may have different sampling frequencies. Therefore, time windowing is necessary to divide the data into equal-length time series segments, ensuring that the data from each modality are aligned and analyzed on the same time dimension. Time windowing methods can include:

[0144] Fixed time window: A fixed window length (e.g., 1 second, 5 seconds) can be set to ensure that the data is input into the model in a fixed format.

[0145] Sliding time window: The sliding window method (such as a 0.5-second step sliding 5-second window) is used to enhance the continuity of time series features and improve the ability to capture short-term dynamic changes.

[0146] Adaptive window: For specific applications, such as emotion recognition, the window length can be dynamically adjusted based on the rate of change of physiological data. For example, when skin conductivity changes rapidly, the window length is shortened to improve temporal resolution.

[0147] Multimodal data exhibits time dependence and cross-modal correlation; therefore, GRU networks are used to model time series data and extract cross-modal features. GRU is a variant of recurrent neural networks (RNN) suitable for long-sequence learning, effectively capturing the long-term dependencies of time series data while avoiding the gradient vanishing problem of traditional RNNs.

[0148] GRU structure: Includes update gate and reset gate, used to control the degree of retention of past information and integration with current input, improving data correlation across time steps.

[0149] Cross-modal feature extraction: The model input includes a subset of primary physiological indicators (such as EDA and SpO2), a subset of three-dimensional motion trajectories (such as acceleration and angular velocity), and a subset of motion features (such as facial expressions and body posture). GRU learns the temporal dependencies of different modal data through multi-layer encoding. For example, it analyzes the synchronicity between heart rate changes and exercise rhythm to infer the user's exercise fatigue state.

[0150] At the output of GRU, an attention mechanism is introduced to dynamically adjust the weights of different modalities in decision-making and generate a mind-body state mapping matrix.

[0151] Attention parameters: The attention weights for each modality are calculated based on the hidden states of GRU. For example, in a meditation training scenario, the system may pay more attention to heart rate and skin conductivity, while in a sports training scenario, it may pay more attention to the three-dimensional motion trajectory.

[0152] Joint angle threshold: For motion state assessment, safe thresholds for joint angles are stored in the physical and mental state mapping matrix (such as triggering a risk warning when the knee flexion angle exceeds a certain value) for motion standardization assessment.

[0153] Users' physical and mental state can be assessed using a stress index and focus weight.

[0154] Stress Index: Calculated based on physiological signals such as skin conductivity and heart rate variability, reflecting the user's level of tension. For example, elevated skin conductivity may indicate increased stress, while decreased heart rate variability may indicate anxiety.

[0155] Attention weighting: This measure assesses a user's level of concentration using data such as eye tracking and electroencephalography (EEG). For example, in a learning context, stable eye movements and enhanced alpha waves on an EEG indicate a high level of focus.

[0156] Physical and mental state classification: Multi-level classification models (such as SVM, random forest) or clustering algorithms (such as K-means) can be used to classify user states into state labels such as "high stress and low focus" and "low stress and high focus".

[0157] Movement standardization assessment is used to analyze the quality of a user's movement execution and ensure that training movements meet standards.

[0158] Calculate the deviation value: Compare the actual movement trajectory with the standard trajectory to calculate the error in the joint angle. For example, in yoga training, if the user's knee joint angle differs from the standard posture by 10°, the deviation value is 10.

[0159] Normative scoring: A normative score is calculated based on the deviation value. For example, Euclidean distance or dynamic time warping (DTW) can be used to calculate the similarity of the motion trajectory. The scoring range can be set from 0 to 100, where 100 represents perfect conformity to the standard movement.

[0160] The final output is a mapping relationship between mind-body states, used for personalized behavioral guidance. This relationship includes:

[0161] Real-time state vector: includes indicators such as stress index, focus, fatigue level, and motor coordination, serving as a quantitative representation of the user's current state.

[0162] Adaptive optimization: Adjusts training intensity and prompts based on real-time status. For example, if a user's stress level is too high, the system can automatically reduce the training load or provide relaxation suggestions.

[0163] This embodiment, through multimodal data fusion and time-series modeling, can more accurately assess a user's physical and mental state, possesses cross-modal feature extraction capabilities, and can improve state recognition accuracy. Furthermore, through attention mechanisms and personalized state mapping, the evaluation criteria can be dynamically adjusted to achieve more intelligent physical and mental state modeling.

[0164] In one embodiment, S30 includes:

[0165] S301, determine the user's current behavior pattern based on the historical behavior feature vector in the user profile and the real-time state vector in the mapping relationship between the physical and mental states;

[0166] S302, determine a basic three-dimensional scene model that matches the user's current behavior pattern from the three-dimensional scene model library;

[0167] S303, adjust the ambient light and shadow intensity parameters and weather simulation parameters in the basic three-dimensional scene model according to the stress index in the dynamic attention parameters;

[0168] S304, Based on the focus weight in the dynamic attention parameters, generate path markers and synchronized virtual tutor actions in the basic three-dimensional scene model;

[0169] S305, integrate the ambient light and shadow intensity parameters, weather simulation parameters, path markers and synchronized virtual tutor actions to generate an immersive interactive scene;

[0170] S306, apply motion trajectory constraints driven by a physics engine to the virtual objects in the immersive interactive scene.

[0171] In this embodiment, during the construction of the immersive interactive environment, in order to enhance the user's sense of immersion and interactive experience, the system needs to combine a 3D scene model library and adaptively adjust the attention parameters based on the mind-body state mapping relationship. Different users have different behavioral patterns, mind-body states, and levels of focus. Therefore, the interactive scene needs to be adaptive to match the user's current needs, optimize visual, auditory, and interactive feedback, and ensure the natural and smooth interactive experience.

[0172] In virtual environments, a user's interaction patterns are closely related to their long-term behavioral characteristics and current physical and mental state. User profiles contain long-term accumulated behavioral data, such as historical training records, preference choices, movement patterns, and scene usage frequency, while the physical and mental state mapping provides current physiological indicators (such as heart rate and skin conductivity) and movement status (such as limb stability and gait characteristics). To determine a user's current behavioral patterns, it is necessary to comprehensively analyze these two information sources. A behavior prediction model based on Long Short-Term Memory (LSTM) networks can be used, inputting historical behavioral feature vectors and real-time state vectors into the network, and predicting user behavior trends through time series analysis. For example, if a user tends to choose relaxing scenes when under high stress and with low concentration, the system can prioritize recommending environments that suit this state in subsequent interactions.

[0173] The 3D scene model library stores multiple virtual scene templates, such as forests, beaches, snow-capped mountains, and quiet meditation rooms, each with a different atmosphere and suitability. Based on the determined user behavior patterns, the most suitable scene model needs to be selected. For example, under low stress and high focus, clear, high-contrast scenes, such as modern office spaces, are suitable; while under high stress, soft, soothing scenes, such as natural landscapes and yoga studios, are more appropriate. A cosine similarity matching algorithm can be used to compare the user behavior pattern vector with the scene feature vector, calculate the suitability score for each scene, and thus select the optimal matching scene model.

[0174] A user's stress level affects their tolerance to ambient lighting and weather; therefore, the base scene model needs to be adjusted based on the stress level. If the user's stress level is high, the system can reduce the intensity of ambient lighting to make the scene light softer and avoid overstimulation. For example, the proportion of blue light can be reduced, warm-toned light sources can be increased, and the brightness variations of the scene can be adjusted through real-time lighting rendering techniques (such as physically based rendering, PBR) to make it more comfortable. Similarly, weather simulation parameters can also be adjusted according to changes in the stress level. For example, under high stress, the scene can add dynamic elements such as breezes and flowing water to enhance the relaxation effect, while under low stress, it can simulate morning sunlight and high-contrast lighting effects to improve focus. These adjustments can be implemented based on a real-time environment rendering engine (such as Unity HDRP or Unreal Engine Lumen) to make the changes in lighting more natural.

[0175] Attention level weighting determines the user's level of concentration, influencing their need for visual guidance and motion synchronization. In a high-attention state, path markings should be simplified to reduce distractions, such as using low-opacity ground markings or providing guidance only at key nodes. In a low-attention state, path guidance needs to be enhanced, for example, by using flowing animations, increased brightness, or color contrast to make the path more intuitive, ensuring the user can clearly perceive the route. Path marking generation can employ A* search algorithms or Bézier curve-based path optimization algorithms to ensure the path's rationality and naturalness. Furthermore, the virtual instructor's motion synchronization also needs to be dynamically adjusted based on the user's attention level. In a high-attention state, the instructor's actions should minimize detailed intervention, providing only key guidance, while in a low-attention state, the instructor needs to add features such as body magnification, voice prompts, and posture breakdown to enhance the comprehensibility of the interaction. Computer vision posture analysis (such as OpenPose or MediaPipe) can be used to evaluate user actions in real time and dynamically adjust the virtual instructor's feedback based on deviations.

[0176] After adjusting all environmental and interactive elements, these parameters need to be integrated to ultimately generate a personalized, immersive interactive scene for the user. The key to this step is optimizing the scene rendering engine to ensure seamless integration of different adjustments. For example, ambient lighting adjustments should not affect the visibility of path markers, and weather simulations should not interfere with the visibility of the virtual tutor. To achieve this, a hierarchical rendering strategy can be used, dividing different types of interactive elements into different rendering layers, allowing path markers, virtual tutors, and ambient lighting to be adjusted independently. Furthermore, physical interaction optimizations (such as those based on Unity Physics or Havok Physics) can be utilized to ensure that objects in the scene respond correctly to user interactions; for example, environmental elements (such as leaves or lake ripples) should also provide adaptive feedback as the user moves along path markers.

[0177] To enhance the realism of the scene, the movement of all virtual objects must conform to the laws of physics. For example, in a deep relaxation environment, candlelight sways gently with the virtual airflow, creating a soft and natural atmosphere; in a tranquil forest scene, leaves sway softly in the wind, and sunlight filters through them, creating dappled shadows; by a lake or stream, the water surface ripples and undulates in response to the user's footsteps or contact with virtual objects. Such physical effects not only enhance visual realism but also strengthen the user's immersive experience through environmental feedback, allowing them to feel more natural and relaxed in the virtual relaxation environment. These effects can be achieved using rigid body physics engines (such as PhysX, Bullet, or Unreal Engine Chaos). Physics engines calculate the trajectories of objects in the environment, ensuring they conform to natural laws such as gravity, inertia, and air resistance. For example, when a user performs Tai Chi movements, the system can analyze the force of the user's push and calculate the feedback from air resistance, causing leaves or airflow in the virtual scene to exhibit corresponding movement effects. In addition, to ensure the smoothness of the motion trajectory, Bézier curves or Kalman filters can be used to optimize the trajectory, making the interactive experience smoother and more natural.

[0178] This embodiment combines a 3D scene model library with a mapping relationship between mind and body states, enabling the system to dynamically adapt to the user's physiological and psychological states and generate personalized, immersive interactive scenes. It can adjust ambient lighting and weather based on stress levels, providing a more comfortable interactive experience under high stress and clearer visual feedback under low stress. Furthermore, through adaptive adjustments to path markings and virtual tutors, the system can provide more precise interactive guidance, improving user learning efficiency and task completion rates in different states. Simultaneously, the introduction of a physics engine ensures that objects in the virtual scene conform to real physical laws, enhancing immersion and interactive realism, allowing users to more naturally integrate into the virtual world.

[0179] In one embodiment, S40 includes:

[0180] S401, Extract the predefined entity relationship graph structure from the domain knowledge graph, the entity relationship graph structure includes scene element entities, behavior pattern entities, and association attributes between scene element entities and behavior pattern entities;

[0181] S402, extract the current scene element entity set from the immersive interactive scene, the current scene element entity set includes virtual tutor entity, path marker entity and interactive prop entity;

[0182] S403, Based on the semantic similarity matching result between the user's current behavior pattern and the behavior pattern entity in the domain knowledge graph, determine the target interaction pattern entity;

[0183] S404, Based on the association attribute between the scene element entity and the behavior pattern entity, generate scene interaction instruction logic parameters bound to the target interaction pattern entity;

[0184] S405, perform timestamp interpolation alignment and spatial coordinate mapping processing on the scene interaction instruction logic parameters and the current scene element entity set to generate an executable scene interaction instruction queue.

[0185] In this embodiment, in an immersive interactive environment, to ensure users can smoothly interact with the virtual scene, it is necessary to parse the interaction commands in the scene based on a domain knowledge graph. The domain knowledge graph contains rich structured information that can describe the elements, behavioral patterns, and relationships between them in the scene, thereby enabling intelligent and dynamic interaction logic. The key to parsing interaction commands lies in correctly matching the user's current behavioral pattern, combining it with the entity relationships in the scene, and dynamically generating reasonable interaction commands to ensure that the execution of the commands not only conforms to the user's habits but also enhances the immersive experience.

[0186] A knowledge graph is a graph-based data representation method that describes the relationships between different entities. The core of a domain knowledge graph consists of scene element entities, behavior pattern entities, association attributes, and constraint rules. Scene element entities refer to key interactive elements in a virtual scene, such as virtual tutors, path markers, and interactive props. Behavior pattern entities describe the user's interaction methods in different states, such as relaxation, walking, and meditation. Association attributes define the interaction relationships between scene elements and behavior patterns; for example, when a user is performing a walking task, path markers should be highlighted, while in a relaxed state, path markers should be faded. To load the knowledge graph structure, graph databases (such as Neo4j, DGL, and GraphDB) can be used to store and retrieve data, and interaction relationships can be obtained through query languages ​​based on SPARQL or Cypher, providing foundational data for subsequent parsing.

[0187] The elements in the scene change according to the user's interaction state, therefore it is necessary to extract the set of entities existing in the current scene in real time. The extracted entities include, but are not limited to:

[0188] Virtual tutor entity: An intelligent agent used to provide guidance through voice, actions, etc.

[0189] Path marker entity: used to indicate the user's route in the interactive environment;

[0190] Interactive props and entities: such as handheld devices, virtual items, etc., can directly interact with users.

[0191] These entities can be extracted using computer vision-based object detection technologies (such as YOLO and Faster R-CNN), combined with the scene management functions of real-time rendering engines (such as Unity and Unreal Engine) to determine the active interactive objects in the current frame and construct a real-time scene element list.

[0192] User behavior patterns are the core factor determining interaction logic, and different user behavior patterns influence their needs for interaction methods. For example, in deep relaxation mode, users require smoother, slower interactions, such as soft background music and slow breathing guidance; while in focused training mode, users may need more efficient and precise instruction feedback, such as clear action instructions and immediate voice prompts. To determine the appropriate interaction pattern, it is necessary to calculate the semantic similarity between the user's current behavior pattern and behavior pattern entities in the knowledge graph. Pre-trained language models (such as BERT and Word2Vec) can be used to calculate the cosine similarity between behavior pattern description texts, or cluster analysis (such as K-means) can be used to categorize historical behavior data to find the behavior pattern entity that best matches the user's current state.

[0193] After determining the target interaction pattern, it is necessary to query the corresponding scene interaction instruction logic parameters from the knowledge graph. These logic parameters include:

[0194] Voice prompt content parameters: These determine the voice commands of the virtual tutor. For example, in a walking task, the tutor may provide suggestions on pacing, while in relaxation mode, the tutor may provide breathing guidance.

[0195] Action trigger condition parameters: Define the key actions in the scene, such as triggering the instructor's action demonstration when the user arrives at a specific path marker;

[0196] Item interaction permission parameters: These parameters control whether users can interact with certain items. For example, in certain modes, virtual devices may need to be locked to prevent accidental operation.

[0197] The generation of these logical parameters can be based on knowledge graph query mechanisms (such as GNN graph neural network inference) or by using rule engines (such as Drools) to parse predefined interaction rules, ensuring that the generated interaction instructions meet the user's current needs.

[0198] To ensure the accuracy of the interaction, it is necessary to ensure that the generated interaction commands match the scene in both time and space dimensions:

[0199] Timestamp interpolation alignment: Due to potential time synchronization discrepancies between different sensor data (such as user actions, voice input, etc.), it is necessary to interpolate and align interactive commands to ensure that commands are triggered at the appropriate time. Kalman filtering or linear interpolation can be used to correct timestamps, making the temporal distribution of commands smoother.

[0200] Spatial coordinate mapping: Different scene elements may be located in different spatial positions. For example, the coordinates of path markers may need to match the user's location to ensure that the user sees the correct guidance. A spatial transformation matrix can be used to transform the coordinates of scene elements, keeping the interaction logic consistent with the user's real-time location.

[0201] Finally, the processed interactive commands are grouped into an ordered command queue to ensure that the order of execution and triggering conditions meet expectations. For example, if a user stays at a path marker for more than a set time, the next command in the queue will trigger a voice prompt from the virtual tutor, prompting the user to continue.

[0202] This embodiment utilizes a domain knowledge graph to parse scene interaction commands, ensuring the rationality and adaptability of the interaction logic. Different user behavior patterns can be dynamically matched to the most suitable interaction method, making the immersive experience more natural and personalized. The introduction of the knowledge graph significantly reduces the hard coding of interaction rules, giving the system stronger scalability and automatic reasoning capabilities. Furthermore, through timestamp interpolation alignment and spatial coordinate mapping, the accuracy of interaction commands is ensured, improving the smoothness of user interaction and the timeliness of feedback in immersive scenarios.

[0203] In one embodiment, the above-mentioned S50 includes:

[0204] S501, fills in missing values ​​and standardizes the historical behavior feature vectors in the user profile to generate standardized historical behavior feature vector data;

[0205] S502, Analyze the weight allocation coefficient of the real-time state vector based on the stress index and focus weight of the real-time state vector;

[0206] S503, Multiply the weight allocation coefficient with the real-time state vector to generate weighted data of the real-time state vector;

[0207] S504, input the standardized data of the historical behavior feature vector and the weighted data of the real-time state vector into the multimodal fusion model based on the attention mechanism to generate the fused feature vector;

[0208] S505, Analyze the cosine similarity between the fused feature vector and the feature vector of the candidate guidance content in the guidance content library to obtain the matching score;

[0209] S506, in the guidance content library, candidate guidance content with matching scores exceeding a preset score threshold is selected, and a comprehensive decision parameter including priority labels and output confidence is generated according to the matching scores of the candidate guidance content in descending order. Each entry in the comprehensive decision parameter is associated with a candidate guidance content.

[0210] In this embodiment, in order to provide accurate personalized guidance content in the personalized interaction system, it is necessary to integrate the historical behavioral characteristics of the user profile with the real-time physical and mental state to generate decision parameters adapted to the current user state. The core of calculating the comprehensive decision parameters lies in how to reasonably combine long-term behavioral preferences and current physiological and psychological states, and optimize the accuracy of content selection through multimodal fusion technology. This process involves several key steps, including data preprocessing, state analysis, weighted calculation, multimodal fusion, and personalized matching.

[0211] User profile data originates from long-term accumulated interaction records, including historical training patterns, preferred scenarios, and past interaction habits. This data may contain missing values ​​or have inconsistent dimensions for different features, thus requiring preprocessing. Missing value imputation can employ interpolation methods based on K-Nearest Neighbors (KNN) or Gaussian Mixture Models (GMM) to fill in missing behavioral records, making the data distribution more complete. Standardization uses Z-score normalization or Min-Max scaling to convert different features to the same dimension, avoiding the influence of different features on subsequent calculations. For example, a user's past dwell time in a meditation scenario might be measured in minutes, while their response time to interactive prompts might be measured in seconds. Standardization allows these data to be calculated on the same scale, improving the comparability of fused features.

[0212] A user's current physiological and psychological state directly impacts their interaction needs; therefore, it's necessary to calculate state weighting coefficients to determine the degree of influence of the current state on comprehensive decision-making. The stress index reflects the user's anxiety level, while the focus weight reflects the user's level of attention. The weighting coefficients can be calculated using a fuzzy logic-based weighting strategy: when the stress index is high, the recommendation weight for high-cognitive-load content is reduced; when the focus weight is high, the weight for deeply interactive content is increased. For example, if a user's stress index is high, the system may reduce the probability of recommending high-information-density content and prioritize recommending relaxation training or simplified interaction tasks.

[0213] The calculated weighting coefficients need to be applied to the real-time state vector to adjust its contribution to subsequent fusion decisions. This can be achieved using the Hadamard product (element-wise multiplication) method, applying the weighting coefficients to each dimension of the state vector. For example, if the stress index has a weight of 0.7 and the focus weight is 0.3, the calculated weighted data will enhance the perception of stress states while appropriately considering the user's focus level, ensuring that subsequent decision-making logic can address the user's current stress situation without ignoring their attention level.

[0214] In the multimodal fusion stage, historical behavioral data and current state data need to be deeply integrated to obtain a more accurate representation of the user's state. Traditional linear weighting methods struggle to capture complex interaction patterns; therefore, an attention-based multimodal fusion model is employed. The model can learn the correlation between historical data and real-time state using a Transformer architecture or a Bi-LSTM+Self-Attention structure. For example, if a user has historically preferred a specific type of training pattern, and the current state matches this preference, the model automatically increases the weight of that behavioral feature, ensuring that the fused feature vector better reflects the user's current needs.

[0215] The fused user feature vector needs to be matched with candidate content in the guidance content library to select personalized guidance content suitable for the current user's state. Cosine similarity can be used to measure the similarity between user features and the features of each candidate content. For example, if the user's current state has a high similarity to meditation-related guidance content, then that type of content will have a higher matching score. The matching score measures the suitability of the content and serves as the basis for subsequent selection.

[0216] The matching score threshold is used to filter out low-relevance content that does not match the user's state. For example, if the preset threshold is 0.6, only candidate content with a cosine similarity greater than 0.6 will be retained. The filtered candidate content is then sorted in descending order of matching score, and priority labels and output confidence scores are generated.

[0217] Priority tags: used to distinguish the recommendation order of different content; the higher the score, the higher the priority of the content.

[0218] Output confidence score: The calculation method can use Softmax normalization to ensure that the confidence scores of all recommended content are distributed between 0 and 1. For example, if the matching score of a certain piece of guidance is 0.85, while the matching scores of other content are lower, then the confidence score of that content is higher, and the system will prioritize its output during the interaction process.

[0219] Ultimately, the comprehensive decision parameters include all selected candidate guidance content, with each item associated with a corresponding guidance content feature, providing data support for subsequent personalized recommendations.

[0220] This embodiment, by fusing long-term behavioral data from user profiles with real-time physiological states, can generate personalized decision parameters that better meet the user's current needs. Historical behavioral data provides insights into the user's long-term preferences, while real-time status data ensures that recommended content adapts to the user's current physiological and psychological state. Employing a multimodal fusion model based on an attention mechanism, the association between the two can be automatically learned, improving the accuracy of personalized recommendations. Furthermore, by calculating the matching score using cosine similarity and combining it with a dynamic threshold filtering mechanism, irrelevant content can be effectively filtered out, ensuring that recommended content is both accurate and matches the user's state, thus enhancing the intelligence level of the interactive experience.

[0221] In one embodiment, S60 includes:

[0222] S601, extract the matching score, priority label and output confidence from the comprehensive decision parameters, and extract the voice prompt content parameters, action trigger condition parameters and prop interaction permission parameters from the scene interaction instructions;

[0223] S602, perform feature alignment processing on the matching degree score in the comprehensive decision parameters and the action triggering condition parameters in the scene interaction instructions to generate a comprehensive matching degree score vector;

[0224] S603, Analyze the cosine similarity between the comprehensive matching score vector and the scene adaptation vector of the candidate guidance content in the guidance content library to obtain the scene interaction feature similarity score;

[0225] S604, Based on the linear weighted result of the output confidence score and the scene interaction feature similarity score, generate a real-time comprehensive score for each candidate guidance content;

[0226] S605, Based on the pressure index of the real-time state vector, generate an adaptive threshold, filter candidate guidance content whose real-time comprehensive score exceeds the adaptive threshold, and generate a set of personalized guidance content containing priority tags in descending order of real-time comprehensive score.

[0227] S606, Perform interaction conflict detection on the personalized guidance content set to remove guidance content in the personalized guidance content set that conflicts with the prop interaction permission parameters in the scene interaction instructions, and generate the final personalized guidance content.

[0228] In this embodiment, within the immersive interactive system, to ensure that the recommended personalized guidance content not only matches the user's current state but also aligns with the scene's interaction logic, a filtering process based on the feature similarity of comprehensive decision parameters and scene interaction commands is required. This filtering process involves key steps such as multi-dimensional feature alignment, similarity calculation, weight fusion, threshold adjustment, and conflict detection to guarantee the accuracy, adaptability, and consistency of the recommended content.

[0229] The comprehensive decision parameters are personalized recommendation data generated after previous steps integrate user profiles and real-time status, including:

[0230] Match score: measures how well a user's current state matches a particular piece of guidance content;

[0231] Priority tags: determine the recommendation order, with content of higher scores being output first;

[0232] Output confidence score: This represents the system's level of confidence in the recommended content, and is usually calculated using Softmax normalization.

[0233] Meanwhile, scene interaction commands provide the interaction logic within the scene, including:

[0234] Voice prompt content parameters: Define the voice prompts that need to be triggered in the scene, such as "Please adjust your posture";

[0235] Action trigger condition parameters: describe the prerequisite for triggering a certain interaction, such as "triggered when the user reaches the specified path marker";

[0236] Item interaction permission parameters: specify whether a user can interact with certain interactive items, such as "this item can only be operated by a specific user group".

[0237] These parameters can be extracted through database queries (SQL or NoSQL) or JSON-based API interfaces to ensure that parameters from different sources can be formatted and processed uniformly.

[0238] To ensure that personalized recommendations match the interaction logic in the scenario, feature alignment processing is required:

[0239] Time axis alignment: The data update frequency of different parameters may be different. For example, the comprehensive decision parameters are updated once per second, while the scene interaction commands may be refreshed once every five seconds. Therefore, linear interpolation or time window sliding average methods are needed for time alignment.

[0240] Spatial coordinate alignment: Some interactive commands (such as action triggering conditions) involve specific spatial locations, and it is necessary to map the applicable range of the matching score to the three-dimensional coordinate system of the scene, using coordinate transformation matrix or nearest neighbor search (k-NN) for matching.

[0241] Finally, the feature-aligned data is combined into a comprehensive matching score vector for subsequent calculations.

[0242] To measure the suitability of recommended content to the scenario, it is necessary to calculate the similarity between the overall matching score vector and the scenario suitability vector of the candidate guidance content. The higher the similarity score, the higher the matching degree between the candidate content and the current scenario.

[0243] To balance the impact of user state fit (output confidence) and scene fit (similarity score), a linear weighted calculation is required:

[0244] Real-time composite score = γ × output confidence + (1-γ) × scene interaction feature similarity score

[0245] Where: γ is the balance coefficient, calculated from the stress index and focus weight in the real-time state vector:

[0246] γ=α×(1-P)+β×A

[0247] P is the stress index; a higher value indicates higher user stress, thus reducing the priority of complex interactions. A is the focus weight; a higher value indicates more focused user attention, thus increasing the priority of deep interactions. α and β are adjustable parameters used to control the influence of different state factors.

[0248] This calculation method ensures that the selection of recommended content considers not only the adaptability of the user's state, but also its suitability in the current interaction scenario.

[0249] Different user states have different interaction needs, therefore the filtering threshold needs to be dynamically adjusted:

[0250] Threshold = Tbenchmark - λ × P

[0251] Where: T is the default recommended threshold; λ is the adjustment coefficient, which controls the influence of the pressure index on the threshold.

[0252] When the stress index is high, the threshold is lowered, allowing more suitable guidance content for relaxation training to be selected. When the stress index is low, the threshold is raised, making the recommended content more accurate. Finally, the selected content is sorted in descending order of real-time comprehensive score to generate a personalized guidance content set, with priority tags added.

[0253] To ensure that recommended content does not conflict with scene interaction rules, conflict detection is required:

[0254] Item permission conflict detection: Check if the recommended content involves restricted items. If the user does not currently have permission to use them, remove the content.

[0255] Time period conflict detection: Ensure that the applicable time of recommended content is consistent with the scene interaction instructions. For example, if a certain guidance content needs to be executed within a specific time period, but the current time is outside the range, it will not be recommended.

[0256] Action logic conflict detection: If the body movements involved in the recommended content conflict with the guiding actions set in the scene, the content will be reordered or replaced.

[0257] Conflict detection can employ rule-based verification (if-else logic) or knowledge graph-based logical reasoning (such as OWL-based reasoning) to ensure that the final output guidance is consistent with the user's state and the scenario logic.

[0258] This embodiment significantly improves the accuracy of personalized guidance content by filtering based on the feature similarity of comprehensive decision parameters and scene interaction commands. This ensures that the recommended content not only matches the user's physical and mental state but also adapts to the interaction needs of the current scene. By calculating a real-time comprehensive score using linear weighting, it balances user state adaptability and scene suitability, making the recommended content more personalized and real-time. Adaptive threshold filtering ensures dynamic adjustment of the recommendation strategy under different states, improving the flexibility of the recommended content. Furthermore, the interaction conflict detection mechanism effectively avoids conflicts between recommended content and scene rules, ensuring the smoothness and consistency of the interaction process and improving the intelligence level of the immersive experience.

[0259] In one embodiment, a multimodal perception-based guidance device is provided, which corresponds one-to-one with the multimodal perception-based guidance method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the guidance device based on multimodal perception of the present invention. The modules include: physiological and motor data acquisition module 10, mind-body state modeling module 20, immersive scene generation module 30, interactive instruction parsing module 40, personalized behavior analysis module 50, guidance content matching module 60, and personalized content output module 70. Detailed descriptions of each functional module are as follows:

[0260] The physiological and motion data acquisition module 10 collects users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset;

[0261] The mind-body state modeling module 20 inputs the multimodal body dataset into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0262] The immersive scene generation module 30, based on a 3D scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, generates an immersive interactive scene containing environmental simulation elements.

[0263] The interaction instruction parsing module 40 parses the scene interaction instructions in the immersive interaction scene based on the entity association relationships in the domain knowledge graph.

[0264] Personalized behavior analysis module 50 integrates historical behavior feature vectors from user profiles with real-time state vectors from the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0265] The guidance content matching module 60 selects personalized guidance content from the guidance content library based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions.

[0266] The personalized content output module 70 outputs the personalized guidance content.

[0267] In one embodiment, the physiological and motor data acquisition module 10 is specifically used for:

[0268] By collecting users' skin conductivity and blood oxygen saturation data through wearable smart devices, a first physiological index subset is generated;

[0269] The acceleration data of the user's torso and the angular velocity data of the limb joints are collected by the inertial measurement unit to generate a subset of three-dimensional motion trajectories;

[0270] The system uses optical motion capture devices to collect data on changes in facial expressions and amplitude of limb movements, and generates a subset of motion features.

[0271] The data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset are time-stamp aligned.

[0272] Signal filtering is performed on the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset after timestamp alignment to generate a standardized multimodal body dataset.

[0273] In one embodiment, the mind-body state modeling module 20 is specifically used for:

[0274] The first physiological index subset, the three-dimensional motion trajectory subset, and the action feature subset in the multimodal body dataset are divided into temporal segments according to a preset time window;

[0275] The time-series segment is input into a mind-body association model constructed based on a gated loop unit, and cross-modal association features between the first physiological index subset, the three-dimensional motion trajectory subset, and the action feature subset are extracted through the mind-body association model.

[0276] Based on the cross-modal association features, a mind-body state mapping matrix containing dynamic attention parameters and joint angle thresholds is generated in the output layer of the mind-body association model.

[0277] Based on the stress index and focus weight in the dynamic attention parameters, the user's physical and mental state level is determined.

[0278] Based on the deviation between the joint angle threshold and the motion trajectory data in the subset of the three-dimensional motion trajectory, a motion standardization evaluation coefficient is generated.

[0279] The physical and mental state levels and the behavioral norms assessment coefficients are integrated into a physical and mental state mapping relationship that includes real-time state vectors.

[0280] In one embodiment, the immersive scene generation module 30 is specifically used for:

[0281] The user's current behavior pattern is determined based on the historical behavioral feature vector in the user profile and the real-time state vector in the mapping relationship between the physical and mental states.

[0282] Determine a basic 3D scene model that matches the user's current behavior pattern from the 3D scene model library;

[0283] Based on the stress index in the dynamic attention parameters, adjust the ambient light and shadow intensity parameters and weather simulation parameters in the basic 3D scene model;

[0284] Based on the focus weight in the dynamic attention parameters, path markers and synchronized virtual tutor actions are generated in the basic 3D scene model.

[0285] By integrating the ambient light and shadow intensity parameters, weather simulation parameters, path markers, and synchronized virtual tutor actions, an immersive interactive scene is generated;

[0286] Apply physics engine-driven motion trajectory constraints to virtual objects in the immersive interactive scene.

[0287] In one embodiment, the interactive instruction parsing module 40 is specifically used for:

[0288] Extract the predefined entity relationship graph structure from the domain knowledge graph. The entity relationship graph structure includes scene element entities, behavior pattern entities, and the association attributes between scene element entities and behavior pattern entities.

[0289] Extract the current scene element entity set from the immersive interactive scene, the current scene element entity set including virtual tutor entity, path marker entity and interactive prop entity;

[0290] Based on the semantic similarity matching results between the user's current behavior pattern and the behavior pattern entities in the domain knowledge graph, the target interaction pattern entity is determined.

[0291] Based on the association attributes between the scene element entities and the behavior pattern entities, generate scene interaction instruction logic parameters that are bound to the target interaction pattern entity;

[0292] The scene interaction command logic parameters are timestamped and aligned with the current scene element entity set, and spatial coordinates are mapped to generate an executable scene interaction command queue.

[0293] In one embodiment, the personalized behavior analysis module 50 is specifically used for:

[0294] Missing values ​​are filled and standardized in the historical behavior feature vectors of user profiles to generate standardized historical behavior feature vector data.

[0295] Based on the stress index and focus weight of the real-time state vector, analyze the weight allocation coefficient of the real-time state vector;

[0296] The weight allocation coefficient is multiplied by the real-time state vector to generate weighted real-time state vector data.

[0297] The standardized data of the historical behavior feature vectors and the weighted data of the real-time state vectors are input into a multimodal fusion model based on an attention mechanism to generate a fused feature vector;

[0298] The cosine similarity between the fused feature vector and the feature vectors of candidate guidance content in the guidance content library is analyzed to obtain a matching score;

[0299] Candidate guidance content with matching scores exceeding a preset score threshold is filtered from the guidance content library. Based on the matching scores of the candidate guidance content in descending order, a comprehensive decision parameter containing priority labels and output confidence is generated. Each entry in the comprehensive decision parameter is associated with a candidate guidance content.

[0300] In one embodiment, the guidance content matching module 60 is specifically used for:

[0301] Extract the matching score, priority label and output confidence from the comprehensive decision parameters, and extract the voice prompt content parameters, action trigger condition parameters and prop interaction permission parameters from the scene interaction instructions;

[0302] The matching score in the comprehensive decision parameters is aligned with the action triggering condition parameters in the scene interaction instructions to generate a comprehensive matching score vector.

[0303] The cosine similarity between the comprehensive matching score vector and the scene adaptability vector of the candidate guidance content in the guidance content library is analyzed to obtain the scene interaction feature similarity score.

[0304] A real-time comprehensive score for each candidate guidance content is generated based on the linear weighted result of the output confidence score and the scene interaction feature similarity score.

[0305] An adaptive threshold is generated based on the pressure index of the real-time state vector. Candidate guidance content with real-time comprehensive scores exceeding the adaptive threshold is filtered out, and a set of personalized guidance content containing priority tags is generated in descending order of real-time comprehensive scores.

[0306] Interaction conflict detection is performed on the personalized guidance content set to remove guidance content that conflicts with the prop interaction permission parameters in the scene interaction instructions, and the final personalized guidance content is generated.

[0307] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements functions or steps on the server side of a multimodal perception-based guidance method.

[0308] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements a user-side function or step of a multimodal perception-based guidance method.

[0309] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0310] Collect users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset;

[0311] The multimodal body dataset is input into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0312] Based on a 3D scene model library, and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated.

[0313] Based on the entity relationships in the domain knowledge graph, the scene interaction commands in the immersive interactive scene are parsed.

[0314] By integrating the historical behavioral feature vectors in the user profile with the real-time state vectors in the mapping relationship between the physical and mental states, comprehensive decision parameters are generated.

[0315] Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, personalized guidance content is selected from the guidance content library;

[0316] Output the personalized guidance content.

[0317] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0318] Collect users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset;

[0319] The multimodal body dataset is input into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0320] Based on a 3D scene model library, and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated.

[0321] Based on the entity relationships in the domain knowledge graph, the scene interaction commands in the immersive interactive scene are parsed.

[0322] By integrating the historical behavioral feature vectors in the user profile with the real-time state vectors in the mapping relationship between the physical and mental states, comprehensive decision parameters are generated.

[0323] Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, personalized guidance content is selected from the guidance content library;

[0324] Output the personalized guidance content.

[0325] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0326] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0327] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0328] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A guidance method based on multimodal perception, characterized in that, Includes the following steps: Collect users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset; The multimodal body dataset is input into a pre-trained mind-body association model to generate a mind-body state mapping relationship. This includes: dividing the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset in the multimodal body dataset into temporal segments according to a preset time window; the three-dimensional motion trajectory subset is generated based on trunk acceleration data and angular velocity data of limb joints; the action feature subset is generated based on facial expression change data and limb movement amplitude data; inputting the temporal segments into a mind-body association model constructed based on a gated recurrent unit; extracting cross-modal association features between the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset through the mind-body association model; generating a mind-body state mapping matrix containing dynamic attention parameters and joint angle thresholds at the output layer of the mind-body association model based on the cross-modal association features; determining the user's mind-body state level based on the stress index and focus weight in the dynamic attention parameters; generating a motion standardization evaluation coefficient based on the deviation between the joint angle threshold and the motion trajectory data in the three-dimensional motion trajectory subset; and integrating the mind-body state level and the motion standardization evaluation coefficient into a mind-body state mapping relationship containing a real-time state vector. Based on a 3D scene model library, and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated. Based on the entity relationships in the domain knowledge graph, the scene interaction commands in the immersive interactive scene are parsed. By integrating the historical behavioral feature vectors in the user profile with the real-time state vectors in the mapping relationship between the physical and mental states, comprehensive decision parameters are generated. Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, personalized guidance content is selected from the guidance content library; Output the personalized guidance content.

2. The guidance method based on multimodal perception as described in claim 1, characterized in that, Collect users' physiological index data and 3D motion trajectory data to generate a multimodal body dataset, including: By collecting users' skin conductivity and blood oxygen saturation data through wearable smart devices, a first physiological index subset is generated; The acceleration data of the user's torso and the angular velocity data of the limb joints are collected by the inertial measurement unit to generate a subset of three-dimensional motion trajectories; The system uses optical motion capture devices to collect data on changes in facial expressions and amplitude of limb movements, and generates a subset of motion features. The data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset are time-stamp aligned. Signal filtering is performed on the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset after timestamp alignment to generate a standardized multimodal body dataset.

3. The guidance method based on multimodal perception as described in claim 1, characterized in that, Based on a 3D scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated, including: The user's current behavior pattern is determined based on the historical behavioral feature vector in the user profile and the real-time state vector in the mapping relationship between the physical and mental states. Determine a basic 3D scene model that matches the user's current behavior pattern from the 3D scene model library; Based on the stress index in the dynamic attention parameters, adjust the ambient light and shadow intensity parameters and weather simulation parameters in the basic 3D scene model; Based on the focus weight in the dynamic attention parameters, path markers and synchronized virtual tutor actions are generated in the basic 3D scene model. By integrating the ambient light and shadow intensity parameters, weather simulation parameters, path markers, and synchronized virtual tutor actions, an immersive interactive scene is generated; Apply physics engine-driven motion trajectory constraints to virtual objects in the immersive interactive scene.

4. The guidance method based on multimodal perception as described in claim 1, characterized in that, Based on the entity relationships in the domain knowledge graph, the scene interaction commands in the immersive interactive scene are parsed, including: Extract the predefined entity relationship graph structure from the domain knowledge graph. The entity relationship graph structure includes scene element entities, behavior pattern entities, and the association attributes between scene element entities and behavior pattern entities. Extract the current scene element entity set from the immersive interactive scene, the current scene element entity set including virtual tutor entity, path marker entity and interactive prop entity; Based on the semantic similarity matching results between the user's current behavior pattern and the behavior pattern entities in the domain knowledge graph, the target interaction pattern entity is determined. Based on the association attributes between the scene element entities and the behavior pattern entities, generate scene interaction instruction logic parameters that are bound to the target interaction pattern entity; The scene interaction command logic parameters are timestamped and aligned with the current scene element entity set, and spatial coordinates are mapped to generate an executable scene interaction command queue.

5. The guidance method based on multimodal perception as described in claim 1, characterized in that, By fusing historical behavioral feature vectors from user profiles with real-time state vectors from the aforementioned physical and mental state mapping relationship, comprehensive decision parameters are generated, including: Missing values ​​are filled and standardized in the historical behavior feature vectors of user profiles to generate standardized historical behavior feature vector data. Based on the stress index and focus weight of the real-time state vector, analyze the weight allocation coefficient of the real-time state vector; The weight allocation coefficient is multiplied by the real-time state vector to generate weighted real-time state vector data. The standardized data of the historical behavior feature vectors and the weighted data of the real-time state vectors are input into a multimodal fusion model based on an attention mechanism to generate a fused feature vector; The cosine similarity between the fused feature vector and the feature vectors of candidate guidance content in the guidance content library is analyzed to obtain a matching score; Candidate guidance content with matching scores exceeding a preset score threshold is filtered from the guidance content library. Based on the matching scores of the candidate guidance content in descending order, a comprehensive decision parameter containing priority labels and output confidence is generated. Each entry in the comprehensive decision parameter is associated with a candidate guidance content.

6. The guidance method based on multimodal perception as described in claim 1, characterized in that, Based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, personalized guidance content is selected from the guidance content library, including: Extract the matching score, priority label and output confidence from the comprehensive decision parameters, and extract the voice prompt content parameters, action trigger condition parameters and prop interaction permission parameters from the scene interaction instructions; The matching score in the comprehensive decision parameters is aligned with the action triggering condition parameters in the scene interaction instructions to generate a comprehensive matching score vector. The cosine similarity between the comprehensive matching score vector and the scene adaptability vector of the candidate guidance content in the guidance content library is analyzed to obtain the scene interaction feature similarity score. A real-time comprehensive score for each candidate guidance content is generated based on the linear weighted result of the output confidence score and the scene interaction feature similarity score. An adaptive threshold is generated based on the pressure index of the real-time state vector. Candidate guidance content with real-time comprehensive scores exceeding the adaptive threshold is filtered out, and a set of personalized guidance content containing priority tags is generated in descending order of real-time comprehensive scores. Interaction conflict detection is performed on the personalized guidance content set to remove guidance content that conflicts with the prop interaction permission parameters in the scene interaction instructions, and the final personalized guidance content is generated.

7. A guidance device based on multimodal perception, characterized in that, The multimodal perception-based guidance device includes: The physiological and motion data acquisition module collects users' physiological index data and three-dimensional motion trajectory data to generate a multimodal body dataset. The mind-body state modeling module inputs the multimodal body dataset into a pre-trained mind-body association model to generate a mind-body state mapping relationship. This includes: dividing the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset from the multimodal body dataset into temporal segments according to a preset time window; the three-dimensional motion trajectory subset is generated based on trunk acceleration data and angular velocity data of limb joints; the action feature subset is generated based on facial expression change data and limb movement amplitude data; inputting the temporal segments into a mind-body association model constructed based on a gated recurrent unit; extracting cross-modal association features between the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset through the mind-body association model; generating a mind-body state mapping matrix containing dynamic attention parameters and joint angle thresholds at the output layer of the mind-body association model based on the cross-modal association features; determining the user's mind-body state level based on the stress index and focus weight in the dynamic attention parameters; generating a motion standardization evaluation coefficient based on the deviation between the joint angle threshold and the motion trajectory data in the three-dimensional motion trajectory subset; and integrating the mind-body state level and the motion standardization evaluation coefficient into a mind-body state mapping relationship containing a real-time state vector. The immersive scene generation module, based on a 3D scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, generates an immersive interactive scene containing environmental simulation elements. The interaction instruction parsing module parses the scene interaction instructions in the immersive interaction scene based on the entity association relationships in the domain knowledge graph; The personalized behavior analysis module integrates historical behavior feature vectors from user profiles with real-time state vectors from the mapping relationship between physical and mental states to generate comprehensive decision parameters. The guidance content matching module selects personalized guidance content from the guidance content library based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions. The personalized content output module outputs the personalized guidance content.

8. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal perception-based guidance program stored in the memory and executable on the processor, wherein the multimodal perception-based guidance program, when executed by the processor, implements the steps of the multimodal perception-based guidance method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a multimodal perception-based guidance program, which, when executed by a processor, implements the steps of the multimodal perception-based guidance method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Personnel decision-making ability evaluation method and system based on multi-modal data

    CN117875751A

  • Personalized rehabilitation plan generation method and system based on multi-modal data

    CN119446411A