Guidance method and device based on multi-modal perception, equipment and medium

Through multimodal data collection and physical and mental state modeling, combined with immersive scene interaction and knowledge graph analysis, personalized guidance content is generated, which solves the problem of insufficient recognition of individual physiological and psychological states in existing technologies, and improves the personalized and immersive interactive experience in the fields of medical health, financial technology and mental health care.

CN120611149AActive Publication Date: 2025-09-09PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510701696.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-09
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing technologies in the fields of healthcare, financial technology, and mental health care have difficulty accurately identifying an individual's physical and psychological state, resulting in insufficient personalized guidance and immersive interactive experience.

Method used

By collecting the user's physiological indicator data and three-dimensional motion trajectory data, a multimodal body data set is generated. The pre-trained mind-body association model is input to generate a mind-body state mapping relationship. Combined with the three-dimensional scene model library and domain knowledge graph, an immersive interactive scene is generated. The user portrait and real-time state vector are integrated to generate personalized guidance content.

Benefits of technology

It achieves accurate perception of the user's physiological state and behavioral characteristics, improves the accuracy and immersion of personalized guidance, adapts to the behavioral patterns of different users, and enhances the effect of intelligent guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611149A_ABST
    Figure CN120611149A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as medical health fund, fusion science and technology and mental health recuperation, and discloses a guidance method, device, equipment and medium based on multi-modal perception.The guidance method comprises the steps that physiological index data and motion trail data are collected, and a multi-modal body data set is generated; inputting a pre-trained mind and body association model to generate a mind and body state mapping relation; combining the three-dimensional scene model library and the dynamic attention parameters to generate an immersive interaction scene, and analyzing a scene interaction instruction based on a domain knowledge graph; and fusing the user portrait and the real-time state vector, generating a comprehensive decision parameter, and screening personalized guidance content from the guidance content library according to the feature similarity between the comprehensive decision parameter and the scene interaction instruction and outputting the personalized guidance content. According to the method, through multi-modal data acquisition and body and mind state modeling, accurate perception of the physiological state and behavior characteristics of the user is realized, and the accuracy of personalized guidance is improved in combination with immersive scene interaction and knowledge graph analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a guidance method, device, equipment and storage medium based on multimodal perception. Background Art

[0002] In the application and development of intelligent interactive systems, although multimodal perception technology, immersive interactive systems and artificial intelligence have made important progress in many industries, existing technologies still have many shortcomings, especially in the fields of medical health, financial technology business and mental health care. Intelligent systems still have great limitations in simulating complex scenarios, accurately identifying individual behavior patterns and providing personalized guidance, making it difficult to meet the in-depth needs of different fields.

[0003] In the field of medical and health business, existing technologies are insufficient in intelligent health monitoring and behavioral intervention. The current health management system mainly relies on user self-input or static health data analysis, and lacks an integrated understanding of the user's multimodal physiological data (such as heart rate, skin conductivity, movement patterns, etc.) and psychological state, resulting in an incomplete assessment of the user's health status. In addition, existing technologies also have limitations in rehabilitation training and psychological intervention. For example, during the rehabilitation process, the patient's movement behavior and psychological state are closely related, but existing technologies are difficult to accurately capture the individual's movement trajectory and emotional fluctuations, making it difficult to formulate personalized intervention plans. In addition, in mental health intervention, it is difficult for existing intelligent systems to combine the dynamic relationship between physiology and psychology, and it is impossible to accurately identify the patient's emotional state and behavioral patterns, resulting in low accuracy and real-time performance of personalized treatment recommendations. Although existing immersive therapies provide a partially immersive environment, they lack in-depth interaction with the patient's physiological and psychological feedback, resulting in limited efficacy.

[0004] In the field of mental health care, existing technologies mainly focus on the cognitive level processing of information and lack an in-depth understanding of the individual's body and mind integration. For example, in the process of mental health care, the individual's breathing rhythm and body posture are closely related to the psychological state. This body-mind interaction can directly affect the relaxation effect and emotional regulation. However, existing intelligent technologies have difficulty in capturing and modeling this body-mind interaction relationship, and only stay at the superficial guidance, such as playing relaxing music or providing fixed meditation guidance, but cannot provide accurate psychological care assistance based on the individual's real-time state. In addition, the current intelligent interactive system cannot effectively simulate real-life care scenes, such as a quiet forest walk, a soothing lakeside meditation, or a quiet rest space, making the user's experience in the virtual environment relatively fragmented and lacking immersion and realism.

[0005] In the fintech sector, existing intelligent analytics systems primarily rely on static data for analyzing financial behavior, making it difficult to accurately identify users' true decision preferences and risk tolerance. Current credit assessment and financial recommendation systems primarily base their calculations on users' historical transaction records, financial data, and market trends, while ignoring multimodal information such as their physiological state and behavioral patterns. For example, during the financial decision-making process, a user's emotional state, focus, and stress level can directly influence investment behavior and risk appetite. However, existing technologies struggle to detect and adjust these factors in real time, resulting in insufficient accuracy and personalization in financial decision support systems. Furthermore, in intelligent financial services, systems often fail to accurately understand user intent and needs. Financial advice recommendations are still based on fixed logic and lack the ability to dynamically adjust to the user's current cognitive state. Furthermore, in immersive trading and financial education scenarios, existing virtual systems still rely on static rules and are unable to dynamically optimize the learning experience based on user behavioral feedback and psychological state. This results in a lack of effective immersion and interaction in the financial decision-making and learning process. Summary of the Invention

[0006] The main purpose of the present invention is to provide a guidance method, device, equipment and storage medium based on multimodal perception, aiming to solve the technical problem that the existing technology lacks accurate perception of the individual's physical and mental state and personalized interaction optimization, and is difficult to provide adaptive immersive guidance.

[0007] To achieve the above objectives, the present invention provides a guidance method based on multimodal perception, comprising:

[0008] Collect the user's physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0009] Inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0010] Based on the three-dimensional scene model library and in combination with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene including environmental simulation elements is generated;

[0011] Parsing scene interaction instructions in the immersive interaction scene according to entity association relationships in the domain knowledge graph;

[0012] Fusing the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0013] Based on the feature similarity between the comprehensive decision parameter and the scenario interaction instruction, screening personalized guidance content from the guidance content library;

[0014] The personalized guidance content is output.

[0015] Furthermore, to achieve the above-mentioned object, the present invention provides a guidance device based on multimodal perception, comprising:

[0016] Physiological and motion data acquisition module, which collects the user's physiological index data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0017] a mind-body state modeling module, which inputs the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0018] An immersive scene generation module generates an immersive interactive scene including environmental simulation elements based on a three-dimensional scene model library and in combination with dynamic attention parameters in the mind-body state mapping relationship;

[0019] An interaction instruction parsing module, which parses the scene interaction instructions in the immersive interaction scene according to the entity association relationship in the domain knowledge graph;

[0020] A personalized behavior analysis module that integrates the historical behavior feature vectors in the user portrait with the real-time state vectors in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0021] A guidance content matching module, which selects personalized guidance content from a guidance content library based on the feature similarity between the comprehensive decision parameters and the scenario interaction instructions;

[0022] The personalized content output module outputs the personalized guidance content.

[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a guidance program based on multimodal perception stored in the memory and executable on the processor. When the guidance program based on multimodal perception is executed by the processor, the steps of the guidance method based on multimodal perception as described above are implemented.

[0024] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a guidance program based on multimodal perception is stored. When the guidance program based on multimodal perception is executed by a processor, the steps of the guidance method based on multimodal perception as described above are implemented.

[0025] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health, finance, technology, and mental health care. It discloses a guidance method based on multimodal perception, including: collecting physiological indicator data and three-dimensional motion trajectory data of users to generate a multimodal body data set; inputting a pre-trained mind-body association model to generate a mind-body state mapping relationship; based on the three-dimensional scene model library and the dynamic attention parameters in the mind-body state mapping relationship, generating an immersive interactive scene containing environmental simulation elements; parsing the scene interaction instructions in the immersive interactive scene according to the entity association relationship in the domain knowledge graph; fusing the historical behavior feature vector in the user portrait with the real-time state vector in the mind-body state mapping relationship to generate a comprehensive decision parameter; based on the feature similarity between the comprehensive decision parameter and the scene interaction instruction, screening personalized guidance content from the guidance content library and outputting the guidance content. The present invention realizes accurate perception of the user's physiological state and behavioral characteristics through multimodal data collection and mind-body state modeling, and improves the accuracy of personalized guidance by combining immersive scene interaction with knowledge graph analysis. The generation and matching optimization of comprehensive decision parameters enable the interactive content to adapt to the behavioral patterns of different users, thereby improving the personalization and immersion of intelligent guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0027] Figure 1 Schematic diagram of an application environment of a guidance method based on multimodal perception in an embodiment of the present invention;

[0028] Figure 2 1 is a flow chart of an embodiment of a guidance method based on multimodal perception according to the present invention;

[0029] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a guidance device based on multimodal perception according to the present invention;

[0030] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0031] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0033] The guidance method based on multimodal perception provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can collect the user's physiological indicator data and three-dimensional motion trajectory data through the user terminal to generate a multimodal body data set; input a pre-trained mind-body association model to generate a mind-body state mapping relationship; based on the three-dimensional scene model library and the dynamic attention parameters in the mind-body state mapping relationship, generate an immersive interactive scene containing environmental simulation elements; according to the entity association relationship in the domain knowledge graph, parse the scene interaction instructions in the immersive interactive scene; fuse the historical behavior feature vector in the user portrait with the real-time state vector in the mind-body state mapping relationship to generate a comprehensive decision parameter; based on the feature similarity between the comprehensive decision parameter and the scene interaction instruction, filter personalized guidance content from the guidance content library and output the guidance content. The present invention realizes accurate perception of the user's physiological state and behavioral characteristics through multimodal data collection and mind-body state modeling, combines immersive scene interaction with knowledge graph analysis, and improves the accuracy of personalized guidance. The generation and matching optimization of comprehensive decision parameters enable the interactive content to adapt to the behavioral patterns of different users, thereby improving the personalization and immersion of intelligent guidance. Among them, the user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server side can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0034] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a multimodal perception-based guidance method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0035] like Figure 2 As shown, the guidance method based on multimodal perception proposed by the present invention includes the following steps:

[0036] S10, collecting the user's physiological index data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0037] In this embodiment, the core goal of collecting the user's physiological indicator data and three-dimensional motion trajectory data and generating a multimodal body data set is to fully perceive the user's physiological state and movement pattern, and build complete physical and mental state information to support subsequent behavioral modeling and immersive interaction. The collection of physiological indicator data comes from a variety of wearable smart devices, including smart bracelets, smart clothing, ear-worn monitoring devices, etc. These devices can monitor the user's heart rate, skin conductivity, blood oxygen saturation and other key physiological parameters in real time. Heart rate monitoring is used to reflect the user's autonomic nervous system activity, skin conductivity is used to perceive the user's emotional fluctuations, and blood oxygen saturation can be used to assess the user's respiratory status and body oxygen supply. In addition to basic physiological parameters, EEG activity data can also be obtained in combination with EEG sensors to more accurately portray the user's psychological reactions in different states.

[0038] The collection of three-dimensional motion trajectory data involves inertial measurement units, optical motion capture devices, and computer vision analysis technology. The inertial measurement unit includes an accelerometer and a gyroscope, which can obtain the user's posture information, movement direction, and acceleration changes. The optical motion capture device captures the user's movements through multiple high-precision cameras and can identify the detailed features of limb movements, such as joint angles, movement amplitude, and gait characteristics. In addition, computer vision analysis technology uses deep learning models to segment, track, and identify user movements, making the analysis of complex movements more accurate. The construction of a three-dimensional motion trajectory not only includes the trajectory of body movement, but can also be combined with data from force feedback sensors to obtain the force conditions of the user when performing specific movements, thereby further optimizing the accuracy of motion analysis.

[0039] Based on the collected physiological indicator data and three-dimensional motion trajectory data, data preprocessing and fusion are required to ensure data consistency and accuracy. First, data from different sources need to be timestamped so that physiological data and motion data can be analyzed on the same timeline. Second, signal filtering is required to remove noise interference and improve data reliability. Filtering algorithms can use Kalman filtering, low-pass filtering, etc. to ensure that sensor data remains highly stable despite high-frequency jitter. Furthermore, during the fusion of physiological and motion data, data format standardization is required to support subsequent deep learning model modeling and analysis. Normalization methods can be based on maximum and minimum data normalization or z-score normalization to eliminate the impact of different measurement units on data analysis.

[0040] Wearable devices can collect physiological data based on photoplethysmography (PPE) sensing technology, which detects minute changes in blood flow to obtain highly accurate heart rate and blood oxygen data. Furthermore, electrophysiological signals can be acquired through skin electrodes to analyze changes in the user's sympathetic nerve activity. For motion data collection, high-precision inertial measurement units (IMUs) combined with computer vision technology can improve the accuracy of motion capture. For example, when a user performs specific yoga poses, the system can detect the user's body posture through an IMU and correct detection errors using an optical camera to provide more accurate posture data.

[0041] During the data fusion process, different processing strategies can be adopted according to different application scenarios. In high-precision scenarios (such as rehabilitation training), a multimodal fusion algorithm based on Bayesian optimization can be used to perform weighted processing on data from different sources to improve the accuracy of the fused data. In real-time interactive scenarios (such as immersive virtual experiences), recursive neural networks can be used to perform time series modeling on the data to ensure the timeliness and consistency of the data. In terms of data transmission, edge computing technology can be combined to complete some data preprocessing tasks on the wearable device side, reducing data transmission delays and improving the system's real-time response capabilities.

[0042] Example: In the healthcare sector, remote patient monitoring and rehabilitation training guidance can be performed. For example, for patients recovering from surgery, gait data and physiological parameters can be collected to assess their recovery and adjust rehabilitation training plans based on data changes. If an unstable gait and an abnormally elevated heart rate are detected, it can be inferred that the patient may be fatigued or unwell, and the system can provide rest recommendations or adjust training intensity.

[0043] In terms of mental health monitoring, physiological and motion data can be combined to analyze the user's emotional state. For example, by monitoring the user's heart rate variability, skin conductivity, and micro-expression characteristics, the user's anxiety level can be assessed and relaxation training or psychological intervention guidance can be provided when appropriate. Furthermore, in anxiety treatment, three-dimensional motion data can be used to assess the user's breathing rhythm, guiding the user to adjust their breathing pattern and improve their mental state.

[0044] In the financial sector, users' financial decision support systems can be optimized. For example, when a user makes a large transaction decision, the system can assess the user's decision-making state by combining physiological indicators (such as stress index) with operational behaviors (such as mouse tracking and keyboard input patterns). If the user is detected to be under high stress, the system can appropriately postpone transaction confirmation or provide a cooling-off period to reduce the risk of impulsive decision-making. Furthermore, in financial education scenarios, combining user attention data with cognitive load analysis can dynamically adjust the pace and difficulty of teaching content to improve learning efficiency.

[0045] By jointly collecting and integrating physiological and motion data, we can fully perceive the user's physical and mental state and accurately analyze their behavior patterns. Compared to single data collection methods, this multimodal perception technology can provide more complete and fine-grained user status information, providing a solid data foundation for subsequent physical and mental state modeling, immersive scene generation, and personalized interaction.

[0046] S20, inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0047] In this embodiment, the core goal of inputting a multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship is to use multimodal data to model the individual's mind-body state, extract the correlation features between physiological indicators, movement patterns and psychological states, and build an accurate mind-body interaction relationship. The mind-body association model is a time series analysis model built based on deep learning methods, which can perform feature extraction, state prediction and pattern analysis on the collected physiological data and movement data. The model adopts neural network structures such as recurrent neural networks (RNN), long short-term memory networks (LSTM) or gated recurrent units (GRU) to ensure the temporal dependency of the data, and introduces an attention mechanism to enable the model to focus on key physiological and movement features.

[0048] Before inputting a multimodal body dataset, the data needs to be preprocessed, including timestamp alignment, data normalization, outlier detection, and noise removal. Timestamp alignment refers to the synchronous processing of data from different data sources to ensure that all data can be modeled within the same time window. Data normalization can use maximum and minimum normalization or standardization to reduce the impact between different data dimensions. Outlier detection focuses on possible short-term distortion or loss of sensor data, using sliding window statistical analysis or autoencoder methods for anomaly detection and data completion. Noise removal can improve data stability through filtering methods such as low-pass filtering or wavelet denoising technology.

[0049] After preprocessing, the input data is divided into time-series data segments and fed into the mind-body association model. This model employs a multi-layer neural network structure, with the input layer receiving a subset of physiological indicators, a subset of three-dimensional motion trajectories, and a subset of motion features. Through a cross-modal feature fusion layer, a convolutional neural network (CNN) or self-attention mechanism is used to extract correlation features between data from different modalities, such as the synchronization between breathing rhythm and gait, and the coordination between heart rate changes and limb movements. During deep learning training, the model can be optimized through supervised or self-supervised learning methods. Training data can come from large-scale health databases, exercise behavior datasets, or data collected from specific experiments.

[0050] At the model's output, a mind-body state mapping is generated, consisting of dynamic attention parameters and mind-body state assessment parameters. Dynamic attention parameters measure an individual's reliance on different physiological and motor characteristics at a specific point in time. For example, during meditation, the system might focus more on the user's breathing depth, while during running training, it might focus more on heart rate changes. Mind-body state assessment parameters include stress index, concentration weight, emotional stability, and movement norm assessment coefficients. These parameters can be used for individual state analysis and subsequent interactive decision-making.

[0051] Through the deep integration of multimodal data and mind-body correlation modeling, we can accurately portray the user's physical and mental state, addressing the shortcomings of existing technologies in perceiving individual physical and mental states. This technology not only analyzes an individual's physiological state but also integrates it with exercise behavior for a comprehensive assessment, improving our understanding of individual states.

[0052] S30, generating an immersive interactive scene including environmental simulation elements based on the three-dimensional scene model library and in combination with the dynamic attention parameters in the mind-body state mapping relationship;

[0053] In this embodiment, based on the three-dimensional scene model library and combined with the dynamic attention parameters in the physical and mental state mapping relationship, the core goal of generating an immersive interactive scene containing environmental simulation elements is to enhance the individual's behavioral guidance effect through an immersive virtual environment, make the interactive experience more realistic, and adaptively adjust the environmental characteristics according to the user's physical and mental state, thereby improving the user's concentration, comfort and interaction matching.

[0054] The 3D scene model library stores 3D models of multiple virtual environments, including buildings, natural scenes, interactive props, avatars, lighting systems, and sound elements. These models can be predefined scenes, such as a mental health clinic, gym, or medical rehabilitation room, or they can be dynamically built through procedural generation to customize personalized environments based on user behavior.

[0055] The dynamic attention parameters in the mind-body state mapping relationship are derived from the user's multimodal physiological data (heart rate, stress index, concentration weight, movement status, etc.) and are used to adjust the dynamic interactive elements of the virtual environment. For example, when the user's stress index is detected to be high, the brightness of the ambient lighting can be reduced and visual distractions can be reduced to create a more relaxing interactive atmosphere. When the user's concentration weight is detected to be low, the user's attention can be improved by enhancing voice guidance and adding visual guidance signs.

[0056] The core elements of three-dimensional environment simulation include: lighting system, weather simulation, scene element interaction, virtual character synchronization, etc. The lighting system uses global illumination calculation (Global Illumination), which can adjust the intensity of ambient lighting according to time, space, and user status. For example, in a meditation scene, the lighting can present warm and soft tones, while in a sports training scene, the lighting can be more dynamic. Weather simulation is used to enhance the realism of the environment. For example, during rehabilitation training, weather changes can increase the user's sense of immersion and affect the training rhythm. Scene element interaction means that objects in the virtual environment can form real-time feedback with the user's movements, such as the virtual mentor character can synchronize the user's movement rhythm, or when the user approaches a certain interaction point, the system automatically triggers relevant prompts.

[0057] By building an environment based on a 3D scene model library, we can provide users with a highly immersive interactive experience, making behavioral guidance more realistic and intuitive. By combining the mapping relationship between physical and mental states, we can achieve personalized scene adjustments, allowing the virtual environment to dynamically adapt to changes in the user's state, improving user focus and interactive experience.

[0058] S40, parsing the scene interaction instructions in the immersive interaction scene according to the entity association relationship in the domain knowledge graph;

[0059] In this embodiment, the core goal of parsing scene interaction instructions in immersive interactive scenarios based on entity associations in a domain knowledge graph is to enable the system to understand and execute interaction logic that aligns with specific contexts, ensuring that the interactive experience in the virtual scene intelligently adapts to the user's state and behavior patterns. Traditional interactive systems rely on predefined rules for scene feedback and lack intelligent dynamic adaptation capabilities. However, methods based on domain knowledge graphs can make interactive content more precise and intelligent through semantic understanding, entity relationship reasoning, and context matching.

[0060] A domain knowledge graph is a knowledge system used for structured storage and representation of a certain domain. Its core elements include entities, entity attributes, and entity relationships. In immersive interactive scenarios, the construction of a knowledge graph involves factors such as scene elements, user behavior patterns, environmental variables, and interactive instructions. For example, in an intelligent rehabilitation training scenario, scene elements may include "rehabilitation equipment," "sports instructors," and "real-time feedback mechanisms"; user behavior patterns may involve "gait training," "strength training," and "balance exercises"; environmental variables may include "light intensity," "background noise," and "time factors"; and interactive instructions may involve "triggering specific training programs," "adjusting training difficulty," and "providing real-time voice feedback."

[0061] The key technical links for parsing scene interaction instructions include loading the knowledge graph, extracting scene elements, matching interaction patterns, generating instruction logic, and processing spatiotemporal alignment. Loading the knowledge graph refers to calling the stored knowledge structure for subsequent query and reasoning. The extraction of scene elements involves computer vision, object recognition, or user input analysis to determine the interactive elements in the current virtual environment. For example, in a virtual yoga training scene, the system may detect that "the user is currently in a tree pose," "the virtual instructor is providing movement guidance," and "meditation music is playing in the background." These elements need to be structured and stored in the knowledge graph for subsequent interactive logic reasoning.

[0062] Matching interaction patterns involves calculating the semantic similarity between the user's current behavior pattern and the standard behavior pattern in the knowledge graph to find the interaction method that best suits the current situation. For example, in a smart fitness scenario, if the user's gait data matches the "jogging mode" in the knowledge graph to a high degree of match, the system can automatically recommend the best action adjustment suggestions for jogging training. The generation of instruction logic depends on the entity association relationship in the knowledge graph. For example, if the user selects the "strength training" mode in the rehabilitation training scenario, and the knowledge graph shows that "strength training" is associated with "heart rate monitoring", the system can automatically enable the heart rate monitoring function.

[0063] Finally, spatiotemporal alignment ensures that scene interaction commands are triggered at the appropriate time and location. For example, in a virtual gym environment, if a user approaches the "dumbbell training area," the system can trigger the "upper limb muscle strengthening training" command based on knowledge graph reasoning and play a demonstration video of the action after the user enters the specific location.

[0064] Through interactive analysis based on domain knowledge graphs, intelligent, personalized, and immersive interactive experiences can be achieved, enabling the virtual environment to adapt to user behavior patterns and provide accurate real-time feedback. It can dynamically adjust interaction logic and optimize scene responses based on user status, thereby improving the naturalness and intelligence of interactions.

[0065] S50, fusing the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate a comprehensive decision parameter;

[0066] In this embodiment, the core goal of generating comprehensive decision parameters by integrating the historical behavioral feature vectors in the user portrait with the real-time state vectors in the physical and mental state mapping relationship is to accurately model the user's physical and mental state and behavioral preferences, and to generate personalized decisions through the fusion of multimodal data to improve the adaptability and intelligence in the interactive scenario.

[0067] User profiles are personalized feature sets built from historical data, encompassing a user's long-term behavior patterns, preferences, interaction habits, physiological state trends, and other information. This data can be obtained from multiple sources, including past training records, psychological state monitoring data, and scenario-based interaction behavior data, and is structured and stored using high-dimensional feature vectors. For example, in a health management scenario, a user profile might include information such as exercise frequency, sleep quality, mood swings, and heart rate trends.

[0068] The real-time state vector in the mind-body state mapping relationship is a dynamic feature set derived by modeling current physiological indicators, movement status, psychological state, and other data. Compared to the long-term behavioral patterns in user profiles, the real-time state vector can reflect the user's immediate state changes, such as their current stress level, fatigue level, and concentration.

[0069] During the data fusion process, the user profile's historical behavioral feature vectors must first be standardized and missing values ​​filled to ensure data consistency and integrity. For example, if certain features in the user profile have missing values, interpolation methods or collaborative filling strategies based on similar users can be used to complete them. Furthermore, to eliminate scale differences between different data sources, Z-score standardization or Min-Max normalization methods can be used to ensure that all feature values ​​are distributed within the same numerical range.

[0070] Secondly, the real-time state vector needs to be matched with the historical behavioral feature vector to ensure data alignment. During processing, a time window sliding method can be used to dynamically smooth the real-time state vector and generate a time-weighted feature vector to enhance the temporal consistency of the data. For example, in a sports training scenario, the moving average of the real-time state vector can be calculated based on the physiological data of the past 30 seconds to reduce the noise interference caused by short-term data fluctuations.

[0071] Then, a multimodal fusion network based on the attention mechanism is used to deeply fuse the historical behavioral feature vector of the user portrait with the real-time state vector. The multimodal fusion network can learn the weight relationship between different feature modes. For example, when the stress state is high, it pays more attention to physiological data, while when the state is stable, it pays more attention to historical behavioral patterns. For example, during meditation training, if the system detects that the user's historical training records show "long-term high stress state" and the current real-time state vector shows "faster breathing rate and increased heart rate", the system can increase the duration of the meditation guidance content or reduce the training intensity to adapt to the user's current state.

[0072] The integrated decision parameters after fusion include:

[0073] Match score: used to measure the similarity between the user profile and the current status.

[0074] Priority tag: used to sort different training plans and guidance content to ensure that high-priority content is recommended first.

[0075] Output confidence: used to evaluate the credibility of the current recommendation. When the confidence is low, multiple interactive options can be provided to enhance the adaptability of the system.

[0076] By integrating user profiles and real-time state vectors, more precise personalized decision-making can be achieved, enabling the system to dynamically adjust interaction content based on the user's long-term behavior patterns and immediate state. This allows for optimized interaction strategies across different application scenarios. For example, in sports training, training intensity can be adjusted based on the user's physical fatigue; in mental health management, targeted relaxation training can be provided based on the user's emotional state; and in smart investment advisory systems, investment recommendation strategies can be optimized based on the user's cognitive load, reducing the risk of impulsive decision-making under high-stress conditions.

[0077] S60, filtering personalized guidance content from a guidance content library based on feature similarity between the comprehensive decision parameter and the scenario interaction instruction;

[0078] In this embodiment, the core goal of selecting personalized guidance content from the guidance content library based on the feature similarity between comprehensive decision parameters and scene interaction instructions is to provide content recommendations that closely match the user's state and interaction needs, ensuring that users receive guidance information that best suits their current state and behavior patterns in an immersive interactive environment. Traditional content recommendation methods typically rely on predefined rules or static matching based on historical behavior, making it difficult to dynamically adapt to the user's real-time physical and mental state. This method, however, achieves real-time adaptation of personalized recommendations by calculating the feature similarity between comprehensive decision parameters and scene interaction instructions.

[0079] Comprehensive decision parameters are decision data generated by integrating the historical behavioral feature vectors in the user profile with the real-time state vectors in the physical and mental state mapping relationship. They include metrics such as matching scores, priority labels, and output confidence. These parameters quantify the user's long-term preferences and immediate state, enabling the recommendation system to tailor recommendations based on the user's current physical and mental state. For example, in a mental health training scenario, comprehensive decision parameters could include information such as the user's emotional stability index, anxiety level, and meditation training preferences to support personalized recommendations.

[0080] Scenario interaction instructions are dynamic interaction data parsed based on the domain knowledge graph, reflecting the interaction rules, trigger conditions, and environmental feedback mechanisms in the current scenario. For example, in a virtual fitness training environment, scenario interaction instructions might include "provide relaxation training suggestions when the user reaches a certain heart rate threshold" or "provide correction prompts when the user makes an incorrect movement."

[0081] Feature similarity is used to measure the degree of match between comprehensive decision parameters and scene interaction instructions. It typically uses cosine similarity, vector distance calculation, or a deep learning-based matching model. For example, if the matching score of the comprehensive decision parameters is highly similar to the action trigger conditions in the scene interaction instructions, the system can prioritize recommended guidance content related to that interaction rule.

[0082] During the calculation process, the feature vectors of the comprehensive decision parameters and scene interaction instructions must first be vectorized to ensure that they are in the same feature space. For example, word vector encoding, feature embedding, and normalization can be used to convert data from different modalities into numerical representations of the same dimension.

[0083] Secondly, a matching strategy based on feature similarity is adopted to calculate the similarity score between the comprehensive decision parameter vector and the scene interaction instruction vector.

[0084] Finally, based on the similarity calculation results, the most matching content in the guidance content library is selected, and the output confidence is combined to adjust the priority of the recommended content. For example, if a training solution has a high match and a high confidence level in the user's historical preferences, then this solution should be given priority.

[0085] By calculating the similarity between decision parameters and the characteristics of scene interaction instructions, accurate and highly personalized content recommendations can be achieved, enabling the system to dynamically adapt to the user's real-time status and enhance the intelligent level of the interactive experience. It can also adaptively adjust recommendation strategies to provide the most appropriate guidance content in different situations.

[0086] S70: Output the personalized guidance content.

[0087] In this embodiment, the core goal of delivering personalized guidance content is to provide highly tailored guidance information based on the user's individual characteristics and real-time state, thereby enhancing the accuracy and effectiveness of the immersive interactive experience. Personalized guidance content includes not only multimodal output forms such as text information, voice commands, video demonstrations, and real-time feedback, but also a dynamic content adjustment mechanism to ensure that the output guidance content is consistent with the user's physical and mental state, behavioral patterns, and interactive environment.

[0088] Personalized guidance content is the best guidance solution selected based on the aforementioned comprehensive decision parameters and the feature similarity of scenario interaction instructions. Its core features include:

[0089] Adaptive content generation: Guidance content should be dynamically adjusted based on the user’s current state. For example, in virtual rehabilitation training, the system can adjust the rhythm and intensity of training guidance based on the user’s fatigue state.

[0090] Multimodal information fusion: Personalized guidance content can be delivered through a combination of text, voice, video, and interactive animation to enhance user understanding. For example, in a smart fitness system, voice prompts and real-time video demonstrations can be provided to ensure users are performing exercises correctly.

[0091] Real-time interactive optimization: After the guidance content is output, the system can make adjustments based on the user's feedback and physiological data. For example, during meditation training, if it detects that the user's heart rate is decreasing slowly, the system can increase the duration of deep breathing guidance to enhance the relaxation effect.

[0092] During content output, it is necessary to ensure intelligent adaptation of output format, feedback method, and interaction logic:

[0093] Intelligent output format adaptation: Different users may have different ways of receiving information. The system can adjust the output format based on their historical preferences. For example, some users prefer text descriptions, while others prefer video demonstrations. The system can automatically select the optimal output format.

[0094] Intelligent optimization of feedback methods: When users follow the guidance content, the system can adjust the feedback strategy in real time based on physiological sensors and behavioral analysis. For example, in a virtual fitness environment, when a user fails to complete an action correctly, the system can immediately issue a voice prompt and provide a real-time correction video to guide the user to make adjustments.

[0095] Personalized interaction logic: The system can adjust the interaction process according to the user's real-time status. For example, in emotion management training, if it detects that the user is in a high-stress state, the system can automatically reduce complex training instructions and instead provide simpler and more intuitive relaxation exercises.

[0096] Example: In the healthcare sector, an immersive behavioral guidance method based on multimodal perception is used to provide precise rehabilitation training content recommendations to meet the personalized rehabilitation guidance needs of rehabilitation patients. The system first collects the patient's physiological indicator data, including heart rate, blood oxygen saturation, and skin conductivity, through wearable smart devices (such as smart bracelets and smart insoles). Simultaneously, an inertial measurement unit (IMU) monitors the patient's three-dimensional motion trajectory, such as gait rhythm, joint angular velocity, and posture angle. Furthermore, an optical motion capture device records the patient's limb range of motion and posture changes, forming a multimodal body data set. This data undergoes timestamp alignment and signal filtering to eliminate time errors between acquisition devices and remove abnormal signals. After acquiring the patient's multimodal body data, the system inputs it into a pre-trained mind-body correlation model to extract the patient's mind-body state mapping relationship. Based on historical training data, this model uses a gated recurrent unit (GRU) to learn the correlation features between physiological indicators, movement patterns, and psychological state. The model generates a state mapping matrix at the output layer, which contains stress index and concentration weights, and is used to assess the patient's current physiological and psychological state. At the same time, combined with the patient's movement trajectory deviation, the movement standardization evaluation coefficient is calculated to evaluate the accuracy of the training movement.

[0097] In order to enhance the immersive experience of rehabilitation training, the system creates a virtual interactive environment for patients that meets their rehabilitation training needs based on a three-dimensional scene model library. The system adjusts the ambient light intensity and spatial layout in the training scene based on the mapping relationship between the patient's physical and mental state. For example, if the patient's stress index is high, the system will reduce the stimulating factors in the environment, such as reducing the brightness of the light and reducing the background noise; if the concentration weight is high, the system will generate clearer path markings in the scene to guide the patient to complete the training task. In addition, the virtual instructor will synchronize with the patient's movement rhythm, provide real-time action demonstrations and voice guidance, and help patients improve the effectiveness of rehabilitation training. The physical engine constraints ensure that the virtual objects in the training scene have real motion feedback, such as simulating the patient's interaction with auxiliary training equipment in the virtual scene.

[0098] During training, the system uses the domain knowledge graph to parse scenario interaction instructions and adjust the interaction logic based on the patient's current behavior pattern. The domain knowledge graph predefines training steps, rules for using rehabilitation equipment, and information linking patient status and movement patterns. For example, if a patient's knee flexion angle does not reach the expected range, the system prompts the patient to adjust the movement through voice; if the patient's execution time for a training movement exceeds the set threshold, the system will dynamically adjust the training rhythm and provide appropriate rest prompts.

[0099] The generation of personalized training guidance is based on the patient's historical training data and current physical and mental state. The system analyzes the patient's performance in different training tasks and calculates the stress index and concentration weight of the real-time state vector to adjust the recommendation weight of personalized training suggestions. Through a multimodal fusion model based on the attention mechanism, the system matches the patient's training behavior characteristics with the rehabilitation guidance content library, calculates the matching score, and filters the training content with higher priority. For example, if the patient's matching score is high and the concentration weight is high, the system may recommend more challenging training content; otherwise, it recommends a more gentle training mode.

[0100] Finally, based on the feature similarity between the comprehensive decision parameters and the scene interaction instructions, the system selects personalized guidance content suitable for the patient's current state from the training content library. The matching score and the scene interaction feature similarity score are calculated through linear weighting to determine the priority of personalized training content, and the recommendation threshold is dynamically adjusted to adapt to the training needs of patients in different states. If certain training tasks conflict with the scene interaction logic (for example, the patient cannot currently use certain rehabilitation equipment), the system will automatically eliminate the conflicting content and output the final personalized training guidance to ensure the safety and efficiency of the rehabilitation training process.

[0101] In the fintech sector, this can be used for intelligent investment advisory, risk management, and personalized financial decision-making support systems. The system first uses multimodal sensing technology to collect physiological indicators (such as heart rate, pupil dilation, and skin conductivity) to assess emotional fluctuations. Simultaneously, it combines the user's financial behavior data (such as transaction records, portfolio adjustment frequency, and risk preference test results) to generate a multimodal financial behavior dataset.

[0102] This data is processed by a pre-trained financial behavior analysis model to generate a user's risk tolerance assessment. This model uses historical trading data and market volatility, combined with the user's physiological feedback (such as emotional reactions to market fluctuations), to calculate a stress index and focus weighting. If a user exhibits a high stress index during periods of high market volatility, the system can reduce the weight of recommended high-risk investments to prevent impulsive investments during periods of emotional instability. If a user's focus weighting is high, more complex investment strategies, such as multi-factor portfolio optimization or quantitative trading strategies, can be recommended.

[0103] To enhance the investment experience, the system constructs a virtual financial analysis environment based on a three-dimensional market scenario model library. For example, in the virtual market hall, users can view historical return curves for different investment portfolios in real time and adjust investment proportions through interactive data visualization tools. If the user's stress index is high, the system can adjust the scene's visual elements, such as reducing the frequency of market fluctuation information or adopting a softer color scheme to reduce visual burden. If the user's focus is high, the system can generate key analysis areas within the scene, such as highlighting key economic indicators that influence investment decisions.

[0104] Personalized financial decisions are generated based on the user's historical investment behavior and the current market environment. The system leverages domain knowledge graphs to analyze market interactions and adjust recommended strategies based on the user's current investment behavior patterns. For example, if a user's trading frequency increases and their sentiment fluctuates significantly, the system may recommend rebalancing their portfolio to reduce risk exposure. If a user's investment in high-risk assets increases, the system may trigger a risk warning and provide alternative investment recommendations based on the user's behavioral data.

[0105] Ultimately, the system, based on comprehensive decision-making parameters and market interaction logic, selects personalized investment recommendations from a library of financial products that match the user's current situation. A linearly weighted calculation is performed using the matching score and the market interaction feature similarity score to ensure that the recommended investment options are both consistent with the user's risk tolerance and adaptable to the current market environment.

[0106] In the field of mental health care, this technology can be applied to personalized meditation, relaxation guidance, and emotional regulation assistance systems. The system first collects the user's physiological indicators (such as heart rate variability, respiratory rate, and brain wave data) through wearable devices, and combines them with body movement data (such as sitting posture, walking rhythm, and body movement amplitude) to form a multimodal mental health behavior dataset.

[0107] These data are processed by a pre-trained psychological state model to generate a mapping relationship between the user's physical and mental state. The model learns the physiological and psychological characteristics of different psychological states through long-term accumulation of conditioning data. For example, the system can identify the user's relaxation depth based on brain wave data and calculate the user's concentration weight in combination with heart rate variability. If the user's concentration is low, the system may recommend simpler breathing guidance exercises, such as abdominal breathing or rhythmic breathing; if the user's stress index is high, the system may guide the user into a more soothing conditioning mode, such as mindful breathing, progressive muscle relaxation, or mind-body scanning exercises.

[0108] Based on a library of three-dimensional psychological conditioning scene models, the system can construct virtual relaxation spaces, such as tranquil lakeside settings, forest trails, or quiet rest rooms, to enhance the immersive psychological conditioning experience. The system dynamically adjusts the scene's atmosphere based on the user's physical and mental state, for example, playing soft background music and reducing visual stimulation during periods of anxiety. During periods of high concentration, the system enhances visual focus, such as guiding the user to gaze at a virtual candle flame or a slowly moving point of light. A virtual instructor can synchronize with the user's breathing rhythm and provide real-time voice guidance, such as "Take a deep breath and exhale slowly" or "Focus on each breath," ensuring a deep experience for the user.

[0109] Ultimately, based on comprehensive decision-making parameters and psychological well-being scenario logic, the system selects personalized guidance content from the psychological well-being guidance library that suits the user's current state. For example, when the system detects that the user has entered a deep state of relaxation, it may reduce the interruption of voice prompts and adjust the cadence of the guidance content. Through intelligent psychological well-being guidance, users can more efficiently enter a deep state of relaxation and experience a variety of psychological well-being methods in a virtual environment.

[0110] Through intelligent and personalized guidance content output, the system ensures the accuracy of the interactive experience, allowing users to receive feedback that best suits their current state. It can dynamically adjust guidance content based on the user's real-time status, achieving a more interactive and personalized user experience.

[0111] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health, finance, technology, and mental health care. It discloses a guidance method based on multimodal perception, including: collecting physiological indicator data and three-dimensional motion trajectory data of users to generate a multimodal body data set; inputting a pre-trained mind-body association model to generate a mind-body state mapping relationship; generating an immersive interactive scene containing environmental simulation elements based on the dynamic attention parameters in the three-dimensional scene model library and the mind-body state mapping relationship; parsing the scene interaction instructions in the immersive interactive scene according to the entity association relationship in the domain knowledge graph; fusing the historical behavior feature vector in the user portrait with the real-time state vector in the mind-body state mapping relationship to generate a comprehensive decision parameter; based on the feature similarity between the comprehensive decision parameter and the scene interaction instruction, screening personalized guidance content from the guidance content library and outputting the guidance content. The present invention realizes accurate perception of the user's physiological state and behavioral characteristics through multimodal data collection and mind-body state modeling, combines immersive scene interaction with knowledge graph analysis, and improves the accuracy of personalized guidance. The generation and matching optimization of the comprehensive decision parameter enable the interactive content to adapt to the behavior patterns of different users, improving the personalization and immersion of intelligent guidance.

[0112] In one embodiment, the above S10 includes:

[0113] S101, collecting skin conductivity data and blood oxygen saturation data of a user through a wearable smart device to generate a first physiological indicator subset;

[0114] S102, collecting acceleration data of the user's torso and angular velocity data of the limb joints through an inertial measurement unit to generate a three-dimensional motion trajectory subset;

[0115] S103, collecting the user's facial expression change data and body movement amplitude data through an optical motion capture device to generate a motion feature subset;

[0116] S104, performing timestamp alignment processing on the data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset;

[0117] S105 , performing signal filtering processing on the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset after the timestamp alignment processing to generate a standardized multimodal body data set.

[0118] In this embodiment, the purpose of collecting the user's physiological indicator data and 3D motion trajectory data is to construct a multimodal body dataset to accurately capture the user's physiological state, motion patterns, and behavioral characteristics. Multimodal data, integrating different types of physiological signals and motion characteristics, helps improve the accuracy of individual state recognition and provides data support for subsequent physical and mental state modeling and personalized guidance.

[0119] Skin conductivity (Electrodermal Activity, EDA) and blood oxygen saturation (SpO2) are important indicators for measuring the state of the human autonomic nervous system. Skin conductivity can reflect the activity level of the sympathetic nervous system and is closely related to mood swings, stress levels, and anxiety levels. Blood oxygen saturation is used to monitor blood oxygen content and can reflect the user's respiratory condition, fatigue level, etc. Wearable smart devices (such as smart bracelets and smart rings) are typically equipped with EDA sensors and optical blood oxygen sensors, which can collect this physiological data in real time.

[0120] During data collection, a combination of continuous and intermittent measurement can be employed. For example, in a static state, data can be collected every 30 seconds; during exercise or intense emotional fluctuations, the sampling frequency can be increased to once per second. Furthermore, to ensure data stability, an adaptive data fusion algorithm can be used to correct data based on environmental factors such as temperature and humidity, preventing external interference from affecting skin conductivity.

[0121] An inertial measurement unit (IMU), typically consisting of sensors such as accelerometers, gyroscopes, and magnetometers, can be used to detect a user's movement patterns. Trunk acceleration data reflects the user's overall movement trends, such as gait, walking speed, and exercise intensity. Angular velocity data from limb joints can be used to analyze a user's gestures, body coordination, and movement habits.

[0122] During data collection, IMU sensors are deployed at multiple locations (such as the wrist, ankle, and waist) to ensure the integrity of the three-dimensional trajectory. After data collection, Kalman filtering or complementary filtering can be used for data fusion to reduce noise and improve accuracy. Furthermore, six-degree-of-freedom (6-DoF) or nine-degree-of-freedom (9-DoF) posture estimation algorithms can be used to calculate information such as the user's posture angle and rotation angle for more accurate motion trajectory modeling.

[0123] Optical motion capture devices (such as RGB cameras, depth cameras, and infrared cameras) can identify a user's facial expressions, body movements, and posture. Facial expression data, including micro-expression features of the eyebrows, eyes, and mouth, can be used to infer a user's emotional state, such as anxiety, joy, and fatigue. Body movement amplitude data can be used to analyze user gesture interaction, posture adjustment, and movement coordination.

[0124] During data collection, a multi-view camera solution is required to ensure accurate recognition of both facial and body movements. For facial expression recognition, deep learning algorithms such as OpenFace and MediaPipe Face Mesh can be used to extract facial key points, and then classify the user's emotional state using support vector machines (SVMs) or convolutional neural networks (CNNs). For body movement recognition, skeletal tracking algorithms such as OpenPose and MediaPipe Pose can be used to perform 3D reconstruction of the user's body key points and analyze characteristics such as movement amplitude and movement smoothness.

[0125] Due to the different sampling frequencies and data formats of different sensors, the raw data usually has time synchronization deviations, and timestamp alignment is required. Time alignment methods include:

[0126] Interpolation-based methods: For low-sampling-rate data (such as EDA and SpO2), linear interpolation or spline interpolation can be used to align them with high-frequency data (such as IMU data).

[0127] Time window-based method: Set a fixed time window (such as 100ms) and map all sensor data to the same time dimension to ensure synchronous data processing.

[0128] Method based on dynamic time warping (DTW): When there is nonlinear time deviation in time series data, the DTW algorithm can be used to align different modal data to keep them consistent in the time dimension.

[0129] Raw sensor data may contain noise and outliers, so signal filtering is required to improve data quality. Signal filtering can include the following methods:

[0130] Low-pass filtering (Butterworth filter, mean filter): Removes high-frequency noise and improves signal smoothness. For example, in IMU data processing, it removes rapidly jittering high-frequency signals to improve the continuity of motion trajectories.

[0131] Median filtering: Used to remove short-term abnormal data points. For example, in EDA signal processing, it can filter out transient spike noise to ensure the stability of skin conductivity data.

[0132] Autoregressive Moving Average (ARMA) filtering: In the process of multimodal data fusion, current values ​​can be predicted based on historical data trends to improve data robustness.

[0133] The standardized multimodal body dataset is the basic data source for subsequent physical and mental state mapping and personalized behavior guidance, and can be used in scenarios such as user state analysis and personalized training recommendations.

[0134] By integrating physiological signals, motion data, and movement characteristics, this embodiment can more accurately characterize the user's physical and mental state, making personalized behavioral guidance more scientific and efficient. This can be used in fields such as sports training, health management, and psychological intervention, improving the accuracy of user status monitoring and supporting intelligent optimization of personalized guidance plans.

[0135] In one embodiment, the above S20 includes:

[0136] S201, dividing the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset in the multimodal body dataset into time series segments according to a preset time window;

[0137] S202, inputting the time series segments into a mind-body association model constructed based on a gated recurrent unit, and extracting cross-modal association features between the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset through the mind-body association model;

[0138] S203, based on the cross-modal association features, generating a mind-body state mapping matrix including dynamic attention parameters and joint angle thresholds at the output layer of the mind-body association model;

[0139] S204, determining the user's physical and mental state level based on the stress index and concentration weight in the dynamic attention parameter;

[0140] S205, generating a motion normativeness evaluation coefficient based on a deviation value between the joint angle threshold and the motion trajectory data in the three-dimensional motion trajectory subset;

[0141] S206: Integrate the physical and mental state levels and the action normativeness evaluation coefficients into a physical and mental state mapping relationship including a real-time state vector.

[0142] In this embodiment, the core goal of inputting a multimodal body dataset into a pre-trained mind-body association model and generating a mind-body state mapping is to establish a correlation between the user's physiological state, movement patterns, and behavioral characteristics to achieve personalized mind-body state assessment. Traditional human state analysis typically relies solely on single-modal data, such as physiological signals or movement trajectories, while ignoring the potential relationships between multimodal information. This method extracts cross-modal association features through a deep learning model and combines it with time series analysis to generate a mind-body state mapping to improve the accuracy and real-time performance of state assessment.

[0143] Multimodal data is usually collected as a continuous time series, but the sampling frequencies of different modalities may be different. Therefore, time windowing is required to divide the data into time series segments of equal length to ensure that the data of each modality is aligned and analyzed on the same time dimension. Time windowing can include:

[0144] Fixed time window: You can set a fixed window length (such as 1 second, 5 seconds) to ensure that the data remains in a fixed format when it is input into the model.

[0145] Sliding time window: A sliding window method (such as a 5-second window with a step size of 0.5 seconds) is used to enhance the continuity of time series features and improve the ability to capture short-term dynamic changes.

[0146] Adaptive Window: For specific applications, such as emotion recognition, the window length can be dynamically adjusted based on the rate of change of physiological data. For example, when skin conductivity changes rapidly, the window length is shortened to improve temporal resolution.

[0147] Multimodal data has temporal dependencies and cross-modal correlations, so a GRU network is used to model time series data and extract cross-modal features. The GRU is a recurrent neural network (RNN) variant suitable for learning long sequences. It can effectively capture the long-term dependencies of time series data while avoiding the vanishing gradient problem of traditional RNNs.

[0148] GRU structure: Contains update gate and reset gate, which are used to control the degree of fusion of past information and current input, and improve the correlation of data across time steps.

[0149] Cross-modal feature extraction: The model input includes a subset of primary physiological indicators (such as EDA and SpO2), a subset of three-dimensional motion trajectories (such as acceleration and angular velocity), and a subset of motion features (such as facial expressions and body posture). The GRU learns the temporal dependencies between different modal data through multi-layer encoding. For example, it analyzes the synchronization between heart rate changes and exercise rhythm to infer the user's exercise fatigue status.

[0150] At the output of GRU, an attention mechanism is introduced to dynamically adjust the weights of different modal data in decision-making and generate a physical and mental state mapping matrix.

[0151] Attention parameters: The attention weight of each modality is calculated based on the hidden state of the GRU. For example, in a meditation training scenario, the system may pay more attention to heart rate and skin conductivity, while in a sports training scenario, it may pay more attention to the three-dimensional motion trajectory.

[0152] Joint angle threshold: For motion state assessment, the safety threshold of the joint angle is stored in the physical and mental state mapping matrix (for example, a risk warning is triggered when the knee flexion angle exceeds a certain value) for motion normative assessment.

[0153] The user's physical and mental state can be graded and assessed using the Stress Index and Focus Weight.

[0154] Stress Index: Calculated based on physiological signals such as skin conductivity and heart rate variability, it reflects the user's level of stress. For example, an increase in skin conductivity may indicate increased stress, while a decrease in heart rate variability may indicate anxiety.

[0155] Concentration Weight: This is calculated using data such as eye tracking and electroencephalogram (EEG) signals to assess the user's concentration. For example, in a learning scenario, stable eye movements and increased alpha waves in the EEG indicate a high level of concentration.

[0156] Physical and mental state classification: A multi-level classification model (such as SVM, random forest) or clustering algorithm (such as K-means) can be used to divide the user's state into state labels such as "high stress and low concentration" and "low stress and high concentration".

[0157] Movement standardization assessment is used to analyze the quality of the user's exercise execution to ensure that the training movements meet the standards.

[0158] Deviation Calculation: Compare the actual motion trajectory with the standard trajectory and calculate the error in joint angles. For example, in yoga training, if the user's knee angle deviates by 10° from the standard posture, the deviation value is 10.

[0159] Normativity Score: Calculates a normativity score for the action based on the deviation value. For example, the similarity of the motion trajectory can be calculated using Euclidean distance or Dynamic Time Warping (DTW). The score range can be set from 0 to 100, with 100 representing perfect conformance to the standard action.

[0160] The final output is the mapping relationship between physical and mental states, which is used for personalized behavior guidance. This relationship includes:

[0161] Real-time state vector: includes indicators such as stress index, concentration, fatigue level, and movement coordination, which serve as a quantitative representation of the user's current state.

[0162] Adaptive Optimization: Adjusts training intensity and prompts based on real-time status. For example, if a user's stress index is too high, the system can automatically reduce the training load or provide relaxation suggestions.

[0163] This embodiment uses multimodal data fusion and time series modeling to more accurately assess a user's physical and mental state. It also features cross-modal feature extraction, improving state recognition accuracy. Furthermore, through an attention mechanism and personalized state mapping, evaluation criteria can be dynamically adjusted, enabling more intelligent physical and mental state modeling.

[0164] In one embodiment, the above S30 includes:

[0165] S301, determining the user's current behavior pattern based on the historical behavior feature vector in the user portrait and the real-time state vector in the physical and mental state mapping relationship;

[0166] S302, determining a basic three-dimensional scene model that matches the current behavior pattern of the user from the three-dimensional scene model library;

[0167] S303, adjusting the ambient light intensity parameters and weather simulation parameters in the basic three-dimensional scene model according to the pressure index in the dynamic attention parameter;

[0168] S304, generating path markings and synchronized virtual tutor actions in the basic three-dimensional scene model according to the concentration weight in the dynamic attention parameter;

[0169] S305, integrating the environmental light and shadow intensity parameters, weather simulation parameters, path markings, and synchronized virtual instructor movements to generate an immersive interactive scene;

[0170] S306: Apply motion trajectory constraints driven by a physics engine to the virtual objects in the immersive interactive scene.

[0171] In this embodiment, during the construction of an immersive interactive environment, to enhance the user's sense of immersion and interactive experience, the system integrates a 3D scene model library and adaptively adjusts attention parameters based on the mind-body state mapping relationship. Different users have different behavioral patterns, mental and physical states, and levels of focus. Therefore, the interactive scene needs to be adaptable to match the user's current needs, optimizing visual, auditory, and interactive feedback to ensure a natural and smooth interactive experience.

[0172] In a virtual environment, the user's interaction method is closely related to their long-term behavioral characteristics and current physical and mental state. The user portrait contains long-term accumulated behavioral data, such as historical training records, preference choices, movement patterns, scene usage frequency, etc., while the physical and mental state mapping relationship provides current physiological indicators (such as heart rate, skin conductivity) and movement status (such as limb stability, gait characteristics). In order to determine the user's current behavior pattern, it is necessary to comprehensively analyze these two information sources. A behavior prediction model based on a long short-term memory network (LSTM) can be used to input the historical behavior feature vector and the real-time state vector into the network, and predict the user's behavior trend through time series analysis. For example, if the user tends to choose a relaxing scene when he is under high pressure and low concentration, the system can give priority to recommending an environment that adapts to this state in subsequent interactions.

[0173] The 3D scene model library stores multiple virtual scene templates, such as forests, beaches, snow-capped mountains, and quiet meditation rooms. Each scene has a different atmosphere and adaptability. Based on the determined user behavior patterns, the most suitable scene model needs to be selected. For example, in low-stress, high-focus states, clear, high-contrast scenes, such as modern office spaces, are suitable; while in high-stress states, softer, soothing scenes, such as natural scenery and yoga studios, are more suitable. The cosine similarity matching algorithm can be used to compare the user behavior pattern vector with the scene feature vector, calculate the adaptability score of each scene, and select the best matching scene model.

[0174] The user's stress index affects their tolerance for ambient lighting and weather, so the basic scene model needs to be adjusted based on the stress index. If the user's stress index is high, the system can reduce the ambient lighting intensity to make the scene lighting softer and avoid overstimulation. For example, the proportion of blue light can be reduced, warm light sources can be increased, and the brightness and darkness of the scene can be adjusted through real-time lighting rendering technology (such as physically based rendering PBR) to make it more comfortable. Similarly, weather simulation parameters can also be adjusted as the stress index changes. For example, in a high-stress state, the scene can add dynamic elements such as breeze and flowing water to enhance the relaxing effect, while in a low-stress state, the morning sun and high-contrast light and shadow effects can be simulated to improve concentration. These adjustments can be implemented based on real-time environment rendering engines (such as Unity HDRP or Unreal Engine Lumen) to make the light and shadow changes more natural.

[0175] The focus weight determines the user's level of attention, influencing their need for visual guidance and motion synchronization. During high-focus states, path markings should be simplified to reduce distractions, such as using low-transparency ground markings or providing guidance only at key junctures. During low-focus states, path guidance needs to be enhanced, for example, through animation, flow effects, brightness enhancement, or color contrast to make the path more intuitive and ensure the user can clearly perceive the route. Path markings can be generated using the A* search algorithm or a Bezier curve-based path optimization algorithm to ensure the path's rationality and naturalness. Furthermore, the virtual tutor's motion synchronization needs to be dynamically adjusted based on the user's focus. During high-focus states, the tutor's movements should minimize detail and provide only key guidance. During low-focus states, the tutor should incorporate features such as body amplification, voice prompts, and posture decomposition to enhance interaction comprehensibility. Computer vision posture analysis (such as OpenPose or MediaPipe) can be used to evaluate user movements in real time and dynamically adjust the virtual tutor's feedback based on deviations.

[0176] After completing the adjustments to all environmental and interactive elements, these parameters need to be integrated to ultimately generate an immersive interactive scene personalized for the user. The key to this step is the optimization of the scene rendering engine to ensure that different adjustment factors can be seamlessly combined. For example, ambient lighting adjustments cannot affect the visibility of path markers, and weather simulations cannot interfere with the visibility of virtual mentors. To this end, a hierarchical rendering strategy can be used to divide different types of interactive elements into different rendering layers, so that path markers, virtual mentors, and ambient lighting can be adjusted independently. In addition, physical interaction optimization (such as based on Unity Physics or Havok Physics) can be used to ensure that objects in the scene can correctly respond to user interactions. For example, when the user moves along the path markers, environmental elements (such as leaves and lake ripples) can also make adaptive feedback.

[0177] To enhance the realism of the scene, the movement of all virtual objects must conform to the laws of physics. For example, in a deep relaxation environment, candlelight gently sways with the flow of virtual air, creating a soft and natural atmosphere. In a tranquil forest scene, leaves sway gently in the wind, and sunlight filters through them, creating dappled light and shadows. At the edge of a lake or stream, the water ripples and ripples in response to the user's footsteps or the contact of virtual objects. These physical effects not only enhance visual realism but also enhance the user's immersive experience through environmental feedback, making the user feel more natural and relaxed in the virtual wellness environment. These effects can be achieved using rigid-body physics engines such as PhysX, Bullet, or Unreal Engine Chaos. Physics engines calculate the motion of objects in the environment, ensuring that they conform to the laws of nature, such as gravity, inertia, and air resistance. For example, when a user performs Tai Chi movements, the system can analyze the force of the user's palm push and calculate the air resistance feedback, causing the leaves or airflow in the virtual scene to move accordingly. In addition, in order to ensure the smoothness of the motion trajectory, Bezier curves or Kalman filters can be used to optimize the trajectory to make the interactive experience smoother and more natural.

[0178] In this embodiment, by combining the three-dimensional scene model library and the mapping relationship between physical and mental states, the system can dynamically adapt to the user's physiological and psychological state and generate a personalized immersive interactive scene. It can adjust the ambient lighting and weather according to the stress index, so that users can get a more comfortable interactive experience when they are under high pressure and get clearer visual feedback when they are under low pressure. In addition, through the adaptive adjustment of path markings and virtual tutors, the system can provide more accurate interactive guidance and improve the user's learning efficiency and task completion in different states. At the same time, the introduction of the physics engine enables objects in the virtual scene to conform to real physical laws, improve the sense of immersion and interactive authenticity, and enable users to integrate into the virtual world more naturally.

[0179] In one embodiment, the above S40 includes:

[0180] S401, extracting an entity relationship graph structure predefined in the domain knowledge graph, wherein the entity relationship graph structure includes scene element entities, behavior pattern entities, and association attributes between scene element entities and behavior pattern entities;

[0181] S402, extracting a current scene element entity set from the immersive interactive scene, wherein the current scene element entity set includes a virtual tutor entity, a path marking entity, and an interactive prop entity;

[0182] S403, determining a target interaction pattern entity based on a semantic similarity matching result between the user's current behavior pattern and the behavior pattern entity in the domain knowledge graph;

[0183] S404, generating a scene interaction instruction logic parameter bound to the target interaction mode entity according to the association attribute between the scene element entity and the behavior mode entity;

[0184] S405 , performing timestamp interpolation alignment and space coordinate mapping processing on the scene interaction instruction logic parameters and the current scene element entity set to generate an executable scene interaction instruction queue.

[0185] In this embodiment, in an immersive interactive environment, in order to ensure that users can interact smoothly with the virtual scene, it is necessary to parse the interaction instructions in the scene based on the domain knowledge graph. The domain knowledge graph contains rich structured information that can describe the elements, behavior patterns, and the relationships between them in the scene, thereby realizing intelligent and dynamic interaction logic. The key to parsing interaction instructions is to correctly match the user's current behavior pattern, combine the entity relationships in the scene, and dynamically generate reasonable interaction instructions to ensure that the execution of the instructions is in line with the user's habits and enhances the immersive experience.

[0186] The knowledge graph is a data representation method based on a graph structure that can describe the relationship between different entities. The core of the domain knowledge graph is composed of scene element entities, behavior pattern entities, association attributes, constraint rules, etc. Scene element entities refer to the key interactive elements in the virtual scene, such as virtual mentors, path markers, interactive props, etc.; behavior pattern entities describe the user's interaction methods in different states, such as relaxation, walking, meditation, etc.; association attributes define the interaction relationship between scene elements and behavior patterns. For example, when the user is performing a walking task, the path marker should be highlighted, while in the relaxed state, the path marker should be faded. In order to load the knowledge graph structure, a graph database (such as Neo4j, DGL, GraphDB) can be used to store and retrieve data, and the interaction relationship can be obtained through a query language based on SPARQL or Cypher to provide basic data for subsequent analysis.

[0187] The elements in the scene will change with the user's interaction state, so it is necessary to extract the entity set in the current scene in real time. The extracted entities include but are not limited to:

[0188] Virtual tutor entity: an intelligent entity that provides guidance through voice, movement, etc.

[0189] Path marker entity: used to indicate the user's route in the interactive environment;

[0190] Interactive prop entities: such as handheld devices, virtual items, etc., which can directly interact with users.

[0191] These entities can be extracted by using computer vision-based target detection technology (such as YOLO and FasterR-CNN), combined with the scene management function of real-time rendering engines (such as Unity and Unreal Engine) to determine the interactive objects that are active in the current frame and build a list of real-time scene elements.

[0192] The user's behavior pattern is the core factor that determines the interaction logic, and the behavior patterns of different users will affect their needs for interaction methods. For example, in deep relaxation mode, users need smoother and slower interactions, such as soft background music and slow breathing guidance; while in focus training mode, users may need more efficient and accurate instruction feedback, such as clear action instructions and immediate voice prompts. In order to determine the appropriate interaction mode, it is necessary to calculate the semantic similarity between the user's current behavior pattern and the behavior pattern entity in the knowledge graph. You can use pre-trained language models (such as BERT, Word2Vec) to calculate the cosine similarity between behavior pattern description texts, or classify historical behavior data through cluster analysis (such as K-means) to find the behavior pattern entity that best matches the user's current state.

[0193] After determining the target interaction mode, you need to query the scene interaction instruction logic parameters corresponding to the mode from the knowledge graph. These logic parameters include:

[0194] Voice prompt content parameters: Determines the voice instructions of the virtual instructor. For example, in the walking task, the instructor may provide pace suggestions, while in relaxation mode, the instructor may provide breathing guidance.

[0195] Action trigger condition parameters: define key actions in the scene, such as triggering the instructor's action demonstration when the user reaches a specific path marker;

[0196] Prop interaction permission parameters: used to control whether users can interact with certain props. For example, in a specific mode, the virtual device may need to be locked to prevent accidental operation.

[0197] The generation of these logical parameters can be based on the query mechanism of the knowledge graph (such as GNN graph neural network reasoning), or use a rule engine (such as Drools) to parse predefined interaction rules to ensure that the generated interaction instructions meet the user's current needs.

[0198] To ensure the accuracy of the interaction, it is necessary to ensure that the generated interaction instructions match the scene in both time and space dimensions:

[0199] Timestamp interpolation and alignment: Because different sensor data (such as user actions and voice input) may have time synchronization deviations, interactive commands need to be interpolated and aligned to ensure that the commands are triggered at the appropriate time. Kalman filtering or linear interpolation can be used to correct timestamps to make the time distribution of commands smoother.

[0200] Spatial coordinate mapping: Different scene elements may be in different spatial locations. For example, the coordinates of path markers may need to match the user's location to ensure that the user sees the correct guidance. The spatial transformation matrix can be used to transform the coordinates of scene elements so that the interaction logic is consistent with the user's real-time location.

[0201] Finally, the processed interactive instructions are combined into an ordered instruction queue to ensure that the order and trigger conditions of the interaction execution meet the expected conditions. For example, if the user stays at a path marker for more than a set time, the next instruction in the queue will trigger a voice prompt from the virtual instructor, prompting the user to continue forward.

[0202] This embodiment can ensure the rationality and adaptability of the interaction logic by parsing scene interaction instructions based on the domain knowledge graph. The behavior patterns of different users can be dynamically matched to the most suitable interaction method, making the immersive experience more natural and personalized. The introduction of the knowledge graph greatly reduces the hard coding of interaction rules, giving the system stronger scalability and automatic reasoning capabilities. In addition, through timestamp interpolation alignment and spatial coordinate mapping, the accuracy of interaction instructions is ensured, and the user's interaction fluency and feedback timeliness in immersive scenes are improved.

[0203] In one embodiment, the above S50 includes:

[0204] S501, performing missing value filling and standardization processing on the historical behavior feature vector in the user portrait to generate standardized data of the historical behavior feature vector;

[0205] S502, analyzing a weight distribution coefficient of the real-time state vector according to the stress index and the concentration weight of the real-time state vector;

[0206] S503, performing a product operation on the weight distribution coefficient and the real-time state vector to generate real-time state vector weighted data;

[0207] S504, inputting the normalized data of the historical behavior feature vector and the weighted data of the real-time state vector into a multimodal fusion model based on an attention mechanism to generate a fusion feature vector;

[0208] S505, analyzing the cosine similarity between the fused feature vector and the candidate guidance content feature vectors in the guidance content library to obtain a matching score;

[0209] S506, screening candidate guidance contents whose matching scores exceed a preset score threshold in the guidance content library, and generating a comprehensive decision parameter including a priority label and an output confidence according to the matching scores of the candidate guidance contents in descending order, wherein each entry in the comprehensive decision parameter is associated with a candidate guidance content.

[0210] In this embodiment, in order to provide accurate personalized guidance content, the personalized interaction system integrates the user's historical behavioral characteristics with their real-time physical and mental state to generate decision parameters tailored to the user's current state. The core of calculating these comprehensive decision parameters lies in rationally combining long-term behavioral preferences with current physiological and psychological states, and optimizing the accuracy of content screening through multimodal fusion technology. This process involves multiple key steps, including data preprocessing, state analysis, weighted calculation, multimodal fusion, and personalized matching.

[0211] User portrait data comes from long-term accumulated interaction records, including historical training patterns, preferred scenarios, past interaction habits, etc. These data may have missing values, or the dimensions of different features may be inconsistent, so preprocessing is required. Missing value filling can use interpolation methods based on K-nearest neighbor (KNN) or Gaussian mixture model (GMM) to fill in missing behavior records and make the data distribution more complete. Standardization processing uses Z-score normalization or Min-Max scaling to convert different features into the same dimension to avoid the impact of different features on subsequent calculations. For example, the user's previous stay time in the meditation scene may be measured in minutes, while the response time to the interaction prompt may be measured in seconds. Standardization can enable these data to be calculated on the same scale, improving the comparability of fused features.

[0212] The user's current physical and mental state has a direct impact on their interaction needs, so it's necessary to calculate a weight distribution coefficient for their state to determine the degree of influence their current state has on the overall decision-making process. The stress index reflects the user's level of anxiety, while the concentration weight reflects the user's level of focus. The weight distribution coefficient can be calculated using a fuzzy logic-based weighting strategy: when the stress index is high, the recommendation weight for content with high cognitive load is reduced; when the concentration weight is high, the weight for content with deep interaction is increased. For example, if the user's stress index is high, the system may reduce the probability of recommending content with high information density and prioritize relaxation exercises or simplified interaction tasks.

[0213] The calculated weight distribution coefficients need to be applied to the real-time state vector to adjust its contribution to subsequent fusion decisions. The weight distribution coefficients can be applied to each state vector dimension using the Hadamard product (element-by-element multiplication). For example, if the stress index has a weight of 0.7 and the concentration has a weight of 0.3, the calculated weighted data will enhance the perception of the stress state while appropriately considering the user's concentration state, ensuring that subsequent decision logic can both respond to the user's current stress state and not ignore their attention level.

[0214] During the multimodal fusion phase, historical behavior data must be deeply integrated with current status data to achieve a more accurate representation of the user's state. Traditional linear weighting methods struggle to capture complex interaction patterns, so a multimodal fusion model based on an attention mechanism is employed. This model uses a Transformer architecture or a Bi-LSTM + Self-Attention structure to learn the correlation between historical data and real-time status. For example, if a user has historically preferred a specific type of training mode and their current state matches this preference, the model automatically increases the weight of that behavior feature, ensuring that the fused feature vector better meets the user's current needs.

[0215] The fused user feature vector needs to be matched against candidate content in the guidance content library to select personalized guidance content suitable for the user's current state. Cosine similarity can be used to measure the similarity between the user's features and the features of each candidate content. For example, if the user's current state has a high similarity to meditation guidance content, the matching score for this type of content is high. This matching score is used to measure the suitability of the content and serves as a basis for subsequent screening.

[0216] The matching score screening threshold is used to filter out low-relevance content that does not match the user's status. For example, if the preset threshold is 0.6, only candidate content with a cosine similarity greater than 0.6 will be retained. The filtered candidate content is sorted in descending order according to the matching score, and a priority label and output confidence are generated:

[0217] Priority label: used to distinguish the recommendation order of different content. The higher the score, the higher the priority.

[0218] Output confidence: Softmax normalization can be used to calculate the confidence of all recommended content, so that the confidence of all recommended content is distributed between 0 and 1. For example, if the matching degree of a certain guidance content is 0.85, while the matching degree of other content is lower, the confidence of this content is higher and the system will give it priority in the interaction process.

[0219] Finally, the comprehensive decision parameters include all the screened candidate guidance contents, and each entry is associated with the corresponding guidance content features, providing data support for subsequent personalized recommendations.

[0220] This embodiment can generate personalized decision parameters that better meet the user's current needs by fusing the long-term behavioral data of the user portrait with the real-time physiological state. Historical behavioral data provides the user's long-term preferences, while real-time status data ensures that the recommended content can adapt to the user's current physiological and psychological conditions. The use of a multimodal fusion model based on the attention mechanism can automatically learn the relationship between the two and improve the accuracy of personalized recommendations. In addition, by calculating the matching score through cosine similarity and combining it with a dynamic threshold screening mechanism, it can effectively filter out irrelevant content, ensuring that the recommended content is both accurate and in line with the user's status, thereby improving the intelligence level of the interactive experience.

[0221] In one embodiment, the above S60 includes:

[0222] S601, extracting the matching score, priority label, and output confidence from the comprehensive decision parameters, and extracting the voice prompt content parameters, action trigger condition parameters, and prop interaction permission parameters from the scene interaction instruction;

[0223] S602, performing feature alignment processing on the matching score in the comprehensive decision parameter and the action trigger condition parameter in the scene interaction instruction to generate a comprehensive matching score vector;

[0224] S603, analyzing the cosine similarity between the comprehensive matching score vector and the scene adaptation vector of the candidate guidance content in the guidance content library to obtain a scene interaction feature similarity score;

[0225] S604: Generate a real-time comprehensive score for each candidate guidance content based on a linear weighted result of the output confidence and the scene interaction feature similarity score;

[0226] S605: Generate an adaptive threshold based on the stress index of the real-time state vector, screen candidate guidance content with real-time comprehensive scores exceeding the adaptive threshold, and generate a personalized guidance content set including priority tags in descending order of the real-time comprehensive scores;

[0227] S606: Perform interaction conflict detection on the personalized guidance content set to remove guidance content in the personalized guidance content set that conflicts with the prop interaction permission parameter in the scene interaction instruction, and generate final personalized guidance content.

[0228] In this embodiment, in an immersive interactive system, to ensure that personalized guidance recommendations are both relevant to the user's current state and compatible with the scenario interaction logic, a screening process is performed based on the feature similarity between comprehensive decision parameters and scenario interaction instructions. This screening process involves key steps such as multi-dimensional feature alignment, similarity calculation, weight fusion, threshold adjustment, and conflict detection to ensure the accuracy, adaptability, and consistency of recommended content.

[0229] Comprehensive decision parameters are personalized recommendation data generated by integrating user profiles and real-time status in the previous steps, including:

[0230] Matching score: measures the degree of fit between the user's current state and a certain guidance content;

[0231] Priority tag: determines the order of recommendation, with content with higher scores being output first;

[0232] Output confidence: Indicates the system's trust in the recommended content, usually calculated through Softmax normalization.

[0233] At the same time, scene interaction instructions provide the interaction logic in the scene, including:

[0234] Voice prompt content parameters: define the voice prompts that need to be triggered in the scene, such as "Please adjust your sitting posture";

[0235] Action trigger condition parameters: describe the prerequisites for triggering an interaction, such as "trigger when the user reaches a specified path marker";

[0236] Prop interaction permission parameters: Specifies whether users can interact with certain interactive props, such as "this prop is only allowed to be operated by a specific user group."

[0237] These parameters can be extracted through database queries (SQL or NoSQL) or data retrieval through API interfaces based on JSON parsing, ensuring that parameters from different sources can be uniformly formatted.

[0238] To ensure that the personalized recommendation content matches the interaction logic in the scene, feature alignment is required:

[0239] Timeline alignment: Different parameters may have different data update frequencies. For example, comprehensive decision parameters may be updated once per second, while scene interaction instructions may be refreshed every five seconds. Therefore, linear interpolation or time window sliding average methods are required for time alignment.

[0240] Spatial coordinate alignment: Some interactive instructions (such as action trigger conditions) involve specific spatial locations. The applicable range of the matching score needs to be mapped to the three-dimensional coordinate system of the scene, and matching is performed using a coordinate transformation matrix or nearest neighbor search (k-NN).

[0241] Finally, the data after feature alignment are combined into a comprehensive matching score vector for subsequent calculations.

[0242] To measure the suitability of the recommended content for the scenario, we need to calculate the similarity between the comprehensive matching score vector and the scenario suitability vector of the candidate guidance content. The higher the similarity score, the more compatible the candidate content is with the current scenario.

[0243] In order to balance the impact of user state adaptation (output confidence) and scene adaptation (similarity score), a linear weighted calculation is required:

[0244] Real-time comprehensive score = γ × output confidence + (1-γ) × scene interaction feature similarity score

[0245] Where: γ is the balance coefficient, which is calculated by the pressure index and concentration weight in the real-time state vector:

[0246] γ=α×(1-P)+β×A

[0247] P is the stress index. A higher value indicates higher user stress, which reduces the priority of complex interactions. A is the focus weight. A higher value indicates more focused user attention, which increases the priority of deep interactions. α and β are adjustable parameters used to control the impact of different state factors.

[0248] Through this calculation method, it can be ensured that the selection of recommended content not only considers the adaptability of the user status, but also considers its adaptability in the current interaction scenario.

[0249] Different user states have different interaction requirements, so the filtering threshold needs to be adjusted dynamically:

[0250] Threshold = Tbase - λ × P

[0251] Where: Tbase is the default recommended threshold; λ is the adjustment coefficient, which controls the impact of the pressure index on the threshold.

[0252] When the stress index is high, the threshold is lowered, allowing more guidance content suitable for relaxation training to be screened out. When the stress index is low, the threshold is raised, making the recommended content more precise. Finally, the screened content is sorted in descending order by real-time comprehensive score to generate a personalized guidance content collection, and priority tags are added.

[0253] To ensure that the recommended content does not conflict with the scene interaction rules, conflict detection is required:

[0254] Item permission conflict detection: Checks whether the recommended content involves restricted items. If the user does not currently have permission to use it, the content is removed.

[0255] Time period conflict detection: Ensures that the applicable time of the recommended content is consistent with the scene interaction instructions. For example, if a certain instruction content needs to be executed within a specific time period, but the current time is out of range, it will not be recommended;

[0256] Action logic conflict detection: If the body movements involved in the recommended content conflict with the guided actions set in the scene, the content will be reordered or replaced.

[0257] Conflict detection can adopt rule-based verification (if-else logic) or knowledge graph-based logical reasoning (such as OWL-based reasoning) to ensure that the final output guidance content is consistent with the user status and scenario logic.

[0258] This embodiment can significantly improve the accuracy of personalized guidance content by screening based on the feature similarity of comprehensive decision parameters and scene interaction instructions, so that the recommended content not only meets the user's physical and mental state, but also adapts to the interaction needs of the current scene. By calculating the real-time comprehensive score through linear weighting, the user state adaptability and scene adaptability are balanced, making the recommended content more personalized and real-time. Adaptive threshold screening ensures dynamic adjustment of recommendation strategies under different states, improving the flexibility of recommended content. In addition, the interactive conflict detection mechanism can effectively avoid conflicts between recommended content and scene rules, ensure the smoothness and consistency of the interaction process, and improve the intelligence level of the immersive experience.

[0259] In one embodiment, a guidance device based on multimodal perception is provided, and the guidance device based on multimodal perception corresponds one-to-one with the guidance method based on multimodal perception in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of a multimodal perception-based guidance device according to the present invention. These modules include a physiological and motion data acquisition module 10, a physical and mental state modeling module 20, an immersive scene generation module 30, an interactive instruction parsing module 40, a personalized behavior analysis module 50, a guidance content matching module 60, and a personalized content output module 70. Each functional module is described in detail below:

[0260] The physiological and motion data acquisition module 10 collects the user's physiological index data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0261] The body-mind state modeling module 20 inputs the multimodal body data set into a pre-trained body-mind association model to generate a body-mind state mapping relationship;

[0262] An immersive scene generation module 30 generates an immersive interactive scene including environmental simulation elements based on a three-dimensional scene model library and in combination with dynamic attention parameters in the mind-body state mapping relationship;

[0263] An interaction instruction parsing module 40 parses the scene interaction instructions in the immersive interaction scene according to the entity association relationship in the domain knowledge graph;

[0264] The personalized behavior analysis module 50 integrates the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0265] A guidance content matching module 60 selects personalized guidance content from a guidance content library based on the similarity between the comprehensive decision parameters and the characteristics of the scene interaction instructions;

[0266] The personalized content output module 70 outputs the personalized guidance content.

[0267] In one embodiment, the physiological and motion data acquisition module 10 is specifically configured to:

[0268] Collecting the user's skin conductivity data and blood oxygen saturation data through a wearable smart device to generate a first physiological indicator subset;

[0269] The user's torso acceleration data and limb joint angular velocity data are collected through an inertial measurement unit to generate a three-dimensional motion trajectory subset;

[0270] The user's facial expression change data and body movement amplitude data are collected through an optical motion capture device to generate a motion feature subset;

[0271] Performing timestamp alignment processing on the data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset;

[0272] Signal filtering is performed on the first physiological indicator subset, three-dimensional motion trajectory subset, and action feature subset after timestamp alignment to generate a standardized multimodal body dataset.

[0273] In one embodiment, the physical and mental state modeling module 20 is specifically configured to:

[0274] Dividing the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset in the multimodal body dataset into time series segments according to a preset time window;

[0275] Inputting the time series segments into a mind-body association model constructed based on a gated recurrent unit, and extracting cross-modal association features between the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset through the mind-body association model;

[0276] Based on the cross-modal correlation features, generating a mind-body state mapping matrix including dynamic attention parameters and joint angle thresholds at the output layer of the mind-body correlation model;

[0277] Determining the user's physical and mental state level based on the stress index and concentration weight in the dynamic attention parameter;

[0278] generating a motion normativeness evaluation coefficient based on a deviation value between the joint angle threshold and the motion trajectory data in the three-dimensional motion trajectory subset;

[0279] The physical and mental state levels and the action normative evaluation coefficients are integrated into a physical and mental state mapping relationship including a real-time state vector.

[0280] In one embodiment, the immersive scene generation module 30 is specifically configured to:

[0281] Determining the user's current behavior pattern based on the historical behavior feature vector in the user portrait and the real-time state vector in the mapping relationship between the physical and mental state;

[0282] Determining a basic three-dimensional scene model that matches the user's current behavior pattern from the three-dimensional scene model library;

[0283] Adjusting the ambient light intensity parameters and weather simulation parameters in the basic three-dimensional scene model according to the pressure index in the dynamic attention parameter;

[0284] generating path markings and synchronized virtual tutor actions in the basic three-dimensional scene model according to the concentration weight in the dynamic attention parameter;

[0285] The environmental light and shadow intensity parameters, weather simulation parameters, path markings, and synchronized virtual instructor movements are integrated to generate an immersive interactive scene;

[0286] Applying motion trajectory constraints driven by a physics engine to virtual objects in the immersive interactive scene.

[0287] In one embodiment, the interaction instruction parsing module 40 is specifically configured to:

[0288] Extracting an entity relationship graph structure predefined in the domain knowledge graph, wherein the entity relationship graph structure includes scene element entities, behavior pattern entities, and association attributes between scene element entities and behavior pattern entities;

[0289] Extracting a current scene element entity set from the immersive interactive scene, the current scene element entity set including a virtual mentor entity, a path marking entity, and an interactive prop entity;

[0290] Determine the target interaction pattern entity based on the semantic similarity matching result between the user's current behavior pattern and the behavior pattern entity in the domain knowledge graph;

[0291] Generate scene interaction instruction logic parameters bound to the target interaction mode entity according to the association attributes between the scene element entity and the behavior mode entity;

[0292] The scene interaction instruction logic parameters are aligned with the current scene element entity set through timestamp interpolation and spatial coordinate mapping to generate an executable scene interaction instruction queue.

[0293] In one embodiment, the personalized behavior analysis module 50 is specifically configured to:

[0294] Perform missing value filling and standardization on the historical behavior feature vectors in the user portrait to generate standardized data of the historical behavior feature vectors;

[0295] Analyzing a weight distribution coefficient of the real-time state vector according to the stress index and the concentration weight of the real-time state vector;

[0296] Performing a product operation on the weight distribution coefficient and the real-time state vector to generate real-time state vector weighted data;

[0297] Inputting the normalized data of the historical behavior feature vector and the weighted data of the real-time state vector into a multimodal fusion model based on an attention mechanism to generate a fusion feature vector;

[0298] Analyzing the cosine similarity between the fused feature vector and the candidate guidance content feature vectors in the guidance content library to obtain a matching score;

[0299] Candidate guidance contents whose matching scores exceed a preset score threshold are screened in the guidance content library, and a comprehensive decision parameter including a priority label and an output confidence is generated according to the matching scores of the candidate guidance contents in descending order, wherein each entry in the comprehensive decision parameter is associated with a candidate guidance content.

[0300] In one embodiment, the guidance content matching module 60 is specifically configured to:

[0301] Extracting the matching score, priority label, and output confidence from the comprehensive decision parameters, and extracting the voice prompt content parameters, action trigger condition parameters, and prop interaction permission parameters from the scene interaction instructions;

[0302] Performing feature alignment processing on the matching score in the comprehensive decision parameter and the action trigger condition parameter in the scene interaction instruction to generate a comprehensive matching score vector;

[0303] Analyzing the cosine similarity between the comprehensive matching score vector and the scene adaptation vector of the candidate guidance content in the guidance content library to obtain a scene interaction feature similarity score;

[0304] Generating a real-time comprehensive score for each candidate guidance content according to a linear weighted result of the output confidence and the scene interaction feature similarity score;

[0305] generating an adaptive threshold based on the stress index of the real-time state vector, screening candidate guidance content whose real-time comprehensive scores exceed the adaptive threshold, and generating a personalized guidance content set including priority tags in descending order of the real-time comprehensive scores;

[0306] Interaction conflict detection is performed on the personalized guidance content set to remove guidance content in the personalized guidance content set that conflicts with the prop interaction permission parameter in the scene interaction instruction, and generate final personalized guidance content.

[0307] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of a guidance method based on multimodal perception.

[0308] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a guidance method based on multimodal perception.

[0309] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0310] Collect the user's physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0311] Inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0312] Based on the three-dimensional scene model library and in combination with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene including environmental simulation elements is generated;

[0313] Parsing scene interaction instructions in the immersive interaction scene according to entity association relationships in the domain knowledge graph;

[0314] Fusing the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0315] Based on the feature similarity between the comprehensive decision parameter and the scenario interaction instruction, screening personalized guidance content from the guidance content library;

[0316] The personalized guidance content is output.

[0317] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0318] Collect the user's physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body data set;

[0319] Inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship;

[0320] Based on the three-dimensional scene model library and in combination with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene including environmental simulation elements is generated;

[0321] Parsing scene interaction instructions in the immersive interaction scene according to entity association relationships in the domain knowledge graph;

[0322] Fusing the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters;

[0323] Based on the feature similarity between the comprehensive decision parameter and the scenario interaction instruction, screening personalized guidance content from the guidance content library;

[0324] The personalized guidance content is output.

[0325] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0326] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0327] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0328] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A guidance method based on multimodal perception, characterized in that: The following steps are involved: Collect the user's physiological indicator data and three-dimensional motion trajectory data to generate a multimodal body data set; Inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship; Based on the three-dimensional scene model library and in combination with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene including environmental simulation elements is generated; Parsing scene interaction instructions in the immersive interaction scene according to entity association relationships in the domain knowledge graph; Fusing the historical behavior feature vector in the user portrait with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters; Based on the feature similarity between the comprehensive decision parameter and the scenario interaction instruction, screening personalized guidance content from the guidance content library; The personalized guidance content is output.

2. The multimodal perception-based guidance method according to claim 1, wherein: Collect the user's physiological indicator data and 3D motion trajectory data to generate a multimodal body data set, including: Collecting the user's skin conductivity data and blood oxygen saturation data through a wearable smart device to generate a first physiological indicator subset; The user's torso acceleration data and limb joint angular velocity data are collected through an inertial measurement unit to generate a three-dimensional motion trajectory subset; The user's facial expression change data and body movement amplitude data are collected through an optical motion capture device to generate a motion feature subset; Performing timestamp alignment processing on the data in the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset; Signal filtering is performed on the first physiological indicator subset, three-dimensional motion trajectory subset, and action feature subset after timestamp alignment to generate a standardized multimodal body dataset.

3. The multimodal perception-based guidance method according to claim 1, wherein: Inputting the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship, including: Dividing the first physiological indicator subset, the three-dimensional motion trajectory subset, and the motion feature subset in the multimodal body dataset into time series segments according to a preset time window; Inputting the time series segments into a mind-body association model constructed based on a gated recurrent unit, and extracting cross-modal association features between the first physiological indicator subset, the three-dimensional motion trajectory subset, and the action feature subset through the mind-body association model; Based on the cross-modal correlation features, generating a mind-body state mapping matrix including dynamic attention parameters and joint angle thresholds at the output layer of the mind-body correlation model; Determining the user's physical and mental state level based on the stress index and concentration weight in the dynamic attention parameter; generating a motion normativeness evaluation coefficient based on a deviation value between the joint angle threshold and the motion trajectory data in the three-dimensional motion trajectory subset; The physical and mental state levels and the action normative evaluation coefficients are integrated into a physical and mental state mapping relationship including a real-time state vector.

4. The multimodal perception-based guidance method according to claim 1, wherein: Based on the three-dimensional scene model library and combined with the dynamic attention parameters in the mind-body state mapping relationship, an immersive interactive scene containing environmental simulation elements is generated, including: Determining the user's current behavior pattern based on the historical behavior feature vector in the user portrait and the real-time state vector in the mapping relationship between the physical and mental state; Determining a basic three-dimensional scene model that matches the user's current behavior pattern from the three-dimensional scene model library; Adjusting the ambient light intensity parameters and weather simulation parameters in the basic three-dimensional scene model according to the pressure index in the dynamic attention parameter; generating path markings and synchronized virtual tutor actions in the basic three-dimensional scene model according to the concentration weight in the dynamic attention parameter; The environmental light and shadow intensity parameters, weather simulation parameters, path markings, and synchronized virtual instructor movements are integrated to generate an immersive interactive scene; Applying motion trajectory constraints driven by a physics engine to virtual objects in the immersive interactive scene.

5. The multimodal perception-based guidance method according to claim 1, wherein: Parsing the scene interaction instructions in the immersive interaction scene according to the entity association relationship in the domain knowledge graph includes: Extracting an entity relationship graph structure predefined in the domain knowledge graph, wherein the entity relationship graph structure includes scene element entities, behavior pattern entities, and association attributes between scene element entities and behavior pattern entities; Extracting a current scene element entity set from the immersive interactive scene, the current scene element entity set including a virtual mentor entity, a path marking entity, and an interactive prop entity; Determine the target interaction pattern entity based on the semantic similarity matching result between the user's current behavior pattern and the behavior pattern entity in the domain knowledge graph; Generate scene interaction instruction logic parameters bound to the target interaction mode entity according to the association attributes between the scene element entity and the behavior mode entity; The scene interaction instruction logic parameters are aligned with the current scene element entity set through timestamp interpolation and spatial coordinate mapping to generate an executable scene interaction instruction queue.

6. The multimodal perception-based guidance method according to claim 1, wherein: The historical behavior feature vector in the user profile is integrated with the real-time state vector in the physical and mental state mapping relationship to generate comprehensive decision parameters, including: Perform missing value filling and standardization on the historical behavior feature vectors in the user portrait to generate standardized data of the historical behavior feature vectors; Analyzing a weight distribution coefficient of the real-time state vector according to the stress index and the concentration weight of the real-time state vector; Performing a product operation on the weight distribution coefficient and the real-time state vector to generate real-time state vector weighted data; Inputting the normalized data of the historical behavior feature vector and the weighted data of the real-time state vector into a multimodal fusion model based on an attention mechanism to generate a fusion feature vector; Analyzing the cosine similarity between the fused feature vector and the candidate guidance content feature vectors in the guidance content library to obtain a matching score; Candidate guidance contents whose matching scores exceed a preset score threshold are screened in the guidance content library, and a comprehensive decision parameter including a priority label and an output confidence is generated according to the matching scores of the candidate guidance contents in descending order, wherein each entry in the comprehensive decision parameter is associated with a candidate guidance content.

7. The multimodal perception-based guidance method according to claim 1, wherein: Based on the feature similarity between the comprehensive decision parameter and the scenario interaction instruction, personalized guidance content is screened from the guidance content library, including: Extracting the matching score, priority label, and output confidence from the comprehensive decision parameters, and extracting the voice prompt content parameters, action trigger condition parameters, and prop interaction permission parameters from the scene interaction instructions; Performing feature alignment processing on the matching score in the comprehensive decision parameter and the action trigger condition parameter in the scene interaction instruction to generate a comprehensive matching score vector; Analyzing the cosine similarity between the comprehensive matching score vector and the scene adaptation vector of the candidate guidance content in the guidance content library to obtain a scene interaction feature similarity score; Generating a real-time comprehensive score for each candidate guidance content according to a linear weighted result of the output confidence and the scene interaction feature similarity score; generating an adaptive threshold based on the stress index of the real-time state vector, screening candidate guidance content whose real-time comprehensive scores exceed the adaptive threshold, and generating a personalized guidance content set including priority tags in descending order of the real-time comprehensive scores; Interaction conflict detection is performed on the personalized guidance content set to remove guidance content in the personalized guidance content set that conflicts with the prop interaction permission parameter in the scene interaction instruction, and generate final personalized guidance content.

8. A guidance device based on multimodal perception, characterized in that: The multimodal sensing-based guidance device includes: Physiological and motion data acquisition module, which collects the user's physiological index data and three-dimensional motion trajectory data to generate a multimodal body data set; a mind-body state modeling module, which inputs the multimodal body data set into a pre-trained mind-body association model to generate a mind-body state mapping relationship; An immersive scene generation module generates an immersive interactive scene including environmental simulation elements based on a three-dimensional scene model library and in combination with dynamic attention parameters in the mind-body state mapping relationship; An interaction instruction parsing module, which parses the scene interaction instructions in the immersive interaction scene according to the entity association relationship in the domain knowledge graph; A personalized behavior analysis module that integrates the historical behavior feature vectors in the user portrait with the real-time state vectors in the physical and mental state mapping relationship to generate comprehensive decision parameters; A guidance content matching module, which selects personalized guidance content from a guidance content library based on the feature similarity between the comprehensive decision parameters and the scenario interaction instructions; The personalized content output module outputs the personalized guidance content.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a multimodal perception-based guidance program stored in the memory and capable of running on the processor. When the multimodal perception-based guidance program is executed by the processor, the steps of the multimodal perception-based guidance method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a guidance program based on multimodal perception, and when the guidance program based on multimodal perception is executed by the processor, the steps of the guidance method based on multimodal perception as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intelligent activity control method and system based on feature matching

    CN121052612A

  • Network security personalized teaching method and system based on SLT algorithm

    CN121213320A

  • Data processing method and device for virtual scene

    CN121371614A

  • Block chain financial identity authentication system

    CN121389090A

  • Progressive interaction guiding method and system based on behavior analysis and fuzzy matching

    CN121455587A