An interaction model training method and device based on extended reality and sensor analysis

By using an interactive model training method based on extended reality and sensor analysis, a dynamic virtual world is constructed using EEG and physiological behavior data, and training parameters are adaptively adjusted. This solves the problem of individual real-time state adaptation of interactive strategies in existing technologies, and improves the speed and flexibility of interactive response.

CN121502363BActive Publication Date: 2026-05-12SHANGHAI SHULI INTELLIGENT TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI SHULI INTELLIGENT TECH CO LTD
Filing Date
2026-01-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing interaction strategies lack real-time individual state adaptation in multi-dimensional interaction scenarios. Relying on preset rules and single-dimensional data makes it difficult to capture dynamic changes in complex scenarios, resulting in insufficient response speed and flexibility.

Method used

By using an interactive model training method based on extended reality and sensor analysis, a dynamic virtual world is constructed using EEG data and physiological behavior data, and training parameters are adaptively adjusted to achieve personalized interactive training effects.

Benefits of technology

It improves the speed and flexibility of user interaction in complex environments, enables individualized and refined adjustment of the intensity and dosage of interactive training, and enhances the effectiveness of interactive training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502363B_ABST
    Figure CN121502363B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electroencephalogram monitoring and data processing, in particular to an interactive model training method and device based on extended reality and sensing analysis, which comprises the following steps: constructing a dynamic virtual world according to virtual world generation parameters; acquiring electroencephalogram data of a user when the user is performing interactive training in the dynamic virtual world, and at least one of physiological sensing data and behavior sensing data; and determining the interactive training effect of the user; in a cyclic iteration process, when the interactive training effect does not satisfy a convergence condition, adjusting the virtual world generation parameters according to a preset prediction model, and continuously constructing a new dynamic virtual world according to the adjusted virtual world generation parameters until the interactive training effect satisfies the convergence condition, ending the cyclic iteration process, and obtaining an interactive model through training. The interactive training scheme based on virtual reality and electroencephalogram can trigger the interactive response of a user in an immersive and personalized environment, so that the effect of 'virtual adjustment' is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electroencephalogram (EEG) signal monitoring and data processing technology, and in particular to an interactive model training method and device based on extended reality and sensor analysis. Background Technology

[0002] In multi-dimensional interaction scenarios, individuals need to cope with potential threats through dynamic strategies, which typically integrate two core mechanisms: cognitive regulation and state response. Cognitive regulation focuses on optimizing behavioral patterns through experience learning and risk assessment to reduce adverse exposures; state response relies on real-time state feedback to adjust the intensity of responses. In existing technologies, the synergy of these two mechanisms mainly relies on preset rules or static models: cognitive regulation generates strategies based on a fixed experience base, lacking adaptation to the individual's real-time state; state response relies on single-dimensional data (such as simple physiological indicators), making it difficult to capture dynamic changes in complex scenarios.

[0003] Existing interaction strategy construction technologies generally suffer from core limitations such as reliance on human experience for parameter adjustment, data utilization being limited to a single modality, and a disconnect between scenario simulation and reality. These problems result in significant deficiencies in the response speed, flexibility, and individualized adaptation of existing interaction strategies, making it difficult to meet the demand for efficient interaction strategies in complex environments. Summary of the Invention

[0004] The purpose of this application is to provide an interactive model training method and device based on extended reality and sensor analysis. This method can trigger user interaction responses in an immersive and personalized environment based on virtual reality and EEG interactive training schemes, thereby achieving a "virtual regulation" effect. Furthermore, it adaptively adjusts the parameters in the interactive training according to relevant neurological, physiological, and behavioral indicators of the interactive responses, achieving the effect of adjusting the "interactive training dosage."

[0005] In some embodiments, this application provides an interactive model training method based on extended reality and sensor analysis. The method includes: determining virtual world generation parameters and user feature parameters; constructing a dynamic virtual world based on the virtual world generation parameters; acquiring EEG data, physiological sensor data, and at least one of behavioral sensor data, when the user performs interactive training in the dynamic virtual world based on the user feature parameters; determining the user's interactive training effect based on the EEG data and at least one of the physiological sensor data and behavioral sensor data; during the iterative process, when the interactive training effect does not meet the convergence condition, adjusting the virtual world generation parameters according to a preset prediction model, and continuing to construct a new dynamic virtual world based on the adjusted virtual world generation parameters until the interactive training effect meets the convergence condition, ending the iterative process and training an interactive model.

[0006] In some embodiments, determining the user's interactive training effect based on the EEG data, combined with at least one of physiological sensor data and behavioral sensor data, includes:

[0007] Based on the trigger tags set in the experimental paradigm when the event occurs, continuous EEG data is divided into multiple trials; each trial covers EEG data before and for a period of time after the event begins; the EEG data from multiple trials are superimposed and averaged to obtain event-related potentials; the positive and negative characteristics of the event-related potentials in different time windows characterize the neural activity characteristics of the electrode region.

[0008] Collect physiological sensor data of users before, after, or during interactive training, and determine the changes and differences in physiological indicators of users before, after, or during interactive training based on the corresponding detection algorithm.

[0009] And / or,

[0010] Collect behavioral sensor data or scale data of users before, after, or during interactive training, and determine the changes and differences in the user's cognitive process before, after, or during interactive training based on the corresponding response theory and its cognitive significance; determine the user's interactive training effect based on the event-related potential, the changes and differences in the physiological indicators and / or the changes and differences in the cognitive process.

[0011] In some embodiments, adjusting the virtual world generation parameters according to a preset prediction model and continuing to construct a new dynamic virtual world according to the adjusted virtual world generation parameters includes: inputting the virtual world generation parameters into a preset prediction model, predicting the interaction training effect according to the preset prediction model; when the difference between the predicted interaction training effect and the user's current actual interaction training effect meets a preset condition, predicting new virtual world generation parameters according to the preset prediction model, and continuing to construct a new dynamic virtual world according to the new virtual world generation parameters.

[0012] In some embodiments, the step of determining virtual world generation parameters includes: constructing an interactive training paradigm based on the interactive training sampling time and condition combinations; obtaining the interactive response intensity distribution of each group of subjects when receiving interactive training with corresponding condition combinations at the corresponding interactive training sampling time based on the interactive training paradigm; wherein the condition combinations include situational feature parameters and negative feature parameters in the virtual world generation parameters; obtaining the baseline value of the virtual world generation parameters based on the interactive response intensity distribution; defining safety constraints, and under the safety constraints, iteratively finding the parameter corresponding to the strongest interactive training effect based on the baseline value of the virtual world generation parameters, and using it as the final virtual world generation parameters.

[0013] In some embodiments, the step of determining the virtual world generation parameters includes: obtaining the interaction response intensity distribution characterizing the user's interaction training effect; obtaining the first condition combination with the strongest interaction training quantification index based on the interaction response intensity distribution; removing negative condition combinations from the first condition combination to obtain a target safety condition combination; using the target safety condition combination as the baseline value of the virtual world generation parameters; initializing a stochastic process model using a clustering algorithm, based on the target safety condition combination as the objective function and each safety constraint condition; predicting the virtual world generation parameters based on the stochastic process model; wherein the stochastic process model characterizes a subset of condition combinations selected from the target safety condition combination and is used as a data subset to fit a model describing the safety condition combinations.

[0014] In some embodiments, predicting the virtual world generation parameters based on the stochastic process model includes: determining the upper confidence bound of each security constraint based on the acquisition function and the security constraints, and predicting the next virtual world generation parameters; wherein the upper confidence bound of the security constraints is used to constrain the next acquisition point to achieve the target interactive training effect while meeting the security constraints; in each iterative cycle, if the predicted next virtual world generation parameters satisfy the security constraints, the objective function dataset is updated based on the predicted next virtual world generation parameters, and the stochastic process function is refitted based on the updated objective function dataset until the target convergence condition is met, thus obtaining the final virtual world generation parameters.

[0015] In some embodiments, removing negative condition combinations from the first condition combination to obtain a target safe condition combination includes: structuring the feedback information of the subjects in the interaction training paradigm according to a pre-trained language model to obtain negative response vectors; clustering the negative response vectors to obtain negative clusters; extracting negative response vectors from the top multiple condition combinations in the interaction response intensity distribution; matching the extracted negative response vectors with the negative clusters; if the extracted negative response vectors successfully match the negative clusters, obtaining a matching condition combination; removing the matching condition combination from the top multiple condition combinations in the interaction response intensity distribution to obtain a first condition combination; if the negative response vectors do not successfully match the negative clusters, using the top multiple condition combinations in the interaction response intensity distribution as the first condition combination; and continuing to remove specified condition combinations from the first condition combination to obtain a target safe condition combination; wherein the specified condition combination is a predefined condition combination that causes the user to have a negative response.

[0016] In some embodiments, the step of constructing the security constraints includes: determining a first security constraint based on the distance relationship between the user's negative response vector and the negative cluster, wherein the first security constraint is related to grouped negative information; performing emergent processing on the response vectors corresponding to the specified condition combination to expand the discrete or continuous point-like response vectors to obtain a continuous spatial response vector forming a semantic cluster around the point-like vectors; and determining a second security constraint based on the distance relationship between the user's negative response vector and the continuous spatial response vector, wherein the second security constraint is related to personalized negative information.

[0017] This application also provides an interactive model training device based on extended reality and sensor analysis, the device comprising:

[0018] The parameter determination module is used to determine the virtual world generation parameters and user characteristic parameters;

[0019] The world generation module is used to construct a dynamic virtual world based on the virtual world generation parameters.

[0020] The virtual reality module is used to acquire EEG data of a user during interactive training in the dynamic virtual world based on the user's characteristic parameters through an EEG acquisition unit, and to acquire at least one of physiological sensing data and behavioral sensing data through a physiological sensing data acquisition unit and a behavioral sensing data acquisition unit.

[0021] The calculation module is used to determine the user's interactive training effect based on the acquired EEG data, combined with at least one of physiological sensor data and behavioral sensor data.

[0022] The adjustment module is used to adjust the virtual world generation parameters according to the preset prediction model when the interaction training effect does not meet the convergence condition during the iterative process, and continue to construct a new dynamic virtual world according to the adjusted virtual world generation parameters until the interaction training effect meets the convergence condition, at which point the iterative process ends and the interaction model is trained.

[0023] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the interactive model training method based on extended reality and sensor analysis provided in any of the above embodiments.

[0024] In the above embodiments, this application constructs a dynamic virtual world corresponding to a three-dimensional dynamic world by acquiring virtual world generation parameters and user feature parameters. Then, through simultaneous multi-channel acquisition of neural, physiological, and behavioral data, the user's electroencephalogram (EEG) data and related behavioral and physiological data are simultaneously acquired during interactive training to achieve time-locked analysis of neural activity and interactive responses, establishing a causal relationship between stimulus, neural response, and interactive response. Furthermore, through adaptive adjustment of interactive training parameters, based on the neural and interactive related indicators output by the adjustment response monitoring module, the virtual world generation parameters for the next round of training are automatically adjusted, achieving individualized and refined adjustment of the intensity and dosage of interactive training, thereby maximizing the induction effect of the interactive training response. Attached Figure Description

[0025] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:

[0026] Figure 1 This is a flowchart illustrating an interactive model training method based on extended reality and sensor analysis, provided in one embodiment of this application.

[0027] Figure 2 This is a schematic diagram of an interactive model training device based on extended reality and sensor analysis provided in one embodiment of this application;

[0028] Figure 3 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0030] The technical solutions of the various embodiments of this application can be combined with each other, but only if they are based on the ability of a person skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this application.

[0031] This solution is not intended to obtain disease diagnosis results or health status. It is merely a method for processing virtual world generation parameters, generating virtual worlds, and acquiring users' EEG data, physiological sensor data, and / or behavioral sensor data in virtual worlds. All steps are performed by computer or other devices for information processing.

[0032] It should be fully understood that the user information involved in this application (including but not limited to user physiological information, user personal information, etc.) is information and data authorized by the user or fully authorized by all parties. The use of user information shall comply with the privacy policies and practices of the industry that are generally considered to meet or exceed the requirements for maintaining user privacy. The collection, use and processing of related data shall comply with relevant laws, regulations and standards, and provide corresponding operation access points for users to choose to authorize or refuse.

[0033] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0034] First, it should be noted that when facing external threats, an organism typically relies on an interactive response system composed of a cognitive organism and a state-response organism. The cognitive organism focuses on recognizing, assessing, and regulating behavior in response to external threats. For example, individuals reduce their exposure to potentially negative environments through cognitive processes such as risk perception, stress adaptation, and avoidance behavior, and can form stable interactive response strategies through learning and conditioned reflexes. The state-response mechanism mainly relies on real-time state feedback to adjust the intensity of the response. This application aims to provide a method and apparatus for training interactive models based on extended reality and sensor analysis. The method provided by this application enables users to achieve dual interactive training in cognitive regulation and state response, improving their adaptability to external challenges. In the following embodiments, the application of the method provided by this application in the field of state-response interaction (e.g., virtual sports skill training scenarios, virtual high-risk operation training scenarios, etc.) is used as an example for illustration. However, in other embodiments, the method provided by this application can also be used in the field of cognitive regulation interaction, which will not be elaborated here.

[0035] In some embodiments, see Figure 1 This application provides a method for training an interactive model based on extended reality and sensor analysis, the method comprising:

[0036] Step S101: Determine the virtual world generation parameters and user characteristic parameters.

[0037] Step S102: Construct a dynamic virtual world based on the virtual world generation parameters.

[0038] Step S103: Obtain at least one of the following: electroencephalogram (EEG) data of the user during interactive training in the dynamic virtual world based on the user characteristic parameters, as well as physiological sensor data and behavioral sensor data.

[0039] Step S104: Determine the user's interactive training effect based on the EEG data, combined with at least one of physiological sensor data and behavioral sensor data.

[0040] Step S105: During the iterative process, when the interaction training effect does not meet the convergence condition, the virtual world generation parameters are adjusted according to the preset prediction model, and a new dynamic virtual world is constructed according to the adjusted virtual world generation parameters until the interaction training effect meets the convergence condition, at which point the iterative process ends and the interaction model is trained.

[0041] This application also provides an interactive model training device based on extended reality and sensor analysis, the device comprising:

[0042] The parameter determination module is used to determine the virtual world generation parameters and user characteristic parameters;

[0043] The world generation module is used to construct a dynamic virtual world based on the virtual world generation parameters.

[0044] The virtual reality module is used to acquire EEG data of a user during interactive training in the dynamic virtual world based on the user's characteristic parameters through an EEG acquisition unit, and to acquire at least one of physiological sensing data and behavioral sensing data through a physiological sensing data acquisition unit and a behavioral sensing data acquisition unit.

[0045] The calculation module is used to determine the user's interactive training effect based on the acquired EEG data, combined with at least one of physiological sensor data and behavioral sensor data.

[0046] The adjustment module is used to adjust the virtual world generation parameters according to the preset prediction model when the interaction training effect does not meet the convergence condition during the iterative process, and continue to construct a new dynamic virtual world according to the adjusted virtual world generation parameters until the interaction training effect meets the convergence condition, at which point the iterative process ends and the interaction model is trained.

[0047] In some embodiments, the virtual world generation parameters include at least one of contextual characteristic parameters and negative characteristic parameters. Contextual characteristic parameters may be crowd characteristic parameters, and negative characteristic parameters may be negatively related parameters, including one or more of negative parameters, negative cognitive parameters, or negative behavioral parameters. User characteristic parameters include at least one of walking speed and exposure time. Physiological sensor data includes one or more of heart rate, skin conductance, respiratory rate, blood pressure, and skin temperature, while behavioral sensor data includes the user's reaction time and accuracy to a specified task.

[0048] Specifically, this application includes a parameter determination module, which is used to determine virtual world generation parameters and user characteristic parameters. The virtual world generation parameters include context characteristic parameters and negative characteristic parameters.

[0049] This application includes a context feature module, which is used to set the context feature module in the subsequent world generation module. In some embodiments, the context feature module may specifically be a crowd feature module, and the following embodiments will be described using the crowd feature module as an example.

[0050] See Figure 2 The crowd feature module is used to set crowd feature parameters for subsequent world generation modules. In some embodiments, this module may include one or more of six core parameters, which may include: crowd density, crowd distance, crowd mood, crowd flow, crowd interaction, and affected proportion. Crowd density sets the density of the crowd in the generated dynamic world; a higher crowd density means more pedestrians per unit area. Crowd distance sets the distance between crowds in the generated dynamic world; a smaller crowd distance means pedestrians are closer to the user. Crowd mood sets the mood of the crowd in the generated dynamic world; a more positive mood indicates a more active and optimistic outlook. Crowd flow sets the flow of crowds in the generated dynamic world; stronger crowd flow means more dynamic and unpredictable crowd movements. Crowd interaction sets the interactivity of the crowd in the generated dynamic world; stronger interactivity means a higher probability of interaction between the crowd and the user. The affected proportion sets the proportion of the crowd in the generated dynamic world that receives preset negative information; a higher affected proportion means more people possess the negative information set in the negative feature module. It should be noted that in different business scenarios, the audience characteristic parameters and contextual characteristic parameters can be flexibly selected and set according to the specific business scenario, and are not limited to the above six characteristic parameters.

[0051] This application also includes a negative feature module, which is used to set negative feature parameters in the subsequent world generation module. In some embodiments, this module may include one or more of two core parameters, which may include negative typical features and negative feature severity. Negative typical features are responsible for setting the negative typical features of the people in the generated dynamic world. For example, the goal is to trigger user interaction with negative information; in some embodiments, negative typical features may include negative typical features, such as sneezing and runny nose. Negative feature severity is responsible for setting the severity of the negative typical features; a higher value indicates a more severe negative effect. It should be noted that in different business scenarios, negative feature parameters can be flexibly selected and set according to the specific business scenario, and are not limited to the two feature parameters mentioned above.

[0052] This application also includes a user parameter module, which is used to set user characteristic parameters for users to complete experimental paradigms in virtual reality. In some embodiments, this module includes two core parameters: movement speed and exposure time. Movement speed sets the user's movement speed when completing the experimental paradigm in virtual reality; a faster movement speed allows the user to control their movement more quickly. Exposure time sets the user's exposure time when completing the experimental paradigm in virtual reality; a longer exposure time requires the user to be exposed in virtual reality for a longer period. It should be noted that user parameters can be flexibly selected and set according to specific business scenarios, and are not limited to the two characteristic parameters mentioned above.

[0053] This application also includes a world generation module for generating a continuous, dynamic world. In some embodiments, this module includes two cue word units and a world model (e.g., a pre-defined WorldGen unit). The cue word units generate cue words based on parameters passed from the crowd feature module and the negative feature module, as well as a pre-tuned large language model. The WorldGen unit generates an explorable, interactive, and highly consistent 3D world based on the cue words. It should be noted that the world generation module can be flexibly selected and set according to specific business scenarios, and is not limited to the two feature parameters mentioned above.

[0054] In some embodiments, this application further includes a virtual reality module, which is used to present a dynamic world to a user wearing a virtual reality device and to complete the interactive training experimental paradigm and data collection within this module. In some embodiments, this module may include three units: an experimental paradigm unit, an EEG data acquisition unit, and a physiological and behavioral sensor data acquisition unit. The experimental paradigm unit allows the user to complete the interactive training program with a predetermined number of trials, duration, and intervals. The EEG data acquisition unit collects the user's EEG data during the interactive training process. The physiological and behavioral sensor data acquisition unit collects the user's physiological and behavioral sensor data before, after, or during the interactive training.

[0055] In some embodiments, this application further includes a training effect monitoring module. In some embodiments, the training effect monitoring module is used to monitor the user's interactive responses during and after interactive training from multiple aspects such as neurometrics, physiological and behavioral indicators.

[0056] In some embodiments, this application further includes an adjustment module. The adjustment module is used to automatically adjust the core parameters of the next round of interactive training based on the output of the training effect monitoring module.

[0057] In the above embodiments, this application uses parameterized crowd and negative feature modeling. Based on the crowd feature module (context feature module) and negative feature module, factors such as crowd density, crowd distance, typical negative features, and negative severity are quantified into adjustable parameters and used as input variables for virtual world generation, achieving precise control over the ecological validity and stimulus intensity of the exposure scenario. Then, based on the large language model and the preset WorldGen dynamic world generation, the pre-trained large language model generates prompts containing scenario parameters, and the preset WorldGen model generates a highly consistent and explorable three-dimensional dynamic world, thereby enhancing the realism and immersion of the exposure scenario. Next, through simultaneous acquisition via neural and interactive dual channels, the user's EEG data and behavioral and physiological data are simultaneously collected during interactive training to achieve time-locked analysis of neural activity and interactive responses, establishing a causal relationship between stimulus, neural response, and interactive response. Through adaptive adjustment of interactive training parameters, based on the neural and interactive related indicators output by the training effect monitoring module, the crowd and negative feature parameters for the next round of training are automatically adjusted to achieve individualized and refined adjustment of the interactive training intensity and "dosage," thereby maximizing the induction effect of the interactive training response. Furthermore, this application ultimately generates an interaction model, which users can use for interactive training. The parameters of the training model can also be updated and optimized in real time based on the user's state to adapt to the user's constantly changing state and improve the flexibility of the interaction model.

[0058] In some embodiments, this application can construct an interactive training paradigm based on the interactive training sampling time and condition combinations; obtain the interactive response intensity distribution of each group of subjects when receiving interactive training with corresponding condition combinations at the corresponding interactive training sampling time based on the interactive training paradigm; wherein, the condition combinations include situational feature parameters and negative feature parameters in the virtual world generation parameters; obtain the baseline value of the virtual world generation parameters based on the interactive response intensity distribution; define safety constraints, and under the safety constraints, iteratively find the parameter corresponding to the strongest interactive training effect based on the baseline value of the virtual world generation parameters, and use it as the final virtual world generation parameters.

[0059] The method provided in this application specifically includes the following steps:

[0060] Step 1: Construct an interactive training paradigm. Specifically, construct an interactive training paradigm based on the interactive training sampling time and condition combinations; obtain the interactive response intensity distribution of each group of subjects when receiving interactive training with corresponding condition combinations at the corresponding interactive training sampling time based on the interactive training paradigm; wherein, the condition combinations include situational feature parameters and negative feature parameters in the virtual world generation parameters.

[0061] In practice, several participants are recruited and divided into N groups (N is a positive integer greater than or equal to 1). Each group consists of M participants (M is a positive integer greater than or equal to 1, e.g., M = 30). The sampling time T and different condition combinations C can be used as variables among the participants. The condition combinations are derived from multiple population characteristics and negative severity levels, conforming to the constraints of a negative model with bilinear incidence rates. See the interactive training paradigm below. Each independent number represents a group of participants, and each group receives an interactive training paradigm constructed from specific condition combinations at a specific sampling time. This yields the EEG characteristics and the distribution of interactive response intensity for each group when receiving the interactive training paradigm with that condition combination at that time, and the participants' feedback information is recorded. The interaction response intensity is obtained based on EEG signals and physiological and behavioral sensor data. Feedback information is obtained jointly from behavioral and physiological sensor data, EEG signals, and user subjective evaluation. Thus, based on traditional indicator data, EEG signals are used to assist in the fine quantification of interaction response intensity, and physiological and behavioral response indicators and EEG signals are used to more accurately quantify negative reactions in user subjective evaluation. The interaction training paradigm is shown below.

[0062]

[0063] In some embodiments, when the sample size is sufficient, condition combinations can be used as variables. When the sample size is insufficient, sampling time can be used as an within-subjects variable, as shown below. If the sample size is too small, for example, only within-subjects variables at sampling time T1 are available, then sampling time can be used as an within-subjects variable.

[0064]

[0065] It is understandable that the above embodiments are merely alternatives for situations where the sample size is too small. In reality, to ensure the effectiveness of the final training results, the number of subjects in each group is usually quite large, i.e., N is relatively large. However, a potential problem when sampling subjects in this case is that each sampling requires data collection for a specific sampling time T and condition combination C to obtain the distribution of EEG features and interaction response intensity. Some subjects in certain groups may not be able to complete the effective sampling process at all sampling times due to various reasons. On the one hand, in some embodiments, only comprehensive and effective sampled data can be selected as the basis for subsequent interactive training. On the other hand, from the perspective of fully utilizing data resources, missing value imputation can be performed on the currently incomplete data. However, the imputation targets are limited to data with fewer than a preset number of missing values ​​after the number of samplings. That is, missing value imputation also has limitations; not all missing values ​​need to be imputed. Imputation methods can be implemented using generative adversarial networks, etc., which will not be elaborated here.

[0066] Step 2: Initialize the crowd feature module. Obtain the baseline values ​​of the context feature parameters in the virtual world generation parameters based on the interaction response intensity distribution.

[0067] Specifically, based on negative characteristics with overt negative symptoms and infectivity, appropriate environmental parameters are determined using correlation analysis methods, such as the analytic hierarchy process (AHP). Typical environmental parameters may include crowd density, crowd distance, crowd mood, crowd flow, crowd interaction, and the proportion of people affected. It should be noted that the physical environment itself is not considered as an environmental parameter to be adjusted; rather, it is treated as a fixed value, such as the physical environment of a commercial street. It should also be noted that in some embodiments, when the scenario feature parameters include crowd feature parameters, this application is used to explore the influence of crowd-related environmental parameters on the intensity of human interaction and EEG characteristics. This is because the complex nonlinear relationship reflected by changes in human interaction intensity and EEG characteristics caused by a single physical environment is difficult to capture effectively, affecting the accuracy of interactive training based on crowd features and causing disturbances to the fine-tuning of crowd features, which is detrimental to subsequent interactive training. Therefore, this application sets crowd density, crowd distance, crowd mood, crowd flow, crowd interaction, and the proportion of people affected as baseline values ​​for initialization. Specifically, the values ​​of the above-mentioned crowd features are explained as follows:

[0068] Population density, ranging from 0 to 10, represents the number of people per unit area.

[0069] Crowd distance, with a value ranging from 0 to 5, represents the initial unit distance from the user when the crowd is refreshed.

[0070] The crowd's mood, ranging from -1 to 1, represents the emotional state of the crowd. Negative values ​​represent negative emotions (panic, anger), 0 represents neutral emotions, and positive values ​​represent positive emotions (smile, friendliness). The magnitude of the absolute value represents the degree of emotional expression.

[0071] Crowd movement, with a value ranging from 0 to 1, represents the overall dynamics and predictability of the crowd. 0 represents that the crowd is in a static state, while 1 represents that the crowd is highly mobile and exhibits randomness in its movement.

[0072] Crowd interaction, with a value ranging from 0 to 1, represents the probability of a crowd interacting with a user (eye contact, nodding, conversation, etc.). 0 means that the crowd does not interact with the user at all, while 1 means that every individual in the crowd will interact with the user.

[0073] The affected proportion, ranging from 0 to 1, represents the proportion of the population with typical negative characteristics. 0 means that no one in the population exhibits negative characteristics, while 1 means that everyone in the population has negative characteristics.

[0074] The setting of the baseline value of the crowd characteristics depends on the dataset obtained in step 1 above. Based on the dataset provided above, the combination of conditions with the strongest interaction response intensity is selected from the dataset, and the set value of the crowd characteristics corresponding to this combination of conditions is used as the baseline value of the current user's crowd characteristics, as shown in the following formula, where the following parameters represent the crowd characteristic parameters P. crowd The composition includes population density D density Crowd distance D distance Emotions of the crowd emotion Crowd Flow M mobility Crowd Interaction interaction The proportion of the population affected (R) infection And provide the initial baseline value.

[0075]

[0076] However, it should be understood and easy to see that the combination of conditions with the strongest interaction response is actually only a suboptimal solution that approximates the optimal combination of conditions. In this regard, this application also proposes an optimization algorithm to optimize the combination of conditions with the strongest interaction response (i.e., the suboptimal solution) selected at the time based on the existing dataset, so as to more accurately approximate the optimal solution under the ideal condition. The specific steps of the optimization algorithm are shown in step 3.

[0077] Step 3: Initialize the negative feature module. Obtain the baseline values ​​of the negative feature parameters in the virtual world generation parameters based on the interaction response intensity distribution.

[0078] The negative typical symptoms and negative severity parameters are set as baseline values. The baseline values ​​are set based on the database obtained in step 1, selecting the combination of conditions with the strongest interaction intensity. The baseline values ​​are set as shown in the following formula, where the following parameters represent the negative characteristic parameter P. disease The components include negative typical symptoms S symptom And the negative severity parameter L level And provide the initial baseline value:

[0079]

[0080] The set values ​​of the negative characteristics are used as the baseline values ​​of the current user's negative characteristics. The values ​​of the negative characteristics are explained as follows:

[0081] Negative typical symptoms, described in text (e.g., sneezing, red nose, etc. when corresponding to negative symptoms). In some embodiments, this parameter does not participate in regulation, but its manifestation is controlled by the severity of the negative; that is, negative typical symptoms are an explicit representation of the severity of the negative.

[0082] The severity of the negative impact ranges from 0 to 1, indicating the degree of severity of the typical negative symptoms. 0 represents no typical negative symptoms, while 1 represents the full manifestation of typical negative symptoms.

[0083] In order to optimize the combination of conditions that has the strongest interaction response intensity (i.e., the suboptimal solution), this application regards crowd density, crowd distance, crowd emotion, crowd flow, crowd interaction, affected proportion and negative severity as factors affecting the intensity of user interaction response. However, these factors are only typical examples and not absolute. These factors can also be added or removed according to actual needs.

[0084] Furthermore, it is important to emphasize that although this application introduces EEG signals to optimize the interactive training effect, this does not mean that the most important aspect of this solution is to obtain the optimal solution in terms of performance. The optimization process of this application is carried out under strict safety constraints. When the optimal solution is applied to the user, the strong negative cognitive reaction may cause many adverse reactions. This application focuses on gradually conducting multiple interactive training sessions for the user while avoiding excessive negative reactions from the user. By introducing EEG signals, it reduces the accuracy sacrifice caused by safety constraints to a certain extent. Conventional optimization processes often focus on pursuing the optimization of performance indicators. For example, in the field of small object detection, the optimization process aims to comprehensively detect all small objects in the captured image without strict safety constraints. This is one of the important differences between this application and conventional practices.

[0085] Step 4: Initialize user parameters.

[0086] Specifically, the travel speed and exposure time parameters are set as baseline values. The baseline values ​​also rely on the database obtained in step 1. The combination of conditions with the strongest interaction intensity is selected, and the user characteristic settings within this combination are used as the baseline values ​​for the current user's user characteristics. The interpretation of the user characteristic values ​​is as follows:

[0087] Movement speed, ranging from 0 to 5, represents the unit distance the user moves forward each time they input a forward command.

[0088] Exposure time, ranging from 5 to 15, represents the unit of time a user spends in each round of interactive training.

[0089] The baseline value is set as shown in the following formula, where the following parameters represent the user feature parameters P. user The components include the speed of travel V speed and exposure time T exposure And provide the initial baseline value:

[0090]

[0091] Specifically, walking speed and exposure time can also be considered as factors affecting the intensity of user interaction responses, and the condition combination can be fine-tuned based on the conditions shown in step 3. However, it should be noted that whether or not to include walking speed and exposure time in the condition combination depends on the dataset obtained from the initial sampling of subjects. If the same walking speed and exposure time are uniformly specified when building the dataset, then including walking speed and exposure time in the condition combination is inappropriate. If walking speed and exposure time are included in the dataset building process, then walking speed and exposure time can also be used as variables in the condition combination, but this will complicate the optimization problem, and the effect on the final interaction training effect is questionable. Considering that including walking speed and exposure time as factors in building the dataset will lead to a relatively high cost in building the dataset, some embodiments of this application set walking speed and exposure time as adjustable fixed values. For example, walking speed can be the average walking speed of normal humans, and exposure time can be set according to the professional opinions of doctors / physiologists. The issue of walking speed and exposure time does not need to be considered when solving for the optimal solution in step 3.

[0092] Step 5: Generate a dynamic virtual world. After completing all initializations, a pre-trained large language model (LLM) for generating prompt words is used. Eight parameters from the crowd feature module and negative feature module are input to generate a prompt word containing eight parameters. As shown in the formula below, parameter p is input into the large language model (LLM) to generate the prompt word "Prompt".

[0093]

[0094] Example prompts: Crowd density 3, Crowd distance 2, Crowd mood -0.5, Crowd flow 0.2, Crowd interaction 0.3, Affected proportion 0.4, Typical negative symptoms include sneezing and red nose, Negative severity 0.7. This can be expressed in Chinese as: A dynamic virtual world is generated: On a dilapidated shopping mall walkway, a group of pedestrians is walking towards the target origin. This group of pedestrians occupies one unit area, with 3 people in each unit area. The initial distance of this group of pedestrians from the target origin is 2 meters. The overall mood of this group of pedestrians is moderately negative (emotional intensity is measured in 0-1, approximately 0.5). This group of pedestrians moves slowly and with a certain regularity (movement randomness is measured in 0-1, approximately 0.2). Approximately 30% of this group of pedestrians will occasionally interact with the user (randomly manifested as eye contact, nodding, or reaching out). Approximately 40% of this group of pedestrians exhibit negative symptoms. Typical negative symptoms include sneezing and red nose. Those exhibiting negative symptoms within this group of pedestrians display fairly similar negative typical symptoms (similarity to typical symptoms is measured on a scale of 0 to 1, approximately 0.7). Realistically render these characteristics in a 3D world, ensuring scene consistency.

[0095] Step 6: Generate a 3D world representation based on design parameters.

[0096] Specifically, the prompt words generated in step 3 are input into the world model (e.g., the preset WorldGen) to generate a dynamic world based on design parameters. The dynamic world and its objects are then output as a three-dimensional world representation, and the changes of the objects over time are recorded.

[0097] Step 7: Construct the virtual reality interface. Specifically, the dynamic world representation generated in Step 4 is constructed into a virtual reality interface and bound to the virtual reality device's turning and forward movement parameters.

[0098] As shown in the formula below, WorldGen(Prompt) represents the generation of a continuous, interactive, dynamic world W based on a given Prompt using a world generation model; Render3D(W) represents a 3D rendering or decoding function that visualizes W. These represent static scene geometry (such as terrain, buildings, and skyboxes) in a 3D environment, and the dynamic changes of objects over time. VRBuilder represents a builder module responsible for integrating the rendered 3D environment and objects into the VR environment; V UI This refers to the visual interactive screen that users ultimately see when they put on VR glasses; This indicates the device's turning and forward movement, as well as the turning parameters; "build" refers to the two-way binding of the user's operation commands (related parameters for turning the head and moving forward) with the VR interface to achieve interaction.

[0099]

[0100]

[0101] Step 8: Connect the virtual reality interface to the virtual reality device worn by the user. Place an EEG device on the user that can collect multi-channel EEG signals. Connect a physiological and behavioral data acquisition device to the user; this device is used to collect the user's physiological and behavioral sensory data before and after interactive training. Complete the assembly and functional testing of all devices.

[0102] Step 9: Implement the interactive training paradigm. Users need to walk straight in a constructed dynamic world, from the starting point to the destination. Users can observe left and right, and the walking speed is determined by user parameters. The duration of each round of interactive training is determined by user parameters. It includes several trials, each trial representing a wave of people conforming to predetermined parameters walking towards the user. The interval between trials is set to pseudo-random based on the average of 5 seconds. During the interactive training, the user's EEG data is collected. Physiological and behavioral sensory data of the user are collected before and after the start of interactive training.

[0103] Step 10: Determine the user's interactive training effect based on the EEG data, combined with at least one of physiological sensor data and behavioral sensor data.

[0104]

[0105]

[0106] In the above formula, EEG ERP (t) represents the average event-related EEG signal, t represents the event window duration, and N represents the total number of events. i It is the i-th event out of N events; F EEG It is for EEG ERP (t) Extracted features, FeatureExtract is for EEG ERP (t) The method for extracting features is detailed in the examples below and will not be repeated here. C △ The difference in physiological or behavioral indicators before and after training is used to characterize the training effect. (C) T0 C represents the first collection of physiological or behavioral sensor data. T1 This represents physiological or behavioral sensor data collected a second time.

[0107] In some embodiments, one or more of the following can be collected: electroencephalogram (EEG) data, cognitive data, and behavioral sensor data, to determine the user's interactive training effect.

[0108] In some embodiments, determining the user's interactive training effect based on the EEG data, combined with at least one of physiological sensor data and behavioral sensor data, includes: dividing continuous EEG data into multiple trials based on trigger tags set in the experimental paradigm when an event occurs; wherein each trial covers EEG data before and for a period of time after the event begins; and averaging the EEG data from multiple trials to obtain event-related potentials; wherein the positive and negative characteristics of the event-related potentials within different time windows characterize the neural activity characteristics of the electrode regions.

[0109] Specifically, based on EEG data during user interaction training, EEG responses dependent on events (such as people approaching or someone sneezing in the crowd) are calculated, and then the responses to all events are averaged to obtain the EEG characteristics corresponding to the interaction response.

[0110] Specifically, the acquired EEG data underwent preprocessing. First, baseline correction was performed using signals from -2 to 0 seconds prior to the event (e.g., a crowd approaching, someone in the crowd exhibiting negative characteristics), and an average rereference method was employed to reduce data drift. Subsequently, to remove potential power frequency interference during acquisition, a 50 Hz notch filter was used, along with a 1–45 Hz bandpass filter to extract effective EEG features related to the interaction response. Next, based on the trigger labels set in the experimental paradigm (e.g., a crowd approaching can be considered an event), the continuous EEG data was divided into several trials, each covering 0.2 seconds before the event and 5 seconds after. Specifically, features were extracted using the FeatureExtract algorithm, and the EEG data from these trials were superimposed and averaged to obtain event-related potentials (ERPs), which include positive and negative features and peak values. The positive and negative features of EEPs within different time windows reflect relatively stable neural activity in specific electrode regions and are important indicators of cognitive processing.

[0111] In the above embodiments, the focus is on a series of event-related potential components related to the affected population: (1) 80–120 ms, positive / negative, parietal / occipital lobe related electrodes, reflecting early sensory processing and attention capture; (2) 180–250 ms, negative, frontal lobe related electrodes, reflecting threat detection; (3) 250–500 ms, positive, frontal lobe related electrodes, reflecting attention allocation to salient stimuli; (4) 400–800 ms, positive, parietal lobe related electrodes, reflecting emotion-related processing. For the above characteristics, features are extracted within their corresponding time windows and electrode regions: (1) average amplitude, the average potential within the time window, reflecting the overall corresponding cognitive process processing level; (2) peak amplitude, the maximum or minimum potential within the time window, reflecting the overall corresponding cognitive process processing degree; (3) latency, reflecting the overall corresponding cognitive process processing speed.

[0112] In the above embodiments, the EEG feature extraction method, which uses event-related potentials (ERPs) as the primary feature extraction method, can reflect more dimensions of information related to the predicted interactive responses and has a high temporal resolution. It can dynamically adjust based on the individual's dynamic changes in perception, attention, emotion, and motivation during interactive training. After the database is established and a large amount of data is collected, EEG features can be extracted to predict the user's interactive responses, thereby enabling real-time monitoring of potential user interactive responses during interactive training without collecting other physiological and behavioral data.

[0113] Specifically, the method also includes analyzing physiological and behavioral sensor data samples collected before and after user interaction training. Based on corresponding detection algorithms, the changes and differences in users' physiological indicators before, after, or during interaction training are determined. Based on corresponding response theories and their cognitive significance, the changes and differences in users' cognitive processes before, after, or during interaction training are determined. Specifically, in some embodiments, physiological sensor data corresponding to users before, after, or during interaction training is collected, and the changes and differences in users' physiological indicators before, after, or during interaction training are determined based on corresponding detection algorithms. Corresponding behavioral sensor data or scale data of users before, after, or during interaction training are collected, and the changes and differences in users' cognitive processes before, after, or during interaction training are determined based on corresponding response theories and their cognitive significance.

[0114] In the above embodiments, the corresponding data type can be adaptively selected according to different scenarios and needs to evaluate the user interaction training effect, improving the accuracy and flexibility of the evaluation of the user interaction training effect. This also allows the method provided in this application to adapt to more interactive training business scenarios. For example, in some embodiments, only event-related potentials (ERPs) can be acquired to determine the user's interactive training effect; in other embodiments, only the physiological sensor data can be acquired, and the user's interactive training effect can be determined based on changes and differences in physiological indicators; and in still other embodiments, only the user's behavioral sensor data can be acquired, and the user's interactive training effect can be determined based on changes and differences in the cognitive processes corresponding to the behavioral sensor data. In another embodiment, the method can also determine the user's interactive training effect based on any two or more combinations of EEG data, physiological sensor data, or behavioral sensor data. It should be noted that EEG data is generally a necessary data type in different scenarios, while physiological sensor data and behavioral sensor data can be adaptively selected according to different scenarios.

[0115] Step 11: Update, adjust, and build a new dynamic virtual world.

[0116] In some embodiments, adjusting the virtual world generation parameters according to a preset prediction model and continuing to construct a new dynamic virtual world according to the adjusted virtual world generation parameters includes: inputting the virtual world generation parameters into a preset prediction model, predicting the interaction training effect according to the preset prediction model; when the difference between the predicted interaction training effect and the user's current actual interaction training effect satisfies a preset convergence condition, predicting new virtual world generation parameters according to the preset prediction model, and continuing to construct a new dynamic virtual world according to the new virtual world generation parameters.

[0117] Specifically, the interaction response indicators based on EEG, physiology, and behavior obtained in step 8 are compared with the response indicators in the database constructed in step 1. The difference compared to the standard interaction response is calculated. This difference can be measured by weighting the changes in EEG characteristics and the changes in cell proportions. When this difference falls outside -2 standard deviations of the corresponding distribution, the regulation module is started.

[0118] Specifically, based on the database collected in step 1, a mapping relationship is established between different combinations of conditions (parameter vectors) and interaction response features (feature vectors). This mapping relationship can be a generalized linear model, which includes an input user feature vector (for example, the nine parameters in the following formula, such as the population density D). density Crowd distance D distance Emotions of the crowd emotion Crowd Flow M mobilityCrowd Interaction interaction The proportion of the population affected (R) infection Negative level L level speed of travel V speed Exposure duration T exposure This yields a predicted interactive training effect value (0,1), where + and - represent increases and decreases, respectively.

[0119]

[0120] The interaction response features measured in step 8 are input into the pre-established mapping model to obtain the predicted conditional combination, which serves as the basis for adjustment in the next training iteration. For example, if, compared to the norm, the user's first-round interaction response falls outside -2 standard deviations of the corresponding conditional interaction response distribution in the database, it indicates that the user's interaction response after the first round of interaction training is insufficient to reach the expected standard level. In this case, the adjustment module will calculate the predicted conditional combination based on the current interaction response (i.e., the optimal solution calculated in step 3). Compared to the user's actual conditional combination, the core parameters in the crowd feature module, negative feature module, and user parameter module may have one or more of the following differences: need to increase crowd density (+), need to decrease crowd distance (-), need to negativeize crowd emotions (-), need to enhance crowd flow (+), need to increase the probability of crowd interaction (+), need to increase the proportion of the crowd affected (+), need to increase the severity of negativity (+), need to slow down the walking speed (-), need to increase the exposure time (+). After obtaining the adjustment directions represented by + and -, the conditional combination corresponding to the user's next round of interaction training needs to be adjusted based on the initial optimal solution. At this point, the objective function remains unchanged, but considering the differences among individual users, a dynamic safety domain corresponding to this round can be generated based on the interaction response intensity and response vector of the user's output in each round. This continuously corrects the understanding of the safety boundary and avoids the problems of poor training effect or poor progress that may be caused by the static safety domain.

[0121] Specifically, when the user's reaction vector does not indicate any risk (i.e., the user's reaction vector meets all safety constraints and the user's subjective feeling is normal), all applicable condition combinations during the user interaction training process are added to the current user's condition combination safety seed set. This reuses the above optimization method, adaptively updates the personalized condition combination safety seed set according to individual differences, and refits the stochastic process model according to the stochastic process model update method shown in the above steps, outputting the condition combination vector for the next round of interaction training. This allows users to conduct refined and efficient interaction training while ensuring safety.

[0122] In some embodiments, the step of determining the virtual world generation parameters includes: obtaining the interaction response intensity distribution characterizing the user's interaction training effect; obtaining the first condition combination with the strongest interaction training quantification index based on the interaction response intensity distribution; removing negative condition combinations from the first condition combination to obtain a target safety condition combination; using the target safety condition combination as the baseline value of the virtual world generation parameters; initializing a stochastic process model using a clustering algorithm, based on the target safety condition combination as the objective function and each safety constraint condition; predicting the virtual world generation parameters based on the stochastic process model; wherein the stochastic process model characterizes a subset of condition combinations selected from the target safety condition combination and is used as a data subset to fit a model describing the safety condition combinations.

[0123] In some embodiments, removing negative condition combinations from the first condition combination to obtain a target safe condition combination includes: structuring the feedback information of the subjects in the interaction training paradigm according to a pre-trained language model to obtain negative response vectors; clustering the negative response vectors to obtain negative clusters; extracting negative response vectors from the top multiple condition combinations in the interaction response intensity distribution; matching the extracted negative response vectors with the negative clusters; if the extracted negative response vectors successfully match the negative clusters, obtaining a matching condition combination; removing the matching condition combination from the top multiple condition combinations in the interaction response intensity distribution to obtain a first condition combination; if the negative response vectors do not successfully match the negative clusters, using the top multiple condition combinations in the interaction response intensity distribution as the first condition combination; and continuing to remove specified condition combinations from the first condition combination to obtain a target safe condition combination; wherein the specified condition combination is a predefined condition combination that causes the user to have a negative response.

[0124] Specifically, using a pre-trained deep learning model (such as BERT or its variants ERNIE, RoBERTa), the feedback information from participants in the current dataset is structured to obtain negative response vectors. These negative response vectors are then clustered using k-means to obtain clusters. The system then determines whether the negative response vectors corresponding to the top-ranking conditional combinations match any of the clusters. Matching methods can include calculating the Euclidean distance between vectors, or calculating the distance between each point in the cluster and the negative response vectors corresponding to multiple conditional combinations to determine a match. Alternatively, several points can be randomly selected from the clusters (ensuring global coverage and avoiding getting stuck in local selection) to calculate the distances between these points and the negative response vectors corresponding to multiple conditional combinations. If a match exists, it is recorded as a matching conditional combination. This matching conditional combination is then removed from the top-ranking conditional combinations (the first conditional combination) to obtain conditional combination A (the first conditional combination). If no match exists, the top-ranking conditional combinations are directly used as conditional combination A. Eliminate the specified condition combination from condition combination A. This specified condition combination is one that occurs in a small number of cases but must be strictly avoided to prevent users from having corresponding negative reactions. It can be given by a doctor or physiologist. This will result in condition combination B (target safety condition combination).

[0125] At this point, the search space corresponding to condition combination B (target safe condition combination) can be regarded as a safe domain, while the search space corresponding to clusters and specified condition combinations is regarded as an unsafe domain. Furthermore, the safe domain and the unsafe domain are not adjacent. Next, we will actively search for the optimal solution that is both safe and high-performance based on the safe domain, rather than simply searching for the point with the best performance.

[0126] In some embodiments, the step of constructing the security constraints includes: determining a first security constraint based on the distance relationship between the user's negative response vector and the negative cluster, wherein the first security constraint is related to grouped negative information; performing emergent processing on the response vectors corresponding to the specified condition combination to expand the discrete or continuous point-like response vectors to obtain a continuous spatial response vector forming a semantic cluster around the point-like vectors; and determining a second security constraint based on the distance relationship between the user's negative response vector and the continuous spatial response vector, wherein the second security constraint is related to personalized negative information.

[0127] Specifically, we define the combination of conditions to be optimized as x = {crowd density, crowd distance, crowd mood, crowd flow, crowd interaction, affected proportion, and severity of negative impact}, define the objective function f(x) to maximize the interaction training intensity, and define the safety constraints as follows:

[0128] (a) The distance between the user's negative reaction vector (User_action) and the cluster (clustering_action) is less than the threshold 1, i.e., Distance(User_action, clustering_action) < threshold 1.

[0129] Furthermore, it can also include emergent processing of the response vectors corresponding to specified condition combinations. Emergent processing is used to reveal condition combinations that also present high risk in the vicinity of the specified condition combination. This includes expanding the discrete or continuous point-like response vectors corresponding to the specified condition combination to obtain continuous spatial response vectors forming semantic clusters around the point-like vectors, thereby expanding the boundary of the corresponding specified response vectors. The degree of expansion is negatively correlated with the severity of the response, which is obtained by a doctor or physiologist after assessing the specified condition combination and can be considered known information. Thus, another safety constraint is defined as follows: the distance between the user's negative response vector (User_action) and the continuous spatial response vector (Specified_boundary_action) is less than a threshold of 2, and Distance(User_action, Specified_boundary_action) is less than threshold2.

[0130] Thus, based on Formula 2 above, we can select several safety seeds with higher safety 2 from the combination of conditions B. Then, we run several safety seeds to calculate the corresponding objective function value (i.e., the fmax(x) function value) and all safety constraint values ​​(i.e., the calculated distance).

[0131] In some embodiments, the step of determining the virtual world generation parameters includes: obtaining the interaction response intensity distribution characterizing the user's interaction training effect; obtaining the first condition combination with the strongest interaction training quantification index based on the interaction response intensity distribution; removing negative condition combinations from the first condition combination to obtain a target safety condition combination; using the target safety condition combination as the baseline value of the virtual world generation parameters; initializing a stochastic process model using a clustering algorithm, based on the target safety condition combination as the objective function and each safety constraint condition; predicting the virtual world generation parameters based on the stochastic process model; wherein the stochastic process model characterizes a subset of condition combinations selected from the target safety condition combination and is used as a data subset to fit a model describing the safety condition combinations.

[0132] Specifically, based on the aforementioned safety seeds, a stochastic process model is initialized for the objective function and each safety constraint. The initial data consists of several safety seeds, the corresponding objective function values, and all safety constraint values. The initialization method can be k-means clustering initialization, utilizing all current condition combinations (including condition combination A, cluster, specified condition combination, and condition combination B) to fit f(x) and a stochastic process model for each safety constraint. These stochastic process models will provide the predicted value (mean μ(·)) and uncertainty (standard deviation σ(·)) for any point x. The fitting process relies on the expectation-maximization algorithm for iterative optimization.

[0133] In some embodiments, predicting the virtual world generation parameters based on the stochastic process model includes: determining the upper confidence bound of each security constraint based on the acquisition function and the security constraints, and predicting the next virtual world generation parameters; wherein the upper confidence bound of the security constraints is used to constrain the next acquisition point to achieve the target interactive training effect while meeting the security constraints; in each iterative cycle, if the predicted next virtual world generation parameters satisfy the security constraints, the objective function dataset is updated based on the predicted next virtual world generation parameters, and the stochastic process function is refitted based on the updated objective function dataset until the target convergence condition is met, thus obtaining the final virtual world generation parameters.

[0134] Specifically, an optimistic confidence safety set Ŝ can be defined based on the safety constraints. The confidence safety set Ŝ is a conservative estimate of the true safety set S. It is defined as all points x that satisfy the upper confidence bound UCB_i(x) <= 0 (for all i) of the safety constraints. The calculation method for the upper confidence bound of the safety constraints includes calculating the upper confidence bound UCB_i(x) = μ{g_i}(x) + β × σ{g_i}(x) for each safety constraint, where β is a hyperparameter controlling the confidence level, usually β = 3 corresponds to a 99.7% confidence level, μ{g_i}(x) represents the predicted mean or posterior mean of the i-th safety constraint function g at the input point x, and σ{g_i}(x) represents the predicted standard deviation of the i-th safety constraint function g at the input point x. If UCB_i(x) <= 0, it means that there is a high confidence that point x satisfies the i-th constraint.

[0135] Next, the next evaluation point x_next is selected. This selection process depends on the acquisition function α(x). Conventional acquisition functions can be implemented using upper confidence bounds (UCB) and expected improvement (EI) functions. However, UCB requires manual tuning of the balancing parameters and may be overly conservative, leading to local optima. EI, on the other hand, may over-rely on the current optimum and is highly sensitive to the choice of the initial point, resulting in a "plateau" waste problem. To address this, this application further proposes an acquisition function optimization method: x_next = argmax_{x∈Ŝ}α(x), where argmax is a mathematical operator representing the search for the x in the set S^ that maximizes the function α(x). Optimizing the acquisition function using this method restricts the optimization range to the confidence-safe set Ŝ, thus ensuring that the new point x_next (the next point) not only has the potential to improve performance but also has a high confidence level to guarantee its safety, aligning with the conservative exploration expectation of this scheme.

[0136] However, compared to EI, UCB is more suitable for optimization using the above optimization methods. The reason is that in conservative exploration scenarios with defined confidence and safety sets, UCB's exploration strategy is more suitable. It can actively select points close to the safety boundary and find the optimal solution within the confidence and safety set more intuitively and efficiently. UCB, on the other hand, is not interested in actively exploring and expanding the boundary of the safety set. Furthermore, UCB avoids the "plateau" waste problem of EI. In this scenario, EI does not show a more obvious advantage than UCB and has its inherent disadvantages.

[0137] Furthermore, this application also includes performance evaluation, running experiments on x_next to calculate the corresponding f(x_next) and g_i(x_next).

[0138] Furthermore, this application also includes a safety check on x_next. Due to the uncertainty of the stochastic process model, for safety reasons, before adding x_next to the safe seed dataset for conditional combinations, it is necessary to verify whether g_i(x_next) actually satisfies the safety constraints. If an unexpected violation occurs, the point is discarded, marked as unsafe, and the uncertainty of the corresponding constraint model is increased (e.g., by increasing β).

[0139] Furthermore, this application also includes an extended dataset. (x_next, f(x_next)) is added to the objective function dataset, and (x_next, g_i(x_next)) is added to the corresponding constraint dataset.

[0140] Furthermore, this application also includes updating the model. All stochastic process models are refitted based on the expanded dataset, following the same fitting steps as described above.

[0141] Furthermore, this application also includes repeating the above steps until the maximum number of iterations or the target performance convergence condition is reached.

[0142] Furthermore, this application also includes selecting the point where the objective function f(x) is optimal from all safe seed datasets as the final solution, thereby obtaining a general and sufficiently safe combination of optimal conditions.

[0143] In the above embodiments, an interactive training method is proposed that, based on virtual reality and EEG interactive training schemes, triggers user interactive responses in an immersive and personalized environment, thereby achieving the effect of "virtual regulation." Furthermore, the parameters in the interactive training are adaptively adjusted according to relevant neurological, physiological, and behavioral indicators of the interactive responses, achieving the effect of adjusting the "interactive dosage."

[0144] In another embodiment, the method provided in this application can also be applied in the field of cognitive regulation and interaction, specifically including the following steps:

[0145] Step A: Construct an interactive training paradigm. See step 1 in the above embodiment for details, which will not be repeated here.

[0146] Step B: Initialize the population feature module. Refer to Step 2 in the above embodiment for details. However, the difference between Step B and Step 2 is that Step B does not include the affected proportion from Step 2, but adds a negative evaluation manifestation proportion. The negative evaluation manifestation proportion, ranging from 0 to 1, represents the proportion of the population exhibiting typical cognitive stress-induced behaviors. 0 means that no one in the population exhibits negative behavior, and 1 means that everyone in the population exhibits negative behavior. Other details are as described in Step 2 and will not be repeated here.

[0147] Step C: Initialize the cognitive stress feature module. Set the negative typical behavior and negative severity parameters as baseline values. The baseline value setting depends on the database obtained in Step A, selecting the combination of conditions with the strongest cognitive response intensity. Use the set values ​​of the cognitive stress features within this combination as the baseline values ​​for the current user's cognitive stress features. The interpretation of the cognitive stress feature values ​​is as follows:

[0148] Negative typical behaviors, textual descriptions (e.g., staring intently, covering mouth to laugh, frowning and accusing, etc.). Note that this parameter does not participate in the modulation; rather, its manifestation is controlled by the severity of the negative. That is, negative typical behaviors are the explicit representation of the severity of the negative.

[0149] The severity of the negative behavior ranges from 0 to 1, indicating the severity of the negative typical behavior (such as the duration of staring or the decibel level of the mocking voice). 0 represents no negative typical behavior, while 1 represents the complete manifestation of negative typical behavior.

[0150] For other details, please refer to step 3 in the above embodiments, which will not be repeated here.

[0151] Step D: Refer to step 4 in the above embodiment, which will not be repeated here.

[0152] Step E: Refer to step 5 in the above embodiment, which will not be repeated here.

[0153] Step F: Refer to step 6 in the above embodiment, which will not be repeated here.

[0154] Step G: Refer to step 7 in the above embodiment, which will not be repeated here.

[0155] Step H: Connect the virtual reality interface to the virtual reality device worn by the user. EEG equipment capable of collecting multi-channel EEG signals is worn by the user. A multi-channel physiological sensor data acquisition device (such as ECG and skin conductance monitoring sensors) is connected to the user. This device is used to collect physiological sensor data such as heart rate variability (HRV) and skin conductance response (GSR) before, during, and after cognitive interaction training. Complete the assembly and functional testing of all devices. Present the user with a cognitive state assessment scale, which is a subjective discomfort scale and a state and trait anxiety scale, to obtain the user's cognitive baseline data. Complete the assembly, calibration, and functional testing of all devices and acquisition modules. Introduce the selection response task in the training to the user. When a red dot appears in the upper right corner of the screen, the right-hand reactor needs to be pressed; when a green dot appears, the left-hand reactor needs to be pressed. Other details are as described in step 8 of the above embodiment and will not be repeated here.

[0156] Step I: The paradigm for cognitive regulation interaction training is implemented. Users need to walk straight in a constructed dynamic world, from the starting point to the end point. Users can observe left and right, and the walking speed is determined by user parameters. The duration of each round of cognitive regulation interaction training is determined by user parameters. It includes several trials, each trial representing a group of people meeting predetermined cognitive stress parameters walking towards the user. The interval between trials is set to pseudo-randomization based on a 5-second average. During the cognitive regulation interaction training, the user's EEG data is collected. Before, during, and after the cognitive regulation interaction training, the user's physiological sensory data, such as heart rate variability and skin conductance response, are acquired through a physiological acquisition device. Other details are as described in step 9 of the above embodiment and will not be repeated here.

[0157] Step J: Based on the EEG data collected during the user's cognitive stress resistance training process. Other details are as described in step 10 of the above embodiment and will not be repeated here.

[0158] Step K: Update, adjust, and construct a new dynamic virtual world. Compare the response indicators obtained in Step H based on EEG, physiological indicators, scale scores, and behavioral performance with those in the database constructed in Step A. See Step 11 in the above embodiment for details, which will not be repeated here.

[0159] In the above embodiments, an interactive training method is proposed that can be used for cognitive regulation of interaction (e.g., emergency decision-making scenarios in outdoor adventures, cross-cultural negotiation simulation scenarios in cross-border e-commerce, and decision-making sand table exercises in corporate crisis public relations). Based on virtual reality and EEG interactive training schemes, it triggers user interactive responses in an immersive and personalized environment, thereby achieving the effect of "virtual regulation." Furthermore, it adaptively adjusts the parameters in interactive training according to relevant neurological, physiological, and behavioral indicators of the interactive responses, thereby achieving the effect of adjusting the "interaction dosage."

[0160] For a detailed description of the interactive model training device based on extended reality and sensor analysis provided by this invention, please refer to the previous embodiment of the interactive model training method based on extended reality and sensor analysis, which will not be repeated here.

[0161] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of an interactive model training method based on extended reality and sensor analysis as described in any of the above embodiments.

[0162] This application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of an interactive model training method based on extended reality and sensor analysis as described in any of the above embodiments.

[0163] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the interactive model training method based on extended reality and sensor analysis described in any of the above embodiments.

[0164] It is understood that the computer device provided in this application can be a server, and its internal structure diagram can be as follows: Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores relevant data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the method provided in this application.

[0165] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of a portion of the structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. The computer device may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program may include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0166] It should be understood that the processor mentioned in the embodiments of this application can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0168] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A training method for an interactive model based on extended reality and sensor analysis, characterized in that, The method includes: The virtual world generation parameters and user characteristic parameters are determined. The steps for determining the virtual world generation parameters include: constructing an interaction training paradigm based on the interaction training sampling time and condition combinations; obtaining the interaction response intensity distribution representing the user interaction training effect when each group of subjects receives interaction training with corresponding condition combinations at the corresponding interaction training sampling time, based on the interaction training paradigm; wherein, the condition combinations include situational characteristic parameters and negative characteristic parameters in the virtual world generation parameters; obtaining the first condition combination with the strongest interaction training quantitative index based on the interaction response intensity distribution; removing negative condition combinations from the first condition combination to obtain the target safety condition combination; and using the target safety condition combination as the target safety condition combination. The process involves defining a baseline value for virtual world generation parameters; defining security constraints; iteratively searching for the parameters that provide the strongest interactive training effect based on the baseline value of the virtual world generation parameters, and using these parameters as the final virtual world generation parameters; initializing a stochastic process model based on the target security condition combination as the objective function and each security constraint using a clustering algorithm; predicting the virtual world generation parameters based on the stochastic process model; wherein the stochastic process model represents a subset of condition combinations selected from the target security condition combinations and is used as a data subset to fit a model describing the security condition combinations; and wherein the objective function is used to maximize the interactive training intensity. The step of removing negative condition combinations from the first condition combination to obtain the target safe condition combination includes: structuring the feedback information of the subjects in the interaction training paradigm according to the pre-trained language model to obtain negative response vectors; clustering the negative response vectors to obtain negative clusters; extracting negative response vectors from the top multiple condition combinations in the interaction response intensity distribution; matching the extracted negative response vectors with the negative clusters; if the extracted negative response vectors successfully match the negative clusters, obtaining a matching condition combination; removing the matching condition combination from the top multiple condition combinations in the interaction response intensity distribution to obtain a first condition combination; if the negative response vectors do not successfully match the negative clusters, using the top multiple condition combinations in the interaction response intensity distribution as the first condition combination; and continuing to remove specified condition combinations from the first condition combination to obtain the target safe condition combination; wherein, the specified condition combination is a predefined condition combination that causes the user to have a negative response. The steps for constructing the security constraints include: determining a first security constraint based on the distance relationship between the user's negative response vector and the negative cluster, wherein the first security constraint is related to grouped negative information; performing emergent processing on the response vectors corresponding to the specified condition combination to expand the discrete or continuous point-like response vectors into a continuous spatial response vector that forms a semantic cluster around the point-like vectors; and determining a second security constraint based on the distance relationship between the user's negative response vector and the continuous spatial response vector, wherein the second security constraint is related to personalized negative information. A dynamic virtual world is constructed based on the virtual world generation parameters; Acquire EEG data of the user during interactive training in the dynamic virtual world based on the user characteristic parameters, as well as at least one of physiological sensor data and behavioral sensor data; Based on the EEG data, the user's interactive training effect is determined by combining at least one of physiological sensor data and behavioral sensor data. During the iterative process, when the interactive training effect does not meet the convergence condition, the virtual world generation parameters are adjusted according to the preset prediction model, and a new dynamic virtual world is constructed based on the adjusted virtual world generation parameters until the interactive training effect meets the convergence condition, at which point the iterative process ends and the interactive model is trained.

2. The method according to claim 1, characterized in that, The step of determining the user's interactive training effect based on the EEG data, combined with at least one of physiological sensor data and behavioral sensor data, includes: Based on the trigger tags set in the experimental paradigm when the event occurs, continuous EEG data is divided into multiple trials; each trial covers EEG data before and for a period of time after the event begins; the EEG data from multiple trials are superimposed and averaged to obtain event-related potentials; the positive and negative characteristics of the event-related potentials in different time windows characterize the neural activity characteristics of the electrode region. Collect physiological sensor data of users before, after, or during interactive training, and determine the changes and differences in physiological indicators of users before, after, or during interactive training based on the corresponding detection algorithm. And / or, Collect user behavioral sensor data or scale data before, after, or during interactive training, and determine the changes and differences in the user's cognitive process before, after, or during interactive training based on the corresponding response theory and its cognitive significance. The user's interactive training effect is determined based on the event-related potentials, changes and differences in physiological indicators, and / or changes and differences in cognitive processes.

3. The method according to claim 1, characterized in that, The step of adjusting the virtual world generation parameters according to a preset prediction model, and then constructing a new dynamic virtual world based on the adjusted virtual world generation parameters, includes: The virtual world generation parameters are input into a preset prediction model, and the interactive training effect is predicted based on the preset prediction model. When the difference between the predicted interactive training effect and the user's current actual interactive training effect meets the preset conditions, new virtual world generation parameters are predicted according to the preset prediction model, and a new dynamic virtual world is constructed according to the new virtual world generation parameters.

4. The method according to claim 1, characterized in that, The virtual world generation parameters predicted based on the stochastic process model include: The confidence upper bound of each security constraint is determined based on the acquisition function and the security constraints, and the next virtual world generation parameters are predicted; wherein, the confidence upper bound of the security constraints is used to constrain the next acquisition point to achieve the target interactive training effect under the premise of meeting the security constraints. In each iteration, if the predicted parameters for the next virtual world generation satisfy the safety constraints, the objective function dataset is updated based on the predicted parameters for the next virtual world generation, and the stochastic process function is refitted based on the updated objective function dataset until the objective convergence condition is met, thus obtaining the final virtual world generation parameters.

5. An interactive model training device based on extended reality and sensor analysis, characterized in that, The device includes: The parameter determination module is used to determine virtual world generation parameters and user feature parameters. The steps for determining virtual world generation parameters include: constructing an interaction training paradigm based on the interaction training sampling time and condition combinations; obtaining the interaction response intensity distribution representing the user interaction training effect when each group of subjects receives interaction training with corresponding condition combinations at the corresponding interaction training sampling time based on the interaction training paradigm; wherein, the condition combinations include situational feature parameters and negative feature parameters in the virtual world generation parameters; obtaining the first condition combination with the strongest interaction training quantitative index based on the interaction response intensity distribution; removing negative condition combinations from the first condition combination to obtain the target safety bar. The system combines target security conditions as the baseline values ​​for the virtual world generation parameters. Security constraints are defined, and under these constraints, the system iteratively searches for the parameters that provide the strongest interactive training effect, using these baseline values ​​as the final virtual world generation parameters. A clustering algorithm is used to initialize a stochastic process model based on the target security condition combination as the objective function and each security constraint. The virtual world generation parameters are then predicted using the stochastic process model. The stochastic process model represents a subset of conditions selected from the target security condition combination and is used as a data subset to fit and describe the security conditions. The model of the combination of conditions; wherein the objective function is used to maximize the interaction training intensity; the step of removing negative condition combinations from the first condition combination to obtain the target safe condition combination includes: performing structured processing on the feedback information of the subjects in the interaction training paradigm according to the pre-trained language model to obtain negative response vectors, performing clustering processing on the negative response vectors to obtain negative clusters; extracting negative response vectors from the top multiple condition combinations in the interaction response intensity distribution, matching the extracted negative response vectors with the negative clusters, and if the extracted negative response vectors successfully match the negative clusters, obtaining a matching condition combination, and removing negative condition combinations from the interaction response intensity distribution. The first condition combination is obtained by removing matching condition combinations from the top-ranking condition combinations in the degree distribution. If the negative response vector does not match the negative cluster, the top-ranking condition combinations in the interaction response intensity distribution are used as the first condition combination. The specified condition combination is then removed from the first condition combination to obtain the target safety condition combination. The specified condition combination is a predefined combination that causes the user to have a negative response. The construction step of the safety constraint includes: determining the first safety constraint based on the distance relationship between the user's negative response vector and the negative cluster, wherein the first safety constraint is related to grouped negative information.Emergent processing is performed on the reaction vectors corresponding to the specified combination of conditions to expand the discrete or continuous point-like reaction vectors to obtain a continuous spatial reaction vector that forms a semantic cluster around the point-like vectors. The second security constraint is determined based on the distance relationship between the user's negative reaction vector and the continuous spatial reaction vector, wherein the second security constraint is related to personalized negative information. The world generation module is used to construct a dynamic virtual world based on the virtual world generation parameters. The virtual reality module is used to acquire EEG data of a user during interactive training in the dynamic virtual world based on the user's characteristic parameters through an EEG acquisition unit, and to acquire at least one of physiological sensing data and behavioral sensing data through a physiological sensing data acquisition unit and a behavioral sensing data acquisition unit. The calculation module is used to determine the user's interactive training effect based on the acquired EEG data, combined with at least one of physiological sensor data and behavioral sensor data. The adjustment module is used to adjust the virtual world generation parameters according to the preset prediction model when the interaction training effect does not meet the convergence condition during the iterative process, and continue to construct a new dynamic virtual world according to the adjusted virtual world generation parameters until the interaction training effect meets the convergence condition, at which point the iterative process ends and the interaction model is trained.

6. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the interactive model training method based on extended reality and sensor analysis as described in any one of claims 1 to 4.