Double virtual reality social anxiety intervention system and method based on multi-target tracking normal form
Through the two-person virtual reality social anxiety intervention system based on the multi-target tracking paradigm, the scene adaptability and patient compliance problems of traditional treatment methods are solved, personalized VR treatment is achieved, the accuracy of attention regulation and evaluation is improved, and the effectiveness of treatment is enhanced.
Patent Information
- Application Number
- CN202510732983.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional social anxiety treatment methods have poor adaptability to scenarios, low patient compliance, and insufficient attention regulation and assessment in existing VR systems, making them unable to meet personalized needs.
A two-person virtual reality social anxiety intervention system based on the multi-target tracking paradigm constructs an adaptive attention guidance mechanism through a multi-target dynamic generation and regulation module, a laser interaction and judgment module, a two-person interaction and feedback module, a behavioral data tracking, traceability and visualization module, and a behavioral data tracking and visualization module to achieve two-person collaborative interaction and feedback and conduct multi-dimensional evaluation.
It improves the fun of treatment and patient compliance, realizes dynamic regulation of user attention, provides a comprehensive evaluation system, and improves the accuracy of treatment effects and personalized support.
Smart Images

Figure CN120661808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to virtual reality technology and psychological intervention technology, and in particular to a two-person virtual reality social anxiety intervention system and method based on a multi-target tracking paradigm. Background Art
[0002] Social anxiety disorder has become a common mental health problem. Traditional treatments for social anxiety, such as cognitive behavioral therapy (CBT) and exposure therapy, rely primarily on realistic scenario simulations. In cognitive behavioral therapy, therapists alleviate symptoms by guiding patients to identify and change negative thought patterns and behavioral habits; exposure therapy involves having patients face anxiety-provoking social scenarios to alleviate fear responses. However, these traditional therapies have many limitations. On the one hand, real-life scenarios are less adaptable and difficult to fully fit the specific circumstances of each patient. For example, different patients have different fears and levels of anxiety about social situations, but the simulation scenarios in traditional therapies are often relatively fixed and cannot accurately match individual differences. On the other hand, the treatment process is relatively boring, and patient compliance is generally low. Long conversations and repetitive scenario exercises can easily make patients feel bored, reducing their enthusiasm for participating in treatment, which in turn affects the treatment effect.
[0003] With the development of virtual reality technology, its application in the field of psychotherapy has gradually increased. Existing VR psychological intervention systems can simulate social scenes and provide patients with a certain degree of immersive experience. For example, by creating virtual gatherings, meetings and other scenes, patients can practice social interactions in a virtual environment. However, these systems generally lack a dynamic regulation mechanism for user attention bias. Attention bias refers to the tendency of an individual to pay preferential attention to certain specific information when faced with multiple information. Among patients with social anxiety, there is often a phenomenon of excessive attention to negative social information, and existing systems cannot effectively intervene in this problem. In addition, the evaluation indicators of existing systems are relatively simple, usually only based on the patient's subjective feelings or simple behavioral performance. They cannot fully and accurately reflect the patient's behavioral changes and treatment effects during the treatment process, and it is difficult to meet the needs of personalized treatment.
[0004] The shortcomings of existing traditional social anxiety treatment methods and VR psychological intervention systems are mainly reflected in the following aspects:
[0005] Poor scenario adaptability: The real-life scenario simulation of traditional therapies is difficult to adjust according to the specific circumstances of each patient, and cannot accurately meet the treatment needs of different patients, resulting in limited treatment effects.
[0006] Low patient compliance: The boring nature of traditional therapies makes patients less motivated to participate in treatment, affecting the continuity and ultimate effectiveness of treatment.
[0007] Lack of dynamic regulation mechanism for attention bias: Existing VR psychological intervention systems are unable to dynamically intervene in the problem of social anxiety patients' excessive attention to negative social information, and it is difficult to fundamentally change the patient's attention pattern.
[0008] Single evaluation indicator: The existing system's evaluation method cannot fully reflect the patient's behavioral changes and treatment effects during the treatment process, which is not conducive to the timely adjustment and optimization of the treatment plan.
[0009] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0010] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a two-person virtual reality social anxiety intervention system and method based on a multi-target tracking paradigm.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] In a first aspect of the present invention, a two-person virtual reality social anxiety intervention system based on a multi-target tracking paradigm comprises:
[0013] A multi-target dynamic generation and control module dynamically generates positive targets in virtual scenes to guide attention and establish positive emotional associations, as well as negative distractors to simulate social interference and train attention inhibition. Distractor parameters, such as the number and frequency of distractors, are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism.
[0014] High-precision laser interaction and judgment module, which is an optional module, determines the hit or mishit status based on the coordinate distance between the laser beam and the target object, providing accurate interactive feedback;
[0015] The two-player collaborative interaction and feedback module uses a real-time network communication protocol to synchronize two players in a virtual scene. It also triggers multi-sensory feedback mechanisms such as particle effects, sound effects, and score-linked feedback through collaborative tasks to enhance user participation motivation.
[0016] The behavioral data tracking and visualization module collects user operation data sets in real time, such as the number of target hits, reaction time, and number of mishits, and generates quantitative evaluation reports through a multi-dimensional visualization interface. For example, multi-dimensional data visualization is performed through heat maps, hit rate curves, and HUD interfaces to build a closed-loop evaluation system.
[0017] In a second aspect of the present invention, a two-person virtual reality social anxiety intervention method based on the system comprises:
[0018] Dynamic generation and regulation of multiple targets: Based on a random algorithm, positive targets are dynamically generated in virtual scenes to guide attention and establish positive emotional associations, as well as negative distractors to simulate social interference information and train attention inhibition. Distractor parameters are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism.
[0019] Laser interaction determination: This is an optional step that triggers hit or miss status feedback by comparing the distance between the laser beam contact point and the target object coordinates;
[0020] Two-player collaborative interaction and feedback: Based on real-time network communication protocols, two players can synchronize virtual scenes and trigger multi-sensory feedback mechanisms through collaborative tasks to enhance user participation motivation.
[0021] Closed-loop evaluation of behavioral data: Collect user operation data sets in real time, generate quantitative evaluation reports through a multi-dimensional visualization interface, and build a closed-loop evaluation system.
[0022] The present invention has the following beneficial effects:
[0023] The present invention proposes a two-person virtual reality social anxiety intervention system based on a multi-target tracking paradigm, which overcomes the problems of poor scene adaptability of traditional VR exposure therapy, low patient compliance, and insufficient attention regulation and evaluation of existing VR systems. Specifically, the present invention solves the problem of poor scene adaptability of traditional social anxiety treatment methods, constructs a highly flexible and accurate VR training system to meet the personalized treatment needs of different patients. Improve patient compliance during the treatment process, increase the fun of treatment through a gamified interactive mechanism, and enhance patient enthusiasm for participating in treatment. Realize dynamic regulation of user attention bias, and utilize the system's multi-target dynamic generation and regulation functions to guide patients to focus on positive goals and reduce attention to negative distractions. Establish a comprehensive and scientific evaluation system, and through behavioral data tracking and visualization functions, accurately reflect the patient's behavioral changes and treatment effects during the treatment process in real time, providing strong support for subsequent adjustments to the treatment plan.
[0024] The core of this invention is to combine the multi-target tracking paradigm with two-person VR interaction to construct a social anxiety intervention system that dynamically regulates attention bias. Based on the system design of the multi-target tracking paradigm, this invention includes technical modules such as dynamic generation and regulation of multiple targets, high-precision laser interaction and judgment (optional), two-person collaborative interaction and feedback, and behavioral data tracking and visualization. Compared with traditional technologies, the advantages of this invention are specifically reflected in the following:
[0025] 1. Multi-target dynamic control mechanism: By generating dynamic interactive scenarios between positive targets and negative distractors, and leveraging randomized algorithms, physics engines, and real-time performance evaluation, the system precisely guides user attention (e.g., by adjusting target attributes, motion trajectories, and the number of distractors, forcing users to actively filter out negative information). This dynamic control mechanism, specifically targeting the attention bias of social anxiety patients, guides patients to change their attention patterns through the flexible setting and adjustment of positive targets and negative distractors.
[0026] 2. Two-person collaborative interaction design: Introducing real-time network synchronization technology (WebSocket protocol) to build a virtual scene of two-person collaboration or competition, enhancing user participation motivation through interactive feedback (particle effects, sound effects, score linkage), and breaking through the boring limitations of traditional single-person VR treatment. Compared with traditional social anxiety treatment methods, the highly flexible and accurate VR training system constructed by the present invention using virtual reality technology can be personalized according to the conditions of different patients, improving the adaptability of the scene. At the same time, the gamification interaction mechanism increases the fun of the treatment and effectively improves patient compliance.
[0027] 3. Data-driven closed-loop evaluation system: Through multi-dimensional behavioral data tracking (number of hits, reaction time, heat map distribution, etc.) and visualization, a personalized treatment closed loop of "intervention-feedback-adjustment" is formed, addressing the single-item evaluation problem of existing systems. This comprehensive behavioral data evaluation system covers multiple key data indicators and enables real-time data tracking and visualization.
[0028] In summary, the present invention systematically solves the problems of scene adaptability, attention intervention, user compliance, and efficacy evaluation in the treatment of social anxiety through the coordinated cooperation of three technical elements: dynamic scene generation, two-person interactive motivation, and real-time data feedback. Compared with existing VR psychological intervention systems, the present invention realizes the dynamic regulation of user attention bias and can specifically intervene in the problem of social anxiety patients' excessive attention to negative social information. Moreover, through a comprehensive and scientific evaluation system, it can more accurately reflect the patient's treatment effect, provide strong support for the adjustment of treatment plans, and thus improve the effectiveness and scientific nature of the treatment.
[0029] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is an architectural diagram of a two-person virtual reality social anxiety intervention system based on a multi-target tracking paradigm according to an embodiment of the present invention.
[0031] Figure 2 This is an overall flow chart of a two-person virtual reality social anxiety intervention method based on a multi-target tracking paradigm according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.
[0033] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0034] The present invention proposes a two-person virtual reality social anxiety intervention system and method based on a multi-target tracking paradigm, aiming to solve the problems of poor scene adaptability of traditional VR exposure therapy, low patient compliance, and insufficient attention regulation and evaluation of existing VR systems. Specifically: The present invention solves the problem of poor scene adaptability of traditional social anxiety treatment methods, constructs a highly flexible and accurate VR training system to meet the personalized treatment needs of different patients. Improve patient compliance during the treatment process, increase the fun of treatment through a gamified interactive mechanism, and enhance patient enthusiasm for participating in treatment. Realize dynamic regulation of user attention bias, utilize the system's multi-target dynamic generation and regulation functions to guide patients to focus on positive goals and reduce attention to negative distractions. Establish a comprehensive and scientific evaluation system, and through behavioral data tracking and visualization functions, accurately reflect the patient's behavioral changes and treatment effects during the treatment process in real time, and provide strong support for subsequent adjustments to the treatment plan.
[0035] See Figure 1An embodiment of the present invention provides a two-person virtual reality social anxiety intervention system based on a multi-target tracking paradigm, comprising a multi-target dynamic generation and regulation module, a two-person collaborative interaction and feedback module, and a behavioral data tracking and visualization module. The multi-target dynamic generation and regulation module dynamically generates positive targets in a virtual scene to guide attention focus and establish positive emotional associations, as well as negative distractors to simulate social interference information and train attention inhibition (positive targets and negative distractors can be generated based on random algorithms and preset rules). Distractor parameters, such as the number and frequency of distractors, are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism. The two-person collaborative interaction and feedback module uses a real-time network communication protocol to synchronize the two-person virtual scene and triggers multi-sensory feedback mechanisms such as particle effects, sound effects, and score-linked feedback through collaborative tasks to enhance user participation motivation. The behavioral data tracking and visualization module collects user operation data sets in real time, such as the number of target hits, reaction time, and number of mishits, and generates quantitative evaluation reports through a multi-dimensional visualization interface. For example, multi-dimensional data visualization is performed through heat maps, hit rate curves, and a head-up display (HUD) interface, establishing a closed-loop evaluation system. Furthermore, the system can also include a high-precision laser interaction and judgment module, which is an optional module that can determine the hit or mishit status based on the coordinate distance between the laser beam and the target object, and provide accurate interactive feedback.
[0036] In some embodiments, the system is developed based on the Unity 3D engine, integrating a physics engine (such as the Rigidbody component) and a path planning tool (such as the NavMeshAgent component) to achieve random movement and intelligent obstacle avoidance of the target object.
[0037] In some embodiments, the multi-target dynamic generation and control module includes: a random algorithm generation unit, which generates differentiated targets and distractions according to regional probability based on the Poisson disk distribution algorithm and configurable parameter objects such as ScriptableObject configuration files; a dynamic difficulty adaptation unit, which dynamically adjusts the number and frequency of distractions according to the user's hit rate and reaction time through a dynamic performance evaluator such as a PerformanceEvaluator script to achieve "ability-difficulty" matching.
[0038] In some embodiments, the high-precision laser interaction and judgment module includes: a laser emission and contact detection unit, which emits laser rays and detects contact coordinates through a ray collision detection method such as the Physics.Raycast method; a state machine judgment unit, which triggers a hit or mishit state machine through distance threshold comparison and provides feedback on the correctness of the operation.
[0039] In some embodiments, the two-player collaborative interaction and feedback module includes: a network synchronization unit, which realizes real-time transmission and delay compensation of handle motion data and virtual character position based on the WebSocket protocol and Kalman filter prediction algorithm; a collaborative task processing unit, which triggers particle special effects and sound effects of two-player synchronous hits through multimodal feedback components such as ParticleSystem components and AudioSource components, and updates the HUD interface score.
[0040] In some embodiments, the behavioral data tracking and visualization module includes: a sliding window calculation unit, which uses a sliding window algorithm (the window size is, for example, 10 operations) to calculate the average reaction time in real time; a heat map generation unit, which maps the hit coordinates to color intensity through a shader tool such as a ShaderGraph shader to generate an attention distribution heat map.
[0041] In some embodiments, in the multi-target dynamic generation and regulation module: positive targets use bright colors and animal images, and negative distractors use dark colors and warning signs, thereby strengthening attention guidance through differences in visual attributes.
[0042] In some embodiments, in the two-player collaborative interaction and feedback module: the virtual character control processes the handle accelerometer data through a low-pass filtering algorithm, smoothly converts it into motion instructions, and controls the character movement and perspective rotation.
[0043] In some embodiments, the network synchronization in the two-person collaborative interaction step further includes a timestamp compensation processing mechanism, which performs the following operations in sequence:
[0044] (a) Delay calculation: Calculate the network transmission delay based on the timestamp generated by the data packet at the sending end and the local timestamp when it is received by the receiving end;
[0045] (b) Predicted position correction: When the delay exceeds a preset threshold, the predicted position at the next moment is calculated based on the predicted position and predicted speed at the current moment, combined with the delay, to correct the target object's position deviation caused by network delay;
[0046] (c) Trajectory smoothing calibration: Within the sending timestamp interval of two adjacent frames of data, the actual position of the sender at the two timestamps is linearly interpolated according to the current rendering time of the receiver to generate a smooth transition trajectory to eliminate prediction errors and ensure visual coherence.
[0047] In some embodiments, the network synchronization in the two-person collaborative interaction step further includes a collaborative processing mechanism that performs the following operations in sequence:
[0048] (a) Real-time data transmission and timestamp: Operational data packets are transmitted in real time through a full-duplex communication protocol, and a timestamp generated by the sender is appended to each data packet;
[0049] (b) Motion parameter prediction rendering: The current position and speed of virtual objects are predicted based on historical motion data. When real-time data has not arrived, the predicted position is used for rendering to avoid image stagnation.
[0050] (c) Collaborative correction of delay error: After receiving real data with timestamps, the following collaborative operations are performed: the network delay is calculated, and if it exceeds the threshold, the target position is corrected based on the predicted speed; within the timestamp interval of adjacent data packets, a smooth motion trajectory is generated through position interpolation to eliminate the visual jump between the predicted position and the actual position.
[0051] In some embodiments, in the behavior data tracking and visualization module: after the training is completed, a time series visualization unit such as the DataVisualizer.GenerateLineChart() method is called to generate a hit rate time curve report, where the horizontal axis is the training time and the vertical axis is the hit rate percentage.
[0052] See Figure 2 The embodiment of the present invention further provides a two-person virtual reality social anxiety intervention method based on the system, comprising:
[0053] Dynamic generation and regulation of multiple targets: Based on a random algorithm, positive targets are dynamically generated in virtual scenes to guide attention and establish positive emotional associations, as well as negative distractors to simulate social interference information and train attention inhibition. Distractor parameters are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism.
[0054] Laser interaction determination: This is an optional step that triggers hit or miss status feedback by comparing the distance between the laser beam contact point and the target object coordinates;
[0055] Two-player collaborative interaction and feedback: Based on real-time network communication protocols, two players can synchronize virtual scenes and trigger multi-sensory feedback mechanisms through collaborative tasks to enhance user participation motivation.
[0056] Closed-loop evaluation of behavioral data: Collect user operation data sets in real time, generate quantitative evaluation reports through a multi-dimensional visualization interface, and build a closed-loop evaluation system.
[0057] The core of this invention is to combine the multi-target tracking paradigm with two-person VR interaction to build a social anxiety intervention system that dynamically regulates attention bias. Specifically, it is manifested in:
[0058] 1. Multi-target dynamic control mechanism: By generating dynamic interaction scenarios between positive targets and negative distractions, and utilizing randomized algorithms, physics engines, and real-time performance evaluation, it achieves precise guidance of user attention (e.g., by adjusting target attributes, motion trajectories, and the number of distractions, forcing users to actively filter out negative information).
[0059] 2. Two-player collaborative interaction design: Introduce real-time network synchronization technology (such as the WebSocket protocol) to build a virtual scene of two-player collaboration or competition. Enhance user participation motivation through interactive feedback (particle effects, sound effects, score linkage), and break through the boring limitations of traditional single-player VR therapy.
[0060] 3. Data-driven closed-loop evaluation system: Through multi-dimensional behavioral data tracking (number of hits, reaction time, heat map distribution, etc.) and visual presentation, a personalized treatment closed loop of "intervention-feedback-adjustment" is formed to solve the problem of single evaluation in the existing system.
[0061] In summary, the present invention systematically solves the problems of scene adaptability, attention intervention, user compliance and efficacy evaluation in the treatment of social anxiety through the coordinated cooperation of the three technical features of dynamic scene generation, two-person interactive motivation, and real-time data feedback.
[0062] The following further describes the specific embodiments of the present invention, system construction examples, and its working principles and working processes.
[0063] Multi-objective dynamic generation and control module:
[0064] Based on the multi-target tracking paradigm, positive targets and negative distractions are generated through a random algorithm, combined with a physics engine to achieve dynamic movement and intelligent obstacle avoidance, and the number and frequency of distractions are dynamically adjusted according to the user's real-time performance to guide attention towards positive targets.
[0065] The multi-target dynamic generation and control module uses a multi-target tracking paradigm to generate positive targets and negative distractors based on a random algorithm and preset rules when the virtual scene is initialized. The CuePhaseController script can be used to control the target object's flashing animation and dynamically adjust the duration according to the training difficulty. The DynamicMovementSystem component and the NavMeshAgent component can be used to achieve the target object's random movement and intelligent obstacle avoidance based on the physics engine. At the same time, the number and frequency of distractors are dynamically adjusted according to the user's real-time performance (hit rate, reaction time), forming an attention guidance mechanism with adaptive "ability-difficulty" matching.
[0066] High-precision laser interaction and judgment module:
[0067] By using the VR handle's laser beam to contact the interactive surface, the touch point coordinates are compared with the target object coordinates to accurately determine the hit / mishit status and provide feedback on the correctness of the operation.
[0068] As an optional module, the high-precision laser interaction and judgment module uses the built-in laser emission device and touch point detection sensor of the VR controller. It can emit laser rays through Physics.Raycast, calculate the distance between the touch point coordinates and the target object coordinates, and compare it with the preset threshold, triggering the hit / mishit state machine to achieve accurate interaction judgment.
[0069] Two-person collaborative interaction and feedback module:
[0070] The WebSocket protocol is used to achieve two-player virtual scene synchronization, support handle motion control of virtual characters and collaborative task interaction, and enhance collaborative experience and participation motivation through particle effects, sound effects and score linkage.
[0071] The two-player collaborative interaction and feedback module uses the WebSocket protocol to achieve real-time transmission of handle motion data, click operations, etc. between the client and the server. It reduces network latency through Kalman filter prediction and timestamp compensation algorithms to ensure synchronization of two-player virtual scenes. Users control the movement and perspective adjustment of virtual characters through handle accelerometer data, and interact with scripted interactive objects. When two people hit the target synchronously, particle effects, sound effects, and HUD interface score updates are triggered. If they miss, the error type is recorded and the target object is reset, thereby enhancing participation motivation through social incentives.
[0072] Behavioral data tracking and visualization module:
[0073] Real-time tracking of key data such as the number of hits and reaction time is achieved through the HUD interface, which displays and generates heat maps and hit rate curve reports in real time to build a multi-dimensional evaluation system.
[0074] Behavioral data tracking and visualization module: Through data embedding, key data such as the number of target hits, average reaction time, and number of false hits are collected in real time, and dynamic indicators are calculated and stored using the sliding window algorithm; the HUD interface is built through Unity Canvas, and the data is displayed in real time with the TextMeshPro component. With the help of ShaderGraph, a heat map is generated to intuitively present the attention distribution. After the training is completed, the DataVisualizer.GenerateLineChart() method is called to generate a hit rate time curve report, forming a multi-dimensional training evaluation system.
[0075] Timestamp compensation algorithm:
[0076] During two-person synchronous interaction, network latency can prevent both parties from seeing the same state of an object in real time, impacting the collaborative experience and the accuracy of evaluation data. To address this, this paper proposes a data synchronization and position correction algorithm based on timestamp compensation. By adding precise timestamps to operational events, this algorithm achieves end-to-end time flow calibration and motion restoration.
[0077] 1. Get network delay
[0078] Δt=T r -T s
[0079] Among them, T s Indicates the timestamp of the data generated at the sending end, T r It indicates the local timestamp when the data is received by the receiver, and Δt is the delay time generated during network transmission.
[0080] 2. State prediction and position correction (triggered when Δt>50ms)
[0081]
[0082] Predicted information based on current location and speed prediction information (calculated by the Kalman filter), and the target's position at time t+Δt after the delay is predicted using the network delay Δt. This method can correct for the target's position deviation caused by network delay, making the target state seen by both ends more consistent.
[0083] 3. Linear interpolation calibration (eliminating prediction errors)
[0084]
[0085] T2 and T1 are the sending timestamps of two adjacent frames of data. x2 and x1 are the sending end’s timestamps at T2 and T1.
[0086] The actual position at time t is the current rendering time at the receiving end (between T2 and T1). When the real data arrives, a smooth transition trajectory is generated through linear interpolation to eliminate the jump between the predicted position and the actual position and ensure visual coherence.
[0087] Steps to implement the method / algorithm:
[0088] (1) Multi-objective dynamic control
[0089] Multi-target dynamic control includes scene initialization and target generation, cues and prompts to guide attention, dynamic tracking, and difficulty adaptation. Initialization creates a foundation for attention guidance through the generation of differentiated targets, cues and prompts enhance target recognition, and dynamic tracking adjusts difficulty in real time based on user performance, forming a closed loop of "stimulus generation - guided recognition - adaptive adjustment."
[0090] Scene initialization and target generation:
[0091] Call the SceneGenerator script to generate positive targets and negative distractors in the virtual scene; customize the generation rules through the ScriptableObject configuration file, including:
[0092] Spatial distribution: For example, the probability of generating positive targets in the center area is 70%, and the probability of generating negative distractors in the edge area is 70%;
[0093] Visual attributes: Positive targets use bright tones (such as orange Color32(255,165,0,255)) and animal images, while negative distractors use dark tones (such as gray Color32(100,100,100,255)) and warning signs.
[0094] Based on the emotional priming effect and attention bias theory in psychology, the differences between positive and negative stimuli are constructed through visual symbols and cognitive associations. The reference design standards for positive targets and negative distractors are given from the dimensions in Table 1 below:
[0095] Table 1
[0096]
[0097]
[0098] Cues guide attention:
[0099] Implementation: Call the CuePhaseController script (inherited from MonoBehaviour) and trigger the target object's flashing animation through the InvokeRepeating method; dynamically adjust the phaseDuration variable (initial value 3 seconds) and flashing frequency according to the training difficulty level (beginner / intermediate / advanced):
[0100] Dynamic tracking and difficulty adaptation:
[0101] Implementation: Enable the DynamicMovementSystem component, assign physical properties to the target object through the Rigidbody component, and use the AddForce method to achieve random movement; plan the path through the SetDestination method of the NavMeshAgent component, and detect the ObstacleAvoidanceType parameter to achieve intelligent obstacle avoidance; after every 10 target interactions, call the PerformanceEvaluator script to evaluate performance: if the hit rate is ≥80% and the reaction time is <1.5 seconds, add 2 distractors and increase their frequency by 20%; otherwise, remove 1 distractor and reduce its frequency by 10%.
[0102] (2) Two-person collaborative interaction
[0103] Two-person collaborative interaction includes network connection and data synchronization, virtual character control and interaction, and collaborative feedback and task processing.
[0104] First, a network connection is established to ensure synchronization of the two-player scene, then the virtual character is controlled through the handle data to achieve interaction, and finally, feedback is triggered based on the collaborative results to strengthen the motivation for participation, forming a complete process of "communication foundation-interaction execution-result feedback".
[0105] Network connection and data synchronization:
[0106] Implementation: Build a communication framework based on the WebSocket-sharp library, create the WebSocketClient class on the client, and the WebSocketServer class on the server; encapsulate the handle motion data in JSON format (such as {"type":"move","x":0.5,"y":0.3,"z":0.2}) and transmit it in real time through full-duplex communication; use the Kalman filter algorithm to predict the next position of the local virtual character (based on the previous three frames of data); if the network delay is greater than 50ms, trigger position compensation through timestamp comparison.
[0107] Virtual character control and interaction:
[0108] Implementation: Use the Input.GetAxis method of the CameraController script to obtain joystick input to control the rotation and movement of the virtual character. The handle accelerometer data is low-pass filtered to remove noise and then converted into virtual character movement instructions. Event linkage with scripted interactive objects (such as clicking a virtual button to trigger a task) is supported.
[0109] Collaborative feedback and task processing:
[0110] Implementation: When two players simultaneously hit the target within the preset time difference, the ParticleSystem component is called to play particle effects, the success sound effect is played through AudioSource, and the HUD interface score (Text component) and task progress (Slider component, refresh every 0.1 seconds) are updated. If the target is not hit or is accidentally hit, the error type is recorded (such as "accidentally hit the distraction" or "missed the target"), the HUD error prompt text and warning sound effect are triggered, and the target is regenerated after 2 seconds.
[0111] (3) Behavioral data processing
[0112] Behavioral data processing includes real-time data collection, indicator calculation and storage, as well as visualization and report generation.
[0113] Raw data is collected through point embedding, evaluation indicators are formed through algorithm calculation, and the results are finally output in a visual form, building an evaluation closed loop of "data collection-analysis and processing-result display".
[0114] Real-time data collection:
[0115] Implementation method: Use DataTracker script to record operation data in OnClick, OnMove and other event callbacks, and use Dictionary<string,object> Storage; Collection parameters include: number of targets hit, hit accuracy, average reaction time, hit rate change over time, number of false hits.
[0116] Indicator calculation and storage:
[0117] Implementation method: Use sliding window algorithm (window size 10 operations) to calculate the average reaction time in real time, and calculate the hit rate by accumulating the number of hits and the total number of interactions; store the original data and calculation results in the dynamic behavior database.
[0118] Visualization and report generation:
[0119] Implementation method: Use the TextMeshProUGUI component to display key indicators (such as the number of hits and accuracy) in real time on the HUD interface; use ShaderGraph to write a shader, convert the hit coordinates into UV coordinates, and map the color intensity (0-255) according to the number of hits to generate a heat map; after training, call the DataVisualizer.GenerateLineChart() method to generate a hit rate time curve report based on the System.Drawing library (horizontally in minutes, vertically in percentage, in PNG format).
[0120] In some embodiments, the present invention provides a two-person virtual reality social anxiety intervention system and method based on a multi-target tracking paradigm, specifically comprising:
[0121] Multi-objective dynamic generation and control
[0122] Generation phase: When the synchronized virtual scene is initialized, the system uses a random algorithm to generate positive targets and negative distractors based on the preset rules of the multi-target tracking paradigm. The preset rules cover the distribution patterns of targets and distractors in the scene, generation probabilities, and other content. For example, it can be set to generate positive targets with a higher probability in specific areas of the scene, while negative distractors are randomly generated in other areas. The appearance (such as positive targets can be designed as cute animal images, and negative distractors are designed as patterns with warning signs) of these targets and distractors, shapes (such as circles, squares, etc.), colors (positive targets use bright and warm tones, and negative distractors use dark tones) and other attributes can be flexibly adjusted according to different training needs to adapt to diverse treatment scenarios.
[0123] Cue phase: During the cue phase, the system calls the CuePhaseController script (the CuePhaseController script is a code module written in the system program, which is used to control the animation performance of the target object in the cue phase) and activates the flashing animation of the target object by modifying the animation-related parameters. The duration of the animation is dynamically configured by the phaseDuration variable (the phaseDuration variable is a parameter in the script used to set the animation duration, and its value can be changed in real time according to the training difficulty and patient response). The training difficulty can be divided into different levels such as elementary, intermediate, and advanced, and different levels correspond to different animation duration ranges. When the patient performs well in the primary training phase, the system can appropriately shorten the animation duration and increase the training difficulty.
[0124] Tracking Phase: During the tracking phase, the DynamicMovementSystem component (a functional module responsible for implementing object motion in the system, developed based on a physics engine and capable of simulating real-world physical motion) is enabled to achieve random object motion based on the physics engine. Simultaneously, the NavMeshAgent component (a path planning tool that helps objects avoid obstacles and plan reasonable movement paths in the virtual scene) is used for path planning, enabling intelligent collision avoidance. For example, if a target object or distractor approaches a virtual wall or other obstacle in the scene during movement, the NavMeshAgent component automatically adjusts its direction of motion to ensure realistic and smooth object movement. The number and frequency of distractors are dynamically adjusted based on the participant's real-time performance level. The system monitors data such as the number of target hits and average reaction time in real time. When a patient has a high hit rate and a short reaction time, the number and frequency of distractors are appropriately increased; otherwise, the number and frequency of distractors are reduced, precisely matching the training difficulty to individual ability.
[0125] 2. High-precision laser interaction and judgment
[0126] Hardware: The system's laser interaction module consists of a virtual reality controller (such as the Pico 4 Pro's VR controller) and an interactive surface. The controller has a built-in laser emitter and touch point detection sensors. The laser emitter emits laser beams for interaction, while the touch point detection sensors detect touch points generated when the laser beams come into contact with the interactive surface.
[0127] Interaction determination process: When the user operates the handle, the laser beam contacts the interactive surface to generate a touch point. The system calls to continuously obtain the first distance between the coordinates of the VR handle laser contact point and the preset target coordinates, and compares this distance with the preset first threshold (such as 10 pixels) through a script. This comparison process is implemented through conditional judgment statements. If the first distance is less than or equal to the first threshold, the user operation is judged to be correct and the target is successfully hit; otherwise, it is judged as an error, possibly due to accidentally hitting a distraction or not hitting the target at all.
[0128] 3. Two-person collaborative interaction and feedback
[0129] Network Synchronization: The system uses real-time network communication technologies, such as the WebSocket protocol (a protocol for full-duplex communication over a single TCP connection, enabling real-time data transmission between the client and server), to ensure synchronized interaction between two users in the virtual scene. Once a stable connection is established between the client and server, the user's operational data (such as controller movement data, target click operations, etc.) is transmitted to the server in real time over the network. The server then forwards this data to the other user's client. To minimize the impact of network latency on the interactive experience, the system employs prediction and compensation algorithms. For example, when a user operates a controller to move a virtual character, the client predicts the character's next position based on the controller's movement trends, displays it locally, and simultaneously sends the operational data to the server. After receiving this data, the server updates the other user's client. If network latency occurs, the system compensates for the character's position based on previous motion data, ensuring that both users see the same virtual scene.
[0130] Virtual character control: Simulating human perspective, users can move freely in virtual training scenes. The system creates a custom perspective control script, in which parameters such as the rotation object (setting the virtual character's head or body as the rotation object), rotation speed (setting the appropriate rotation speed according to actual needs, such as 30 degrees per second), and movement speed (such as 2 meters per second) are precisely set. Users obtain the motion data of the handle through the accelerometer of the handle, which can monitor the acceleration changes of the handle in all directions in real time. The system converts the acceleration data of the handle into movement instructions for the virtual character, thereby controlling the direction and speed of the virtual character in the scene. For example, when the user tilts the handle forward, the accelerometer detects the corresponding acceleration change, and the system moves the virtual character forward according to the preset conversion rules. Users can also interact with scripted interactive objects in the scene in a rich manner. These interactive objects have pre-programmed interaction logic, such as clicking on a specific object to trigger a corresponding event or task.
[0131] Virtual character triggering behavior: In the virtual scene, users can not only see each other's virtual images, but also get real-time feedback on the results of the interaction. For two-person collaborative tasks, the system establishes a special feedback mechanism. When two users hit the target at the same time within a preset time difference (1 second), the system enhances the user's operation experience through particle effects (such as generating colorful particle effects at the target hit location) and sound feedback mechanism (playing a cheerful success prompt sound), while updating the score and HUD (Head-Up Display) interface information, and displaying the user's score and task progress in real time on the HUD interface. If the hit is missed or mishit, the system records the error type (such as accidentally hitting a distraction, not hitting the target, etc.) and triggers a corresponding error prompt (displaying eye-catching error prompt text on the HUD interface and playing a warning sound effect at the same time). The target will be regenerated after a short delay (2 seconds) to ensure the continuity and challenge of the training.
[0132] 4. Behavioral data tracking and visualization
[0133] Data Tracking: During training, the system uses data tracking to track key metrics such as the number of target hits, average reaction time, hit rate over time, and the number of mishits in real time. The data tracking script runs continuously in the background. Every time a user performs an action (such as clicking a target or moving the controller), the script records the relevant data and calculates metrics such as average reaction time and hit rate based on a pre-set algorithm.
[0134] Table 2
[0135]
[0136]
[0137] Data Visualization: A HUD interface is constructed using Unity Canvas, and the TextMeshPro component is used to display data in real time. The TextMeshPro component displays data in the HUD interface in clear, aesthetically pleasing text format, making it easy for users to view at any time. ShaderGraph (a tool for creating and editing shaders) is also used to generate RenderTexture visualization charts based on hit coordinates, such as heat maps. When generating heat maps, ShaderGraph uses the color depth and density of the hit coordinates to visually display the distribution of user behavior data. Darker colors and higher density indicate areas where users hit their targets more frequently, while darker colors indicate areas with fewer hits. After training, the system calls the DataVisualizer.GenerateLineChart() method to generate a detailed hit rate time curve report. This method organizes and analyzes the hit rate data recorded during training to plot a curve showing the hit rate over time, providing comprehensive and scientific data support for subsequent effect evaluation and solution adjustments.
[0138] Examples
[0139] System construction and environment configuration
[0140] This system is developed based on the Unity 3D game engine and runs on a 64-bit Windows 10 operating system. The hardware requirements include an Intel Core i7-12700K processor, 16GB or more of RAM, and an NVIDIA GeForce RTX 3060 or higher graphics card. The Pico 4 Pro is used as the VR interaction device, with its 6DoF (six degrees of freedom) tracking technology accurately capturing user movements. The server uses an Alibaba Cloud ECS compute instance with four cores and 8GB of RAM. Network communication uses the WebSocket protocol for low-latency data transmission.
[0141] Implementation of multi-objective dynamic generation and control modules
[0142] Generation phase: Create a SceneGenerator script in the Unity project and implement the spatial distribution of the target and distractors based on the Poisson disk distribution algorithm. Create a configuration file through ScriptableObject to customize the probability of generating positive targets (such as 70% in the center area and 30% in the edge area) and the probability of generating negative distractors (and vice versa). The appearance of the target object is loaded through the SpriteRenderer component with preset image resources, and the color is adjusted using the Color32 structure. For example, the base color of the positive target is set to new Color32(255,165,0,255) (orange), and the base color of the negative distractor is set to new Color32(100,100,100,255) (gray).
[0143] Cue Phase: The CuePhaseController script inherits from MonoBehaviour and implements the target's flashing animation via the InvokeRepeating method. In the Update function, the flashing frequency is dynamically adjusted based on the phaseDuration variable (initialized to 3 seconds) and the current training difficulty level. For example, at the beginner difficulty level, the color transparency switches every 0.5 seconds, while at the intermediate difficulty level, this is shortened to 0.3 seconds, and at the advanced difficulty level, it's 0.1 seconds.
[0144] Tracking phase: The DynamicMovementSystem component is based on the Unity physics engine. It assigns physical properties to the target and distractors through the Rigidbody component and uses the AddForce method to achieve random movement. The SetDestination method of the NavMeshAgent component is used for path planning, and intelligent obstacle avoidance is achieved by detecting the ObstacleAvoidanceType parameter (set to QualitySettings.defaultAvoidance). The dynamic adjustment logic of distractors is implemented through the PerformanceEvaluator script. A performance evaluation is performed every 10 target interactions. If the hit rate is ≥80% and the reaction time is <1.5 seconds, the number of distractors is increased by 2, and the frequency of occurrence is increased by 20%; otherwise, the number of distractors is reduced by 1, and the frequency is reduced by 10%.
[0145] High-precision laser interaction and judgment module
[0146] Hardware Adaptation: The Pico 4Pro controller is connected to the Unity project via the SteamVR plugin. The controller's laser emitter corresponds to the TriggerPressed event of the SteamVR_Controller.Device class, and touch point detection relies on the SteamVR_TrackedObject component to obtain controller position and orientation information.
[0147] Interaction determination: In the InteractionManager script, use the Physics.Raycast method to launch a laser ray, with the ray originating at the controller position and directed in the direction of the controller. Use the Vector3.Distance method to calculate the distance between the laser contact point and the target center point, and compare it with a preset threshold (10 pixels, approximately 0.02 meters in Unity units). The hit determination logic uses a state machine design, defining four states: Idle, Aiming, Hit, and Miss. State transitions and feedback are implemented using switch-case statements.
[0148] Two-person collaborative interaction and feedback module
[0149] Network synchronization: Build a communication framework based on the WebSocket-sharp library. Create the WebSocketClient class on the client and the WebSocketServer class on the server. User operation data is encapsulated in JSON format. For example, controller movement data is encapsulated as {"type":"move","x":0.5,"y":0.3,"z":0.2}. The prediction algorithm uses a Kalman filter to predict the next position based on the controller movement data from the previous three frames. The compensation algorithm uses timestamp comparison. If network latency exceeds 50ms, position compensation is triggered to adjust the avatar to the correct position.
[0150] Avatar Control: The CameraController script implements view control, using the Input.GetAxis method to obtain joystick input to control the avatar's rotation (sensitivity set to 30° / second) and movement (speed set to 2 meters / second). Accelerometer data is processed using a low-pass filter algorithm to remove high-frequency noise and ensure smooth transitions between motion commands.
[0151] Interactive Feedback: Collaborative task feedback uses the ParticleSystem component to implement particle effects, playing preset colorful particle effects upon successful hits. Audio feedback uses the AudioSource component to load the success.wav and error.wav audio files. The HUD interface is updated through the UGUI system, using a Text component to display the score and a Slider component to show task progress, refreshing the data every 0.1 seconds.
[0152] Behavioral data tracking and visualization module
[0153] Data Tracking: DataTracker script through Dictionary<string,object> The data structure stores key data and records operation information in event callback functions such as OnClick and OnMove. For example, when a target hit event is triggered, {"event":"hit","time":Time.time,"targetId":"target_01"} is stored in the dictionary. Average reaction time is calculated using a sliding window algorithm with a window size of 10 operations, and the average is updated in real time.
[0154] The specific basis for selecting the sliding window size:
[0155] Matching system paradigm: Traditional attention training paradigms often use 10 trials as the basic statistical unit. The multi-target tracking module of this system also uses 10 rounds of interaction as a training cycle. Therefore, the selection of the window is best synchronized with the training cycle.
[0156] Computational efficiency: The average computation of 10 operations is small, which is suitable for resource constraints of real-time rendering scenes (each frame update takes <1ms in Unity).
[0157] Response speed: If the window is too large, the indicator will lag (for example, if the window is 100 operations, it will take 2-3 minutes to accumulate data). If it is too small, it will be greatly affected by accidental errors (for example, if the window is 3 operations, it will be easily disturbed by a single false click).
[0158] Data Visualization: The HUD interface uses the TextMeshPro plugin to create the TextMeshProUGUI component, which enables real-time display by binding data fields. Heatmap generation uses a custom shader written with ShaderGraph to convert hit coordinates to UV coordinates and map color intensity (0-255) based on the number of hits. After training, the DataVisualizer class uses the System.Drawing library to generate a hit rate time curve report, with training time (minutes) plotted on the horizontal axis and hit rate percentage on the vertical axis. The report is exported in PNG format.
[0159] In a specific example, the detailed working process is as follows:
[0160] Multi-objective dynamic generation and control
[0161] a. Generation Phase: Based on the Poisson disk distribution algorithm (to avoid overlapping targets) and the ScriptableObject configuration file, positive targets (such as orange animal images) and negative distractors (such as gray warning patterns) are generated according to regional probability. Visual attributes such as color and shape are used to enhance the difference between positive and negative stimuli.
[0162] b. Cue phase: Use the CuePhaseController script to control the target object's flashing animation. Dynamically change the stimulation duration by adjusting the phaseDuration variable (linked to the training difficulty) to guide the user to quickly locate the target.
[0163] c. Tracking phase: The physics engine (Rigidbody component) is used to achieve random movement of the target object, and the NavMeshAgent component is used to avoid obstacles. The number and frequency of distractions are dynamically adjusted based on the user's real-time performance (hit rate, reaction time), forming an adaptive "ability-difficulty" match.
[0164] 2. High-precision laser interaction and judgment
[0165] a. Hardware interaction: A VR controller (such as Pico 4 Pro) emits laser rays, obtains the controller's position and orientation through the SteamVR plug-in, and uses the Physics.Raycast method to detect collisions between the ray and the target object.
[0166] b. Logic judgment: Calculate the distance between the laser contact point and the target coordinates, compare it with the preset threshold (10 pixels), and trigger the hit / miss state machine (Idle→Aiming→Hit / Miss) to achieve precise operational feedback.
[0167] 3. Two-person collaborative interaction
[0168] a. Network Synchronization: Controller motion data (JSON format) is transmitted over the WebSocket protocol. A Kalman filter is used to predict the positions of local virtual characters and objects, and timestamps are used to compensate for network latency to ensure consistency in the two-player scene. The WebSocket protocol establishes a real-time data transmission channel between the client and server, transmitting information such as object motion data with timestamps. The Kalman filter predicts the current object position based on previous data. When WebSocket transmission latency causes data to arrive late, the client first renders using the predicted position to avoid image lag. After receiving the real data, the delay Δt is calculated based on the timestamp. If Δt exceeds a threshold (50ms), a timestamp compensation algorithm is used to correct the position based on the uniform motion assumption using the predicted speed and delay time. Linear interpolation is then used to smoothly transition the predicted position to the real position to correct the error. In this process, WebSocket provides the transmission foundation, the Kalman filter predicts motion parameters, and the timestamp compensation algorithm combines the real data with the correction. These three elements work together to minimize the impact of network latency on two-player VR interaction, ensuring virtual scene synchronization and smooth interaction. (1. WebSocket protocol provides real-time data channel → 2. Kalman filter predicts object position and velocity → 3. Timestamp compensation uses real data and predicted parameter calibration → Low-latency two-person synchronization)
[0169] b. Interactive feedback: When two players hit the target simultaneously, particle effects and sound effects are triggered, and the HUD interface score is updated. If they miss, the error type is recorded, the target generation is reset, and the motivation for collaboration is strengthened.
[0170] 4. Data tracking and visualization
[0171] a. Data collection: Use the DataTracker script to record operational events (such as hit time and target ID) in real time, use the sliding window algorithm to calculate the average reaction time, and build a dynamic behavior database.
[0172] b. Visual presentation: The HUD interface displays indicators such as the number of hits and accuracy in real time; ShaderGraph generates a heat map to intuitively display attention distribution; after training, a hit rate time curve is generated to provide a quantitative basis for efficacy evaluation.
[0173] In summary, the multi-target dynamic control mechanism in the present invention includes dynamic generation logic for positive targets and negative distractors (such as a random generation algorithm based on probability distribution), a clue prompt mechanism (such as flashing animation or sound effect guidance) and real-time difficulty adjustment (such as increasing or decreasing the number of distractors according to the hit rate). In terms of real-time interaction and data transmission, network communication technology is used to achieve synchronization of operations of two-person (or multi-person) virtual scenes (such as real-time transmission of handle motion data) and feedback linkage (such as collaborative hits to trigger special effects or score sharing) to ensure the real-time and consistency of interaction. A behavioral data closed-loop system tracks key behavioral data such as the number of target hits, reaction time, hit rate, etc., and realizes data feedback through a visual interface (such as HUD real-time display or curve report generation after training) to support treatment effect evaluation and program adjustment.
[0174] Alternative embodiment:
[0175] Multi-target generation algorithm expansion: Replace target generation with Gaussian distribution or fractal algorithms to simulate more complex social scenarios (such as random movement in dense crowds). Introduce reinforcement learning algorithms to automatically optimize target generation rules based on user historical data (such as adjusting color / shape to specific patient preferences).
[0176] Interaction upgrades: Adding gesture recognition (such as Vive gesture tracking) or eye control (such as Tobii eye trackers) to replace laser controller interaction and adapt to different hardware configurations. Integrating haptic feedback (such as vibration controllers) provides tactile stimulation when hitting the target, enhancing immersion.
[0177] Network architecture optimization: Adopt a distributed server architecture (such as AWS GlobalAccelerator) to support more users to collaborate online simultaneously (extended to multi-person group therapy scenarios).
[0178] Hardware and Interaction Methods: VR devices can use VR headsets and controllers, or be adapted for AR devices. Interaction Hardware Type: The laser controller can be replaced with a gesture recognition sensor, eye tracking device, or brain-computer interface, as long as the functional requirement of "user interaction with virtual objects" is met. Server and Network Protocol: The WebSocket protocol can be replaced with other real-time communication protocols; servers can be migrated from cloud platforms to local deployments, requiring only low latency and reliability of data transmission.
[0179] Elements for constructing virtual scenes:
[0180] Scene themes and visual styles: Social anxiety intervention scenes (such as parties and meetings) can be replaced with other themes (such as natural landscapes and science fiction spaces) according to training objectives without affecting the core logic of multi-target interaction.
[0181] Appearance attributes of target objects and distractors: Positive targets can be designed as icons, text, or dynamic models (such as animals, geometric shapes), and negative distractors can be adjusted in color, shape, or sound effects (such as warning sounds, noises). Only the distinction between "positive and negative stimuli" needs to be retained.
[0182] Physics engine and motion algorithm: The Unity physics engine can be replaced with Unreal Engine or other engines; the NavMeshAgent path planning can be replaced with other algorithms or preset obstacle avoidance strategies, only needing to achieve the effect of "random movement and obstacle avoidance" of the target object.
[0183] In summary, the present invention solves the problems of poor scene adaptability, low patient compliance, and lack of dynamic regulation and comprehensive assessment of attention bias in traditional social anxiety treatment methods, by constructing a two-person VR social anxiety intervention system based on a multi-target tracking paradigm, and utilizing technologies such as multi-target dynamic generation and regulation, and high-precision laser interaction and judgment.
[0184] This system design, based on a multi-target tracking paradigm, includes key technical modules such as dynamic multi-target generation and control, high-precision laser interaction and judgment, two-person collaborative interaction and feedback, and behavioral data tracking and visualization. This dynamic control mechanism targets the attentional bias of social anxiety patients, guiding them to change their attention patterns through flexible setting and adjustment of positive targets and negative distractions. A comprehensive behavioral data assessment system encompasses multiple key data indicators and enables real-time data tracking and visualization.
[0185] Compared to traditional social anxiety treatments, this invention utilizes virtual reality technology to create a highly flexible and precise VR training system that can be customized to suit individual patient needs, improving adaptability. Furthermore, the gamified interactive mechanism increases the fun of treatment and effectively improves patient compliance.
[0186] Compared to existing VR psychological intervention systems, this invention dynamically regulates user attention bias, enabling targeted intervention for social anxiety patients who are overly preoccupied with negative social information. Furthermore, through a comprehensive and scientific evaluation system, it more accurately reflects the patient's treatment outcome, providing strong support for adjusting treatment plans and thus enhancing the effectiveness and scientific nature of treatment.
[0187] The core advantages of the present invention include:
[0188] 1. Scenario adaptability and personalization:
[0189] The dynamic generation mechanism supports flexible configuration of target attributes, distribution, and motion trajectories, enabling customized training scenarios tailored to individual patients' fears (e.g., crowd density, social distance), breaking through the limitations of traditional VR fixed scenes. Compared to existing systems, which only offer 3-5 static scene templates (e.g., fixed party scenes) and lack the ability to dynamically adjust stimulation parameters, this new approach mitigates the scenario requirements of exposure therapy through virtual reality attention training, adapting to individual differences.
[0190] 2. Accuracy of attention intervention:
[0191] Through a dynamic game of "positive goal guidance + negative distraction," users are forced to actively shift their attention. Combined with real-time difficulty adjustment (such as increasing or decreasing the number of distractions), this technology achieves an upgrade from "passive exposure" to "active regulation." Compared to existing technologies: early eye tracking only monitored attention distribution without intervention; fixed prompts (such as flashing arrows) lacked dynamic adaptability; this invention forms a closed-loop control loop for attention training through multi-target interaction.
[0192] 3. Improved user compliance:
[0193] The two-player collaborative mode introduces elements of competition and cooperation (such as reward feedback for synchronized hits), combined with gamification (particle effects, sound effects, and a scoring system), transforming the tedious exposure therapy into an immersive interactive experience. Compared to existing technologies: Traditional VR therapy relies primarily on repetitive single-player exercises, resulting in high patient dropout rates; this new approach significantly enhances participation through social incentives.
[0194] 4. Scientific nature of the evaluation system:
[0195] By integrating operational data (number of hits, reaction time), spatial distribution data (heat map) and time series data (hit rate curve), a multi-dimensional evaluation model is constructed to replace traditional subjective scales and simple behavioral statistics.
[0196] Application scenarios of the present invention include:
[0197] Expanding treatment targets: Designing specialized scenarios (e.g., classroom presentations, business negotiations) for specific populations (e.g., adolescents, professionals). Expanding treatment to other psychological disorders, such as dynamic exposure therapy for post-traumatic stress disorder (PTSD) (distracting trauma-related attention through multi-target interaction).
[0198] Training Mode: Added "Adversarial Mode": Two users control a positive target and a negative distractor, respectively, strengthening attention training through competition. Introducing NPC characters (like virtual social partners) to engage in natural language conversations with users, combined with multi-target tracking tasks, to simulate real-life social pressure scenarios.
[0199] Deepening of data dimensions:
[0200] Integrated physiological indicator monitoring: Connect a heart rate sensor (such as Empatica E4) or a galvanic skin response (GSR) device to monitor physiological indicator levels in real time and correlate and analyze them with behavioral data.
[0201] Augmented reality (AR) integration: superimpose virtual targets onto real environments (such as offices and shopping malls) to achieve "mixed reality exposure therapy" and improve the generalization of training.
[0202] This paper has developed a virtual reality training system based on the attention regulation mechanism. Although this paper is mainly aimed at the field of social anxiety intervention, its core technology mechanism (multi-target attention regulation in virtual reality) has broad applicability. The following are some of the more extensive application scenarios:
[0203] 1. Education
[0204] a. Specialized attention training: Designed for children with ADHD or learning disabilities, multi-target tracking games (such as chasing colorful targets and avoiding obstacles) improve concentration and reaction speed through dynamic difficulty adjustment.
[0205] b. Collaborative Learning Platform: Build puzzle-solving or experimental scenarios (such as jointly assembling virtual mechanical parts) based on two-player VR interaction, and cultivate teamwork and problem-solving skills through task division and real-time communication.
[0206] 2. Vocational training
[0207] a. High-voltage scenario simulation:
[0208] Aviation: Pilots' anti-interference training, superimposing virtual events such as engine failure alarms and sudden weather changes during multi-target tracking missions, to train their ability to make quick decisions and allocate resources.
[0209] Medical field: Simulating emergency room scenarios, medical staff need to simultaneously track patients' vital signs (positive goals) and interference information (such as equipment alarms) to improve emergency response efficiency.
[0210] 3. Neuroscience and rehabilitation
[0211] a. Basic research tools: Combined with brain imaging technologies (such as fMRI and EEG), this tool analyzes the neural activity patterns of brain attention allocation during multi-target tracking tasks, providing data support for cognitive science research.
[0212] b. Assisted rehabilitation treatment program: Design customized training tasks (such as tracking virtual objects of a specific color) for patients with brain injury and stroke, evaluate the progress of attention recovery through data visualization, and assist in formulating personalized rehabilitation plans.
[0213] 4. Entertainment and gaming industry
[0214] a. Social Competitive Games: Develop collaborative, multi-objective competitive games for two or more players (e.g., collecting resources together, fighting enemy distractions), leveraging real-time interaction and feedback mechanisms to enhance player immersion and social interaction.
[0215] b. Immersive escape experience: Combining dynamic scene generation with two-player collaborative puzzle solving (such as synchronously operating mechanisms and avoiding virtual obstacles), the game enhances repeatability and challenge through random target changes and plot linkage.
[0216] 5. Industry and Manufacturing
[0217] a. Skills training system: Assembly line workers use VR to simulate complex work scenarios (such as simultaneously monitoring the status of multiple devices and sorting parts) and utilize multi-target tracking training to improve their multi-tasking capabilities and operational accuracy.
[0218] b. Remote collaborative maintenance: Engineers collaborate with remote colleagues to repair equipment in a virtual environment, improving maintenance efficiency and collaboration fluency through dynamic marking (positive goals) and fault prompts (negative distractions).
[0219] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0220] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0221] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.
[0222] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc or a read-only optical disc (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0223] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0224] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0226] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc. Various media that can store program codes.
[0227] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0228] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0229] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0230] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0231] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.
Claims
1. A two-person virtual reality social anxiety intervention system based on a multi-target tracking paradigm, characterized by: include: A multi-target dynamic generation and control module dynamically generates positive targets in virtual scenes to guide attention and establish positive emotional associations, as well as negative distractors to simulate social interference and train attention inhibition. Distractor parameters, such as the number and frequency of distractors, are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism. High-precision laser interaction and judgment module, which is an optional module, determines the hit or mishit status based on the coordinate distance between the laser beam and the target object, providing accurate interactive feedback; The two-player collaborative interaction and feedback module uses a real-time network communication protocol to synchronize two players in a virtual scene. It also triggers multi-sensory feedback mechanisms such as particle effects, sound effects, and score-linked feedback through collaborative tasks to enhance user participation motivation. The behavioral data tracking and visualization module collects user operation data sets in real time, such as the number of target hits, reaction time, and number of mishits, and generates quantitative evaluation reports through a multi-dimensional visualization interface. For example, multi-dimensional data visualization is performed through heat maps, hit rate curves, and HUD interfaces to build a closed-loop evaluation system.
2. The system according to claim 1, wherein The multi-objective dynamic generation and control module includes: A random algorithm generation unit, based on the Poisson disk distribution algorithm and configurable parameter objects, generates differentiated targets and distractors according to regional probability; The dynamic difficulty adaptation unit dynamically adjusts the number and frequency of distractions based on the user's hit rate and reaction time through a dynamic performance evaluator to achieve "ability-difficulty" matching.
3. The system according to claim 1, wherein: The high-precision laser interaction and determination module includes: A laser emission and contact point detection unit, which emits laser rays and detects the coordinates of the contact points by using a ray collision detection method; The state machine judgment unit triggers the hit or mis-hit state machine by comparing the distance threshold and provides feedback on the correctness of the operation.
4. The system according to claim 1, wherein: The two-person collaborative interaction and feedback module includes: The network synchronization unit, based on the WebSocket protocol and Kalman filter prediction algorithm, realizes the real-time transmission and delay compensation of the handle motion data and the virtual character position; The collaborative task processing unit triggers the particle effects and sound effects of two-player simultaneous hits through the multimodal feedback component, and updates the HUD interface score; Preferably, in the two-player collaborative interaction and feedback module: the virtual character control processes the handle accelerometer data through a low-pass filtering algorithm, smoothly converts it into motion instructions, and controls the character movement and perspective rotation.
5. The system according to any one of claims 1 to 4, characterized in that The behavior data tracking and visualization module includes: Sliding window calculation unit, which uses sliding window algorithm to calculate the average response time in real time; The spatial attention mapping unit maps hit coordinates to color intensity through the shader editing tool to generate an attention distribution heat map.
6. The system according to any one of claims 1 to 4, characterized in that In the multi-objective dynamic generation and control module: Positive targets use bright colors and animal images, while negative distractors use dark colors and warning signs, strengthening attention guidance through differences in visual attributes.
7. The system according to any one of claims 1 to 4, characterized in that The network synchronization in the two-person collaborative interaction step further includes a timestamp compensation processing mechanism, which performs the following operations in sequence: (a) Delay calculation: Calculate the network transmission delay based on the timestamp generated by the data packet at the sending end and the local timestamp when it is received by the receiving end; (b) Predicted position correction: When the delay exceeds a preset threshold, the predicted position at the next moment is calculated based on the predicted position and predicted speed at the current moment, combined with the delay, to correct the target object's position deviation caused by network delay; (c) Trajectory smoothing calibration: Within the sending timestamp interval of two adjacent frames of data, the actual position of the sender at the two timestamps is linearly interpolated according to the current rendering time of the receiver to generate a smooth transition trajectory to eliminate prediction errors and ensure visual coherence.
8. The system according to any one of claims 1 to 4, characterized in that The network synchronization in the two-person collaborative interaction step further includes a collaborative processing mechanism that performs the following operations in sequence: (a) Real-time data transmission and timestamp: Operational data packets are transmitted in real time through a full-duplex communication protocol, and a timestamp generated by the sender is appended to each data packet; (b) Motion parameter prediction rendering: The current position and speed of virtual objects are predicted based on historical motion data. When real-time data has not arrived, the predicted position is used for rendering to avoid image stagnation. (c) Collaborative correction of delay errors: After receiving the real data with timestamps, the following collaborative operations are performed: Calculate the network delay and, if it exceeds the threshold, correct the target position based on the predicted speed; Within the timestamp interval of adjacent data packets, a smooth motion trajectory is generated by position interpolation to eliminate the visual jump between the predicted position and the actual position.
9. The system according to any one of claims 1 to 8, characterized in that Optionally, the system is developed based on the Unity 3D engine, integrating a physics engine and path planning tools to achieve random movement of the target object and intelligent obstacle avoidance; Optionally, in the behavior data tracking and visualization module: after the training is completed, the time series visualization unit is called to generate a hit rate time curve report, where the horizontal axis is the training time and the vertical axis is the hit rate percentage.
10. A two-person virtual reality social anxiety intervention method based on the system according to any one of claims 1 to 9, characterized in that: include: Dynamic generation and regulation of multiple targets: Based on a random algorithm, positive targets are dynamically generated in virtual scenes to guide attention and establish positive emotional associations, as well as negative distractors to simulate social interference information and train attention inhibition. Distractor parameters are adjusted based on the user's real-time performance, forming an adaptive attention guidance mechanism. Laser interaction determination: This is an optional step that triggers hit or miss status feedback by comparing the distance between the laser beam contact point and the target object coordinates; Two-player collaborative interaction and feedback: Based on real-time network communication protocols, two players can synchronize virtual scenes and trigger multi-sensory feedback mechanisms through collaborative tasks to enhance user participation motivation. Closed-loop evaluation of behavioral data: Collect user operation data sets in real time, generate quantitative evaluation reports through a multi-dimensional visualization interface, and build a closed-loop evaluation system.