Method for enhancing multi-user coexistence sense and behavior perception in VR collaborative environment

By constructing an AI-powered intelligent vision adaptation system and a multimodal feedback mechanism, the problem of imbalance between global and local perception in VR collaboration was solved, achieving an efficient multi-user collaborative experience and immersive perception.

CN121785469APending Publication Date: 2026-04-03HEBEI GEO UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing VR collaboration technologies suffer from imbalances in global and local perception and a single perception dimension in large-scale scenes, leading to decreased collaboration efficiency and weakened immersion. They also lack AI-driven dynamic field-of-view adaptation and multimodal feedback mechanisms.

Method used

We will build an AI-powered intelligent vision adaptation system to identify scenes and user intentions in real time, automatically switch between global and local vision, and establish a collaborative feedback system of vision, audio, and touch. Through high-precision data collection and synchronization, we will achieve multi-dimensional information collaborative transmission.

Benefits of technology

It significantly enhances the sense of multi-user coexistence and behavior awareness, improves the smoothness of collaboration and immersion, and reduces the delay and perception error of field-of-view switching operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785469A_ABST
    Figure CN121785469A_ABST
Patent Text Reader

Abstract

The invention provides a method for enhancing multi-user coexistence sense and behavior perception in a VR cooperation environment, and relates to the technical field of VR cooperation. The method for enhancing the multi-user coexistence feeling and behavior perception in the VR cooperative environment comprises the following steps of S1, hardware model selection, S2, a data acquisition layer, S3, a data synchronization layer, S4, an AI intelligent WIM optimization layer, S5, a WIM core interaction layer, S6, a multi-mode feedback layer and S7, an output layer. Aiming at the defects of global and local perception imbalance and single perception dimension in large-scene cooperation of the existing VR cooperation technology, an AI intelligent view adaptation system is constructed, scenes, tasks and interaction intentions are recognized in real time, a global overhead view and a local close-up view are automatically switched, the contradiction between global and local perception is solved, and the real-time interaction of the scene, the task and the interaction intention is realized. Through the synergistic effect of the AI dynamic view and the multi-modal feedback, the single dependence of visual feedback is broken, the information of the collaboration party is accurately transmitted, the coexistence sense and behavior perception ability of multiple users are remarkably enhanced, and the collaboration fluency and immersion are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of VR collaboration technology, specifically a method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaboration environment. Background Technology

[0002] In VR collaborative environments, a series of technical means have been developed to enhance the sense of coexistence and behavioral awareness among multiple users.

[0003] Virtual avatars are a very common approach, creating digital representations of users in virtual space to facilitate interaction. This primarily relies on the following technologies: first, skeletal animation-based avatar representation technology, which captures basic head and hand movements using headsets, controllers, etc., and maps them to corresponding parts of the virtual avatar; second, basic spatial positioning technology, utilizing SLAM (Simultaneous Localization and Mapping) or cell tower triangulation to determine the user's position in virtual space, enabling synchronized spatial movement of the avatar; and third, simple interactive feedback technology, conveying interactive behavior through visual highlights (such as the glowing effect when the avatar's hand touches an object) or single sound effects (such as a "click" sound). TechViz's multi-channel software, with its virtual avatars like VictoR and AmbeR, uses photogrammetry to capture real models and create highly realistic virtual avatars. These avatars feature complete skinning and digital skeletons bound to geometric models, driven by 52 joints, resulting in extremely natural and fluid movements that help users recognize each other and collaborate in virtual environments. The paper "Techviz: Advantages and Importance of Virtual Avatars in Human-Computer Collaboration" elaborates on the application of such virtual avatars in simulating operator behavior, conducting accessibility tests, collision detection, and ergonomic assessments. By allowing users to control virtual avatars, collaborative scenarios such as the adaptability of multiple operators in confined spaces can be verified in virtual environments.

[0004] Sharing cues is another important technology, enhancing mutual behavioral awareness by providing users with specific information about other collaborators. Some studies, such as the paper "The Effects of Sharing Awareness Cues in Collaborative Mixed Reality," have explored the impact of different combinations of virtual perceptual cues on mixed reality collaboration. User studies compared different combinations of three cues: field of view (FOV) frustum, gaze ray, and head gaze ray. The results showed that these perceptual cues significantly improved user performance, usability, and subjective preferences. This means that with such shared cues, users can better understand the attention direction and actions of their collaborators, thereby improving collaboration efficiency and immersion.

[0005] Furthermore, position tracking and motion capture technologies play a crucial role in VR multi-user collaboration. Devices track the user's position and movements in real-world space in real time and synchronously map them onto their virtual avatar in virtual space, achieving synchronized movements. Common devices include inertial motion capture suits and finger motion capture gloves, which can accurately collect information such as the user's torso posture, limb movement trajectories, and finger joint angles, with sampling frequencies up to 120Hz, ensuring no detail of the movement is missed.

[0006] Some solutions also enhance the sense of coexistence among multiple users by generating "external perspectives." For example, a "panoramic small window for collaborators" can be set at the edge of the user's head-mounted display interface to display the virtual environment from the perspective of the collaborators in real time; or a "third-party overhead view" can be used to generate a dynamic overhead view in the virtual space, marking the position coordinates and action trajectories of all collaborators (such as highlighting the hand operation area), so that users can intuitively observe the overall collaborative layout.

[0007] The existing technology has the following main problems:

[0008] (1) Imbalance between global and local perception in large-scale collaborative scenarios, lack of AI-driven dynamic vision adaptation mechanism.

[0009] The main issue is the contradiction between global location awareness and the acquisition of local interactive details in large-scale scenarios, coupled with a lack of AI-driven dynamic field-of-view adaptation capabilities. In large-scale VR collaborative scenarios (such as virtual factory equipment maintenance, cross-regional virtual engineering collaboration, and multi-person guided tours of large virtual exhibition halls), existing technologies cannot simultaneously meet the core requirements of "global location awareness" and "acquisition of local interactive details," leading to decreased collaboration efficiency and a fragmented interactive experience. On the one hand, when the VR scene space is large, users can only obtain a "rough coordinate position" of themselves and their collaborators through basic spatial positioning technologies (SLAM or base station positioning), and cannot intuitively grasp the distribution of all collaborators in the global scene. For example, in a virtual factory maintenance scenario, users find it difficult to quickly determine "whether a collaborator is in the equipment area on another floor" or "whether they are close to the core components that need to be operated together," requiring them to frequently switch their own perspective or ask about the location of collaborators, increasing communication costs. On the other hand, when users need to understand the local interaction details of the collaborator (such as the disassembly steps of a collaborator on a precision device or the parameter adjustment operation of a virtual part), existing external perspective solutions (such as a fixed panoramic window of the collaborator or a static third-party overhead view) cannot accurately focus on key interaction areas: if a panoramic window is used, the window can only display a broad view around the collaborator in a large scene, and cannot zoom in to show details such as "the angle of the finger turning the virtual screw" or "the process of adjusting the parameters of the device panel"; if a third-party overhead view is used, although the global position can be seen, the viewing height is fixed and cannot be automatically zoomed in to the local operation area of ​​the collaborator. Users need to manually adjust the viewing zoom ratio, which is cumbersome and easy to miss key interaction nodes.

[0010] More importantly, existing technologies lack AI-powered intelligent judgment and dynamic view switching mechanisms, failing to automatically adapt the "global scene view" and "local detail view" according to the needs of the collaborative scenario and the user's interaction intent. For example, in virtual engineering collaboration, when a user is in the "planning the overall collaboration route" stage, they need to see the global location distribution of all collaborators in the large scene. Existing technologies require users to manually turn on the overhead view. However, when the user switches to the "assisting collaborators in resolving local equipment failures" stage, they need to focus on the collaborators' hand operation details. Existing technologies still require users to manually turn off the global view and open a close-up window of the collaborator's local area. The entire process relies on active user operation, which not only interrupts the collaboration process but may also cause the collaboration rhythm to break down due to operation delays. Furthermore, the existing external view switching logic is "fixed trigger" (such as the user pressing a specific button on the controller to switch), which cannot intelligently determine the switching time based on scene characteristics and interaction behavior. For example, when the system detects that "multiple collaborators are moving towards the same key device at the same time", it cannot automatically switch from the global view to the local detail view of that device; when it detects that "collaborators are scattered in different areas of a large scene and operating independently", it cannot automatically switch back to the global view to present the overall distribution, resulting in the view adaptation always lagging behind the actual collaboration needs.

[0011] The core reason for this shortcoming lies in the fact that the external field of view design of existing technologies follows the logic of "static function setting" and fails to build an intelligent association system of "scene-intent-field of view". First, it lacks a scene feature recognition module, making it impossible to determine the size of the VR scene (such as large / small scene) and the type of collaborative task (such as global planning / local operation) in real time. Second, it lacks the ability to analyze user interaction intent, making it impossible to infer whether the user needs "global position information" or "local detail information" based on the user's action trajectory (such as whether they are continuously paying attention to a collaborator or moving towards a specific device) and device data (such as the direction of gaze of the headset and the frequency of controller operation). Third, it lacks a dynamic field of view rendering and switching engine, making it impossible to adjust the field of view range, scaling ratio and presentation form in real time based on AI judgment results (such as using an overhead view for the global view and a follow-up close-up view for the local detail view). Ultimately, this leads to the inability to resolve the contradiction between global and local information acquisition in large scenes, limiting the smoothness and accuracy of multi-user collaboration.

[0012] (2) The lack of multimodal feedback leads to a blurred sense of spatial presence, and visually dependent interaction disrupts the sense of collaborative immersion.

[0013] In the existing VR collaboration technology system, the perception feedback mechanism presents the core problem of "single visual dominance and lack of multimodal collaboration". It relies too much on visual signals to transmit collaborative information and has not built an audio and tactile feedback system that is strongly related to spatial location and action intensity. As a result, users cannot accurately perceive the spatial location and interaction details of collaborators, which ultimately weakens the "real coexistence" and collaborative immersion among multiple users.

[0014] From the perspective of the limitations of audio feedback, existing solutions can only achieve "basic sound effect output without spatial attributes," and cannot convey the spatial location information of collaborators through auditory means. In current technology, the voice communication or interactive sound effects of collaborators (such as "click sounds" and "grab sounds" when operating virtual objects) mostly adopt a "mono stereo" playback mode, that is, regardless of whether the collaborator is to the left, right, in front, or behind the user in the virtual space, the voice / sound effect heard by the user has no directional difference; even if some solutions support basic "left and right channel differentiation," they cannot adjust the attenuation of sound effects according to the distance between the collaborator and the user—for example, when the collaborator speaks at 1 meter away from the user, the volume and clarity heard by the user are exactly the same as when the collaborator speaks at 10 meters away, and it is impossible to judge the distance of the collaborator by hearing. This "audio feedback without spatial attributes" means that in complex VR scenarios (such as multi-room virtual offices and large virtual workshops), users can only find the location of collaborators visually, and cannot "locate by sound" as in real-world scenarios. This not only increases the cost of spatial cognition, but also makes it easy to miss the interaction signals of collaborators due to the distraction of visual attention (such as when collaborators call out in the user's blind spot).

[0015] At the haptic feedback level, existing technologies exhibit a "fixed and unrelated" feedback pattern, failing to match the intensity of the collaborator's actions and resulting in a break in the transmission of interaction details. Current mainstream VR devices (such as controllers) provide only "single vibrations of preset intensity." For example, when a collaborator lightly touches a virtual button, presses a virtual switch forcefully, or lightly grasps a virtual part or grips a virtual tool, the user receives a uniform frequency and intensity vibration signal, making it impossible to distinguish the varying degrees of intensity in the collaborator's actions. This "undifferentiated haptic feedback" prevents the transmission of crucial action intentions during collaboration. For instance, in a virtual device assembly scenario, "lightly tightening a screw to calibrate the position" and "forcefully tightening a screw to fix the component" are two completely different operational intentions. Existing technologies can only judge this by visually observing the screw's rotation. If the user's gaze is not focused on that area, they cannot perceive the collaborator's operational progress and required force, easily leading to erroneous operations such as "the collaborator has already tightened the screw, but the user is still trying to assist with force," thus reducing collaboration efficiency.

[0016] Furthermore, existing technologies lack a coordinated feedback logic encompassing "visual-audio-tactile" feedback. The three sensory dimensions operate independently, further exacerbating the ambiguity of spatial presence. In real-world collaborative scenarios, humans form a complete spatial understanding of their collaborators by integrating multi-dimensional information such as "seeing the other party's location," "hearing the other party's voice direction," and "feeling the other party's tactile force." However, in current VR collaboration, visual feedback (such as avatar actions), audio feedback (such as voice without direction), and tactile feedback (such as fixed vibration) are unrelated. For example, the avatar shows the collaborator operating to the user's left, but the audio still plays from directly in front. Tactile feedback is not synchronized with the avatar's actions. This "perceptual information conflict" interferes with the user's spatial cognitive judgment, preventing them from accurately establishing the correspondence between "virtual avatar - collaborator's real actions - spatial location." Ultimately, this results in a disconnect where "the collaborator exists in the virtual space, but lacks a real sense of presence."

[0017] The root cause of this shortcoming lies in the fact that the feedback design logic of existing technologies focuses on "basic interactive notification" rather than "immersive sensory enhancement." On the one hand, at the hardware level, there is a lack of dedicated devices that support multimodal feedback, such as audio output modules that cannot achieve 3D spatial sound effects, and haptic feedback components (such as force feedback gloves and vibration vests) that cannot precisely control vibration intensity and frequency. On the other hand, at the software level, a "perceptual data association model" has not been established. The spatial location data of collaborators (such as coordinates and distance) is not linked with the orientation parameters of the audio system (such as channels and volume attenuation coefficients), nor is the motion intensity data of collaborators (such as grip strength and pressing pressure) bound with the feedback parameters of the haptic system (such as vibration frequency and intensity). As a result, multimodal feedback is always in a state of "independent and uncoordinated" operation, which fails to meet the user's need for precise perception of the spatial presence of collaborators. Summary of the Invention

[0018] To address the shortcomings of existing technologies, this invention provides a method for enhancing multi-user coexistence and behavioral perception in VR collaborative environments. Addressing the deficiencies of current VR collaborative technologies in large-scene collaboration, such as the imbalance between global and local perception and the single perception dimension, this invention constructs an AI intelligent field-of-view adaptation system. This system identifies scenes, tasks, and interaction intentions in real time, automatically switching between a global overview and a close-up view to resolve the contradiction between global and local perception. Simultaneously, a multimodal collaborative feedback system integrating vision, audio, and touch is established, linking spatial location, sound effect parameters, motion intensity, and feedback parameters. Combined with high-precision motion mapping, this enables the collaborative transmission of multi-dimensional perceptual information. Ultimately, through the synergistic effect of AI dynamic field of view and multimodal feedback, the reliance on single visual feedback is broken, accurately conveying information to collaborators, significantly enhancing multi-user coexistence and behavioral perception capabilities, and improving the smoothness and immersion of collaboration.

[0019] To achieve the above objectives, the present invention provides the following technical solution: a method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaborative environment, specifically comprising the following steps:

[0020] S1. Hardware Selection

[0021] The selection includes VR headsets, interactive controllers, audio devices, haptic enhancement devices, computing devices, and positioning systems to support the operation of the system and realize various functions;

[0022] S2. Data Acquisition Layer

[0023] S201. User location collection: Collect the user's coordinates and movement trajectory in the main scene;

[0024] S202. Motion capture, capturing head-mounted display gaze direction, handle posture, and glove pressure;

[0025] S203. Scene acquisition: Acquire the number of objects, occlusion status, and workspace location;

[0026] S204. Collaborative data collection: Collect the location and operation status of collaborating parties;

[0027] S3. Data Synchronization Layer

[0028] S301. Utilizes the Photon Unity Networking framework to optimize synchronization strategies for "multi-user collaboration" requirements;

[0029] S302. Core data synchronization;

[0030] S303. Auxiliary data synchronization;

[0031] S304. Conflict arbitration mechanism: When multiple users operate on the same object in WIM at the same time, the "timestamp priority" principle is adopted and an "operation lock" prompt is generated in WIM.

[0032] S305. Disconnection Recovery: If a user loses connection, their WIM avatar will be marked as gray and semi-transparent. After the connection is restored, the core data during the disconnection period will be resent to ensure uninterrupted collaboration.

[0033] S4. AI Intelligent WIM Optimization Layer

[0034] S401. Through the scene feature recognition module, the task complexity is automatically determined based on the number of objects and occlusion status in the scene, and the corresponding WIM initial rendering parameters are output.

[0035] S402. Through the user intent parsing module, visual, operational, and voice data are integrated, and user needs are determined through the "random forest + rule engine" model, and corresponding user need tags are output.

[0036] S403. Through the WIM intelligent adaptation module, based on scene characteristics and user intent, the scaling ratio, viewing angle and occlusion handling of WIM are dynamically adjusted to achieve an upgrade from static to dynamic intelligence and solve the problem of global-local perception imbalance in large scenes.

[0037] S5.WIM Core Interaction Layer

[0038] S501. Basic synchronous interaction;

[0039] S502. AI-enhanced interaction;

[0040] S503.WIM position adaptive function;

[0041] S6. Multimodal Feedback Layer

[0042] S601.3D spatial audio dynamically adjusts the direction and volume of audio based on the position of collaborators / target objects in WIM;

[0043] S602. Motion-associated haptic feedback, where haptic feedback intensity / frequency is bound to WIM operation type;

[0044] S603. Enhanced visual feedback: WIM and the main scene visual effects are synchronized to enhance interactive perception;

[0045] S7. Output Layer

[0046] S701. Headset output, main scene full-screen display, WIM overlaid in the form of a semi-transparent floating window;

[0047] S702. Audio output: 3D spatial sound effects are output through headphones, with the orientation strictly matching the location of the collaborating party / target in WIM;

[0048] S703. Haptic output, with handle / glove vibration synchronized with WIM operation and collaborator actions.

[0049] Preferably, in step S2, the sampling frequency for user position acquisition is 120Hz, the coordinates are based on Unity WorldSpace, the direction accuracy for motion acquisition gaze is ±0.5°, the controller posture sampling is 120Hz, the glove force detection accuracy is ±0.1N, and the scene acquisition classifies the scene into three categories: Simple, Medium, and Complex according to the task complexity. Simple has a few objects and no occlusion, Medium has a medium number of objects and partial occlusion, and Complex has multiple objects and full occlusion. The synchronization frequency for collaborative data acquisition is 120Hz, and the operation status is determined by the controller action.

[0050] Preferably, in step S302, the core data synchronization includes user location, WIM scaling parameters, and collaborator operation status.

[0051] Preferably, in S303, the auxiliary data synchronization includes gaze direction and glove force data.

[0052] Preferably, step S501 specifically includes the following steps: object operation synchronization: when the user drags and rotates a spherical target object in WIM, the corresponding object in the main scene moves synchronously.

[0053] Collaborator avatar synchronization: The position and actions of collaborators' avatars in WIM are synchronized 1:1 with the main scene, including head rotation, hand movement, and object placement.

[0054] Preferably, step S502 specifically includes the following steps: one-click view jump: when the user clicks on the collaborator's avatar in WIM, the main scene view automatically jumps to the vicinity of the collaborator;

[0055] Task progress indicators: AI analyzes the collaborator's operation status in real time.

[0056] Preferably, step S503 specifically includes the following steps: binding the non-dominant hand: the WIM moves with the non-dominant hand, a "widget-like" design to ensure availability anytime, anywhere;

[0057] Fixed floating window: WIM is fixed in the upper right corner of the main scene to avoid obstructing the core operations of the main scene.

[0058] This invention provides a method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaborative environment. It has the following beneficial effects:

[0059] 1. This invention provides a method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaborative environment. This method features an AI-driven WIM dynamic field-of-view adaptation mechanism. By constructing a closed-loop intelligent system of "scene feature recognition - user intent parsing - WIM parameter adaptation," it upgrades the external field of view of WIM from static to dynamic. The system can automatically categorize task complexity based on scene complexity (such as the number of objects and occlusion status) and output initial WIM rendering parameters. Simultaneously, by integrating visual, operational, and audio data, it analyzes user needs in real time and dynamically adjusts the WIM's scaling ratio, viewing angle, and occlusion handling rules. This allows for automatic switching between global and local fields of view in large-scale collaborative scenes, reducing operational interruptions and communication costs, and improving the smoothness of collaboration.

[0060] 2. This invention provides a method for enhancing the sense of coexistence and behavioral perception among multiple users in a VR collaborative environment. This method features a collaborative mechanism between Web Interaction Model (WIM) and multimodal feedback. By establishing a strong correlation model of "WIM interaction - multimodal feedback," spatial location and operation intensity data in WIM are transformed into 3D spatial audio, motion-related haptic feedback, and visual enhancement feedback. Through 3D spatial audio linkage, "sound positioning" is achieved; through motion-related haptic linkage, operational details are conveyed; and through visual enhancement linkage, the main scene's visual effects are triggered synchronously. This "visual-audio-haptic" three-dimensional perception system significantly enhances users' spatial presence and behavioral perception capabilities in VR collaboration, reducing misunderstandings of action intentions.

[0061] 3. This invention provides a method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaborative environment. This method incorporates WIM's AI-enhanced interactive features, adding three new AI-enhanced interactive functions to the existing WIM "synchronized main scene operation": "one-click perspective jump," "intelligent task progress prompts," and "adaptive display mode." The one-click perspective jump function allows users to automatically jump to the collaborator's perspective after clicking on the collaborator's avatar in WIM; the intelligent task progress prompt function uses AI to analyze the collaborator's operation status and synchronously displays the task progress in both WIM and the main scene; the adaptive display mode supports two modes: "binding the non-dominant hand" and "fixed floating window," allowing users to switch according to their needs and improving the efficiency and ease of use of WIM for collaborative assistance. Attached Figure Description

[0062] Figure 1 This is a schematic diagram of the technical solution of the present invention;

[0063] Figure 2 This is a schematic diagram of the core design concept of WIM in this invention;

[0064] Figure 3 This is a schematic diagram of the external view structure of the present invention;

[0065] Figure 4 This is a schematic diagram illustrating the present invention's method of locating and placing objects using an external viewpoint;

[0066] Figure 5 This is a flowchart illustrating an application example of the present invention. Detailed Implementation

[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] Example 1

[0069] like Figure 1 As shown in the figure, this embodiment of the invention provides a method for enhancing the sense of coexistence and behavior awareness among multiple users in a VR collaborative environment, specifically including the following steps:

[0070] S1. Hardware Selection

[0071] The selection includes VR headsets, interactive controllers, audio devices, haptic enhancement devices, computing devices, and positioning systems to support the system's operation and implement various functions. Specific hardware types and models are shown in the table below:

[0072] Hardware Category Model / Specification effect VR headset Meta Quest 2 The main scene display and WIM support gaze tracking (for AI intent parsing). Interactive controller Quest 2 controllers (6DoF tracking, vibration feedback) The vibration intensity of the WIM and main scene can be adjusted (linked to multimodal haptic feedback). audio equipment Bose QuietComfort Earbuds II (Supports 3D Spatial Audio) Output directional sound effects based on the location of the collaborating party in WIM. Tactile enhancement devices Manus Prime X Force Feedback Gloves (5-finger independent vibration, 0-10N force detection) The interaction with WIM provides vibration feedback corresponding to the force applied, enhancing the realism of the interaction. computing devices Desktop PC (i7-7700k, 16GB RAM, GTX 1080 Ti) Running an AI model using Unity + Photon Networking (intent parsing, WIM adaptation) Positioning system Quest 2 has built-in SLAM (static positioning ±1mm, dynamic positioning ±3mm). User location data is collected and synchronized to AI for collaborator location calibration in WIM.

[0073] S2. Data Acquisition Layer

[0074] Specifically, as the input source of the system, the data acquisition layer accurately acquires three-dimensional data of "scene-user-collaboration", which must simultaneously meet the task requirements (such as collaborative object finding) and the refined requirements of AI / multimodal.

[0075] S201. User location acquisition: Acquires the user's coordinates and movement trajectory in the main scene. The sampling frequency is 120Hz, and the coordinates are based on Unity World Space.

[0076] Specifically, by sampling at high frequency, the system can capture the user's position changes in the virtual scene in real time and accurately, providing accurate location information for subsequent interaction analysis and task execution, and ensuring that the system can dynamically adjust the task execution strategy and feedback content according to the user's position.

[0077] S202. Motion capture: captures head-mounted display gaze direction, handle posture, and glove force, with a direction accuracy of ±0.5°, handle posture sampling at 120Hz, and glove force detection accuracy of ±0.1N.

[0078] Specifically, high-precision motion capture can accurately capture the user's interaction intentions and operation details, determine the user's focus by the direction of gaze, and determine the user's operation force and gestures by the handle posture and glove force, thereby achieving a more natural and smooth human-computer interaction experience and providing rich motion data for AI systems to analyze and make decisions.

[0079] S203. Scene Acquisition: Acquire the number of objects, occlusion status, and workspace location. Based on task complexity, the scene is classified into three categories: Simple, Medium, and Complex. Simple has a small number of objects and no occlusion; Medium has a medium number of objects and partial occlusion; Complex has multiple objects and full occlusion.

[0080] Specifically, by collecting and classifying scenes in detail, the system can adjust the data collection strategy and focus according to the complexity of different tasks. In Simple scenes, the focus is on quickly collecting and processing information of a small number of objects, while in Complex scenes, more refined processing of the occlusion relationships and spatial distribution of objects is required to ensure the smooth execution of tasks and the accurate judgment of the AI ​​system.

[0081] S204. Collaborative data acquisition: Acquires the location and operation status of collaborating parties, with a synchronization frequency of 120Hz. The operation status is determined by the handheld action.

[0082] Specifically, the collection of collaborative data enables collaborative interaction among multiple users. By synchronizing the location and operational status of collaborators in real time, the system can promptly adjust task allocation and collaboration strategies to ensure the efficient completion of collaborative tasks. When one party is in a "searching" state, the system can provide guidance and prompts to other collaborators based on its location and operational status. When one party has "found" the target object, the system can promptly notify other collaborators and adjust the task objective, achieving seamless connection and efficient collaboration in the process.

[0083] S3. Data Synchronization Layer

[0084] Specifically, the main goal of the data synchronization layer is to ensure the consistency of multi-user WIM (Work in Motion) interactions, ensuring that each user can obtain the status and operation information of other users in real time and accurately in a multi-user collaborative environment, thereby achieving seamless collaborative interaction.

[0085] S301. Utilizes the Photon Unity Networking framework to optimize synchronization strategies for "multi-user collaboration" requirements;

[0086] Specifically, by adopting the PUN framework, data synchronization between multiple users can be handled efficiently, ensuring a stable interactive experience even in complex network environments. Optimizing synchronization strategies can effectively avoid data conflicts and latency, improving overall system performance and user experience.

[0087] S302. Core data synchronization, including user location, WIM zoom parameters, and collaborator operation status;

[0088] Specifically, this data is synchronized via the UDP protocol at a sampling frequency of 120Hz to ensure real-time performance (latency ≤20ms). The UDP protocol's low latency makes it suitable for data transmission with high real-time requirements. Real-time synchronization of user locations allows other users to promptly understand their companions' location changes, while synchronization of WIM scaling parameters ensures all users collaborate from the same perspective and operating scale. Synchronization of collaborators' operational status allows team members to stay informed about task progress and current operational status, thus achieving efficient collaborative work.

[0089] S303. Auxiliary data synchronization, including gaze direction and glove force data;

[0090] Specifically, this data is synchronized via the TCP protocol at a sampling frequency of 60Hz to ensure reliability (loss rate ≤0.1%) and avoid multimodal feedback errors. The TCP protocol is characterized by high reliability, making it suitable for scenarios with high data integrity requirements. Synchronizing gaze direction and glove force data provides users with richer interactive feedback. The gaze direction determines the user's focus, while the glove force data provides tactile feedback, thereby enhancing the user's immersion and interactive experience.

[0091] S304. Conflict arbitration mechanism: When multiple users operate on the same object in WIM at the same time, the "timestamp priority" principle is adopted to retain the operation within the latest 1ms and generate an "operation lock" prompt (red border) in WIM.

[0092] Specifically, this mechanism effectively resolves conflicts in multi-user interactions, ensuring that only one user's action is effective at a time, thus avoiding confusion and errors caused by conflicting operations. The red-bordered indicator visually informs other users that the object is currently locked, guiding them to perform other operations or wait, thereby improving the efficiency and stability of collaboration.

[0093] S305. Disconnection Recovery: If a user loses connection, their WIM avatar will be marked as gray and semi-transparent. After the connection is restored, the core data during the disconnection period will be resent to ensure uninterrupted collaboration.

[0094] Specifically, in a multi-user collaborative environment, network instability or device malfunction may cause users to lose connection. By marking the avatar of the disconnected user as gray and semi-transparent, other users can intuitively understand the user's status. When the user reconnects, the system automatically resends the core data from the period of disconnection, ensuring that the user can quickly recover to the state before the disconnection and continue to participate in collaborative tasks, thereby guaranteeing the continuity and integrity of collaboration.

[0095] S4. AI Intelligent WIM Optimization Layer

[0096] S401. Through the scene feature recognition module, the task complexity is automatically determined based on the number of objects and occlusion status in the scene, and the corresponding initial WIM rendering parameters are output. Specifically:

[0097] Input: The number of objects and their occlusion status obtained from the scene data acquisition layer.

[0098] Judgment logic: Automatically classify task complexity (Simple, Medium, Complex) based on preset thresholds.

[0099] Output: Initial WIM rendering parameters. For example, for the Complex scene, the initial scaling ratio is 1:20 to ensure the complete maze is displayed; for the Simple scene, the initial scaling ratio is 1:30 to reduce redundant field of view.

[0100] In practical applications, such as in a virtual factory scene, if there are many objects and significant occlusion (such as equipment, walls, etc.), it is classified as a Complex scene; if there are fewer objects and less occlusion, it is classified as a Simple scene. Based on different levels of complexity, corresponding initial WIM rendering parameters are output to ensure that WIM can adapt to the scene requirements in its initial state.

[0101] S402. Through the user intent parsing module, it integrates visual, operational, and voice data, uses a "random forest + rule engine" model to determine user needs, and outputs corresponding user need tags. Specifically:

[0102] Input data (three-dimensional fusion):

[0103] Visual intent: The duration of time the head-mounted display gazes at a certain area of ​​WIM (>2 seconds → local requirement, >3 seconds → global requirement).

[0104] Operation Intent: Handle Actions (Pinch WIM with two fingers → Global zoom, Click a point in WIM → Local focus).

[0105] Voice intent: Supports commands such as "view the whole picture" and "zoom in on user B" (recognition accuracy ≥95%, response time ≤500ms, using the lightweight speech recognition model Whisper Tiny).

[0106] Intent determination model: The model adopts a fusion model of "random forest + rule engine" with 1000+ user operation samples (labeled as "global needs", "local focus - collaborator" and "local focus - target"). The determination accuracy is ≥92% and the inference time is ≤10ms (without affecting real-time performance).

[0107] Output: User requirement tags (e.g., "Local Focus - Collaborator A", "Global Observation").

[0108] In practical applications, if a user stares at a certain area of ​​WIM for an extended period (more than 3 seconds), it indicates that the user needs global information; users can also clearly express their needs through controller operations (such as pinching with two fingers) or voice commands (such as "see the global view"). Through the "random forest + rule engine" model, this module can quickly and accurately output user need labels, providing a basis for subsequent WIM adaptation.

[0109] S403. Through the WIM intelligent adaptation module, based on scene characteristics and user intent, the scaling ratio, viewing angle and occlusion handling of WIM are dynamically adjusted to achieve an upgrade from static to dynamic intelligence and solve the problem of global-local perception imbalance in large scenes.

[0110] The dynamic scaling strategy is shown in the table below:

[0111] User needs Scaling range Example (Medium scenario) Global observation 1:30~1:50 Displaying the complete VR scene, collaborating avatars, and marking their positions and movement trajectories (dashed lines for the last 3 seconds). Local focus (collaborating party) 1:10~1:8 Zoom in on the collaborator's area to a 30×30cm field of view to display the collaborator's hand movements (such as fingers manipulating objects). Local focus (target) 1:5~1:3 Zoom in on the target object to a 10×10cm field of view to display detailed object markers (resolves occlusion issues).

[0112] To address the issue of "occlusion causing the target / collaborator to be invisible," AI uses a "depth sorting algorithm" to identify key objects (collaborators, target objects) that are occluded in WIM. It then uses "semi-transparent red + outline stroke" (50% brightness, 2px stroke width) to ensure that key objects can be clearly identified in WIM even in complex scenes (full occlusion).

[0113] In actual use, when users need to observe the entire scene, WIM automatically adjusts to a scaling ratio of 1:30 to 1:50 to display the complete VR scene and the positions and movement trajectories of collaborators. When users need to focus on a specific area, WIM automatically adjusts to a scaling ratio of 1:10 to 1:8 or 1:5 to 1:3 based on the target object (collaborator or target object), magnifying key areas and displaying details. Furthermore, to address occlusion issues, AI uses a depth sorting algorithm to identify occluded key objects and processes them with semi-transparent red and outline strokes, ensuring that users can clearly see key objects under any circumstances.

[0114] S5.WIM Core Interaction Layer

[0115] Specifically, the WIM core interaction layer serves as a "bridge between users and the system," embodying both WIM's core value and enhancing collaboration efficiency through AI. The external view plays a crucial role in computer-mediated collaboration, especially in Collaborative Virtual Environments (CVEs), where monitoring the actions and operations of collaborators is paramount. The external perspective provides users with target points for determining the location and actions of their partners, enhancing spatial awareness and social presence within CVEs. Furthermore, interacting through the external view helps users efficiently complete search and operational tasks remotely. Figure 2 The WIM section indicated in the image provides a viewpoint of the complete VR scene. Through WIM, users can view the entire VR scene and see the location of their companions, as well as any changes in the scene (including user movement and object changes). Furthermore, when an object is far from the user, the user can directly interact with the mirrored object in the WIM, and the object in the VR scene will change synchronously. This improves collaboration efficiency and the sense of collaborative coexistence.

[0116] S501. Basic Synchronous Interaction

[0117] Synchronized object manipulation: When a user drags and rotates a spherical target object in WIM, the corresponding object in the main scene moves synchronously with a positional deviation of ≤2mm and a delay of ≤20ms.

[0118] Collaborator avatar synchronization: The position and actions of collaborators' avatars in WIM are synchronized 1:1 with the main scene, including head rotation, hand movement, and object placement, ensuring that users can keep track of collaborators' dynamics in real time through WIM.

[0119] S502. AI-Enhanced Interaction

[0120] One-click view jump: When a user clicks on a collaborator's avatar in WIM, the main scene view automatically jumps to the vicinity of the collaborator, within 3 meters, with the view facing the collaborator, avoiding the time-consuming process of the user having to manually move to find the collaborator;

[0121] Task progress indicators: AI analyzes the collaborator's operation status in real time. For example, when a collaborator completes "place the target in the waiting area," the corresponding workspace in WIM automatically changes color. Figure 3 and Figure 4 As shown;

[0122] S503.WIM Position Adaptive Function

[0123] Binding to the non-dominant hand: WIM moves with the non-dominant hand, a "widget-like" design that ensures it is available anytime, anywhere;

[0124] Fixed floating window: WIM is fixed in the upper right corner of the main scene, occupying 15% of the field of view, with an adjustable transparency of 30%-70%, to avoid obstructing the core operations of the main scene;

[0125] S6. Multimodal Feedback Layer

[0126] Specifically, it is mainly used to compensate for the shortcomings of visual feedback alone. By linking WIM interaction with "visual-audio-tactile" sensations, it enhances the user's "spatial presence" of collaborators and fully responds to the need for "multimodal collaborative feedback".

[0127] S601.3D spatial audio dynamically adjusts the direction and volume of audio based on the position of collaborators / target objects in WIM;

[0128] S602. Motion-associated haptic feedback, where haptic feedback intensity / frequency is bound to WIM operation type;

[0129] S603. Enhanced visual feedback: WIM and the main scene visual effects are synchronized to enhance interactive perception;

[0130] The specific feedback types and their corresponding implementation logic and technical parameters are shown in the table below:

[0131] Feedback type Implementation logic (strongly correlated with WIM interaction) Technical specifications (to ensure a smooth experience) 3D spatial audio Based on the location of collaborators / target objects in WIM, dynamically adjust the audio's orientation and volume: - Orientation: Left side of WIM → Left channel volume +40%, Right side → Right channel +40%; - Distance: For every 1m increase in distance in WIM, the volume decreases by 5% (stabilizing at 50% after 10m). Sampling rate 48kHz, channel separation ≥80dB, audio latency ≤30ms (synchronized with WIM operation) Motion-related tactile sensation The intensity / frequency of haptic feedback is tied to the type of WIM operation to avoid the monotony of "fixed vibration" in the paper: - Clicking a WIM object: the handle vibrates once (intensity 20%, duration 100ms); - Dragging a WIM object: the vibration frequency increases with the dragging speed (10Hz→50Hz); - Collaborator grasping forcefully: the force feedback glove vibrates all fingers (intensity 80%, lasting until the grasping action ends). Vibration intensity adjustment range 0-100%, frequency range 10-100Hz, tactile delay ≤20ms (synchronized with WIM motion). Visual Augmented Feedback WIM's visual effects are synchronized with the main scene, enhancing interactive perception: - When a collaborator raises their hand in WIM: the avatar's hand in the main scene is highlighted in green (60% brightness); - When a target object is clicked in WIM: the target object in the main scene flashes (2Hz, green); - When a task is completed: WIM flashes green across the entire screen (1Hz, lasting 2 seconds) + main scene fireworks effect. The highlight color is configurable (default is green), the blinking frequency is 2Hz, and the transparency is 50% (it does not affect the observation of the main scene).

[0132] S7. Output Layer

[0133] S701. Headset output, main scene full-screen display, WIM overlaid in the form of a semi-transparent floating window, with transparency adjustable from 30% to 70%, supports "one-click hide / show" (e.g., "menu button + down button" on the controller), to avoid obstructing the core operation of the main scene;

[0134] S702. Audio output: 3D spatial sound effects are output through headphones. The orientation strictly matches the location of the collaborator / target in WIM. For example, if the collaborator in WIM is 3m to the left front, the user's left ear can clearly perceive the "sound source to the left front".

[0135] S703. Haptic output: The vibration of the handle / glove is synchronized with the operation of WIM and the actions of the collaborator. For example, when the collaborator presses a virtual button in WIM, the user's handle vibrates synchronously, with the intensity matching the pressure applied by the collaborator.

[0136] Example 2

[0137] The following section will present an application example in a specific usage environment, based on the content of this application:

[0138] Using the task of "two people collaboratively finding a target spherical object and placing it into the activated workspace" (Medium scene: 20 objects, 10 partially occluded) as an example, the complete workflow of the system is demonstrated, showcasing the value of AI and multimodal computing. The completion process is as follows: Figure 5 As shown.

[0139] 1. Scene initialization (0-5 seconds):

[0140] Data acquisition layer: Acquires the positions of 20 spherical objects (including 10 partially occluded ones) and 2 active workspace coordinates (fixed positions in the paper);

[0141] AI scene recognition: It is determined to be a Medium scene, and the WIM initial scaling is 1:30, with semi-transparent red highlights occluding objects;

[0142] Output layer: The head-mounted display shows the main scene (a 6×6m maze + a WIM floating window in the upper right corner (50% transparency), in which two user avatars (blue cubes) are located at random positions in the maze).

[0143] 2. Collaborative object finding (5-30 seconds):

[0144] User A stares at User B's avatar in WIM for more than 2 seconds (visual intent), while the controller hovers over the avatar (operational intent).

[0145] AI intent analysis: Determined to be "local focus on user B", output WIM adaptation command - zoom in on user B's area to 1:10, showing that user B is touching a partially occluded spherical object (the object is highlighted in red in WIM).

[0146] Multimodal feedback: 1) Audio: User B is on the left side of WIM → the left channel outputs "object touch sound" (volume 40% higher than the right channel); 2) Haptic: User A clicks on the object in WIM → the controller vibrates once (intensity 20%, duration 100ms); 3) Visual: The object in WIM flashes (2Hz), and the main scene synchronously generates a yellow arrow pointing to the object;

[0147] Synchronous interaction: When user B grabs the object in the main scene, user B's avatar hand in WIM synchronously displays the "grabbing action" and marks the object's surface with numbers (such as "5").

[0148] 3. Place synchronously (30-50 seconds):

[0149] User B places an object into the waiting area of ​​the Workspace. The AI ​​detects the operation status → the Workspace in WIM changes color, and a pop-up window in the main scene prompts "User B is ready".

[0150] User A observed the prompt through WIM and gave the voice command "View the whole picture" → AI switched WIM to a 1:30 scale to show the positions of both parties (both were near the active Workspace).

[0151] After communication between the two parties, the object is placed in the placement area simultaneously, triggering multimodal feedback: 1) Audio: Dual-channel output of "task completed sound effect" (volume 100%); 2) Haptic: The controller vibrates twice (intensity 50%), and the gloves vibrate once in all fingers; 3) Visual: WIM full-screen green flashing, and the main scene displays a "task completed" pop-up window.

[0152] Example 3

[0153] The core objective of this application is to solve the problems of "imbalance between global and local perception in large-scale VR collaborative scenes" and "lack of multimodal feedback". Regarding the specific content of this application, there are the following alternative solutions that can achieve the same inventive purpose. All alternative solutions retain the principle of "unchanged core functions and optimized technical path" to ensure that they do not deviate from the inventive objective.

[0154] 1. An alternative to the AI-driven WIM dynamic field-of-view adaptation mechanism

[0155] This WIM adaptation mechanism, based on a "rule engine + user presets" approach (without an AI model), does not rely on AI models such as "random forests." It achieves dynamic WIM adaptation through preset scene rules and user-defined thresholds, reducing reliance on computing resources and making it suitable for low-configuration hardware scenarios. This solution incurs no AI model training costs, offers faster response times (≤5ms), and its adaptation logic is transparent and debuggable. However, rules require manual maintenance and cannot adapt to complex scenarios without preset rules (such as "20 objects + dynamic occlusion"). Users must manually adjust thresholds, making it less flexible than AI-based solutions.

[0156] This external view solution (non-WIM carrier) uses "multi-view switching + gesture triggering" to replace WIM's dynamic scaling with "multiple preset views + gesture triggering." By pre-generating global / local views, users can switch between them using gestures, reducing real-time rendering pressure. This solution offers fast view loading speed (≤10ms), no real-time scaling calculations, and strong hardware compatibility. However, the fixed views cannot adapt to dynamic scene changes (e.g., the view needs to be reloaded after a collaborator moves), and gesture recognition carries the risk of false triggering (false touch rate approximately 5%).

[0157] 2. Alternative solutions for the collaborative linkage mechanism between WIM and multimodal feedback

[0158] This audio feedback solution, based on "main scene coordinates + pre-calculated sound effects library" (non-WIM linkage), does not rely on WIM's coordinate data. Instead, it directly uses the real-time coordinates of collaborators / targets in the main scene to call the pre-calculated sound effects library, achieving 3D spatial audio and simplifying the data link. This solution has a short data link (main scene → audio engine), reducing latency to ≤20ms and avoiding WIM coordinate synchronization errors. However, it requires pre-calculating a large number of sound effect parameters and cannot adapt to virtual operations in WIM (such as no audio feedback when dragging a target in WIM).

[0159] This haptic solution, based on "action threshold + fixed feedback template" (non-intensity-related), does not rely on WIM's operation intensity data. It triggers a fixed haptic template based on the "presence / absence / type" of actions in the main scene, simplifying the haptic mapping logic. It is suitable for basic controllers without force detection (such as the standard Quest2 controller). This solution eliminates the need for a force sensor, reducing hardware costs (eliminating the need for Manus gloves) and is simple to implement. However, it cannot transmit differences in action intensity (e.g., consistent feedback between a light grip and a firm grip), increasing the misunderstanding rate of collaborative intentions by approximately 15%.

[0160] 3. Alternative solutions for WIM's AI-enhanced interactive features

[0161] This perspective-jumping solution (non-AI enhanced) uses voice commands and fixed shortcut keys instead of AI-driven one-click jumps. Users actively trigger the jump, reducing the complexity of intent parsing. This solution has no AI intent parsing errors, achieves 100% trigger accuracy, and has a low learning curve. However, users need to memorize commands / shortcut keys, and the number of steps increases in multi-target scenarios (ID input is required).

[0162] This task feedback solution, based on a "progress bar + text prompt" (not WIM color-changing), does not rely on the color-changing prompts in the WIM Workspace. Instead, it conveys collaboration progress through a progress bar and text pop-ups in the main scene, reducing visual complexity. This solution provides intuitive prompts, avoids WIM visual obstruction, and allows users to focus their attention more effectively. However, text pop-ups may obscure main scene operations, and progress bar information may be redundant in multi-party collaboration scenarios.

[0163] All alternatives achieve the invention's objective of "enhancing the sense of multi-user coexistence and behavioral awareness." The applicable scenarios and core differences between the different solutions are shown in the table below:

[0164] Innovation Alternative Solution Types Applicable Scenarios Core advantages Core limitations Dynamic field of view adaptation Rule engine + user preset Low-configuration hardware, static scene No AI cost, fast response No dynamic adaptation capability Dynamic view adaptation Multi-view switching + gesture triggering Static scenes, limited hardware resources Fast loading and strong compatibility Fixed view, risk of accidental touch Multimodal linkage Main scene coordinates + pre-calculated sound effects library No WIM scenarios, pursuing low latency Short link, low latency No WIM virtual operation feedback Multimodal linkage Motion threshold + fixed tactile template Basic handle, non-high-precision collaboration Low hardware cost and simple logic No intensity difference feedback AI-enhanced interaction Voice / shortcut key perspective jump Multi-target scenarios, pursuing high accuracy No errors, low learning cost Instructions need to be memorized AI-enhanced interaction Progress bar + text prompt The main scene has intensive operations; avoid WIM occlusion. Intuitive and unobstructed Risk of text obscuring

[0165] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for enhancing the sense of coexistence and behavioral awareness among multiple users in a VR collaborative environment, characterized in that, Specifically, the following steps are included: S1. Hardware Selection The selection includes VR headsets, interactive controllers, audio devices, haptic enhancement devices, computing devices, and positioning systems to support the operation of the system and realize various functions; S2. Data Acquisition Layer S201. User location collection: Collect the user's coordinates and movement trajectory in the main scene; S202. Motion capture, capturing head-mounted display gaze direction, handle posture, and glove pressure; S203. Scene acquisition: Acquire the number of objects, occlusion status, and workspace location; S204. Collaborative data collection: Collect the location and operation status of collaborating parties; S3. Data Synchronization Layer S301. Utilizes the Photon Unity Networking framework to optimize synchronization strategies for "multi-user collaboration" requirements; S302. Core data synchronization; S303. Auxiliary data synchronization; S304. Conflict arbitration mechanism: When multiple users operate on the same object in WIM at the same time, the "timestamp priority" principle is adopted and an "operation lock" prompt is generated in WIM. S305. Disconnection Recovery: If a user loses connection, their WIM avatar will be marked as gray and semi-transparent. After the connection is restored, the core data during the disconnection period will be resent to ensure uninterrupted collaboration. S4. AI Intelligent WIM Optimization Layer S401. Through the scene feature recognition module, the task complexity is automatically determined based on the number of objects and occlusion status in the scene, and the corresponding WIM initial rendering parameters are output. S402. Through the user intent parsing module, visual, operational, and voice data are integrated, and user needs are determined through the "random forest + rule engine" model, and corresponding user need tags are output. S403. Through the WIM intelligent adaptation module, based on scene characteristics and user intent, the scaling ratio, viewing angle and occlusion handling of WIM are dynamically adjusted to achieve an upgrade from static to dynamic intelligence and solve the problem of global-local perception imbalance in large scenes. S5.WIM Core Interaction Layer S501. Basic synchronous interaction; S502. AI-enhanced interaction; S503.WIM position adaptive function; S6. Multimodal Feedback Layer S601.3D spatial audio dynamically adjusts the direction and volume of audio based on the position of collaborators / target objects in WIM; S602. Motion-associated haptic feedback, where haptic feedback intensity / frequency is bound to WIM operation type; S603. Enhanced visual feedback: WIM and the main scene visual effects are synchronized to enhance interactive perception; S7. Output Layer S701. Headset output, main scene full-screen display, WIM overlaid in the form of a semi-transparent floating window; S702. Audio output: 3D spatial sound effects are output through headphones, with the orientation strictly matching the location of the collaborating party / target in WIM; S703. Haptic output, with handle / glove vibration synchronized with WIM operation and collaborator actions.

2. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: In S2, the sampling frequency for user position acquisition is 120Hz, the coordinates are based on Unity World Space, the direction accuracy for motion acquisition gaze is ±0.5°, the controller posture sampling is 120Hz, the glove force detection accuracy is ±0.1N, and the scene acquisition classifies the scene into three categories: Simple, Medium, and Complex according to the task complexity. Simple has a few objects and no occlusion, Medium has a medium number of objects and partial occlusion, and Complex has multiple objects and full occlusion. The synchronization frequency for collaborative data acquisition is 120Hz, and the operation status is determined by the controller action.

3. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: In S302, the core data synchronization includes user location, WIM scaling parameters, and collaborator operation status.

4. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: In S303, the auxiliary data synchronization includes gaze direction and glove force data.

5. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: S501 specifically includes the following steps: object operation synchronization: when the user drags and rotates a spherical target object in WIM, the corresponding object in the main scene moves synchronously. Collaborator avatar synchronization: The position and actions of collaborators' avatars in WIM are synchronized 1:1 with the main scene, including head rotation, hand movement, and object placement.

6. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: S502 specifically includes the following steps: one-click view jump: when the user clicks on the collaborator's avatar in WIM, the main scene view automatically jumps to the vicinity of the collaborator; Task progress indicators: AI analyzes the collaborator's operation status in real time.

7. The method for enhancing multi-user coexistence and behavioral awareness in a VR collaborative environment according to claim 1, characterized in that: The S503 specifically includes the following steps: binding the non-dominant hand: the WIM moves with the non-dominant hand, a "widget-like" design that ensures it is available anytime, anywhere; Fixed floating window: WIM is fixed in the upper right corner of the main scene to avoid obstructing the core operations of the main scene.

Citation Information

Patent Citations

  • Projection projection in virtual environment

    CN114174960A

  • VR teaching experience enhancement system and method

    CN120523333A

  • Collaborative Educational E-Learning Multi and Single Device, Supplemental Pedagogical Data Management UX / UI System Technology Platform Using Immersive Interactive Mixed Reality

    US20210027645A1