Virtual reality multi-person interaction method and system
Through the combination of distributed architecture and edge computing, the problem of insufficient scalability and real-time in virtual reality multi-person interaction is solved. AES-256-GCM encryption and biometric desensitization technology is adopted to achieve efficient user privacy protection and a stable multi-person interaction environment.
Patent Information
- Application Number
- CN202510408618.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There are problems in the existing virtual reality multi-person interaction methods and systems that have low scalability and real-time virtual reality interaction, as well as insufficient user data security.
The distributed architecture is combined with edge computing, and the environment is synchronized initialized through spatial anchor calibration and dynamic object loading, physical rules and interactive logic are configured, communication data is encrypted using AES-256-GCM and biometrics are converted into irreversible hash values. Combined with multimodal auditing and hierarchical response mechanisms, users' privacy data security and real-time behavior supervision are realized.
It significantly improves the scalability and real-time nature of virtual reality multi-person interaction, reduces latency issues, and effectively improves user privacy data security and the stability of the interactive environment.
Smart Images

Figure CN120337284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual reality interaction, and specifically to a method and system for multi-person interaction in virtual reality. Background Technique
[0002] Virtual reality systems are high-tech in the field of graphics and images that have emerged in recent years. Virtual reality technology encompasses computers, electronic information, and simulation technology. Its basic implementation method is for a computer to simulate a virtual environment to give people a sense of immersion in the environment. With the continuous development of social productivity and science and technology, the demand for VR technology in all walks of life is increasing day by day. At the same time, VR technology has also made great progress and has gradually become a new field of science and technology.
[0003] After retrieval, a method and system for multi-person interaction in virtual reality are disclosed in the invention patent with the Chinese patent publication number CN111782054A. The method for multi-person interaction in virtual reality in this invention patent forms a complete virtual reality technology by setting a central processing unit, a database, an operation center, wearable devices, a computer, and a three-dimensional scanning device, and at the same time can save and store past big data in real time to extend the data usage time;
[0004] However, in this method for multi-person interaction in virtual reality, only a single central processing unit is relied on to process data, which easily limits scalability and real-time performance. At the same time, only relying on a basic database for storage leads to the problem of low data security. Therefore, a method and system for multi-person interaction in virtual reality are proposed. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the deficiencies of the prior art, the present invention provides a method and system for multi-person interaction in virtual reality, which have the advantages of combining a distributed architecture with edge computing to improve the scalability and real-time performance of virtual reality interaction, and combining multiple encryption protection means to improve the security of user privacy data, and solves the problems that the existing method and system for multi-person interaction in virtual reality have low scalability and real-time performance of virtual reality interaction, and cannot effectively improve the security of user data in the above background technique.
[0007] (2) Technical Solutions
[0008] To achieve the purpose of combining a distributed architecture with edge computing to improve the scalability and real-time performance of virtual reality interaction, and combining multiple encryption protection means to improve the security of user privacy data, the present invention provides the following technical solutions: A method for multi-person interaction in virtual reality includes the following specific steps:
[0009] S1. Scene analysis definition: Analyze user requirements according to the type of interaction target to define the scene scale and plan the configuration of hardware devices;
[0010] S2. Interactive scene establishment: Based on spatial anchor calibration and dynamic object loading, achieve environmental synchronization initialization, and configure physical rules and interaction logics;
[0011] S3. Multi-person synchronous connection: Through connection establishment and session management, execute the data synchronization mechanism;
[0012] S4. Real-time interaction implementation: According to user input processing and motion capture, perform real-time motion synchronization and interaction logic processing;
[0013] S5. Experience optimization processing: Protect user data privacy and security, and supervise content security and user behavior.
[0014] Preferably, the type of interaction target includes industrial design, game competition, virtual meeting, and hybrid mode. Determining the user scale based on the type of interaction target, the steps for realizing scene analysis definition include: 1) Configure a unified input protocol and rendering parameters according to the type of terminal device; 2) Deploy a 5G private network to ensure that the bandwidth is greater than 100 Mbps and the latency is less than 50 ms, enable QoS priority division, and optimize the network environment; 3) Establish an identity identifier through biometrics and cross-platform accounts, scan facial features according to a depth camera, and preset a skeleton binding and motion capture template to generate a personalized virtual image.
[0015] Preferably, the steps for establishing the interactive scene in step S2 include:
[0016] 1) Establish a shared coordinate system: a. Based on the Lighthouse positioning system, the base station emits laser and infrared light, calculates the 6DoF pose of the built-in sensors of the device, and uses the headset SLAM algorithm to construct a spatial grid in real time to generate the origin of the shared coordinate system; b. Generate spatial anchors by scanning the environment and upload them to the server for initializing the primary user. By scanning the same physical markers or feature points, align the coordinate systems of secondary users based on the ICP algorithm to match point clouds;
[0017] 2) Data loading and synchronization: a. Pre-load static models of buildings and terrains, adopt the HLOD technology to dynamically switch the level of detail according to the distance, calculate the global illumination offline, synchronize the light map hash value in real time and verify the consistency, and update the environment in real time; b. Assign a unique UUID to each dynamic object, define the synchronization fields through Protobuf, and synchronize the objects within and adjacent to the user's field of view based on Octree spatial division to achieve segmented loading;
[0018] 3) Rule Configuration and Interaction Feedback: a. Parametrically configure the physics engine based on scene accuracy and scale parameters, and apply for an ownership token according to the Token Passing mechanism when the user grabs an object, which will be automatically released when it times out; b. When multiple people interact simultaneously, adopt a distance priority strategy. When multiple people apply forces to the same object simultaneously, the server synthesizes a vector force, expressed as: where the weight w i is dynamically calculated based on the relative position between the user and the object, and F i is the force application data for each person; c. Build a multi-modal feedback channel based on haptic and spatial audio feedback.
[0019] Preferably, the specific steps for multiple people to synchronously connect in step S3 include:
[0020] 1) Connection Establishment: a. When the user initiates a connection request, select the nearest edge server node for service matching according to the IP geographical location; b. Define room parameters, including interaction mode, maximum number of people, and privacy settings. The server reserves bandwidth and calculates resources to achieve resource pre-allocation; c. Encrypt the invitation and perform identity verification when crossing platforms;
[0021] 2) Data Synchronization: a. Collect the head pose, handle input, and object state through the client for each frame to complete data acquisition; b. Perform data compression and encapsulation processing, including coordinate compression and rotation compression;
[0022] 3) Consistency Assurance: a. Predict the future position based on speed and acceleration, expressed as:
[0023] where P current is the current position, and P predicted is the predicted future position;
[0024] b. Send the operation to the server through the client for simulation effects. The server verifies the legality of the operation. If it is invalid, it broadcasts a rollback instruction to restore the state of the client. Use timestamps to mark the operation and only roll back the operation for the affected entities;
[0025] 4) Network Optimization: a. Give priority to using the UDP protocol to transmit real-time data, cooperate with the RUDP protocol to achieve partial reliable transmission, add 20% redundant packets and allow lossless recovery at a 10% packet loss rate to achieve forward error correction. When the frame rate is insufficient, reuse the previous frame image to reduce motion blur to optimize the rendering layer; b. The client performs cubic spline interpolation processing on the received discrete state data to generate a smooth motion trajectory. Based on Kalman filtering processing, reduce the positioning jitter through the state estimation model, expressed as:
[0026] where K k is the Kalman gain, which dynamically adjusts the weights of the predicted value and the measured value;
[0027] c. Set the server to save the full-scene snapshot every 5s, load the most recently saved snapshot after disconnection and reconnection, and supplement the missing status through the operation log, and preferentially restore the high-priority objects within the user's field of view.
[0028] Preferably, the steps for realizing real-time interaction in step S4 include:
[0029] 1) Input processing and motion capture: a. Based on the head-mounted device, obtain the 6DoF pose through the inertial measurement unit IMU, ensuring that the sampling rate ≥ 1000Hz; b. Based on the handle device, capture the finger bending degree through the capacitive sensor and transmit data to the force feedback trigger; c. Based on the full-body motion capture device, transmit 19 bone node data at a frequency of 30Hz, realize the data acquisition and fusion of multi-modal input devices, and perform standardization processing; d.
[0030] 2) Motion synchronization and interaction logic: a. Based on the AABB bounding box, quickly screen candidate objects. When sending a grab request to the server, attach the timestamp and operation hash value through precise Mesh collision detection. If multiple users request simultaneously, allocate them in timestamp order;
[0031] b. According to the Verlet integration algorithm, locally simulate and predict the throwing trajectory, expressed as:
[0032] x n+1 = 2x n - x n-1 + a n Δt 2 , after receiving the predicted trajectory, recalculate and verify through the PhysX engine. If the deviation is greater than 10cm, force correction;
[0033] c. Based on the operation conversion algorithm, realize multi-person collaborative editing, including: when user A executes operation O A locally, generate an operation log. Before the server broadcasts O A , first apply the conflict resolution rule. After user B receives O A , adjust the local operation order through the conversion function T(O B , O A );
[0034] 3) Voice interaction: a. Calculate the sound source direction and attenuation through the Steam Audio SDK device and based on the head-related transfer function HRTF; b. Control the facial muscles through 52 Blend Shapes of mixed deformations in the FACS system, and extract the ARKit data stream and MLP neural network predicted mouth shape parameters as data sources.
[0035] 4) Performance optimization feedback: Render the central viewing area at native resolution, and use 1 / 4 resolution for the peripheral viewing area, combined with temporal anti-aliasing to smooth the edges.
[0036] Preferably, the steps for experience optimization in step S5 include:
[0037] 1) Privacy and security protection: a. Use the AES-256-GCM algorithm to encrypt all end-to-end communication data; b. Use the double-ratchet protocol to dynamically update the session key to prevent the leakage of historical data; c. Convert biometric data such as iris / fingerprint into irreversible hash values for biometric desensitization; d. Separate the user's real identity from the virtual character ID through an intermediate mapping layer to ensure that behavioral data cannot be traced back.
[0038] 2) Content and behavior supervision: Build a real-time multi-modal audit system and establish a hierarchical response mechanism, including: a. Automatically mute the violator and record the log after triggering a first-level violation warning; b. Force the virtual character to be transparent and limit the interaction ability for 30 minutes after triggering a second-level violation warning; c. Immediately kick the user out of the session and freeze the account for review after triggering a third-level violation warning.
[0039] 3) Comfort optimization: a. When fast movement is detected, automatically reduce the peripheral viewing area by 20%, and add a dashboard in the moving scene to provide a visual anchor for dynamic viewing control;
[0040] b. Use the IMU data of the headset to predict the poses of the next 2 frames for motion prediction compensation, and pre-render the compensated images to reduce the dizziness caused by motion blur. The formula is expressed as:
[0041] where θ current is the current motion angle, θ predicted is the predicted future angle, ω is the angular velocity, and α is the angular acceleration;
[0042] 4) Visual protection: a. Dynamically adjust the color temperature according to the usage duration, reducing it from 6500K to 4000K to filter blue light; b. Automatically match the virtual and real light intensities through an ambient light sensor; c. The system enforces a rest mechanism according to the usage duration.
[0043] A virtual reality multi-person interaction system includes a terminal management module for providing an audiovisual experience and capturing limb movements and facial expressions to interact with the virtual environment;
[0044] A network architecture module for adopting a distributed architecture for the central server to achieve spatio-temporal consistency of the state and event synchronization mechanism;
[0045] A software function module for generating and simulating a dynamic environment according to an environment engine, and providing user authentication management and social functions;
[0046] A technology optimization module for providing network optimization strategies to achieve AI enhancement and pipeline rendering optimization.
[0047] Preferably, the terminal management module includes a head-mounted display device, an input device, and an auxiliary sensing device, and configures a network infrastructure for the terminal management module. The network architecture module includes a synchronization engine and a distributed service architecture layer. The software function module includes a user management system, an environment engine, and a social interaction unit. The technology optimization module includes a network optimization unit, an AI enhancement unit, and a rendering pipeline optimization unit.
[0048] (III) Beneficial effects
[0049] Compared with the prior art, the present invention provides a method and system for virtual reality multi-person interaction, which has the following beneficial effects:
[0050] 1. For the method and system for virtual reality multi-person interaction, by adopting a distributed server architecture and dynamically allocating the nearest edge node according to the user's geographical location, dynamic load balancing can be achieved, QoS priority division can be enabled, and UDP+RUDP protocols are used to achieve low-latency transmission, which can significantly reduce latency problems.
[0051] 2. For the method and system for virtual reality multi-person interaction, by encrypting communication data using AES-256-GCM end-to-end, dynamically updating keys in combination with the double ratchet protocol, and converting biometric features into irreversible hash values, the security of user privacy data can be effectively improved, and real-time behavior supervision effects can be achieved in combination with multi-modal auditing and hierarchical response. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the modules of the virtual reality multi-person interaction system of the present invention;
[0053] Figure 2 It is a flowchart for implementing the virtual reality multi-person interaction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Please refer to Figure 1-2, a virtual reality multi-person interaction method, comprising the following specific steps:
[0056] S1. Scenario analysis and definition: Analyze user needs according to the interaction target type to define the scenario scale and plan and configure hardware equipment;
[0057] S2. Interaction scene establishment: based on spatial anchor point calibration and dynamic object loading, the environment is synchronized and initialized, and the configuration value physical rules and interaction logic are realized;
[0058] S3. Multi-person synchronous connection: Execute data synchronization mechanism by establishing connection and session management;
[0059] S4. Real-time interaction implementation: Based on user input processing and motion capture, real-time action synchronization and interaction logic processing are performed;
[0060] S5. Experience optimization processing: protect user data privacy and security, and supervise content security and user behavior.
[0061] Furthermore, the interaction target types include industrial design, game competition, virtual meeting and hybrid mode. The user scale is determined based on the interaction target type. The steps to implement the scenario analysis definition include: 1) configuring unified input protocols and rendering parameters according to the terminal device type; 2) deploying 5G private networks to ensure that the bandwidth is greater than 100Mbps and the delay is less than 50ms, enabling QoS priority division, and optimizing the network environment; 3) establishing identity through biometrics and cross-platform accounts, scanning facial features based on the depth camera, and presetting bone binding and motion capture templates to generate personalized virtual images.
[0062] Specifically, configuring unified input protocols and rendering parameters according to the terminal device type helps ensure the compatibility and consistency of different devices in interactive scenarios, reduces problems caused by device differences, and further improves the stability of user experience;
[0063] Deploy 5G private networks and make clear requirements for bandwidth, latency, etc., and enable QoS priority division to provide a stable and high-speed network environment for various types of interactive targets, meet the needs of industrial design, game competitions and other scenarios with high network requirements, ensure the smoothness of real-time interaction, and avoid data transmission jams and delays;
[0064] By establishing identity through biometrics and cross-platform accounts, using a depth camera to scan facial features and preset related templates to generate a personalized virtual image, the uniqueness and security of the user's identity can be enhanced and the personalized needs of different users can be met.
[0065] Furthermore, the step of establishing the interactive scene in step S2 includes:
[0066] 1) Establish a shared coordinate system: a. Based on the Lighthouse positioning system, the base station emits laser and infrared light to calculate the 6DoF pose of the built-in sensors of the device. The head-mounted display SLAM algorithm is used to construct a spatial grid in real time to generate the origin of the shared coordinate system; b. Spatial anchors are generated by scanning the environment and uploaded to the server for initialization of the primary user. By scanning the same physical markers or feature points, the secondary user coordinate system alignment is achieved based on the ICP algorithm to match the point cloud.
[0067] 2) Data loading and synchronization: a. Pre-load static models of buildings and terrain, and use the HLOD technology to dynamically switch the level of detail according to the distance. Calculate the global illumination offline, synchronize the light map hash value in real time and verify the consistency to update the environment in real time; b. Assign a unique UUID to each dynamic object, define the synchronization fields through Protobuf, and synchronize the objects within and adjacent to the user's field of view based on the Octree space division to achieve sharded loading.
[0068] 3) Rule configuration and interaction feedback: a. Parametrically configure the physics engine based on the scene accuracy and scale, and apply for the ownership token when the user grabs an object according to the Token Passing mechanism, and it will be automatically released when the time limit is exceeded; b. When multiple people interact simultaneously, adopt the distance priority strategy. When multiple people apply forces to the same object at the same time, the server synthesizes the vector force, which is expressed as: where the weight w i is dynamically calculated based on the relative position of the user and the object, and F i is the force application data of each person; c. Construct a multi-modal feedback channel based on tactile and spatial audio feedback.
[0069] Specifically, using the Lighthouse positioning system and the head-mounted display SLAM algorithm to generate the origin of the shared coordinate system, and at the same time achieving user coordinate system alignment through spatial anchors and the ICP algorithm, can accurately determine the spatial position and direction, provide a unified and accurate spatial reference for the interaction scene, and ensure the spatial consistency during multi-person interaction;
[0070] For static models, use the HLOD technology to dynamically switch the level of detail, combined with offline calculation of global illumination and real-time synchronization of light map hash values, which can ensure the quality and effect of model display. For dynamic objects, sharded loading is achieved by assigning UUIDs, defining synchronization fields, and Octree space division, and relevant objects can be synchronized in real time according to the user's field of view, improving the pertinence and real-time performance of data loading;
[0071] The Token Passing mechanism can effectively manage object ownership and avoid conflicts when multiple people operate the same object at the same time. The distance priority strategy and vector force synthesis algorithm can reasonably handle the force application during multi-person interaction, making the interaction more in line with real logic and enhancing the fairness and rationality of the interaction. It builds a multimodal feedback channel based on tactile and spatial audio feedback, which can provide users with a richer and more realistic interactive experience from multiple sensory dimensions, increase the immersion and fun of the interaction, and facilitate the acquisition of more comprehensive perceptual information.
[0072] Furthermore, the specific steps of multi-person synchronous connection in step S3 include:
[0073] 1) Connection establishment: a. When a user initiates a connection request, the nearest edge server node is selected for matching service based on the IP geographic location; b. Room parameters are defined, including interaction mode, maximum number of people, and privacy settings. The server reserves bandwidth and calculates resources to implement resource pre-allocation; c. Encrypt invitations and perform identity authentication when crossing platforms;
[0074] 2) Data synchronization: a. Collect head posture, handle input and object status in each frame through the client to complete data collection; b. Compress and encapsulate the data, including coordinate compression and rotation compression;
[0075] 3) Consistency guarantee: a. Predict the future position based on speed and acceleration, expressed as:
[0076] Where P current is the current position, P predicted For predicted future positions;
[0077] b. The client sends the operation to the server to simulate the effect. The server verifies the legitimacy of the operation. If it is invalid, it broadcasts a rollback instruction to restore the client state. The operation is marked with a timestamp and only the affected entity is rolled back.
[0078] 4) Network optimization: a. Use UDP protocol to transmit real-time data first, cooperate with RUDP protocol to achieve partial reliable transmission, add 20% redundant packets and allow lossless recovery at 10% packet loss rate to achieve forward error correction, reuse the previous frame when the frame rate is insufficient, reduce motion blur to optimize the rendering layer; b. The client performs cubic spline interpolation processing on the received discrete state data to generate a smooth motion trajectory, based on Kalman filtering processing, and reduces positioning jitter through the state estimation model, which is expressed as:
[0079] Where K k Kalman gain dynamically adjusts the weights of predicted values and measured values;
[0080] c. The server saves the full-scene snapshot every 5 seconds, loads the most recently saved snapshot after disconnection and reconnection, and supplements the missing status through the operation log, giving priority to restoring high-priority objects within the user's field of view.
[0081] Specifically, selecting the nearest edge server node according to the IP geographical location can effectively reduce network latency, improve connection speed and service quality, define room parameters in advance and perform resource pre-allocation to ensure the reasonable utilization of resources and stable operation during multi-person interaction, and enhance the security and reliability of the connection according to encrypted invitations and cross-platform authentication;
[0082] By predicting future positions, it is possible to respond to possible changes in advance and improve the fluency of interaction; the server's verification and rollback mechanism for operation legality, as well as the use of timestamps, ensure the consistency of all client states, avoid state chaos caused by operation errors, and guarantee the normal progress of multi-person interaction.
[0083] Furthermore, the steps for realizing real-time interaction in step S4 include:
[0084] 1) Input processing and motion capture: a. Based on the head-mounted device, obtain the 6DoF pose through the inertial measurement unit IMU, ensuring a sampling rate ≥ 1000Hz; b. Based on the handle device, capture the finger bending degree through the capacitive sensor and transmit data to the force feedback trigger; c. Based on the full-body motion capture device, transmit 19 skeletal node data at a frequency of 30Hz to realize the data acquisition and fusion of multi-modal input devices and perform standardized processing; d.
[0085] 2) Motion synchronization and interaction logic: a. Based on the AABB bounding box, quickly filter candidate objects, perform precise Mesh collision detection, and attach the timestamp and operation hash value when sending a grab request to the server. If multiple users request simultaneously, it will be allocated in timestamp order;
[0086] b. According to the Verlet integration algorithm, locally simulate and predict the throwing trajectory, expressed as:
[0087] x n+1 =2x n -x n-1 +a n Δt 2 , after receiving the predicted trajectory, recalculate and verify through the PhysX engine. If the deviation is greater than 10 cm, it will be forced to correct;
[0088] c. Based on the operation conversion algorithm, realize multi-person collaborative editing, including: when user A performs operation O A locally, generate an operation log. Before the server broadcasts O A , first apply the conflict resolution rule. User B receives O AAfter that, the local operation order is adjusted through the conversion function T(O B ,O A ).
[0089] 3) Voice interaction: a. Calculate the sound source direction and attenuation through the Steam Audio SDK device and based on the head-related transfer function HRTF; b. Control the facial muscles through 52 Blend Shapes of mixed deformations in the FACS system, and extract the ARKit data stream and the MLP neural network to predict the mouth shape parameters as the data source;
[0090] 4) Performance optimization feedback: The central visual field area is rendered at the native resolution, and the peripheral visual field area is rendered at 1 / 4 resolution and combined with temporal anti-aliasing to smooth the edges.
[0091] Specifically, in terms of action synchronization, the AABB bounding box is combined with precise Mesh collision detection to improve the efficiency and accuracy of collision detection. The processing of multiple people's simultaneous requests is allocated in timestamp order to ensure the fairness of interaction. At the same time, the combination of the Verlet integration algorithm and the PhysX engine can more accurately simulate and verify the throwing trajectory, enhancing the realism of interaction;
[0092] The operation conversion algorithm enables multi-person collaborative editing. Through conflict resolution rules and operation order adjustment, conflicts during multi-person operations are effectively avoided, improving the fluency of collaborative interaction. Differentiated rendering is performed according to the importance of the visual field area. The central visual field area uses the native resolution to ensure the picture quality of the key area, and the peripheral visual field area uses a lower resolution and is combined with temporal anti-aliasing processing. Without affecting the main visual experience, the rendering pressure of the system can be reduced, achieving a balance between picture quality and performance.
[0093] Furthermore, the steps of experience optimization processing in step S5 include:
[0094] 1) Privacy and security protection: a. Use the AES-256-GCM algorithm to encrypt all end-to-end communication data; b. Use the double-ratchet protocol to dynamically update the session key to avoid the leakage of historical data; c. Convert biometric data such as iris / fingerprint into irreversible hash values to desensitize biometric features; d. Separate the user's real identity from the virtual character ID through an intermediate mapping layer to ensure that behavior data cannot be traced back;
[0095] 2) Content and behavior supervision: Build a real-time multi-modal audit system and establish a hierarchical response mechanism, including: a. Automatically mute the violator and record the log after triggering a first-level violation warning; b. Force the virtual character to be transparent and limit the interaction ability for 30 minutes after triggering a second-level violation warning; c. Immediately kick the violator out of the session and freeze the account for review after triggering a third-level violation warning;
[0096] 3) Comfort optimization: a. When rapid movement is detected, automatically reduce the visual field edge by 20%, and add a dashboard in the moving scenario to provide a visual anchor for dynamic visual field control;
[0097] b. Use the head-mounted display IMU data to predict the future 2-frame postures for motion prediction compensation. By pre-rendering the compensated images, reduce the dizziness caused by motion blur. The formula is as follows:
[0098] where θ current is the current motion angle, θ predicted is the predicted future angle, ω is the angular velocity, and α is the angular acceleration;
[0099] 4) Visual protection: a. Dynamically adjust the color temperature according to the usage duration, reducing it from 6500K to 4000K for blue light filtering; b. Automatically match the virtual and real light intensities through the ambient light sensor; c. The system enforces a rest mechanism according to the usage duration.
[0100] Specifically, combine multiple encryption and protection means. Ensure the security of communication data through the AES-256-GCM algorithm, dynamically update the key according to the double ratchet protocol to prevent data leakage, and use biometric desensitization processing and the separation of identity and role IDs to comprehensively protect user privacy, prevent user information from being illegally obtained and tracked, and enhance users' trust in the system;
[0101] The real-time multi-modal audit system and the hierarchical response mechanism can ensure the effective handling of violations, help maintain a good interaction environment and order. Through measures such as dynamic visual field control and motion prediction compensation, and adjusted according to the user's motion state, it can effectively reduce the dizziness caused by rapid movement and motion blur, and improve the comfort and immersion of users when experiencing in the virtual environment.
[0102] In summary, for the method and system of virtual reality multi-person interaction, by adopting a distributed server architecture and dynamically allocating the nearest edge node according to the user's geographical location, it can achieve dynamic load balancing, enable QoS priority division, and use the UDP+RUDP protocol to achieve low-latency transmission, which can significantly reduce the latency problem. By encrypting the communication data with AES-256-GCM end-to-end, combining the double ratchet protocol to dynamically update the key, and converting the biometric features into irreversible hash values, it can effectively improve the security of user privacy data, and combine multi-modal audit and hierarchical response to achieve the effect of real-time behavior supervision.
[0103] All relevant modules involved in this system are hardware system modules or functional modules that combine computer software programs or protocols in the prior art with hardware. The computer software programs or protocols themselves involved in this functional module are all well-known technologies to those skilled in the art and are not the improvements of this system. The improvement of this system lies in the interaction relationship or connection relationship between modules, that is, the overall structure of the system is improved to solve the corresponding technical problems to be solved by this system.
[0104] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for virtual reality multi-person interaction, characterized in that, It includes the following specific steps: S1. Scenario analysis and definition: Analyze user requirements according to the type of interaction target to define the scenario scale, and plan and configure hardware devices; S2. Establishment of interaction scenario: Based on spatial anchor calibration and dynamic object loading, realize environmental synchronization initialization, and configure physical rules and interaction logic; S3. Multi-person synchronous connection: Through connection establishment and session management, execute the data synchronization mechanism; S4. Real-time interaction implementation: According to user input processing and action capture, perform real-time action synchronization and interaction logic processing; S5. Experience optimization processing: Protect user data privacy and security, and supervise content security and user behavior.
2. The method for virtual reality multi-person interaction according to claim 1, wherein: The types of interaction targets include industrial design, game competition, virtual meeting, and hybrid mode. The user scale is determined based on the type of interaction target. The steps to realize scenario analysis and definition include: 1) Configure a unified input protocol and rendering parameters according to the type of terminal device; 2) Deploy a 5G private network to ensure that the bandwidth is greater than 100 Mbps and the latency is less than 50 ms, enable QoS priority division, and optimize the network environment; 3) Establish an identity identifier through biometrics and cross-platform accounts, scan facial features according to the depth camera, and preset bone binding and action capture templates to generate a personalized virtual image.
3. A method for virtual reality multi-person interaction according to claim 1, characterized in that, The steps for establishing the interaction scenario in step S2 include: 1) Establish a shared coordinate system: a. Based on the Lighthouse positioning system, the base station emits laser and infrared light to calculate the 6DoF pose of the built-in sensors of the device, and uses the headset SLAM algorithm to construct a spatial grid in real time to generate the origin of the shared coordinate system; b. Generate spatial anchors by scanning the environment and upload them to the server for initializing the primary user. By scanning the same physical markers or feature points, the coordinate systems of secondary users are aligned based on the ICP algorithm to match the point cloud; 2) Data loading and synchronization: a. Pre-load static models of buildings and terrains, use the HLOD technology to dynamically switch the level of detail according to the distance, calculate the global illumination offline, synchronize the light map hash value in real time and verify the consistency, and update the environment in real time; b. Assign a unique UUID to each dynamic object, define the synchronization fields through Protobuf, and synchronize the objects within and adjacent to the user's field of view based on the Octree space division to achieve segmented loading; 3) Rule Configuration and Interaction Feedback: a. Parametrically configure the physics engine based on scene accuracy and scale parameters, and apply for an ownership token according to the TokenPassing mechanism when the user grabs an object, which will be automatically released when the timeout occurs; b. When multiple people interact simultaneously, adopt a distance priority strategy. When multiple people apply forces to the same object simultaneously, the server synthesizes a vector force, expressed as: where the weight w i is dynamically calculated from the relative position between the user and the object, and F i is the force application data for each person; c. Construct a multi-modal feedback channel based on tactile and spatial audio feedback.
4. A method for virtual reality multi-person interaction according to claim 1, characterized in that, The specific steps for multi-person synchronous connection in step S3 include: 1) Connection establishment: a. When the user initiates a connection request, select the nearest edge server node according to the IP geographical location to match the service; b. Define room parameters, including interaction mode, maximum number of people, and privacy settings, and the server reserves bandwidth and calculates resources for resource pre-allocation; c. Encrypt the invitation and perform identity verification when crossing platforms; 2) Data synchronization: a. Collect the head pose, handle input, and object state by the client for each frame to complete data acquisition; b. Compress and encapsulate the data, including coordinate compression and rotation compression; 3) Consistency guarantee: a. Predict the future position according to the speed and acceleration, expressed as: Where P current is the current position, and P predicted is the predicted future position; b. Send operations from the client to the server for simulation effects. The server verifies the legality of the operations. If invalid, it broadcasts a rollback instruction to restore the client's state. Use timestamps to mark operations and only roll back operations on affected entities; 4) Network optimization: a. Prioritize using the UDP protocol to transmit real-time data, and cooperate with the RUDP protocol to achieve partial reliable transmission. Add 20% redundant packets and allow lossless recovery at a 10% packet loss rate to implement forward error correction. When the frame rate is insufficient, reuse the previous frame to reduce motion blur and optimize the rendering layer; b. The client performs cubic spline interpolation on the received discrete state data to generate a smooth motion trajectory. Based on Kalman filtering, reduce positioning jitter through a state estimation model, expressed as: x k = F k x k-1 + B k u k + K k (z k - H k x k-1 ) where K k is the Kalman gain, dynamically adjusting the weights of the predicted value and the measured value; c. Set the server to save a full-scene snapshot every 5 seconds. After reconnecting after a disconnection, load the most recently saved snapshot and supplement the missing state through the operation log, and prioritize restoring high-priority objects within the user's field of view.
5. A method for virtual reality multi-person interaction according to claim 1, characterized in that, The steps for real-time interaction in step S4 include: 1) Input processing and action capture: a. Based on the head-mounted device, obtain the 6DoF pose through the inertial measurement unit (IMU), ensuring a sampling rate ≥ 1000Hz; b. Based on the handle device, capture the finger bend through a capacitive sensor and transmit data to the force feedback trigger; c. Based on the full-body motion capture device, transmit 19 bone node data at a frequency of 30Hz to achieve data collection and fusion of multi-modal input devices and perform normalization processing; d. 2) Action synchronization and interaction logic: a. Based on the AABB bounding box, quickly filter candidate objects. When sending a grab request to the server, attach a timestamp and an operation hash value through precise Mesh collision detection. If multiple users request simultaneously, allocate them in timestamp order; b. According to the Verlet integration algorithm, locally simulate and predict the throwing trajectory, expressed as: x n+1 = 2x n -x n-1 + a n Δt 2 , after receiving the predicted trajectory, recalculate and verify through the PhysX engine. If the deviation is greater than 10 cm, force correction; c. Implement multi - person collaborative editing based on the operation transformation algorithm, including: When user A executes operation O locally A an operation log is generated, and the server broadcasts O A Before that, conflict resolution rules are applied first. When user B receives O A after that, the local operation order is adjusted through the transformation function T(O B , O A ); 3) Voice interaction: a. Calculate the sound source direction and attenuation through the Steam Audio SDK device and based on the head-related transfer function (HRTF); b. Control facial muscles through 52 Blend Shapes of mixed deformations in the FACS system, and extract the ARKit data stream and MLP neural network predicted mouth shape parameters as data sources; 4) Performance optimization feedback: Render operations at the native resolution for the central viewing area, and use a 1 / 4 resolution for the peripheral viewing area, and cooperate with temporal anti-aliasing to smooth the edges.
6. A method for virtual reality multi-person interaction according to claim 1, characterized in that, The steps for experience optimization processing in step S5 include: 1) Privacy and security protection: a. Use the AES-256-GCM algorithm to encrypt all end-to-end communication data; b. Use the double-ratchet protocol to dynamically update the session key to avoid leakage of historical data; c. Convert biometric data such as iris / fingerprint into irreversible hash values for biometric desensitization; d. Separate the user's real identity from the virtual character ID through an intermediate mapping layer to ensure that behavioral data cannot be traced back; 2) Content and Behavior Supervision: Build a real-time multi-modal review system and establish a hierarchical response mechanism, including: a. Automatically mute violators and record logs after triggering a first-level violation warning; b. Force the virtual character to be transparent and limit the interaction ability for 30 minutes after triggering a second-level violation warning; c. Immediately kick the user out of the session and freeze the account for review after triggering a third-level violation warning. 3) Comfort Optimization: a. When rapid movement is detected, automatically reduce the visual field edge by 20%, and add a dashboard in the moving scene to provide a visual anchor for dynamic visual field control. b. Use the IMU data of the headset to predict the poses of the next 2 frames for motion prediction compensation, and reduce the dizziness caused by motion blur through pre-rendered compensation images. The formula is as follows: where θ current is the current motion angle, θ predicted is the predicted future angle, ω is the angular velocity, and α is the angular acceleration; 4) Visual Protection: a. Dynamically adjust the color temperature according to the usage duration, reducing it from 6500K to 4000K for blue light filtering; b. Automatically match the virtual and real light intensities through an ambient light sensor; c. The system enforces a rest mechanism according to the usage duration.
7. A virtual reality multi-person interaction system, characterized in that, It includes a terminal management module for providing an audio-visual experience and capturing body movements and facial expressions to achieve interaction with the virtual environment. A network architecture module for adopting a distributed architecture for the central server to achieve spatio-temporal consistency of the status and event synchronization mechanism. A software function module for dynamically generating and simulating the environment according to the environment engine, providing user authentication management and social functions. A technology optimization module for providing network optimization strategies to achieve AI enhancement and pipeline rendering optimization.
8. A method for virtual reality multi-person interaction according to claim 7, characterized in that, The terminal management module includes a headset device, an input device, and an auxiliary perception device, and configures network infrastructure for the terminal management module. The network architecture module includes a synchronization engine and a distributed service architecture layer. The software function module includes a user management system, an environment engine, and a social interaction unit. The technology optimization module includes a network optimization unit, an AI enhancement unit, and a rendering pipeline optimization unit.
Citation Information
Patent Citations
Virtual reality multi-person interaction method and system
CN111782054A
Cited By
Virtual laboratory simulation system and method based on VR technology
CN120704540A
A virtual laboratory simulation system and method based on VR technology
CN120704540B
WebXR-based immersive interactive interface generation system and method
CN121209868A