Multi-modal perception fusion mr counter-play interaction method and system

By using multimodal perception fusion technology, combining eye-tracking and gesture data streams with real-world environmental data to generate virtual weapons, the problem of personalized generation in existing MR combat interactions has been solved, achieving an efficient and immersive virtual object generation experience.

CN121060068BActive Publication Date: 2026-02-24SHIYOU (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511216916.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-02-24
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing MR combat interaction technologies struggle to achieve highly personalized and creative weapon generation for users. The existing interaction paradigm leads to a loss of immersion and an inefficient generation process, failing to meet users' high expectations for personalized configurations and instant generation.

Method used

By using multimodal perception fusion technology, eye-tracking data streams and hand gesture data streams are used for intent activation and spatial calibration of generation. Material context sampling is performed by combining real-world 3D mesh data. Virtual weapon primitives are selected and instantiated, and modular selection and parameter fine-tuning are performed through gaze guidance and gesture drive.

Benefits of technology

It achieves a closed-loop, immersive generation experience from the initial intention to the detailed configuration. The user's gaze and hands serve as a unified multi-dimensional input channel, enabling efficient and personalized generation of virtual objects, enhancing immersion and natural interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121060068B_ABST
    Figure CN121060068B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal perception fusion MR battle interaction method and system, relates to the technical field of MR interaction, and first dynamically captures the generation intention of a user in a real space by cooperatively analyzing the eye movement focus and the hand posture of the user, then further fuses the line-of-sight direction and the gesture instruction, carries out context sampling on the real environment physical material concerned by the user, and seamlessly migrates the texture and the texture of the real world to attribute configuration of a virtual object. Based on this, the instantiation of a virtual object core framework and the intuitive selection of modular components are completed by using a refined gesture driving mechanism, the modular locking is carried out by line-of-sight guidance, and finally, the parameterized fine adjustment of the locked module is carried out through a continuous gesture. In this way, the line-of-sight and the two hands of the user can be taken as unified multi-dimensional input channels, and a closed-loop and immersive generation experience from intention generation, material association, form construction to detail configuration is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of MR interaction technology, and more specifically, to a multimodal perception fusion MR battle interaction method and system. Background Technology

[0002] With the rise of the metaverse concept and the rapid development of mixed reality (MR) technology, the boundaries between virtual and reality are becoming increasingly blurred, bringing users an unprecedented immersive experience. Especially in the gaming and entertainment field, MR combat interaction, with its unique ability to perceive the real environment and overlay virtual content, is becoming a core direction for the next generation of human-computer interaction. To achieve truly immersive interaction in MR combat, the system needs to accurately and in real-time understand the user's complex intentions and seamlessly transform these intentions into dynamic virtual content. This requires MR systems to possess powerful multimodal perception and fusion capabilities, integrating various information flows such as eye movements, gestures, and environmental data to simulate natural and intuitive human interaction methods, thereby breaking free from the constraints of traditional controllers or menu-based operations and providing users with unprecedented freedom and creativity.

[0003] Existing technologies attempt to improve the naturalness and efficiency of interaction by integrating eye-tracking and gesture recognition. For example, some solutions utilize eye contact as a quick target selection tool, followed by gestures for confirmation or simple grasping and moving operations. However, most of these existing multimodal interaction solutions remain at a relatively shallow command execution level. When faced with tasks requiring highly personalized and complex attribute configurations for virtual objects, such as creating a unique weapon in a combat game, the demand for such high personalization and creativity far exceeds the scope of simple commands. Users cannot accurately convey a complex generation intention encompassing multiple dimensions of information, including style, structure, and function, simply through a combination of gaze and gestures. Furthermore, existing interaction paradigms often force users to make selections in cumbersome menu interfaces or execute a series of rigid and fragmented preset gestures. This not only severely undermines immersion but also makes the weapon generation process inefficient and lacking in creativity, failing to meet users' high expectations for personalized configuration and instant generation.

[0004] Therefore, there is an urgent need for an optimized multimodal perception fusion MR battle interaction method and system. Summary of the Invention

[0005] This application is made in order to solve the above-mentioned technical problems.

[0006] According to one aspect of this application, a multimodal perception fusion-based MR combat interaction method is provided, comprising: performing intent activation and generation space calibration based on eye-tracking data stream and hand gesture data stream to obtain generation anchor points and switching the system state to generation mode; performing environmental material context sampling based on eye-tracking data stream and hand gesture data stream, combined with real-world 3D mesh data, to obtain a material profile; selecting and instantiating weapon primitives based on hand gesture data stream, generation anchor points, and material profiles to obtain a basic virtual weapon model; performing gaze-guided modular selection on the basic virtual weapon model based on eye-tracking data stream to obtain selected module IDs; and performing gesture-driven parameterized fine-tuning on the selected module IDs based on continuous gesture data stream to obtain a modified virtual weapon model.

[0007] According to another aspect of this application, a multimodal perception fusion MR combat interaction system is provided, comprising: an anchor point generation module, used to perform intent activation and generation space calibration based on eye-tracking data stream and hand posture data stream to obtain generation anchor points and switch the system state to generation mode; a material configuration module, used to perform environmental material context sampling based on eye-tracking data stream and hand posture data stream, combined with real environment 3D mesh data to obtain a material configuration file; a weapon primitive instantiation module, used to select and instantiate weapon primitives based on hand posture data stream, generation anchor points and material configuration files to obtain a basic virtual weapon model; a modular selection module, used to perform gaze-guided modular selection on the basic virtual weapon model based on eye-tracking data stream to obtain the selected module ID; and a parametric fine-tuning module, used to perform gesture-driven parametric fine-tuning on the selected module ID based on continuous gesture data stream to obtain a modified virtual weapon model.

[0008] Compared with existing technologies, this application provides a multimodal perception fusion-based MR combat interaction method and system. First, it dynamically captures the user's generative intent in real space by collaboratively analyzing the user's eye-tracking focus and hand gestures. Then, it further integrates gaze direction and gesture commands, performing contextual sampling of the physical materials of the real environment that the user is interested in, seamlessly transferring the textures and feel of the real world to the attribute configuration of virtual objects. Based on this, a refined gesture-driven mechanism is used to instantiate the core framework of virtual objects and intuitively select modular components. Then, gaze guidance is used for modular locking, and finally, continuous gestures are used to fine-tune the parameters of the locked modules. In this way, the user's gaze and hands are used as a unified multi-dimensional input channel, achieving a closed-loop, immersive generative experience from intent generation, material association, form construction to detailed configuration. Attached Figure Description

[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 This is a flowchart of a multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0011] Figure 2 This is a data flow diagram of a multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0012] Figure 3 This is a flowchart of sub-step S1 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0013] Figure 4 This is a flowchart of sub-step S2 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0014] Figure 5 This is a flowchart of sub-step S3 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0015] Figure 6 This is a flowchart of sub-step S33 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0016] Figure 7 This is a flowchart of sub-step S5 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application.

[0017] Figure 8 This is a block diagram of a multimodal perception fusion MR battle interaction system according to an embodiment of this application. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] To address the problems mentioned above in the background technology, this application proposes a multimodal perception fusion MR battle interaction method. Figure 1This is a flowchart of a multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 2 This is a data flow diagram of a multimodal perception fusion MR battle interaction method according to an embodiment of this application. For example... Figure 1 and Figure 2 As shown, the multimodal perception fusion MR combat interaction method includes the following steps: S1, based on eye-tracking data stream and hand gesture data stream, perform intent activation and generation space calibration to obtain generation anchor points and switch the system state to generation mode; S2, based on eye-tracking data stream and hand gesture data stream, and combined with real environment 3D mesh data, perform environmental material context sampling to obtain material configuration files; S3, based on hand gesture data stream, generation anchor points, and material configuration files, perform weapon primitive selection and instantiation to obtain a basic virtual weapon model; S4, based on eye-tracking data stream, perform gaze-guided modular selection on the basic virtual weapon model to obtain the selected module ID; S5, based on continuous gesture data stream, perform gesture-driven parameterized fine-tuning on the selected module ID to obtain the modified virtual weapon model.

[0020] In the aforementioned multimodal perception fusion-based MR combat interaction method, step S1 involves performing intent activation and spatial calibration based on eye-tracking data streams and hand gesture data streams to obtain generation anchor points and switch the system state to generation mode. It should be understood that single-modal interaction data cannot accurately capture the user's deep intent and spatial positioning needs, while eye-tracking data streams can reflect the user's visual focus, and hand gesture data streams can reflect the intention of operation direction. Based on this, this application performs intent activation determination on eye-tracking data streams and hand gesture data streams to confirm the user's generation needs, and uses a spatial calibration algorithm to map visual attention and operation direction to three-dimensional spatial coordinates to generate precise generation anchor points. Simultaneously, it triggers the system state to switch from standby or interaction mode to generation mode, providing spatial reference and mode support for subsequent content generation.

[0021] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S1 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 3 As shown, step S1 includes: S11, performing gesture detection on the hand posture data stream to determine whether the first condition is met; S12, performing gaze detection on the eye movement data stream to determine whether the second condition is met; S13, in response to the simultaneous satisfaction of the first and second conditions, switching the system state to the generation mode; S14, in response to the simultaneous satisfaction of the first and second conditions, creating the generation anchor point on the palm of the left hand.

[0022] Specifically, in step S11, gesture detection is performed on the hand posture data stream to determine whether a first condition is met. The first condition is that a hand presents a tray gesture and remains stable for more than a threshold time. In a specific example of this application, the threshold time is set to 0.5 seconds. This allows for the confirmation of the user's true intention through continuous operation for a certain period, reducing the probability of erroneous operation, while ensuring the timely response of the system after the user's operation, maintaining the natural fluency of the interaction, and meeting the dual requirements of accuracy of intent recognition and real-time interaction in multimodal perception fusion scenarios. Specifically, by analyzing the hand posture data stream in real time, it is identified whether a tray gesture that meets the preset standard appears and whether the gesture remains stable for more than a threshold time. This serves as a hand signal verification for the user's generated intent, improving the anti-interference capability of intent recognition, ensuring the validity and reliability of hand input information, and providing a reliable input basis for subsequent system state switching.

[0023] Specifically, in one possible embodiment, step S11 is implemented as follows: First, a hand posture data stream is received, which contains a three-dimensional coordinate time series of each joint of the user's left hand. Then, the data stream is parsed frame by frame, and the coordinate data of the palm root, thumb root, little finger root, and wrist joint of the left hand in each frame are extracted. By calculating the vector connecting the palm root and the wrist joint, it is determined whether the palm is in an upward posture, and the angle between the vector and the vertical direction is less than 15 degrees. At the same time, the relative positions of the fingertips and palm roots are analyzed to confirm that the five fingers are in a naturally extended state, that is, the distance between the fingertips and the palm root is within a preset range, and the angle between the fingers is maintained between 30 and 60 degrees, thereby identifying the tray gesture characteristics. Subsequently, posture stability monitoring is activated, and the Euclidean distance change of each joint coordinate in 10 consecutive frames is continuously calculated. When the position change of all joints is less than 0.02 meters, the posture is determined to be in a stable state. At the same time, a timer is started, and when the duration of the stable state reaches 0.5 seconds, it is determined that the first condition is met.

[0024] Specifically, in step S12, gaze detection is performed on the eye-tracking data stream to determine whether a second condition is met. The second condition is that the user's gaze remains continuously on the center area of ​​the palm of the tray gesture for a duration exceeding a threshold time. In essence, this application continuously tracks the eye-tracking data stream to determine whether the user's gaze is continuously focused on the center area of ​​the palm of the tray gesture and whether the dwell time exceeds a threshold, thereby verifying the user's focus and authenticity of intent and providing a second modal verification for intent activation. This accurately captures the user's gaze focus state and forms a collaborative verification mechanism with hand posture signals, significantly improving the accuracy of intent recognition. It ensures that only when the user actively focuses their attention on the operation area is the signal considered valid input, reducing interference from non-intent operations.

[0025] Specifically, in one possible embodiment, step S12 is implemented as follows: First, an eye-tracking data stream is acquired, which includes the gaze direction vectors of the user's eyes and the pupil center position information. Then, combined with the coordinates of the left palm root and fingertips extracted from the left hand posture data stream, the three-dimensional spatial range of the left palm center area is calculated as a cuboid region extending 0.1 meters towards the fingertips and 0.08 meters laterally, with the palm root as the origin. Next, the gaze direction vector in the eye-tracking data stream is converted into a spatial ray. By detecting the geometric intersection of the ray with the palm center area, it is determined whether the gaze is pointing towards this area. When the intersection of the ray and the area is detected, a dwell timer is started, and the position of the gaze ray is continuously tracked. If the ray always falls within the palm center area in the subsequent 50 frames and does not deviate by more than 0.05 meters, the timer's accumulated duration reaches 0.5 seconds, confirming that the second condition is met.

[0026] Specifically, in step S13, in response to the simultaneous fulfillment of the first and second conditions, the system state is switched to generation mode. It should be understood that the simultaneous fulfillment of multimodal conditions is a clear manifestation of the user's generation intent, requiring a state switch to activate subsequent functional modules related to virtual weapon generation to support further user operations. Specifically, after the dual conditions are verified, the system is switched from non-generation mode to generation mode, activating the runtime environment for subsequent functions such as material sampling and weapon primitive instantiation, providing functional support for the creation of virtual weapons. This achieves a seamless connection from intent confirmation to function activation, ensuring the continuity and real-time nature of the generation process.

[0027] Specifically, in one possible embodiment, step S13 is implemented as follows: When the determination results of the first and second conditions are simultaneously satisfied, the state monitoring module sends a mode switching trigger signal to the system control unit. Upon receiving the signal, the low-power process in the current standby mode is immediately terminated, and standby-related resources are released. Subsequently, the mode configuration interface is called to load the runtime environment corresponding to the generation mode, including activating the eye-tracking-gesture collaborative processing thread, starting the real-time call channel for the 3D mesh data of the environment, and initializing the index service of the virtual weapon arsenal. At the same time, the mode flag in the status register is updated from "STANDBY" to "GENERATION", and a mode-ready command is sent to the material configuration module and the weapon primitive instantiation module, so that each module enters the standby state. Finally, a green progress bar is displayed on the edge of the user's field of vision through the head-mounted display device to complete the animation, accompanied by a short prompt sound, to inform the user that the generation mode has been entered.

[0028] Specifically, in step S14, in response to the simultaneous fulfillment of the first and second conditions, a generation anchor point is created on the palm of the left hand. It should be understood that the instantiation of a virtual weapon requires a clear spatial coordinate reference. The left hand, as a stable physical carrier, provides a fixed spatial reference for association with the user's limbs, ensuring that the virtual model has a stable attachment point in real space and avoiding drift or confusion in the spatial positioning of the virtual object. Specifically, a generation anchor point is established on the palm of the left hand where the conditions are met. By recording the three-dimensional spatial coordinates of this anchor point, a precise spatial positioning reference is provided for the subsequent instantiation of the virtual weapon model, enabling the virtual model to maintain a stable spatial association with the user's limbs. The obtained generation anchor point is accurately created on the palm of the left hand, and its three-dimensional transformation matrix is ​​recorded by the system in real time, providing a reliable world coordinate reference for the subsequent rendering of the virtual weapon model. This achieves precise integration of the virtual object with real space, laying a spatial foundation for subsequent model interaction.

[0029] Specifically, in one possible embodiment, step S14 is implemented as follows: When the first and second conditions are simultaneously met, the latest frame of left hand joint coordinate data is extracted from the left hand posture data stream, including the three-dimensional coordinates of the palm root joint, index finger root joint, and middle finger root joint. Using a preset palm coordinate calculation algorithm, with the palm root joint as the reference point and combining the average coordinates of the index finger root and middle finger root joints, the three-dimensional spatial coordinates of the left palm are calculated: X = palm root X + 0.3 × (index finger root X + middle finger root X) / 2. The Y and Z coordinates are derived similarly. These coordinates are then determined as the initial three-dimensional coordinates for generating the anchor point, and a dynamic association mechanism between the anchor point and the left hand joint is established. Simultaneously, a real-time anchor point coordinate update thread is started. Each frame, based on the position changes of the left hand joints in the hand posture data stream, the palm coordinates are recalculated and the position of the generated anchor point is updated, ensuring that the anchor point is always spatially synchronized with the left palm. Finally, the coordinate data of the generated anchor point is stored in a shared memory area, and a confirmation message indicating successful anchor point creation is returned to the anchor point generation module. This anchor point will subsequently serve as the spatial reference for the instantiation of the virtual weapon model.

[0030] In the aforementioned multimodal perception fusion-based MR combat interaction method, step S2 involves sampling environmental material context based on eye-tracking data streams and hand gesture data streams, combined with real-world 3D mesh data, to obtain a material profile. It should be understood that the real-world 3D mesh data contains environmental geometric structure information, and combining these three can overcome the limitations of a single data modality in acquiring material information. Specifically, this application analyzes eye-tracking data streams to determine the environmental material areas that the user is interested in, combines this with hand gesture data streams to define sampling boundaries, extracts material features based on real-world 3D mesh data, and integrates multi-source data through a context association algorithm to obtain a material profile, providing data support for material matching of generated content.

[0031] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S2 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 4 As shown, step S2 includes: S21, emitting a ray based on the eye-tracking data stream; S22, determining the intersection point between the ray and the real environment 3D mesh data; S23, extracting texture and color information near the intersection point from the real environment 3D mesh data; S24, in response to detecting that the hand gesture data stream is a sampling gesture, packaging the texture and color information near the intersection point into the material configuration file.

[0032] Specifically, in step S21, a ray is emitted based on the eye-tracking data stream. It should be understood that the ray maps the user's gaze onto the three-dimensional space of the real environment, providing a spatial correlation basis for subsequent material sampling point localization. Specifically, this application generates a corresponding spatial ray based on the real-time changes in the eye-tracking data stream, enabling the ray to follow the user's gaze movement accurately in real time, precisely matching the direction of the user's gaze in the real environment, achieving seamless mapping between the gaze direction and the three-dimensional spatial ray, ensuring that the ray accurately points to the real environment area of ​​interest to the user, and providing a reliable spatial path for subsequent intersection point calculation.

[0033] Specifically, in one possible embodiment, step S21 is implemented as follows: First, the eye-tracking data stream is preprocessed by removing noise data caused by blinking through a filtering algorithm, retaining valid gaze frames. Combining user head posture tracking data, including head rotation angle and translation vector, the coordinates of the pupil centers are transformed to three-dimensional positions in the world coordinate system. Using the center of the left pupil as the ray origin reference, the position of the right pupil center is calibrated to determine the spatial origin of the ray. Subsequently, the gaze direction vector of each frame is extracted and transformed from the user's local eye coordinate system to the world coordinate system using a coordinate transformation matrix to obtain the ray's direction vector. Finally, based on real-time updated eye-tracking data, the ray's origin and direction are refreshed every 10 milliseconds, generating a continuous spatial ray. This ray accurately reflects the user's gaze trajectory in the real environment, providing a dynamic tracking spatial path for subsequent intersection detection.

[0034] Specifically, step S22 involves determining the intersection point between the ray and the 3D mesh data of the real environment. It should be understood that the specific material sampling location cannot be determined solely by the ray. Therefore, this application calculates the geometric intersection point between the ray and the 3D mesh data of the real environment to accurately locate the spatial coordinates of the physical surface of the real environment that the user's line of sight is focused on. These coordinates strictly correspond to the physical surface viewed by the user in the real environment, providing a clear spatial reference for subsequently extracting texture and color information from this location, thus ensuring the spatial accuracy of material sampling.

[0035] Specifically, in one possible embodiment, step S22 is implemented as follows: First, load real-world 3D mesh data, which consists of millions of triangular faces, each containing three vertex coordinates and face normal information. Then, initiate a ray-mesh intersection detection process. Using the ray's origin and direction vector as input, a spatial partitioning retrieval algorithm is employed to first determine the mesh partitions the ray might traverse, narrowing the detection range. Within the target partition, traverse all triangular faces and determine whether the ray intersects a face using geometric calculations: calculate the intersection point between the ray and the plane containing the face, and then use the barycentric coordinate method to determine if the intersection point is located inside the face. For all detected valid intersection points, calculate their Euclidean distance from the ray's origin, and select the closest intersection point as the final result. If the ray does not intersect any face, extend the ray length to a preset maximum value and re-detect. The preset maximum ray length extension is 5 meters to ensure that valid intersection points are still obtained when the user's line of sight is pointing towards the distant environment.

[0036] Specifically, step S23 involves extracting texture and color information near the intersection point from the real-world 3D mesh data. Specifically, this application, based on the intersection point coordinates, parses and obtains texture details of the area surrounding the intersection point from the real-world 3D mesh data, such as texture patterns, texture density, texture features related to surface roughness, and color attributes, such as RGB color values, color saturation, and brightness distribution, providing raw visual feature data for the material configuration of the virtual weapon. The successfully extracted texture and color information, which accurately reflects the visual characteristics of the physical surface near the intersection point, has high fidelity and can accurately reproduce the visual texture of materials in the real environment, providing high-quality raw data support for the subsequent generation of material configuration files.

[0037] Specifically, in one possible embodiment, step S23 is implemented as follows: After determining the triangular facet where the intersection point is located, the texture mapping data associated with the facet is retrieved from the real environment's 3D mesh data, including the texture atlas index, the texture coordinates corresponding to the facet vertices, and the texture sampling method parameters. Based on the 3D coordinates of the intersection point on the facet, its corresponding 2D texture coordinates are calculated through interpolation. Centered on these texture coordinates, a square area with a side length of 64 pixels is extracted from the texture atlas as the sampling range. All pixel data within this area are extracted, including the RGB values ​​and Alpha channel information of each pixel, thereby obtaining texture details near the intersection point, such as the direction of wood grain and the particle distribution of stone. Simultaneously, the RGB mean value of all pixels within the sampling range is calculated, and the color values ​​are corrected in conjunction with ambient lighting data to obtain color information reflecting the real visual perception, including the dominant hue, color saturation, and brightness distribution characteristics. For texture edge areas, an edge extension algorithm is used to avoid discontinuities in the sampled data.

[0038] Specifically, in step S24, in response to detecting that the hand gesture data stream is a sampling gesture, the texture information and color information near the intersection point are packaged into the material configuration file. Specifically, after recognizing the user's sampling gesture, this application structures and packages the previously extracted texture and color information near the intersection point according to a preset data format to form a material configuration file that can be directly called by the virtual weapon model. This file fully contains the texture and color data near the intersection point, and its format is compatible with the subsequent rendering system of the virtual weapon model. It can be accurately and efficiently applied to the material configuration of weapon primitives, giving the virtual weapon a visual texture consistent with the sampling area of ​​the real environment, achieving seamless integration of virtual and real materials, and enhancing the realism and immersion of the virtual weapon.

[0039] Specifically, in one possible embodiment, step S24 is implemented as follows: When a rapid pinching motion is detected between the tips of the right thumb and index finger, i.e., the distance between the two fingertips shortens from 0.1 meters to within 0.02 meters, and then the hand is released, i.e., the distance returns to more than 0.1 meters, this is determined to be a sampling gesture. Upon detecting the sampling gesture, the previously extracted texture and color information near the intersection point is immediately retrieved. The texture information is converted into a MIP texture chain according to a specified format, containing six different resolution levels, and the texture repetition pattern and filtering method are recorded. The color information is converted into a standardized parameter set, including the baseline values ​​of the RGB three channels, color attenuation coefficients, and ambient light reflectivity. Finally, these data are encapsulated in a preset binary format, and a data header containing the sampling timestamp and intersection coordinates, along with a checksum, are added to generate a material configuration file. This file is stored in the system's temporary cache and marked as "pending application" for subsequent use during weapon primitive instantiation.

[0040] In the aforementioned multimodal perception fusion-based MR combat interaction method, step S3 involves selecting and instantiating weapon primitives based on hand posture data streams, generated anchor points, and material configuration files to obtain a basic virtual weapon model. It should be understood that driving the selection and instantiation of weapon primitives through hand posture data streams ensures that the virtual weapon's grip position and movement trajectory are completely synchronized with the user's hand, avoiding a disconnected interaction caused by a fixed model. Generating anchor points establishes a spatial mapping relationship between the virtual weapon and the user's hand joints, ensuring the model maintains physical stability during movement. The material configuration file defines the visual and physical attributes of each weapon component, ensuring that the resulting basic virtual weapon model conforms to the characteristics of a real weapon in terms of appearance and mechanical performance. This constructs a basic virtual weapon model that is physically consistent with the user's hand posture in real time, providing an operable platform for subsequent modular adjustments.

[0041] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S3 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 5 As shown, step S3 includes: S31, in response to the hand gesture data stream being a preset category gesture, retrieving the basic 3D model corresponding to the preset category gesture from the virtual weapon library; S32, applying the material configuration file to the basic 3D model to obtain configured model data; S33, rendering the configured model data above the generated anchor point to obtain the basic virtual weapon model.

[0042] Specifically, in step S31, in response to a hand gesture data stream being a preset category gesture, a basic 3D model corresponding to the preset category gesture is retrieved from the virtual weapon library. Specifically, by recognizing the preset category gesture in the hand gesture data stream, a precise association is established between the gesture and the basic 3D model in the virtual weapon library. This successfully retrieves the basic 3D model matching the user's preset gesture from the virtual weapon library. This model possesses clear structural divisions and editable attributes, providing a morphological basis for subsequent material applications and spatial instantiation. This achieves natural interaction and precise response in weapon type selection, maintaining the continuity and immersion of the interaction process.

[0043] Specifically, in one possible embodiment, step S31 is implemented as follows: First, a continuous data stream of hand postures is received, which includes the three-dimensional coordinates and angle change sequences of each joint of the left and right hands. A preset category gesture is defined as follows: "The left hand is raised, palm facing upwards, fingers naturally spread, the line connecting the palm heel and wrist joint forms an angle of 15-25 degrees with the horizontal plane, the right thumb and middle finger touch to form a ring, the other three fingers are naturally extended, and the right wrist joint is on the same horizontal plane as the left wrist joint with a distance of 0.2-0.3 meters between them." The data stream is then parsed frame by frame to extract the left palm orientation vector, the distance between the right thumb and middle finger tips, and the relative position parameters of the two wrist joints. When 25 consecutive frames satisfy the above posture characteristics, and the inter-frame variation of the joint coordinates is less than 0.012 meters, indicating a stable posture, the preset category gesture is confirmed. A retrieval command is then sent to the virtual weapon arsenal, containing the category identifier corresponding to the gesture. After receiving the instruction, the virtual weapon arsenal retrieves the internally set gesture-model mapping table, determines that the corresponding basic 3D model is the rifle basic model, reads the 3D data of the model from the storage directory, including vertex coordinate set, triangular facet topology, basic material channel information and model skeleton data, loads it into the system running memory and returns a confirmation signal that loading is complete.

[0044] Specifically, in step S32, the material configuration file is applied to the base 3D model to obtain configured model data. Specifically, texture information, color information, and other data in the material configuration file are mapped to the surface attribute channels of the base 3D model, such as diffuse reflection and specular highlights. This allows the base 3D model to visually exhibit texture characteristics consistent with the sampled area of ​​the real environment, forming configured model data that combines structural form and material attributes. This provides model data with complete visual attributes for subsequent spatial rendering, enhancing the realism and immersion of the virtual weapon.

[0045] Specifically, in one possible embodiment, step S32 is implemented as follows: First, the mesh data of the basic 3D model is acquired. This data includes vertex information, face indexes, and material attribute binding relationships for each component of the model, with each component corresponding to an independent material channel. Simultaneously, the previously generated material configuration file is called, and the texture and color information contained in the file are parsed. Based on the material channel identifiers of each component of the model, a mapping relationship between the material configuration file and the model components is established. For example, texture information is assigned to the diffuse channel of the gun barrel, and the ambient light reflectance parameter in the color information is applied to the specular channel of the grip. For the gun barrel, the sampled texture data is mapped to each vertex of the gun barrel mesh through texture coordinate mapping, ensuring that the texture naturally extends with the curvature of the gun barrel. For the grip, its base color is adjusted according to the color reference value, and the color change pattern under different lighting angles is set in conjunction with the color attenuation coefficient. During application, the material attributes of the model need to be consistent to ensure that the texture stretching ratio, color parameters, and model geometry match, avoiding texture distortion or color distortion. After processing, configured model data containing complete material attributes is generated and stored in the model cache.

[0046] Specifically, in step S33, the configured model data is rendered above the generation anchor point to obtain a basic virtual weapon model. Specifically, the configured model data is spatially instantiated based on the 3D transformation matrix of the generation anchor point using a rendering engine. The configured model data is successfully rendered above the generation anchor point as a basic virtual weapon model. This model maintains a stable positional association with the generation anchor point in space, possesses a clear visual form and realistic material representation, and can be tracked by the system in real time and respond to subsequent modular selection operations. This completes the crucial transformation of the virtual weapon from data configuration to spatial representation, providing users with a virtual interactive object that can be intuitively perceived and operated.

[0047] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S33 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 6 As shown, step S33 includes: S331, inputting the configured model data into the rendering engine to obtain a basic virtual weapon model instance; S332, obtaining the three-dimensional transformation matrix of the generated anchor point; S333, setting the three-dimensional transformation matrix as the world coordinate transformation of the basic virtual weapon model instance; S334, adding the basic virtual weapon model instance to the scene rendering queue.

[0048] More specifically, step S331 involves inputting the configured model data into the rendering engine to obtain a basic virtual weapon model instance. Specifically, this application utilizes the rendering engine to perform format parsing, resource loading, and instantiation operations on the configured model data, generating a basic virtual weapon model instance containing complete rendering parameters. This instance includes complete geometric topology and material mapping relationships, and can respond to subsequent coordinate transformations and rendering commands, enabling the model data to be rendered in real time, thus laying the foundation for accurate spatial positioning and visualization of the model.

[0049] Specifically, in one possible embodiment, step S331 is implemented as follows: First, the rendering engine's model loading interface is called, and the configured model data is passed into the interface. The configured model data includes the weapon's vertex coordinate set, triangular facet topology, material parameters for each partition, and bone binding information. After the rendering engine receives the data, it starts the data parsing process, first normalizing the vertex coordinates and converting the local coordinates into an engine-compatible floating-point format. Then, it parses the triangular facet index to construct the geometric topology of the mesh. Subsequently, it loads the material parameters, binds the corresponding texture atlas and color channels to each mesh partition, and sets rendering attributes such as diffuse and specular highlights. At the same time, the engine verifies the bone binding information to ensure that the weight mapping between joints and the mesh is correct. After parsing is completed, the engine performs an instantiation operation to create an independent rendering instance object for the model. This object contains complete geometric data, material configuration, and rendering status markers, and can respond to subsequent coordinate transformations and rendering commands, ultimately generating a basic virtual weapon model instance.

[0050] More specifically, step S332 involves obtaining the three-dimensional transformation matrix of the generated anchor point. It should be understood that the spatial position, rotation angle, and scaling ratio of the generated anchor point need to be quantitatively represented using a three-dimensional transformation matrix. This matrix is ​​the core data describing the spatial state of the anchor point in the world coordinate system, providing a precise benchmark for the spatial positioning of the virtual weapon model instance. Since the spatial placement of virtual objects relies on explicit coordinate transformation parameters to match the real environment, the purpose of this step is to extract the three-dimensional transformation matrix corresponding to the generated anchor point from the system. This matrix contains translation, rotation, and scaling components, which can completely describe the absolute position and attitude of the anchor point in world space, accurately reflecting the spatial state of the generated anchor point. As a reference standard for the spatial positioning of the virtual weapon model instance, it provides a quantitative basis for the subsequent coordinate binding of the model instance, ensuring that the spatial positioning of the model instance remains strictly consistent with the anchor point.

[0051] Specifically, in one possible embodiment, step S332 is implemented as follows: First, the real-time data buffer of the anchor point generation module is accessed, which stores the dynamic spatial information of the generated anchor point. The real-time three-dimensional coordinates of the left palm based on the world coordinate system are extracted from the buffer, including position parameters in the X, Y, and Z axes. Simultaneously, the rotation information of the anchor point is obtained, including rotation angles around the pitch (X-axis), yaw (Y-axis), and roll (Z-axis), calculated from the relative posture of the left wrist joint and the palm heel. Furthermore, the scaling factor of the anchor point is extracted, with a default value of 1.0, as the anchor point size is fixed. Then, a matrix calculation function is called to convert the position parameters into translation components, the rotation angles into rotation matrices, and the scaling factor into scaling components. The translation, rotation, and scaling components are then combined into a 4x4 three-dimensional transformation matrix through matrix multiplication. This matrix fully describes the position, posture, and size of the generated anchor point in world space and is stored in the system's shared memory for subsequent steps.

[0052] More specifically, in step S333, the three-dimensional transformation matrix is ​​set as the world coordinate transformation of the basic virtual weapon model instance. Specifically, the three-dimensional transformation matrix for generating the anchor point is used as the world coordinate transformation parameter for the basic virtual weapon model instance, ensuring that the position of the basic virtual weapon model instance in world space maintains a preset spatial relationship with the generated anchor point on the left palm, such as above the anchor point. This creates a stable positional association between the virtual model and the generated anchor point in space, achieving real-time tracking of the virtual model and the user's limb movements, avoiding spatial drift of the virtual object, and ensuring that the virtual weapon has a predictable spatial attachment position in the real environment.

[0053] Specifically, in one possible embodiment, step S333 is implemented as follows: First, the 3D transformation matrix for generating the anchor point is read from shared memory, and the attribute interface of the basic virtual weapon model instance is obtained. The world coordinate transformation parameters of the model instance are accessed through this interface; these parameters, by default, include initial position, rotation, and scaling information. Then, the coordinate transformation assignment function is called, and the translation component in the 3D transformation matrix, corresponding to the coordinates of the left palm, is written into the position parameters of the model instance, making the spatial origin of the model instance coincide with the anchor point. Then, the rotation component in the matrix, corresponding to the posture angle of the left hand, is written into the rotation parameters of the model instance, ensuring that the model's orientation is consistent with the left palm, while the scaling component remains unchanged. After the assignment is completed, the coordinate update event of the model instance is triggered, causing the model instance to apply the new world coordinate transformation in real time, ensuring that its position and posture in world space are strictly related to the generated anchor point and change synchronously with the movement or rotation of the anchor point.

[0054] More specifically, step S334 involves adding the basic virtual weapon model instance to the scene rendering queue. Specifically, the basic virtual weapon model instance with its coordinates already set is added to the scene rendering queue, making it part of the scene rendering process. Through frame-by-frame processing by the rendering engine, it is calculated and output in real time during each frame rendering. Users can intuitively observe a virtual weapon with realistic materials and a stable spatial position. This weapon blends naturally with the real environment and can respond to subsequent interactive operations, completing a closed loop from model instantiation to visualization.

[0055] Specifically, in one possible embodiment, step S334 is implemented as follows: First, the rendering priority flag of the basic virtual weapon model instance is obtained, set to high by default, as the weapon is the core interactive object. Its instance status is then checked, requiring it to be in a "ready" state, meaning the coordinate transformation settings have been completed. Next, the queue addition function is called, passing the model instance's reference pointer to the queue, along with its rendering priority and update frequency parameters, consistent with the scene frame rate to ensure real-time performance. Upon receiving the model instance, the rendering queue inserts it into the corresponding position according to priority, with high-priority objects placed at the front of the rendering process. The queue also records the model instance's dependencies, such as its association with left-hand pose data, ensuring that the world coordinate transformation of the model instance has been updated according to the latest anchor point matrix before each frame is rendered. After addition, the queue returns a confirmation signal to the system, indicating that the model instance has been included in the scene rendering process and will be rendered and output in real-time in subsequent frames.

[0056] In the aforementioned multimodal perception fusion-based MR combat interaction method, step S4 involves performing a gaze-guided modular selection of the basic virtual weapon model based on eye-tracking data streams to obtain the selected module ID. It should be understood that eye-tracking, as a natural and low-latency pointing method, can accurately locate the user's attention focus, avoiding the disruption of immersion caused by traditional physical interaction or menu selection. Specifically, this application continuously tracks the user's eye-tracking data stream to identify the dwell position of the gaze on the basic virtual weapon model. When the gaze dwells on a specific module area for a certain duration reaches a threshold, the module is determined to be the target of the user's intended modification, and the corresponding module ID is output. This achieves contactless, highly accurate module positioning, simplifies the selection process, enhances the intuitiveness and immersion of the interaction, and provides clear target guidance for parameterized fine-tuning.

[0057] Specifically, in one possible embodiment, step S4 is implemented as follows: First, modular structure data of the basic virtual weapon model is acquired. This model is pre-divided into four independent functional modules: barrel, main body, grip, and scope. Each module corresponds to a preset three-dimensional spatial region. Using the model's local coordinate system as a reference, the spatial boundaries of each module are defined by the vertex coordinate range. For example, the barrel module extends from the muzzle vertex to the vertex range where it connects to the main body. Simultaneously, eye-tracking data streams are received. These data streams contain the user's eye gaze direction vectors and the three-dimensional coordinates of the pupil center in the world coordinate system. The eye-tracking data streams are analyzed in real time. Each frame, the gaze direction vector is converted into a spatial ray. Starting from the pupil center, the ray extends along the gaze direction and, combined with the world coordinate transformation parameters of the basic virtual weapon model, the ray is transformed from the world coordinate system to the model's local coordinate system. Subsequently, the geometric intersection relationship between the transformed ray and the spatial regions of each module is detected. By calculating the intersection point of the ray with the module boundary box, it is determined whether the gaze falls within a certain module region. When the system detects that the user's gaze remains continuously within the spatial area of ​​the barrel module, a timer is started. Simultaneously, the system continuously tracks the ray position in subsequent frames. If the ray does not deviate from the area for 30 consecutive frames, and the timer's cumulative duration reaches 0.5 seconds, it is determined that the user intends to select that module. At this point, the system retrieves the corresponding identifier information for the barrel module from the model's modular configuration table, generates the selected module ID, and stores this ID in the interaction state buffer for subsequent parameter fine-tuning steps.

[0058] In the aforementioned multimodal perception fusion-based MR combat interaction method, step S5 involves performing gesture-driven parameter fine-tuning on the selected module ID based on a continuous gesture data stream to obtain a modified virtual weapon model. It should be understood that continuous gestures provide continuous, dynamic control signals, suitable for achieving fine-grained, real-time parameter adjustments, and are more in line with natural interaction logic compared to discrete commands or numerical inputs. Therefore, this application drives the dynamic changes of module parameters through a continuous gesture data stream, ultimately generating a modified virtual weapon model that meets the user's precise needs. The parameter adjustment process is responsive and transitions naturally, significantly improving the user's creative control over the virtual weapon and enhancing the interactive immersive experience.

[0059] In particular, in one specific embodiment, Figure 7 This is a flowchart of sub-step S5 of the multimodal perception fusion MR battle interaction method according to an embodiment of this application. Figure 7As shown, step S5 includes: S51, using the selected module ID as the query key, retrieving a preset module-control parameter mapping database to obtain an activity control mapping; S52, extracting, filtering, and normalizing the continuous gesture data stream to obtain the final parameter value; S53, based on the final parameter value and the activity control mapping, performing parameter-driven real-time model deformation and rendering update on the basic virtual weapon model to obtain the modified virtual weapon model.

[0060] Specifically, in step S51, the selected module ID is used as the query key to retrieve a preset module-control parameter mapping database to obtain the activity control mapping. It should be understood that different modules have different adjustable parameters, such as length, angle, and curvature, and the association logic between these parameters and gesture control signals differs. Therefore, it is necessary to accurately locate the specific control rules through the module ID to ensure that gesture operations can accurately act on the target parameters. Specifically, this application, based on the selected module ID, retrieves the list of adjustable parameters corresponding to that module and the mapping relationship between each parameter and gesture signals from the preset database, such as the ratio of length change corresponding to the opening and closing amplitude of the gesture. This obtains the activity control mapping matching the selected module, clearly defining the adjustable parameter types, parameter ranges, and corresponding gesture control logic of the module, ensuring that subsequent gesture operations can be accurately associated with the target parameters, and providing a precise control basis for parameterized fine-tuning.

[0061] Specifically, in one possible embodiment, step S51 is implemented as follows: First, the retrieval interface of the module-control parameter mapping database is called, and the module ID is passed to the interface. The database stores the association information between each module and the adjustable parameters. After receiving the retrieval request, the entry corresponding to the grip module is located through index matching. This entry contains the adjustable parameters of the grip, including grip length, grip diameter, surface curvature, and anti-slip texture depth. It also records the mapping rules between each parameter and continuous gestures, such as the opening and closing range of the thumb and index finger corresponding to the change in grip diameter, the sliding distance of the finger along the grip axis corresponding to the change in grip length, the change in the palm contact angle corresponding to the adjustment of surface curvature, and the change in fingertip pressure corresponding to the change in texture depth. It also includes the adjustment range of each parameter, such as the grip diameter of 0.08-0.12 meters. The database integrates these parameter information, mapping rules, and adjustment ranges into an activity control mapping, and returns it to the system through the interface as the control basis for subsequent parameterized fine-tuning.

[0062] Specifically, step S52 involves extracting, filtering, and normalizing the continuous gesture data stream to obtain the final parameter values. Specifically, this application extracts effective signal components reflecting the user's intent from the continuous gesture data stream, removes noise using filtering algorithms such as moving average filtering, and then normalizes the signals to the effective adjustment range of the module parameters, such as 0-100 corresponding to the minimum and maximum values ​​of the parameters, generating final parameter values ​​that can be directly used for parameter driving. These values ​​maintain consistency with the changing trend of the gesture action and strictly match the parameter adjustment range, effectively avoiding parameter fluctuations caused by noise. This ensures a stable and reliable linear or nonlinear correspondence between gesture operation and parameter adjustment, providing high-quality input for real-time parameter adjustment.

[0063] Specifically, in one possible embodiment, step S52 is implemented as follows: First, a continuous gesture data stream for the scope module is received, which includes a three-dimensional coordinate time series of the right thumb, index fingertip, and palm base. Then, based on the association rule between lens magnification and thumb-index finger pinch amplitude in the activity control mapping, the straight-line distance between the thumb and index fingertip in each frame is extracted as the original gesture signal. Next, a Kalman filter algorithm is used to process the original signal, eliminating high-frequency noise caused by hand tremors through prediction and update steps, making the signal curve smooth and continuous. Subsequently, a normalization operation is performed, mapping the filtered distance signal to the effective adjustment range of the lens magnification, such as 1-5x, where the minimum distance corresponds to 1x and the maximum distance corresponds to 5x. The real-time magnification value is calculated through nonlinear interpolation. After 30 consecutive frames of verification to ensure signal stability, the final parameter value, i.e., the target value of the current lens magnification, is generated.

[0064] Specifically, in step S53, based on the final parameter values ​​and the activity control mapping, the basic virtual weapon model undergoes parameter-driven real-time deformation and rendering updates to obtain the modified virtual weapon model. Specifically, the model deformation nodes corresponding to the final parameter values ​​are determined based on the activity control mapping, such as the barrel extension node and the grip rotation node. The geometric properties of these nodes change in real time through parameter-driven changes, and rendering data is updated synchronously, such as lighting and shadows adapting to the new form. This allows the modified virtual weapon model to be presented instantly. The model's form details highly match the user's gesture operations, significantly improving the real-time nature and immersion of the interaction. It completes the precise transformation of the virtual weapon from a basic form to a personalized form, achieving visual feedback for user operations.

[0065] Specifically, in one possible embodiment, step S53 is implemented as follows: The model deformation node corresponding to the final parameter value is determined based on the activity control mapping, namely the lens mesh group and the telescopic section of the scope. The model deformation interface is called, and the final parameter value is passed in. The interface calculates the scaling factor of the lens mesh based on the magnification value. The higher the magnification, the larger the radial dimension of the lens. Simultaneously, the vertex position of the telescopic section is adjusted, stretching or shrinking along the axial direction to ensure the proportions of the internal optical structure are adapted. During the deformation process, the rendering parameters of the crosshair are updated synchronously, automatically adjusting the thickness according to the magnification. The crosshair is thinner at high magnification, and the coordinates of the reflective texture on the lens surface are corrected to maintain consistent optical texture. After triggering the rendering update mechanism, the rendering engine recalculates the lighting and shadow effects of the deformed scope, including the adjustment of the refraction area caused by the lens size change, the shadow transition at the connection between the scope and the gun body, and the blurring effect at the edge of the field of view at high magnification. The updated model data is written to the rendering buffer after frame synchronization processing, finally presenting the modified virtual weapon model after the lens magnification adjustment.

[0066] Here, in the process of performing preset category gesture recognition based on hand posture data stream and parameterizing the continuous gesture signal of the continuous gesture data stream, hand parameters, such as the relative position, distance, and angle of hand joint coordinates in space, are required. Therefore, instead of treating gesture recognition as an independent classification problem, and parameter mapping as an independent regression problem, the entire hand posture space can be regarded as a high-dimensional gesture manifold, making each specific hand posture a point on this manifold. That is, a single unified model is constructed to output in real time the probability distribution of the current hand posture point on the manifold that is closest to the category anchor point (i.e., what is the probability of an L-shape, what is the probability of a fist, etc.), and at the same time output a continuous parameter vector describing the inherent geometric properties of the posture point (i.e., regardless of what gesture it is, what are its finger spacing, curvature, etc.).

[0067] Therefore, it is desirable to use probability distribution as context to dynamically and weightedly gate or filter continuous parameters. For example, when the system is highly certain that the user is making an L-shaped gesture, it will mainly activate parameters related to the L-shape (such as the thumb-index finger distance); while when the gesture is ambiguous, the weights of all parameters are lower to avoid misoperation.

[0068] In particular, in another possible preferred embodiment, step S5, which involves fine-tuning the selected module ID based on gesture-driven parameters to obtain a modified virtual weapon model using a continuous gesture data stream, is implemented through a multi-task learning neural network model. Specifically, this includes: obtaining a hand gesture input vector from the continuous gesture data stream; inputting the hand gesture input vector into the shared encoder of the multi-task learning neural network model to generate an encoded feature vector; inputting the encoded feature vector into the classification head of the multi-task learning neural network model to obtain a gesture category probability distribution vector; modulating the encoded feature vector based on the gesture category probability distribution vector to generate a modulated feature vector; performing a one-dimensional convolution between the gesture category probability distribution vector and the encoded feature vector to obtain an overall probability distribution embedding feature vector; and adding the modulated feature vector and the overall probability distribution embedding feature vector together and then inputting them into the regression head of the multi-task learning neural network model to obtain the final parameter value.

[0069] Specifically, a multi-task learning neural network model can be constructed, which takes hand joint data as input and includes a shared encoder and independent classification and regression heads. Here, the hand posture input vector is obtained from the continuous gesture data stream. For the model's input data, i.e., the hand posture input vector, the hand posture input vector not only needs to contain the three-dimensional coordinates of the hand joints, but also, to ensure rotation invariance, should include at least one of the following features: normalized coordinates of all hand joints relative to the center of the palm, the distance between hand joints, and the angle between the skeletal vectors formed by the hand joints. In this way, when the user adjusts the grip size of the virtual weapon through continuous gestures, this operation can effectively capture detailed features such as the real-time changes in the distance between the thumb and index finger and the differences in the angle of finger bending, ensuring that the core information related to grip size adjustment in the original gesture data is completely preserved, providing reliable raw data support for subsequent parameter calculations.

[0070] Then, the hand gesture input vector is input to the shared encoder of the multi-task learning neural network model to generate an encoded feature vector. This enables feature reuse and abstract representation, providing a unified feature foundation for subsequent classification and regression tasks. For example, when a user makes a series of gestures to adjust the barrel length of a gun, the shared encoder can transform the original features such as the finger stretching range and wrist movement trajectory into unified encoded features that include the category attribute of the stretching gesture and the specific stretching amount. This allows the classification head and regression head to perform subsequent processing based on the same features, improving the model's processing efficiency and feature utilization.

[0071] Furthermore, for the encoded feature vector obtained by the encoder based on the hand pose input vector, for example denoted as... The encoded feature vector is input into the classification head of the multi-task learning neural network model to obtain the gesture category probability distribution vector, i.e. ,in Let be the probability distribution vector of the gesture category. Indicates the first The probability values ​​for each category are calculated. In other words, the network structure of the classification head calculates the probability distribution of the current gesture belonging to each preset gesture category. This provides clear category context guidance for subsequent parameter modulation, clarifying the primary operation type of the current gesture. For example, when the probability value of a pinch gesture is significantly higher than other categories, it can provide a clear category indication for adjusting the lens magnification, ensuring that the direction of subsequent parameter adjustments is consistent with the user's gesture intent.

[0072] Then, based on the gesture category probability distribution vector, the encoded feature vector is modulated to generate a modulated feature vector, i.e.: ;in, Represents the modulated feature vector. Represents the encoded feature vector. Indicates the first The probability values ​​of each category, Represents element-wise multiplication. This represents element-wise addition. Indicating in the index Summation operation on, For the first Each category of modulation hyperparameters will The values ​​are modulated to a positive-zero range to make the weighted feature values ​​meaningful. That is, the probabilities of each class output by the classification head are used. As a gating signal, the cross-entropy information of the class is used to constrain the encoding features through the inherent distribution consistency constraint. Each class probability corresponds to a parameter that drives the feature distribution change of the encoded feature vector, which is used to drive the distribution focusing of the encoded features based on the predetermined class. For example, for the class of pistol, the expected encoded feature distribution focuses on the thumb-index finger distance (barrel length), and for the class of sword, the expected encoded feature distribution focuses on the wrist pitch angle (blade angle), etc.

[0073] Furthermore, for the modulated vector The individual category probability modulation attribute is then used to perform a one-dimensional convolution between the gesture category probability distribution vector and the encoded feature vector to obtain the overall probability distribution embedding feature vector. ,in, This is a one-dimensional convolution operation. A feature vector is embedded into the overall probability distribution. Finally, the modulated feature vector and the embedded feature vector of the overall probability distribution are summed and then input into the regression head of the multi-task learning neural network model to obtain the final parameter values.

[0074] In other words, considering the continuity of human movements, a smooth transition from a relaxed gesture to an L-shaped gesture, followed by fine-tuning the distance between the thumb and forefinger, constitutes a coherent action. Independent, classification, and regression models disrupt this natural flow of movement and fail to reuse information. For example, if the gesture category is a large L-shaped gesture, both the L-shaped category information and the large parameter information are simultaneously encoded in the joint coordinates. Therefore, by coupling the classification probability distribution with the encoded features both locally and globally, one-step generation and fine-tuning can be achieved. For instance, a user can directly make a large L-shaped gesture, which simultaneously identifies the pistol category and sets its barrel length parameter to a large initial value. Furthermore, the user can control multiple parameters through subtle changes in the gesture; for example, while maintaining the L-shape, the weapon color can be controlled by adjusting the curvature of the middle finger. Moreover, due to the introduction of the category probability distribution, even if the user's gesture is not perfectly aligned, it will not directly fail to be recognized but will output a low-confidence probability distribution, further resulting in low activation weights for all parameters. For example, this can make the weapon model present a blurred outline, thus guiding the user to make a clearer gesture.

[0075] In summary, the MR battle interaction method based on multimodal perception fusion, as described in this application, is elucidated. First, it dynamically captures the user's generation intention in real space by collaboratively analyzing the user's eye-tracking focus and hand gestures. Then, it further integrates gaze direction and gesture commands, performing contextual sampling of the physical materials of the real environment that the user is interested in, seamlessly transferring the textures and feel of the real world to the attribute configuration of virtual objects. Based on this, a refined gesture-driven mechanism is used to instantiate the core framework of virtual objects and intuitively select modular components. Then, gaze guidance is used for modular locking, and finally, continuous gestures are used to fine-tune the parameters of the locked modules. In this way, the user's gaze and hands can be used as a unified multi-dimensional input channel, achieving a closed-loop, immersive generation experience from intention generation, material association, form construction to detailed configuration.

[0076] Figure 8 This is a block diagram of a multimodal perception fusion MR battle interaction system according to an embodiment of this application. Figure 8As shown, the MR combat interaction system 100 based on the multimodal perception fusion according to an embodiment of this application includes: an anchor point generation module 110, used to perform intent activation and generation space calibration based on eye-tracking data stream and hand posture data stream to obtain generation anchor points and switch the system state to generation mode; a material configuration module 120, used to perform environmental material context sampling based on eye-tracking data stream and hand posture data stream, combined with real environment 3D mesh data to obtain a material configuration file; a weapon primitive instantiation module 130, used to select and instantiate weapon primitives based on hand posture data stream, generation anchor points and material configuration files to obtain a basic virtual weapon model; a modular selection module 140, used to perform eye-tracking data stream-based gaze-guided modular selection of the basic virtual weapon model to obtain the selected module ID; and a parametric fine-tuning module 150, used to perform gesture-driven parametric fine-tuning of the selected module ID based on continuous gesture data stream to obtain a modified virtual weapon model.

[0077] As described above, the multimodal perception fusion MR battle interaction system 100 according to the embodiments of this application can be implemented in various wireless terminals, such as servers with multimodal perception fusion MR battle interaction algorithms. In one possible implementation, the multimodal perception fusion MR battle interaction system 100 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the multimodal perception fusion MR battle interaction system 100 can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the multimodal perception fusion MR battle interaction system 100 can also be one of many hardware modules of the wireless terminal.

[0078] Alternatively, in another example, the multimodal perception fusion MR battle interaction system 100 and the wireless terminal can also be separate devices, and the multimodal perception fusion MR battle interaction system 100 can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.

[0079] Here, those skilled in the art will understand that the specific operations of each step in the above-described multimodal perception fusion MR battle interaction system have been referenced above. Figures 1 to 7 The multimodal perception fusion MR battle interaction method has been described in detail in the previous section, and therefore, its repeated description will be omitted.

Claims

1. A multi-modal perception fusion-based MR battle interaction method, characterized in that, The method comprises the following steps: Based on the eye movement data stream and the hand gesture data stream, the intention is activated and the space calibration is generated to obtain the generation anchor point and switch the system state to the generation mode, including: the hand gesture data stream is detected to determine whether the first condition is met; the line of sight detection is performed on the eye movement data stream to determine whether the second condition is met; the system state is switched to the generation mode in response to the first condition and the second condition being met at the same time; the generation anchor point is created in the palm center of the left hand in response to the first condition and the second condition being met at the same time; wherein the first condition is that one hand presents a tray gesture and remains stable for more than a threshold time, and the second condition is that the user's line of sight stays in the palm center area of the tray gesture for more than the threshold time; Based on the eye movement data stream and the hand gesture data stream, and combined with the real environment three-dimensional grid data, the environment material context sampling is performed to obtain the material configuration file; Based on the hand gesture data stream, the generation anchor point and the material configuration file, the weapon primitive selection and instantiation are performed to obtain the basic virtual weapon model; Based on the eye movement data stream, the line of sight guided modular selection is performed on the basic virtual weapon model to obtain the selected module ID; Based on the continuous gesture data stream, the gesture driven parameterized fine tuning is performed on the selected module ID to obtain the modified virtual weapon model.

2. The MR counter-play interaction method of multi-modal perception fusion according to claim 1, characterized in that, Based on the eye movement data stream and the hand gesture data stream, and combined with the real environment three-dimensional grid data, the environment material context sampling is performed to obtain the material configuration file, comprising: Based on the eye movement data stream, a ray is emitted; Determine the intersection between the ray and the real environment three-dimensional grid data; Extract the texture information and color information near the intersection from the real environment three-dimensional grid data; In response to detecting that the hand gesture data stream is a sampling gesture, the texture information and color information near the intersection are packaged into the material configuration file.

3. The MR counter-play interaction method of multi-modal perception fusion according to claim 1, characterized in that, Based on the hand gesture data stream, the generation anchor point and the material configuration file, the weapon primitive selection and instantiation are performed to obtain the basic virtual weapon model, comprising: In response to the hand gesture data stream being a preset category gesture, a basic three-dimensional model corresponding to the preset category gesture is retrieved from a virtual weapon library; The material configuration file is applied to the basic three-dimensional model to obtain configured model data; The configured model data is rendered above the generation anchor point to obtain the basic virtual weapon model.

4. The MR counter-play interaction method of multi-modal perception fusion according to claim 3, characterized in that, The configured model data is rendered above the generation anchor point to obtain the basic virtual weapon model, comprising: The configured model data is input into a rendering engine to obtain a basic virtual weapon model instance; A three-dimensional transformation matrix of the generation anchor point is obtained; The three-dimensional transformation matrix is set as the world coordinate transformation of the basic virtual weapon model instance; The basic virtual weapon model instance is added to a scene rendering queue.

5. The MR counter-play interaction method of multi-modal perception fusion according to claim 1, characterized in that, Based on the continuous gesture data stream, the gesture driven parameterized fine tuning is performed on the selected module ID to obtain the modified virtual weapon model, comprising: The selected module ID is used as a query key to retrieve a preset module control parameter mapping database to obtain an active control mapping; extracting, filtering and normalizing the continuous gesture data stream to obtain final parameter values; based on the final parameter values and the activity control mapping, performing parameter-driven model real-time deformation and rendering update on the base virtual weapon model to obtain the modified virtual weapon model.

6. A multi-modal perceptual fusion MR interactive combat system, characterized in that, comprise: an anchor generation module for generating an anchor point and switching the system state to a generation mode based on an eye movement data stream and a hand pose data stream, comprising: performing gesture detection on the hand pose data stream to determine whether a first condition is met; performing gaze detection on the eye movement data stream to determine whether a second condition is met; in response to the first condition and the second condition being met simultaneously, switching the system state to the generation mode; in response to the first condition and the second condition being met simultaneously, creating the anchor point at the center of the palm of the left hand; wherein the first condition is that one hand presents a tray gesture and remains stable for more than a threshold time, and the second condition is that the user's gaze stays in the palm center area of the tray gesture for more than the threshold time; a material configuration module for sampling the context of the environment material based on the eye movement data stream and the hand pose data stream, and combining real environment three-dimensional grid data to obtain a material configuration file; a weapon primitive instantiation module for selecting and instantiating a weapon primitive based on the hand pose data stream, the anchor point and the material configuration file to obtain a base virtual weapon model; a modular selection module for performing line-of-sight guided modular selection on the base virtual weapon model based on the eye movement data stream to obtain a selected module ID; a parameterized fine-tuning module for performing gesture-driven parameterized fine-tuning on the selected module ID based on the continuous gesture data stream to obtain a modified virtual weapon model.

Citation Information

Patent Citations

  • Action triggering interaction method for virtual object in real person and MR virtual environment

    CN114327072A

  • VR large-space positioning interaction system based on multi-modal perception

    CN120469587A