XR Virtual-Real Interactive Control Device with Optical Motion Capture and Inertial Navigation Fusion

By combining a metasurface polarization coding marker with a micro inertial measurement unit, optical motion capture and inertial navigation fusion is achieved, solving the problems of marker identity confusion and inertial measurement drift in traditional optical tracking systems, and improving the accuracy and robustness of XR virtual-real interaction.

CN121704706BActive Publication Date: 2026-04-21SICHUAN WUTONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN WUTONG TECH CO LTD
Filing Date
2026-02-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Traditional optical tracking systems suffer from marker confusion and trajectory jumps, while inertial measurement integral drift leads to attitude estimation deviations, and visual-inertial fusion systems experience performance degradation in sparse texture environments.

Method used

A combination of a metasurface polarization coding marker array and a micro inertial measurement unit array is used. A hardware synchronization controller is used to synchronize polarization decoding and inertial data to construct a visual-inertial joint optimization map for six-degree-of-freedom pose fusion estimation.

Benefits of technology

Achieving stable tracking in occluded and fast-moving scenarios improves the accuracy, robustness, and real-time response of XR virtual-real interaction, solves the problems of marker identity confusion and trajectory jump, and reduces computational load and system latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121704706B_ABST
    Figure CN121704706B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of virtual reality technology, specifically relating to an XR virtual-real interaction control device that integrates optical motion capture and inertial navigation. It includes: a metasurface polarization-encoded marker array, distributed and installed at preset tracking point positions on the user's hand and head, used to generate reflected light carrying polarization identification codes when receiving illumination light; a miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, used to collect acceleration and angular velocity data at each tracking point; and a near-eye display device, including an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head areas, and the polarization camera module receives the reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization images. This invention achieves stable tracking in occlusion and rapid motion scenarios through tightly coupled visual-inertial fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual reality technology, specifically relating to an XR virtual-real interaction control device that integrates optical motion capture and inertial navigation. Background Technology

[0002] Optical tracking schemes estimate pose by observing markers on the user's body or directly recognizing the human silhouette using a camera. Active optical tracking systems deploy infrared LEDs as markers on the user's surface; the camera detects the image position of the LEDs and calculates the three-dimensional coordinates of the markers using triangulation. Passive optical tracking systems use retroreflective spheres or patches as markers; these markers do not emit light themselves but reflect illumination back towards the camera. Regardless of whether the scheme is active or passive, the identification of markers relies on their geometric arrangement in space, i.e., identification is based on the relative positional patterns of a group of markers. This geometric pattern-based identification method has significant drawbacks: when multiple markers overlap or their geometric configuration changes during movement, identification can easily become confused or lost, leading to tracking interruptions or trajectory jumps. Furthermore, traditional optical markers only provide intensity information, offering limited information dimensionality and making it difficult to accommodate rich coding requirements.

[0003] Inertial tracking schemes measure motion acceleration and angular velocity using microelectromechanical accelerometers and gyroscopes, and obtain position and attitude estimates through integration. Inertial measurement units (IMUs) offer advantages such as high sampling frequency, immunity to obstruction, and no need for external infrastructure, making them particularly suitable for capturing transient details of fast-moving motion. However, an inherent drawback of inertial measurement is integration drift. The zero-bias error of the gyroscope, after time integration, causes the attitude estimate to continuously deviate from the true value, while the zero-bias error of the accelerometer, after quadratic integration, causes the position estimate to diverge with a quadratic velocity. Without external correction information, pure inertial tracking accumulates unacceptable errors within seconds to tens of seconds.

[0004] To overcome the limitations of single sensors, visual-inertial fusion schemes combine optical observation with inertial measurement. Optical observation provides an absolute position reference to eliminate inertial integration drift, while inertial measurement provides high-frequency motion information to fill in the gaps between optical frames and enhance robustness to occlusion. Existing visual-inertial fusion systems mostly use natural feature point tracking or artificially marked point tracking as the visual front end. Natural feature point tracking relies on the texture structure of corners, edges, etc., in the scene, and its performance degrades significantly in environments with sparse or repetitive textures. Although artificially marked point tracking can provide stable and reliable visual observation, it faces the problem of difficult identification, as mentioned earlier. Summary of the Invention

[0005] Therefore, the main objective of this invention is to provide an XR virtual-real interaction control device that integrates optical motion capture and inertial navigation. It solves the problem of marker point identity confusion in traditional optical tracking by utilizing the intrinsic physical properties of metasurface polarization encoding, eliminates the delay error of software timestamps through hardware-level time synchronization, achieves stable tracking in occlusion and fast-moving scenarios through visual-inertial tight coupling fusion, and reduces computational load and system latency by directly determining gesture state using geometric features, thereby improving the accuracy, robustness and real-time response capability of XR virtual-real interaction.

[0006] The technical solution adopted in this invention is as follows:

[0007] The XR virtual-real interaction control device integrating optical motion capture and inertial navigation includes: a metasurface polarization-encoded marker array, distributed and installed at preset tracking point positions on the user's hand and head, used to generate reflected light carrying polarization identification codes when receiving illumination light; a miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, used to collect acceleration and angular velocity data at each tracking point; and a near-eye display device, including an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head area, and the polarization camera module receives the reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization... The system includes: an image processing module; a hardware synchronization controller for synchronously triggering data acquisition from the polarization camera module and the micro inertial measurement unit array; a polarization decoding module for performing polarization state analysis and marker identification on multi-channel polarization images, outputting the image coordinates and identification of each metasurface polarization encoding marker; a visual-inertial fusion positioning module for calculating the six-degree-of-freedom pose fusion estimation results of the user's hand and head based on the image coordinates and identification output by the polarization decoding module and the acceleration and angular velocity data acquired by the micro inertial measurement unit array; and an interactive control command generation module for parsing the user's interaction intent and generating XR interactive control commands based on the six-degree-of-freedom pose fusion estimation results.

[0008] Furthermore, each metasurface polarization coding marker in the metasurface polarization coding marker array includes a transparent substrate and a metal nanoantenna array covering the surface of the transparent substrate. The major axis orientation angle of each nanoantenna unit in the metal nanoantenna array is distributed according to a preset spatial arrangement rule, so that when linearly polarized illumination light is incident, each nanoantenna unit generates anisotropic phase delay and amplitude modulation of the incident light, thereby making the reflected light present an elliptical polarization state as the polarization identity code of each metasurface polarization coding marker.

[0009] Furthermore, each metasurface polarization coding marker in the metasurface polarization coding marker array is distributed and installed on the back of the user's hand, each finger joint, and the frontal and temporal sides of the head; each micro inertial measurement unit in the micro inertial measurement unit array includes a triaxial microelectromechanical accelerometer and a triaxial microelectromechanical gyroscope.

[0010] Furthermore, the illumination module is a ring-shaped near-infrared light-emitting diode illumination array, used to generate linearly polarized illumination light with a wavelength of 850 nanometers; the polarization camera module includes a focal plane polarization sensor, on which a micro polarizer array is integrated. The micro polarizer array is arranged periodically in 2×2 pixel units, and the four pixels in each 2×2 pixel unit respectively cover the micro polarizers along the light transmission axis in the 0-degree, 45-degree, 90-degree, and 135-degree directions.

[0011] Furthermore, the hardware synchronization controller generates periodic synchronization pulse signals, which simultaneously trigger the exposure start of the polarization camera module and the data latching of all micro inertial measurement units in the micro inertial measurement unit array, so that the polarization image frame and the inertial measurement data frame have the same acquisition time mark.

[0012] Furthermore, the polarization decoding module performs polarization state analysis as follows: for each 2×2 pixel unit in the multi-channel polarization image, it reads the gray values ​​of the 0-degree channel, 45-degree channel, 90-degree channel, and 135-degree channel pixels within the 2×2 pixel unit; it adds the gray values ​​of the 0-degree channel and 90-degree channel pixels to obtain a first intensity sum; it subtracts the gray values ​​of the 0-degree channel and 90-degree channel pixels to obtain a first intensity difference; it subtracts the gray values ​​of the 45-degree channel and 135-degree channel pixels to obtain a second intensity difference; it divides the first intensity difference by the first intensity sum to obtain a first normalized polarization component; it divides the second intensity difference by the first intensity sum to obtain a second normalized polarization component; and it adds the square of the first normalized polarization component and the square of the second normalized polarization component to obtain the linear polarization degree value.

[0013] Furthermore, the polarization decoding module employs the Bongaley spherical projection matching method when performing marker identification. This includes: traversing the linear polarization degree values ​​of all 2×2 pixel units, retaining pixel units with linear polarization degree values ​​higher than 0.6 as candidate polarization response regions, performing octet analysis on the candidate polarization response regions and merging them into polarization response patches, calculating the centroid position of each polarization response patch as its image coordinates; for each polarization response patch, taking the mean of the first normalized polarization component of all pixel units within the polarization response patch as the horizontal polarization feature value, taking the mean of the second normalized polarization component of all pixel units within the polarization response patch as the diagonal polarization feature value, and using the horizontal polarization feature value and the diagonal polarization feature value as the feature point coordinates on the Bongaley spherical equatorial circle; and then... The equatorial circle of the Bongare spherical surface is divided into 36 angular sectors, each covering a central angle range of 10 degrees. Standard feature points from a pre-stored polarization identity coding library are categorized according to their respective angular sectors to form a sector index table. The polar angle of the feature points of the current polarization response patch relative to the center of the Bongare spherical equatorial circle is calculated, and the target angular sector into which the polar angle falls is determined. All standard feature points within the target angular sector and its two adjacent angular sectors are extracted from the sector index table as a candidate matching set. The arc length distance between the feature points of the current polarization response patch and each standard feature point in the candidate matching set on the Bongare spherical equatorial circle is calculated. The metasurface polarization coding marker identity identifier corresponding to the standard feature point with the smallest arc length distance is selected as the identification result of the current polarization response patch.

[0014] Furthermore, the visual-inertial fusion positioning module performs inertial pre-integration on the acceleration and angular velocity data acquired by the micro inertial measurement unit array between adjacent frames to obtain the relative attitude change and relative velocity change between adjacent frames. The visual-inertial fusion positioning module constructs a visual-inertial joint optimization graph structure, which includes pose nodes and marker nodes. The pose nodes represent the position and attitude of the user's hand or head in the world coordinate system at each frame acquisition time, and the marker nodes represent the three-dimensional position of each metasurface polarization coding marker in the world coordinate system. The visual-inertial joint optimization graph structure includes inertial constraint edges and visual constraint edges. The inertial constraint edges connect the pose nodes of adjacent frames and carry the relative attitude change and relative velocity change, while the visual constraint edges connect the pose nodes and marker nodes and carry the image coordinates output by the polarization decoding module. The visual-inertial fusion positioning module performs iterative optimization on the visual-inertial joint optimization graph structure to obtain the six-degree-of-freedom pose fusion estimation result.

[0015] Furthermore, the interactive control command generation module calculates the hand movement direction vector and movement speed scalar based on the three-dimensional position changes of each metasurface polarization coding marker on the hand in multiple consecutive frames. The interactive control command generation module determines the current gesture state based on the relative position changes of the metasurface polarization coding marker corresponding to each fingertip relative to the metasurface polarization coding marker corresponding to the back of the hand. The current gesture state includes open state, clenched fist state, pinched state, and pointing state.

[0016] Furthermore, the interaction control command generation module calculates the midpoint position of the user's eyes based on the head pose in the six-degree-of-freedom pose fusion estimation result, and emits a gaze ray based on the gaze direction. It then performs an intersection test with the bounding boxes of each interactive virtual object in the virtual scene, and determines the interactive virtual object whose intersection point is closest to the midpoint of the user's eyes as the gaze focus object. The interaction control command generation module generates XR interaction control commands based on the current gesture state and the gaze focus object. The XR interaction control commands include object selection commands, object drag commands, object release commands, and menu call commands.

[0017] By employing the above technical solutions, this invention achieves the following beneficial effects: This invention uses metasurface polarization coding technology to replace the traditional geometric pattern recognition method. Each metasurface polarization coding marker generates a unique elliptic polarization output through the differentiated design of its metal nanoantenna array. This polarization output is an intrinsic physical property of the marker rather than an external spatial arrangement relationship. Therefore, even when multiple markers overlap or their relative positions change, each marker can still be accurately identified through its polarization characteristics, effectively solving the problems of marker point confusion and trajectory jumps in traditional optical tracking systems. The matching method based on Bongale spherical projection reduces high-dimensional polarization features to a one-dimensional polar angle space for rapid retrieval. Combined with the pre-classification strategy of the sector index table, this significantly improves the computational efficiency of identity recognition.

[0018] This invention uses a hardware synchronization controller to generate periodic synchronization pulse signals, which simultaneously trigger the exposure start of the polarization camera module and the data latching of the micro inertial measurement unit array. This eliminates the inherent scheduling delay and jitter error of the software timestamp scheme at the hardware level, enabling the polarization image frame and the inertial measurement data frame to be precisely aligned on the time axis, laying a high-quality data foundation for subsequent visual-inertial fusion positioning.

[0019] This invention constructs a visual-inertial joint optimization graph structure, which unifies the image coordinate observation output by polarization decoding and the relative motion constraints obtained by inertial pre-integration as constraint edges in the factor graph. Through iterative optimization, the pose state of each frame and the three-dimensional position of each marker point are estimated simultaneously. This fully leverages the absolute positioning capability of visual observation and the high-frequency motion capture capability of inertial measurement, achieving complementary advantages of the two information sources. Even in scenarios where the marker is partially occluded or the image is blurred due to rapid movement, it can still maintain stable and reliable tracking performance. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the nano-antenna array structure and polarization modulation principle of the metasurface polarization coding marker provided in an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of the micro-polarizer array structure of the focal plane polarization sensor provided in an embodiment of the present invention;

[0022] Figure 3 A comparison curve of visual-inertial fusion positioning performance provided for embodiments of the present invention;

[0023] Figure 4 The geometric feature parameters of gesture recognition and the state recognition curve provided in the embodiments of the present invention. Detailed Implementation

[0024] The XR virtual-real interaction control device integrating optical motion capture and inertial navigation includes: a metasurface polarization-encoded marker array, distributed and installed at preset tracking point positions on the user's hand and head, used to generate reflected light carrying polarization identification codes when receiving illumination light; a miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, used to collect acceleration and angular velocity data at each tracking point; and a near-eye display device, including an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head area, and the polarization camera module receives the reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization... The system includes: an image processing module; a hardware synchronization controller for synchronously triggering data acquisition from the polarization camera module and the micro inertial measurement unit array; a polarization decoding module for performing polarization state analysis and marker identification on multi-channel polarization images, outputting the image coordinates and identification of each metasurface polarization encoding marker; a visual-inertial fusion positioning module for calculating the six-degree-of-freedom pose fusion estimation results of the user's hand and head based on the image coordinates and identification output by the polarization decoding module and the acceleration and angular velocity data acquired by the micro inertial measurement unit array; and an interactive control command generation module for parsing the user's interaction intent and generating XR interactive control commands based on the six-degree-of-freedom pose fusion estimation results.

[0025] The construction of metasurface polarization-encoded marker arrays is based on the precise control of electromagnetic wave polarization states by optical metasurfaces. Each metasurface polarization-encoded marker consists of two parts: a transparent substrate and a metal nanoantenna array covering the surface of the transparent substrate. The transparent substrate is made of fused silica glass or sapphire, which has excellent light transmittance in the near-infrared band, stable refractive index, and low coefficient of thermal expansion. The thickness of the transparent substrate is typically set between 0.5 mm and 1 mm to ensure sufficient mechanical strength to withstand bending stress during wear without increasing the user's burden due to excessive thickness. Before the metal nanoantenna array is deposited, the surface of the transparent substrate undergoes chemical mechanical polishing to reduce the surface roughness to below 0.5 nm, thereby avoiding the influence of substrate surface defects on the subsequent nanostructure processing accuracy.

[0026] Metal nanoantenna arrays are fabricated on a transparent substrate using electron beam lithography and metal lift-off processes. Specifically, a 200 nm thick layer of polymethyl methacrylate (PMMA) positive photoresist is first spin-coated onto the transparent substrate. Then, an electron beam exposure system is used to expose the nanoantenna units point-by-point according to a pre-designed pattern. After exposure, development is performed to form a patterned photoresist mask of the nanoantenna units. Next, an electron beam evaporation process is used to deposit a 50 nm thick gold film across the entire substrate surface. Gold exhibits a strong plasmon resonance response at a wavelength of 850 nm, effectively enhancing the interaction between light and the nanostructure. Finally, the metal layer in the masked area is lifted off by immersion in acetone solution, preserving the metal nanoantenna array structure complementary to the photoresist pattern.

[0027] Each nanoantenna element is a rectangular rod with a major axis dimension between 180 nm and 220 nm and a minor axis dimension between 60 nm and 80 nm. This size range allows the nanoantenna elements to generate localized surface plasmon resonances under 850 nm wavelength illumination, and the resonance enhancement effect significantly improves the modulation capability of the nanoantenna elements to the polarization state of the incident light. When linearly polarized light is incident on the nanoantenna element, the electric field components along the major axis and the electric field components along the minor axis respectively excite plasmon oscillations in different modes. Due to the difference in resonance frequencies in the two directions, a phase delay occurs between the two orthogonally polarized components in the reflected light. The phase delay depends on the ratio of the major and minor axes of the nanoantenna element and the wavelength of the incident light. Simultaneously, the amplitudes of the two orthogonally polarized components exhibit different degrees of attenuation due to the anisotropic absorption characteristics. The combined effect of the phase delay and amplitude difference causes the originally linearly polarized incident light to become elliptically polarized light after reflection.

[0028] refer to Figure 1The lower part of the image shows the stacked structure of a metasurface polarization-encoded marker, consisting of a transparent substrate and a metal nanoantenna array covering the surface of the substrate. The transparent substrate, appearing as a gray rectangular area, is made of fused silica glass or sapphire, exhibiting excellent light transmittance in the near-infrared band. The metal nanoantenna array, located above the transparent substrate, is composed of multiple rectangular rod-shaped nanoantenna elements arranged periodically. The image clearly shows nine nanoantenna elements, each a slender, golden rectangle with a different long-axis orientation angle. From left to right, the long-axis orientation angles of the nanoantenna elements increase sequentially; three representative angle values ​​are marked in the image: 0 degrees, 45 degrees, and 90 degrees. This differentiated long-axis orientation angle arrangement is key to achieving polarization encoding. Different metasurface polarization-encoded markers form their own unique polarization characteristics, distinguishing them from other markers, through the unique combination of the long-axis orientation angles of their nanoantenna elements. The figure also indicates that the typical size range of the nanoantenna unit is 180 to 220 nanometers. This size allows the nanoantenna unit to produce a localized surface plasmon resonance effect when illuminated by light with a wavelength of 850 nanometers. Figure 1 The upper part illustrates the optical process of polarization modulation.

[0029] A beam of linearly polarized light is incident on the surface of a metal nanoantenna array from the upper left of the diagram, pointing diagonally downwards. The incident beam is indicated by a blue arrow, whose direction indicates the propagation direction of the light. Several horizontal bidirectional arrows are drawn beside the incident beam to indicate that the electric field vibration direction of the incident light is horizontally linearly polarized. The wavelength of the incident light is marked as 850 nm, corresponding to the near-infrared band. When linearly polarized illumination light is incident on the metal nanoantenna array, each nanoantenna element produces different degrees of phase delay and amplitude modulation for the two orthogonal polarization components of the incident light. This anisotropic optical response originates from the different plasmon resonance characteristics of the nanoantenna elements along their long and short axes. After polarization modulation by the metal nanoantenna array, the polarization state of the reflected light changes from linear polarization to elliptic polarization, indicated by a red arrow, and the reflected beam exits from the surface of the metal nanoantenna array diagonally upwards. Figure 1 The invention demonstrates, in an intuitive graphical manner, the physical process by which a metasurface polarization-encoded marker converts incident linearly polarized light into reflected light with a specific elliptic polarization state. This polarization state conversion process forms the physical basis for the marker's identity encoding in this invention.

[0030] Different metasurface polarization-encoded markers form their unique polarization identity codes through differentiated long-axis orientation angles of their nanoantenna elements. Within a single metasurface polarization-encoded marker, the metal nanoantenna array is divided into several sub-regions. Within each sub-region, the long-axis orientation angles of the nanoantenna elements remain consistent, while the long-axis orientation angles between different sub-regions increase or decrease according to specific rules. For example, the effective optical area of ​​a single metasurface polarization-encoded marker can be divided into 16 sub-regions (4×4). The long-axis orientation angles of the nanoantenna elements in each sub-region can be set to a sequence of combinations from 0 degrees, 10 degrees, 20 degrees up to 150 degrees. Different combinations correspond to different polarization identity codes. When illumination light shines on the entire metasurface polarization-encoded marker, the elliptically polarized reflected light generated by each sub-region superimposes to form a composite light field with a specific spatial polarization distribution. The polarization characteristics of this composite light field are the unique identifier that distinguishes this metasurface polarization-encoded marker from other markers.

[0031] The placement of the metasurface polarization-encoded marker array is ergonomically optimized to ensure comprehensive coverage of the key degrees of freedom of hand and head movement. In the hand region, metasurface polarization-encoded markers are fixed at the center of the back of the hand, the metacarpophalangeal joint of the thumb, the interphalangeal joint of the thumb, the metacarpophalangeal joint of the index finger, the proximal interphalangeal joint of the index finger, the distal interphalangeal joint of the index finger, and the corresponding joints of the middle, ring, and little fingers, totaling 15 to 21 tracking points. The metasurface polarization-encoded marker at the center of the back of the hand characterizes the overall position and posture of the hand, while the markers at each finger joint reconstruct the bending angle and opening / closing state of the fingers. In the head region, metasurface polarization-encoded markers are fixed at three locations: the center of the forehead, the left temple, and the right temple. The rigid geometric relationship formed by these three points uniquely determines the six degrees of freedom pose of the head in three-dimensional space. The metasurface polarization coding marker is fixed to the skin or the surface of the wearable fabric using a medical-grade silicone adhesive patch or an adjustable elastic band. The thickness of the silicone adhesive patch is controlled to be within 0.3 mm to reduce interference with the user's tactile sense.

[0032] In one alternative embodiment, the transparent substrate of the metasurface polarization coding marker uses a flexible polyimide film instead of rigid quartz glass. The thickness of the polyimide film can be reduced to 50 micrometers, allowing the metasurface polarization coding marker to adhere to skin surfaces with high curvature, such as finger joints. The metal nanoantenna array on the flexible substrate is mass-produced using nanoimprint lithography, which significantly reduces manufacturing costs and increases production capacity compared to electron beam lithography. The nanoimprint process first etches a groove pattern complementary to the target nanoantenna array onto the surface of a silicon mold. Then, a flexible substrate coated with UV-curable resin is pressed onto the silicon mold surface. After UV irradiation to cure the resin, the mold is demolded. Finally, a metal film is sputtered onto the resin pattern surface to form the nanoantenna structure.

[0033] In another alternative implementation, the material of the metal nanoantenna array is replaced with silver or aluminum instead of gold. Silver has lower plasmon resonance loss than gold in the visible to near-infrared band, resulting in higher polarization modulation efficiency. However, silver has poor chemical stability, requiring a 5- to 10-nanometer-thick silicon dioxide protective film to prevent oxidation. Aluminum is significantly cheaper than both gold and silver, and its plasmon resonance wavelength extends to the ultraviolet band, providing design flexibility for systems using shorter wavelength illumination.

[0034] Each micro inertial measurement unit (IMU) in the micro inertial measurement unit array is physically adjacent to its corresponding metasurface polarization coding marker and fixed by a rigid connector. The rigid connector is made of carbon fiber reinforced plastic or titanium alloy with an elastic modulus higher than 100 gigapascals, ensuring that the metasurface polarization coding marker and the micro inertial measurement unit do not undergo relative displacement or rotation during movement, thus guaranteeing a strict spatial correspondence between their output pose information. The mass of the rigid connector is controlled to within 0.5 grams to avoid affecting the user's natural hand movements due to excessive added mass.

[0035] Each miniature inertial measurement unit (MMU) integrates a triaxial MEMS accelerometer and a triaxial MEMS gyroscope. The triaxial MEMS accelerometer operates based on the principle of capacitive differential sensing, containing a mass block micro-machined from silicon, suspended from a fixed frame by a flexible cantilever beam. When the MMU accelerates, the mass block displaces relative to the fixed frame due to inertia. This displacement alters the capacitance gap between the mass block and adjacent fixed electrodes, resulting in a change in the differential capacitance value. The detection circuit converts this change in differential capacitance into a voltage output proportional to the acceleration. The triaxial MEMS accelerometer is set to ±16 times the acceleration due to gravity, meeting the requirements for measuring peak acceleration generated during rapid hand movements. Its noise density is below 100 microgravity RHz, enabling it to distinguish weak acceleration signals generated by minute finger tremors.

[0036] A triaxial MEMS gyroscope detects angular velocity based on the Coriolis effect. The driving mass inside the gyroscope vibrates at a fixed frequency along one axis under electrostatic actuation. When the gyroscope rotates around a sensing axis perpendicular to the driving axis, the driving mass experiences a Coriolis force, generating a vibration response along a third orthogonal axis. The amplitude of this vibration response is proportional to the input angular velocity. By arranging the driving mass and the detection structure along the three mutually orthogonal sensing axes, the angular velocity components around all three axes can be measured simultaneously. The triaxial MEMS gyroscope has a range of ±2000 degrees per second, covering the range of angular velocities during rapid head rotation. Its angular random walk coefficient is less than 0.005 degrees per square root of second, maintaining high attitude estimation accuracy even after short-time integration.

[0037] The miniature inertial measurement unit (MMU) also integrates a temperature sensor and digital signal processing circuitry. The temperature sensor monitors the operating temperature of the microelectromechanical (MEMS) sensitive structure in real time, while the digital signal processing circuitry performs online calibration of the raw outputs of the accelerometer and gyroscope based on a pre-calibrated temperature compensation coefficient, eliminating zero-bias drift and sensitivity variations caused by changes in ambient temperature. The miniature MMU's package size is controlled within 3 mm × 3 mm × 1 mm, and the power consumption of a single unit is less than 5 milliwatts, making it suitable for battery-powered wearable applications. The miniature MMU communicates with external systems via a serial peripheral interface or integrated circuit bus, with a configurable data output rate between 100 Hz and 1000 Hz. The higher output rate enables the capture of transient details of rapid hand movements.

[0038] In one alternative implementation, the micro inertial measurement units (IMUs) in the micro IMU array are interconnected in series via a flexible printed circuit board (PCB). The PCB uses a 25-micrometer-thick polyimide substrate, on which 0.1-millimeter-wide copper wires are printed for transmitting power and data signals. The flexible PCB can freely deform to accommodate finger bends, without mechanically restricting the range of motion of the hand joints. Each micro IMU shares the same data bus and is addressed via pre-assigned device addresses, reducing the number of signal lines and simplifying the wiring complexity of the wearable device.

[0039] In another alternative implementation, the miniature inertial measurement unit integrates a low-power Bluetooth wireless communication chip. Each miniature inertial measurement unit is independently packaged and powered by a built-in miniature lithium polymer battery. Wireless communication completely eliminates wired connections between wearable components, further enhancing the user's freedom of movement. The Bluetooth chip supports high-speed data transmission with connection intervals as low as 7.5 milliseconds, meeting the bandwidth requirements for real-time inertial data transmission. The miniature lithium polymer battery has a capacity of 20 mAh and can operate continuously for over 8 hours at a 100 Hz data output rate.

[0040] The illumination module and polarization camera module installed on the near-eye display device constitute the optical observation system of the metasurface polarization-encoded marker array. The illumination module adopts a ring-shaped near-infrared light-emitting diode (NIR) illumination array, which consists of 8 to 16 NIR LEDs evenly distributed along the edge of the near-eye display device's frame. The center emission wavelength of each NIR LED is 850 nm, and the full width at half maximum (FWHM) of the spectrum is controlled within 30 nm. The narrow-band emission characteristics ensure that the illumination light energy is concentrated near the designed operating wavelength of the metasurface polarization-encoded marker, thereby maximizing the polarization modulation efficiency.

[0041] Each near-infrared LED is equipped with a collimating lens and a linear polarizer in front of its light-emitting surface. The collimating lens converges the divergent beam emitted by the LED chip into a near-collimated beam with a half-divergence angle of approximately 15 degrees. The transmission axis of the linear polarizer is uniformly set to the horizontal direction, ensuring that the illumination light emitted by all near-infrared LEDs has the same linear polarization state. After the horizontally linearly polarized illumination light illuminates the surface of the metasurface polarization-encoded marker, it is converted into elliptically polarized reflected light carrying a polarization identity code through an anisotropic polarization modulation by the metal nanoantenna array. The total optical power output of the illumination module is set between 50 milliwatts and 200 milliwatts. Within a typical interaction distance of 0.3 meters to 1 meter between the user's hand and the near-eye display device, the illuminance on the surface of the metasurface polarization-encoded marker can reach 100 lux to 500 lux, which is sufficient to form a high signal-to-noise ratio image within the short exposure time of the polarization camera module.

[0042] In one alternative implementation, the illumination module uses a vertical-cavity surface-emitting laser (VCSEL) array instead of a near-infrared light-emitting diode (LED) array. The spectral linewidth of a VCSEL is only 0.5 to 1 nanometer, much narrower than that of an LED, which further improves the polarization response consistency of the metasurface polarization-encoded marker. The emitting area of ​​the VCSEL is less than 100 square micrometers, and when combined with a microlens array, a highly collimated illumination beam can be achieved, reducing ambient stray light interference caused by illumination beam divergence.

[0043] The core component of the polarization camera module is a focal plane polarization sensor. The photosensitive surface of the focal plane polarization sensor is composed of a pixel array with a resolution of 1280×1024 or higher and a pixel pitch of 5 to 10 micrometers. A micro-polarizer array is directly integrated above the photosensitive surface. The micro-polarizer array is strictly aligned with the pixel array, with each micro-polarizer unit covering exactly one pixel. The micro-polarizer array is arranged periodically in 2×2 pixel units. Within each 2×2 pixel unit, the micro-polarizer covered by the top-left pixel has a transmission axis along the 0-degree direction, the micro-polarizer covered by the top-right pixel has a transmission axis along the 45-degree direction, the micro-polarizer covered by the bottom-left pixel has a transmission axis along the 90-degree direction, and the micro-polarizer covered by the bottom-right pixel has a transmission axis along the 135-degree direction.

[0044] refer to Figure 2The left side of the image shows a micro-polarizer array region consisting of 16 pixels arranged in a 4x4 grid. This array region is composed of four 2x2 pixel units. Each pixel is represented by a square block, and pixels with different polarization directions are distinguished by different colors. The legend in the upper right corner of the image illustrates the correspondence between color and polarization channel: red corresponds to the 0-degree channel, cyan to the 45-degree channel, blue to the 90-degree channel, and green to the 135-degree channel. A short line segment is drawn inside each pixel block, and the orientation of the line segment indicates the direction of the light transmission axis of the micro-polarizer covering that pixel. Observing the array region on the left side, it can be seen that the micro-polarizer array is arranged periodically with 2x2 pixels as the basic repeating unit. Within each 2x2 pixel unit, the top left pixel covers the micro-polarizer along the 0-degree light transmission axis, the top right pixel covers the micro-polarizer along the 45-degree light transmission axis, the bottom left pixel covers the micro-polarizer along the 90-degree light transmission axis, and the bottom right pixel covers the micro-polarizer along the 135-degree light transmission axis.

[0045] A bracket label is also drawn on the left side of the image to indicate the range of a single 2x2 pixel unit. Figure 2 The lower right corner displays a magnified view of a single 2x2 pixel unit. The four pixel squares are presented in a larger size, with the corresponding polarization angle values ​​directly labeled inside each square: 0 degrees, 45 degrees, 90 degrees, and 135 degrees. Below the magnified view are also the grayscale value symbols for each pixel. , , and These four grayscale values ​​are the raw input data for subsequent polarization state analysis calculations. A gray curve with an arrow connects the left array area to the magnified view in the lower right corner, indicating the correspondence between the two. Figure 2 The labels above indicate the names of the micro-polarizer array and the focal plane polarization sensor, meaning the micro-polarizer array is directly integrated onto the photosensitive surface of the focal plane polarization sensor. Through this micro-polarizer array structure, the focal plane polarization sensor can simultaneously acquire image information from four polarization channels in a single exposure, providing a hardware foundation for real-time polarization state analysis and metasurface polarization coding marker identification.

[0046] When elliptically polarized reflected light is incident on a focal plane polarization sensor, each pixel only receives the polarization component passing through the micro-polarizer above it. Because the projection intensity of elliptically polarized light differs in different polarization directions, the grayscale values ​​output by the four pixels within the same 2×2 pixel unit differ. Assume the amplitudes of the electric field components of the incident elliptically polarized light in the 0-degree and 90-degree directions are respectively... and Furthermore, there is a phase difference between the two components. Then the theoretical response grayscale value of each polarization channel pixel can be expressed as: , , , ,in The grayscale value of the 0-degree channel pixel. This represents the grayscale value of the 45-degree channel pixels. This represents the grayscale value of a 90-degree channel pixel. This represents the grayscale value of a 135-degree channel pixel. The split-plane polarization sensor simultaneously acquires images from four polarization channels in a single exposure. Compared to the traditional time-division acquisition method using rotating polarizers, the split-plane polarization sensor eliminates the impact of motion blur on polarization measurement accuracy, making it particularly suitable for real-time tracking applications in dynamic scenes.

[0047] An imaging lens and a near-infrared bandpass filter are mounted in front of the focal plane polarization sensor. The imaging lens has a focal length set between 6 mm and 12 mm and a field of view covering 90 degrees to 120 degrees, enabling simultaneous observation of all metasurface polarization-coded markers located in the user's hand and head areas within a single image. The center transmission wavelength of the near-infrared bandpass filter matches the emission wavelength of the illumination module, with a passband width of 40 nm to 60 nm and an out-of-band rejection ratio higher than optical density 4. This effectively blocks ambient light interference in the visible light band, ensuring that the focal plane polarization sensor responds only to the near-infrared light emitted by the illumination module.

[0048] The imaging frame rate of the polarization camera module is set between 60 and 120 frames per second. A higher frame rate reduces the displacement of the metasurface polarization coded markers between adjacent frames, thus alleviating the computational burden of marker association matching in subsequent tracking algorithms. The exposure time of the focal plane polarization sensor is adaptively adjusted according to the illumination intensity and scene dynamic range, typically ranging from 1 to 5 milliseconds. The four-channel polarization images output by the polarization camera module are transmitted to the processing unit via a Gigabit Ethernet interface or a Universal Serial Bus interface for subsequent polarization decoding and pose estimation.

[0049] In one alternative implementation, the polarization camera module employs a binocular stereo configuration. Two focal plane polarization sensors are mounted on the left and right sides of the near-eye display device, respectively, with a baseline distance set to 60 mm to 80 mm to approximate the human eye's interpupillary distance. This binocular configuration not only expands the effective field of view to reduce the probability of the metasurface polarization coding marker being occluded, but also enables the direct acquisition of the metasurface polarization coding marker's three-dimensional spatial coordinates through triangulation, providing stronger geometric constraints for visual-inertial fusion positioning.

[0050] In another alternative implementation, the polarization camera module integrates an event camera sensor. Each pixel of the event camera sensor independently and asynchronously detects local brightness changes and outputs a timestamped event stream, achieving a time resolution on the order of microseconds, far exceeding that of traditional frame cameras. When the metasurface polarization-encoded marker undergoes rapid movement, the event camera sensor can capture the motion trajectory with extremely low latency, complementing the focal plane polarization sensor and thus improving the system's tracking robustness in high-speed motion scenarios.

[0051] The core function of the hardware synchronization controller is to ensure precise alignment of the polarization camera module and the micro inertial measurement unit array on the time axis. This synchronization mechanism is crucial for the accuracy of subsequent visual-inertial fusion positioning. Since the polarization camera module and the micro inertial measurement unit array operate independently at different sampling frequencies, any deviation in their data acquisition times will introduce temporal misalignment errors during the fusion process. This error is particularly significant in high-speed motion scenarios. For example, if a user's hand moves at a speed of 2 meters per second, a 10-millisecond time deviation between the polarization image frame and the inertial measurement data frame will result in a spatial position error of 20 millimeters. This magnitude of error is sufficient to cause gesture recognition failure or virtual object positioning misalignment.

[0052] refer to Figure 3 , Figure 3The chart contains two sub-plots, one above the other, showing how the position estimation error and attitude estimation error change over time. In the upper sub-plot, the horizontal axis represents time in seconds, ranging from 0 to 30 seconds, and the vertical axis represents position error in millimeters, ranging from 0 to 200 millimeters. Three curves are plotted in the sub-plot, corresponding to three positioning schemes: pure inertial tracking, pure visual tracking, and visual-inertial fusion. The pure inertial tracking curve shows a clear quadratic growth trend, with the position error approaching 200 millimeters at 30 seconds. This characteristic reflects the inherent integral drift problem of inertial measurement; the accelerometer zero-bias error, after double integration, causes the position estimate to diverge with a quadratic velocity. The pure visual tracking curve remains at a relatively low level overall but exhibits significant periodic fluctuations. More notably, sharp error jumps occur at approximately 8 seconds, 15 seconds, and 22 seconds, corresponding to scenarios where the metasurface polarization coding marker is occluded, marked with gray vertical dashed lines and labeled with the word "occlusion" in the figure. The visual-inertial fusion curve remained stable and at its lowest level throughout the entire 30-second time period. It exhibited neither the cumulative drift problem of pure inertial tracking nor the error jumps seen in pure visual tracking at occlusion moments, demonstrating the advantages of complementary fusion of visual and inertial information. The horizontal axis of the sub-graph below also represents time, ranging from 0 to 30 seconds, while the vertical axis represents attitude error in degrees, ranging from 0 to 3 degrees. The variation patterns of the three curves are similar to those in the sub-graph above: the attitude error of pure inertial tracking shows a linear growth trend, reflecting the continuous deviation of the attitude estimate from the true value due to time integration of the gyroscope's zero bias; the attitude error of pure visual tracking also shows jumps at occlusion moments; and the attitude error of visual-inertial fusion remains consistently at a low level of around 0.1 degrees. Figure 3 By comparing quantitative experimental curves, the significant advantages of the visual-inertial fusion positioning scheme in terms of accuracy and robustness compared to the single-sensor scheme are intuitively demonstrated.

[0053] The hardware synchronization controller employs a master-slave triggering architecture to achieve time synchronization across multiple sensors. Internally, the controller integrates a high-precision crystal oscillator as the system's master clock source. The crystal oscillator has a nominal frequency of 10 MHz, a frequency stability better than ±20 ppm, and negligible frequency drift during long-term operation at room temperature. Based on the master clock source, the hardware synchronization controller generates periodic synchronization pulse signals via a programmable frequency divider. The frequency of these pulse signals can be configured to 60 Hz, 90 Hz, or 120 Hz, maintaining consistency with the target frame rate of the polarization camera module. The rising edge of the synchronization pulse signal serves as a global time reference, simultaneously triggering the exposure start of the polarization camera module and the data latching of all micro-inertial measurement units in the micro-inertial measurement unit array.

[0054] The synchronization pulse signal is distributed to each sensor unit through a low-impedance transmission line. Upon receiving the synchronization pulse signal, the external trigger input port of the polarization camera module immediately activates the electronic rolling shutter or global shutter exposure sequence of the focal plane polarization sensor. Each micro inertial measurement unit (IMU) in the array latches its current acceleration and angular velocity measurements into its internal register at the rising edge of the synchronization pulse signal, and then reads them sequentially via the data bus. To eliminate signal arrival time deviations caused by differences in transmission line length, the hardware synchronization controller incorporates adjustable delay circuits on each transmission line, with a delay adjustment accuracy of 10 nanoseconds, ensuring that the timing deviation of the synchronization pulse signal arriving at each sensor unit is less than 1 microsecond.

[0055] The hardware synchronization controller also handles timestamp marking and distribution. Whenever a synchronization pulse signal is emitted, the internal counter of the hardware synchronization controller increments and generates a 64-bit unsigned integer as the global timestamp of the current frame. This timestamp uses the oscillation period of the master clock source as the smallest unit of time, corresponding to a time resolution of 100 nanoseconds. The global timestamp is transmitted to the subsequent processing unit along with the polarization image data and inertial measurement data, enabling precise pairing of data frames from different sensors using the timestamp.

[0056] In one alternative implementation, the hardware synchronization controller employs a network time protocol or a precision time protocol to achieve distributed time synchronization. When each micro inertial measurement unit in the micro inertial measurement unit array operates independently wirelessly, it cannot receive synchronization pulse signals through a physical trigger line. In this case, the wireless communication chip built into each micro inertial measurement unit exchanges timestamps bidirectionally with the time server on the near-eye display device, calculates and compensates for the deviation between the local clock and the server clock, and achieves sub-millisecond software time synchronization.

[0057] In another alternative implementation, the hardware synchronization controller integrates a Global Navigation Satellite System (GNSS) timing receiver. The GNSS timing receiver extracts the Coordinated Universal Time (UTC) reference from the satellite signals and outputs a second pulse signal to calibrate the long-term drift of the local crystal oscillator. This approach is particularly suitable for large-scale outdoor tracking scenarios, where multiple near-eye display devices can achieve cross-device data synchronization and collaborative interaction based on a unified global time reference.

[0058] The polarization decoding process first performs polarization state analysis on the multi-channel polarization image output by the focal plane polarization sensor, extracting physical quantities characterizing the polarization state from the original pixel grayscale values. As mentioned earlier, the micro-polarizer array of the focal plane polarization sensor is arranged periodically in 2×2 pixel units; therefore, polarization state analysis uses 2×2 pixel units as the basic processing unit. For each 2×2 pixel unit, the grayscale value of the 0-degree channel pixel is read. 45-degree channel pixel grayscale value 90-degree channel pixel grayscale value With 135-degree channel pixel grayscale value ,in This represents the response value of the pixel covered by the micropolarizer along the 0-degree direction of the light transmission axis. This represents the response value of the pixel covered by the micropolarizer along the 45-degree direction of the light transmission axis. This represents the response value of the pixel covered by the micropolarizer along the 90-degree direction of the light transmission axis. This represents the response value of the pixel covered by the micropolarizer along the 135-degree direction of the light transmission axis.

[0059] Based on the grayscale values ​​of the four polarization channels mentioned above, Stokes parameters are calculated to fully describe the polarization state of the incident light. The Stokes parameters consist of four components, denoted as follows: , , and ,in Indicates total light intensity. This represents the intensity difference between horizontal and vertical polarization. This indicates the intensity difference between 45-degree polarization and 135-degree polarization. This represents the intensity difference between left-handed and right-handed circular polarization. Because the focal plane polarization sensor only has a linear polarization sensing element and not a circular polarization sensing element, therefore... Since the component cannot be directly measured, it is set to zero in this embodiment. By and Add them together to get the result. This calculation is based on the principle of energy conservation of orthogonal polarization components, where the sum of the polarization components in the two orthogonal directions of 0 degrees and 90 degrees is equal to the total energy of the incident light. By and Subtraction yields the result, i.e. ,when A positive value indicates that the polarization principal axis is biased in the horizontal direction, while a negative value indicates that the polarization principal axis is biased in the vertical direction. By and Subtraction yields the result, i.e. ,when A positive value indicates that the polarization principal axis is biased at 45 degrees, while a negative value indicates that the polarization principal axis is biased at 135 degrees.

[0060] To eliminate the influence of incident light intensity variations on polarization characteristic comparison, it is necessary to... and Normalization is performed. The first normalized polarization component... Defined as Divide by ,Right now Its value ranges from -1 to +1. The second normalized polarization component... Defined as Divide by ,Right now Its value ranges from -1 to +1. Normalization ensures that the polarization feature depends only on the polarization state itself and is independent of light intensity, thus maintaining stable identity recognition performance even under conditions of uneven illumination or sensor gain drift.

[0061] linear polarization degree Used to measure the proportion of linearly polarized components in incident light, its calculation method is as follows: The square of and The square root of the sum of the squares is... The linear polarization degree ranges from 0 to 1. When the value is 0, it indicates that the incident light is completely unpolarized. A value of 1 indicates that the incident light is fully linearly polarized. The reflected light generated by the metasurface polarization-encoded marker has a high degree of linear polarization due to the anisotropic modulation of the metal nanoantenna array, typically between 0.7 and 0.95. In contrast, ambient stray light tends to be unpolarized due to multiple diffuse reflections, and its degree of linear polarization is usually below 0.3. Based on this difference, the response region of the metasurface polarization-encoded marker can be segmented from the background by setting a linear polarization degree threshold.

[0062] After polarization state analysis, the output results are used for metasurface polarization-encoded marker identification. The identification process begins with high polarization region extraction, iterating through the linear polarization values ​​of all 2×2 pixel units in the multi-channel polarization image, and marking pixel units with linear polarization values ​​higher than 0.6 as candidate polarization response regions. The linear polarization threshold of 0.6 balances detection sensitivity and anti-interference capability: a threshold that is too low may lead to some stray ambient light being misjudged as marker responses, while a threshold that is too high may miss markers with low polarization modulation efficiency. Eight-connected domain analysis is then performed on the candidate polarization response regions, checking whether the eight adjacent pixel units of each candidate pixel unit are also candidate regions, merging all interconnected candidate pixel units into an independent polarization response patch. Eight-connected domain analysis can aggregate scattered candidate pixels into complete marker response regions, while also distinguishing multiple adjacent but unconnected marker response regions.

[0063] For each polarization response patch, the centroid positions of all pixels within the patch are calculated as the image coordinates of the polarization response patch. The centroid positions are calculated using a gray-level weighted method, based on the total light intensity of all pixels within the polarization response patch. As a weighting factor, the x-coordinate of the centroid is equal to the x-coordinate of all pixel units and their corresponding... Sum of products divided by all The sum of the ordinates of the centroids equals the sum of the ordinates of all pixel units and their corresponding coordinates. Sum of products divided by all The sum of the two. Compared with the geometric center, the gray-weighted centroid can more accurately locate the optical center of the polarization response patch, especially when the patch shape is irregular or partially obscured.

[0064] The identification of metasurface polarization-encoded markers employs the Bongaley spherical projection matching method. The Bongaley sphere is a geometric representation space describing polarization states, with each point on the sphere corresponding to a unique polarization state. For a system configuration with linearly polarized illumination and linearly polarized sensitive detection, the polarization state only varies along the equatorial circle of the Bongaley sphere; therefore, identification can be simplified to a one-dimensional matching problem on the equatorial circle. The polarization features of each polarization response patch are projected onto the equatorial circle of the Bongaley sphere. Specifically, the first normalized polarization component of all pixel units within the polarization response patch is taken. The arithmetic mean of the values ​​is used as the horizontal polarization characteristic value. Take the second normalized polarization component of all pixel units within the polarization response patch. The arithmetic mean is used as the diagonal polarization characteristic value. ,by As the first Cartesian coordinate component on the equatorial circle of the Bangalore sphere, As the second Cartesian coordinate component on the equatorial circle of the Bongale sphere, the polarization characteristics of the polarization response patch are mapped to a feature point on the equatorial circle of the Bongale sphere.

[0065] During system initialization, the polarization identity codes of all deployed metasurface polarization coding markers are calibrated and converted into standard feature points on the Bangalai spherical equatorial circle. These standard feature points are stored in a polarization identity code library. To accelerate the matching process, the Bangalai spherical equatorial circle is divided into 36 angular sectors, each covering a central angle range of 10 degrees. The sector numbers increase sequentially from 0 to 35, with sector 0 corresponding to a polar angle range of 0 to 10 degrees, sector 1 corresponding to a polar angle range of 10 to 20 degrees, and so on. Each standard feature point in the polarization identity code library is categorized according to the angular sector to which its polar angle on the Bangalai spherical equatorial circle belongs, forming a sector index table. The sector index table's data structure is a hash table or an ordered array, where the key is the angular sector number, and the value is the identity of all standard feature points falling within the corresponding angular sector and their associated metasurface polarization coding markers.

[0066] For the characteristic points of the current polarization response patch, first calculate the polar angle of the characteristic point relative to the center of the equatorial circle of the Bongaley sphere. The polar angle is calculated by taking the diagonal polarization characteristic value. With horizontal polarization eigenvalues The ratio is taken as the arctangent, i.e. and according to and The sign maps the results to a full circumference ranging from 0 degrees to 360 degrees. (Based on polar angle) The numerical value determines the target angle sector into which the feature point falls, for example, when When the angle is 47 degrees, the target angle sector is sector 4. All standard feature points within the target angle sector and its two adjacent angle sectors are extracted from the sector index table as a candidate matching set. The reason for introducing adjacent angle sectors is that the polar angle measurement of feature points has some noise. If the true polar angle of a feature point happens to be near the boundary of two adjacent angle sectors, searching only the target angle sector may lead to missing correct matches.

[0067] Nearest neighbor identification is performed in the candidate matching set. The arc distance on the equatorial circle of the Bangalore sphere between the feature points of the current polarization response patch and each standard feature point in the candidate matching set is calculated. The arc distance is calculated as the absolute value of the difference between the polar angles of the two feature points. When the difference exceeds 180 degrees, 360 degrees is subtracted from the difference to select the shorter arc path. The metasurface polarization coding marker identity corresponding to the standard feature point with the smallest arc distance is selected as the identification result of the current polarization response patch. If the minimum arc distance still exceeds the preset matching tolerance threshold, which is usually set to 5 degrees, the current polarization response patch is determined not to belong to any known metasurface polarization coding marker and may be a high polarization interference source in the environment, and is therefore removed.

[0068] After identifying all polarization response patches in the current frame, the observation set of the metasurface polarization-encoded markers for the current frame is obtained. Each record in the observation set of the metasurface polarization-encoded markers for the current frame contains the identifier of a metasurface polarization-encoded marker and the image coordinates of the polarization response patch. The image coordinates are in pixels, with the origin located at the upper left corner of the multi-channel polarization image.

[0069] In one alternative implementation, the polarization decoding process employs a lookup table to accelerate polarization state analytical calculations. This is done pre-calculated for all possible polarization states within the grayscale range of 0 to 255. , , , Combine, calculate the corresponding , and The values ​​are stored in a three-dimensional lookup table. In real-time processing, the polarization parameters are obtained by directly accessing the lookup table using the grayscale values ​​of the four channels as indexes, avoiding floating-point division and square root operations, and significantly reducing computational latency.

[0070] In another alternative implementation, the polarization decoding process is executed in parallel on the graphics processor. The 2×2 pixel units of a multi-channel polarized image are independent of each other, making them naturally suitable for massively parallel processing. Deploying the polarization resolution kernel within the computation shader of the graphics processor and executing it concurrently in 2×2 pixel units can complete polarization resolution of megapixel-level images in sub-millisecond time.

[0071] The goal of visual-inertial fusion localization is to combine visual observation information provided by a polarization camera module with inertial measurement information provided by a miniature inertial measurement unit array to estimate the six-degree-of-freedom (DOF) pose of a user's hand and head in the world coordinate system. The six DDF includes three translational and three rotational degrees of freedom, describing the position and orientation of the tracked target in three-dimensional space. Relying solely on visual observation or inertial measurement has limitations: visual observation fails when the metasurface polarization-encoded marker is occluded or motion-blurred, while inertial measurement accumulates attitude errors after long-term integration due to gyroscope zero-bias drift. Visual-inertial fusion achieves high-precision, robust real-time pose estimation by complementing the advantages of both information sources.

[0072] Visual-inertial fusion localization first performs inertial pre-integration on the acceleration and angular velocity data acquired by the miniature inertial measurement unit array between two adjacent frames. The core idea of ​​inertial pre-integration is to integrate all inertial measurement data between two adjacent keyframe moments into a single relative motion constraint, without relying on the absolute pose estimation at each keyframe moment. This characteristic allows the inertial pre-integration results to be reused during pose optimization iterations, avoiding the computational overhead of re-integrating the inertial data in each iteration.

[0073] Let the first Frame and the The frame acquisition times are respectively and During this time interval, the miniature inertial measurement unit outputs acceleration measurements at a fixed frequency. With angular velocity measurement value ,in The output of the triaxial accelerometer is represented by a three-dimensional vector. The output of the three-axis gyroscope is represented by a three-dimensional vector. The acceleration measurement includes the gravitational acceleration component, the actual motion acceleration component, the accelerometer bias, and the measurement noise. The angular velocity measurement includes the actual angular velocity, the gyroscope bias, and the measurement noise.

[0074] The inertial pre-integration operation outputs three quantities: relative rotation increment. Relative velocity increment relative position increment Relative rotation increment It is a 3×3 rotation matrix describing the rotation from the first... Frame body coordinate system to the first The attitude change in the frame's body coordinate system. Relative velocity increment. For a three-dimensional vector, describe the first... Velocity change represented in the frame body coordinate system. Relative position increment. For a three-dimensional vector, describe the first... The positional changes represented in the frame body coordinate system.

[0075] The specific calculation process of inertial pre-integration adopts a discrete-time recursive method. The relative rotation increment is initialized as an identity matrix, the relative velocity increment as a zero vector, and the relative position increment as a zero vector. For each inertial measurement sampling point within the time interval, the corrected angular velocity is first obtained by subtracting the gyroscope's zero-bias estimate from the measured angular velocity. The corrected angular velocity is then multiplied by the sampling interval to obtain a small rotation angle vector. This small rotation angle vector is converted into a small rotation matrix using the Rodrigues formula. The current relative rotation increment is then multiplied on the right by this small rotation matrix to obtain the updated relative rotation increment. Next, the corrected acceleration is obtained by subtracting the accelerometer's zero-bias estimate from the measured acceleration. The current relative rotation increment is then applied to the corrected acceleration to obtain the current relative rotation increment at the [missing value]. The acceleration vector represented in the frame body coordinate system is multiplied by the sampling interval and accumulated to the relative velocity increment. The current relative velocity increment is then multiplied by the sampling interval and accumulated to the relative position increment. The above steps are repeated until all inertial measurement sampling points within the time interval have been processed.

[0076] Visual-inertial fusion localization constructs a visual-inertial joint optimization graph structure to estimate the pose state and 3D position of each metasurface polarization-encoded marker at each frame. The visual-inertial joint optimization graph structure is a factor graph representation containing two types of nodes and two types of factors. The first type of node is the pose node, where each pose node corresponds to the position and orientation of the user's hand or head in the world coordinate system at a given frame acquisition time. The state variables of the pose node include a 3D position vector, a rotation quaternion or rotation matrix, and a 3D velocity vector. The second type of node is the marker node, where each marker node corresponds to the 3D position of a metasurface polarization-encoded marker in the world coordinate system.

[0077] The first type of factor is the inertial constraint edge, which connects pose nodes in two adjacent frames and carries the relative rotation increment, relative velocity increment, and relative position increment obtained from inertial pre-integration as constraint quantities. The residual of the inertial constraint edge is defined as the difference between the relative motion estimated from the adjacent pose node estimates and the inertial pre-integration constraint quantities. The second type of factor is the visual constraint edge, which connects pose nodes and marker nodes and carries the image coordinates of the metasurface polarization-coded marker output from polarization decoding as the observation. The residual of the visual constraint edge is defined as the difference between the reprojected pixel coordinates estimated from the pose node estimates and the marker node estimates and the actual observed image coordinates. The calculation process of the reprojected pixel coordinates is as follows: first, the 3D position of the marker node in the world coordinate system is transformed to the camera coordinate system through rotation and translation of the pose node, and then projected to the pixel coordinate system through the camera intrinsic parameter matrix.

[0078] The solution for the joint visual-inertial optimization graph structure employs either the Gauss-Newton iterative method or the Levenberg-Marquardt method. In each iteration, the inertial residuals of all inertial constraint edges and the visual residuals of all visual constraint edges are calculated based on the current estimates of all pose nodes and marker nodes. The Jacobian matrix is ​​calculated for each residual with respect to the relevant node state variables; the elements of the Jacobian matrix represent the partial derivatives of the residuals with respect to the state variables. All Jacobian matrices and residual vectors are assembled into a sparse linear system of equations. The coefficient matrix of the system is an approximation of the Hessian matrix, and the right-hand side is the negative gradient vector. Solving the linear system yields the state increments for each node, which are then superimposed onto the current node estimate. For rotated state variables, the increment superposition requires a manifold update method to maintain the orthogonality of the rotation matrix or the unit norm constraint of the quaternions. Repeat the iteration until the L2 norm of the residual or the L2 norm of the state increment is less than the preset convergence threshold. A typical convergence threshold is set to a residual L2 norm change rate of less than 0.001 or a state increment L2 norm of less than 0.0001.

[0079] The solution for the joint visual-inertial optimization graph structure employs a sliding window strategy to limit computational complexity. The sliding window maintains the pose nodes of the most recent few frames and all observed marker nodes within the window; the window size is typically set to 10 to 20 frames. When a new frame arrives, the pose nodes and related constraints of the new frame are added to the sliding window, while the oldest frame is removed. The visual constraint edges between the pose nodes and marker nodes of the removed frame are converted into prior constraints on the remaining nodes through a marginalization operation, implemented using the Schur complement method. The sliding window strategy ensures that the computational cost of each optimization solution is proportional to the window size rather than the cumulative number of frames, meeting the requirements of real-time processing.

[0080] The visual-inertial fusion positioning outputs a six-DOF pose fusion estimation result. This result includes the three-dimensional position coordinates of the user's hands and head in the world coordinate system, as well as their pose expressed as a rotation matrix or Euler angles. Under good operating conditions, the accuracy of the position coordinates can reach the millimeter level, and the pose accuracy can reach the 0.1 metric level. The six-DOF pose fusion estimation result is output at the same frame rate as the polarization camera module, i.e., 60 to 120 frames per second.

[0081] In one alternative implementation, visual-inertial fusion localization employs an extended Kalman filter (EPF) instead of factor graph optimization for state estimation. The EPF models the pose state as a Markov process, propagating the state mean and covariance forward using inertial pre-integration results in the prediction step, and refining the state estimate using visual observations in the update step. The EPF has lower computational complexity than factor graph optimization, making it suitable for computationally limited embedded platforms.

[0082] In another alternative implementation, visual-inertial fusion localization introduces loop closure detection and global optimization to eliminate drift errors accumulated over long periods of operation. Loop closure detection identifies whether the user has returned to a previously visited spatial location by comparing the observation patterns of the metasurface polarization-coded markers in the current frame with those in historical frames. When a loop closure is detected, loop closure constraint edges are added to the visual-inertial joint optimization graph structure, connecting the pose nodes of the current frame and those of historical frames. Subsequently, global pose graph optimization is performed to distribute the accumulated error across the entire trajectory.

[0083] The interactive control command generation process parses the user's interaction intent based on the six-DOF pose fusion estimation results and maps it into control commands executable by the XR system. The interaction intent parsing includes two parallel processing flows: gesture state determination and gaze focus object determination.

[0084] Gesture state determination is based on the 3D position changes of various metasurface polarization-encoded markers on the hand across multiple consecutive frames. First, the overall movement direction vector and movement velocity scalar of the hand are calculated. The reference point for the overall hand is selected as the metasurface polarization-encoded marker corresponding to the center of the back of the hand. The movement direction vector is equal to the 3D position of the back of the hand in the current frame minus the 3D position of the back of the hand in the previous frame, normalized to a unit vector. The movement velocity scalar is equal to the Euclidean norm of the difference between the 3D positions of the back of the hand in the current frame and the previous frame, divided by the inter-frame time interval. The unit of the movement velocity scalar is meters per second, with a typical hand movement speed range of 0 to 3 meters per second.

[0085] Subsequently, the relative positional changes of the metasurface polarization coding markers corresponding to each fingertip relative to the metasurface polarization coding markers on the back of the hand were calculated. A local coordinate system for the hand was established with the three-dimensional position of the back of the hand as the origin. The coordinate vectors of each fingertip in the local coordinate system reflected the degree of finger extension and the angle of opening and closing. The finger extension was defined as the Euclidean distance from the fingertip to the center of the back of the hand. The extension reached its maximum value when the finger was fully extended and approached its minimum value when the finger was bent and clenched into a fist. The thumb-index finger pinch was defined as the Euclidean distance between the tips of the thumb and index finger. The pinch was close to zero when the thumb and index finger were pinched together and increased when the thumb and index finger were separated.

[0086] The current gesture state is determined based on the aforementioned geometric features. The open state is determined when the extension of all fingers is greater than the extension threshold and the pinching degree of the thumb and index finger is greater than the pinching degree threshold. The extension threshold is typically set to 0.8 times the palm width, and the pinching degree threshold is typically set to 30 mm. The fist state is determined when the extension of all fingers is less than 0.5 times the extension threshold. The pinching state is determined when the pinching degree of the thumb and index finger is less than 0.5 times the pinching degree threshold, while the extension of the remaining fingers is greater than 0.5 times the extension threshold. The pointing state is determined when the extension of the index finger is greater than the extension threshold, while the extension of the remaining fingers is less than 0.5 times the extension threshold. If the geometric features of the current frame do not meet any of the above state determination conditions, the gesture state determination result of the previous frame is maintained to avoid state jitter.

[0087] The focus of gaze is determined based on the head pose from the six-DOF pose fusion estimation and the gaze direction output by the eye-tracking unit built into the near-eye display device. First, the position of the user's binocular midpoint in the world coordinate system is calculated based on the head pose. The spatial offset of the user's binocular midpoint relative to the head tracking point is measured and stored during the system calibration phase. This spatial offset is a fixed three-dimensional vector. Applying the rotation matrix of the head pose to the spatial offset vector and then adding the translation vector of the head pose yields the position of the user's binocular midpoint in the world coordinate system.

[0088] The gaze direction output by the eye-tracking unit is a unit vector, representing the direction of the user's gaze in the head coordinate system. Applying the head pose rotation matrix to the gaze direction vector yields the gaze direction in the world coordinate system. A gaze ray is emitted from the midpoint of the user's eyes along the gaze direction in the world coordinate system; this gaze ray is a semi-infinite straight line originating from the midpoint of the user's eyes and oriented along the gaze direction.

[0089] A gaze ray intersection test is performed on each interactive virtual object in the current virtual scene. Each interactive virtual object has a bounding box geometric representation, typically an axis-aligned bounding box or an oriented bounding box. The gaze ray intersection test calculates the intersection points of the gaze ray with each face of the bounding box. If an intersection point exists and lies within the positive half-space of the gaze ray, the gaze ray is considered to intersect with the interactive virtual object. For all interactive virtual objects that intersect, the Euclidean distance from each intersection point to the midpoint of the user's eyes is calculated, and the interactive virtual object with the closest distance is determined as the gaze focus object. If no intersecting interactive virtual objects exist, the gaze focus object is empty.

[0090] XR interactive control commands are generated based on the current gesture state and the object being gazed at. The types of interactive control commands include four types: object selection commands, object dragging commands, object release commands, and menu invocation commands.

[0091] When the current gesture state is a pinch state, the focus object exists, and there is no currently selected object, an object selection command is generated. The object selection command carries a unique identifier of the focus object. After receiving the object selection command, the XR scene rendering engine marks the corresponding interactive virtual object as the selected object and visually reflects the selection status to the user through methods such as highlighting a border or color change.

[0092] When the current gesture is a pinch gesture and an object is already selected, an object drag command is generated. This drag command carries the target displacement vector of the selected object, which is equal to the hand movement direction vector multiplied by the hand movement speed scalar, and then multiplied by the frame interval. Upon receiving the drag command, the XR scene rendering engine moves the selected object's 3D position along the target displacement vector by the corresponding distance, achieving a dragging effect where the virtual object follows the user's hand movements.

[0093] When the current gesture changes from a pinched state to an open state and an object is already selected, an object release command is generated. The object release command carries a unique identifier of the selected object. After receiving the object release command, the XR scene rendering engine deselects the corresponding interactive virtual object, and the virtual object remains at the position at the time of release.

[0094] When the current gesture is pointing and the focus object exists, a menu invocation command is generated. The menu invocation command carries a unique identifier of the focus object. After receiving the menu invocation command, the XR scene rendering engine renders the context operation menu at a 3D spatial location near the focus object. The menu content is dynamically generated according to the type of focus object. For example, for 3D model objects, operation options such as rotation, scaling, and deletion are displayed, while for media objects, operation options such as play, pause, and volume adjustment are displayed.

[0095] refer to Figure 4 , Figure 4 Containing three sub-images (top, middle, and bottom), this diagram fully illustrates the processing flow from initial geometric feature extraction to final gesture state determination. The top sub-image shows the curve of finger extension variation. The horizontal axis represents time in seconds, ranging from 0 to 10 seconds, while the vertical axis represents finger extension, expressed normally, ranging from 0 to 1.1. Finger extension is defined as the ratio of the Euclidean distance from the fingertip to the center of the back of the hand to the width of the palm; a higher value indicates a more fully extended finger. The curve exhibits distinct phased changes: from 0 to 1.5 seconds, the finger extension remains at a high level of around 0.85, corresponding to the open finger state; from 1.5 to 3 seconds, the finger extension gradually decreases to around 0.5; from 3 to 5 seconds, the finger extension remains at a low level of around 0.5; from 5 to 6 seconds, the finger extension rapidly recovers to around 0.85; from 6 to 8 seconds, the finger extension is at a mid-level of around 0.65; and from 8 to 10 seconds, the finger extension recovers again to around 0.85. A red horizontal dashed line marks the extension threshold of 0.8; areas above the threshold are filled in green to indicate an extended state, and areas below the threshold are filled in orange to indicate a bent state. The intermediate subplot shows the change in thumb-index finger pinching, with the vertical axis representing the distance between the thumb and index finger in millimeters, ranging from 0 to 90 millimeters. The thumb-index finger distance is defined as the Euclidean distance between the tips of the thumb and index finger; a smaller value indicates a higher degree of pinching. The curve trend shows a negative correlation with the finger extension curve in the upper sub-graph: when finger extension decreases, the distance between the thumb and index finger decreases simultaneously, indicating that the user is performing a pinching motion. A red horizontal dashed line is drawn in the graph to mark the pinching threshold of 30 mm; areas below the threshold are filled in red to indicate a pinched state, and areas above the threshold are filled in cyan to indicate a separated state. The lower sub-graph displays the gesture state recognition results, using colored bars to visually present the gesture state determination results for different time periods. Four gesture states are defined in the graph and distinguished by different colors: green represents an open state, orange represents a clenched fist state, red represents a pinched state, and blue represents a pointing state. Based on the geometric feature curves and preset thresholds in the two upper sub-graphs, the system automatically determines the gesture state for each time period and presents it as continuous colored bars.

[0096] The generated XR interactive control commands are sent to the XR scene rendering engine for execution. The XR scene rendering engine updates the position, posture, or display status of the corresponding interactive virtual objects in the virtual scene according to the received interactive control commands, and renders the updated virtual scene image onto the display screen of the near-eye display device to present to the user, completing a closed-loop interaction from the user's physical action to the virtual scene response.

[0097] In one optional implementation, a gesture state confirmation mechanism is introduced into the interaction control command generation process to avoid false triggers. When a change in gesture state is detected, the corresponding interaction control command is not generated immediately. Instead, the gesture state is allowed to remain stable for several consecutive frames before the state change is confirmed. The number of confirmation frames is typically set to 3 to 5 frames. The gesture state confirmation mechanism can filter out momentary false detections caused by finger tremors or occlusion, improving the stability of the interaction experience.

[0098] In another optional implementation, the interactive control command generation process supports two-handed collaborative interaction. When both hands are pinched together and the focus of both hands is the same, object scaling commands are generated based on the change in distance between the centers of the backs of the hands; increasing distance corresponds to a zoom-in operation, and decreasing distance corresponds to a zoom-out operation. When both hands are pinched together and the focus of both hands is different, simultaneous operation of two independent virtual objects is supported. The two-handed collaborative interaction mode expands the freedom and efficiency of user interaction with the virtual scene.

[0099] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. An XR virtual-real interaction control device integrating optical motion capture and inertial navigation, characterized in that, include: A metasurface polarization-encoded marker array is distributed and installed at preset tracking points on the user's hand and head to generate reflected light carrying polarization identification codes when receiving illumination light. A miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, is used to acquire acceleration and angular velocity data at each tracking point; a near-eye display device includes an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head area, while the polarization camera module receives reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization images; a hardware synchronization controller is used to synchronously trigger data acquisition between the polarization camera module and the miniature inertial measurement unit array; and a polarization decoding module is used to perform polarization state analysis and marker identification on the multi-channel polarization images, outputting the image coordinates and identification of each metasurface polarization-encoded marker. The visual-inertial fusion positioning module is used to calculate the six-degree-of-freedom pose fusion estimation result of the user's hand and head based on the image coordinates and identity identifier output by the polarization decoding module and the acceleration and angular velocity data collected by the micro inertial measurement unit array. The interactive control command generation module is used to parse the user's interaction intent and generate XR interactive control commands based on the six-degree-of-freedom pose fusion estimation results.

2. The apparatus according to claim 1, characterized in that, Each metasurface polarization coding marker in the metasurface polarization coding marker array includes a transparent substrate and a metal nanoantenna array covering the surface of the transparent substrate. The major axis orientation angle of each nanoantenna unit in the metal nanoantenna array is distributed according to a preset spatial arrangement rule, so that when linearly polarized illumination light is incident, each nanoantenna unit generates anisotropic phase delay and amplitude modulation of the incident light, thereby making the reflected light elliptical polarized as the polarization identity code of each metasurface polarization coding marker.

3. The apparatus according to claim 1, characterized in that, Each metasurface polarization coding marker in the metasurface polarization coding marker array is distributed and installed on the back of the user's hand, each finger joint, and the frontal and temporal sides of the head; each micro inertial measurement unit in the micro inertial measurement unit array includes a triaxial microelectromechanical accelerometer and a triaxial microelectromechanical gyroscope.

4. The apparatus according to claim 1, characterized in that, The illumination module is a ring-shaped near-infrared light-emitting diode illumination array, used to generate linearly polarized illumination light with a wavelength of 850 nanometers; the polarization camera module includes a focal plane polarization sensor, and a micro polarizer array is integrated on the photosensitive surface of the focal plane polarization sensor. The micro polarizer array is arranged periodically in 2×2 pixel units, and the four pixels in each 2×2 pixel unit cover the micro polarizers along the light transmission axis in the 0-degree, 45-degree, 90-degree, and 135-degree directions, respectively.

5. The apparatus according to claim 1, characterized in that, The hardware synchronization controller generates periodic synchronization pulse signals, which simultaneously trigger the exposure start of the polarization camera module and the data latching of all micro inertial measurement units in the micro inertial measurement unit array, so that the polarization image frame and the inertial measurement data frame have the same acquisition time mark.

6. The apparatus according to claim 4, characterized in that, The polarization decoding module performs polarization state analysis as follows: For each 2×2 pixel unit in the multi-channel polarization image, the gray values ​​of the 0-degree channel, 45-degree channel, 90-degree channel, and 135-degree channel pixels within the 2×2 pixel unit are read. The gray values ​​of the 0-degree channel and 90-degree channel pixels are added to obtain a first intensity sum. The gray values ​​of the 0-degree channel and 90-degree channel pixels are subtracted to obtain a first intensity difference. The gray values ​​of the 45-degree channel and 135-degree channel pixels are subtracted to obtain a second intensity difference. The first intensity difference is divided by the first intensity sum to obtain a first normalized polarization component. The second intensity difference is divided by the first intensity sum to obtain a second normalized polarization component. The square root of the sum of the first and second normalized polarization components is obtained to obtain the degree of linear polarization.

7. The apparatus according to claim 6, characterized in that, When performing tag identification, the polarization decoding module employs the Bongaret spherical projection matching method, which includes: traversing the linear polarization degree values ​​of all 2×2 pixel units, retaining pixel units with linear polarization degree values ​​higher than 0.6 as candidate polarization response regions, performing octet analysis on the candidate polarization response regions and merging them into polarization response patches, calculating the centroid position of each polarization response patch as its image coordinates; for each polarization response patch, taking the mean of the first normalized polarization component of all pixel units within the polarization response patch as the horizontal polarization feature value, taking the mean of the second normalized polarization component of all pixel units within the polarization response patch as the diagonal polarization feature value, and using the horizontal polarization feature value and the diagonal polarization feature value as the coordinates of the feature point on the equatorial circle of the Bongaret sphere; and then... The equatorial circle is divided into 36 angular sectors, each covering a central angle range of 10 degrees. Standard feature points in the pre-stored polarization identity coding library are categorized according to the angular sector in which they are located to form a sector index table. The polar angle of the feature points of the current polarization response patch relative to the center of the Bangalley spherical equatorial circle is calculated, and the target angular sector in which the polar angle falls is determined. All standard feature points in the target angular sector and the two adjacent angular sectors of the target angular sector are extracted from the sector index table as candidate matching sets. The arc length distance between the feature points of the current polarization response patch and each standard feature point in the candidate matching set on the Bangalley spherical equatorial circle is calculated. The metasurface polarization coding marker identity identifier corresponding to the standard feature point with the smallest arc length distance is selected as the identification result of the current polarization response patch.

8. The apparatus according to claim 1, characterized in that, The visual-inertial fusion positioning module performs inertial pre-integration on the acceleration and angular velocity data acquired by the micro inertial measurement unit array between adjacent frames to obtain the relative attitude and relative velocity changes between adjacent frames. The module constructs a visual-inertial joint optimization graph structure, which includes pose nodes and marker nodes. The pose nodes represent the position and attitude of the user's hand or head in the world coordinate system at each frame acquisition time, and the marker nodes represent the three-dimensional positions of each metasurface polarization-encoded marker in the world coordinate system. The graph also includes inertial constraint edges and visual constraint edges. The inertial constraint edges connect the pose nodes of adjacent frames and carry the relative attitude and relative velocity changes, while the visual constraint edges connect the pose nodes and marker nodes and carry the image coordinates output by the polarization decoding module. The module then performs iterative optimization on the visual-inertial joint optimization graph structure to obtain a six-degree-of-freedom pose fusion estimation result.

9. The apparatus according to claim 1, characterized in that, The interactive control command generation module calculates the hand movement direction vector and movement speed scalar based on the three-dimensional position changes of each metasurface polarization coding marker on the hand in multiple consecutive frames. The interactive control command generation module determines the current gesture state based on the relative position changes of the metasurface polarization coding marker corresponding to each fingertip relative to the metasurface polarization coding marker corresponding to the back of the hand. The current gesture state includes open state, clenched fist state, pinched state, and pointing state.

10. The apparatus according to claim 9, characterized in that, The interaction control command generation module calculates the midpoint position of the user's eyes based on the head pose in the six-degree-of-freedom pose fusion estimation result, and emits a gaze ray based on the gaze direction. It then performs an intersection test with the bounding boxes of each interactive virtual object in the virtual scene, and identifies the interactive virtual object whose intersection point is closest to the midpoint of the user's eyes as the gaze focus object. The interaction control command generation module generates XR interaction control commands based on the current gesture state and the gaze focus object. The XR interaction control commands include object selection commands, object drag commands, object release commands, and menu call commands.

Citation Information

Patent Citations

  • Large-view-field high-resolution imaging system based on polarization regulation and control

    CN121454738A

  • Integrated metasurfaces for free-space wavefront generation with complete amplitude, phase, and polarization control

    US20230367144A1