A public space non-contact interaction anti-interference and master control right locking method based on space-time double-flow feature fusion

By fusing spatiotemporal dual-stream features and employing a dynamic hysteresis arbitration mechanism, the instability of interaction in non-contact gesture tracking systems in multi-user scenarios is resolved. This enables accurate evaluation of interaction intent and control locking in complex environments, thereby improving the user experience.

CN122239944APending Publication Date: 2026-06-19GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2026-03-30
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing contactless gesture tracking systems struggle to accurately identify legitimate interaction intentions in multi-user concurrent scenarios, leading to unstable interfaces, frequent focus jumps, easy loss of control, and poor user experience.

Method used

A spatiotemporal dual-stream feature fusion method is adopted, which utilizes Bi-GRU network and VGG network combined with dynamic hysteresis arbitration mechanism. Hand movements are analyzed by temporal matrix and depth image tensor to construct two-dimensional temporal matrix and local depth map. Combined with dual-stream neural network and dynamic hysteresis arbitration mechanism, intelligent arbitration and interference isolation for multiple targets are achieved.

Benefits of technology

Accurately assessing interaction intent in complex environments, isolating unauthorized interference, firmly locking in control, ensuring smooth and stable interaction, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_15
    Figure SMS_15
  • Figure SMS_29
    Figure SMS_29
Patent Text Reader

Abstract

This invention discloses a method for anti-interference and control locking of non-contact interaction in public spaces based on spatiotemporal dual-stream feature fusion, belonging to the field of extended reality and human-computer interaction technology. Addressing the problems of interaction focus conflicts and control confusion caused by multiple user gesture intrusions in public spaces, as well as the inability of single-modal recognition to distinguish continuous interaction intentions and poor anti-interference performance, this invention adopts a spatiotemporal dual-stream feature fusion architecture: Bi-GRU is used to evaluate the temporal continuity of hand skeletons, VGG16 extracts spatial hierarchical features from the depth map, and MLP is used to calculate the legality score of interaction intentions, achieving control arbitration and interference data shielding. This invention can lock the initial master user's gesture data stream and filter interference targets at the lower level, making it suitable for XR holographic display of cultural relics in public spaces such as museums and science and technology museums, ensuring stable and continuous non-contact gesture interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction and computer vision interdisciplinary technology, specifically involving non-contact gesture / posture interaction control technology in public spaces, and particularly involving a non-contact interaction method in public spaces that achieves environmental interference suppression, accurate identification of interactive targets, and exclusive locking of master control through spatiotemporal dual-stream feature fusion. Background Technology

[0002] With the rapid development of extended reality (XR) technology and the continuous evolution of human-computer interaction (HCI) concepts, interaction methods are undergoing a profound transformation from traditional physical contact to more natural and intuitive non-contact spatial interaction. Among numerous three-dimensional spatial interaction methods, dynamic gesture recognition technology based on machine vision and depth sensors, with its advantages of zero contact, high immersion, and low learning cost, is gradually becoming the core carrier for building modern high-level digital experiences. Especially in open public spaces such as digital exhibitions in museums and interactive science museums, in order to balance the safety protection of physical cultural relics with the in-depth exploration experience of visitors, the use of gesture interaction technology for the holographic display and structural analysis of three-dimensional holographic models of cultural relics has become an important development trend in the industry. Such non-contact interaction systems typically rely on high-precision hand tracking peripherals, such as the LeapMotion Controller, to capture the operator's hand skeletal joint coordinate sequence and spatial depth information in real time, enabling viewers to examine the microscopic details of cultural heritage from all angles using natural gestures that conform to human intuition, such as grasping, rotating, and translating. This interactive paradigm not only breaks the visual limitations of traditional physical display cases and greatly expands the dimensions of information transmission, but also lays a solid technical foundation for creating a highly interactive and participatory modern public digital cultural space.

[0003] However, when such visual tracking-based contactless interactive systems are widely deployed in crowded open public spaces, their robustness and interaction stability in multi-user concurrent scenarios face extremely severe challenges. Existing gesture tracking devices and their underlying driver frameworks are mostly designed based on an ideal operating environment from a single user's perspective. The system kernel lacks an effective mechanism for identifying intent and arbitrating ownership when multiple targets appear simultaneously. In actual public display applications, when the first user is dominating the interactive interface, such as continuously and meticulously rotating and observing a 3D model, due to the openness of the physical space, it is highly likely that the arms of other onlookers will unintentionally or intentionally intrude into the sensor's effective recognition field of view. At this time, the system's underlying model often cannot intelligently determine which hand target is the truly legitimate control subject, leading to severe confusion of multi-target feature data and pollution of the feature space. This phenomenon of "sudden intrusion of heterogeneous targets" directly causes frequent focus jumps in the interactive interface, severe key bindings in control commands, and even the complete collapse of the system's control logic, forcibly interrupting the main user's continuous operation and greatly damaging the continuity of the immersive experience and the quality of the exhibition. Most existing solutions rely solely on the spatial location of a single frame for filtering, failing to combine the temporal continuity of the user's hand movements (time dimension) with the hierarchical relationship of physical depth (spatial dimension) for deep feature mining. This results in the inability to accurately extract the master control signal with "truly legitimate interaction intent" from the underlying data stream.

[0004] Therefore, how to overcome the inherent limitations of existing devices in single-person, single-machine mode, and explore an intelligent arbitration mechanism that can accurately assess the intensity of interaction intentions of all parties, isolate and block illegal intrusion interference in real time, and firmly lock the control of the first user in a complex and dynamic multi-person viewing environment has become an urgent need and core driving force for promoting the application of contactless spatial interaction technology to large-scale public scenarios. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies in contactless interaction in public spaces, such as poor anti-interference capabilities, easy loss of control, and unsatisfactory user experience, this invention provides a method for anti-interference and control locking in contactless interaction in public spaces based on spatiotemporal dual-stream feature fusion. This invention constructs a two-dimensional temporal matrix and a depth image tensor, utilizes a dual-stream neural network to jointly evaluate interaction intent, and combines a dynamic hysteresis arbitration mechanism to achieve data anti-channeling in complex environments and smooth driving of the 3D engine.

[0006] The present invention provides a method for non-contact interactive anti-interference and master control locking in public spaces based on spatiotemporal dual-stream feature fusion, with the following steps:

[0007] Step 1: Use a non-contact infrared depth sensor to concurrently capture multiple targets' raw heterogeneous data streams within the same time window. The data streams include three-dimensional spatial coordinates and raw global depth maps.

[0008] Step 2: The system scans the effective field of view at a fixed sampling rate. At any time, if detected Individual hand target set First, the dominant hand attribute (left or right hand) and the three-dimensional coordinates of the wrist root are extracted for each hand target. Then, a lifecycle identifier is assigned to each hand based on spatial connectivity analysis. Subsequently, based on human biomechanical constraints, homogeneous entity clustering is performed: if there are two mutually exclusive targets (one on the left and one on the right) in the field of view, and the relative distance between their wrists and the angle between their forearm vectors is within a preset torso span threshold, then the two targets are bound to the same user identifier. The "two-handed collaboration group" is used; for targets that do not pair successfully, a separate user identifier is assigned independently. The system uses user identifiers. Independent isolated data channels are established for basic units, and dual-stream heterogeneous data within the same "two-handed collaboration group" are transmitted synchronously in the same user channel;

[0009] Step 3: For the target Extracting its 5 fingertips in the first Frame's three-dimensional coordinates They are then concatenated into a single-frame feature vector. The system traces back 150 consecutive frames along the time axis to construct a two-dimensional time series matrix. ;

[0010] Step 4: Based on Calculate the 2D bounding box from the bone projection points, and crop the local depth map from the global depth map. Then, it is resampled to 128×128 pixels, and then depth gradient Min-Max normalization is performed:

[0011]

[0012] in, and These are the minimum and maximum depth thresholds for the effective field of view; they are then converted into single-channel grayscale image tensors. Output For heterogeneous data sources with uniform format ;

[0013] Step 5: Expanded into a sequence vector set at time steps, and then input into a fully connected layer for dimensionality upgrading. ;

[0014] Step 6: Extract the feature vector of a single frame The input is fed into a Bidirectional Gated Recurrent Unit (Bi-GRU) network. The initial forward and backward hidden states are both set to zero vectors, i.e. , .

[0015] At time step The forward network computes the forward update gate in ascending time order. With forward reset door :

[0016]

[0017] Physical defense logic: For users who are consistently observing cultural relics, the historical trajectory is continuous. The system tends to retain memories; when the interfering arm suddenly extends, the coordinates change drastically. The system was forcibly activated to disconnect historical connections, which was determined to be a trajectory fault.

[0018] Then the forward candidate states are calculated. With the final forward hidden state :

[0019]

[0020] Simultaneously, the backward network performs symmetric computation in reverse temporal order (from frame 150 to frame 1) to obtain the backward update gate. Reverse reset door Backward candidate state And the final backward hidden state :

[0021]

[0022] Step 7: After calculation, extract the final forward hidden state of frame 150. Final backward hidden state with frame 1 The features are concatenated to construct a global temporal feature containing the complete context. :

[0023]

[0024] Subsequently, dimensionality reduction is performed using a fully connected layer to output the final temporal coherence feature vector. ;

[0025] Step 8: Convert the grayscale image tensor Input the VGG network to extract spatial features, where the first... Convolutional layers calculate high-frequency gradient features:

[0026]

[0027] in This is the feature activation map of the previous layer; features closest to the sensor are continuously preserved through a 2×2 max pooling layer.

[0028]

[0029] Physical defense logic: In the depth map, brighter pixels represent those closer to the sensor. Max pooling mathematically preserves the physical bumps closest to the sensor, automatically filtering out interfering depth information from distant or edge-related areas.

[0030] Step 9: After multiple convolutional blocks, the deep feature map is flattened, and spatial depth-level feature vectors are calculated through fully connected layers. ;

[0031] Step 10: and Connected in the channel dimension Input MLP hidden layer:

[0032]

[0033]

[0034] Step 11: By activating the sigmoid function of a single neuron in the output layer, the high-dimensional features are compressed into scalars, and the scalar score is output. :

[0035]

[0036] This step and the steps described above and Both are pre-trained learnable parameter matrices and bias vectors;

[0037] Step 12: Obtain the set of concurrent target scores Perform exponential moving average dejitting:

[0038]

[0039] The initial time hour, ;

[0040] Step 13: The system upgrades the evaluation dimension to "user entity ( Extracting the comprehensive interaction intent score for each user. (If it is a two-handed collaborative group, the highest score will be taken). The system then dynamically switches the arbitration mode based on the number of active UIDs in the current field of view: when there is only a single user in the field of view, the computing power bypass optimization is triggered, and the master control can be established quickly by simply releasing the threshold; when multiple users interfere in the field of view, the dual threshold hysteresis is fully activated, and new users need to break through the higher preemption threshold to take over, while the original master user only needs to hold the lower release threshold to maintain the lock, thereby building an absolute asymmetric anti-interference barrier;

[0041] Step 14: Based on the arbitration result, generate a user ( () is a Boolean control mask set of dimension . For any detected hand target As long as it belongs to the current master user Within the set, its physical coordinates are preserved; otherwise, a low-level scalar multiplication filter is performed to return it to zero.

[0042]

[0043] Physical shielding logic: Data from non-master users (whether it's one hand or two hands) is completely muted here. The judgment criteria are... At the hierarchical level, the system ultimately outputs clean driver data to the rendering engine from only the legitimate user. This allows the system to support multi-user exclusivity and interference resistance while seamlessly supporting advanced two-handed interactive gestures at the underlying level.

[0044] In step 2 above, the fixed sampling rate is 60fps; in step 3, the system traces back 150 consecutive frames along the time axis, corresponding to the construction of a historical time series data window with a length of 2.5 seconds.

[0045] The minimum depth threshold of the effective field of view in step 4 above With maximum depth threshold The effective physical distance range for non-contact interaction is set; while reducing dimensionality, the max pooling layer continuously extracts and retains the convex feature points that are closest to the infrared depth sensor in terms of physical distance, so as to filter out environmental depth interference from afar.

[0046] In step 8 above, the VGG network uses single-channel adaptive convolution to extract grayscale tensor features. The feature extraction process includes five consecutive convolutional block operations, and each convolutional block is followed by the max pooling layer for depth feature filtering.

[0047] In step 13 above, the parameters of the dual-threshold hysteresis circuit are configured as a set of asymmetric thresholds, including a preemption threshold and a release threshold, and the preemption threshold is strictly greater than the release threshold. When a new target appears in the set of concurrent targets, it can only take control if its interaction intent legitimacy score is higher than the preemption threshold. If a master user has been established, as long as its interaction intent legitimacy score does not fall below the release threshold, the system maintains the locked state of the master user, thereby forming an anti-interference protection zone.

[0048] In step 15 above, converting the unique legal coordinates into the rendering engine coordinate system specifically means mapping the original right-handed coordinate system spatial data of the infrared depth sensor to the left-handed coordinate system of the 3D rendering engine through a 3D affine transformation.

[0049] The standard three-dimensional spatial coordinates and control event package described above are output to the upper-level display system to smoothly drive the rotation, scaling and dissection display of the three-dimensional model of digital cultural heritage in non-contact interactive scenarios in public spaces. Beneficial effects of this invention patent

[0050] This invention innovatively introduces a spatiotemporal dual-stream feature fusion architecture. In the temporal dimension, the reset gate mechanism of the Bi-GRU network can sensitively capture and sever historical connections of sudden interference trajectories (such as a passerby suddenly reaching out). In the spatial dimension, a VGG network combined with a max-pooling layer continuously retains the nearest physical protrusion features (the master's fingertip) while filtering out distant background. The combination of these two approaches maps high-dimensional features to an interaction intent legitimacy score, fundamentally solving the blindness of traditional sensors that "grab anything that's a hand."

[0051] To address the issue of frequent focus switching caused by multiple viewers, this invention designs a master control arbitration mechanism consisting of "EMA smoothing + dual threshold hysteresis". By setting unequal preemption and release thresholds, the system provides the current master user with extremely high anti-interference protection. Even if the master's hand experiences a brief tremor or is partially obscured, the control state can be forcibly maintained, ensuring an absolutely smooth and fluid digital cultural heritage display process.

[0052] After establishing the master user, this invention directly generates a Boolean control mask set at the underlying level and performs scalar multiplication to zero the physical data of non-master users, achieving hardware-level data anti-interference. Simultaneously, addressing the underlying differences between sensors (right-handed) and rendering engines (left-handed), it incorporates standard coordinate system affine transformation logic, outputting absolutely pure control event packages that can smoothly drive the rotation, scaling, and anatomical display of various 3D digital artifacts in a plug-and-play manner. Attached Figure Description

[0053] Figure 1The main flowchart of a non-contact interactive anti-interference and master control locking method for public spaces based on spatiotemporal dual-stream feature fusion provided by the present invention is shown. Detailed Implementation

[0054] As shown in the attached diagram of the specification, a non-contact interactive anti-interference and master control locking method for public spaces based on spatiotemporal dual-stream feature fusion includes the following steps in its overall execution process:

[0055] System hardware deployment: In this embodiment, the non-contact infrared depth sensor uses a Leap Motion controller, which is hidden beneath the artifact display case or within the inclined platform, forming an inverted cone-shaped effective interactive field of view. The upper-level display system is developed based on 3D rendering engines such as Unity, and is used for high-precision rendering of 3D models of digital cultural heritage such as bronzes and ceramics.

[0056] The method of this invention is executed by a background control computer and specifically includes the following core steps:

[0057] S101: Process begins.

[0058] S102: Employs a non-contact infrared depth sensor to concurrently capture raw heterogeneous data streams from multiple targets within the same time window. The data streams include three-dimensional spatial coordinates and raw global depth maps.

[0059] S103: The system continuously scans the effective field of view at a fixed high sampling rate (e.g., 60fps).

[0060] S104: Intent detection and determination. Determine the current time. Was it detected? A set of hand targets (of which) If concurrent targets are detected, proceed to step S105; if no target is detected or only a single target is detected, return to step S103 to continue scanning.

[0061] S105: Based on spatial connectivity analysis, assign a unique lifecycle identifier to each hand within the field of view. Furthermore, an independent, isolated data channel is established at the underlying level to prevent cross-contamination of heterogeneous data. Subsequently, the system distributes the data to a spatiotemporal dual-stream architecture for parallel processing.

[0062] Temporal characteristic flow branches (S106-S107):

[0063] S106: Extract the three-dimensional spatial coordinates of the target fingertip and construct a two-dimensional temporal matrix by traversing 150 consecutive frames along the time axis. .

[0064] S107: Input the time series matrix into the Bi-GRU network. Through the update and reset gate mechanisms of the forward and backward networks, the historical connections of sudden interference trajectories (such as a suddenly intruding arm) are accurately captured and severed, and finally the temporal coherence feature vector is extracted. .

[0065] Spatial characteristic flow branches (S108-S109):

[0066] S108: Based on the two-dimensional bounding box of the hand, a local depth map (ROI) is cropped from the global depth map, resampled and the depth gradient is normalized by Min-Max, and then converted into a single-channel grayscale image tensor.

[0067] S109: Input the grayscale tensor into the VGG convolutional neural network, combine multiple convolutional blocks and max pooling layers to continuously retain the physical protrusion features closest to the sensor (such as the fingertips of the controller's hand) and filter out distant background interference, extracting spatial depth-level feature vectors. .

[0068] S110: Multimodal fusion computing. and Connected in the channel dimension The data is then fed into the hidden layer of a multilayer perceptron (MLP) for cross-feature mining.

[0069] S111: By using the Sigmoid activation function in the output layer, the high-dimensional features are compressed and mapped to scalars, and the legality score of the interaction intent of each target is calculated. Its range is (0,1).

[0070] S112: Perform exponential moving average (EMA) de-jitter processing on the obtained concurrent target score set to smooth instantaneous characteristic fluctuations and output smoothed scores. .

[0071] S113: User (UID)-based mastery arbitration. The system first assesses the environmental complexity within the current field of view and executes dynamic branch routing based on the total number of valid UIDs.

[0072] (1) Branch A (Single-person rapid response mode): If it is determined that there is only a single user in the current field of view (total number of UIDs = 1), the system triggers the computing power bypass optimization logic. At this time, the complex preemption game is skipped, and only it is determined whether the comprehensive score of the UID is higher than the release threshold (such as 0.40). If so, the user is directly granted or maintained control. This logic not only saves a lot of underlying computing overhead, but also lowers the starting threshold for users to get started for the first time - users do not need to deliberately make extremely exaggerated actions to rush for a high score preemption threshold, they only need to show a passable interaction intention to naturally take over control.

[0073] (2) Branch B (Multi-user High-Prevention and Anti-Interference Mode): If an intruder (total number of UIDs ≥ 2) enters the field of view, the system fully activates the dual-threshold hysteresis. At this time, the system strictly compares the intent strength of new and old users: the score of the new UID of the passerby or intruder must forcibly break through the extremely high preemption threshold (e.g., 0.75) to seize control; while the established master user, even if the intention score is temporarily lowered due to slight adjustments and rotations of both hands, as long as the highest score does not fall below the release threshold (e.g., 0.40), the system will lock it firmly in the master position. This asymmetric arbitration mechanism fundamentally eliminates the system focus flickering and loss of control caused by unintentional waving of bystanders.

[0074] S114: Generate a Boolean control mask set based on the arbitration result, and perform low-level scalar multiplication filtering on all target coordinates. Data from non-master controllers is completely zeroed out and muted here.

[0075] S115: Affine Transformation and Output of Coordinate System. Considering the differences between hardware sensors and upper-level systems, the uniquely valid coordinates are transformed into a standard 3D rendering engine coordinate system through scaling, rotation, and translation matrix operations. For example, right-handed coordinates are mapped to the left-handed coordinate system of Unity3D or Unreal Engine 5. Finally, standard 3D spatial coordinates and control event packages are output to the underlying event system, smoothly driving real-time rotation, scaling, and internal structural analysis of 3D models of digital cultural heritage in the Unity scene, providing visitors with an extremely stable and interference-resistant user experience.

[0076] S116: Process complete.

Claims

1. A method for non-contact interactive anti-interference and master control locking in public spaces based on spatiotemporal dual-stream feature fusion, characterized in that, The steps of this method are as follows: Step 1: Use a non-contact infrared depth sensor to concurrently capture multiple targets' raw heterogeneous data streams within the same time window. The data streams include three-dimensional spatial coordinates and raw global depth maps. Step 2: The system scans the effective field of view at a fixed sampling rate. At any time, if detected Individual hand target set Instead of blindly isolating all targets, the system introduces a clustering mechanism for homologous entities. The system extracts the dominant hand attribute (left / right hand) and wrist joint spatial coordinates for each target. If two hands, belonging to the left and right hands respectively, are detected, and their spatial physical distance is within a reasonable range of shoulder width and arm span for a normal adult, the system clusters them at the underlying level into a "two-handed coordination group" belonging to the same operator, and assigns a unified user identifier. For isolated hand targets that cannot be paired according to the above rules, assign a separate target. The final system is based on Establish independent, isolated data channels for each dimension; Physical defense and collaborative logic: In public spaces, it is very easy for two tourists standing in adjacent positions to reach out their hands at the same time. If only spatial distance is relied upon, it is easy to misjudge the adjacent hands of the two people as the same person. This step requires that the collaborative group must meet the physical mutual exclusion property of "one left and one right", which fundamentally eliminates the absurd situation of "two right hands" being bound to the same master. This logic enables the system to perfectly support advanced complex gestures such as two-hand zoom and dual-axis rotation of a single user while filtering malicious interference. Step 3: For the target Extracting its 5 fingertips in the first Frame's three-dimensional coordinates Concatenate them into a single-frame feature vector The system traces back 150 consecutive frames along the time axis to construct a two-dimensional time series matrix. ; Step 4: Based on Calculate the 2D bounding box from the bone projection points, and crop the local depth map from the global depth map. Then, it is resampled to 128×128 pixels, and then depth gradient Min-Max normalization is performed: in, and These are the minimum and maximum depth thresholds for the effective field of view; they are then converted into single-channel grayscale image tensors. Output For heterogeneous data sources with uniform format ; Step 5: Expanded into a sequence vector set at time steps, and then input into a fully connected layer for dimensionality upgrading. ; Step 6: Extract the feature vector of a single frame The input is fed into a Bidirectional Gated Recurrent Unit (Bi-GRU) network; both the initial forward and backward hidden states are set to zero vectors, i.e. , ; At time step The forward network computes the forward update gate in ascending time order. With forward reset door : Physical defense logic: For users who are consistently observing cultural relics, the historical trajectory is continuous. The system tends to retain memories; when the interfering arm suddenly extends, the coordinates change drastically. The system was forcibly activated to disconnect historical connections, which was determined to be a trajectory fault. Then the forward candidate states are calculated. With the final forward hidden state : Simultaneously, the backward network performs symmetric computation in reverse temporal order (from frame 150 to frame 1) to obtain the backward update gate. Reverse reset door Backward candidate state And the final backward hidden state : Step 7: After calculation, extract the final forward hidden state of frame 150. Final backward hidden state with frame 1 The features are concatenated to construct a global temporal feature containing the complete context. : Subsequently, dimensionality reduction is performed using a fully connected layer to output the final temporal coherence feature vector. ; Step 8: Convert the grayscale image tensor Input the VGG network to extract spatial features, where the first... Convolutional layers calculate high-frequency gradient features: in This is the feature activation map of the previous layer; the features closest to the sensor are continuously preserved through a max pooling layer. Physical defense logic: Brighter pixels in the depth map represent those closer to the sensor. Max pooling mathematically preserves the physical protrusions closest to the sensor (such as the fingertips of the controller's hand), automatically filtering out interfering depth information from distant or edge locations. Step 9: After multiple convolutional blocks, the deep feature map is flattened, and spatial depth-level feature vectors are calculated through fully connected layers. ; Step 10: and Connected in the channel dimension Input MLP hidden layer: Step 11: By activating the sigmoid function of a single neuron in the output layer, the high-dimensional features are compressed into scalars, and the scalar score is output. : This step and the steps described above and Both are learnable parameter matrices and bias vectors obtained through pre-training; Step 12: Obtain the set of concurrent target scores Perform exponential moving average dejitting: The initial time hour, ; Step 13: Using user identifiers As an independent evaluation unit, calculate the comprehensive interaction intent score for each user. If the user channel only contains one hand, then If the user channel contains a "two-handed coordination group", then the higher intent score of the two hands is taken as the representative. ; Then, a user-level arbitration of control is conducted: 1) If there is only a single user in the current field of view (i.e., valid) (Total number is 1), the system performs computing power bypass optimization, only determining this... If the threshold is exceeded, the master status is established or maintained directly; otherwise, the master status is released. 2) If multiple users are concurrently accessing the current field of view (i.e., effective) total If a new user's overall score exceeds the preemption threshold, then a dual-threshold hysteresis mechanism is activated for anti-interference game theory: the new user seizes control when their overall score exceeds the preemption threshold; if a controlling user has already been established... As long as the user's overall score does not fall below the release threshold, the system will forcibly maintain the user's full control lock status. Step 14: Generate a user-level Boolean control mask set for all hand targets within the field of view. Filtering based on physical coordinates: All physical data of non-master users are completely muted at this moment, and the final output is the three-dimensional spatial coordinates of the gesture of the only legitimate user (including one hand or two hands in a legitimate pair), which serves as the sole driving data for system interaction.

2. The method according to claim 1, characterized in that, In step 2, the fixed sampling rate is 60fps; in step 3, the system traces back 150 consecutive frames along the time axis, corresponding to the construction of a historical time series data window with a length of 2.5 seconds.

3. The method according to claim 1, characterized in that, The minimum depth threshold of the effective field of view in step 4 With maximum depth threshold The effective physical distance range for non-contact interaction is set; while reducing dimensionality, the max pooling layer continuously extracts and retains the convex feature points that are closest to the infrared depth sensor in terms of physical distance, so as to filter out environmental depth interference from afar.

4. The method according to claim 1, characterized in that, In step 8, the VGG network uses single-channel adaptive convolution to extract grayscale tensor features. The feature extraction process includes five consecutive convolutional block operations, and each convolutional block is followed by the max pooling layer for depth feature filtering.

5. The method according to claim 1, characterized in that, In step 13, the parameters of the dual-threshold hysteresis circuit are configured as a set of asymmetric thresholds, including a preemption threshold and a release threshold, and the preemption threshold is strictly greater than the release threshold. When a new target appears in the set of concurrent targets, it can only take control if its interaction intent legitimacy score is higher than the preemption threshold. If a master user has been established, as long as its interaction intent legitimacy score does not fall below the release threshold, the system maintains the locked state of the master user, thereby forming an anti-interference protection zone.

6. The method according to claim 1, characterized in that, The method further includes a step of converting the unique legal gesture's three-dimensional spatial coordinates into the underlying rendering engine's coordinate system, specifically comprising: using a three-dimensional affine transformation to map the original coordinate system spatial data of the infrared depth sensor to the target coordinate system of the three-dimensional rendering engine; the coordinate transformation formula is: in , , These are the scaling matrix, rotation matrix, and translation vector, respectively. Output standard 3D spatial coordinates and control event packages to the underlying event system.

7. The method according to any one of claims 1 to 6, characterized in that, The standard three-dimensional spatial coordinates and control event package are output to the upper-level display system to smoothly drive the rotation, scaling and dissection display of the three-dimensional model of digital cultural heritage in non-contact interactive scenarios in public spaces.