A method, device and equipment for controlling an intelligent remote controller based on gesture recognition
Through technologies such as space-time dimension exchange and probability superposition verification, the problems of insufficient fusion of multi-sensors and solidification of communication modes in gesture recognition are solved, the recognition accuracy and communication efficiency are improved, and the efficient control of the intelligent remote control in complex environments is realized.
Patent Information
- Application Number
- CN202510887253.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing gesture recognition technology has problems such as insufficient fusion depth of multi-sensors, limited feature expression capabilities, simple identification decision-making mechanism, and solidified communication mode selection, resulting in poor recognition accuracy and user experience.
By integrating space-time dimension exchange, probability superposition verification, dynamic hierarchical reasoning and dual-mode communication competition technologies, a complete technical chain from multimodal perception to deep feature extraction to intelligent verification decision-making and adaptive communication transmission is built to improve gesture recognition accuracy and enhance environmental adaptability.
It improves the accuracy and anti-interference ability of gesture recognition, optimizes communication efficiency, and enhances the reliability and adaptability of intelligent remote control in diverse environments.
Smart Images

Figure CN120386456B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent interaction technology, and in particular to an intelligent remote control method, device and equipment based on gesture recognition. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and the Internet of Things (IoT) technologies, the popularity of smart home devices has continued to rise, and users' demands for a more user-friendly interactive experience have also grown. Gesture recognition, as a natural and intuitive method of human-computer interaction, enables device control by capturing and analyzing user hand movements, and has become a key development direction in smart remote control technology. Traditional gesture recognition methods primarily rely on single sensor technologies, such as camera-based visual recognition or inertial measurement unit-based motion sensing. These methods suffer from unstable recognition accuracy and weak anti-interference capabilities in complex environments.
[0003] While existing multi-sensor gesture recognition technology has improved recognition performance to a certain extent, it still faces numerous technical bottlenecks: The multimodal sensor data fusion strategy is simplistic and crude, lacking in-depth feature correlation analysis, making it difficult to fully exploit the complementary advantages of different sensors; gesture feature extraction methods are simplistic, ignoring the complex temporal and spatial correlations of gestures, resulting in insufficient feature expression; the recognition decision-making mechanism lacks adaptive capabilities, unable to dynamically adjust the recognition strategy based on signal quality and environmental complexity; and the communication mode selection mechanism is rigid, unable to intelligently match gesture type and communication environment status. These issues severely limit the recognition accuracy and user experience of smart remote controls in diverse application scenarios, necessitating the development of a new generation of gesture recognition technology with deep fusion, intelligent decision-making, and adaptive optimization capabilities. Summary of the Invention
[0004] This invention provides a method, device, and equipment for controlling an intelligent remote control based on gesture recognition. These methods aim to address key technical issues in existing gesture recognition technology, such as insufficient multi-sensor fusion depth, limited feature expression capabilities, simple recognition decision-making mechanisms, and rigid communication mode selection. By integrating innovative technologies such as spatiotemporal dimension exchange, probabilistic superposition verification, dynamic hierarchical reasoning, and dual-mode communication competition, a complete technology chain is constructed, from multimodal perception to deep feature extraction, intelligent verification and decision-making, and adaptive communication transmission. This approach improves gesture recognition accuracy, enhances environmental adaptability, and intelligently optimizes communication efficiency, resulting in an intelligent remote control solution featuring deep perception, intelligent reasoning, autonomous decision-making, and dynamic optimization.
[0005] A first aspect of the present invention provides a method for controlling an intelligent remote controller based on gesture recognition, comprising the following steps:
[0006] Acquire multimodal sensor data, perform feature extraction on the multimodal sensor data to obtain gesture spatiotemporal features, and perform spatiotemporal dimension exchange on the gesture spatiotemporal features to generate exchange dimension features;
[0007] Acquire a candidate gesture probability superposition state based on the exchange dimension feature, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate a reversal verification parameter based on the semantic negation hypothesis;
[0008] Performing weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and performing probability aggregation on the corrected probability state to determine a target gesture;
[0009] Extracting verification feature data based on the target gesture, inputting the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification, and outputting a gesture recognition result;
[0010] Detect the current communication environment state, build the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment state, aggregate the dual-mode communication competition state into the optimal communication mode at the moment the command is sent, and complete the intelligent control of the remote control.
[0011] A second aspect of the present invention provides an intelligent remote control device based on gesture recognition, comprising:
[0012] a feature processing module, configured to acquire multimodal sensor data, perform feature extraction on the multimodal sensor data to acquire spatiotemporal features of gestures, and perform spatiotemporal dimension exchange on the spatiotemporal features of gestures to generate exchange dimension features;
[0013] a probability construction module, configured to obtain a candidate gesture probability superposition state based on the exchange dimension feature, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis;
[0014] A state correction module, configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine a target gesture;
[0015] a hierarchical reasoning module, configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification, and output a gesture recognition result;
[0016] The communication control module detects the current communication environment status, constructs the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment status, and at the same time establishes a parallel processing mechanism for conflicting decisions between selecting Star Flash and selecting UWB. At the moment the command is sent, the dual-mode communication competition state is aggregated into the optimal communication mode to complete the intelligent control of the remote control.
[0017] The third aspect of the present invention proposes a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of a smart remote control control method based on gesture recognition disclosed in the first aspect are implemented.
[0018] The beneficial effects of the present invention are reflected in the following points: First, through innovative spatiotemporal dimension exchange and probabilistic superposition state construction technology, a deep feature extraction framework including dimensional decomposition, cross-mapping, and association matrix reconstruction is established. Combined with the semantic negation hypothesis and inversion verification mechanism, the core technical difficulties of the traditional method of single feature expression and lack of verification mechanism are solved, and a technological breakthrough is achieved from the shallow spatiotemporal representation to the deep association representation of gesture features, which improves the recognition accuracy, feature differentiation ability and verification reliability of complex gestures, enabling the system to accurately identify gesture movements with subtle differences and provide reliable verification guarantees.
[0019] Secondly, through an innovative dual verification mechanism that combines probability superposition verification with classification model verification, a multi-level verification system is constructed that includes probability statistical verification, classification algorithm verification, and signal confidence assessment. This breaks through the technical bottleneck of insufficient reliability of the single verification mechanism of the existing method, and realizes the technological leap from traditional single recognition to dual verification recognition, improving recognition accuracy, verification reliability and anti-interference ability, and providing multiple guarantees and reliability verification for complex gesture recognition.
[0020] Finally, through dual-mode communication competition and instantaneous convergence optimization technology, the intelligent selection, dynamic weight allocation and real-time switching capabilities of Star Flash and UWB communication modes are realized, which solves the key problems of traditional communication methods such as single mode selection, poor environmental adaptability and low switching efficiency, and forms a new generation of communication control scheme with environmental perception and intelligent decision-making characteristics, which improves the communication reliability, transmission efficiency and environmental adaptability of smart remote controls in diverse environments.
[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein illustrate specific examples of the technical solutions described in the present invention, and together with the specific implementation methods constitute a part of the specification, and are used to explain the technical solutions, principles and effects of the present invention.
[0023] Unless otherwise specified or defined, the same reference numerals in different drawings represent the same or similar technical features, and the same or similar technical features may also be represented by different reference numerals.
[0024] Figure 1 The present invention is a flowchart of a method for controlling an intelligent remote controller based on gesture recognition.
[0025] Figure 2 This is a structural block diagram of an intelligent remote control device based on gesture recognition in the present invention.
[0026] Figure 3 It is a structural schematic diagram of a computer device of the present invention. DETAILED DESCRIPTION
[0027] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0029] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] As used in this specification and the appended claims, the term "if" can be interpreted as meaning "uponce" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]," depending on the context.
[0031] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0033] The technical solutions of the embodiments of this application are introduced below.
[0034] like Figure 1 As shown, an embodiment of the present invention provides a method for controlling an intelligent remote controller based on gesture recognition, comprising the following steps S110 to S150:
[0035] Step S110 , acquiring multimodal sensor data, performing feature extraction on the multimodal sensor data to acquire gesture spatiotemporal features, and performing spatiotemporal dimension exchange on the gesture spatiotemporal features to generate exchange dimension features.
[0036] Specifically, multimodal sensor data acquisition relies on a high-precision multi-sensor fusion module that integrates a Time of Flight (ToF) camera, millimeter-wave radar, and an Inertial Measurement Unit (IMU) within the smart remote control. Each sensor is connected via a high-speed data bus, enabling synchronous data transmission and interaction. The ToF camera utilizes a 940nm near-infrared light source, features a depth perception range of 0.2-3.0 meters, a frame rate of 30fps, and a resolution of 64×64 pixels, capturing the three-dimensional contour and depth information of the hand. The millimeter-wave radar operates at 60GHz, has a detection range of 1.5 meters, an angular resolution of 5 degrees, and a velocity resolution of 0.1m / s. It is specifically designed to detect micro-gesture signals, including subtle movements such as finger sliding, clicking, and hovering. The IMU integrates a three-axis accelerometer and a three-axis gyroscope, with an acceleration measurement range of ±16g, an angular velocity measurement range of ±2000° / s, and a sampling frequency of 1kHz, enabling real-time tracking of the remote control's spatial posture changes and motion trajectory. The hardware layout of the multi-sensor fusion module has been carefully optimized. The ToF camera and millimeter-wave radar are integrated into the top area of the smart remote control, covering a 120-degree fan-shaped detection area in front, ensuring comprehensive and accurate capture of gestures in front. The UWB ultra-wideband module operates in the 6.5GHz frequency band, and the Star Flash communication module operates in the 2.4GHz frequency band. The two use an RF switch to achieve time-sharing multiplexing of the antenna. They are located at the bottom or side of the smart remote control to ensure stable transmission and reception of communication signals. Multimodal sensor data preprocessing uses advanced technologies such as spatiotemporal alignment, noise filtering, and data synchronization to add a unified timestamp to the ToF depth map, millimeter-wave radar point cloud, and IMU six-axis data. The PTP precise time protocol is used to synchronize the clocks of each sensor via a high-speed data bus with microsecond accuracy. Through multimodal data fusion processing, a synchronized multimodal sensor data stream containing complete gesture motion information is ultimately generated.
[0037] Feature extraction is performed on multimodal sensor data to obtain spatiotemporal features of gestures. The feature extraction process employs a submodal processing strategy, with specialized feature extraction algorithms designed for the data characteristics of different sensors. Feature extraction from ToF camera data utilizes a 3D keypoint detection algorithm. A deep learning model identifies the three-dimensional coordinates of 21 keypoints, including the palm center, fingertips, and joints. Each keypoint consists of three coordinate components, X, Y, and Z, forming a 63-dimensional hand geometric feature vector. The keypoint detection process first preprocesses the 64×64 pixel depth map, including noise filtering, edge enhancement, and depth correction. Spatial geometric features are then extracted using a 3D convolutional neural network. The network architecture comprises multiple layers of 3D convolutional layers, which progressively extract a hierarchical representation, from low-level edge features to high-level semantic features. Feature extraction from millimeter-wave radar data focuses on analyzing micro-motion frequencies and motion patterns. A short-time Fourier transform (SFT) is used to convert the time-domain radar signal into a time-frequency spectrum. Frequency-domain features such as peak frequency, frequency bandwidth, spectral centroid, and spectral roll-off are extracted from these features. These features can effectively distinguish different types of micro-motion gestures, such as fast swipes, slow taps, and hovering. IMU data feature extraction combines time-domain and frequency-domain analysis methods. In the time domain, statistical features of acceleration and angular velocity within a sliding window are calculated, including mean, variance, peak value, and zero-crossing rate. In the frequency domain, a FFT transform is used to obtain the power spectrum density, extracting the energy ratio and distribution characteristics of the main frequency components. Multimodal feature fusion technology is used to encode the coordinates of ToF keypoints, millimeter-wave radar frequency-domain features, and IMU motion statistics into a unified multi-channel tensor representation, forming a gesture spatiotemporal feature that contains complete spatiotemporal information.
[0038] In some embodiments, the performing spatiotemporal dimension exchange on the gesture spatiotemporal features to generate exchange dimension features includes: decomposing the gesture spatiotemporal features into a time dimension component and a space dimension component; cross-mapping the time dimension component with the space dimension component; reconstructing a dimension association matrix based on the cross-mapping result; and generating exchange dimension features using the dimension association matrix.
[0039] First, the spatiotemporal features of the gesture obtained above are decomposed into temporal and spatial components. The temporal component is extracted by sliding a one-dimensional convolution kernel along the time axis with a kernel size of 5 and a stride of 1, specifically extracting local patterns and changing trends in the time series. Global average pooling is used to compress the spatial features at each time step into scalar values, forming a pure time series. This series reflects the temporal evolution of the gesture, including temporal phases such as the onset, development, climax, and end of the gesture. The spatial component is extracted by independently processing the spatial features at each time step using a two-dimensional convolution with a kernel size of 3×3, specifically extracting spatial geometric structure and shape patterns. Temporal average pooling eliminates temporal variations while retaining pure spatial distribution information. This component reflects the spatial geometric characteristics and shape attributes of the gesture, such as the hand contour, key point distribution, and the geometry of the motion trajectory. In complex gesture recognition tasks, when a user draws a circle in the air, the temporal component captures the periodic velocity variation pattern, with peaks occurring at the turning points of the circle. The spatial component extracts the geometric features of the circular trajectory, including key parameters such as the center position, radius, and shape regularity. Based on dimensional decomposition, independent feature representations for the temporal and spatial components are generated.
[0040] Next, the temporal and spatial components are cross-mapped. This cross-mapping process uses an attention mechanism to calculate the correlation weights between the temporal and spatial components, generating a spatiotemporal correlation matrix through matrix multiplication. Each element of the correlation matrix represents the strength of the correlation between a specific time step and a specific spatial location. Larger values indicate stronger correlation, reflecting the importance of that spatiotemporal location combination for gesture recognition. To enhance the representation of correlation, a learnable mapping function is introduced to perform nonlinear transformations on the temporal and spatial components, respectively. This mapping function employs a multi-layer perceptron architecture with two hidden layers and a Reluctant Unit (ReLU) activation function, enabling it to learn complex nonlinear spatiotemporal correlation patterns. The cross-mapping process also considers spatiotemporal correlation patterns at different scales. Using a multi-scale cross-attention mechanism, correlation weights are calculated across different time windows and spatial neighborhoods, capturing spatiotemporal interaction patterns from fine-grained to coarse-grained. The multi-scale correlation weights are weightedly fused to generate the final cross-mapping result, with the weight coefficients optimized through back-propagation learning. In the "tap to confirm" gesture, cross-mapping reveals a strong correlation between the time of finger press and the spatial location of the contact area. The correlation weight peaks at the moment of contact and decreases significantly during the non-contact period. By using cross-mapping processing, the association relationship between the time dimension component and the space dimension component is established.
[0041] The dimensional correlation matrix is then reconstructed based on the cross-mapping results. The dimensional correlation matrix consists of four sub-matrix blocks: the temporal autocorrelation matrix, the spatial autocorrelation matrix, the time-to-space correlation matrix, and the space-to-time correlation matrix. The temporal autocorrelation matrix is obtained by calculating the similarity between features at different time steps, using the cosine similarity metric to reflect the intrinsic structure and periodicity of the time series. The spatial autocorrelation matrix calculates the correlation between features at different spatial locations, taking into account the combined influence of spatial proximity and feature similarity, and reflecting the geometric structure and continuity of the spatial distribution. The time-to-space correlation matrix and the space-to-time correlation matrix are generated from the cross-mapping results described above and normalized to ensure numerical stability and comparability. The dimensional correlation matrix is reconstructed using block matrix concatenation to form a complete description of the spatiotemporal dimensional correlations. To improve the matrix's expressive power, low-rank decomposition techniques are introduced to reduce redundant information. The correlation matrix is decomposed using singular value decomposition, retaining the primary singular value components to achieve dimensionality compression and noise filtering. In the analysis of the "slide to turn a page" gesture sequence, the dimensional correlation matrix exhibits a distinct diagonal block structure. The temporal autocorrelation matrix exhibits high correlation at the start and end of the slide, reflecting the consistency of the gesture. The spatial autocorrelation matrix shows strong correlation between adjacent positions along the slide path, demonstrating the continuity of the trajectory. Reconstructing the dimensional correlation matrix yields a structured matrix representation containing complete spatiotemporal interaction information.
[0042] Finally, the dimensional correlation matrix is used to generate exchanged dimension features. Matrix multiplication is used to multiply the dimensional correlation matrix by the flattened representation of the original spatiotemporal features to obtain preliminary exchanged dimension features. To maintain the physical meaning and geometric structure of the features, a dimension permutation operation is introduced, swapping the positions of the temporal and spatial dimensions. This allows the original temporal features to be expanded in the spatial dimension, and the spatial features to be extended in the temporal dimension. The permutation operation is implemented through tensor reshaping and dimension transposition, changing the dimensional organization of the feature tensor and providing the model with different perspectives on the features. The exchanged dimension features are also adaptively modulated through a gating mechanism. The gating unit learns the importance weight of each dimensional feature. The gating weights are generated using a sigmoid activation function to ensure that the weights are between 0 and 1. The final exchanged dimension features are obtained by weighted fusion of the original features and the exchanged features, achieving adaptive feature selection and fusion. In the "rotation gesture" recognition task, the exchanged dimension features map the angular velocity variation pattern in the original time series to a geometric description of the spatial rotation trajectory, enabling the model to understand the temporal rotation law from a spatial geometric perspective. Combined with the dimension permutation processing technique, exchanged dimension features with enhanced expressive power are obtained.
[0043] Step S120 , obtaining a candidate gesture probability superposition state based on the exchange dimension feature, constructing a semantic negation hypothesis for the candidate gesture probability superposition state, and generating an inversion verification parameter based on the semantic negation hypothesis.
[0044] Specifically, the deterministic gesture features are converted into a probabilistic superposition state, and a semantic negation hypothesis is constructed for reverse verification. Through the dual mechanisms of probabilistic superposition and semantic negation, the robustness and accuracy of gesture recognition are improved.
[0045] In some embodiments, the obtaining of the candidate gesture probability superposition state based on the exchange dimension feature includes: performing probability distribution analysis on the exchange dimension feature to obtain probability component data; constructing a multi-state energy field based on the probability component data; analyzing the multi-state energy field to generate energy distribution parameters; and generating the candidate gesture probability superposition state through the energy distribution parameters.
[0046] First, the swapped-dimensional features are analyzed for probability distribution to obtain probability component data. Using a Bayesian statistical framework, each dimension in the swapped-dimensional features is treated as a random variable, and its probability distribution characteristics are fitted using maximum likelihood estimation. For the hand keypoint features extracted by the ToF camera, a multivariate Gaussian distribution is used to model the uncertainty of their spatial coordinates. The covariance matrix reflects the accuracy and correlation of keypoint detection. For the millimeter-wave radar's micro-motion frequency characteristics, a mixture Gaussian model is used to describe the frequency distribution characteristics of different gestures, with each Gaussian component corresponding to a typical gesture pattern. For the IMU's motion trajectory characteristics, a Markov model is used to characterize the temporal transition probabilities of motion states, and the state transition matrix describes the transition relationship between different motion phases. The probability distribution analysis process also considers the effects of sensor noise and environmental interference, using a noise model to modify the theoretical distribution parameters to ensure the accuracy and robustness of the probability modeling. When the IMU detects an acceleration greater than 1.5g or the millimeter-wave radar detects a hand micro-motion speed greater than 0.5m / s, the variance parameters of the probability distribution are dynamically adjusted to accommodate the uncertainty characteristics of highly dynamic gestures. Using probability distribution analysis techniques, probability component data with statistical characteristics is generated.
[0047] Next, a multi-state energy field is constructed based on the probability component data, converting the probabilistic information into a representation in energy space. The multi-state energy field is constructed by mapping the probability distribution corresponding to each candidate gesture into a potential energy distribution in energy space. Using kernel density estimation, a Gaussian kernel function is used to generate a continuous energy distribution in feature space, using the probability component data as the core. The kernel function's bandwidth parameter is adaptively adjusted based on the variance of the feature dimension to ensure smoothness and local preservation of the energy field. Feature regions with high probability density have lower corresponding potential energy values, indicating a more stable gesture state. Regions with lower probability density have higher potential energy values, indicating an unstable or transitional state. The multi-state energy field also incorporates a multi-scale energy hierarchy, constructing corresponding potential energy distributions at different feature scales. The coarse-scale potential energy field captures the overall pattern of the gesture, while the fine-scale potential field describes local, detailed variations. The multi-state nature of the energy field is reflected in the fact that the same feature space location may correspond to multiple energy states, reflecting the ambiguity and ambiguity in gesture recognition. In the "wave goodbye" gesture recognition, the multi-state energy field can simultaneously describe two different energy states: a fast wave and a slow wave. Each energy state corresponds to a different probability weight and energy level. After potential energy field construction and processing, a multi-state energy distribution representation is formed.
[0048] The multi-state energy field is then deeply analyzed to generate energy distribution parameters. Energy distribution parameter extraction utilizes the partition function theory from statistical physics, quantifying the global and local characteristics of the energy distribution by calculating the statistical characteristics of the potential energy field. Global energy parameters, including total energy, average energy, energy variance, and energy entropy, reflect the overall distribution characteristics and stability of the potential energy field. Local energy parameters, including the modulus, direction, and rate of change of the energy gradient, are obtained through gradient analysis of the potential energy field, describing the local variation characteristics and distribution of stable points in the energy field. Stable point identification utilizes potential energy surface analysis, calculating the critical points of the potential energy field to identify energy minima, maxima, and saddle points. These key points correspond to different gesture states and transition paths. Energy barrier calculation assesses the difficulty of transitions between different gesture states. High energy barriers indicate high discrimination between gestures, while low energy barriers indicate areas prone to misidentification. Dynamic energy analysis considers temporal evolution and tracks the evolution of gestures through energy changes over time. In the "click to confirm" gesture analysis, the energy distribution parameters show a clear minimum energy value at the moment of finger contact, corresponding to a stable click state. However, higher energy values in the transition phase before and after contact reflect the instability of the gesture. Using an energy field analysis algorithm, we obtain energy distribution parameters that describe the potential energy distribution characteristics.
[0049] Finally, the energy distribution parameters are used to generate a probabilistic superposition state of candidate gestures. This probabilistic superposition state generation utilizes a quantum probability theory framework, expanding classical probabilities into complex probability amplitudes. This allows for quantum superposition and interference effects between different gesture states. Each candidate gesture corresponds to a probability amplitude. The square of the modulus of the probability amplitude represents the classical probability of the gesture, and the phase of the probability amplitude reflects the coherence relationship between gestures. The superposition state is constructed through linear combination, with the superposition coefficient determined by the aforementioned energy distribution parameters. Low-energy states receive a larger superposition weight, while high-energy states receive a smaller weight. The phase relationship is determined based on the semantic similarity and motion continuity between gestures. Similar gestures have similar phases, while significantly different gestures have phase differences that are nearly orthogonal. The probabilistic superposition state also considers temporal evolution, describing the temporal evolution of the superposition state using a discretized form of the Schrödinger equation. The evolution operator is determined by the Hamiltonian constructed from the energy distribution parameters. The collapse mechanism of the superposition state simulates the quantum measurement process. When external observation intervenes, the superposition state collapses to a definitive gesture recognition result. In the recognition of complex gesture sequences like "slide + click," the probabilistic superposition state can simultaneously maintain the probability amplitudes of both the slide and click gestures. This coherent superposition of probability amplitudes at the critical moment of transition provides more accurate recognition. Leveraging the quantum probabilistic superposition mechanism, a probabilistic superposition state representation of candidate gestures is constructed.
[0050] In some embodiments, constructing a semantic negation hypothesis for the candidate gesture probability superposition state includes: performing semantic reverse parsing on the candidate gesture probability superposition state; constructing an antonym semantic set based on the semantic reverse parsing result; and generating a semantic negation hypothesis using the antonym semantic set.
[0051] First, semantic reverse parsing is performed on the probabilistic superposition state of candidate gestures to extract the gesture's reverse semantic features. Based on contrastive learning theory, semantic reverse parsing enhances recognition discriminative capabilities by analyzing the "non-features" of each candidate gesture. The reverse parsing process constructs a negative sample space and systematically generates a corresponding negative representation for each positive gesture category. For the "right swipe" gesture, for example, its reverse semantics include all non-right swipe action patterns, including "left swipe," "up swipe," "down swipe," and "stand still." Reverse parsing employs feature inversion and semantic flipping to mathematically transform the feature vectors of the positive gesture and generate the corresponding reverse feature representation. For hand trajectory features captured by the ToF camera, reverse parsing generates reverse trajectory patterns through operations such as trajectory inversion, direction flipping, and velocity inversion. For motion acceleration features detected by the IMU, reverse parsing constructs reverse motion patterns through sign flipping, amplitude inversion, and frequency inversion. Semantic reverse parsing also considers reverse features in the temporal dimension, extracting temporal reverse semantics through methods such as reversing the time series, time scale transformation, and causal flipping. In the multimodal feature space, reverse parsing achieves a comprehensive semantic reversal through geometric operations such as symmetric transformation, complementary projection, and orthogonal decomposition of the feature space. Based on the reverse parsing algorithm, complete semantic reverse parsing result data is generated.
[0052] Next, an antonym set is constructed based on the semantic reverse parsing results. Using a hierarchical organizational structure, the antonym set is categorized and organized according to multiple dimensions, including semantic level, action type, and feature dimension. Top-level semantic antonyms include oppositions between basic action types, such as "still vs. moving," "approaching vs. moving away," and "pressing vs. lifting." Mid-level semantic antonyms involve oppositions between specific gestures, such as "left swipe vs. right swipe," "slide up vs. slide down," and "clockwise vs. counterclockwise." Bottom-level semantic antonyms address oppositions within specific action details, such as "fast vs. slow," "large vs. small," and "continuous vs. discontinuous." The antonym set also includes temporal antonyms, describing temporal oppositions within action sequences, such as "fast then slow vs. slow then fast," "crescendo vs. diminuendo," and "periodic vs. aperiodic." The construction of semantic sets incorporates ontological knowledge representation methods, describing the logical structure and constraints of antonym relationships through semantic networks. Each antonym semantic node contains attribute information such as semantic identification, feature description, opposition strength, and applicable conditions. The strength of antonym relationships is quantified through similarity metrics and contrast calculations. Strong oppositions have high contrast values, while weak oppositions have low contrast. In constructing the antonyms of the "clench fist + release" compound gesture, the antonym semantic set can identify various opposition patterns, such as "constantly clenching the fist," "constantly releasing the fist," and "relaxing the fist first, then clenching the fist," providing a rich semantic foundation for subsequent negation hypothesis generation. Through semantic organization and relationship modeling, a complete knowledge system of antonym semantic sets is established.
[0053] Finally, semantic negation hypotheses are generated using the antonym semantic set. Using logical reasoning and hypothesis testing, negation propositions are constructed based on the oppositional relationships within the antonym semantic set. For each candidate gesture, the system generates a series of negation hypotheses, describing the criteria for what the gesture "is not." Negation hypotheses are constructed using predicate logic, taking the form of a logical expression such as "NOT (gesture X has feature Y)." The hypothesis generation process considers multiple levels of negation, including categorical negation, feature negation, temporal negation, and combined negation. Categorical negation hypotheses address basic gesture types, such as "the current gesture is not a click" or "the current gesture is not a slide," representing high-level semantic negations. Feature negation hypotheses address specific feature attributes, such as "the hand movement speed does not exceed a threshold" or "the movement trajectory is not circular," representing detailed feature negations. Temporal negation hypotheses focus on the temporal characteristics of the action, such as "the action duration is not less than a minimum value" or "the action frequency is not within a specific range," reflecting temporal constraints. Combined negation hypotheses involve the joint negation of multiple features, constructing compound negation conditions through logical connectives. The generation of negative hypotheses also incorporates uncertainty quantification. Each negative hypothesis is associated with a confidence score, reflecting the reliability of the negative judgment. The confidence calculation takes into account factors such as the strength of feature evidence, the success rate of historical verification, and the clarity of the semantic opposition. In generating negative hypotheses for the "rotation gesture," the system constructed multiple negative hypotheses, including "not linear motion," "not static," and "not unidirectional motion." Each hypothesis has a corresponding confidence score, providing a basis for subsequent verification. Using logical reasoning and confidence calculation, the systematic generation of semantic negative hypotheses is achieved.
[0054] In some embodiments, generating the inversion verification parameters based on the semantic negation hypothesis includes: converting the semantic negation hypothesis into verification constraints; analyzing the verification constraints to obtain reverse verification indicators; performing parameter quantization processing based on the reverse verification indicators to generate inversion verification parameters.
[0055] First, semantic negation hypotheses are converted into verification constraints, establishing an operational verification framework. Constraint conversion utilizes a mapping method from symbolic logic to numerical logic. Using a predicate logic parser and mathematical expression generator, the logically formalized negation hypotheses are converted into numerically constrained inequalities or equalities. The conversion process consists of four processing stages: syntax parsing, semantic analysis, mathematical modeling, and constraint generation. Syntax parsing identifies the logical structure and keywords in the negation hypothesis, while semantic analysis extracts the meaning of the negation relationship and the constraint range. For categorical negation hypotheses, the converter uses a probability threshold mapping mechanism to convert semantic negation into numerical constraints on classification probabilities. For example, "gesture is not a click" is converted into a numerical constraint "click category probability is less than a threshold" through a probability inversion algorithm. The threshold is dynamically determined based on historical statistical data and error rate analysis. For feature negation hypotheses, the converter uses a feature space boundary definition method to convert feature attribute negation into feature value range constraints. For example, "motion speed does not exceed a threshold" is converted into an interval constraint "speed feature value is less than or equal to an upper bound" through velocity feature analysis. The upper bound is determined through kinematic analysis and sensor accuracy assessment. For timing negation assumptions, the converter utilizes temporal constraint modeling techniques to convert time-related negation conditions into constraints based on temporal parameters. For example, "action duration must be no less than a minimum value" is converted to a "time length greater than or equal to a lower bound" constraint through temporal analysis. The time bounds are derived from motion pattern analysis and user behavior statistics. The constraints are deeply integrated into the multi-sensor signal fusion verification mechanism, establishing a dynamic mapping between sensor state and constraint strength. When the IMU detects acceleration greater than 1.5g, the motion intensity assessment algorithm increases the verification strength of the corresponding constraint for the "non-stationary assumption," automatically tightening the verification threshold. When the ToF camera detects a hand within 1.5m, the distance sensing mechanism activates the "entering the interaction zone" constraint, enabling close-range precision verification. The constraints also incorporate a complex logical combination mechanism. Through Boolean algebra operations and logical expression optimization, logical connectives such as AND, OR, and NOT are used to construct multi-level, complex constraints, reflecting the collaborative verification relationships and interdependencies between multiple negation assumptions. The dynamic constraint adjustment mechanism utilizes adaptive parameter control technology. Based on real-time sensor signal quality assessment, ambient noise level detection, and recognition accuracy monitoring, it adaptively modifies constraint parameters and verification thresholds to ensure verification accuracy and real-time performance. After a complete mathematical transformation, a complete and directly computable verification constraint system is formed.
[0056] Next, the verification constraints are analyzed in depth to obtain reverse verification indicators. Reverse verification indicator extraction utilizes a multi-dimensional constraint satisfaction problem-solving method. Through constraint propagation algorithms, linear programming solvers, and heuristic search techniques, key characteristics of the constraints, such as satisfiability, consistency, and redundancy, are analyzed to extract verification-related quantitative indicators. Satisfiability indicator evaluation is implemented through a constraint solving engine. An algorithm combining backtracking search and constraint propagation is used to evaluate whether a set of constraints has a feasible solution. The feasible region of the constraint space is calculated using the linear programming simplex method or interior point method to quantify the degree of satisfiability of the constraint system and the size of the solution space. Consistency indicator detection utilizes conflict analysis and logical reasoning techniques. Through a constraint conflict detection algorithm and a logical consistency verifier, logical conflicts and contradictions between constraints are identified. A constraint dependency graph is constructed using graph theory methods to identify conflicting constraint pairs and conflict propagation paths. The overall consistency score and local conflict intensity of the constraints are calculated. Redundancy analysis utilizes constraint independence testing and sensitivity analysis. Constraint removal testing and impact calculations are used to analyze the independence and necessity of constraints. Gradient analysis and partial derivative calculations are used to determine the contribution and influence weight of each constraint to the overall verification performance. Verification strength quantification utilizes constraint tightness analysis and coverage assessment techniques. The geometric distance of the constraint boundaries and the volume ratio of the constraint space are calculated to assess the verification capability and strength of the constraints. The tightness of the constraint boundaries and the coverage of the feasible region are used to quantify the rigor and comprehensiveness of the verification. Verification robustness evaluation utilizes perturbation analysis and stability testing. Noise injection and parameter perturbation experiments are used to assess the sensitivity of constraints to sensor noise and environmental disturbances. Monte Carlo simulations and stability analysis are used to calculate the stability coefficient and anti-interference capability score of the constraint verification. Timeliness evaluation utilizes computational complexity analysis and performance testing techniques. Through algorithmic complexity theory analysis and actual runtime measurements, the computational complexity and response time of constraint verification are evaluated to ensure the feasibility and response speed requirements of real-time verification. In the verification analysis of the complex "circle in the air + click" gesture, reverse validation metrics showed high consistency scores between the periodic constraints of the circle and the instantaneous constraints of the click, indicating that the two constraints logically support each other. However, there was a certain logical tension between the continuity constraint of the circle and the discrete constraint of the click, requiring time-series segmented validation and action phase division to resolve the constraint conflict. A multi-dimensional constraint analysis algorithm was used to obtain comprehensive and quantitative reverse validation metric data.
[0057] Finally, parameter quantization is performed based on the reverse validation indicators to generate reverse validation parameters. Parameter quantization utilizes a multi-criteria decision analysis approach. Through indicator weight calculation and numerical fusion techniques, the multi-dimensional reverse validation indicators are converted into standardized numerical parameters that can be directly used for validation calculations. The quantization process begins with indicator normalization. Using normalization techniques such as minimum-maximum normalization and standardized scaling, validation indicators with different dimensions, numerical ranges, and distribution characteristics are uniformly mapped to a standard interval, ensuring comparability, numerical stability, and computational accuracy across different indicators. Satisfiability indicators are normalized to the [0, 1] interval using a linear mapping function. Consistency indicators undergo logarithmic transformation to eliminate numerical deviations. Redundancy indicators achieve interval normalization using a piecewise linear function. The weight assignment method uses the Analytic Hierarchy Process (AHP) to determine the importance and priority ranking of different verification indicators. By constructing an indicator importance hierarchy and calculating a judgment matrix, the satisfiability indicator receives a weight of 0.35 due to its direct impact on the verification results; the consistency indicator receives a weight of 0.30 due to its logical integrity requirement; the verification strength indicator receives a weight of 0.20 due to its constraint assessment; the redundancy indicator receives a weight of 0.10 due to its auxiliary role; and the verification robustness indicator receives a weight of 0.05. Parameter fusion is achieved through weighted linear combination and nonlinear fusion functions, combining multiple verification indicators into a comprehensive verification parameter based on their weights and correlations. The fusion process uses a combination of weighted arithmetic and geometric mean methods to consider the correlation and complementarity between indicators and avoid information redundancy and bias amplification. The verification threshold is set using statistical analysis of historical data. By analyzing the success rate and error rate of verification results, the optimal judgment threshold and decision boundary are determined to balance the sensitivity and specificity of verification. The standard verification threshold is set at 0.75, the strict verification threshold is set at 0.85, and the relaxed verification threshold is set at 0.65. The adaptive adjustment mechanism dynamically adjusts parameter values and threshold settings based on real-time verification performance monitoring and accuracy statistics. When the verification accuracy drops by more than 5%, the verification threshold is appropriately tightened by 0.05. When the false alarm rate rises by more than 10%, the verification conditions are appropriately relaxed by lowering the threshold by 0.03. The adjustment range is determined through threshold sensitivity analysis. The time decay factor of the parameter adopts an exponential decay function. Considering the timeliness and dynamic changes of verification information, the time decay coefficient is set to 0.95. Newer verification evidence is given a higher influence weight through the time weight function, and the influence of historical verification information decreases exponentially with the time interval, ensuring the timeliness and adaptability of the parameters. In the multimodal sensor fusion verification, through sensor reliability assessment and verification contribution analysis, the verification parameter weight of the ToF camera is set to 0.4, the verification parameter weight of the millimeter-wave radar is set to 0.35, and the verification parameter weight of the IMU is set to 0.25. The weight distribution fully reflects the reliability differences, accuracy levels, and applicable scenarios of different sensors in gesture verification.Quantization processing technology and parameter optimization methods are used to accurately generate inversion verification parameters.
[0058] Step S130 , weighting the candidate gesture probability superposition state based on the inversion verification parameter to generate a modified probability state, and performing probability aggregation on the modified probability state to determine the target gesture.
[0059] Specifically, the verification strength of the inverted verification parameters is used to dynamically adjust the probability weights of candidate gestures, effectively suppressing unreliable probability components of failed verification while enhancing highly reliable probability components of successful verification. The weighted probability state is gradually converged through a multi-stage aggregation algorithm, ultimately determining a unique target gesture recognition result from multiple competing candidates.
[0060] In some embodiments, the weight correction of the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state includes: using the inversion verification parameter to generate a weight correction coefficient; weight adjustment of the candidate gesture probability superposition state based on the weight correction coefficient to obtain adjusted probability data; analyzing the adjusted probability data to generate a stability evaluation result; and generating a corrected probability state based on the stability evaluation result.
[0061] First, the inverted verification parameters are used to generate weight correction coefficients. An adaptive mapping function is used to map the numerical range of the inverted verification parameters to the weight correction coefficient range. For gesture candidates with high verification strength, the corresponding correction coefficients are close to 1.0, maintaining their original probability weights. For candidates with low verification strength, the correction coefficients are significantly less than 1.0, significantly reducing their probability weights. The calculation of the correction coefficients takes into account the fusion of multi-sensor verification results. When the IMU detects acceleration greater than 1.5g and the millimeter-wave radar detects hand micro-movements greater than 0.5m / s, the correction coefficients for dynamic gestures receive an additional boost. When the ToF camera detects a hand entering the interaction area at a distance of less than 1.5m, the correction coefficients for static and fine manipulation gestures receive a corresponding boost. The correction coefficients also incorporate a time decay mechanism, whereby recent verification evidence receives higher correction strength, while the correction effect of historical verification information decays exponentially over time. The nonlinear correction function uses a sigmoid activation to ensure that the correction coefficients vary within a reasonable range, avoiding system instability caused by extreme values. Multimodal correction coefficients are weighted and fused to generate the final combined correction coefficient. The ToF camera verification result is weighted at 0.4, the millimeter-wave radar verification result is weighted at 0.35, and the IMU verification result is weighted at 0.25. In the "rotation gesture" verification, when the verification parameters of all sensors support the rotation hypothesis, the combined correction coefficient reaches 0.92, significantly increasing the probability weight of the rotation gesture. A complete set of weighted correction coefficients is ultimately generated through a verification strength mapping algorithm.
[0062] The candidate gesture probability superposition state is then weighted based on the weight correction coefficient to obtain the adjusted probability data. The adjustment process uses element-wise multiplication, multiplying the probability amplitude of each gesture candidate by the corresponding correction coefficient to achieve precise adjustment of the probability weights. This adjustment process maintains the quantum coherence properties of the probability superposition state, modifying only the modulus of the probability amplitude while preserving the phase relationship, ensuring that the coherent interference effect between gestures is preserved. For complex probability amplitudes, the weight adjustment applies to both the real and imaginary parts, maintaining a constant phase angle through proportional scaling. Normalization ensures that the adjusted probability superposition state satisfies the probability conservation condition, with the sum of the squared moduli of all probability amplitudes equal to 1. A dynamic adjustment strategy adaptively adjusts the adjustment strength based on real-time sensor signal quality. Standard adjustment mode is used when the sensor signal confidence level is above 85%, while conservative adjustment mode is activated when the confidence level is below 85%, reducing the aggressiveness of the adjustment. Multi-scale adjustment considers verification results from different time windows: short-term window verification results are used for rapid response, while long-term window verification results are used for stability assurance. The adjustment process also incorporates a robustness protection mechanism. When the correction coefficient for a gesture candidate is too small, the minimum probability weight is retained to prevent complete loss of valid information. In the weight adjustment of the complex "slide + click" gesture sequence, the probability weight of the slide phase receives a correction factor of 0.85 based on motion continuity verification results, while the click phase receives a correction factor of 0.78 based on pressure sensing verification results. The adjusted probability data more accurately reflects the actual execution status of the action. This precise weight adjustment process generates verified and corrected adjusted probability data.
[0063] The adjusted probability data is then analyzed in depth to generate stability assessment results. A multi-dimensional analysis framework is used to comprehensively evaluate the stability of the adjustment effect from the perspectives of uniformity, concentration, variability, and temporal consistency of the probability distribution. Distribution uniformity analysis evaluates the distribution of probability weights across gesture candidates by calculating the entropy and Gini coefficient of the probability distribution. High entropy values indicate a relatively uniform probability distribution, resulting in greater recognition uncertainty; low entropy values indicate a concentration of probabilities among a small number of candidates, resulting in relatively certain recognition results. Concentration analysis evaluates the prominence of the dominant gesture by calculating the ratio of the maximum probability value to the second-highest probability value. A high ratio indicates a clear dominant gesture and good recognition stability; a low ratio indicates intense competition among multiple candidates, requiring further verification. Variability analysis evaluates the dispersion of the probability distribution by calculating the probability variance and standard deviation. Low variability indicates a stable probability distribution, while high variability indicates significant fluctuation. Temporal consistency analysis evaluates the temporal stability of recognition results by calculating the correlation of probability distributions within consecutive time windows. High correlation indicates consistent recognition results over time, while low correlation indicates temporal fluctuations. Confidence interval calculation quantifies the uncertainty of probability estimates. Confidence intervals for probability estimates are calculated using bootstrapping or Bayesian inference, providing a reliable reference for decision-making. Stability threshold assessment determines the criteria for stability assessment based on historical statistical data. When multiple stability indicators exceed the threshold, the state is considered stable; otherwise, it is marked as unstable and requires additional verification. In the stability assessment of the "clench and release" gesture, the probability distribution entropy value during the clench phase is 0.3, indicating high recognition certainty. The probability distribution entropy value during the release phase rises to 0.7, reflecting increased uncertainty during the transition period, prompting adjustments to subsequent verification strategies. Comprehensive stability assessment results are established based on multidimensional stability analysis.
[0064] Finally, a revised probability state is generated based on the stability assessment results. Appropriate processing strategies are implemented based on the different stability assessment results. When the stability assessment indicates high stability, the adjusted probability data is retained as the revised probability state to ensure rapid output of stable identification results. When the stability assessment indicates moderate stability, a probability smoothing algorithm is activated to smooth the time series of the adjusted probability data, using methods such as sliding average or exponential smoothing to reduce probability fluctuations and improve state stability. When the stability assessment indicates low stability, the probability reconstruction mechanism is activated, combining historical probability state information with the current adjustment results for probability fusion. The probability state is re-estimated using methods such as Bayesian updating or Kalman filtering. The revised probability state also includes confidence annotations, with each probability component associated with a corresponding confidence score, reflecting the reliability of the probability estimate. An anomaly detection mechanism monitors the rationality of the revised probability state and triggers a revalidation process when unusual patterns in the probability distribution appear. A state cache mechanism stores the most recent revised probability state sequence, providing historical data support for subsequent time series analysis and trend prediction. In actual applications of smart remote controls, the revised probability state assigns a probability weight of 0.87 to the "swipe right" gesture and a probability weight of 0.13 to the "stationary state," with confidence levels of 0.92 and 0.78, respectively. This provides a reliable data foundation for subsequent probability aggregation and final decision-making. Combined with a stability assessment mechanism, a fully calibrated revised probability state is obtained.
[0065] The corrected probability state is probabilistically aggregated to determine the target gesture. Probabilistic aggregation employs a multi-stage convergence algorithm to gradually converge the multiple probability components of the quantum superposition state into a single, deterministic recognition result. The aggregation process begins with probability sorting, arranging all gesture candidates from high to low according to their corrected probability weights. This identifies the primary candidate with the highest probability and the second-highest probability alternative candidate. Competition analysis assesses the probability gap between the primary and alternative candidates. When the primary candidate's probability advantage exceeds a preset threshold, it is identified as the target gesture. When the probability gap is smaller, a refined verification process is initiated. Refined verification further distinguishes competing candidates by increasing verification dimensions and improving verification accuracy. It also utilizes deep features and temporal correlation information from multi-sensor signals for supplementary verification. Semantic consistency verification ensures that the aggregated results adhere to the logical constraints of the gesture semantics. The semantic knowledge base verifies the rationality and interpretability of the recognition results. Temporal coherence verification assesses the coherence of the current recognition result with the historical recognition sequence to avoid logically discontinuous gesture transitions. The aggregation algorithm also incorporates a decision confidence calculation, assigning a confidence score to the final recognition result to reflect the reliability of the aggregated decision. Adaptive threshold adjustment dynamically adjusts the aggregation threshold based on real-time recognition results, raising it for high-precision scenarios and lowering it for real-time response. In a smart remote control application, when a "swipe right" target gesture is recognized with a confidence level of 0.89, the corresponding device control command is immediately generated, completing the complete control chain from gesture recognition to device response. A probabilistic aggregation algorithm is used to determine a unique target gesture recognition result.
[0066] Step S140 : re-extracting verification feature data based on the target gesture result, inputting the verification feature data into a dynamic hierarchical reasoning model including a primary classification model and a secondary classification model for classification verification, and outputting a gesture recognition result.
[0067] Verification feature data is re-extracted based on the target gesture results. A differentiated feature extraction strategy is used for secondary verification of the target gestures derived from probabilistic aggregation. This strategy re-examines the different dimensional features of the original sensor data to provide independent validation for the classification model. The key areas and dimensions for feature extraction are determined based on the type, confidence, and execution parameters of the target gesture results. The feature reconstruction process also extracts key motion parameters, including IMU acceleration intensity features, millimeter-wave radar hand micro-motion velocity features, multi-sensor signal confidence indicators, and ToF camera hand distance features. For dynamic gestures such as "swiping" and "waving," the focus is on extracting IMU motion trajectory features and millimeter-wave radar micro-motion frequency features to enhance the representation of temporal dynamic information. For static gestures such as "clicking" and "hovering," the focus is on extracting ToF camera spatial positioning features and shape contour features to enhance the accuracy of spatial geometric information. The feature reconstruction process retraces the original multimodal sensor data and accurately extracts the target gesture's temporal window and spatial range, ensuring high consistency between the verification feature data and the identified target. The IMU acceleration feature reflects the intensity of gestures, the millimeter-wave radar micromotion velocity feature describes the velocity changes of hand movements, the multi-sensor confidence index assesses the reliability of feature quality, and the ToF distance feature determines the spatial relationship between the hand and the device. When the target gesture is a "swipe right," the focus is on extracting the IMU acceleration change sequence, the millimeter-wave radar horizontal motion spectrum, and the ToF camera's hand trajectory key points during the swipe period. The corresponding acceleration intensity, micromotion velocity, confidence, and distance parameters are simultaneously calculated to form a dedicated feature set for swipe verification. The verification feature data also includes quality assessment metrics such as feature completeness, time synchronization, and sensor consistency. Through feature reconstruction, high-quality verification feature data containing motion parameters is generated.
[0068] In some embodiments, the verification feature data is input into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification, and a gesture recognition result is output, including: when the verification feature data meets the preset classification conditions, the primary classification model is triggered to perform preliminary classification to obtain a preliminary classification result; when the verification feature data does not meet the preset classification conditions, the secondary classification model is triggered to perform in-depth analysis to obtain a deep classification result; and a gesture recognition result is generated based on the above classification results.
[0069] When the verification feature data meets the preset classification conditions, the primary classification model is triggered to perform preliminary classification and obtain the preliminary classification results. The preset classification conditions include any one of the following conditions in the verification feature data: IMU acceleration intensity greater than 1.5g, millimeter-wave radar hand micro-motion speed greater than 0.5m / s, multi-sensor signal confidence greater than 85%, and ToF camera hand distance greater than 1.5m. The primary classification model uses a lightweight 1D-CNN architecture to specifically handle high-dynamic gesture recognition tasks that meet the classification conditions, achieving fast response while ensuring recognition accuracy. The model inputs the IMU time domain features and millimeter-wave radar frequency domain energy features in the verification feature data, and the feature dimensions are compressed to 128 dimensions to optimize computational efficiency. The network architecture consists of two one-dimensional convolutional layers and one fully connected layer. The first convolutional layer uses 64 kernels with a kernel size of 5 to extract local temporal patterns in the validation feature data. A max pooling layer performs downsampling to reduce computational overhead. The second convolutional layer uses 128 kernels with a kernel size of 3 to extract higher-level abstract features. The final fully connected layer contains 256 neurons and outputs four basic gesture classification results: left swipe, right swipe, tap, and pause. A softmax activation function converts the network output into a probability distribution, with each gesture category corresponding to a probability value, forming a complete gesture probability vector. The ReLU activation function enhances the network's nonlinear representation and prevents the vanishing gradient problem. The model is trained using the cross-entropy loss function and the Adam optimizer with a learning rate of 0.001 and a batch size of 32. In practical applications, for example, when the validation feature data indicates an IMU acceleration of 1.5g, the first-level classification model is triggered, meeting the classification criteria. The model can quickly identify the "right swipe" category with a confidence level of 0.94 and a response time of less than 10 milliseconds. After lightweight network processing, preliminary classification results are quickly obtained.
[0070] Then, when the verification feature data does not meet the preset classification conditions, the secondary classification model is triggered to perform in-depth analysis and obtain the in-depth classification results. When the verification feature data does not meet any of the aforementioned preset classification conditions, a more complex secondary classification model needs to be enabled. The secondary classification model uses the spatiotemporal frequency fusion network ST-FFN architecture to synchronously process the ToF depth map sequence, millimeter wave radar spectrum map, and IMU trajectory features in the verification feature data to output a composite gesture semantic result. The ST-FFN model contains three specialized feature extraction branches and a fusion layer. The ToF branch uses a 3D convolutional neural network to extract the spatiotemporal features of the hand depth map in the verification feature data. The input is a 16-frame 64×64 pixel depth map sequence. The first layer of 3D convolution contains 16 output channels, the convolution kernel size is 3×3×3, the stride is 1, and the padding method is the same, followed by ReLU activation and 3D maximum pooling; the second layer of 3D convolution contains 32 output channels, the convolution kernel size is 3×3×3, and the pooling size is 2×2×2; the third layer of 3D convolution contains 64 output channels, the convolution kernel size is 3×3×3, and finally outputs a 64-dimensional spatiotemporal feature vector through global average pooling to preserve the spatiotemporal continuity of hand movement. The radar branch uses a 2D convolutional neural network to process the micro-Doppler spectrograms in the verification feature data. The first layer of 2D convolution has 32 output channels, a 3×3 kernel size, followed by ReLU activation and 2×2 max pooling. The second layer has 64 output channels, and the third layer has 128 output channels. Finally, global average pooling is used to output a 128-dimensional spectral feature vector. The IMU branch uses a long short-term memory (LSTM) network to encode the six-axis motion trajectory in the verification feature data. The input is a time series consisting of three-axis acceleration and three-axis angular velocity. The LSTM has 64 hidden units and outputs a 64-dimensional trajectory encoding feature. The fusion layer weights the three features together using a modality-level channel attention mechanism, resulting in a total feature dimension of 256. The attention weight generation network consists of a fully connected layer (FC(256→128), ReLU activation, FC(128→3), and a softmax layer to generate the trimodal weight distribution. Dropout = 0.3 is used in the final classification layer to prevent overfitting. When processing a "swipe + tap" compound gesture, the ST-FFN model accurately identifies the trajectory characteristics of the swipe phase and the pressure changes during the tap phase when the verification feature data indicates a slow, complex motion at close range. Based on the deep fusion network architecture, it achieves deep classification results for complex gestures.
[0071] Finally, gesture recognition results are generated based on the aforementioned classification results. The result generation process first identifies the classification path to determine whether the current recognition result originates from the primary or secondary classification model, and records the triggering conditions and processing path information. For the preliminary results of the primary classification model, probability distribution analysis is used to ensure recognition quality. The result is adopted when the confidence score of the most probable gesture category exceeds 0.85 and the difference between it and the second-most probable category is greater than 0.3. Otherwise, a result stability check is performed. For the in-depth analysis results of the secondary classification model, a semantic rationality check is performed to ensure the logical consistency and feasibility of the composite gesture. A result fusion mechanism handles special cases where both the primary and secondary models produce valid results. The final recognition result is determined through a weighted fusion or voting mechanism, with weight allocation taking into account the model's applicability and the complexity of the current scenario. A temporal consistency check assesses the coherence of the current recognition result with the historical recognition sequence, filtering out unreasonable recognition jumps using state transition probabilities and gesture grammar rules. The output contains complete information, including gesture category, confidence score, execution parameters, and timestamp. Gesture categories use standardized encoding, confidence scores reflect recognition reliability, and execution parameters include quantitative metrics such as the gesture's spatial extent, movement speed, and duration. An anomaly detection mechanism monitors the rationality of recognition results and triggers re-recognition when unusual patterns occur. In a scenario where a smart remote control controls a TV, when a "swipe right" gesture is recognized with a confidence level of 0.91, a complete recognition result is generated, including the gesture category "RIGHT_SWIPE," a confidence level of 0.91, a swipe distance of 15 cm, and a swipe speed of 0.8 m / s. This supports the subsequent generation of device control commands. Through a unified result generation mechanism, combining probabilistic aggregation verification with classification model validation, reliable final gesture recognition results are output.
[0072] Step S150, detect the current communication environment state, construct the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment state, aggregate the dual-mode communication competition state into the optimal communication mode at the moment the command is sent, and complete the intelligent control of the remote control.
[0073] Specifically, the system detects the current communication environment status and constructs a dual-mode communication state for Starflash and UWB based on gesture recognition results and the communication environment status. Communication environment status detection utilizes multi-sensor fusion monitoring technology to assess the operating environment quality of both Starflash and UWB communication modes in real time. Starflash environment detection obtains real-time Starflash communication performance parameters through signal strength scanning, network congestion analysis, and frequency band occupancy monitoring. A signal strength exceeding -70dBm and a network latency below 50ms are considered excellent. UWB environment detection assesses the environmental adaptability of UWB communication through impulse response measurement, multipath interference analysis, and positioning accuracy verification. A suitable environment is determined when the positioning error is less than 10cm and the signal-to-noise ratio is greater than 15dB. Gesture recognition result analysis extracts communication requirement characteristics based on the gesture category, confidence level, and execution parameters output by gesture recognition. Simple gestures such as "click" require low-latency and fast response, while complex gestures such as "slide + rotate" require high-bandwidth parameter transmission. And sophisticated gestures such as "fine-tune" require high-precision positioning support. The dual-mode communication competition state is established by cross-analyzing the environmental adaptability of Starflash and UWB with gesture communication requirements through performance matching calculations, generating dynamic competition weights. When the gesture is a "precise positioning operation" and the UWB environment is excellent, UWB receives a higher competition weight. When the gesture is a "batch device control" and the Starflash network is stable, Starflash gains an advantage. The competition equilibrium point is determined through real-time weighted calculations, and the weight distribution is dynamically adjusted based on environmental changes, demand changes, and historical performance. By matching environmental perception with demand, a dynamically balanced dual-mode communication competition state is established.
[0074] In some embodiments, aggregating the dual-mode communication contention into the optimal communication mode at the moment of instruction sending includes: detecting an instruction sending trigger signal; performing instantaneous convergence on the dual-mode communication contention based on the trigger signal to obtain convergence result data; determining a unique optimal solution through the convergence result data; and converting the unique optimal solution into the optimal communication mode.
[0075] First, the command transmission trigger signal is detected, establishing a time synchronization mechanism for communication mode decisions. Trigger signal detection utilizes a multi-level monitoring framework consisting of primary, auxiliary, and forced triggers. This framework monitors user gesture recognition completion events, remote control status changes, and external control command events in real time. The gesture recognition completion trigger, serving as the primary trigger, monitors state changes in gesture recognition results. This involves three key detection criteria: confidence threshold detection, gesture category confirmation, and execution parameter verification. A primary trigger signal is generated when the recognition confidence exceeds 0.85, the gesture category is confirmed, and the execution parameters are complete. The remote control status trigger, serving as an auxiliary trigger, monitors power, connection, and operating mode changes, including transitions from standby to active, disconnected to connected, and silent to responsive. State change detection is implemented through both hardware interrupts and software polling. The external control trigger, serving as a forced trigger, responds to user physical button presses, voice commands, or other interactions, providing a forced trigger mechanism for emergency control scenarios. External triggers have the highest priority and can interrupt the current processing flow. Trigger signal priority management utilizes a preemptive scheduling strategy to ensure that multiple triggers are processed according to the preset priority when they occur simultaneously. Gesture recognition triggers have the highest priority, followed by state triggers. External control triggers can preempt all other triggers in emergencies. Trigger time window control uses a sliding time window mechanism to prevent frequent trigger events from affecting stability. The minimum trigger interval is set at 50 milliseconds. Consecutive trigger events are time-merged, and similar triggers are processed only once within the time window. Signal strength assessment determines trigger reliability through a multi-dimensional analysis of trigger event confidence, stability, and consistency. Only trigger signals with strength exceeding a dynamically adjusted threshold initiate subsequent convergence processing. In a TV remote control scenario, when the user completes a "swipe right" gesture and the recognition confidence reaches 0.91, the gesture recognition completion event immediately generates a trigger signal, initiating the convergence process for communication mode selection. Multi-level monitoring technology is utilized to accurately capture command transmission trigger signals.
[0076] Then, based on the trigger signal, the dual-mode communication competition state is instantaneously converged, and the convergence result data is obtained. The instantaneous convergence algorithm acts on the established dual-mode communication competition state, using a strategy that combines forced convergence and natural convergence to converge the competition state between star flash and UWB from an uncertain coexistence state to a certain single selection state. The convergence process is achieved through a competition state resolution mechanism, which includes three stages: competition state breaking, direction selection, and state solidification. When the trigger signal arrives, the dynamic equilibrium state of the dual-mode communication competition state is instantly broken, forcing the competition state to converge to a single mode. The convergence algorithm adopts the principle of minimizing the potential energy of the competition state. By calculating the potential energy distribution of star flash and UWB in the current competition state, it analyzes the competition intensity, stability, and convergence trend of star flash and UWB in the dual-mode communication competition state, and selects the communication mode with the lowest potential energy and the most stable as the convergence target. Convergence speed for contention is achieved through multiple control mechanisms, including convergence time window control, convergence acceleration factor adjustment, and convergence interrupt protection. The standard convergence time is 10 milliseconds, ensuring that contention completes its transition from contention to determinism before instructions are issued. Convergence stability is ensured through a continuous contention sampling verification mechanism, including sampling frequency control, sampling consistency verification, and statistical analysis of sampling results. Convergence is confirmed when consecutive sampling results point to the same communication mode with increasing confidence. The contention conflict resolution mechanism specifically addresses situations of intense contention. When dual-mode communication contention reaches a state of equilibrium, external decision-making factors such as environmental priorities, historical preferences, and energy consumption considerations are introduced to force a break in the contention balance, ensuring that the convergence process produces clear and actionable results. Convergence monitoring utilizes real-time trajectory tracking technology to record the complete transition of dual-mode communication contention from contention to determinism, generating comprehensive convergence result data including convergence direction, convergence speed, stability indicators, and confidence distribution. When the trigger signal activates the convergence algorithm, the dual-mode communication contention quickly converges from a state of coexistence between Starflash and UWB to a state where Starflash is the preferred choice. Convergence data shows that Starflash is the only choice with a convergence confidence level of 0.92. This instantaneous convergence mechanism rapidly resolves the uncertainty of the dual-mode communication contention and generates convergence data.
[0077] The convergence result data is then used to determine the unique optimal solution, enabling a clear choice of communication mode. The optimal solution determination process employs a strategy that combines result verification and exception handling. The convergence direction, convergence speed, stability indicators, and confidence distribution in the convergence result data are analyzed to extract the final selection result and reliability assessment for the competitive convergence. Convergence result verification is achieved through a multi-dimensional quality inspection mechanism, including convergence confidence inspection, stability indicator verification, consistency analysis, and timeliness assessment. When the convergence confidence exceeds 0.8, the stability indicators meet the requirements, and the results are consistent, the convergence result is confirmed to be the optimal solution. The convergence exception handling mechanism specifically addresses abnormal situations such as insufficient convergence confidence, oscillation, or inconsistent results. It activates an alternative decision-making mechanism that includes historical data analysis, user preference matching, and environmental adaptability reassessment for supplementary analysis. The optimal solution is determined through historical performance data statistics, user preference setting queries, and real-time environmental re-detection. The optimal solution consistency check utilizes a timing analysis method to assess the coherence and rationality of the current choice with historical decision patterns. Through technical means such as decision history review, pattern recognition analysis, and abnormal change detection, frequent communication mode switching can be prevented from impacting stability and user experience. Decision timeliness control utilizes parallel processing and priority scheduling mechanisms to ensure the optimal solution determination process is completed within 5 milliseconds, meeting the stringent response requirements of real-time control. This includes technical support such as fast path optimization, concurrent verification processing, and an interrupt response mechanism. Decision result records utilize a structured storage method to store the optimal solution selection criteria, confidence information, decision path, and timestamps, supporting subsequent performance analysis, optimization adjustments, and fault diagnosis. Convergence results data demonstrates a convergence confidence level of 0.92 for the Star Flash mode and excellent stability indicators, confirming that Star Flash communication is the only optimal solution, eliminating the need for additional decision analysis. Based on convergence results data analysis, the single optimal solution is quickly determined.
[0078] Finally, the unique optimal solution is converted to the optimal communication mode, completing the communication configuration and command transmission of the smart remote control. Communication mode conversion utilizes a layered conversion architecture, encompassing three conversion levels: hardware abstraction layer encapsulation, protocol parameter configuration, and interface adaptation control. This maps the abstract optimal solution into specific communication hardware configuration and protocol parameters. Star Flash communication mode configuration encompasses four configuration dimensions: frequency band selection, power control, encoding scheme, and transmission protocol settings. Communication parameters are optimized for optimal performance based on the current network environment and device capabilities. Frequency band selection utilizes dynamic spectrum sensing technology to avoid interfering bands, and power control employs adaptive power management to balance transmission range and energy efficiency. UWB communication mode configuration encompasses four core configuration items: pulse parameters, time domain synchronization, spatial positioning, and anti-interference settings, ensuring stable communication in complex electromagnetic environments. Pulse parameter optimization enhances signal penetration, while time domain synchronization ensures the temporal accuracy of data transmission. Mode switching control utilizes unified interface standards and switching protocols for seamless switching. This includes four switching phases: switch preparation, state save, mode activation, and state restore, minimizing data loss and latency during the switching process. Switching time is kept within 5 milliseconds to ensure a continuous user experience. Communication protocol adaptation utilizes a multi-layered protocol conversion mechanism to ensure that gesture recognition results are correctly encoded into control command formats recognizable by the target device. Supporting multiple downstream communication protocols, including infrared, Bluetooth, and WiFi, protocol conversion includes format adaptation, parameter mapping, and validation steps. Command encapsulation utilizes a standardized data structure to convert gesture recognition semantics into standardized device control commands, including complete information such as command type, parameter value, priority, and timestamp. The command format is scalable to accommodate the control needs of different devices. Transmission quality assurance is achieved through multiple mechanisms, including error detection, retransmission mechanisms, acknowledgement, and timeout handling, ensuring reliable command transmission. A maximum retransmission count of three and a timeout of 100 milliseconds are set. User feedback and alternative handling options are provided in the event of transmission failure. A status feedback mechanism utilizes a bidirectional communication architecture to monitor the success of command transmission and device response, providing users with real-time feedback on operation results, including transmission confirmation, execution status, and error notifications. In the TV volume control app, when the StarFlash communication mode is selected, a "swipe right" gesture is converted into a standard TV control command, "volume +1," and sent to the smart TV device via the StarFlash protocol, achieving accurate volume control response. Combining hardware abstraction and protocol adaptation enables intelligent remote control.
[0079] In order to implement the above method embodiment corresponding to the intelligent remote control method based on gesture recognition to achieve the corresponding functions and technical effects. Figure 2 , Figure 2The following is a block diagram of a smart remote control device 200 based on gesture recognition according to an embodiment of the present application. For ease of explanation, only the parts related to the present embodiment are shown. The smart remote control device 200 based on gesture recognition according to an embodiment of the present application includes:
[0080] The feature processing module 201 is used to obtain multimodal sensor data, perform feature extraction on the multimodal sensor data to obtain gesture spatiotemporal features, and perform spatiotemporal dimension exchange on the gesture spatiotemporal features to generate exchange dimension features;
[0081] A probability construction module 202 is configured to obtain a candidate gesture probability superposition state based on the exchange dimension feature, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis;
[0082] A state correction module 203 is configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine a target gesture;
[0083] A hierarchical reasoning module 204 is configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification, and output a gesture recognition result;
[0084] The communication control module 205 detects the current communication environment status, constructs the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment status, and at the same time establishes a parallel processing mechanism for the conflicting decisions of selecting Star Flash and selecting UWB. At the moment the command is sent, the dual-mode communication competition state is aggregated into the optimal communication mode to complete the intelligent control of the remote control.
[0085] The aforementioned gesture recognition-based intelligent remote control device 200 can implement the gesture recognition-based intelligent remote control method of the aforementioned method embodiment. The optional options in the aforementioned method embodiment also apply to this embodiment and will not be described in detail here. The remaining contents of the present application embodiment can be referenced to the contents of the aforementioned method embodiment and will not be further described in this embodiment.
[0086] like Figure 3 As shown, the third embodiment of the present invention further provides a computer device, including a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302, characterized in that when the processor 302 executes the program, the steps of the intelligent remote control control method based on gesture recognition described in the first embodiment of the present invention are implemented.
[0087] The purpose of the above embodiments is to exemplify and deduce the technical solution of the present invention, and to fully describe the technical solution, purpose and effect of the present invention. Its purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosed content of the present invention, and it does not limit the scope of protection of the present invention.
[0088] The above embodiments are not exhaustive and may include many other embodiments not listed above. Any replacements and improvements made without violating the concept of the present invention are within the scope of protection of the present invention.
Claims
1. A method for controlling an intelligent remote controller based on gesture recognition, characterized in that: include: Acquire multimodal sensor data, perform feature extraction on the multimodal sensor data to obtain gesture spatiotemporal features, and perform spatiotemporal dimension exchange on the gesture spatiotemporal features to generate exchange dimension features; Acquire a candidate gesture probability superposition state based on the exchange dimension feature, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate a reversal verification parameter based on the semantic negation hypothesis; Performing weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and performing probability aggregation on the corrected probability state to determine a target gesture; Extracting verification feature data based on the target gesture, inputting the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification, and outputting a gesture recognition result; Detect the current communication environment state, build the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment state, aggregate the dual-mode communication competition state into the optimal communication mode at the moment the command is sent, and complete the intelligent control of the remote control.
2. The method according to claim 1, characterized in that The performing spatiotemporal dimension exchange on the spatiotemporal feature of the gesture to generate an exchange dimension feature includes: Decomposing the gesture spatiotemporal features into a time dimension component and a space dimension component; Cross-mapping the time dimension component and the space dimension component; reconstructing a dimensional correlation matrix based on the cross-mapping result; The dimension correlation matrix is used to generate exchange dimension features.
3. The method according to claim 1, characterized in that The obtaining of a candidate gesture probability superposition state based on the exchange dimension feature includes: Performing probability distribution analysis on the exchange dimension feature to obtain probability component data; Constructing a multi-state energy field based on the probability component data; Analyzing the multi-state energy field to generate energy distribution parameters; A candidate gesture probability superposition state is generated using the energy distribution parameters.
4. The method according to claim 1, wherein The constructing of a semantic negation hypothesis for the candidate gesture probability superposition state includes: Performing semantic reverse parsing on the candidate gesture probability superposition state; Constructing an antonym semantic set based on the semantic reverse parsing result; A semantic negation hypothesis is generated using the antonym semantic set.
5. The method according to claim 1, wherein The generating of the inversion verification parameter based on the semantic negation hypothesis includes: converting the semantic negation assumption into a verification constraint; Analyze the verification constraint conditions to obtain reverse verification indicators; Parameter quantization processing is performed based on the reverse verification indicator to generate a reverse verification parameter.
6. The method according to claim 1, characterized in that The step of performing weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state includes: generating a weight correction coefficient using the inversion verification parameter; Performing weight adjustment on the candidate gesture probability superposition state based on the weight correction coefficient to obtain adjusted probability data; Analyzing the adjusted probability data to generate a stability assessment result; A revised probability state is generated based on the stability assessment result.
7. The method according to claim 1, characterized in that The step of inputting the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification and outputting a gesture recognition result includes: When the verification feature data meets the preset classification conditions, the first-level classification model is triggered to perform preliminary classification to obtain a preliminary classification result; When the verification feature data does not meet the preset classification conditions, the secondary classification model is triggered to perform in-depth analysis to obtain a deep classification result; Generate gesture recognition results based on the above classification results.
8. The method according to claim 1, characterized in that Aggregating the dual-mode communication contention into the optimal communication mode at the moment of sending the instruction includes: The detection instruction sends a trigger signal; Performing instantaneous convergence on the dual-mode communication contention based on the trigger signal to obtain convergence result data; Determine a unique optimal solution through the convergence result data; The unique optimal solution is converted into an optimal communication mode.
9. An intelligent remote control device based on gesture recognition, characterized in that: include: a feature processing module, configured to acquire multimodal sensor data, perform feature extraction on the multimodal sensor data to acquire spatiotemporal features of gestures, and perform spatiotemporal dimension exchange on the spatiotemporal features of gestures to generate exchange dimension features; a probability construction module, configured to obtain a candidate gesture probability superposition state based on the exchange dimension feature, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; A state correction module, configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine a target gesture; a hierarchical reasoning module, configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical reasoning model comprising a primary classification model and a secondary classification model for classification verification, and output a gesture recognition result; The communication control module is used to detect the current communication environment state, build the Star Flash and UWB dual-mode communication competition state based on the gesture recognition result and the communication environment state, aggregate the dual-mode communication competition state into the optimal communication mode at the moment the command is sent, and complete the intelligent control of the remote control.
10. A computer device, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
FMCW radar gesture recognition method based on spatial-temporal feature sequence
CN117828468A
Dynamic gesture recognition method and system, electronic equipment and computer readable storage medium
CN118196879A