Intelligent remote controller control method, device and equipment based on gesture recognition
Patent Information
- Application Number
- CN202510887253.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
Smart Images

Figure CN120386456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent interaction technology, and particularly to an intelligent remote control method, device and equipment based on gesture recognition. Background Art
[0002] In recent years, with the booming development of artificial intelligence and Internet of Things technologies, the popularity of smart home devices has been continuously increasing, and users' requirements for device interaction experiences have also been growing day by day. As a natural and intuitive human-computer interaction method, gesture recognition realizes device control by capturing and analyzing users' hand movements, and has become an important direction for the development of intelligent remote control technology. Traditional gesture recognition methods mainly rely on single-sensor technologies, such as vision recognition based on cameras or motion perception based on inertial measurement units, and have problems such as unstable recognition accuracy and weak anti-interference ability in complex environments.
[0003] Although existing multi-sensor gesture recognition technologies have improved recognition performance to a certain extent, there are still many technical bottlenecks: the multi-modal sensor data fusion strategy is simple and rough, lacking in-depth feature correlation analysis, and it is difficult to fully exploit the complementary advantages between different sensors; the gesture feature extraction method is single, ignoring the complex correlation of gesture movements in the spatio-temporal dimension, resulting in insufficient feature expression ability; the recognition decision-making mechanism lacks adaptability and cannot dynamically adjust the recognition strategy according to signal quality and environmental complexity; the communication mode selection mechanism is rigid and cannot be intelligently matched according to gesture types and communication environment states. These problems severely limit the recognition accuracy and user experience of intelligent remote controls in diverse application scenarios, and there is an urgent need to develop a new generation of gesture recognition technologies with capabilities of deep fusion, intelligent decision-making, and adaptive optimization. Summary of the Invention
[0004] The present invention provides an intelligent remote control method, device and equipment based on gesture recognition, aiming to solve key technical problems in existing gesture recognition technologies, such as insufficient depth of multi-sensor fusion, limited feature expression ability, simple recognition decision-making mechanism, and rigid communication mode selection. By integrating innovative technologies such as spatio-temporal dimension exchange, probability superposition verification, dynamic hierarchical reasoning, and dual-mode communication race, a complete technical chain from multi-modal perception to deep feature extraction, then to intelligent verification decision-making and adaptive communication transmission is constructed, achieving an improvement in gesture recognition accuracy, an enhancement in environmental adaptability, and an intelligent optimization of communication efficiency, and forming an intelligent remote control solution with features of deep perception, intelligent reasoning, autonomous decision-making, and dynamic optimization.
[0005] In the first aspect of the present invention, an intelligent remote control method based on gesture recognition is proposed, including the following steps: Obtain multi-modal sensor data, extract features from the multi-modal sensor data to obtain gesture spatio-temporal features, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features; Based on the exchanged dimension features, obtain the candidate gesture probability superposition state, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; Based on the inversion verification parameter, perform weight correction on the candidate gesture probability superposition state to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine the target gesture; Extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification, and output the gesture recognition result; Detect the current communication environment state, construct a StarFlash and UWB dual-mode communication race condition based on the gesture recognition result and the communication environment state, and aggregate the dual-mode communication race condition into an optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
[0006] A second aspect of the present invention proposes an intelligent remote control device based on gesture recognition, including: A feature processing module, configured to obtain multi-modal sensor data, extract features from the multi-modal sensor data to obtain gesture spatio-temporal features, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features; A probability construction module, configured to obtain the candidate gesture probability superposition state based on the exchanged dimension features, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; A state correction module, configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine the target gesture; A hierarchical inference module, configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification, and output the gesture recognition result; A communication control module, which detects the current communication environment state, constructs a StarFlash and UWB dual-mode communication race condition based on the gesture recognition result and the communication environment state, and simultaneously establishes a contradictory decision parallel processing mechanism for selecting StarFlash and selecting UWB, and aggregates the dual-mode communication race condition into an optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
[0007] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of an intelligent remote control method based on gesture recognition disclosed in the first aspect are implemented.
[0008] The beneficial effects of the present invention are reflected in the following aspects: First, through innovative techniques of spatio-temporal dimension exchange and probability superposition state construction, a deep feature extraction framework including dimension decomposition, cross mapping, and correlation matrix reconstruction is established. Combining semantic negation hypothesis and inversion verification mechanism, it solves the core technical problems of single feature expression and lack of verification mechanism in traditional methods, realizes the technical breakthrough from shallow spatio-temporal representation of gesture features to deep correlation representation, improves the recognition accuracy, feature discrimination ability, and verification reliability of complex gestures, enabling the system to accurately recognize subtle gesture actions and provide reliable verification guarantees.
[0009] Second, through a dual verification mechanism combining innovative probability superposition verification and classification model verification, a multi-level verification system including probability statistical verification, classification algorithm verification, and signal confidence evaluation is constructed. It breaks through the technical bottleneck of insufficient reliability of the single verification mechanism in existing methods, realizes the technical leap from traditional single recognition to dual verification recognition, improves the recognition accuracy, verification reliability, and anti-interference ability, and provides multiple guarantees and reliability verification for complex gesture recognition.
[0010] Finally, through dual-mode communication race and instantaneous convergence optimization technology, the intelligent selection, dynamic weight allocation, and real-time switching capabilities of StarFlash and UWB communication modes are realized, solving the key problems of single mode selection, poor environmental adaptability, and low switching efficiency in traditional communication methods, forming a new generation of communication control scheme with environmental perception and intelligent decision-making characteristics, and improving the communication reliability, transmission efficiency, and environmental adaptability of the intelligent remote control in diverse environments.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings here show specific examples of the technical solutions of the present invention and form part of the specification together with the specific implementation manners, and are used to explain the technical solutions, principles, and effects of the present invention.
[0013] Unless otherwise specified or defined, in different drawings, the same reference numerals represent the same or similar technical features. For the same or similar technical features, different reference numerals may also be used for representation.
[0014] Figure 1It is a schematic flowchart of a method for controlling an intelligent remote control based on gesture recognition according to the present invention.
[0015] Figure 2 It is a block diagram of the structure of a device for controlling an intelligent remote control based on gesture recognition according to the present invention.
[0016] Figure 3 It is a schematic diagram of the structure of a computer device according to the present invention. Detailed implementation manners
[0017] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0018] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0019] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0020] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0021] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0022] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0023] The technical solutions of the embodiments of this application will be introduced below.
[0024] As Figure 1 shown, an embodiment of the present invention provides an intelligent remote control method based on gesture recognition, including the following steps S110 - step S150: Step S110, obtain multi-modal sensor data, extract features from the multi-modal sensor data to obtain gesture spatio-temporal features, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features.
[0025] Specifically, the multi-modal sensor data acquisition relies on a high-precision multi-sensor fusion module. This module integrates a ToF camera, a millimeter-wave radar, and an IMU inertial measurement unit inside the smart remote control. The sensors are connected through a high-speed data bus to achieve synchronous data transmission and interaction. The ToF camera uses a 940nm near-infrared light source, with a depth perception range of 0.2 - 3.0 meters, a frame rate of 30fps, and a resolution of 64×64 pixels, capturing the three-dimensional contour and depth information of the hand. The millimeter-wave radar operates at a frequency of 60GHz, with a detection range of 1.5 meters, an angular resolution of 5 degrees, and a velocity resolution of 0.1m / s, specifically for detecting micro hand gesture signals, including subtle movements such as finger sliding, clicking, and hovering. The IMU inertial measurement unit integrates a three-axis accelerometer and a three-axis gyroscope, with an acceleration measurement range of ±16g, an angular velocity measurement range of ±2000° / s, and a sampling frequency of 1kHz, tracking the spatial attitude changes and movement trajectories of the remote control in real time. The hardware layout of the multi-sensor fusion module is carefully optimized. The ToF camera and the millimeter-wave radar are integrated in the top area of the smart remote control, covering a 120° fan-shaped detection area in the front to ensure comprehensive and accurate capture of front gestures. The UWB ultra-wideband module operates in the 6.5GHz band, and the SparkLink communication module operates in the 2.4GHz band. The two achieve time-division multiplexing of the antenna through a radio frequency switch, and are set at the bottom or side position of the smart remote control to ensure stable transmission and reception of communication signals. The multi-modal sensor data preprocessing uses advanced technologies such as spatio-temporal alignment, noise filtering, and data synchronization. A unified timestamp is added to the ToF depth map, millimeter-wave radar point cloud, and IMU six-axis data. The clocks of the sensors are synchronized through the high-speed data bus using the PTP precise time protocol, with an accuracy reaching the microsecond level. Through multi-modal data fusion processing, a synchronous multi-modal sensor data stream containing complete gesture motion information is finally generated.
[0026] Feature extraction is performed on multi-modal sensor data to obtain gesture spatio-temporal features. The feature extraction process adopts a sub-modal processing strategy, and specialized feature extraction algorithms are designed according to the data characteristics of different sensors. For the feature extraction of ToF camera data, a 3D key-point detection algorithm is used. The three-dimensional coordinates of 21 key points such as the palm center, fingertips, and joints are identified through a deep learning model. Each key point contains three coordinate components of X, Y, and Z, forming a 63-dimensional hand geometry feature vector. In the key-point detection process, the depth map of 64×64 pixels is first preprocessed, including noise filtering, edge enhancement, and depth correction, and then the spatial geometry features are extracted through a 3D convolutional neural network. The network structure includes multiple 3D convolutional layers, and hierarchical representations from low-level edge features to high-level semantic features are extracted layer by layer. The feature extraction of millimeter-wave radar data focuses on the analysis of micro-motion frequency and motion pattern. The time-domain radar signal is converted into a time-frequency spectrogram through the short-time Fourier transform, and frequency-domain features such as peak frequency, frequency bandwidth, spectral centroid, and spectral roll-off are extracted from it. These features can effectively distinguish different types of gesture micro-motion patterns, such as fine movements like fast sliding, slow clicking, and hovering. The feature extraction of the IMU inertial measurement unit data combines time-domain and frequency-domain analysis methods. The statistical features of acceleration and angular velocity within a sliding window are calculated in the time domain, including mean, variance, peak value, zero-crossing rate, etc. In the frequency domain, the power spectral density is obtained through the FFT transform, and the energy ratio and distribution characteristics of the main frequency components are extracted. Through the multi-modal feature fusion technology, the ToF key-point coordinates, millimeter-wave radar frequency-domain features, and IMU motion statistical features are encoded into a unified multi-channel tensor representation, forming gesture spatio-temporal features containing complete spatio-temporal information.
[0027] In some embodiments, the spatio-temporal dimension exchange of the gesture spatio-temporal features to generate exchange dimension features includes: decomposing the gesture spatio-temporal features into a time dimension component and a space dimension component; performing cross mapping on the time dimension component and the space dimension component; reconstructing a dimension correlation matrix based on the cross mapping result; and generating exchange dimension features using the dimension correlation matrix.
[0028] First, decompose the previously obtained spatio-temporal features of the gesture into time-dimensional components and space-dimensional components. For the extraction of time-dimensional components, a one-dimensional convolutional kernel slides along the time axis with a kernel size of 5 and a stride of 1, specifically extracting local patterns and changing trends of the time series. By global average pooling, the spatial features at each time step are compressed into scalar values, forming a pure time series that reflects the evolutionary characteristics of the gesture action in the time dimension, including time-phase features such as the start, development, climax, and end of the action. For the extraction of space-dimensional components, two-dimensional convolution is used to independently process the spatial features at each time step. The kernel size is 3×3, specifically extracting spatial geometric structures and shape patterns. By temporal average pooling, the time-varying factors are eliminated, and pure spatial distribution information is retained. This component reflects the spatial geometric features and shape attributes of the gesture, such as the hand contour, key-point distribution, and geometric shape of the motion trajectory. In a complex gesture recognition task, when the user performs the "draw a circle in the air" action, the time-dimensional component can capture the periodic speed change pattern, with peaks appearing at the turning points of the circle drawing; the space-dimensional component can extract the geometric features of the circular trajectory, including key parameters such as the center position, radius size, and shape regularity. Based on the dimension decomposition processing, independent time-dimensional component and space-dimensional component feature representations are generated.
[0029] Secondly, perform cross-mapping on the time-dimensional component and the space-dimensional component. The cross-mapping uses an attention mechanism to calculate the correlation weights between the time component and the space component, and generates a spatio-temporal correlation matrix through matrix multiplication. Each element of the correlation matrix represents the correlation strength between a specific time step and a specific spatial position. The larger the value, the stronger the correlation, reflecting the importance of this spatio-temporal position combination for gesture recognition. To enhance the expressive ability of the correlation, a learnable mapping function is introduced to perform non-linear transformations on the time and space components respectively. The mapping function adopts a multi-layer perceptron structure with two hidden layers, and the activation function is ReLU, which can learn complex non-linear spatio-temporal correlation patterns. The cross-mapping process also considers spatio-temporal correlation patterns at different scales. Through a multi-scale cross-attention mechanism, correlation weights are calculated in different time windows and spatial neighborhoods to capture spatio-temporal interaction patterns from fine-grained to coarse-grained. The multi-scale correlation weights are fused through weighting to generate the final cross-mapping result, and the weight coefficients are learned and optimized through backpropagation. In the "click to confirm" gesture, the cross-mapping can find that the time point when the finger presses down has a strong correlation with the spatial position of the contact area, and the correlation weight reaches a peak at the moment of contact and significantly decreases during non-contact times. Using the cross-mapping processing, the correlation relationship representation between the time-dimensional component and the space-dimensional component is established.
[0030] Then, based on the cross - mapping results, the dimension correlation matrix is reconstructed. The dimension correlation matrix contains four sub - matrix blocks: the time auto - correlation matrix, the space auto - correlation matrix, the time - to - space correlation matrix, and the space - to - time correlation matrix. The time auto - correlation matrix is obtained by calculating the similarity between features at different time steps, using cosine similarity metric, which reflects the internal structure and periodic characteristics of the time series. The space auto - correlation matrix calculates the correlation of features at different spatial positions, considering the combined influence of spatial proximity and feature similarity, and reflects the geometric structure and continuity characteristics of the spatial distribution. The time - to - space correlation matrix and the space - to - time correlation matrix are generated from the aforementioned cross - mapping results and are normalized to ensure numerical stability and comparability. The reconstruction of the dimension correlation matrix adopts the method of block matrix splicing to form a complete spatio - temporal dimension correlation description. To improve the expressive ability of the matrix, low - rank decomposition technology is introduced to reduce redundant information. The correlation matrix is decomposed by singular value decomposition, and the main singular value components are retained to achieve dimension compression and noise filtering. In the analysis of the "swipe - to - flip" gesture sequence, the dimension correlation matrix shows an obvious diagonal block structure. The time auto - correlation matrix has high correlation at the beginning and end of the swipe, reflecting the consistency of the start and end of the action; the space auto - correlation matrix shows strong correlation at adjacent positions on the swipe path, reflecting the continuity of the trajectory. Under the action of the dimension correlation matrix reconstruction, a structured matrix representation containing complete spatio - temporal interaction information is generated.
[0031] Finally, the exchanged - dimension features are generated using the dimension correlation matrix. Through matrix multiplication operation, the dimension correlation matrix is multiplied by the flattened representation of the original spatio - temporal features to obtain the preliminary exchanged features. To maintain the physical meaning and geometric structure of the features, a dimension permutation operation is introduced to exchange the positions of the time dimension and the space dimension, so that the original time features are expanded in the space dimension and the space features are extended in the time dimension. The permutation operation is achieved through tensor reshaping and dimension transposition, which changes the dimension organization mode of the feature tensor and provides different perspectives for the model to observe features. The exchanged - dimension features are also adaptively modulated through a gating mechanism. The gating unit learns the importance weights of each dimension feature, and the gating weights are generated through the sigmoid activation function to ensure that the weight values are between 0 and 1. The final exchanged - dimension features are obtained by weighted fusion of the original features and the exchanged features, realizing adaptive feature selection and fusion. In the "rotating gesture" recognition task, the exchanged - dimension features map the angular velocity change pattern in the original time series to the geometric description of the spatial rotation trajectory, enabling the model to understand the rotation law in time from the spatial geometric perspective. Combining the dimension exchange processing technology, exchanged - dimension features with enhanced expressive ability are obtained.
[0032] Step S120: Based on the exchanged - dimension features, obtain the candidate gesture probability superposition state, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis.
[0033] Specifically, convert the deterministic gesture features into a probability superposition state, and at the same time construct a semantic negation hypothesis for reverse verification. Through the dual mechanisms of probability superposition and semantic negation, improve the robustness and accuracy of gesture recognition.
[0034] In some embodiments, obtaining the candidate gesture probability superposition state based on the exchanged dimension features includes: performing probability distribution analysis on the exchanged dimension features to obtain probability component data; constructing a multi-situation energy field based on the probability component data; analyzing the multi-situation energy field to generate energy distribution parameters; generating the candidate gesture probability superposition state through the energy distribution parameters.
[0035] First, perform probability distribution analysis on the exchanged dimension features to obtain probability component data. Using the Bayesian statistical framework, regard each feature dimension in the exchanged dimension features as a random variable, and fit its probability distribution characteristics through the maximum likelihood estimation method. For the hand key point features extracted by the ToF camera, use a multivariate Gaussian distribution to model the uncertainty of the spatial coordinates, and the covariance matrix reflects the accuracy and mutual correlation of key point detection. For the micro-motion frequency features of the millimeter-wave radar, use a mixture Gaussian model to describe the frequency distribution characteristics of different gesture actions, and each Gaussian component corresponds to a typical gesture pattern. For the motion trajectory features of the IMU, use a Markov model to describe the temporal transition probability of the motion state, and the state transition matrix describes the conversion relationship between different motion stages. The probability distribution analysis process also considers the influence of sensor noise and environmental interference, and corrects the theoretical distribution parameters through the noise model to ensure the accuracy and robustness of probability modeling. When the IMU detects that the acceleration is greater than 1.5g or the millimeter-wave radar detects that the hand micro-motion speed is greater than 0.5m / s, dynamically adjust the variance parameter of the probability distribution to adapt to the uncertainty characteristics of high-dynamic gestures. Use the probability distribution analysis technology to generate probability component data containing statistical characteristics.
[0036] Next, a multi-situation energy field is constructed based on the probability component data, converting the probability information into a representation form in the energy space. The multi-situation energy field is constructed by mapping the probability distribution corresponding to each candidate gesture into the potential energy distribution in the energy space. Using the kernel density estimation method, with the probability component data as the core, a continuous energy distribution is generated in the feature space through the Gaussian kernel function. The bandwidth parameter of the kernel function is adaptively adjusted according to the variance of the feature dimension to ensure the smoothness and local preservation of the energy field. For the feature region with a high probability density, the corresponding potential energy value is low, indicating a strong stability of the gesture state; for the region with a low probability density, the potential energy value is high, indicating an unstable or transitional state. The multi-situation energy field also introduces a multi-scale energy hierarchy, constructing the corresponding potential energy distribution at different feature scales. The coarse-scale potential energy field captures the overall pattern of the gesture, and the fine-scale potential energy field describes the local detail changes. The multi-state characteristic of the energy field is reflected in that multiple energy states may correspond to the same position in the feature space, reflecting the ambiguity and polysemy in gesture recognition. In the recognition of the "wave goodbye" gesture, the multi-situation energy field can simultaneously describe two different energy states of a fast wave and a slow wave, and each energy state corresponds to different probability weights and energy levels. After the construction process of the potential energy field, an energy distribution representation with multi-state characteristics is formed.
[0037] Then, a deep analysis is performed on the multi-situation energy field to generate energy distribution parameters. The extraction of energy distribution parameters adopts the partition function theory in statistical physics, quantifying the global and local characteristics of the energy distribution by calculating the statistical features of the potential energy field. The global energy parameters include total energy, average energy, energy variance, and energy entropy, reflecting the overall distribution characteristics and stability of the potential energy field. The local energy parameters are obtained through the gradient analysis of the potential energy field, including the magnitude, direction, and change rate of the energy gradient, describing the local change characteristics and the distribution of stable points of the energy field. The stable point identification adopts the potential energy surface analysis method, determining the minimum energy point, maximum energy point, and saddle point by calculating the critical points of the potential energy field. These key points correspond to different gesture states and transition paths. The energy barrier calculation evaluates the conversion difficulty between different gesture states. A high energy barrier indicates a high discrimination between gestures, and a low energy barrier indicates an area where misrecognition is likely to occur. The dynamic energy analysis considers the time evolution characteristics, tracking the development trajectory of the gesture movement through the energy changes in the time series. In the analysis of the "click to confirm" gesture, the energy distribution parameters show that there is an obvious minimum energy at the moment of finger contact, corresponding to the stable click state, while the energy value is high in the transitional stage before and after contact, reflecting the instability of the action. Through the energy field analysis algorithm, the energy distribution parameters describing the characteristics of the potential energy distribution are obtained.
[0038] Finally, a candidate gesture probability superposition state is generated through the energy distribution parameter. The generation of the probability superposition state adopts the quantum probability theory framework, extends the classical probability to the complex probability amplitude, and allows quantum superposition and interference effects between different gesture states. Each candidate gesture corresponds to a probability amplitude. The square of the modulus length of the probability amplitude represents the classical probability of the gesture, and the phase of the probability amplitude reflects the coherence relationship between gestures. The construction of the superposition state is achieved through linear combination, and the superposition coefficient is determined by the aforementioned energy distribution parameter. The low-energy state obtains a larger superposition weight, and the high-energy state has a smaller weight. The setting of the phase relationship is based on the semantic similarity and action continuity between gestures. Similar gestures have similar phases, and gestures with large differences have phase differences close to orthogonality. The probability superposition state also considers the time evolution characteristics, describes the time evolution of the superposition state through the discretized form of the Schrödinger equation, and the evolution operator is determined by the Hamiltonian composed of the energy distribution parameter. The collapse mechanism of the superposition state simulates the quantum measurement process. When external observation intervenes, the superposition state collapses into a definite gesture recognition result. In the recognition of the complex gesture sequence "slide + click", the probability superposition state can simultaneously maintain the probability amplitudes of both the slide and click gestures, and provide a more accurate recognition judgment through the coherent superposition of the probability amplitudes at the critical moment of action conversion. With the help of the quantum probability superposition mechanism, a representation of the candidate gesture probability superposition state is constructed.
[0039] In some embodiments, constructing a semantic negation hypothesis for the candidate gesture probability superposition state includes: performing semantic reverse parsing on the candidate gesture probability superposition state; constructing an antonym semantic set based on the semantic reverse parsing result; and generating a semantic negation hypothesis using the antonym semantic set.
[0040] First, the semantic reverse parsing is performed on the candidate gesture probability superposition state to extract the reverse semantic features of the gesture. The semantic reverse parsing is based on the contrastive learning theory, and enhances the discrimination ability of recognition by analyzing the "non-features" of each candidate gesture. The reverse parsing process constructs a negative sample space, and systematically generates its corresponding negative representation for each positive gesture category. Taking the "swipe right" gesture as an example, its reverse semantics include all action patterns that are not "swipe right", such as "swipe left", "swipe up", "swipe down", and "stay still". The reverse parsing adopts the methods of feature inversion and semantic flipping, and performs mathematical transformation on the feature vector of the positive gesture to generate the corresponding reverse feature representation. For the hand trajectory features captured by the ToF camera, the reverse parsing generates reverse trajectory patterns through operations such as trajectory inversion, direction flipping, and speed inversion. For the motion acceleration features detected by the IMU, the reverse parsing constructs reverse motion patterns through methods such as sign flipping, amplitude inversion, and frequency inversion. The semantic reverse parsing also considers the reverse features in the time dimension, and extracts the time reverse semantics through methods such as reverse playback of the time series, time scale transformation, and causality flipping. In the multi-modal feature space, the reverse parsing realizes the comprehensive flipping of semantics through geometric operations such as symmetric transformation, complementary projection, and orthogonal decomposition of the feature space. Based on the reverse parsing algorithm, complete semantic reverse parsing result data is generated.
[0041] Next, an antonym semantic set is constructed based on the semantic reverse parsing results. Using a hierarchical organizational structure, the reverse semantics are classified and organized according to multiple dimensions such as semantic levels, action types, and feature dimensions. The top-level semantic antonyms include the opposition relationships of basic action types, such as "static vs. motion", "approach vs. away", "press vs. lift", etc. The middle-level semantic antonyms involve the opposition relationships of specific gesture actions, such as "swing left vs. swing right", "slide up vs. slide down", "clockwise vs. counterclockwise", etc. The bottom-level semantic antonyms are refined to the opposition features of action details, such as "fast vs. slow", "large amplitude vs. small amplitude", "continuous vs. discontinuous", etc. The antonym semantic set also includes temporal antonym relationships, which describe the temporal opposition patterns of action sequences, such as "fast first then slow vs. slow first then fast", "gradually strengthen vs. gradually weaken", "periodic vs. non-periodic", etc. The construction of the semantic set introduces the ontology knowledge representation method, and describes the logical structure and constraint conditions of the antonym relationship through a semantic network. Each antonym semantic node contains attribute information such as semantic identification, feature description, opposition intensity, and applicable conditions. The intensity of the antonym relationship is quantified through similarity measurement and contrast calculation. Strong opposition relationships have high contrast values, while weak opposition relationships have lower contrast values. In the antonym construction of the "clench fist + release" composite gesture, the antonym semantic set can identify various opposition patterns such as "always clench fist", "always release", "release first then clench fist", etc., providing a rich semantic basis for subsequent negative hypothesis generation. Through semantic organization and relationship modeling, a complete knowledge system of the antonym semantic set is established.
[0042] Finally, semantic negation hypotheses are generated using the antonym semantic set. By adopting the methods of logical reasoning and hypothesis testing, negative propositions are constructed based on the opposition relationships in the antonym semantic set. For each candidate gesture, the system generates a series of negation hypotheses that describe the criteria for judging what the gesture "is not". The construction of the negation hypotheses is represented using predicate logic in the form of logical expressions like "NOT(Gesture X has Feature Y)". The hypothesis generation process considers multiple levels of negation relationships, including different types such as category negation, feature negation, temporal negation, and combined negation. Category negation hypotheses target the basic types of gestures, such as high-level semantic negations like "the current gesture is not a click action" and "the current gesture is not a swipe action". Feature negation hypotheses target specific feature attributes, such as detailed feature negations like "the hand movement speed does not exceed the threshold" and "the movement trajectory is not circular". Temporal negation hypotheses focus on the temporal characteristics of the action, such as temporal constraint negations like "the action duration is not less than the minimum value" and "the action frequency is not within a specific range". Combined negation hypotheses involve the joint negation of multiple features, constructing compound negation conditions through logical connectives. The generation of the negation hypotheses also introduces uncertainty quantification, with each negation hypothesis associated with a confidence score that reflects the reliability of the negative judgment. The confidence calculation considers factors such as the strength of the feature evidence, the success rate of historical verification, and the clarity of semantic opposition. In the generation of negation hypotheses for the "rotation gesture", the system constructs multiple negation hypotheses such as "not a linear motion", "not a static state", and "not a unidirectional motion", and each hypothesis has a corresponding confidence score, providing a basis for judgment in the subsequent verification process. Through logical reasoning and confidence calculation, the systematic generation of semantic negation hypotheses is completed.
[0043] In some embodiments, generating reverse verification parameters based on the semantic negation hypotheses includes: converting the semantic negation hypotheses into verification constraint conditions; analyzing the verification constraint conditions to obtain reverse verification indicators; and performing parameter quantification processing based on the reverse verification indicators to generate reverse verification parameters.
[0044] First, convert the semantic negation hypothesis into verification constraint conditions and establish an operable verification framework. The constraint condition conversion adopts the mapping method from symbolic logic to numerical logic. Through the predicate logic parser and the mathematical expression generator, the logical-form negation hypothesis is converted into numerical constraint inequalities or equations. The conversion process includes four processing stages: syntax parsing, semantic analysis, mathematical modeling, and constraint generation. Syntax parsing identifies the logical structure and keywords in the negation hypothesis, and semantic analysis extracts the meaning and constraint scope of the negation relationship. For the category negation hypothesis, the converter converts the semantic negation into a numerical constraint condition of the classification probability through the probability threshold mapping mechanism. For example, "the gesture is not a click" is converted into a numerical constraint of "the click category probability is less than the threshold" through the probability inversion algorithm, and the threshold is dynamically determined through historical statistical data and error rate analysis. For the feature negation hypothesis, the converter adopts the method of defining the feature space boundary to convert the negation relationship of the feature attributes into the range constraint of the feature values. For example, "the movement speed does not exceed the threshold" is converted into an interval constraint of "the speed feature value is less than or equal to the upper bound" through the speed feature analysis, and the upper bound value is determined through kinematic analysis and sensor accuracy evaluation. For the temporal negation hypothesis, the converter uses the time constraint modeling technology to convert the time-related negation conditions into the constraint conditions of time parameters. For example, "the action duration is not less than the minimum value" is converted into a time constraint of "the time length is greater than or equal to the lower bound" through the temporal analysis, and the time boundary is obtained through the action pattern analysis and user behavior statistics. The setting of the constraint conditions deeply considers the fusion verification mechanism of multi-sensor signals, and establishes a dynamic mapping relationship between the sensor state and the constraint strength. When the IMU detects that the acceleration is greater than 1.5g, through the motion intensity evaluation algorithm, the corresponding constraint conditions require an increase in the verification strength of the "non-stationary hypothesis", and the verification threshold is automatically tightened. When the ToF camera detects that the hand distance is less than 1.5m, through the distance perception mechanism, the relevant constraint conditions of "entering the interaction area" are activated, and the close-range precise verification mode is enabled. The constraint conditions also include a complex logical combination relationship construction mechanism. Through Boolean algebra operations and logical expression optimization, logical connectives such as AND, OR, and NOT are used to construct multi-level composite constraints, reflecting the collaborative verification relationship and mutual dependence among multiple negation hypotheses. The dynamic adjustment mechanism of the constraint conditions adopts the adaptive parameter control technology. According to the real-time sensor signal quality evaluation, environmental noise level detection, and recognition accuracy monitoring, the constraint parameters and verification thresholds are adaptively modified to ensure the accuracy and real-time requirements of the verification. After the complete mathematical conversion process, a complete and directly computable verification constraint condition system is formed.
[0045] Next, conduct an in-depth analysis of the verification constraints to obtain reverse verification metrics. The extraction of reverse verification metrics adopts a multi-dimensional constraint satisfaction problem-solving method. Through constraint propagation algorithms, linear programming solvers, and heuristic search techniques, analyze key characteristics such as the satisfiability, consistency, and redundancy of the constraints, and extract quantitative metrics related to verification. The satisfiability metric is evaluated through a constraint-solving engine, using an algorithm that combines backtracking search and constraint propagation to evaluate whether there is a feasible solution in the set of constraints, calculating the feasible region of the constraint space through the linear programming simplex method or interior point method, and quantifying the satisfiability degree of the constraint system and the size of the solution space. The consistency metric is detected using conflict analysis and logical reasoning techniques. Through constraint conflict detection algorithms and logical consistency verifiers, identify whether there are logical conflicts and contradictions between the constraints. Construct a constraint dependency graph through graph theory methods to identify conflicting constraint pairs and conflict propagation paths, and calculate the overall consistency score and local conflict intensity of the constraints. The redundancy metric analysis adopts constraint independence testing and sensitivity analysis methods. Through constraint deletion tests and influence degree calculations, analyze the independence and necessity of the constraints. Through gradient analysis and partial derivative calculations, determine the contribution degree and influence weight of each constraint to the overall verification effect. The verification strength metric is quantified using constraint tightness analysis and coverage evaluation techniques. By calculating the geometric distance of the constraint boundaries and the volume ratio of the constraint space, evaluate the verification ability and constraint strength of the constraints. Quantify the strictness and comprehensiveness of the verification through the tightness of the constraint boundaries and the coverage of the feasible region. The verification robustness metric is evaluated using perturbation analysis and stability testing methods. Through noise injection and parameter perturbation experiments, evaluate the sensitivity of the constraints to sensor noise and environmental perturbations. Through Monte Carlo simulation and stability analysis, calculate the stability coefficient and anti-interference ability score of the constraint verification. The timeliness metric is evaluated using computational complexity analysis and performance testing techniques. Through theoretical analysis of algorithm complexity and actual running time measurement, evaluate the computational complexity and response time of the constraint verification, ensuring the feasibility and response speed requirements of real-time verification. In the verification analysis of the complex gesture "drawing a circle in the air + clicking", the reverse verification metrics show that the periodic constraint of the circle-drawing action and the instantaneous constraint of the clicking action have a high consistency score, indicating that the two constraints logically support each other. However, there is a certain logical tension between the continuity constraint of the circle-drawing and the discreteness constraint of the clicking, which needs to be resolved through time-segmented verification and action stage division to address the constraint conflicts. Use the multi-dimensional constraint analysis algorithm to obtain comprehensive and quantitative reverse verification metric data.
[0046] Finally, parameter quantization processing is performed based on the reverse verification metrics to generate reverse verification parameters. The parameter quantization processing adopts the multi-criteria decision analysis method. Through index weight calculation and numerical fusion technology, the multi-dimensional reverse verification metrics are converted into standardized numerical parameters that can be directly used for verification calculations. The quantization process first performs index normalization processing. Using normalization techniques such as min-max normalization and standard scaling, the verification metrics with different dimensions, numerical ranges, and distribution characteristics are uniformly mapped into the standard interval to ensure comparability, numerical stability, and calculation accuracy among different metrics. The satisfiability metric is normalized to the interval [0,1] through a linear mapping function, the consistency metric eliminates numerical bias through logarithmic transformation processing, and the redundancy metric realizes interval standardization through a piecewise linear function. The weight allocation uses the analytic hierarchy process to determine the importance weights and priority rankings of different verification metrics. Through the construction of the index importance hierarchy and judgment matrix calculation, the satisfiability metric obtains a weight coefficient of 0.35 due to its direct impact on the verification result, the consistency metric obtains a weight coefficient of 0.30 due to its logical integrity requirements, the verification strength metric obtains a weight coefficient of 0.20 due to its constraint ability evaluation, the redundancy metric obtains a weight coefficient of 0.10 due to its auxiliary role, and the verification robustness metric obtains a weight coefficient of 0.05. Parameter fusion is achieved through weighted linear combination and non-linear fusion functions. Multiple verification metrics are synthesized into a comprehensive verification parameter according to their weights and correlations. The fusion process uses a method that combines weighted arithmetic mean and geometric mean, considering the correlation and complementarity among metrics to avoid information redundancy and bias amplification. The setting of the verification threshold adopts the historical data statistical analysis method. Through the success rate statistics and error rate analysis of the verification results, the optimal judgment threshold and decision boundary are determined to balance the sensitivity and specificity of verification. The standard verification threshold is set to 0.75, the strict verification threshold is set to 0.85, and the loose verification threshold is set to 0.65. The adaptive adjustment mechanism dynamically corrects the parameter values and threshold settings according to real-time verification effect monitoring and accuracy rate statistics. When the verification accuracy rate drops by more than 5%, the verification threshold is appropriately tightened by 0.05. When the false alarm rate rises by more than 10%, the verification conditions are appropriately relaxed by reducing the threshold by 0.03. The adjustment amplitude is determined through threshold sensitivity analysis. The time decay factor of the parameter adopts an exponential decay function, considering the timeliness and dynamic changes of verification information. The time decay coefficient is set to 0.95. Newer verification evidence obtains a higher influence weight through the time weight function, and the influence of historical verification information decreases exponentially according to the time interval to ensure the timeliness and adaptability of the parameters. In multi-modal sensor fusion verification, through sensor reliability evaluation and verification contribution analysis, the verification parameter weight of the ToF camera is set to 0.4, the verification parameter weight of the millimeter-wave radar is set to 0.35, and the verification parameter weight of the IMU is set to 0.25. The weight allocation fully reflects the reliability differences, accuracy levels, and applicable scenarios of different sensors in gesture verification.Using quantization processing technology and parameter optimization methods, the precise generation of reverse verification parameters is completed.
[0047] Step S130: Based on the reverse verification parameters, perform weight correction on the candidate gesture probability superposition state to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine the target gesture.
[0048] Specifically, use the verification strength of the reverse verification parameters to dynamically adjust the probability weights of candidate gestures, effectively suppressing the unreliable probability components of verification failures, and at the same time enhancing the highly credible probability components of verification passes. The probability state after weight correction gradually converges through a multi-stage aggregation algorithm, and finally determines the unique target gesture recognition result from multiple competing candidates.
[0049] In some embodiments, the performing weight correction on the candidate gesture probability superposition state based on the reverse verification parameters to generate a corrected probability state includes: generating a weight correction coefficient using the reverse verification parameters; performing weight adjustment on the candidate gesture probability superposition state based on the weight correction coefficient to obtain adjusted probability data; analyzing the adjusted probability data to generate a stability evaluation result; and generating a corrected probability state based on the stability evaluation result.
[0050] First, generate a weight correction coefficient using the reverse verification parameters. An adaptive mapping function is used to map the numerical range of the reverse verification parameters to the coefficient interval of weight correction. For gesture candidates with a higher verification strength, the corresponding correction coefficient is close to 1.0, maintaining their original probability weights; for candidates with a lower verification strength, the correction coefficient is significantly less than 1.0, greatly reducing their probability weights. The calculation of the correction coefficient takes into account the fusion of multi-sensor verification results. When the IMU detects an acceleration greater than 1.5g and the millimeter-wave radar detects a hand micro-motion speed greater than 0.5m / s, the correction coefficient of dynamic gestures obtains an additional enhancement factor. When the ToF camera detects that the distance of the hand entering the interaction area is less than 1.5m, the correction coefficients of static gestures and fine operation gestures are correspondingly increased. The correction coefficient also introduces a time decay mechanism, where recent verification evidence obtains a higher correction strength, and the correction impact of historical verification information decays exponentially with time. The non-linear correction function uses the sigmoid activation form to ensure that the correction coefficient changes within a reasonable range and avoid system instability caused by extreme values. The multi-modal correction coefficient is generated through weighted fusion to obtain the final comprehensive correction coefficient. The weight of the ToF camera verification result is 0.4, the weight of the millimeter-wave radar verification result is 0.35, and the weight of the IMU verification result is 0.25. In the "rotation gesture" verification, when the verification parameters of all sensors support the rotation action hypothesis, the comprehensive correction coefficient reaches 0.92, significantly enhancing the probability weight of the rotation gesture. Through the verification strength mapping algorithm, a complete set of weight correction coefficients is finally generated.
[0051] Next, based on the weight correction coefficient, the weight of the candidate gesture probability superposition state is adjusted to obtain the adjusted probability data. The adjustment process uses element-wise multiplication operations, multiplying the probability amplitude of each gesture candidate by the corresponding correction coefficient to achieve precise adjustment of the probability weight. The quantum coherence property of the probability superposition state is maintained during the adjustment process, only the modulus length of the probability amplitude is modified, and the phase relationship remains unchanged, ensuring that the coherent interference effect between gestures is retained. For complex probability amplitudes, the weight adjustment acts on both the real and imaginary parts simultaneously, maintaining a constant phase angle of the complex number by scaling in the same proportion. Normalization ensures that the adjusted probability superposition state satisfies the probability conservation condition, and the sum of the squares of the moduli of all probability amplitudes is equal to 1. The dynamic adjustment strategy adaptively corrects the adjustment intensity according to the real-time sensor signal quality. When the sensor signal confidence is higher than 85%, the standard adjustment mode is adopted. When the confidence is lower than 85%, the conservative adjustment mode is enabled to reduce the aggressiveness of the correction. The multi-scale adjustment considers the verification results of different time windows. The verification results of the short-time window are used for rapid response, and the verification results of the long-time window are used for stability assurance. The adjustment process also introduces a robustness protection mechanism. When the correction coefficient of a gesture candidate is too small, the minimum probability weight is retained to prevent the complete loss of valid information. In the weight adjustment of the complex gesture sequence "slide + click", the probability weight in the sliding stage obtains a correction coefficient of 0.85 according to the motion continuity verification result, and the click stage obtains a correction coefficient of 0.78 according to the pressure sensing verification result. The adjusted probability data more accurately reflects the true execution state of the action. After precise weight adjustment processing, the adjusted probability data with verification and correction is generated.
[0052] Then, conduct in-depth analysis on the adjusted probability data to generate the stability assessment results. Adopt a multi-dimensional analysis framework to comprehensively evaluate the stability of the adjustment effect from aspects such as the uniformity, concentration, variability, and temporal consistency of the probability distribution. The distribution uniformity analysis evaluates the distribution of probability weights among different gesture candidates by calculating the entropy value and Gini coefficient of the probability distribution. A high entropy value indicates a relatively uniform probability distribution and greater recognition uncertainty; a low entropy value indicates that the probability is concentrated in a few candidates and the recognition result is relatively certain. The concentration analysis evaluates the prominence of the dominant gesture by calculating the ratio of the maximum probability value to the second-largest probability value. A high ratio indicates a clear dominant gesture and better recognition stability; a low ratio indicates intense competition among multiple candidates and further verification is required. The variability analysis evaluates the dispersion degree of the probability distribution by calculating the probability variance and standard deviation. Low variability indicates a stable probability distribution, while high variability indicates significant fluctuations. The temporal consistency analysis evaluates the temporal stability of the recognition results by calculating the correlation of the probability distribution within consecutive time windows. High correlation indicates that the recognition results are consistent over time, while low correlation indicates temporal fluctuations in the recognition results. The confidence interval calculation provides a quantification of the uncertainty of the probability estimate. Calculate the confidence interval of the probability estimate through the bootstrap method or Bayesian inference to provide a reliability reference for decision-making. The stability threshold judgment determines the judgment criteria for stability assessment based on historical statistical data. When multiple stability indicators exceed the threshold, it is determined to be in a stable state; otherwise, it is marked as an unstable state and requires additional verification. In the stability assessment of the "fist clenching and releasing" gesture, the entropy value of the probability distribution in the fist clenching stage is 0.3, indicating a relatively high recognition certainty; the entropy value of the probability distribution in the releasing stage rises to 0.7, reflecting an increase in uncertainty during the action transition period. Based on this, adjust the subsequent verification strategy. Establish a comprehensive stability assessment result based on multi-dimensional stability analysis.
[0053] Finally, a corrected probability state is generated based on the stability assessment results. Corresponding processing strategies are adopted according to different results of the stability assessment. When the stability assessment shows high stability, the adjusted probability data is retained as the corrected probability state to ensure the rapid output of stable recognition results. When the stability assessment shows medium stability, a probability smoothing algorithm is enabled to perform temporal smoothing on the adjusted probability data, reducing probability fluctuations through methods such as moving average or exponential smoothing to enhance state stability. When the stability assessment shows low stability, a probability reconstruction mechanism is activated to fuse probabilities by combining historical probability state information and the current adjustment results, and the probability state is re-estimated through methods such as Bayesian update or Kalman filtering. The corrected probability state also includes confidence annotations, and each probability component is associated with a corresponding confidence score, reflecting the reliability of the probability estimate. An anomaly detection mechanism monitors the reasonableness of the corrected probability state and triggers a re-verification process when an abnormal pattern appears in the probability distribution. A state caching mechanism stores the sequence of the most recent corrected probability states, providing historical data support for subsequent temporal analysis and trend prediction. In the practical application of the smart remote control, the corrected probability state assigns a probability weight of 0.87 to the "swipe right" gesture and a probability weight of 0.13 to the "stationary state", with confidence levels of 0.92 and 0.78 respectively, providing a reliable data basis for subsequent probability aggregation and final decision-making. Combining the stability assessment mechanism, a comprehensively calibrated corrected probability state is obtained.
[0054] Perform probability aggregation on the corrected probability state to determine the target gesture. The probability aggregation uses a multi-stage convergence algorithm to gradually converge multiple probability components in the quantum superposition state into a single deterministic recognition result. The aggregation process first performs probability sorting, arranging all gesture candidates in descending order according to the corrected probability weights, and identifying the main candidate with the highest probability and the alternative candidate with the second-highest probability. Competitive analysis evaluates the probability gap between the main candidate and the alternative candidate. When the probability advantage of the main candidate exceeds the preset threshold, it is determined as the target gesture. When the probability gap is small, a refined verification process is initiated. The refined verification further differentiates the competing candidates by increasing the verification dimension and improving the verification accuracy, and uses the deep features and temporal correlation information of multi-sensor signals for supplementary verification. The semantic consistency check ensures that the aggregation result conforms to the logical constraints of gesture semantics, and verifies the rationality and interpretability of the recognition result through the semantic knowledge base. The temporal coherence check evaluates the coherence between the current recognition result and the historical recognition sequence to avoid logically discontinuous gesture jumps. The aggregation algorithm also introduces decision confidence calculation, assigns a confidence score to the final recognition result, and reflects the reliability of the aggregation decision. Adaptive threshold adjustment dynamically corrects the aggregation threshold according to the real-time recognition effect, increases the aggregation threshold in scenarios with high-precision requirements, and appropriately reduces the threshold in scenarios with real-time response requirements. In the application of an intelligent remote control, when the "swipe right" target gesture is recognized and the confidence reaches 0.89, a corresponding device control instruction is immediately generated to realize the complete control link from gesture recognition to device response. According to the probability aggregation algorithm, a unique target gesture recognition result is determined.
[0055] Step S140, re-extract the verification feature data based on the target gesture result, input the verification feature data into a dynamic hierarchical inference model including a first-level classification model and a second-level classification model for classification verification, and output the gesture recognition result.
[0056] Re-extract the verification feature data based on the target gesture result. For the target gesture obtained by probability aggregation, a differential feature extraction strategy is adopted for secondary verification. By reexamining the different-dimensional features of the original sensor data, an independent verification basis is provided for the classification model. Determine the key regions and critical dimensions for feature extraction according to the type, confidence level, and execution parameters of the target gesture result. Determine the key regions and critical dimensions for feature extraction according to the type, confidence level, and execution parameters of the target gesture result. The feature reconstruction process simultaneously extracts key motion parameters, including IMU acceleration intensity features, millimeter-wave radar hand micro-motion speed features, multi-sensor signal confidence indicators, and ToF camera hand distance features. For dynamic gestures such as "swipe" and "wave", focus on extracting the motion trajectory features of the IMU and the micro-motion frequency features of the millimeter-wave radar to enhance the expression of temporal dynamic information; for static gestures such as "click" and "hover", focus on extracting the spatial positioning features and shape contour features of the ToF camera to enhance the accuracy of spatial geometric information. The feature reconstruction process accurately intercepts by tracing back the original multi-modal sensor data according to the time window and spatial range of the target gesture, ensuring a high degree of consistency between the verification feature data and the recognition target. The IMU acceleration feature reflects the motion intensity of the gesture action, the millimeter-wave radar micro-motion speed feature describes the speed change of the hand movement, the multi-sensor confidence indicator evaluates the reliability of the feature quality, and the ToF distance feature determines the spatial relationship between the hand and the device. When the target gesture is "swipe right", focus on extracting the IMU acceleration change sequence during the swipe period, the horizontal motion spectrum of the millimeter-wave radar, and the hand trajectory key points of the ToF camera, and simultaneously calculate the corresponding acceleration intensity, micro-motion speed, confidence level, and distance parameters to form a dedicated feature set for swipe verification. The verification feature data also includes quality evaluation indicators such as feature integrity, time synchronization, and sensor consistency. Through feature reconstruction processing, high-quality verification feature data containing motion parameters is generated.
[0057] In some embodiments, inputting the verification feature data into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification, and outputting a gesture recognition result, includes: triggering the primary classification model to perform a preliminary classification to obtain a preliminary classification result when the verification feature data meets the preset hierarchical condition; triggering the secondary classification model to perform a deep analysis to obtain a deep classification result when the verification feature data does not meet the preset hierarchical condition; generating a gesture recognition result based on the above classification results.
[0058] When the verification feature data meets the preset classification conditions, the first-level classification model is triggered for preliminary classification to obtain the preliminary classification result. The preset classification conditions include any one of the following conditions: the IMU acceleration intensity in the verification feature data is greater than 1.5g, the millimeter-wave radar hand micro-movement speed is greater than 0.5m / s, the multi-sensor signal confidence is greater than 85%, and the hand distance of the ToF camera is greater than 1.5m. The first-level classification model adopts a lightweight 1D-CNN architecture, which is specifically designed to handle high-dynamic gesture recognition tasks that meet the classification conditions, achieving fast response while ensuring recognition accuracy. The model inputs the IMU time-domain features and the millimeter-wave radar frequency-domain energy features in the verification feature data, and compresses the feature dimension to 128 dimensions to optimize the calculation efficiency. The network structure includes two one-dimensional convolutional layers and a fully connected layer. The first convolutional layer uses 64 convolutional kernels with a kernel size of 5 to extract local temporal patterns in the verification feature data; the max-pooling layer performs downsampling to reduce the computational amount; the second convolutional layer uses 128 convolutional kernels with a kernel size of 3 to extract higher-level abstract features; the final fully connected layer contains 256 neurons, and outputs the classification results of four basic gestures: left wave, right wave, click, and still. The network output is converted into a probability distribution through the Softmax activation function, and each gesture category corresponds to a probability value, forming a complete gesture probability vector. The activation function uses ReLU to enhance the non-linear expression ability of the network and prevent the problem of gradient disappearance. The model is trained using the cross-entropy loss function and the Adam optimizer, with the learning rate set to 0.001 and the batch size to 32. In practical applications, for example, when the verification feature data shows that the IMU acceleration reaches 1.5g, the first-level classification model is triggered when the classification conditions are met, and it can quickly recognize the "right wave" category with a confidence of 0.94 and a response time of less than 10 milliseconds. After being processed by the lightweight network, the preliminary classification result is obtained quickly.
[0059] Next, when the verification feature data does not meet the preset classification conditions, the secondary classification model is triggered for in-depth analysis to obtain the in-depth classification result. When the verification feature data does not meet any of the aforementioned preset classification conditions, a more complex secondary classification model needs to be enabled. The secondary classification model adopts the spatio-temporal frequency domain fusion network ST-FFN architecture, synchronously processes the ToF depth map sequence, millimeter-wave radar spectrogram, and IMU trajectory features in the verification feature data, and outputs the composite gesture semantic result. The ST-FFN model contains three specialized feature extraction branches and a fusion layer. The ToF branch uses a 3D convolutional neural network to extract the spatio-temporal features of the hand depth map in the verification feature data. The input is a sequence of 16 frames of 64×64 pixel depth maps. The first layer of 3D convolution contains 16 output channels, the convolution kernel size is 3×3×3, the stride is 1, the padding method is same, followed by ReLU activation and 3D max pooling; the second layer of 3D convolution contains 32 output channels, the convolution kernel size is 3×3×3, and the pooling size is 2×2×2; the third layer of 3D convolution contains 64 output channels, the convolution kernel size is 3×3×3, and finally a 64-dimensional spatio-temporal feature vector is output through global average pooling, retaining the spatio-temporal continuity of hand movement. The radar branch uses a 2D convolutional neural network to process the micro-Doppler spectrogram in the verification feature data. The first layer of 2D convolution contains 32 output channels, the convolution kernel size is 3×3, followed by ReLU activation and 2×2 max pooling; the second layer contains 64 output channels; the third layer contains 128 output channels, and finally a 128-dimensional spectral feature vector is output through global average pooling. The IMU branch uses a long short-term memory network LSTM to encode the six-axis motion trajectory in the verification feature data. The input is a time series containing three-axis acceleration and three-axis angular velocity. The number of LSTM hidden units is 64, and a 64-dimensional trajectory encoding feature is output. The fusion layer weights and splices the three-way features through the modal-level channel attention mechanism. The total feature dimension is 256. The attention weight generation network contains a fully connected layer FC(256→128), ReLU activation, FC(128→3), and a Softmax layer, generating a three-modal weight assignment. Finally, the classification layer uses Dropout = 0.3 to prevent overfitting. When processing the "swipe + click" composite gesture, when the verification feature data shows a low-speed composite action and a short distance, the ST-FFN model can accurately identify the trajectory features in the swipe stage and the pressure change in the click stage. Based on the deep fusion network architecture, the in-depth classification result of complex gestures is obtained.
[0060] Finally, the gesture recognition result is generated based on the above classification results. In the result generation process, the classification path is first identified to determine whether the current recognition result comes from the primary classification model or the secondary classification model, and the triggering conditions and processing path information are recorded. For the preliminary results of the primary classification model, the recognition quality is ensured through probability distribution analysis. When the confidence of the gesture category with the highest probability exceeds 0.85 and the gap with the second-highest probability category is greater than 0.3, this result is adopted; otherwise, the result stability test is carried out. For the in-depth analysis results of the secondary classification model, semantic rationality tests are conducted to ensure the logical consistency and execution feasibility of the composite gesture. The result fusion mechanism deals with the situation where both the primary and secondary models give valid results in special cases, and determines the final recognition result through a weighted fusion or voting mechanism. The weight assignment considers the applicability of the model and the complexity of the current scenario. The temporal consistency test evaluates the coherence of the current recognition result with the historical recognition sequence, and filters out unreasonable recognition jumps through state transition probabilities and gesture grammar rules. The output result includes complete information such as gesture category, confidence score, execution parameters, and timestamp. The gesture category uses a standardized encoding, the confidence score reflects the recognition reliability, and the execution parameters include quantitative indicators such as the spatial range, movement speed, and duration of the gesture. The anomaly detection mechanism monitors the rationality of the recognition result and triggers the re-recognition process when an abnormal pattern appears. In the application scenario of an intelligent remote control for controlling a TV, when the "swipe right" gesture is recognized with a confidence of 0.91, a complete recognition result including gesture category "RIGHT_SWIPE", confidence "0.91", swipe distance "15 cm", and swipe speed "0.8 m / s" is generated to support the generation of subsequent device control instructions. Through the unified result generation mechanism, the dual confirmation of probability aggregation verification and classification model verification is integrated to output a reliable final gesture recognition result.
[0061] Step S150: Detect the current communication environment state, construct a StarFlash and UWB dual-mode communication race condition based on the gesture recognition result and the communication environment state, and aggregate the dual-mode communication race condition into the optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
[0062] Specifically, detect the current communication environment status, and construct a StarFlash and UWB dual-mode communication race based on the gesture recognition result and the communication environment status. The communication environment status detection uses multi-sensor fusion monitoring technology to evaluate the working environment quality of the StarFlash and UWB communication modes in real time. The StarFlash environment detection obtains the real-time performance parameters of StarFlash communication through signal strength scanning, network congestion analysis, and frequency band occupancy monitoring. When the signal strength is higher than -70 dBm and the network delay is lower than 50 milliseconds, it is marked as an excellent environment. The UWB environment detection evaluates the environmental adaptability of UWB communication through pulse response measurement, multipath interference analysis, and positioning accuracy verification. When the positioning error is less than 10 centimeters and the signal-to-noise ratio is greater than 15 dB, it is confirmed as a suitable environment. The analysis of the gesture recognition result extracts the communication requirement features according to the gesture category, confidence level, and execution parameters output by the gesture recognition. Simple gestures such as "click" require low-latency and fast response, composite gestures such as "slide + rotate" require high-bandwidth parameter transmission, and fine gestures such as "fine-tuning" require high-precision positioning support. The construction of the dual-mode communication race performs cross-analysis on the environmental adaptability of StarFlash and UWB and the gesture communication requirements through performance matching degree calculation to generate dynamic competition weights. When the gesture is "precision positioning operation" and the UWB environment quality is excellent, UWB obtains a higher race weight; when the gesture is "batch device control" and the StarFlash network is stable, StarFlash gains an advantageous position. The race balance point is determined through real-time weighted calculation, and the weight distribution is dynamically adjusted according to environmental changes, demand changes, and historical performance. Through environmental perception and demand matching, a dynamically balanced dual-mode communication race is established.
[0063] In some embodiments, aggregating the dual-mode communication race into an optimal communication mode at the instant of instruction sending includes: detecting an instruction sending trigger signal; performing instantaneous convergence on the dual-mode communication race based on the trigger signal to obtain convergence result data; determining a unique optimal solution through the convergence result data; and converting the unique optimal solution into an optimal communication mode.
[0064] First, detect the instruction sending trigger signal and establish a time synchronization mechanism for communication mode decision-making. The trigger signal detection adopts a multi-level monitoring framework, including three levels: main trigger, auxiliary trigger, and forced trigger, which real-time monitors user gesture recognition completion events, remote control status change events, and external control instruction events. The gesture recognition completion trigger serves as the main trigger signal and is achieved by monitoring the status change of the gesture recognition result. Specifically, it includes three detection dimensions: confidence threshold detection, gesture category confirmation, and execution parameter verification. When the recognition confidence exceeds 0.85, the gesture category is determined, and the execution parameters are complete, a main trigger signal is generated. The remote control status trigger serves as the auxiliary trigger signal, monitoring changes in power status, connection status, and working mode, including state transitions such as from standby to active, from disconnected to connected, and from silent to responsive. The status change detection is achieved through a dual mechanism of hardware interruption and software polling. The external control trigger serves as the forced trigger signal, responding to the user's physical button operations, voice commands, or other interaction methods, providing a forced trigger mechanism for emergency control scenarios. The external trigger has the highest priority and can interrupt the current processing flow. The trigger signal priority management adopts a preemptive scheduling strategy to ensure that multiple triggers are processed according to the preset priority when they occur simultaneously. The gesture recognition trigger has the highest priority, followed by the status trigger, and the external control trigger can preempt all other triggers in case of emergency. The trigger time window control prevents overly frequent trigger events from affecting stability through a sliding time window mechanism. The minimum trigger interval is set to 50 milliseconds, and consecutive trigger events are processed with time merging. The same type of trigger is only processed once within the time window. The signal strength evaluation determines the reliability of the trigger by multi-dimensionally analyzing the confidence, stability, and consistency of the trigger event. Only trigger signals with an intensity exceeding the dynamically adjusted threshold can initiate the subsequent convergence process. In the TV remote control scenario, when the user completes the "swipe right" gesture and the recognition confidence reaches 0.91, the gesture recognition completion event immediately generates a trigger signal to initiate the convergence process of communication mode selection. Precise capture of the instruction sending trigger signal is achieved using multi-level monitoring technology.
[0065] Next, based on the trigger signal, instantaneous convergence of the dual-mode communication race condition is performed to obtain the convergence result data. The instantaneous convergence algorithm acts on the established dual-mode communication race condition, adopting a strategy that combines forced convergence and natural convergence, converging the competing state between StarFlash and UWB from an uncertain coexistence state to a definite single-choice state. The convergence process is achieved through a race condition resolution mechanism, which includes three stages: race condition breaking, direction selection, and state solidification. When the trigger signal arrives, the dynamic balance state of the dual-mode communication race condition is instantly broken, and the race condition is forced to converge to a single mode. The convergence algorithm adopts the principle of minimizing race condition potential energy. By calculating the potential energy distribution of StarFlash and UWB in the current race condition, the competition intensity, stability, and convergence trend of StarFlash and UWB in the dual-mode communication race condition are analyzed, and the communication mode with the lowest potential energy and the highest stability is selected as the convergence target. The race condition convergence speed is achieved through a multiple control mechanism, including convergence time window control, convergence acceleration factor adjustment, and convergence interruption protection. The standard convergence time is 10 milliseconds, ensuring that the race condition can complete the complete conversion from competition to determination before the instruction is sent. The convergence stability guarantee is achieved through a continuous multiple race condition sampling verification mechanism, including sampling frequency control, sampling consistency test, and sampling result statistical analysis. When the continuous sampling results point to the same communication mode and the confidence level increases, the convergence is confirmed to be completed. The race condition conflict resolution mechanism specifically deals with the evenly matched situation with intense competition. When the dual-mode communication race condition shows a state of equal strength, external decision-making factors such as environmental priority, historical preference, and energy consumption consideration are introduced to forcefully break the race condition balance, ensuring that the convergence process can produce clear and executable results. The convergence process monitoring uses real-time trajectory tracking technology to record the complete conversion process of the dual-mode communication race condition from the competition state to the determination state, generating comprehensive convergence result data including convergence direction, convergence speed, stability index, and confidence level distribution. When the trigger signal activates the convergence algorithm, the dual-mode communication race condition quickly converges from the competing state of coexistence between StarFlash and UWB to the determined state with StarFlash being preferred. The convergence result data shows that StarFlash becomes the only choice and the convergence confidence level is 0.92. Through the instantaneous race condition convergence mechanism, the uncertainty of the dual-mode communication race condition is quickly resolved and the convergence result data is generated.
[0066] Then, the unique optimal solution is determined through the convergence result data to achieve a clear selection of the communication mode. The optimal solution determination process adopts a strategy combining result verification and exception handling, analyzes the convergence direction, convergence speed, stability index, and confidence distribution in the convergence result data, and extracts the final selection result and reliability assessment of race convergence. The convergence result verification is achieved through a multi-dimensional quality inspection mechanism, including convergence confidence inspection, stability index verification, consistency analysis, and timeliness evaluation. When the convergence confidence exceeds 0.8, the stability index meets the requirements, and the result consistency is good, the convergence result is confirmed as the optimal solution. The convergence exception handling mechanism specifically addresses abnormal situations such as insufficient convergence confidence, oscillation, or inconsistent results, and enables an alternative decision-making mechanism including historical data analysis, user preference matching, and environmental adaptability re-evaluation for supplementary analysis. The optimal solution is determined through means such as historical performance data statistics, user preference setting queries, and real-time environment re-detection. The optimal solution consistency test adopts a time series analysis method to evaluate the coherence and rationality of the current selection with the historical decision-making mode. Through technical means such as decision history review, pattern recognition analysis, and abnormal change detection, frequent communication mode switching is avoided to affect stability and user experience. The decision timeliness control adopts a parallel processing and priority scheduling mechanism to ensure that the optimal solution determination process is completed within 5 milliseconds, meeting the strict response requirements of real-time control, including fast path optimization, concurrent verification processing, interruption response mechanism, and other technical guarantees. The decision result record adopts a structured storage method to save the selection basis, confidence information, decision path, and timestamp of the optimal solution, supporting subsequent performance analysis, optimization adjustment, and fault diagnosis. The convergence result data shows that the convergence confidence of the SparkLink mode is 0.92 and the stability index is good, confirming SparkLink communication as the unique optimal solution without additional decision analysis process. Based on the analysis of the convergence result data, the unique optimal solution is quickly determined.
[0067] Finally, the unique optimal solution is converted into the optimal communication mode to complete the communication configuration and instruction sending of the smart remote control. The communication mode conversion adopts a hierarchical conversion architecture, including three conversion levels: hardware abstraction layer encapsulation, protocol parameter configuration, and interface adaptation control, which maps the abstract optimal solution into specific communication hardware configurations and protocol parameters. The XingShan communication mode configuration includes four configuration dimensions: frequency band selection, power control, coding method, and transmission protocol setting. The communication parameters are optimized according to the current network environment and device capabilities to obtain the best performance. The frequency band selection uses dynamic spectrum sensing technology to avoid interference bands, and the power control uses adaptive power management to balance the transmission distance and energy consumption efficiency. The UWB communication mode configuration covers four core configuration items: pulse parameters, time-domain synchronization, spatial positioning, and anti-interference settings to ensure stable communication in complex electromagnetic environments. The optimization of pulse parameters improves the signal penetration ability, and the time-domain synchronization ensures the time accuracy of data transmission. The mode switching control realizes seamless switching through a unified interface standard and switching protocol, including four switching stages: switching preparation, status saving, mode activation, and status restoration, avoiding data loss and increased latency during the switching process. The switching time is controlled within 5 milliseconds to ensure the continuity of the user experience. The communication protocol adaptation adopts a multi-layer protocol conversion mechanism to ensure that the gesture recognition results can be correctly encoded into the control instruction format recognizable by the target device, supporting multiple downstream communication protocols such as infrared, Bluetooth, and WiFi. The protocol conversion includes processing steps such as format adaptation, parameter mapping, and checksum addition. The instruction encapsulation adopts a standardized data structure to convert the semantic information of gesture recognition into standardized device control instructions, including complete information such as instruction type, parameter value, priority, and timestamp. The instruction format supports extension to meet the control requirements of different devices. The transmission quality assurance is achieved through multiple guarantee mechanisms, including error detection, retransmission mechanism, acknowledgment response, timeout handling, and other technical means to ensure the reliable transmission of instructions. The maximum retransmission times are set to 3 times and the timeout time is set to 100 milliseconds. User feedback and alternative processing solutions are provided in case of transmission failure. The status feedback mechanism adopts a two-way communication architecture to monitor the success status of instruction sending and the device response, providing real-time feedback on the operation results to the user, including feedback information such as sending confirmation, execution status, and error prompt. In the application of TV volume adjustment, after selecting the XingShan communication mode, the "swipe right" gesture is converted into a standard TV control instruction of "volume +1" and sent to the smart TV device through the XingShan protocol to achieve an accurate volume control response. Combining hardware abstraction and protocol adaptation, the intelligent control of the remote control is completed.
[0068] To execute the gesture recognition-based smart remote control method corresponding to the above method embodiments to achieve the corresponding functions and technical effects. Refer to Figure 2 , Figure 2The block diagram of an intelligent remote control device 200 based on gesture recognition provided by an embodiment of the present application is shown. For ease of description, only the parts related to this embodiment are shown. The intelligent remote control device 200 based on gesture recognition provided by the embodiment of the present application includes: A feature processing module 201, configured to obtain multimodal sensor data, extract gesture spatio-temporal features from the multimodal sensor data, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features; A probability construction module 202, configured to obtain a candidate gesture probability superposition state based on the exchanged dimension features, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; A state correction module 203, configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine a target gesture; A hierarchical inference module 204, configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical inference model including a first-level classification model and a second-level classification model for classification verification, and output a gesture recognition result; A communication control module 205, configured to detect the current communication environment state, construct a StarFlash and UWB dual-mode communication race based on the gesture recognition result and the communication environment state, and simultaneously establish a contradictory decision parallel processing mechanism for selecting StarFlash and selecting UWB, and aggregate the dual-mode communication race into an optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
[0069] The above-mentioned intelligent remote control device 200 based on gesture recognition can implement the intelligent remote control method based on gesture recognition in the above method embodiment. The optional items in the above method embodiment are also applicable to this embodiment, which will not be elaborated here. The remaining content of the embodiment of the present application can refer to the content of the above method embodiment, and will not be repeated in this embodiment.
[0070] As Figure 3 shown, the third embodiment of the present invention further provides a computer device, including a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302, wherein the processor 302 implements the steps of the intelligent remote control method based on gesture recognition described in the first embodiment of the present invention when executing the program.
[0071] The purpose of the above embodiments is to exemplarily reproduce and deduce the technical solution of the present invention, and to completely describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to understand the disclosed content of the present invention more thoroughly and comprehensively, and it does not limit the protection scope of the present invention hereby.
[0072] The above embodiments are not exhaustive listings based on the present invention. In addition, there may be multiple other embodiments not listed. Any substitution and improvement made on the basis of not violating the concept of the present invention fall within the protection scope of the present invention.
Claims
1. An intelligent remote control method based on gesture recognition, characterized in that, Including: Obtain multi-modal sensor data, extract features from the multi-modal sensor data to obtain gesture spatio-temporal features, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features; Based on the exchanged dimension features, obtain the candidate gesture probability superposition state, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; Based on the inversion verification parameter, perform weight correction on the candidate gesture probability superposition state to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine the target gesture; Extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification, and output a gesture recognition result; Detect the current communication environment state, construct a StarFlash and UWB dual-mode communication race condition based on the gesture recognition result and the communication environment state, and aggregate the dual-mode communication race condition into an optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
2. The method according to claim 1, wherein The performing spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features includes: Decompose the gesture spatio-temporal features into a time dimension component and a space dimension component; Perform cross mapping on the time dimension component and the space dimension component; Reconstruct a dimension correlation matrix based on the cross mapping result; Generate exchanged dimension features using the dimension correlation matrix.
3. The method according to claim 1, characterized in that The obtaining the candidate gesture probability superposition state based on the exchanged dimension features includes: Perform probability distribution analysis on the exchanged dimension features to obtain probability component data; Construct a multi-situation energy field based on the probability component data; Analyze the multi-situation energy field to generate energy distribution parameters; Generate a candidate gesture probability superposition state through the energy distribution parameters.
4. The method according to claim 1, wherein The constructing a semantic negation hypothesis for the candidate gesture probability superposition state includes: Perform semantic reverse analysis on the candidate gesture probability superposition state; Construct an antonym semantic set based on the semantic reverse analysis result; Generate a semantic negation hypothesis using the antonym semantic set.
5. The method according to claim 1, characterized in that, The generating an inversion verification parameter based on the semantic negation hypothesis includes: Convert the semantic negation hypothesis into a verification constraint condition; Analyze the verification constraint condition to obtain reverse verification indicators; Perform parameter quantization processing based on the reverse verification indicators to generate an inversion verification parameter.
6. The method according to claim 1, characterized in that, The performing weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state includes: Generate a weight correction coefficient using the inversion verification parameter; Perform weight adjustment on the candidate gesture probability superposition state based on the weight correction coefficient to obtain adjusted probability data; Analyze the adjusted probability data to generate a stability evaluation result; Generate a corrected probability state based on the stability evaluation result.
7. The method according to claim 1, characterized in that The inputting the verification feature data into a dynamic hierarchical inference model including a primary classification model and a secondary classification model for classification verification and outputting a gesture recognition result includes: When the verification feature data meets the preset classification condition, trigger the primary classification model to perform primary classification to obtain a primary classification result; When the verification feature data does not meet the preset classification condition, trigger the secondary classification model to perform in-depth analysis to obtain a depth classification result; Generate a gesture recognition result based on the above classification results.
8. The method according to claim 1, characterized in that The aggregating the dual-mode communication race into an optimal communication mode at the moment of instruction sending includes: Detect an instruction sending trigger signal; Based on the trigger signal, instantaneously converge the dual-mode communication race to obtain convergence result data; Determine a unique optimal solution through the convergence result data; Convert the unique optimal solution into an optimal communication mode.
9. An intelligent remote control device based on gesture recognition, characterized in that, Including: A feature processing module, configured to obtain multi-modal sensor data, extract gesture spatio-temporal features from the multi-modal sensor data, and perform spatio-temporal dimension exchange on the gesture spatio-temporal features to generate exchanged dimension features; A probability construction module, configured to obtain a candidate gesture probability superposition state based on the exchanged dimension features, construct a semantic negation hypothesis for the candidate gesture probability superposition state, and generate an inversion verification parameter based on the semantic negation hypothesis; A state correction module, configured to perform weight correction on the candidate gesture probability superposition state based on the inversion verification parameter to generate a corrected probability state, and perform probability aggregation on the corrected probability state to determine a target gesture; A hierarchical reasoning module, configured to extract verification feature data based on the target gesture, input the verification feature data into a dynamic hierarchical reasoning model including a primary classification model and a secondary classification model for classification verification, and output a gesture recognition result; A communication control module, configured to detect the current communication environment state, construct a StarFlash and UWB dual-mode communication race based on the gesture recognition result and the communication environment state, and aggregate the dual-mode communication race into an optimal communication mode at the moment of instruction sending to complete the intelligent control of the remote control.
10. A computer device, characterized in that, Including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
FMCW radar gesture recognition method based on spatial-temporal feature sequence
CN117828468A
Dynamic gesture recognition method and system, electronic equipment and computer readable storage medium
CN118196879A
Gesture control method and system for intelligent wearable device
CN118885078A
Smart home multi-modal human-machine natural interaction system and method thereof
WO2022110564A1
Cited By
Humanoid robot remote control method, device and equipment based on wireless communication
CN120620231A
Multi-mode trajectory prediction photoelectric tracking system
CN121834141A
Three-mode remote control cross-device intelligent control method and system supporting star flash technology
CN121938171A
Three-mode remote control cross-device intelligent control method and system supporting star flash technology
CN121938171B