A viewing interactive method and system based on gesture recognition and a storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-08-11
AI Technical Summary
1.实现了隐私保护型的自然交互:本发明通过惯性传感器和蓝牙定位,不依赖摄像头或麦克风即可识别手势动作,有效保护了用户隐私。相比视觉识别方案,系统端到端延迟为50-100毫秒,性能提升3倍。
Smart Images

Figure CN122053895B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, specifically to a movie-watching interaction method, system, and storage medium based on gesture recognition. Background Technology
[0002] With the increasing functionality of smart TVs, traditional remote control interaction methods can no longer meet users' demands for convenience and immersion. Current contactless interaction solutions mainly include camera-based visual recognition and wearable device-based inertial recognition. Visual solutions are limited by lighting conditions and pose a serious risk of home privacy leaks. Existing inertial recognition technologies based on wristbands or single rings have the following drawbacks: The device cannot distinguish whether the user is in a valid viewing position or is just passing through the living room, causing the device to be in a state of continuous high power consumption monitoring. It is also prone to false triggers caused by capturing the user's daily activities. Existing technologies mostly rely on simple acceleration threshold judgments and cannot distinguish actions with similar signal characteristics (for example, misinterpreting the vibration of "slapping the thigh" as "clapping" or "unintentional waving" as "command gesture"). It lacks in-depth verification of the physical mechanism of the action (such as the hardness of the impact medium and the direction of momentum transfer), resulting in extremely low usability in complex life scenarios.
[0003] This paper proposes a privacy-preserving system that does not require a camera or microphone. The system should reliably determine the user's actual presence while watching the video and accurately recognize everyday gestures such as clapping and tapping, translating these natural actions into interactive television commands. To this end, a gesture recognition-based interactive viewing method, system, and storage medium are presented. Summary of the Invention
[0004] This invention aims to solve the problems of high accidental touch rate and single recognition dimension of existing wearable interactive devices in complex living room environments, and provides a movie-watching interaction method, system and storage medium based on gesture recognition, which achieves accurate capture of user intention through dual-ring linkage.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A gesture recognition-based interactive movie-watching method includes: A dynamic geometric region is set as a presence gate based on preset angle and distance thresholds. When a user wearing two rings enters the presence gate, the user's presence status is established. The base station continuously receives and analyzes the inertial data stream and wireless signal characteristics reported by the rings, and performs filtering and gravity compensation processing on the inertial data stream. A left-hand model is constructed to distinguish between left and right hands based on ring ID and wireless signal characteristics; based on the processed inertial data stream, the vibration frequency, acceleration waveform, and time difference characteristics of the two rings are analyzed in real time to perform gesture recognition, and a gesture action is matched from a predefined gesture dictionary; a feature verification model is constructed to verify the matched gesture, and after verification, the gesture action is mapped to a measure of interactive participation, and the base station sends instructions to the TV through USBHID or virtual serial port; The TV module parses the instructions and converts the interactive participation into program interactive feedback, which includes, but is not limited to, playback control, program interaction, and advertising feedback event reporting.
[0006] Preferably, the presence gate is defined by a preset angle threshold and a distance threshold. The spatial position is calculated based on Bluetooth 6.0 direction measurement and distance measurement, and is obtained in real time through wireless signal interaction between the user's dual smart rings and the TV base station. The presence gate is a dynamic geometric area, which is a fan-shaped or rectangular area in front of the TV base station. The angle threshold can be set to the viewing angle range of the TV screen, and the distance threshold can be set to the typical living room viewing distance. When the user's spatial position meets these threshold conditions, the user's presence status is established as valid viewing, and subsequent gesture recognition and interaction processes are activated. If the position exceeds the threshold, it is determined to be an absence status, and data collection is stopped to ensure privacy protection and data validity.
[0007] Preferably, the determination of the presence gate further includes an adaptive manifold region of irregular shape: During system initialization, a preset generalized sector area is set as the initial gate. As user interaction data is generated, the system records the spatial coordinates of each gesture that is determined to be valid. The coordinates are processed using a clustering algorithm with a time decay factor to dynamically adjust the gate boundary, so that the boundary gradually converges and fits the user's actual viewing activity range. When the user's coordinates are detected to be within the manifold area, the high-frequency gesture sampling mode is activated. If the user is outside the area, it is determined to be an off-screen or invalid area, and the system enters a low-power standby state.
[0008] Preferably, the steps of filtering and gravity compensation processing the inertial data stream include: The inertial data stream originates from the three-axis accelerometer and gyroscope built into the dual smart ring. The collected data includes acceleration vector, angular velocity, and vibration signal. The data stream is filtered by low-pass filtering to remove noise and environmental interference. Gravity compensation processing is performed, and the accelerometer's attitude is corrected using gyroscope data to separate the pure motion acceleration component. A left-hand and right-hand model is constructed, and the ring has a unique ID. The system distinguishes between the left and right hands based on the ID and wireless signal characteristics. The model further integrates time series data to analyze the source of the action and the type of behavior.
[0009] Preferably, matching a gesture from a predefined gesture dictionary specifically includes: Key features are extracted from the processed inertial data stream, including vibration frequency, acceleration waveform, and double-ring time difference. The system uses a pattern matching algorithm to compare the key features with a predefined gesture dictionary. The gesture dictionary consists of short, trigger-type actions, including at least: mutual slapping gestures such as slapping each other once or twice; object-touching gestures such as slapping the leg with one or both hands, or slapping the armrest of a seat with one or both hands; micro-movement gestures such as snapping fingers by rubbing the thumb and middle finger; and air gestures such as clenching a fist forward with one or both hands, or clenching a fist instantly when the palm changes from an open state to a closed state.
[0010] Preferably, a feature verification model is constructed to verify the matched gestures. The model includes a spatial motion module and a signal dynamics module. The spatial motion module is used to analyze the relative trajectory characteristics of the two rings: for mutual patting gestures, it verifies whether they conform to the mutual convergence model; for touching gestures, it verifies whether they conform to the parallel motion model in the same direction; for aerial gestures and micro-motion gestures, it verifies whether they conform to the hovering braking model. The signal dynamics module is used to analyze the frequency and time domain characteristics of inertial signals: for mutual slapping gestures, it matches high-frequency rigid impact characteristics; for touching gestures, it matches damped vibration characteristics accompanied by energy decay; for aerial gestures, it matches non-collision active emergency stop characteristics; for micro-motion gestures, it matches single pulse characteristics under static background; the system outputs the final interactive command only when the output results of the spatial motion module and the signal dynamics module simultaneously meet the consistency of the preset classification.
[0011] Preferably, converting the instructions into playback control and program interaction feedback includes: A mapping logic between gesture commands and system functions is pre-established; for playback control feedback: when the parsed command is a control command, the TV module directly executes playback control operations, including pause, play, fast forward, rewind, or volume adjustment, by calling the underlying player interface of the system; for program interaction feedback: the TV activates a transparent rendering layer to generate dynamic effects and / or drive changes in virtual props in real time based on gesture commands, achieving visual feedback synchronized with the video.
[0012] Preferably, the method further includes reporting advertising feedback events: The TV module maintains a mapping table between playback timeline and content metadata. When it receives an interactive instruction from the base station, the system captures the timestamp of the current playback screen and the corresponding ad creative ID. It associates the interactive engagement metric with the ad creative ID to generate an active attention log. When the interactive engagement metric exceeds a preset threshold, an event marker is triggered, and the TV module uploads the active attention log to the cloud ad server in real time through the network layer.
[0013] Preferably, the system includes: at least one ring containing an inertial sensor and a first wireless transceiver; a base station associated with a television, the base station including an antenna array for performing angle of arrival measurements, a second wireless transceiver, a processor, and an interface for communicating with the television; A memory that stores instructions that, when executed by the processor, cause the system to perform the method.
[0014] Preprocessing module: Based on preset angle and distance thresholds, a dynamic geometric region is set as an presence gate. When a user wearing two rings enters the presence gate, the user's presence status is established. The base station continuously receives and analyzes the inertial data stream and wireless signal characteristics reported by the rings, and performs filtering and gravity compensation processing on the inertial data stream. Gesture matching module: Constructs left and right hand models to distinguish between left and right hands based on ring ID and wireless signal characteristics; Based on the processed inertial data stream, it analyzes vibration frequency, acceleration waveform, and time difference characteristics between the two rings in real time to perform gesture recognition, and matches a gesture action from a predefined gesture dictionary; Constructs a feature verification model to verify the matched gesture, and after verification... Parsing module: Maps the gestures to a measure of interactive engagement. The base station sends instructions to the TV via USBHID or a virtual serial port. The TV module parses the instructions and converts the interactive engagement into program interaction feedback, which includes, but is not limited to, playback control, program interaction, and advertising feedback event reporting.
[0015] Preferably, when the computer program is executed by the processor, it implements the steps of the gesture recognition-based movie-watching interaction method.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Achieves privacy-preserving natural interaction: This invention uses inertial sensors and Bluetooth positioning to recognize gestures without relying on cameras or microphones, effectively protecting user privacy. Compared to visual recognition solutions, the system's end-to-end latency is 50-100 milliseconds, resulting in a 3x performance improvement.
[0017] 2. Improved data collection effectiveness: Through the presence gate mechanism, the system can automatically filter invalid viewing data such as those from leaving the site or being in standby mode. Compared with schemes that do not use presence determination, the proportion of effective data is increased by 35%.
[0018] 3. Expanded richness of interaction parameters: Compared with the discrete buttons of traditional remote controls (which can only express 0 or 1), this system can express more granular interaction commands based on continuous acceleration intensity and trigger frequency parameters, expanding the dimension of interaction parameters from 1-dimensional to 2-dimensional.
[0019] 4. Achieved high-precision time alignment: This invention improves time resolution by 10 times compared to the traditional second-level precision solution by associating millisecond-level timestamps with playback content metadata. The time resolution between two adjacent interactive events is improved from 1 second to 100 milliseconds. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a gesture recognition-based interactive movie-watching method according to the present invention. Figure 2 This is a schematic diagram of the user presence status management process proposed in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structure of a movie-watching interactive system based on gesture recognition according to the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Please see Figures 1 to 3 This invention provides a method, system, and storage medium for interactive movie viewing based on gesture recognition, referring to... Figure 1 Flowchart and Figure 3 The structural diagram and technical solution are as follows: A gesture recognition-based interactive movie-watching method includes: A dynamic geometric region is set as an presence gate based on preset angle and distance thresholds. When a user wearing two rings enters the presence gate, the user's presence status is established. The base station continuously receives and analyzes the inertial data stream and wireless signal characteristics reported by the ring, and performs filtering and gravity compensation processing on the inertial data stream. A left-hand model is constructed to distinguish between left and right hands based on ring ID and wireless signal characteristics; based on the processed inertial data stream, the vibration frequency, acceleration waveform, and time difference characteristics of the two rings are analyzed in real time to perform gesture recognition, and a gesture action is matched from a predefined gesture dictionary; a feature verification model is constructed to verify the matched gesture, and after verification, the gesture action is mapped to a measure of interactive participation, and the base station sends instructions to the TV through USBHID or virtual serial port; The TV module parses the instructions and converts the interactive participation into program interactive feedback, which includes, but is not limited to, playback control, program interaction, and advertising feedback event reporting.
[0023] Example 1 In this embodiment, in a home living room movie-watching scenario, the TV base station is configured at the center directly below the display device as the origin of the spatial coordinate system. The user wears a smart ring on each hand and establishes a real-time two-way wireless communication link with the TV base station using the Bluetooth 6.0 protocol. The high-precision reference signal periodically sent by the TV base station triggers the two smart rings to respond with feedback packets.
[0024] Reference Figure 2 The diagram below illustrates the user presence status management process of the present invention. Specifically, the wireless processing chip built into the TV base station acquires phase difference information containing 128 subcarriers, and performs signal sampling in conjunction with a frequency hopping pulse sequence with an 80 MHz bandwidth. By calculating the flight time of the signal in free space and performing correlation analysis on the sampling points, the initial radial distance between the dual smart rings and the TV base station is obtained.
[0025] Furthermore, the TV base station is configured with an array of four equally spaced dipole antennas to receive the feedback signal from the dual smart rings. The original I / Q signals received by each antenna channel are normalized and calibrated. The phase offset between adjacent antenna elements is calculated and the physical quantity is linearly mapped to a horizontal detection dimension of -90 degrees to +90 degrees, thereby determining the deflection angle of the dual smart rings relative to the center normal of the TV screen.
[0026] Specifically, the system presets an angle threshold of 45 degrees that extends horizontally to the left and right along the axis of the TV screen normal, and a distance threshold of 0.5 meters to 4.5 meters in depth. The physical boundaries of these two dimensions together construct a fan-shaped or rectangular dynamic geometric area in front of the TV, which serves as a presence gate to determine whether the user is in a valid viewing position.
[0027] Furthermore, the system inputs the real-time collected distance sequence and angle mapping sequence into a preset spatial coordinate refined convolutional neural network model. The first convolutional layer of this model contains 32 convolutional kernels of size 3x3, with a stride of 1 and zero padding to maintain the feature map size. The second convolutional layer is configured with 64 convolutional kernels of size 3x3, followed by a 2x2 max pooling layer to extract the local saliency of the positional features. The third layer is a fully connected layer containing 128 neurons, which are processed by the ReLU activation function and output the double-ring spatial three-dimensional coordinates after removing multipath interference.
[0028] Specifically, when the coordinates of the two smart rings are simultaneously within a preset ±45 degree angle threshold and the distance value is within the presence gate of 0.5 meters to 4.5 meters, the logic controller determines that the user's presence status is valid and generates an enable level to be sent to the subsequent gesture recognition processing chip, thereby starting the gesture command capture process based on millimeter-wave radar or image sensor.
[0029] Furthermore, the system dynamically corrects the boundary parameters of the geometric region based on the relative displacement changes between the two smart rings. For example, when the system detects that the distance between the left and right rings fluctuates between 30 cm and 50 cm, it automatically maintains a distance threshold upper limit of 4.5 meters to cover the entire living room depth and adjusts the edge tolerance of the angle threshold in real time according to the change ratio of the ring distance.
[0030] Specifically, the system establishes a position polling mechanism with a period of 20 milliseconds, inputting five sets of continuously acquired coordinate data into a sliding window filter for smoothing. If the coordinates of any ring deviate from the preset angle or distance threshold and continue for more than three periods, the system immediately determines that the user is in an off-site state and cuts off the data inflow of the gesture recognition algorithm module. At the same time, it performs an immediate physical erasure operation on all Bluetooth physical layer feature data in the buffer area.
[0031] Furthermore, a hysteresis comparator mechanism and geomagnetic fingerprint verification are added to the judgment logic of the on-site gate to construct a two-layer buffer area containing "entry threshold" and "exit threshold". Specifically, the system sets the inner ring boundary Rin and the outer ring boundary Rout, where Rin < Rout. When a user enters from the outside, they must break through the smaller Rin to trigger the "present" state; once present, they are only considered "absent" when their position moves outside the larger Rout. Additionally, in the boundary fuzzy zone, the base station records the geomagnetic intensity vector of the spatial point as an auxiliary feature. When the Bluetooth AoA signal fluctuates and causes the positioning to jump near the threshold, the system compares the current geomagnetic reading with the pre-stored "sofa area" geomagnetic fingerprint. If the geomagnetic feature matching degree is higher than 90%, even if the Bluetooth coordinates instantaneously drift out of bounds, the system still maintains the "present" determination, preventing misjudgment caused by signal multipath effects. Through hysteresis logic, it effectively solves the frequent and repeated jumping of the "present / absent" state caused by the user's slight movement at the gate edge, ensuring the stability of the system state; introducing geomagnetic auxiliary verification compensates for the defect that single Bluetooth positioning is vulnerable to interference in a complex indoor electromagnetic environment, ensuring the continuity of the viewing interaction.
[0032] A closed-loop from the original wireless physical layer signal to high-order logical determination is achieved. The input end is the complex-domain signal received by the Bluetooth antenna. After phase calculation and the non-linear mapping of the convolutional neural network, a binary presence control variable is finally output, achieving precise control and privacy protection of the user's state in the viewing scenario.
[0033] In the initialization stage of the TV base station, the system sets a generalized fan-shaped area with a preset angle of 90 degrees and a radius depth of 4.5 meters in the three-dimensional space directly in front of the display device as the initial presence gate.
[0034] Specifically, the user conducts Bluetooth 6.0 wireless communication with the TV base station through dual smart rings worn on both hands. The base station captures the signal phase difference and converts the physical characteristics of 128 subcarriers into angle values in the polar coordinate system. At the same time, it uses a frequency hopping sequence with a bandwidth of 80 megahertz to obtain the flight time data of signal propagation, thereby real-time generating the three-dimensional coordinate points of the dual smart rings in space.
[0035] Furthermore, when the user conducts gesture interaction within the initial presence gate, the system control module real-time records the spatial position information each time a valid gesture instruction is triggered, and stores these four-dimensional data packets containing the abscissa, ordinate, vertical coordinate, and timestamp into the system's temporary feature matrix.
[0036] Specifically, the system calls the pre-constructed adaptive manifold clustering model to process the stored coordinate points. This model adopts a neural network architecture with a three-layer structure, where the input layer receives the four-dimensional position vector, the middle hidden layer contains 64 neurons and uses the Adam optimizer for weight update, and the output layer generates a manifold feature vector representing the current user activity center and boundary range through non-linear mapping.
[0037] Furthermore, during the clustering process, the system assigns a weight coefficient to each coordinate point that decreases over time. This weight coefficient is calculated using an exponential function with a decay constant of 0.01, which reduces the contribution of old coordinate points that are more than 600 seconds away from the current time point to the boundary to less than 10% of the initial value, thereby ensuring that the gate boundary can track the dynamic drift of the user's position.
[0038] Specifically, the system calculates an irregularly shaped envelope based on the weighted coordinate cluster, and converges the initial 90-degree sector region into an adaptive manifold region that closely follows the user's actual operation trajectory. This region forms a high-density sampling core area at the center of the sofa where the user often sits or lies, while physically shrinking the edge areas far from the operation path.
[0039] Furthermore, when the system detects that the real-time spatial coordinates of the dual smart rings cross the boundary of the aforementioned adaptive manifold region and enter the interior, the controller immediately sends a command to the Bluetooth communication chip to increase the signal sampling frequency from the initial 50 Hz to a high-frequency sampling mode of 200 Hz, and simultaneously activates the gesture feature extraction algorithm in the backend.
[0040] Specifically, if the spatial coordinate values fed back by the dual smart rings exceed the boundary range defined by the adaptive manifold region, the system determines that the user has entered an off-field state or is located in an invalid interaction zone. At this time, the logic gate circuit outputs a low-level signal, forcibly switching the Bluetooth receiving link to a low-frequency listening state of 1 Hz, and stopping the allocation of computing power to the three-dimensional coordinate data.
[0041] The execution process forms a closed-loop data flow, starting from the phase signal input of the Bluetooth physical layer, through spatial coordinate calculation and clustering model processing with time factor, outputting dynamically adjusted manifold boundary parameters, and finally controlling the power consumption state and sampling accuracy of the system by comparing the inclusion relationship between the real-time position and the dynamic boundary.
[0042] The TV base station continuously receives raw inertial data streams reported in real time by the dual smart rings worn by the user via a Bluetooth 6.0 communication link. This data stream contains 16-bit high-precision binary signals collected by the built-in three-axis accelerometer, three-axis gyroscope, and piezoelectric vibration sensor.
[0043] Specifically, the system first inputs the acquired raw acceleration and angular velocity vectors into a preset Kalman filter module. This module constructs an observation matrix containing six state variables to characterize the instantaneous velocity and displacement changes of the ring in the spatial coordinate system. The state transition matrix F of the Kalman filter module adopts a standard constant velocity motion model, and the matrix form is a 6×6 block diagonal matrix, where the diagonal blocks are in the form of [1T; 01], T is the sampling period (0.02 seconds), the observation matrix H is a 3×6 matrix, the first three columns are the identity matrix I3 (representing the directly observed position), and the last three columns are the zero matrix (representing the velocity that cannot be directly observed), and the process noise covariance matrix Q... The initialization is diag(0.01,0.01,0.01,0.1,0.1,0.1), where the first three diagonal elements correspond to the position dimension and the last three correspond to the velocity dimension; the measurement noise covariance matrix R is initialized to 0.1×I3, indicating a moderate level of confidence in the accelerometer; the initial state vector X0 is set to zero, and the initial estimation error covariance P0 is set to I6; the system executes a Kalman filter prediction-update loop once every sampling period (50 milliseconds) to achieve real-time noise reduction of the inertial data stream, and recursively suppresses the high-frequency noise generated by the sensor physical layer at 500 samples per second through the preset covariance matrix.
[0044] Furthermore, the system independently extracts the vibration signal from the raw data stream and inputs it into a bandpass filter with a sampling frequency of 1000 Hz. The cutoff frequency of the bandpass filter is set between 20 Hz and 250 Hz to preserve the characteristic frequencies generated by finger tapping or micro-vibration and filter out low-frequency shaking interference generated by human walking.
[0045] Specifically, the filtered vibration signal is sent to the feature mapping unit, which calculates the root mean square value and peak factor of the signal within a 20-millisecond time window, and transforms the continuous vibration waveform into feature scalars representing contact intensity and touch rhythm, providing auxiliary dimensions for subsequent behavior type determination.
[0046] Furthermore, the acceleration data after noise reduction is fed into the gravity compensation calculation unit. This unit uses the real-time angular velocity data output by the three-axis gyroscope to perform quaternion attitude calculation. Through the real-time updated quaternion matrix, the acceleration vector of 9.8 meters per second squared caused by the Earth's gravity is accurately removed from the sensor's composite vector, thereby separating the pure motion acceleration component generated by the actual finger movement. Specifically, the TV base station uses the 48-bit unique identification code ID in the data packet reported by the dual smart rings, combined with the signal arrival phase difference characteristics measured by the Bluetooth 6.0 antenna array, to extract the complex domain I / Q sampling sequence generated by the four equally spaced antenna elements when receiving the same pulse signal, and calculates the phase offset between adjacent antenna elements due to the propagation path difference. This phase offset is then converted into an angle mapping to obtain the spatial azimuth angle corresponding to the identification code ID.
[0047] Furthermore, to ensure the stability of gravity compensation accuracy during long-term use, the system introduces a gyroscope zero-bias adaptive recalibration mechanism. Specifically, the system monitors the measurements of the three-axis accelerometer and the three-axis gyroscope in real time. When the system detects that the user has entered a "stationary state" (defined as: average three-axis acceleration < 0.1g, average three-axis angular velocity < 2° / s, duration > 2 seconds), it automatically triggers the zero-bias recalibration program. During the recalibration process, the system calculates the mean vector of the acceleration measurements during the stationary period. Since this vector is only affected by Earth's gravity, it should always point downwards, and its magnitude should be 9.8 m / s². The system uses this calculated gravity vector as a reference for the "current true gravity" and compares it with the gravity direction estimated by the gyroscope to calculate the gyroscope's zero-bias drift and perform compensation. Through this adaptive recalibration, the system can ensure that the gyroscope's zero-bias drift is corrected in a timely manner.
[0048] Furthermore, the system utilizes a feature classification network with a three-layer structure to perform left-hand or right-hand attribution determination. The first input layer of the model receives an eight-dimensional feature vector consisting of a 48-bit identification code, three sets of phase difference values, and a received signal strength indication. The second hidden layer contains 32 neurons and uses a hyperbolic tangent activation function for non-linear association. The third output layer contains two neurons and uses a Softmax function to calculate the probability score of whether the ID belongs to the left or right hand.
[0049] Specifically, the algorithm logic compares the relative position of the real-time measured azimuth angle value with the normal of the TV center. It accurately matches the real-time data stream corresponding to the identification code with an azimuth angle in the range of -45 degrees to 0 degrees to the preset left-hand motion model, and maps the identification code with an azimuth angle in the range of 0 degrees to +45 degrees to the right-hand motion model. The verification range of the phase difference feature is limited to a tolerance range of ±15 degrees to ensure the identification stability under multipath interference environment.
[0050] Furthermore, in order to ensure the closed-loop nature of data flow, the judgment result is written into the system's task scheduling register in real time, and the acceleration, angular velocity and vibration signal stream sampled 500 times per second under the corresponding identification code is distributed to the independent feature parsing buffer of the left or right hand as the structured data source for the input layer of the subsequent gesture recognition model. Specifically, the system compares the consistency of the phase difference in each frame signal processing cycle. If the phase difference fluctuation measured in five consecutive sampling cycles exceeds 15 degrees, the system starts the self-healing recalibration procedure, which verifies the identity by retrieving the 48-bit unique identification code ID and refreshes the spatial coordinate mapping relationship, thereby ensuring that the physical properties of the left and right hand models remain correct during the dynamic interaction process. Furthermore, the system constructs a multi-dimensional input LSTM long short-term memory network model, which includes one input layer, two hidden layers and one output layer. The input layer is responsible for receiving a fusion vector composed of motion acceleration, angular velocity and vibration feature scalars. The first hidden layer is configured with 128 neurons and the second hidden layer is configured with 64 neurons. The Adam optimizer is used to dynamically adjust the weight gradient.
[0051] Specifically, the model uses a forget gate mechanism inside the hidden layer to perform feature fusion on time-series data within a 500-millisecond step, cross-compares the momentum features of pure motion acceleration with the touch features of vibration signals, calculates the probability distribution value of whether the current action belongs to a click, swipe, or zoom behavior, and generates the final behavior type recognition instruction through Softmax processing in the output layer.
[0052] Furthermore, to adapt to gestures of varying lengths, the system employs an adaptive time window mechanism instead of a fixed 500-millisecond window. Specifically, the system monitors the second derivative of the acceleration signal (jerk, in m / s³) in real time to identify the "gesture start point" and "gesture end point." A gesture is considered to begin when the absolute value of the jerk exceeds the threshold of 1.0 m / s³, and to end when the jerk returns to a baseline value of <0.2 m / s³ and remains above 200 milliseconds. The system uses the start point as a time anchor and collects variable-length data segments ranging from 150 to 350 milliseconds, dynamically adjusting based on the actual gesture trigger duration. For short gestures such as micro-movements, 150 milliseconds are collected to ensure complete capture of the 20-30 millisecond pulse characteristics; for long gestures (such as a forward punch), 350 milliseconds are collected to ensure complete capture of the acceleration and deceleration process. These variable-length sequences are resampled to a unified 224 time points using bilinear interpolation and then input into an LSTM for processing, ensuring that both short and long actions are adequately represented.
[0053] The filtering and gravity compensation processing specifically employs adaptive frequency domain separation to dynamically distinguish between "attitude-maintaining low-frequency gravity components" and "intentional high-frequency motion components." After receiving the inertial data stream, the system does not use a fixed cutoff frequency for filtering. Instead, it calculates the rate of change of variance of the acceleration vector in real time. When the variance is detected to be low (stationary or slightly moving), the filter uses an extremely low cutoff frequency (e.g., 0.5Hz) to strongly suppress noise and accurately lock the direction of the gravity vector for attitude calibration. When a sudden increase in variance is detected (action occurs), the system instantly switches to high-pass mode and removes the currently estimated gravity component from the total acceleration through "spectral subtraction". Furthermore, the system performs "windowed derivative" calculation on the separated motion components, retaining only signal segments with steep rising edges as valid gesture candidates and filtering out slow stretching or picking-up actions. This solves the contradiction between "fast response" and "static stability" in traditional fixed filtering, achieving zero-delay capture of the user's subtle gesture intentions. By extracting the transient characteristics of the signal, it distinguishes between "unconscious limb adjustments" and "conscious control gestures" from the physical level, significantly reducing the accidental touch rate in daily life scenarios. It realizes a closed-loop transformation from the voltage value of the ring's underlying sensor to high-order semantic behavior. The input is a multi-modal raw signal containing acceleration, angular velocity and vibration. After Kalman filtering, bandpass filtering, gravity compensation and nonlinear mapping of LSTM network, the final output is the logical event to control the TV interactive interface, ensuring the real-time and accuracy of the interactive action.
[0054] The system extracts multi-dimensional key features in real time from the pre-processed inertial data stream, including the vibration frequency sequence captured by the piezoelectric sensor in the range of 100 Hz to 500 Hz, the acceleration waveform with amplitude within ±8 gravitational acceleration units output by the triaxial accelerometer, and the time difference between the two rings obtained by comparing the absolute timestamps of the signal peaks generated by the left and right rings.
[0055] Furthermore, the system's four antenna arrays are arranged linearly with equal spacing, and the antenna spacing d is set to λ / 2, where λ is the Bluetooth operating wavelength, approximately 12 cm (based on a 2.4 GHz frequency). The base station's antenna array receives Bluetooth feedback signals from the dual smart rings, and the received signal of each antenna channel... In the complex field, it is represented as: ; Where A represents the amplitude of the signal. It is a cosine function. Let ω be the angular frequency of the signal, and t be time. Let n be the reference phase for the nth antenna. The arrival phase of the target signal; The system calculates the phase difference between adjacent antennas (n and n+1): According to array signal processing theory, the relationship between phase difference and azimuth angle is as follows: ; in It is the arcsine function. λ is the phase difference, λ is the Bluetooth operating wavelength, and d is the spacing. By averaging the three sets of phase difference values generated by the four antenna pairs, the system can obtain a robust azimuth angle estimate θ. The system pre-sets left-hand and right-hand angle mapping rules: if θ∈[-45°,0°], the device ID corresponds to the left hand; if θ∈[0°,45°], it corresponds to the right hand. The system allows short-term fluctuations in the azimuth angle within a tolerance range of ±15° (to compensate for multipath propagation interference). However, if the azimuth angle fluctuation exceeds 15° within five adjacent sampling periods, a recalibration procedure is triggered, the device ID is re-called for identity verification, and the spatial coordinate mapping relationship is refreshed.
[0056] Cross-correlation function analysis with a sliding time window was used to analyze the time difference characteristics of the two rings, verifying the morphological similarity and temporal synchronization of the left and right hand signals in the energy envelope. When constructing the left and right hand models, not only were the timestamps TL and TR of the action triggering time compared, but the acceleration magnitude sequence of the left ring was convolved with the sequence of the right ring to calculate the cross-correlation coefficient. The system set an extremely narrow effective time window (e.g., 15ms) and searched for the maximum peak of the cross-correlation function within this window. If the peak exceeded a preset morphological similarity threshold, such as 0.85, and the phase delay corresponding to the peak was close to zero, it was determined to be a valid "coordinated action of the two rings," such as a high-five. If both hands moved simultaneously, but the waveform envelopes were not correlated, such as one hand scratching the head and the other slapping the leg, the cross-correlation peak would be significantly reduced and thus eliminated by the model. The coordinated intention of the two hands was verified from the perspective of signal morphological isomorphism, which can accurately distinguish between "coincidental simultaneous movement" and "targeted coordinated gestures." By using correlation analysis, the robustness of the system against asymmetric environmental noise (such as unilateral hand colliding with an object) was greatly improved.
[0057] Furthermore, the system transforms the aforementioned multidimensional features into a temporal feature matrix and inputs it into a preset gesture pattern matching model. This model consists of an input layer, two deep convolutional layers, and a fully connected layer. The first convolutional layer contains 16 5x1 convolutional kernels to capture single-axis waveform features, the second convolutional layer contains 32 3x3 convolutional kernels to extract the spatiotemporal correlation features between the two rings, and the pooling layer uses a 2x2 kernel size for feature downsampling.
[0058] Specifically, the pattern matching algorithm measures similarity by calculating the dynamic time-warped distance between the input real-time data stream and the standard templates in the gesture dictionary, denoted as . The specific processing involves constructing a 100x100 Euclidean distance-cost matrix, using a cumulative loss path algorithm to find the minimum cost path, and determining a successful match when the cumulative loss value is below a preset threshold of 150 units. Specifically, the pattern matching algorithm employs a dynamic time warping method, calculated using the following formula: ; in , Let the frequencies of the two gestures at the i-th sampling point be denoted as . , Let be the acceleration at the i-th sampling point.
[0059] In the offline phase, 5000 standard gesture samples were collected from 100 volunteers in different scenarios, and a standard template for the feature sequence of each gesture class was calculated. During real-time recognition, the system calculates the DTW distance between the input real-time data stream and the four standard templates. When the distance is less than 150 units and the Softmax confidence of the corresponding class is greater than 0.8, it is considered a successful match for that gesture class. The distance threshold of 150 is determined based on ROC curve analysis of the validation set, and this threshold can be adjusted within the range of 0-200 according to the cost of false recognition and false negative recognition in actual applications.
[0060] Instead of simple rule mapping, a neural network is used to determine left-hand and right-hand ownership. This is because single azimuth information can fluctuate by ±15 degrees in a multipath propagation environment (indoor reflected signals). The system significantly reduces the impact of a single feature by fusing multiple feature dimensions—48-bit unique ID, 4 sets of phase differences, radio signal strength indicator (RSSI), and signal arrival time difference—for joint discrimination. The nonlinear mapping capability of the neural network makes it superior to linear decision rules when dealing with complex associations of multidimensional features, thereby improving the accuracy of left-hand and right-hand identification from pure rule to multidimensional judgment.
[0061] Specifically, the gesture dictionary maintained by the system is a predefined database containing the following four categories of standard gestures and their characteristic parameter ranges: Mutual slap gestures: trigger duration 20-50 milliseconds, vibration frequency center 150-250 Hz, hand arrival time difference 0-10 milliseconds, acceleration waveform peak 3-8 units of gravitational acceleration, waveform characteristics of two separate Gaussian pulses (corresponding to the two contact points of the left and right hands); Object-touching gestures: trigger duration 50-150 milliseconds, vibration frequency initially 150-200 Hz, later decaying to below 50 Hz, exhibiting damped exponential decay characteristics, damping coefficient (frequency decay rate) 0.8-0.95, acceleration waveform showing a single peak. Or a bimodal pattern, with a peak value of 2-6 gravitational acceleration units; micro-motion gestures: trigger duration 10-25 milliseconds, frequency concentrated at 300-450 Hz (the high-frequency characteristic unique to snapping fingers), acceleration peak value less than 0.5 gravitational acceleration units (distinguishing it from the larger amplitude of other actions), energy pulse width 5-15 milliseconds; air gestures: trigger duration 100-250 milliseconds, acceleration waveform shows a clear three-stage pattern of rising-maintaining-falling (distinguishing it from the rapid pulse of collision-type actions), acceleration at the highest point is 5-8 gravitational acceleration units, and the acceleration rapidly decays to the baseline value during the braking phase (indicating that the gesture stops actively rather than being forced to stop due to a collision).
[0062] Furthermore, the system recognizes the mutual clapping gestures in the gesture dictionary. The action is established by verifying the synchronous pulse waveform with a time difference of less than 10 milliseconds and a vibration frequency center located at 150 Hz. The action is defined as follows: a single clapping of the hands corresponds to an isolated wave with a peak height exceeding 3 units of gravitational acceleration in the acceleration waveform, while a double clapping of the hands corresponds to two consecutive peaks with a time interval of 150 to 300 milliseconds within a 500-millisecond period.
[0063] Specifically, the recognition of touch gestures is based on the extraction of the decay characteristics of the tail of the acceleration waveform. When one or both hands pat the leg, the vibration frequency shows an exponential decay distribution characteristic that rapidly decreases from 200 Hz to 50 Hz due to the energy absorption effect of human tissue. When patting hard objects such as seat armrests, a high-frequency vibration response with a larger amplitude and a duration of only 50 milliseconds is generated.
[0064] Furthermore, for the snapping gesture in micro-motion gestures, the system makes a judgment by detecting the instantaneous energy burst point of the vibration signal after high-pass filtering in the frequency band of 300 Hz to 450 Hz. This action feature does not produce a displacement component of more than 0.5 units of gravitational acceleration on the acceleration waveform, but its energy weight in a specific frequency band accounts for more than 70% of the judgment ratio in the pattern matching cost matrix.
[0065] Specifically, aerial gestures such as a forward punch with one or both hands are characterized by an acceleration waveform that rapidly climbs from a baseline value to 5 gravitational acceleration units within 200 milliseconds, accompanied by a braking peak with an opposite phase. The instantaneous fist clenching action is identified by recognizing a weak tremor signal with a frequency of 40 Hz generated by forearm muscle contraction, combined with the temporal domain feature of the angular velocity vector dropping sharply from 200 degrees per second to 0 degrees within 100 milliseconds.
[0066] The entire pattern matching process forms a closed loop of data flow. The input is a 16-bit binary vector containing vibration, acceleration and absolute timestamp, which is synchronously collected by the two smart rings. After feature space mapping and similarity path search processing by the pattern matching algorithm, a unique gesture type code is finally output from the predefined gesture dictionary and sent to the application control layer of the TV base station to trigger the corresponding interaction command.
[0067] A feature verification model based on the fusion of multi-stream convolutional neural network and long short-term memory network is constructed. The model consists of a spatial motion module and a signal dynamics module that do not interfere with each other. The input end receives the three-dimensional trajectory sequence after coordinate transformation fed back in real time by dual smart rings and the raw high-frequency sampling data stream collected by inertial sensors.
[0068] Furthermore, the spatial motion module employs a 3D convolutional neural network with a three-layer structure to analyze the relative trajectory features of the two smart rings. The first convolutional layer contains 16 3x3x3 convolutional kernels to capture the displacement vector changes in three-dimensional space. The second convolutional layer contains 32 3x3x3 convolutional kernels to extract the relative velocity correlation between the two rings. The third layer is a fully connected layer with 128 neurons, which quantifies the consistency and directionality of their motion by calculating the dot product and cross product of the position vectors of the two rings.
[0069] Specifically, for mutual slapping gestures, the spatial motion module calls the opposing convergence model for verification. This model extracts the instantaneous velocity vectors of the left and right rings in the horizontal and vertical dimensions. When the angle between the two velocity vectors is detected to be in the range of 170 to 190 degrees and the relative displacement decreases at a rate exceeding 0.8 meters per second, the trajectory is determined to conform to the physical characteristics of opposing convergence.
[0070] Furthermore, for touch-related gestures, the spatial motion module uses a parallel motion model in the same direction to verify the trajectory. This model monitors the acceleration components of the two rings in the direction perpendicular to the horizontal plane in real time. When the vertical displacement deviation of the two rings is kept within 5 cm and the correlation coefficient between the motion direction vectors exceeds 0.95, it is determined that it conforms to the parallel motion trajectory characteristics of patting an object with both hands or one hand.
[0071] Specifically, for aerial gestures and micro-motion gestures, the spatial motion module calls the hovering braking model to perform trajectory convergence analysis. This model calculates the displacement amplitude in real time through an energy detector with a sliding window of 50 milliseconds. When the change in three-dimensional coordinates rapidly decays to less than 0.5 centimeters within 100 milliseconds, or when the instantaneous velocity at a specific coordinate point drops to less than 0.05 meters per second, the action is determined to meet the braking characteristics at the end of the interaction.
[0072] Furthermore, the signal dynamics module analyzes the frequency and time domain characteristics of inertial signals through a parallel one-dimensional convolutional neural network branch. This branch contains one convolutional layer with a kernel size of 5 and one LSTM layer with 64 hidden units. The input is the original acceleration and vibration sequence with a sampling rate of 500 Hz, and the output is the feature probability distribution representing the signal dynamics properties.
[0073] Specifically, for mutual slapping gestures, the signal dynamics module extracts the peak factor of the acquired waveform. When the signal rise steepness exceeds 50 units of gravitational acceleration per millisecond and is accompanied by high-frequency rigid impact characteristics with a center frequency of 200 Hz to 400 Hz, it is determined to be a hard contact signal generated by physical collision.
[0074] Furthermore, for touch gestures, the module identifies the damping coefficient of the vibration envelope and calculates the decay rate of the signal energy using the recursive least squares method. When the energy exhibits an exponential decrease within a time window of 50 to 80 milliseconds and is accompanied by damped vibration characteristics with a frequency below 100 Hz, the action is determined to be consistent with the touch response of human tissue or furniture surface.
[0075] Specifically, for air gestures, the signal dynamics module matches an active emergency stop feature. This feature is manifested in the time domain waveform as the acceleration vector rapidly drops back to the reference value after reaching its peak, and there are no secondary oscillation ripples caused by physical impact at the end of the waveform. The energy proportion of its high-frequency components is less than 5%.
[0076] Furthermore, for micro-motion gestures, the module matches the single pulse characteristics under static background conditions. That is, while maintaining the three-axis angular velocity fluctuation of less than 2 degrees per second, it identifies local vibration energy blocks with a pulse width of 10 to 25 milliseconds generated by rapid friction of finger muscles. The power spectral density of this energy block in the 450 Hz frequency band exceeds the background noise by 20 dB.
[0077] The mechanism for the collaborative work of the spatial motion module and the signal dynamics module is that both use the Softmax classifier to perform parallel analysis on four types of gestures: mutual shooting, touching objects, micro-movement, and air gestures, and output the corresponding probability distributions respectively. In the core verification logic, the system executes a strict double verification strategy: first, the gesture category with the highest probability in the two modules is extracted. Only when the categories judged by the two modules are completely consistent and their respective confidence levels reach or exceed 85% will the gesture be finally confirmed as a valid command and the corresponding code will be output. If the above-mentioned category consistency and high confidence conditions cannot be met at the same time, the system will determine the current data frame as a mis-touch or low confidence event and immediately trigger the cleanup protection mechanism. This mechanism will mark all feature data within the current 2000-millisecond sliding window as invalid, clear the feature buffer in memory, record the mis-touch log for offline analysis, and remain in standby state to continue capturing new gesture signals.
[0078] The classification labels of the spatial motion module and the signal dynamics module are finally determined by a logic AND gate circuit. Only when the two modules output the same Softmax classification index for the gesture type within the same time window and their respective confidence scores exceed 0.85 can the system establish the gesture as a valid command and output the final interactive control signal to the TV base station. Otherwise, the data stream is judged as a mis-touch or environmental noise and the data is cleared.
[0079] The TV module pre-establishes a mapping logic table at the system application layer that maps 8-bit gesture type codes to specific function operation codes. This table is stored in the TV's non-volatile flash memory space in the form of a static constant matrix.
[0080] The mapping logic table between gesture commands and system functions was obtained by collecting 10,000 sets of user gesture behavior samples in an offline environment and matching them with high-frequency control intentions in a TV scene. A 128-dimensional gesture feature vector, including the amplitude and frequency of the movement, was extracted and input into a multilayer perceptron model with a 3-layer structure and 512 neurons in the hidden layer. The Softmax function was used to calculate the probability distribution of each gesture feature corresponding to specific functions such as pause, play, or volume adjustment. For each function intention, the system selected a set of gesture features with a classification confidence of over 95% as a standard template and assigned it a unique 8-digit hexadecimal identifier ID. A hash algorithm was used to associate these gesture IDs with the low-level API function entry points of the TV operating system, and specific parameter step values such as 5% volume adjustment or 5000 millisecond progress jump were defined to form a static logical index table.
[0081] Furthermore, the index table is stored in the non-volatile memory of the TV module in the form of a binary file, realizing a closed-loop process of input gesture encoding, retrieval in the cache by address offset, and output as system executable instructions.
[0082] Specifically, regarding playback control feedback, when the function operation code output by the protocol parsing module is determined to be a control instruction, the TV module executes the specific operation by calling the application programming interface in the underlying multimedia framework. For example, after receiving the pause operation code, it immediately sends a suspend signal to the system playback thread to make the decoder stop outputting the H.265 video stream.
[0083] Furthermore, for fast forward or rewind operations, the TV module obtains the current system clock reference value and performs a timestamp offset calculation with a step value of 5000 milliseconds on it. The calculated new timestamp is then reloaded into the player's demultiplexing unit as a jump parameter to achieve precise progress adjustment.
[0084] Specifically, for volume adjustment commands, the system uses a linear step algorithm to correct the current global audio gain variable. This is achieved by accumulating the current volume value with the gain weight carried by the command, and using a smoothing filter with a 5-level moving average coefficient to ensure that the volume change rate is within 10% per second.
[0085] Furthermore, in response to program interaction feedback, the TV module independently activates a hardware-accelerated rendering layer with alpha channel transparency attribute above the main video rendering layer. This layer uses TextureView technology to achieve synchronous visual mapping of real-time gesture actions.
[0086] Specifically, when the system receives a gesture command with a specific intensity factor sent by the base station, the graphics engine in the transparent rendering layer calls the preset particle emitter algorithm, inputs the intensity factor as an independent variable into the emission frequency calculation formula, and generates a dynamic special effects particle cluster at the specified coordinates of the video screen whose coverage area expands as the score increases by calculating the linear proportional coefficient between the score and the emission density.
[0087] Furthermore, the logic driving the changes of virtual props obtains the motion trajectory curvature data output by the gesture recognition module, and maps the rate of change of curvature to the scaling factor or geometric vertex offset of the 3D virtual prop. For example, when a momentary fist-clenching gesture is detected, the system calculates a deformation coefficient between 0.5 and 1.5 based on the peak acceleration and applies it to the model matrix of the virtual prop.
[0088] Specifically, the entire interactive feedback process realizes a closed-loop data flow from HID instruction parsing to graphics rendering output. The input is a binary message sent by the base station containing gesture type and participation score. After routing and distribution by the mapping logic table and processing by the system API or graphics engine, the final output is a change in playback status or interactive visual elements that change in real time on the transparent rendering layer.
[0089] By directly mapping playback control through the underlying interface, millisecond-level response and system stability are ensured for basic operations such as pause and fast forward. At the same time, dynamic effects are superimposed on an independent transparent rendering layer, achieving visual decoupling between interactive feedback and video content, enhancing the immersiveness and fun of watching movies without interfering with the playback.
[0090] Furthermore, upon receiving system broadcast control instructions, the TV module establishes a correspondence between the playback timeline and content metadata by parsing the metadata description file carried in the video stream header. It also uses a hash index structure to associate the millisecond-level timestamp of video playback with a specific advertising material ID. During system initialization, the decoder directly extracts advertising segment information from the video file's metadata, establishing a correspondence between time ranges and advertising content. After entering real-time playback, the system quickly locates the advertising content displayed on the screen in the mapping table based on the display time of each frame obtained from the decoder's clock. Specifically, when the TV module receives a program interaction instruction sent by the base station, the real-time synchronizer inside the system immediately captures the current decoding time position of the media player, such as 360500 milliseconds, and uses this as an index value to retrieve the corresponding advertising material identifier in the mapping table in real time.
[0091] Furthermore, the system invokes a log generation model based on feature embedding logic to perform closed-loop processing on the captured data. This model contains an input layer with an input dimension of 4, which receives timestamps, ad creative IDs, gesture type codes, and interaction engagement scores. Specifically, the model internally embeds a weight calculation algorithm, which calculates the active attention weight of the interaction behavior by multiplying the interaction engagement score by a preset gain coefficient of 0.85 and linearly weighting it in combination with the duration of the gesture action.
[0092] Furthermore, the system encapsulates the above calculation results into an active attention log. The data structure of this log adopts a standardized binary format, including a start code, a 16-bit device serial number, a 32-bit advertising material identifier, an 8-bit attention weight value, and a 16-bit cyclic redundancy check code. Specifically, when the calculated attention weight value exceeds the system's preset 80% threshold, the logic judgment unit will trigger a high-value event marking instruction, changing the 17th status bit in the log file from 0 to 1, to indicate that the user has generated extremely high intensity brand interaction behavior.
[0093] Furthermore, the network layer protocol stack of the TV module stores the generated active attention logs in the sending buffer and calls the secure transmission protocol to establish an asynchronous connection with the remote cloud advertising server via 5G or wireless LAN. Specifically, the network module uploads encrypted log packets from the cache to the cloud server's API interface in real time at a frequency of 1 Hz. The output is a JSON dataset containing user feedback behavior characteristics, thus providing the cloud with the original physical layer evidence for verifying the true reach of advertising.
[0094] To overcome end-to-end signal latency caused by wireless transmission, the system introduces a predictive alignment mechanism. This mechanism automatically compensates for the playback timestamp used for querying by calculating the time difference between the command transmission time and the TV reception time, ensuring that gesture commands accurately correspond to the actual frame seen by the user. Furthermore, the system has intelligent perception capabilities for user playback control behavior: when operations such as fast forward or rewind that cause playback position jumps are detected, the captured gesture commands are marked as invalid and filtered, thus ensuring the rigor of advertising interaction statistics. In the data reporting stage, the system maintains redundant logs containing complete interaction details locally. Before initiating network transmission, the system performs strict data integrity checks, checking the continuity of time records and detecting any sequence loss, to ensure that the data uploaded to the cloud is completely consistent with the local records, effectively preventing data loss or corruption during network transmission.
[0095] Furthermore, to ensure the reliability of ad feedback time alignment, the system employs a three-layer fault tolerance mechanism: (1) Double timestamp verification: The system maintains two independent time bases—the base station clock and the TV player clock. Both timestamps and their time difference are recorded simultaneously during reporting, allowing for lag correction when the cloud aligns.
[0096] (2) Time jump detection: The system monitors the monotonicity of the PTS value. When a forward jump (greater than 1 second) is detected in the PTS, it is determined that the user has performed a fast-forward operation; when the PTS jumps backward, it is determined that the user has rewound. The system marks the interactive events during fast forward / rewound as "invalid" and filters them during statistics to prevent these events from being incorrectly associated with the target advertisement.
[0097] (3) Local log redundancy: The TV maintains a complete local event log, containing all captured interactive events and their timestamps. The log is stored in a circular buffer format, storing a maximum of 48 hours of event records. When uploading over the network, the system verifies the data received from the cloud. If packet loss or out-of-order delivery is detected, retransmission or repair can be performed based on the local log, thereby ensuring the integrity and accuracy of the cloud data.
[0098] It achieves a complete transformation from interactive engagement data output from the base station to cloud feedback logs. The input is the user gesture score obtained through the Bluetooth link. After time-domain retrieval of the TV mapping table and weighted encoding of the log model, a data closed loop is finally formed that can be used for secondary analysis by advertisers.
[0099] A gesture recognition-based interactive viewing system and its storage medium are disclosed. The computer program instructions embedded in the storage medium are executed by a base station processor configured with four equally spaced antenna arrays. The driving system acquires the raw motion information of the two smart rings through an inertial sensor with a sampling frequency of 500 Hz, and uses a convolutional neural network model containing 32 and 64 3x3 convolutional kernels to lock the spatial coordinates of the two rings within a dynamic geometric presence gate with an angle of ±45 degrees and a distance between 0.5 meters and 4.5 meters. Specifically, the base station performs Kalman filtering on the received 16-bit binary data stream with a 6-dimensional state observation matrix, and uses a quaternion attitude matrix to remove the gravitational acceleration component of 9.8 meters per second squared. Combined with a 48-bit unique identification code ID and a phase difference feature of ±15 degrees, a left and right hand behavior model is constructed. Furthermore, the processor inputs the processed acceleration sequence and the vibration signal filtered by a bandpass filter from 20 Hz to 250 Hz into an LSTM long short-term memory network containing 128 neurons and a dropout rate of 0.2. The processor calculates the dynamic time warping similarity of less than 150 distance units between the real-time features and the predefined gesture templates using a pattern matching algorithm. This is used to identify mutual shooting actions that conform to the mutual shooting trajectory, touching actions that conform to the damped vibration characteristics, aerial actions that conform to the active emergency stop characteristics, and micro-movement actions that conform to the single pulse characteristics. Specifically, the system uses an interactive scoring model to weight and fuse normalized intensity factors from 0 to 100 with frequency factors within a 10,000-millisecond sliding window at a ratio of 0.6 to 0.4, generating an interactive engagement measurement data package that conforms to the USBHID human interface standard. The TV module associates the measurement score with the advertising material ID by retrieving the mapping table between the playback timeline and content metadata. When the active attention weight exceeds 80%, a high-value event is triggered, and the generated encrypted log is asynchronously uploaded to the cloud server through the network layer. This achieves a complete technical closed loop from storage medium instruction calls to physical layer feature capture and application layer value feedback.
[0100] This invention's system uses a base station antenna array in conjunction with dual rings to perform high-precision angle of arrival measurement, ensuring the spatial perception accuracy of the "presence gate." It also employs a storage-computation separation architecture to offload complex calculations to the base station, significantly reducing ring power consumption and size while maintaining real-time recognition. The software-encapsulated algorithm, implemented with storage media, not only facilitates cross-platform deployment but also supports flexible expansion of the gesture dictionary and optimization model through software updates, allowing users to continuously obtain functional upgrades and experience optimizations without replacing hardware.
[0101] This embodiment constructs a virtual presence gate through high-precision spatial perception, activating the system only when a user wearing both rings enters the valid viewing area, effectively blocking accidental touches outside the viewing area. The base station performs filtering and gravity compensation on the ring's inertial data, removing the gravity component to obtain pure motion data. The system utilizes a dual-ring collaborative model to deeply analyze vibration frequency, waveform, and time difference characteristics to accurately match gestures and quantify them as an "interaction participation metric." The base station sends instructions to the television through a standard interface, which is then parsed and converted into playback control, program effects, or advertising feedback events, achieving a closed-loop interaction across the entire chain from spatial perception and motion capture to content response.
[0102] All neural network models in the system (convolutional neural network, LSTM, multilayer perceptron) were trained offline on a training set containing over 100,000 standard gesture samples. This training set was obtained from 100 volunteers and covered the following multi-dimensional scenarios to ensure generalization ability: (1) User diversity: aged 18-65, weighing 50-95kg, with a gender ratio of 1:1, including users who have or do not have movie-watching habits; (2) Environmental diversity: 3 common living room layouts, 2 lighting conditions (natural light, artificial light), 5 mainstream TV brands; (3) Diverse clothing: short sleeves, long sleeves, objects around the arms (pillows, remote controls, etc.); (4) Distance diversity: viewing positions at 0.5m, 1.5m, 2.5m, 3.5m, and 4.5m, with 1000 hand gestures collected at each distance; (5) Furniture diversity: 4 types of sofa materials (wood, fabric, leather, and mixed), 2 types of armrest heights; Each volunteer performed at least 50 repetitions of the standard gesture in each scenario, generating a total of 100 × 50 × 20 (number of scenarios) = 100,000 original samples. In addition, the system applied a 3x data augmentation to the training data: random rotation (±10 degrees), noise injection (Gaussian noise σ=0.05), and time stretching (±10%), ultimately forming 300,000 augmented training samples. Training employed a 5-fold cross-validation scheme, training the model on 4 folds of training data and evaluating generalization performance on 5 folds of validation data. Finally, the model performance was validated on a completely independent external test set (5,000 gesture samples from 30 volunteers who did not participate in training), achieving an accuracy of 95.3% and a false touch rate of 3.1%, validating the model's good generalization ability.
[0103] Example 2 To address the complex issue of accidental interaction triggers caused by unconscious physical actions of users in movie-watching scenarios, such as adjusting glasses, scratching, or tidying hair, the system builds and runs a spatiotemporal decoupling filter and a deep intent detection module based on the Transformer architecture. The system synchronously acquires multi-source data streams of the two smart rings within a 2000-millisecond sliding time window via Bluetooth 6.0 communication link and inertial sensor. The data streams include triaxial acceleration and angular velocity signals sampled 500 times per second, as well as a three-dimensional spatial coordinate sequence sampled 50 times per second by the antenna array. These heterogeneous data are spliced into a spatiotemporal tensor with a feature dimension of 128 after resampling. Furthermore, the system inputs the spatiotemporal tensor into a pre-defined 4-head parallel attention encoder. The encoder first maps the feature dimension to a 64-dimensional embedding space through a linear transformation layer. Then, the four independent attention heads perform dot product operations in the spatial and temporal dimensions, respectively. The correlation significance of the motion trajectory is identified by calculating the autocorrelation weight of each sampling point in the feature matrix relative to the global window. The encoder contains a feedforward neural network with 256 hidden neurons and uses layer normalization to suppress non-stationary noise caused by wireless signal jitter. Specifically, the weighted feature vector output by the encoder is fed into the trajectory curvature verification model to perform white-box logic analysis. This model extracts the instantaneous radius of curvature and acceleration vector direction of the motion trajectory by performing second derivative operations on the three-dimensional coordinate sequence within a continuous step of 100 milliseconds. If the algorithm detects that the motion vector of the ring converges towards the three-dimensional center region of the user's head pre-calibrated by the base station, and the magnitude of its displacement vector exhibits an exponential decay with a slope exceeding 0.5 when approaching the endpoint coordinates, that is, when the end velocity decreases by more than 30% compared to the peak, the system logic determines that the action is a typical parasitic non-interactive action such as adjusting glasses or scratching.
[0104] Furthermore, the system projects the motion vector after feature decoupling onto the intent detection space, and uses the vector cosine similarity algorithm to calculate the angle between the main direction vector of the current action and the normal vector of the center of the TV screen, wherein the normal vector is preset by the physical installation angle of the base station antenna array; Specifically, the physical conditions for the system to determine the validity of the interaction intent are strictly limited to a dual-verification range. On the one hand, the angle between the main direction vector of the action and the screen normal is required to be in the quasi-vertical range of 85 to 95 degrees. On the other hand, the power spectral density of the acceleration waveform is extracted using fast Fourier transform, and the energy ratio in the effective interaction frequency band of 5 Hz to 15 Hz is analyzed by bandpass filtering algorithm. Only when the energy weight in this frequency band exceeds 75% of the total energy and the above angle condition is met simultaneously can the system output a high-level intent enable signal to the subsequent gesture recognition engine.
[0105] Furthermore, if the judgment result of any of the above dimensions does not conform to the preset distribution of the intent feature matrix, the system immediately determines that the current action is environmental interference or unconscious limb behavior. At this time, the logic gate circuit will block the parsing process of the gesture command and perform immediate physical erasure of all cached features within the current 2000-millisecond window, thereby suppressing the false trigger rate of interaction in complex dynamic backgrounds to below 3%. Specifically, the entire false trigger suppression process realizes a closed-loop flow from the original multimodal sensor data to semantic intent judgment. The input is high-frequency sampled spatial coordinates and inertial vectors. After the spatiotemporal feature extraction of the Transformer encoder, the second-order dynamic analysis of the trajectory curvature, and the directional verification based on the screen normal, the final output is a binary intent establishment state, ensuring the robustness of the movie-watching interaction system in complex life scenarios.
[0106] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A gesture recognition based viewing interaction method, characterized in that, include: A dynamic geometric region is set as a presence gate based on preset angle and distance thresholds. When a user wearing two rings enters the presence gate, the user's presence status is established. A hysteresis comparator mechanism and geomagnetic fingerprint verification are added to the determination logic of the presence gate to construct a double-layer buffer region containing entry and exit thresholds. The base station continuously receives and analyzes the inertial data stream and wireless signal characteristics reported by the ring, and performs filtering and gravity compensation processing on the inertial data stream. Construct left and right hand models to distinguish between left and right hands based on ring ID and wireless signal characteristics; based on the processed inertial data stream, analyze vibration frequency, acceleration waveform, and time difference characteristics of the two rings in real time to perform gesture recognition, and match a gesture action from a predefined gesture dictionary; A feature verification model is constructed to verify the matched gestures. The feature verification model includes a spatial motion module and a signal dynamics module. The spatial motion module is used to analyze the relative trajectory characteristics of the two rings: for mutual patting gestures, it verifies whether they conform to the mutual convergence model; for touching gestures, it verifies whether they conform to the parallel motion model in the same direction; for aerial gestures and micro-motion gestures, it verifies whether they conform to the hovering braking model. The signal dynamics module is used to analyze the frequency and time domain characteristics of inertial signals: for mutual shooting gestures, it matches the characteristics of high-frequency rigid impact; For touch-based gestures, damped vibration characteristics with energy decay are matched; for air gestures, active emergency stop characteristics without collision are matched; for micro-motion gestures, single pulse characteristics under static background are matched. The mechanism for the collaborative work of the spatial motion module and the signal dynamics module is that both use the Softmax classifier to perform parallel analysis on four types of gestures: mutual shooting, touching objects, micro-movement, and air gestures, and output the corresponding probability distributions respectively. In the core verification logic, the system executes a strict double verification strategy: first, the gesture category with the highest probability in the two modules is extracted. Only when the categories judged by the two modules are completely consistent and their respective confidence levels reach or exceed 85% will the gesture be finally confirmed as a valid command and the corresponding code will be output. After verification, the gesture is mapped to a measure of the execution command, and the base station sends the command to the TV via USBHID or virtual serial port; The TV module parses the instructions and converts them into interactive feedback, including but not limited to playback control, program interaction, and advertising feedback event reporting; the TV module maintains a mapping table between playback timeline and content metadata; when it receives an interactive instruction from the base station, the system captures the timestamp of the current playback screen and the corresponding advertising material ID; The interaction engagement metric is associated with the ad creative ID to generate an active attention log; When the interaction engagement measurement exceeds a preset threshold, an event marker is triggered, and the TV terminal uploads the active attention log to the cloud advertising server in real time through the network layer. In order to ensure the reliability of the advertising feedback time alignment, the system adopts a three-layer fault tolerance mechanism: double timestamp verification, time jump detection, and local log redundancy.
2. The movie-watching interaction method based on gesture recognition according to claim 1, characterized in that, The on-site gate is defined by a preset angle threshold and a distance threshold. The spatial position is calculated based on Bluetooth 6.0 direction measurement and distance measurement, and is obtained in real time through wireless signal interaction between the user's dual smart rings and the TV base station. The presence gate is a dynamic geometric area, which is a fan-shaped / rectangular area in front of the TV base station. The angle threshold can be set to the viewing angle range of the TV screen, and the distance threshold can be set to the viewing distance in the living room. When the user's spatial position meets these threshold conditions, the user's presence status is established as valid viewing, and the subsequent gesture recognition and interaction process is activated. If the location exceeds the threshold, it is determined to be in an off-site state, and data collection is stopped to ensure privacy protection and data validity.
3. The movie-watching interaction method based on gesture recognition according to claim 2, characterized in that, The determination of the presence gate also includes an adaptive manifold region with an irregular shape: During system initialization, a preset generalized sector area is set as the initial gate. As user interaction data is generated, the system records the spatial coordinates of each gesture that is determined to be valid. The coordinates are processed using a clustering algorithm with a time decay factor to dynamically adjust the gate boundary, so that the boundary gradually converges and fits the user's actual viewing activity range. When the user's coordinates are detected to be within the manifold area, the high-frequency gesture sampling mode is activated. If the user's coordinates are outside the area, it is determined to be an off-screen or invalid area, and the system enters a low-power standby state.
4. The movie-watching interaction method based on gesture recognition according to claim 1, characterized in that, The steps of filtering and gravity compensation processing the inertial data stream include: The inertial data stream originates from the three-axis accelerometer and gyroscope built into the dual smart ring. The collected data includes acceleration vector, angular velocity, and vibration signal. The data stream is filtered by low-pass filtering to remove noise and environmental interference. Gravity compensation processing is performed, and the accelerometer's attitude is corrected using gyroscope data to separate the pure motion acceleration component. Independent models for the left and right hands are constructed. The spatial azimuth angle is calculated by combining the unique ID of the ring with the phase difference of the wireless signal, and the devices located in different sectors are accurately mapped to the left or right hand model. Based on this, inertial data streams are integrated to analyze the spatial source and specific behavior type of the action in real time.
5. The movie-watching interaction method based on gesture recognition according to claim 1, characterized in that, The specific steps of matching a gesture action from a predefined gesture dictionary include: Key features are extracted from the processed inertial data stream, including vibration frequency, acceleration waveform, and double-ring time difference. The system uses a pattern matching algorithm to compare the key features with a predefined gesture dictionary. The gesture dictionary consists of short, trigger-type actions, including at least: mutual slapping gestures such as slapping each other once or twice; object-touching gestures such as slapping the leg with one or both hands, or slapping the armrest of a seat with one or both hands; micro-movement gestures such as snapping fingers by rubbing the thumb and middle finger; and air gestures such as clenching a fist forward with one or both hands, or clenching a fist instantly when the palm changes from an open state to a closed state.
6. The movie-watching interaction method based on gesture recognition according to claim 1, characterized in that, The conversion of these instructions into playback control and program interaction feedback includes: A mapping logic between gesture commands and system functions is pre-established; for playback control feedback: when the parsed command is a control command, the TV module directly executes playback control operations, including pause, play, fast forward, rewind, or volume adjustment, by calling the underlying player interface of the system; for program interaction feedback: the TV activates a transparent rendering layer to generate dynamic effects and / or drive changes in virtual props in real time based on gesture commands, achieving visual feedback synchronized with the video.
7. A movie-watching interactive system based on gesture recognition, characterized in that, The system includes: At least one ring comprising an inertial sensor and a first wireless transceiver; a base station associated with a television, the base station comprising an antenna array for performing angle of arrival measurements, a second wireless transceiver, a processor, and an interface for communicating with the television; and a memory storing instructions that, when executed by the processor, cause the system to perform the method of claim 1. Preprocessing module: Based on preset angle and distance thresholds, a dynamic geometric region is set as an presence gate. When a user wearing two rings enters the presence gate, the user's presence status is established. The base station continuously receives and analyzes the inertial data stream and wireless signal characteristics reported by the rings, and performs filtering and gravity compensation processing on the inertial data stream. Gesture matching module: Constructs left and right hand models, distinguishes left and right hands based on ring ID and wireless signal characteristics; Based on the processed inertial data stream, analyzes vibration frequency, acceleration waveform, and double ring time difference characteristics in real time to perform gesture recognition, and matches a gesture action from a predefined gesture dictionary; Constructs a feature verification model to verify the matched gesture, and after verification, maps the gesture action to a measure of interaction participation. Parsing Module: The base station sends instructions to the TV via USBHID or a virtual serial port; the TV module parses the instructions and converts the interactive participation into program interaction feedback, which includes, but is not limited to, playback control, program interaction, and advertising feedback event reporting.
8. A gesture-based interactive viewing storage medium, wherein computer program instructions are stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the gesture recognition-based interactive viewing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Hand gesture interaction device
CN106527670A
Smart ring
CN109085885A
Ring pairing and left and right hand judgment method and system based on multiple users and medium
CN121857980A