Three-dimensional motion data processing method and device and computer readable storage medium
Through technical means such as signal preprocessing of user motion data, hierarchical primitive division and multi-scale window segmentation, the problem of inaccurate complex action recognition in the existing technology is solved, efficient and robust action recognition and instruction generation are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510464892.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
When processing complex and diverse user motion data, the prior art is not very accurate and robust, making it difficult to accurately identify complex continuous actions and generate intelligent instructions.
Three-dimensional motion data processing methods are adopted, including signal preprocessing, hierarchical primitive division, multi-scale window segmentation, multi-level feature extraction, primitive and action-level filtering, timing mode analysis and context processing, and multi-modal collaborative processing is carried out in combination with user physiological state data to realize action recognition and instruction mapping.
It significantly improves the accuracy of user's recognition of complex actions and the anti-interference ability of the system, enhances the ability to distinguish similar actions, realizes a seamless transformation from action recognition to application control, and improves the interactive experience and system practicality.
Smart Images

Figure CN120296675A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of somatosensory device data processing, and in particular, to a three-dimensional motion data processing method, device, and computer-readable storage medium. Background Art
[0002] In recent years, with the popularization of wearable devices, using three-dimensional motion data collected by somatosensory devices (such as smart watches and sports bracelets) for human activity recognition and control has become an important research field. These devices can capture the user's movement trajectory and posture changes through built-in sensors such as multi-axis accelerometers and gyroscopes, providing a rich data basis for applications such as sports health monitoring, human-computer interaction, and game control. Current technical solutions usually involve signal preprocessing of the collected raw motion data, and then using pattern recognition algorithms (such as threshold methods, machine learning methods) to map specific motion patterns to corresponding actions or instructions. However, when dealing with complex and diverse user movements, these existing methods often face problems such as low accuracy and weak robustness, and it is difficult to meet the growing application requirements.
[0003] The main problems existing in the prior art include the following aspects. First, due to the randomness and individual differences of user movements, as well as the possible noise interference of the device itself, the collected raw motion data often contains a large amount of redundant information and noise, and simple threshold judgment or direct mapping easily leads to false triggering and recognition errors. Second, for complex continuous actions or combined actions, existing technical solutions often lack effective modeling and recognition mechanisms, and it is difficult to accurately distinguish the subtle differences between different actions. Especially in the aspect of accurately identifying specific patterns and durations of composite actions and generating intelligent filtering and instructions according to the context environment, the prior art still has obvious deficiencies.
[0004] These problems severely limit the interaction efficiency and user experience of three-dimensional motion data in practical applications, and there is an urgent need for a comprehensive processing method that can optimize the entire process from data collection to instruction mapping. Summary of the Invention
[0005] An embodiment of the present application provides a three-dimensional motion data processing method, aiming to improve the accuracy of recognizing complex user actions.
[0006] To achieve the above object, an embodiment of the present application provides a three-dimensional motion data processing method, including:
[0007] Performing signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device to obtain a preprocessed three-dimensional motion data stream;
[0008] Perform hierarchical primitive partitioning and template generation on the preprocessed three-dimensional motion data stream, and establish a motion primitive library containing multi-level primitive features and transition relationships;
[0009] Apply multi-scale window segmentation to the three-dimensional motion data stream, dynamically adjust the window parameters according to the signal characteristics, and obtain a segmented data set with a hierarchical time structure of micro-window, medium-window, and macro-window;
[0010] Perform multi-level feature extraction corresponding to its time structure on the segmented data set, and perform pattern matching between the extracted multi-level features and the templates in the motion primitive library to obtain a preliminary primitive sequence containing primitive information and confidence;
[0011] Perform two-level filtering at the primitive level and action level on the preliminary primitive sequence, and perform primitive-level and action-level verification based on the primitive features and transition relationships in the motion primitive library to obtain a highly reliable action sequence after verification;
[0012] Apply temporal pattern analysis and context processing to the highly reliable action sequence, perform action recognition with reference to the combination rules in the motion primitive library, and integrate multi-source data to obtain a structured action recognition result;
[0013] Perform instruction mapping conversion on the structured action recognition result to obtain an application control instruction corresponding to the recognized action.
[0014] To achieve the above object, an embodiment of the present application also proposes a three-dimensional motion data processing device, including a memory, a processor, and a three-dimensional motion data processing program stored on the memory and executable on the processor. When the processor executes the three-dimensional motion data processing program, it implements the three-dimensional motion data processing method described in any one of the above.
[0015] To achieve the above object, an embodiment of the present application also proposes a computer-readable storage medium, on which a three-dimensional motion data processing program is stored. When the three-dimensional motion data processing program is executed by a processor, it implements the three-dimensional motion data processing method described in any one of the above.
[0016] The 3D motion data processing method of this application significantly improves the signal-to-noise ratio of the original sensor data and effectively eliminates tremor interference and external noise effects by implementing adaptive signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device. At the same time, the multi-level primitive division technology constructs a structured representation system from micro-primitives to composite primitives, enabling the system to accurately decompose complex actions and capture the composition rules of actions through the primitive transition probability matrix, greatly enhancing the ability to distinguish similar actions and the recognition accuracy of complex actions. In addition, the multi-scale window segmentation architecture composed of micro-windows, medium-windows, and macro-windows realizes the parallel capture of action features at different time scales, solving the technical problem that traditional fixed-window methods are difficult to take into account both instantaneous actions and continuous actions at the same time. The multi-level feature extraction based on this architecture corresponds to each time scale, forming a multi-dimensional feature space from the time-domain statistical features of the micro-window to the temporal pattern features of the macro-window, greatly enriching the information dimension of action representation. In particular, the double-level filtering mechanism at the primitive level and the action level effectively filters out misrecognition results through multiple screening means such as constructing adaptive thresholds, time verification, and energy verification, significantly enhancing the anti-interference ability and recognition stability of the system. At the same time, the multi-level sliding observation window system of this method realizes the continuous monitoring of short-term, medium-term, and long-term action sequences, enabling the system to have the ability to capture behavior patterns across time scales. The multi-modal collaborative processing of user physiological state data and the three-level instruction mapping framework not only achieve seamless conversion from action recognition to application control, but also continuously optimize user-specific parameters through personalized learning algorithms, effectively adapting to the action habits and usage preferences of different users, thus greatly improving the overall interaction experience and system usability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0018] Figure 1 It is a module structure diagram of an embodiment of the 3D motion data processing device of the present invention;
[0019] Figure 2 It is a flowchart of an embodiment of the 3D motion data processing method of the present invention.
[0020] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0022] To better understand the above technical solution, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0023] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The term "comprising" appearing in the text does not exclude the presence of components or steps not listed in the claims. The indefinite article "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a properly programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of "first", "second", and "third", etc. does not denote any order and these words can be interpreted as names.
[0024] As Figure 1 shown, Figure 1 is a schematic structural diagram of a server 1 (also called a three-dimensional motion data processing device) in the hardware operating environment involved in the embodiment solution of the present invention.
[0025] The server in the embodiment of the present invention, such as "Internet of Things devices", intelligent air conditioners with networking functions, intelligent lights, intelligent power supplies, AR / VR devices with networking functions, intelligent speakers, autonomous driving vehicles, PCs, smartphones, tablets, e-book readers, portable computers, and other devices with display functions.
[0026] As Figure 1 shown, the server 1 includes: a memory 11, a processor 12, and a network interface 13.
[0027] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the server 1 in some embodiments, such as the hard disk of the server 1. The memory 11 can also be an external storage device of the server 1 in other embodiments, such as a plug-in hard disk equipped on the server 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0028] Furthermore, the memory 11 can also include an internal storage unit and an external storage device of the server 1. The memory 11 can be used not only to store application software installed on the server 1 and various types of data, such as the code of the three-dimensional motion data processing program 10, etc., but also to temporarily store data that has been output or will be output.
[0029] The processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chips in some embodiments, and is used to run the program code stored in the memory 11 or process data, such as executing the three-dimensional motion data processing program 10, etc.
[0030] The network interface 13 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.
[0031] The network can be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN) and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol and / or Bluetooth communication protocol or a combination thereof.
[0032] Optionally, the server may further include a user interface, which may include a display, an input unit such as a keyboard, and optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be referred to as a display screen or a display unit, and is used to display the information processed in the server 1 and to display a visual user interface.
[0033] Figure 1 Only the server 1 with components 11-13 and the three-dimensional motion data processing program 10 is shown. Those skilled in the art can understand that, Figure 1 The shown structure does not constitute a limitation on the server 1. It may include fewer or more components than shown, or combine certain components, or have different component arrangements. In this embodiment, the processor 12 may be used to call the three-dimensional motion data processing program stored in the memory 11 and perform the following operations:
[0034] Perform signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device to obtain a preprocessed three-dimensional motion data stream;
[0035] Perform hierarchical primitive partitioning and template generation on the preprocessed three-dimensional motion data stream to establish a motion primitive library containing multi-level primitive features and transition relationships;
[0036] Apply multi-scale window segmentation to the three-dimensional motion data stream, dynamically adjust the window parameters according to the signal characteristics, and obtain a segmented data set with a hierarchical time structure of micro-windows, medium-windows, and macro-windows;
[0037] Perform multi-level feature extraction corresponding to its time structure on the segmented data set, and perform pattern matching between the extracted multi-level features and the templates in the motion primitive library to obtain a preliminary primitive sequence containing primitive information and confidence;
[0038] Perform double-level filtering at the primitive level and the action level on the preliminary primitive sequence, and perform primitive-level and action-level verification based on the primitive features and transition relationships in the motion primitive library to obtain a highly reliable action sequence after verification;
[0039] Apply temporal pattern analysis and context processing to the highly reliable action sequence, perform action recognition with reference to the combination rules in the motion primitive library, and integrate multi-source data to obtain a structured action recognition result;
[0040] Perform an instruction mapping conversion on the structured action recognition result to obtain an application control instruction corresponding to the recognized action.
[0041] Based on the hardware architecture of the above three-dimensional motion data processing device, an embodiment of the three-dimensional motion data processing method of the present invention is proposed. Refer to Figure 2 , Figure 2 This is an embodiment of the three-dimensional motion data processing method of the present invention. The three-dimensional motion data processing method includes the following steps:
[0042] S10. Perform signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device to obtain a preprocessed three-dimensional motion data stream. The purpose of this step is to perform preliminary cleaning, fusion, and feature extraction on the original sensor data obtained from the somatosensory device. In some embodiments, step S10 can be implemented through steps S11-S14:
[0043] In step S11, apply an adaptive low-pass filter to the multi-axis acceleration data and angular velocity data to obtain filtered sensor data. An adaptive low-pass filter is a filter that can dynamically adjust the cut-off frequency according to the characteristics of the input signal. Its core lies in being able to retain the effective signal in the data while filtering out high-frequency noise. In practical applications, a Butterworth low-pass filter can be used in combination with signal energy analysis for implementation. For example, for acceleration sensor data with a sampling frequency of 100 Hz, a low-pass filter with an initial cut-off frequency of 10 Hz can be set. However, when it is detected that the user is performing vigorous exercise (such as when the signal energy exceeds a preset threshold), the cut-off frequency can be automatically adjusted to 20 Hz to retain more high-frequency details; when the user is in a stationary or slow-motion state, the cut-off frequency can be reduced to 5 Hz to more effectively suppress noise. The filtering process can be implemented through the following formula: y[n] = b0*x[n] + b1*x[n-1] +... + bM*x[n-M] - a1*y[n-1] -... - aN*y[n-N], where x[n] is the input signal, y[n] is the filtered output signal, and b and a are filter coefficients, and these coefficients are dynamically adjusted according to the current signal characteristics.
[0044] In step S12, dynamic window smoothing processing based on the signal amplitude change rate and duration is performed on the filtered sensor data to obtain the smoothed sensor data. The dynamic window smoothing processing can automatically adjust the smoothing window size according to the change rate of the motion amplitude, which is more suitable for processing human motion signals than fixed window smoothing. For signal segments with a large change rate (such as a quick waving motion), a smaller smoothing window (such as 5 - 10 sampling points) is used to retain the action details; while for signal segments with a small change rate (such as slow walking), a larger smoothing window (such as 20 - 30 sampling points) is used to obtain a smoother signal curve. In actual implementation, the amplitude change rate of the signal can be calculated first: change rate[n] = |signal[n] - signal[n - 1]| / sampling time interval. Then, the smoothing window size is determined according to the change rate: window size = base window size * (1 - α * min(change rate, maximum change rate) / maximum change rate), where α is an adjustment coefficient (usually taken as 0.5 - 0.8), and the maximum change rate is a preset empirical value. Finally, moving average smoothing processing is applied: smoothed signal[n] = (1 / window size) * Σ(signal[n - i]), where i ranges from 0 to window size - 1.
[0045] In step S13, sensor fusion operations based on complementary filtering or Kalman filtering are performed on the smoothed sensor data to obtain the fused three - dimensional motion attitude data. This step aims to effectively fuse the accelerometer and gyroscope data to obtain a more accurate attitude estimate. Here, complementary filtering utilizes the characteristics that the accelerometer is accurate in the low - frequency band and the gyroscope is stable in the high - frequency band, and the fusion is achieved through the following formula: angle estimate[n] = α * (angle estimate[n - 1] + gyroscope angular velocity * dt) + (1 - α) * accelerometer angle, where α is a weight coefficient (usually 0.9 - 0.98), and dt is the sampling time interval.
[0046] For more complex motion scenarios, Kalman filtering can be used for more accurate attitude estimation. Kalman filtering includes two stages: prediction and update. In the prediction stage, the angular velocity data is used to predict the attitude change, and in the update stage, the acceleration data is used to correct the prediction result.
[0047] In step S14, statistical feature parameters are calculated for the fused three - dimensional motion attitude data to obtain the pre - processed three - dimensional motion data stream. This step extracts features from the attitude data to provide a more discriminative feature representation for subsequent action recognition. Commonly used statistical feature parameters include mean, variance, kurtosis, skewness, signal energy, zero - crossing rate, etc. At the same time, the correlation coefficients between different axes can also be calculated to capture the collaborative relationship of multi - dimensional motion:
[0048] It can be understood that through the series processing of the above four steps, the original multi-axis acceleration data and angular velocity data have been comprehensively optimized and enhanced. Adaptive low-pass filtering effectively removes high-frequency noise while retaining the key signal features in different motion states; dynamic window smoothing intelligently adjusts the smoothing degree according to the change of motion amplitude, retaining the details of fast actions and smoothing the fluctuations of slow actions; sensor fusion operation comprehensively utilizes the complementary advantages of the accelerometer and gyroscope, overcomes the limitations of a single sensor, and provides more accurate attitude estimation; the calculation of statistical feature parameters further extracts the essential features of the signal, laying a solid foundation for subsequent primitive division and pattern recognition. This series of processing not only significantly improves the signal-to-noise ratio and stability of the signal, but also enhances the distinguishability between different actions, enabling more accurate recognition of complex and diverse human motions and providing high-quality data input for subsequent action recognition and instruction mapping.
[0049] S20. Perform hierarchical primitive division and template generation on the preprocessed three-dimensional motion data stream, and establish a motion primitive library containing multi-level primitive features and transition relationships. Here, hierarchical means that motion primitives are organized into different levels, from short-term basic motion units to long-term complex motion patterns, forming a structure from simple to complex. A primitive refers to the smallest recognizable unit that constitutes a complex motion. Features describe the attributes of each primitive, such as shape, amplitude, duration, etc. The transition relationship describes the temporal sequence and combination method between different primitives.
[0050] In some embodiments, step S20 can be implemented through steps S21 - S25:
[0051] In step S21, a time segmentation method based on a sliding time window and an amplitude threshold is applied to the preprocessed three-dimensional motion data stream for micro-primitive division, and a micro-primitive set containing the features of short-time basic motion units is established. Micro-primitives are the most basic constituent units in motion data, similar to phonemes or letters in language, usually with a short duration (about 0.2 - 0.5 seconds), representing basic actions such as wrist rotation, lifting, or pressing. Specifically, the sliding window method can be used to segment the data stream, and the window size is usually set to 20 - 50 sampling points (assuming a sampling frequency of 100 Hz). For the data within each window, calculate the signal energy or amplitude change: window energy = Σ(x[i]^2 + y[i]^2 + z[i]^2), where i ranges from 1 to the window size; amplitude change = max(√(x[i]^2 + y[i]^2 + z[i]^2)) - min(√(x[i]^2 + y[i]^2 + z[i]^2)). When the amplitude change exceeds a preset threshold (such as dynamically set according to the user's historical data), mark this window as a potential micro-primitive boundary. In this way, the continuous motion data stream can be segmented into a series of micro-primitives, and each micro-primitive contains a simple atomic action. For example, the "wave upward" action may be segmented into three micro-primitives: "lift the wrist", "move upward", and "stop".
[0052] In step S22, a clustering algorithm based on the dynamic time warping distance and the Euclidean distance is applied to adjacent micro-primitives in the micro-primitive set for clustering, and a basic-primitive set containing medium-duration motion patterns is obtained. Basic primitives are equivalent to combining multiple related micro-primitives into meaningful action units, similar to words in language, usually with a duration of 0.5 - 2 seconds. The dynamic time warping (DTW) distance is used to calculate the similarity between sequences of different lengths: DTW(X,Y) = min{Σd(xi,yj)}, where X and Y are two micro-primitive sequences, and d(xi,yj) is the Euclidean distance between two data points. In practical applications, hierarchical clustering or the K-means++ algorithm can be combined to cluster similar micro-primitive sequences according to the DTW distance. For example, set the DTW distance threshold to 5.0. When the DTW distance between two micro-primitive sequences is less than this threshold, classify them into the same class. In this way, sequences of micro-primitives for multiple "waves upward", although different in speed and amplitude, will be clustered into the same "wave" basic primitive.
[0053] In step S23, the basic primitives in the basic primitive set are combined by applying a combination rule based on sequence alignment and pattern matching to generate a composite primitive set containing long-term complex motion patterns, thereby establishing a hierarchical primitive structure of micro primitives, basic primitives, and composite primitives. The composite primitives represent complete actions with specific semantics, similar to sentences in language, and usually have a duration of 2 - 10 seconds. The sequence alignment technique can identify patterns in the basic primitive sequence. For example, the Smith-Waterman algorithm is used to find locally similar sequences. In this way, patterns such as "wave up" + "pause" + "wave down" can be discovered and defined as the "wave greeting" composite primitive.
[0054] In step S24, temporal alignment processing based on the dynamic time warping algorithm is applied to the samples of each level of primitives to construct primitive templates containing statistical features and distribution information. The primitive templates are the standardized representations of each type of primitive, facilitating subsequent pattern matching. For each type of primitive (such as "wave up"), multiple sample instances are collected and aligned to a reference template through the DTW algorithm. After alignment, the mean and variance at each time point can be calculated to construct a statistical template: template mean[t] = (1 / N) * Σx_i[t], where i ranges from 1 to N; template variance[t] = (1 / N) * Σ(x_i[t] - template mean[t])^2, where N is the number of samples of this type of primitive. The templates generated in this way contain both the typical patterns of the primitives and retain the variation information, capable of adapting to the subtle differences of different users.
[0055] In step S25, perform probability modeling processing based on Hidden Markov Model (HMM) or Gaussian Mixture Model (GMM) on the primitive templates, analyze the temporal transition probabilities between primitive elements at different levels, and obtain a hierarchical motion primitive library containing primitive feature vectors and transition probability matrices. This step establishes the statistical relationship between primitives and provides probabilistic support for subsequent sequence recognition. For the HMM model, states (corresponding to primitives) and observations (corresponding to sensor data) can be defined, and the transition probability matrix A and emission probability matrix B can be estimated through the Baum-Welch algorithm: A[i,j]=P(state j|state i)=frequency of transition from primitive i to primitive j; B[i,k]=P(observation k|state i)=probability that primitive i generates observation k. For example, analyzing historical data may find that after the "raise wrist" primitive, there is an 80% probability of the "move up" primitive, a 15% probability of the "move horizontally" primitive, and a 5% probability of the "move down" primitive. These probability values form the transition probability matrix, providing a statistical basis for subsequent action prediction. Similarly, GMM can be used to simulate the feature distribution of each primitive: p(x)=ΣπkN(x|μk,Σk), where k ranges from 1 to K, πk is the weight of the kth Gaussian component, and N(x|μk,Σk) is a Gaussian distribution with mean μk and covariance Σk. These parameters can be estimated through the Expectation-Maximization (EM) algorithm to establish a feature distribution model for each type of primitive.
[0056] It can be understood that through the hierarchical primitive division and template generation process, a well-structured motion primitive library is established, which includes a multi-level structure from micro-primitives to composite primitives and the transition relationships between primitives. This hierarchical processing method has multiple technical advantages: First, the division of micro-primitives captures the most basic motion features, laying a solid foundation for subsequent analysis; Second, the basic primitives formed by clustering based on micro-primitives improve the system's recognition ability for similar action variants; Third, the construction of composite primitives enables the system to recognize complete actions with semantic meanings; Finally, the introduction of probability models enables the system to handle the uncertainty and ambiguity in action sequences and improves the recognition accuracy of complex continuous actions. This multi-level design is similar to the human cognitive process of actions, which can not only recognize simple basic actions but also understand complex action combinations, greatly enhancing the system's adaptability and recognition accuracy for diverse user behaviors.
[0057] S30. Apply multi-scale window segmentation to the three-dimensional motion data stream, dynamically adjust the window parameters according to the signal characteristics, and obtain a segmented data set with a hierarchical time structure of micro-windows, medium-windows, and macro-windows. The purpose of this multi-scale window segmentation and dynamic adjustment of window parameters is to analyze motion data at different time granularities and flexibly adjust the segmentation strategy according to the characteristics of the signal, so as to better capture motion patterns of different durations and complexities.
[0058] In some embodiments, step S30 can be implemented through steps S31 - S35:
[0059] In step S31, a three - level parallel sliding window processing including a micro - window, a medium - window, and a macro - window is applied to the three - dimensional motion data stream to obtain window sequences at different time scales. The three - level window structure design corresponds to different time - scale characteristics of human motion and can capture both subtle actions and complex combined actions simultaneously.
[0060] Specifically, the micro - window is usually set to 0.1 - 0.5 seconds (about 10 - 50 sampling points, assuming a sampling frequency of 100 Hz), mainly used to capture instantaneous action changes, such as finger tapping, rapid wrist rotation, etc.; the medium - window is usually set to 0.5 - 3 seconds (about 50 - 300 sampling points), used to capture complete basic actions, such as waving, turning in circles, etc.; the macro - window is usually set to 3 - 10 seconds (about 300 - 1000 sampling points), used to capture complex combined action sequences, such as continuous actions like "waving - pausing - turning in circles", etc.
[0061] In step S32, the energy change rate is calculated for the signals within each initial window sequence, and the points where the energy change rate exceeds the adaptive threshold are marked as action transition points to obtain window data containing action transition marks. The energy change rate reflects the severity change of the motion state and is an important indicator for identifying action transitions. For the signals within the window, the energy change rate can be obtained by calculating the signal energy difference between consecutive time points: Energy[t]=x[t]^2 + y[t]^2+z[t]^2; Energy change rate[t]=|Energy[t] - Energy[t - 1]| / time interval. The adaptive threshold can be dynamically set according to the statistical characteristics of historical data. For example: Adaptive threshold = average energy change rate+β*standard deviation of energy change rate, where β is an adjustable sensitivity parameter, usually taking a value of 1.5 - 3. When the energy change rate at a certain moment exceeds this threshold, that moment is marked as an action transition point.
[0062] In step S33, a complexity metric is calculated for the window data containing action transition markers, and the window overlap rate is dynamically adjusted according to the complexity metric to obtain an optimized window coverage scheme with adaptive overlap characteristics. The dynamic adjustment of the window overlap rate here can ensure that key action transitions are not missed while guaranteeing computational efficiency. The complexity metric can be calculated in various ways, such as signal entropy, spectral complexity, or the number of peaks. According to the calculated complexity metric, the window overlap rate can be dynamically adjusted as follows: overlap rate = base overlap rate + γ * min(complexity metric, maximum complexity) / maximum complexity, where γ is an adjustment coefficient (usually 0.3 - 0.6). For example, for a complex "rapid stop after continuous rotation" action, if the calculated complexity metric is relatively high (such as 0.8), the window overlap rate may increase from the base 30% to 70% to ensure accurate capture of the transition process.
[0063] In step S34, based on the optimized window coverage scheme, motion state analysis is performed on the three-dimensional motion data stream, and regions with sparse action transition markers and low complexity metrics are identified as stationary periods or invalid actions and marked as low-priority regions, resulting in a data stream with priority stratification. This step effectively filters out data regions that contribute significantly to action recognition and reduces the computational burden of subsequent processing. Stationary periods or invalid actions can be identified according to the following criteria:
[0064] If (action transition point density < density threshold) AND (average complexity metric < complexity threshold): region type = low-priority region;
[0065] else: region type = high-priority region
[0066] In step S35, separate caching processing is performed on the regions around the action transition points in the three-dimensional motion data stream with priority stratification to enhance the sampling density of high-information-density regions, resulting in a segmented data set containing micro-window, medium-window, and macro-window hierarchical time structures, adaptive overlap characteristics, and priority information. This enhanced processing for key regions further improves the capture accuracy of the action transition process. For each marked action transition point, an extended buffer can be set for high-density sampling: buffer size = base buffer size * (1 + δ * action transition point energy change rate / maximum energy change rate); sampling density = base sampling density * (1 + ε * action transition point complexity / maximum complexity), where δ and ε are adjustable parameters (usually taken as 0.5 - 1.0).
[0067] It can be understood that through the above multi-scale window segmentation processing, the original three-dimensional motion data stream is converted into a structured segmentation data set, which has three key characteristics: First, the multi-scale time structure enables the system to analyze short-term micro-motions and long-term composite motions simultaneously; Second, the adaptive overlap feature ensures that key information is not lost during the action conversion process due to window segmentation; Finally, the priority hierarchy and high-density sampling mechanism significantly improve the processing accuracy of the system for key action regions while reducing the computational resource consumption for invalid data. This intelligent data segmentation method not only improves the accuracy of subsequent feature extraction and action recognition but also optimizes the allocation of computational resources, enabling the system to achieve more efficient action recognition with limited computational capabilities. Especially for resource-constrained wearable devices, this intelligent processing strategy based on data characteristics is particularly important, which can reduce energy consumption and latency while ensuring recognition accuracy.
[0068] S40. Extract multi-level features corresponding to its time structure from the segmented data set, and perform pattern matching between the extracted multi-level features and the templates in the motion primitive library to obtain a preliminary primitive sequence containing primitive information and confidence. The core objective of this step is to extract multi-level features corresponding to its time structure from the data set with micro-window, medium-window, and macro-window level time structures segmented in the previous step, and then perform pattern matching between these features and the templates in the motion primitive library established in step S20, finally obtaining a preliminary primitive sequence, which includes the identified primitive information and the confidence of the match. Here, multi-level feature extraction and pattern matching are adopted to utilize information at different time scales to identify motion primitives at different levels.
[0069] In some embodiments, step S40 can be completed through steps S41 - S45:
[0070] In step S41, extract time-domain statistical features, rate of change, and jerk features from the segmented data at the micro-window level to obtain a micro-window level feature set. Optionally, micro-windows usually capture short-term motion data of 0.1 - 0.5 seconds, mainly reflecting the instantaneous change characteristics of the motion. Time-domain statistical features include statistics such as mean, variance, standard deviation, kurtosis, and skewness. The rate of change feature reflects the degree of rapid change of the signal, including the first derivative (velocity), zero-crossing rate, and the number of change points. The jerk feature is the derivative of acceleration, reflecting the smoothness and abruptness of the action. These features help distinguish different types of micro-action patterns.
[0071] In step S42, frequency domain features, morphological features, and cross-correlation coefficients of triaxial data are extracted from the segmentation data of the middle window level to obtain the middle window level feature set. Optionally, the middle window captures medium-duration motion data of 0.5 - 3 seconds, which can reflect complete basic motion features. The frequency domain features are obtained through the fast Fourier transform (FFT) and include the main frequency, spectral energy distribution, frequency band energy ratio, etc. The morphological features describe the shape characteristics of the signal and include the number of wave peaks, the number of wave valleys, waveform symmetry, etc. The cross-correlation coefficients of triaxial data reflect the correlation between data in different axes and help identify spatial motion patterns.
[0072] In step S43, temporal pattern features, state transition features of primitive connection characteristics, and time evolution features of energy distribution are extracted from the segmentation data of the macro window level to obtain the macro window level feature set. Optionally, the macro window captures long-duration motion data of 3 - 10 seconds, which can reveal the structural characteristics of complex action sequences. The temporal pattern features describe the overall pattern of a long time series and include repetitive patterns, periodicity, and trend. The state transition features of primitive connection characteristics analyze the transition pattern between primitives and identify state transitions by detecting mutation points of energy, speed, or direction. The time evolution features of energy distribution describe the variation law of energy in the time dimension. By analyzing these features, the system can identify complex long-duration action sequences and the transition characteristics between actions.
[0073] In step S44, the matching algorithms corresponding to the templates at the corresponding levels in the motion primitive library are applied to the feature sets at each level to obtain the matching degree scores at each level. Different levels of primitives use different matching algorithms to adapt to their specific time scale and complexity characteristics. Micro primitives use rule matching, mainly based on threshold judgment and simple distance metrics:
[0074] if (|feature value - template feature value| < feature threshold) for all key features: match successfully;
[0075] else: match fails.
[0076] Basic primitives use the dynamic time warping (DTW) algorithm with feature weights, which can handle the same actions performed at different speeds. For example, for the "drawing a circle" primitive, a higher weight may be assigned to the cross-correlation coefficients of the three axes, while for the "swinging a straight line" primitive, a higher weight may be assigned to the change rate of the main axis direction.
[0077] Composite primitives use a combination of the hidden Markov model (HMM) and conditional random field (CRF) for processing, which can model complex temporal dependencies. For example, for the "waving hello" composite primitive, the HMM can model the state transition sequence of "lifting - waving - lowering", while the CRF can integrate context information such as the static state before the action and the posture maintenance after the action.
[0078] In step S45, the Euclidean distance, cosine similarity, and Mahalanobis distance are calculated for the matching degree score, and a weighted fusion score is calculated to obtain a preliminary primitive sequence including primitive type, timestamp, duration, and recognition confidence. The fusion of multiple distance metrics can provide a more comprehensive similarity assessment. The Euclidean distance measures the absolute distance in the feature space. The cosine similarity measures the directional similarity of feature vectors. The Mahalanobis distance takes into account the correlation between features. The final fusion score can be calculated by weighted average: fusion score = w1 * (1 - normalized Euclidean distance) + w2 * cosine similarity + w3 * (1 - normalized Mahalanobis distance), where w1, w2, and w3 are weight coefficients with a sum of 1. For example, for the "waving" primitive, w1 = 0.3, w2 = 0.4, and w3 = 0.3 may be set; for the direction-sensitive "pointing" primitive, the weight of the cosine similarity may be increased, and w1 = 0.2, w2 = 0.6, and w3 = 0.2 may be set. Based on the fusion score, the system assigns a confidence score to each recognized primitive, records its timestamp and duration, and forms a preliminary primitive sequence.
[0079] It can be understood that through the above multi-level feature extraction and pattern matching process, the system converts the original 3D motion data into a structured primitive sequence. This conversion process fully utilizes information at different time scales: micro-window features capture the fine changes of instantaneous actions, middle-window features identify basic action patterns, and macro-window features reveal the long-term action structure and transition relationships. The combined application of multiple matching algorithms adapts to the characteristics of primitives with different complexities, and the fusion of multiple distance metrics improves the accuracy and robustness of the matching. This multi-level, multi-algorithm processing architecture significantly enhances the system's ability to recognize complex actions, especially showing good adaptability when dealing with factors such as speed changes, amplitude differences, and individual habit differences. Compared with traditional single-time-scale feature extraction and simple threshold matching methods, this technical solution can more accurately capture the essential features and temporal relationships of actions, greatly improving the accuracy and robustness of primitive recognition, and laying a solid foundation for subsequent action verification and instruction generation.
[0080] S50. Perform two-level filtering at the primitive level and action level on the preliminary primitive sequence, and perform primitive-level and action-level verification based on the primitive features and transition relationships in the motion primitive library to obtain a verified high-reliability action sequence. The core objective of this step is to refine the preliminary primitive sequence obtained in step S40. Through two-level filtering at the primitive level and action level and verification based on the motion primitive library, possible misidentifications or low-reliability primitives and action sequences are removed, and finally a verified high-reliability action sequence is obtained.
[0081] In some embodiments, step S50 can be implemented relying on steps S51 - S56:
[0082] In step S51, primitive-level filtering is performed on the preliminary primitive sequence, an adaptive confidence threshold is calculated, and the results below the threshold are filtered to obtain the primitive sequence after confidence filtering. This step removes the primitive recognition results with insufficient matching confidence and improves the accuracy of primitive recognition. Different from a fixed threshold, the adaptive confidence threshold is dynamically adjusted according to the current user's action characteristics and environmental conditions and can be calculated in the following way: Adaptive threshold = Base threshold + α * Signal-to-noise ratio factor - β * User proficiency factor, where α and β are adjustment coefficients (usually α ranges from 0.1 to 0.3 and β ranges from 0.05 to 0.15). The signal-to-noise ratio factor reflects the current signal quality, and the user proficiency factor evaluates the user's operation proficiency based on the historical recognition rate.
[0083] In step S52, duration verification is performed on the primitive sequence after confidence filtering, and the time tolerance adaptive algorithm is applied to adjust the effective time range of each type of primitive to obtain the primitive sequence after time verification. This step filters the primitive recognition results with abnormal durations and ensures the rationality of time characteristics. Each type of primitive has its typical duration range. For example, "click" usually lasts from 0.1 to 0.3 seconds, and "wave" usually lasts from 0.8 to 2.0 seconds. The time tolerance adaptive algorithm dynamically adjusts this range according to the user's historical action data and the current environment: Lower limit of effective time = Standard lower limit * (1 - γ * Coefficient of speed variation); Upper limit of effective time = Standard upper limit * (1 + γ * Coefficient of speed variation), where γ is the tolerance coefficient (usually ranging from 0.2 to 0.5), and the coefficient of speed variation reflects the degree of fluctuation of the user's action speed. Through time verification, primitives with durations significantly deviating from the effective range will be filtered out. For example, a "wave" primitive with a duration of only 0.3 seconds or a "click" primitive with a duration of 3.5 seconds, which may be caused by sensor jitter or recognition errors.
[0084] In step S53, the primitive sequence after time verification is subjected to movement amplitude verification. The primitive energy integral is calculated to filter out primitives with abnormal amplitudes, resulting in a primitive sequence after amplitude verification. This step ensures that the recognized primitives have reasonable movement amplitudes and filters out results with abnormal energy characteristics. The primitive energy integral evaluates the movement amplitude by calculating the total signal energy during the duration of the primitive: Energy integral = Σ(x[t]^2 + y[t]^2 + z[t]^2)*dt, where t ranges from the start to the end of the primitive. For each type of primitive, the system establishes a standard energy integral range. For example, the energy integral range for the "tap" primitive may be 5 - 20, while that for the "forceful wave" primitive may be 50 - 150. Based on the user's historical data, the system can adjust these ranges: Adjusted lower limit = Standard lower limit * (1 - δ * User force coefficient); Adjusted upper limit = Standard upper limit * (1 + δ * User force coefficient), where δ is the adjustment coefficient (usually taken as 0.1 - 0.3), and the user force coefficient reflects the intensity of the user's habitual actions. For example, for a user with a habit of using greater force (force coefficient of 0.2), the adjusted energy integral range for the "tap" primitive may be 4 - 24. Through energy integral verification, the system filters out primitives with abnormal amplitudes, such as a "wave" with too low energy (possibly due to an unconscious small-amplitude hand movement) or a "tap" with too high energy (possibly due to sensor jitter caused by an impact).
[0085] In step S54, the primitive sequence after amplitude verification is subjected to action-level filtering. The context-aware time interval evaluation is performed on the primitive sequence to obtain a sequence after time rationality verification. This step enters action-level filtering to verify whether the temporal relationship between primitives is reasonable. The context-aware time interval evaluation takes into account the primitive type and sequence context and evaluates the time interval between consecutive primitives. There are reasonable time interval ranges for different pairs of primitive types. For example, the reasonable interval between the "raise wrist" and "wave" primitives may be 0 - 0.3 seconds, and the reasonable interval between "wave" and "lower wrist" is also 0 - 0.3 seconds. However, if two "wave" primitives are to form a continuous action, the interval should generally be less than 0.2 seconds; if they are two independent waves, the interval should be greater than 1.0 seconds. The system establishes a probability distribution model for the primitive pair time intervals by analyzing historical action sequences. When the detected interval deviates too much from the center of the probability distribution (such as deviating by more than 2 standard deviations), the primitive pair will be marked as suspicious and may be split or merged. For example, if the interval between two "wave" primitives is 0.5 seconds, which neither meets the compact interval for continuous waves nor the sufficient interval for independent waves, the system may regard it as an abnormal situation that requires further verification.
[0086] In step S55, the sequence after time rationality verification is processed by a probabilistic state machine to verify the rationality of the transition probability, and a sequence after state transition verification is obtained. This step utilizes the state transition probability model in the motion primitive library to verify whether the transitions in the primitive sequence conform to statistical laws. The probabilistic state machine is constructed based on the transition probability matrix of the hidden Markov model. Each primitive type corresponds to a state, and the transition probability between states is statistically derived from historical data: P(primitive j|primitive i) = frequency of primitive j following primitive i / total occurrence frequency of primitive i. For a given primitive sequence, its overall transition probability score can be calculated: sequence probability score = ∏P(primitive t|primitive t-1), where t ranges from 2 to the sequence length. The system sets a probability threshold θ. When the probability P(primitive j|primitive i) of a certain transition is lower than θ, this transition will be marked as suspicious. For example, if the probability of "swing upward" followed by "swing downward" is 0.75, while the probability of "swing upward" followed by "swing left" is only 0.05 (lower than the threshold 0.1), then the sequence containing "swing upward -> swing left" will be marked as a suspicious sequence and needs further verification or adjustment. For a suspicious sequence, the system can attempt to adjust the sequence by inserting, deleting, or replacing primitives to increase its transition probability to an acceptable level. For example, adjusting "swing upward -> swing left" to "swing upward -> pause -> swing left". If the probability of this adjusted sequence is significantly increased, the adjusted sequence is adopted.
[0087] In step S56, motion trajectory integration analysis is performed on the sequence after state transition verification to verify spatial consistency, and context enhancement filtering is performed in combination with the user's current physiological state to obtain a highly reliable action sequence containing the type, time information, primitive composition, and reliability score of the verified action. Motion trajectory integration analysis reconstructs the motion trajectory in three-dimensional space by performing double integration on the acceleration data: velocity[t] = velocity[t-1] + acceleration[t] * dt; position[t] = position[t-1] + velocity[t] * dt. For the reconstructed trajectory, the system verifies its spatial consistency and checks whether it conforms to physical laws and the expected action pattern. If the reconstructed trajectory deviates significantly from the expected pattern, this action sequence will be marked as having low reliability or be filtered out. Physiological state context enhancement takes into account the user's current physiological state, such as heart rate, activity level, and fatigue: context feasibility score = g(action type, heart rate, activity level, fatigue,...). For example, for the detected "rapid continuous swing" action, if the user's current heart rate is already high (e.g., >150 bpm) and the continuous activity time is long (e.g., >30 minutes), the feasibility score of this action will be reduced because the probability of performing complex actions in a fatigued state is low. On the contrary, if the user is in a resting state and concentrated (e.g., during screen interaction), the feasibility score of fine actions (such as "precision click") will be increased.
[0088] Finally, based on the multi-level verification results, the system assigns a comprehensive reliability score to each action and records its complete information, including the action type, start and end times, the sequence of constituent primitives, and the reliability score. For example: Action: "Wave goodbye to the right"; Time: starts at 2.5 seconds, ends at 4.8 seconds, lasts for 2.3 seconds; Constituent primitives: "Raise the wrist" -> "Wave to the right" -> "Wave to the left" -> "Lower the wrist"; Reliability score: 0.92 (high reliability).
[0089] It can be understood that through the double-level filtering and verification process of S50, the initially identified primitive sequence has undergone strict multiple verifications, significantly improving the reliability and accuracy of action recognition. The primitive-level filtering verifies the confidence, duration, and movement amplitude of individual primitives, screening out misidentifications and abnormal data; the action-level filtering verifies the rationality of the primitive sequence from the perspectives of time intervals, state transitions, and spatial trajectories, ensuring that the recognized actions conform to human motion laws and statistical patterns. In particular, the introduction of adaptive thresholds and context-aware mechanisms enables the system to flexibly adjust the judgment criteria for different users and environmental conditions, greatly enhancing the robustness and adaptability of action recognition. The context-enhanced filtering combined with the user's physiological state further integrates action recognition with the natural behavior patterns of the human body, further improving the intelligence level of recognition. This multi-level and multi-angle verification mechanism significantly reduces the false recognition rate and missed recognition rate, providing a highly reliable action sequence input for subsequent instruction mapping, which is a key link in achieving precise human-computer interaction.
[0090] S60. Apply temporal pattern analysis and context processing to the high-reliability action sequence, perform action recognition with reference to the combination rules in the motion primitive library, and integrate multi-source data to obtain a structured action recognition result. The core objective of this step is to perform a higher-level analysis on the previously obtained high-reliability action sequence, identify more complete and complex actions through temporal pattern analysis and context processing, and with reference to the combination rules in the motion primitive library, while integrating multi-source data, and finally obtain a structured action recognition result.
[0091] In some embodiments, step S60 can be implemented through steps S61 - S67:
[0092] In step S61, a multi-level sliding observation window system including short-term, medium-term, and long-term is constructed for the high-reliability action sequence to obtain action observation sequences at multiple time scales. The multi-level observation window enables the system to capture action patterns and rules at different time scales simultaneously. The short-term observation window usually covers the action sequence in the recent 5 - 10 seconds, mainly used to capture immediate action transitions and rapid combinations; the medium-term observation window covers the action sequence of 30 seconds - 2 minutes, used to identify complete activity cycles and behavior patterns; the long-term observation window covers the action history of 5 - 30 minutes, used to analyze user behavior habits and long-term activity trends. These three-level windows act on the action sequence in a sliding manner simultaneously. After each new action is recognized, all windows are updated synchronously.
[0093] In step S62, the adaptive dynamic time warping algorithm is applied to process the periodic actions in the action observation sequences at multiple time scales to identify the periodic patterns and frequency characteristics, and the recognition results of the periodic actions are obtained. The adaptive dynamic time warping (ADTW) algorithm is an improved version of DTW, which can automatically adjust the elastic coefficient to adapt to different changes in motion speed. For the identified periodic actions, the system further extracts the frequency characteristics and cycle stability: Main frequency = 1 / average cycle duration; Cycle stability = 1 - (standard deviation of cycle duration / average cycle duration).
[0094] In step S63, the long short-term memory network with a primitive attention layer is applied to process the non-periodic complex actions in the action observation sequences at multiple time scales to capture long-distance temporal dependencies and obtain the recognition results of the non-periodic complex actions. Non-periodic complex actions (such as yoga pose transitions, specific sign languages, etc.) have complex temporal dependencies and require advanced deep learning models for recognition. The long short-term memory (LSTM) network is an effective tool for processing sequence data, and its core lies in being able to "remember" long-distance relevant information: Input gate: it = σ(Wi·[ht-1,xt] + bi); Forget gate: ft = σ(Wf·[ht-1,xt] + bf); Output gate: ot = σ(Wo·[ht-1,xt] + bo); Candidate memory: Memory cell: Hidden state: ht = ot * tanh(ct). Based on the LSTM, adding a primitive attention mechanism further improves the model's ability to focus on key primitives: Attention weight: αt,i = softmax(score(ht, primitive i feature)); Attention context: ct = Σiαt,i * primitive i feature; Output: yt = softmax(Wy·[ht,ct] + by).
[0095] The primitive attention mechanism enables the model to automatically focus on the most relevant primitives in the sequence based on the current state, thereby capturing key information. By combining LSTM and the attention mechanism, the system can effectively identify complex actions with long-range dependencies, such as "unlock gestures", "command actions", or "specific dance moves", even if there are individual differences in the execution speed and style of these actions.
[0096] In step S64, the recognition results of periodic actions and aperiodic complex actions are merged, and the Bayesian network is applied to the merged action sequence to integrate the sensor data collected by the somatosensory device and the user's historical action data, perform action context analysis, and obtain a context-enhanced action interpretation based on the action history. The Bayesian network can effectively model the conditional dependence relationship between variables and provide a probabilistic inference framework for action recognition. The core of the Bayesian network is the inference based on conditional probability: P(action|observation, history) = P(observation|action, history) * P(action|history) / P(observation|history). By establishing a probabilistic graphical model among actions, sensor observations, and historical data, the system can integrate multi-source information for comprehensive inference. This context-enhanced interpretation greatly improves the accuracy of action recognition and the semantic understanding ability, enabling the system to distinguish the different meanings of similar actions in different contexts.
[0097] In step S65, a weighted directed action transition graph is constructed for the context-enhanced action interpretation based on the action history, and the trajectory features of the user on the weighted directed action transition graph are analyzed to obtain a representation of the user's high-order behavior pattern. The action transition graph is an effective tool for understanding the user's behavior pattern and can reveal the rules and user habits in the action sequence. In the weighted directed action transition graph, nodes represent different action types, edges represent the transition relationships between actions, and the weights of the edges reflect the transition frequencies: weight(action i → action j) = transition frequency(action i → action j) / total occurrence times of action i. By analyzing the trajectory features of the user on this graph, the system can extract high-order behavior patterns, such as: path frequency = number of transitions of the user along a specific path / total number of transitions; path specificity = user path frequency / average user path frequency.
[0098] In step S66, the user's physiological state data is collected and preprocessed, and the preprocessed physiological state data is subjected to multimodal collaborative processing with the user's high-order behavior pattern representation to obtain an enhanced physiological state of the user's action understanding. Physiological state data provides important supplementary information for action understanding and can reveal the user's internal state and intentions. Physiological state data may include heart rate, skin conductivity, body temperature, respiratory rate, etc., and this data is collected through additional sensors of the somatosensory device or connected health devices. Multimodal collaborative processing integrates the behavior pattern and physiological state through a fusion model: Joint representation = F([Behavior pattern representation, Physiological state representation]), where F is the fusion function, which can be simple concatenation, weighted average, or a complex attention mechanism. For example, when the system detects that the user performs the "repeated click" action and the heart rate and skin conductivity increase simultaneously, it may be interpreted as "the user anxiously tries to operate"; while the same action in a state of stable heart rate and low skin conductivity may be interpreted as "the user calmly performs routine operations". This multimodal collaborative processing enables the system to go beyond surface behavior, understand the user's emotional state and usage intention, and lay the foundation for more natural human-computer interaction.
[0099] In step S67, based on the enhanced physiological state of the user's action understanding, an incremental learning algorithm is applied to update the specific action model parameters of the user, where the learning rate is dynamically adjusted according to the physiological state indicators, and a structured user action recognition result including action type, time boundary, constituent primitives, confidence score, and behavior pattern label is obtained. Incremental learning can be achieved through methods such as stochastic gradient descent, and the key lies in the dynamically adjusted learning rate: Learning rate = Base learning rate * (1 - ρ * Fatigue index) * (1 + κ * Attention index), where ρ and κ are adjustment coefficients (usually taken as 0.3 - 0.5). The fatigue index and attention index are calculated based on the physiological state data and reflect the user's current learning adaptability: when the user is fatigued, the system slows down the learning speed to avoid overfitting abnormal actions; when the user is concentrated, the system speeds up the learning speed to more quickly adapt to the user's conscious action pattern. Through incremental learning, the system continuously refines and optimizes the action model for a specific user, and finally outputs a highly personalized and structured action recognition result, such as: Action type: "Two-finger zoom"; Time boundary: starts at 15.3 seconds, ends at 16.7 seconds, and lasts for 1.4 seconds; Constituent primitives: "Two-finger touch screen" → "Two fingers separate" → "Two fingers lift"; Confidence score: 0.96 (extremely high reliability); Behavior pattern label: "Fine control mode", "Focused state operation".
[0100] It can be understood that through the timing pattern analysis and context processing of S60, the system has elevated the verified action sequence to a higher level of understanding. The multi-level sliding observation window captures action patterns at different time scales; adaptive DTW and primitive attention LSTM achieve accurate recognition for periodic and aperiodic actions respectively; the Bayesian network integrates sensor data and historical data to achieve context-enhanced understanding; the weighted directed action transition graph reveals the user's high-level behavior patterns; multi-modal collaborative processing incorporates physiological state information into action understanding to reveal the user's internal state; and incremental learning ensures that the system can continuously adapt and optimize the recognition performance for specific users. This series of processes has greatly enhanced the depth and accuracy of the system's understanding of user actions, enabling the recognition results to contain not only the surface information of "what action was done", but also deep semantic information such as "why this action was done" and "the precise meaning of this action". This deep understanding ability is the key foundation for realizing natural and intelligent human-computer interaction, enabling the system to more accurately predict user intentions and provide more precise responses.
[0101] S70. Perform instruction mapping conversion on the structured action recognition result to obtain an application control instruction corresponding to the recognized action. The instruction mapping conversion here aims to establish a bridge from the recognized user action to the operation of a specific application program.
[0102] In some embodiments, step S70 can be implemented through steps S71 - S77:
[0103] In step S71, the structured action recognition result is processed using a three-level instruction mapping framework including a basic mapping layer, a context enhancement layer, and a personalized adaptation layer to obtain a preliminary action-instruction mapping relationship. The three-level framework provides a progressive mapping process from general to personalized, ensuring that the instructions conform to both general standards and personal habits. The basic mapping layer establishes a direct correspondence between action types and basic instructions, similar to a standard dictionary. These mappings are usually stored in a basic mapping table, and the basic instructions can be quickly obtained through table lookup operations. The context enhancement layer integrates the time boundaries and component primitive information of the action, and adjusts the instruction parameters according to the detailed features of the action execution: enhanced instruction = basic instruction.adjust(action duration, component primitive sequence, execution strength). The personalized adaptation layer applies the user's preference settings and historical usage patterns to further customize the instructions: personalized instruction = enhanced instruction.personalize(user preference, historical usage pattern). Through these three levels of processing, the system obtains a preliminary action-instruction mapping relationship, which takes into account both general standards and incorporates context information and personal preferences.
[0104] In step S72, the dynamic instruction synthesis mechanism is applied to the action execution mode feature in the preliminary action-instruction mapping relationship. By constructing a continuous mapping function between the action feature and the instruction parameter, a parameterized instruction synthesis rule is obtained. This step converts the continuous feature of the action into the specific parameter value of the instruction. The dynamic instruction synthesis mechanism establishes a mapping function from the action feature space to the instruction parameter space: Instruction parameter = f(action feature vector). Similarly, for the "rotation to control brightness" function, a mapping from the rotation angle to the brightness change can be established: Percentage of brightness change = min(100, γ * |rotation angle|) * sign(rotation angle), where γ is the sensitivity coefficient. Through these continuous mapping functions, the system can generate parameterized instructions according to the precise execution mode of the action, rather than simple discrete instructions, greatly improving the fineness and naturalness of the interaction.
[0105] In step S73, the parameterized instruction synthesis rule is fused and analyzed with the current application state. By updating the scene-intention-instruction probability graph, a situation-adaptive instruction planning scheme is obtained. This step ensures that the generated instruction is suitable for the current application scenario and user intention. The scene-intention-instruction probability graph is a Bayesian network that describes the conditional dependence relationship among the three: P(instruction|scene, intention) = P(intention|instruction, scene) * P(instruction|scene) / P(intention|scene). The system calculates the posterior probability of each possible instruction according to the currently detected application scenario (such as "music playing", "map navigation", or "document editing") and the inferred user intention (such as "adjust parameters", "turn the page", or "select an item"), and selects the most likely instruction as the output. For example, when the system detects that the user performs a "two-finger spread" action in the map application, the scene is "map navigation", and the inferred intention is "view more areas", the calculation result of the probability graph may show that the probability of the "zoom in on the map" instruction is 0.85, much higher than other possible instructions such as "open a new window" (0.12) or "split the screen" (0.03). Therefore, "zoom in on the map" is selected as the output instruction. Through this situation-aware instruction planning, the system can dynamically adjust the instruction mapping of the same action according to different application scenarios, making the interaction more intuitive and efficient.
[0106] In step S74, a complexity adjustment based on the user proficiency index is applied to the instruction planning scheme adapted to the scenario. The complexity of available functions is dynamically adjusted according to the proficiency level to obtain a progressive complexity adjustment scheme. This step enables the system to adapt to users with different proficiency levels and provide a personalized interaction experience. The user proficiency index is calculated based on multiple factors: Proficiency index = w1 * average action confidence + w2 * success rate + w3 * frequency of using advanced functions + w4 * interaction fluency, where w1 - w4 are weight coefficients. For example, the system may classify users into three proficiency levels: beginner (0 - 0.4), intermediate (0.4 - 0.7), and advanced (0.7 - 1.0). According to the proficiency level, the system dynamically adjusts the complexity of available functions: Adjusted instruction set = basic instruction set + proficiency index * (advanced instruction set - basic instruction set). This progressive complexity adjustment allows users to gradually explore and master more functions as their proficiency improves, avoiding the frustration of beginners facing a complex system while meeting the needs of professional users for efficient tools.
[0107] In step S75, the application state changes and user responses after the progressive complexity adjustment scheme is executed are monitored, a reward signal is calculated, and the mapping weight matrix is updated using the reinforcement learning framework to obtain mapping rule parameters with adaptive optimization characteristics. This step enables the system to learn from user feedback and continuously optimize the instruction mapping relationship. The reinforcement learning framework consists of four core elements: state, action, reward, and policy: State S: a combination of the current application state and the user state; Action A: the instruction mapping selection executed by the system; Reward R: the reward signal calculated based on user responses; Policy π: the mapping policy from state to action. The reward signal can be calculated based on multiple user feedbacks: Reward = λ1 * task completion efficiency + λ2 * user explicit feedback - λ3 * instruction cancellation / undo rate - λ4 * number of repeated attempts, where λ1 - λ4 are weight coefficients. Through reinforcement learning algorithms such as Q - learning, the system updates the weight matrix of the instruction mapping: Q(S,A) = Q(S,A) + η * [R + γ * maxQ(S',A') - Q(S,A)], where η is the learning rate, γ is the discount factor, and S' is the new state after executing action A. Through continuous learning, the system gradually optimizes the mapping policy to make it more in line with user expectations and usage habits.
[0108] In step S76, the user-defined custom action-instruction mapping rules and the mapping rule parameters with adaptive optimization characteristics are integrated to assign differentiated instructions to different execution variants of the same action type, resulting in a complete personalized mapping system that includes both system learning rules and user-defined rules. This step balances system intelligence and user control, ensuring that users maintain direct control over critical functions. The custom mapping rules can be directly set through the user interface. For different execution variants of the same action type, the system assigns differentiated instructions according to the variant characteristics. For example, for the "wave" action: "quick short-distance wave" → "switch to the next item"; "slow long-distance wave" → "scroll through content"; "repeated quick wave" → "fast forward / rewind". This fine-grained differentiation enables users to trigger different functions through subtle action changes, greatly enhancing the interaction efficiency and flexibility.
[0109] In step S77, the complete personalized mapping system undergoes security verification processing to screen out potential conflicting instructions and apply the final mapping transformation, obtaining application control instructions corresponding to the structured action recognition results. The security verification processing includes multiple checks: 1. Conflict detection: Check whether multiple mapping rules lead to ambiguity; 2. Hazardous operation verification: Check whether there is a confirmation mechanism for high-risk operations (such as delete, send, purchase); 3. Frequency limit: Check whether high-frequency repeated instructions are reasonable; 4. Context consistency: Check whether the instructions are consistent with the current application state. Through these security mechanisms, the system screens out potential risky instructions and finally generates safe, reliable, and user-intent-compliant application control instructions. This structured instruction contains information such as operation type, parameters, priority, context, and security settings, and can be precisely executed by the application program to achieve the purpose of users controlling the application through body gestures.
[0110] It can be understood that through the instruction mapping transformation process, the system has successfully converted the recognized user actions into application control instructions, completing a complete closed-loop from action perception to instruction execution. The three-level instruction mapping framework establishes a mapping hierarchy from basic to personalized; the dynamic instruction synthesis mechanism realizes the conversion from continuous action features to precise instruction parameters; the scenario-intent-instruction probability graph ensures the context adaptability of the instructions; the user proficiency-driven complexity adjustment provides a progressive functional experience; the reinforcement learning framework enables the system to continuously optimize from user feedback; the integration of user-defined mapping and system learning rules balances intelligence and controllability; and the security verification mechanism ensures the reliability and security of the instructions. This multi-level, adaptive, user-centered instruction mapping method greatly improves the accuracy, naturalness, and personalization of body gesture interaction, enabling users to efficiently control various applications through intuitive body movements and achieving a more natural and intelligent human-computer interaction experience.
[0111] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which may be any one or any combination of a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, etc. The computer-readable storage medium includes a three-dimensional motion data processing program 10. The specific implementation manners of the computer-readable storage medium of the present invention are substantially the same as those of the above-mentioned three-dimensional motion data processing method and the server 1, and will not be described in detail herein.
[0112] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A three-dimensional motion data processing method, characterized in that, Including: Performing signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device to obtain a preprocessed three-dimensional motion data stream; Performing hierarchical primitive division and template generation on the preprocessed three-dimensional motion data stream to establish a motion primitive library containing multi-level primitive features and transition relationships; Applying multi-scale window segmentation to the three-dimensional motion data stream, dynamically adjusting window parameters according to signal characteristics, and obtaining a segmented data set with a hierarchical time structure of micro-windows, medium-windows, and macro-windows; Performing multi-level feature extraction corresponding to its time structure on the segmented data set, and performing pattern matching between the extracted multi-level features and the templates in the motion primitive library to obtain a preliminary primitive sequence containing primitive information and confidence; Performing double-level filtering at the primitive level and action level on the preliminary primitive sequence, and performing primitive-level and action-level verification based on the primitive features and transition relationships in the motion primitive library to obtain a highly reliable action sequence after verification; Applying time series pattern analysis and context processing to the highly reliable action sequence, performing action recognition with reference to the combination rules in the motion primitive library, and integrating multi-source data to obtain a structured action recognition result; Performing instruction mapping conversion on the structured action recognition result to obtain an application control instruction corresponding to the recognized action.
2. The three-dimensional motion data processing method according to claim 1, wherein Performing signal preprocessing on the multi-axis acceleration data and angular velocity data collected by the somatosensory device to obtain a preprocessed three-dimensional motion data stream, including: Applying an adaptive low-pass filter to the multi-axis acceleration data and angular velocity data to obtain filtered sensor data; Performing dynamic window smoothing processing based on the signal amplitude change rate and duration on the filtered sensor data to obtain smoothed sensor data; Performing sensor fusion operation based on complementary filtering or Kalman filtering on the smoothed sensor data to obtain fused three-dimensional motion attitude data; Calculating statistical feature parameters for the fused three-dimensional motion attitude data to obtain a preprocessed three-dimensional motion data stream.
3. The three-dimensional motion data processing method according to claim 1, characterized in that Performing hierarchical primitive division and template generation on the preprocessed three-dimensional motion data stream to establish a motion primitive library containing multi-level primitive features and transition relationships, including: Applying a time segmentation method based on a sliding time window and amplitude threshold to the preprocessed three-dimensional motion data stream for micro-primitive division, and establishing a micro-primitive set containing short-time basic motion unit features; Applying a clustering algorithm based on dynamic time warping distance and Euclidean distance to adjacent micro-primitives in the micro-primitive set for clustering to obtain a basic primitive set containing medium-duration motion patterns; Applying a combination rule based on sequence alignment and pattern matching to the basic primitives in the basic primitive set for combination to generate a composite primitive set containing long-time complex motion patterns, thereby establishing a hierarchical primitive structure of micro-primitives, basic primitives, and composite primitives; Performing time series alignment processing based on the dynamic time warping algorithm on the samples of each level of primitives to construct a primitive template containing statistical features and distribution information; Perform probability modeling processing based on Hidden Markov Model or Gaussian Mixture Model on the primitive template, analyze the temporal transition probability between different levels of primitives, and obtain a hierarchical motion primitive library containing primitive feature vectors and transition probability matrices.
4. The three-dimensional motion data processing method according to claim 1, wherein, Apply multi-scale window segmentation to the 3D motion data stream, dynamically adjust the window parameters according to the signal characteristics, and obtain a segmented data set with a hierarchical time structure of micro-window, medium-window, and macro-window, including: Apply a three-level parallel sliding window processing including micro-window, medium-window, and macro-window to the 3D motion data stream to obtain window sequences at different time scales; Calculate the energy change rate of the signals within each initial window sequence, and mark the points where the energy change rate exceeds the adaptive threshold as action transition points to obtain window data containing action transition marks; Calculate the complexity index for the window data containing action transition marks, and dynamically adjust the window overlap rate according to the complexity index to obtain an optimized window coverage scheme with adaptive overlap characteristics; Based on the optimized window coverage scheme, perform motion state analysis on the 3D motion data stream, identify the regions with sparse action transition marks and low complexity index as stationary periods or invalid actions, and mark them as low-priority regions to obtain a data stream with priority stratification; Perform separate caching processing on the regions around the action transition points in the 3D motion data stream with priority stratification to enhance the sampling density of high-information-density regions, and obtain a segmented data set containing the hierarchical time structure of micro-window, medium-window, and macro-window, adaptive overlap characteristics, and priority information.
5. The three-dimensional motion data processing method according to claim 1, wherein, Perform multi-level feature extraction corresponding to its time structure on the segmented data set, and perform pattern matching between the extracted multi-level features and the templates in the motion primitive library to obtain a preliminary primitive sequence containing primitive information and confidence, including: Extract time-domain statistical features, change rate, and jerk features from the segmented data at the micro-window level to obtain a micro-window level feature set; Extract frequency-domain features, morphological features, and cross-correlation coefficients of three-axis data from the segmented data at the medium-window level to obtain a medium-window level feature set; Extract temporal pattern features, state transition features of primitive connection characteristics, and time evolution features of energy distribution from the segmented data at the macro-window level to obtain a macro-window level feature set; Apply the matching algorithm for the corresponding level templates in the motion primitive library to each level feature set to obtain the matching degree scores for each level. Among them, micro-primitives use rule matching, basic primitives use the dynamic time warping algorithm with feature weights, and composite primitives use a combination of Hidden Markov Model and conditional random field for processing; Calculate the Euclidean distance, cosine similarity, and Mahalanobis distance for the matching degree scores, and perform weighted fusion score calculation to obtain a preliminary primitive sequence containing primitive type, timestamp, duration, and recognition confidence.
6. The three-dimensional motion data processing method according to claim 1, wherein, Perform a two-level filtering at the primitive level and action level on the preliminary primitive sequence, and perform primitive-level and action-level verification based on the primitive features and transition relationships in the motion primitive library to obtain a verified high-reliability action sequence, including: Perform primitive-level filtering on the preliminary primitive sequence, calculate the adaptive confidence threshold, filter the results below the threshold, and obtain the primitive sequence after confidence filtering; Perform duration verification on the primitive sequence after confidence filtering, apply the time tolerance adaptive algorithm to adjust the effective time range of each type of primitive, and obtain the primitive sequence after time verification; Perform motion amplitude verification on the primitive sequence after time verification, calculate the primitive energy integral, filter the primitives with abnormal amplitudes, and obtain the primitive sequence after amplitude verification; Perform action-level filtering on the primitive sequence after amplitude verification, perform context-aware time interval evaluation on the primitive sequence, and obtain the sequence after time rationality verification; Apply a probabilistic state machine to the sequence after time rationality verification, verify the rationality of the transition probability, and obtain the sequence after state transition verification; Perform motion trajectory integral analysis on the sequence after state transition verification, verify the spatial consistency, and perform context-enhanced filtering in combination with the user's current physiological state to obtain a highly reliable action sequence containing the type, time information, primitive composition, and reliability score of the verified actions; 7. The three-dimensional motion data processing method according to claim 1, wherein Apply temporal pattern analysis and context processing to the highly reliable action sequence, perform action recognition and integrate multi-source data according to the combination rules in the motion primitive library, and obtain a structured action recognition result, including: Construct a multi-level sliding observation window system including short-term, medium-term, and long-term for the highly reliable action sequence to obtain action observation sequences at multiple time scales; Apply the adaptive dynamic time warping algorithm to the periodic actions in the action observation sequences at multiple time scales to identify the periodic patterns and frequency characteristics, and obtain the recognition results of the periodic actions; Apply a long short-term memory network with a primitive attention layer to the non-periodic complex actions in the action observation sequences at multiple time scales, where the attention mechanism is weighted based on the primitive characteristics in the motion primitive library to capture long-distance temporal dependencies, and obtain the recognition results of the non-periodic complex actions; Merge the recognition results of the periodic actions and the recognition results of the non-periodic complex actions, apply a Bayesian network to the merged action sequence, integrate the sensor data collected by the somatosensory device and the user's historical action data, and perform action context analysis to obtain a context-enhanced action interpretation based on the action history; Construct a weighted directed action transition graph for the context-enhanced action interpretation based on the action history, analyze the trajectory characteristics of the user on the weighted directed action transition graph, and obtain the user's high-order behavior pattern representation; Collect and preprocess the user's physiological state data, and perform multi-modal collaborative processing on the preprocessed physiological state data and the user's high-order behavior pattern representation to obtain a user action understanding with enhanced physiological state; Based on the user action understanding with enhanced physiological state, update the specific action model parameters of the user using an incremental learning algorithm, where the learning rate is dynamically adjusted according to the physiological state indicators, and obtain a structured user action recognition result including the action type, time boundary, constituent primitives, confidence score, and behavior pattern label.
8. The three-dimensional motion data processing method according to claim 1, wherein Perform instruction mapping conversion on the structured action recognition result to obtain an application control instruction corresponding to the recognized action, including: Process the structured action recognition result by applying a three-level instruction mapping framework including a basic mapping layer, a context enhancement layer, and a personalized adaptation layer. Among them, the basic mapping layer establishes the correspondence between action types and basic instructions, the context enhancement layer integrates action time boundaries and constituent primitive information, and the personalized adaptation layer applies user preference settings to obtain a preliminary action-instruction mapping relationship; Process the action execution mode characteristics in the preliminary action-instruction mapping relationship by applying a dynamic instruction synthesis mechanism. By constructing a continuous mapping function between action characteristics and instruction parameters, a parameterized instruction synthesis rule is obtained; Conduct a fusion analysis of the parameterized instruction synthesis rule and the current application state. By updating the scenario-intention-instruction probability graph, a situation-adaptive instruction planning scheme is obtained; Apply complexity adjustment based on the user proficiency index to the situation-adaptive instruction planning scheme. The user proficiency index is calculated by analyzing the confidence score and behavior pattern mark in the structured action recognition result. Dynamically adjust the complexity of available functions according to the proficiency level to obtain a progressive complexity adjustment scheme; Monitor the application state changes and user responses after the progressive complexity adjustment scheme is executed, calculate the reward signal, and apply the reinforcement learning framework to update the mapping weight matrix to obtain the mapping rule parameters with adaptive optimization characteristics; Integrate the user-defined custom action-instruction mapping rule and the mapping rule parameters with adaptive optimization characteristics, allocate differentiated instructions to different execution variants of the same action type, and obtain a complete personalized mapping system that includes both system learning rules and user-defined rules; Perform security verification processing on the complete personalized mapping system, screen out potential conflicting instructions, and apply the final mapping conversion to obtain an application control instruction corresponding to the structured action recognition result.
9. A three-dimensional motion data processing device, characterized in that, It includes a memory, a processor, and a three-dimensional motion data processing program stored in the memory and executable on the processor. When the processor executes the three-dimensional motion data processing program, it implements the three-dimensional motion data processing method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, A three-dimensional motion data processing program is stored on the computer-readable storage medium. When the three-dimensional motion data processing program is executed by the processor, it implements the three-dimensional motion data processing method according to any one of claims 1-8.