Self-adaptive human hand trajectory prediction method and system, computer equipment and storage medium
By combining the index gating unit, KAN network and the adaptive manual trajectory prediction method of the corrected Kalman filter, the problems of high computational complexity and insufficient RNN prediction accuracy are solved, and efficient and accurate human-machine collaboration in medical dispensing scenarios are achieved.
Patent Information
- Application Number
- CN202510779714.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the medical dispensing scenario, the Transformer model has high computational complexity and high data demands, making it difficult to meet real-time requirements. However, the RNN-based method lacks prediction accuracy when processing complex hand movements, resulting in low human-computer collaboration efficiency.
Adaptive hand trajectory prediction method is adopted, combined with the index gating unit (EGU) and KAN network, online adaptive optimization is performed through the modified Kalman filter (MKF), which improves the real-time performance of the model and data representation ability, which is especially suitable for the prediction of complex hand movements.
It improves the efficiency and prediction accuracy of human-computer collaboration in medical dispensing scenarios, can accurately capture complex hand motion trajectories, and enhances the adaptability and robustness of the model.
Smart Images

Figure CN120296526A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine intelligence, and particularly relates to an adaptive human hand trajectory prediction method, system, computer device, and storage medium. Background Technique
[0002] In recent years, with the rapid development of robot technology, the production lines in the medical drug dispensing field have become more efficient and flexible. However, this technological progress has also imposed higher requirements on the accuracy and efficiency of human-robot collaboration. In a medical drug dispensing environment, a robotic arm must be able to quickly respond to human actions within a limited working space, and predicting the human hand trajectory is a key step in solving this problem. However, due to the random and non-linear characteristics of human behavior, action prediction still faces significant challenges.
[0003] In the prior art, the main objectives of human action prediction include generating prediction results close to the actual situation and inferring various possible motion trends to cope with the uncertainty of future actions. Common methods are generally divided into probability models and deterministic models. For example, some studies use Gaussian process regression combined with human joint and scene constraints to predict human actions within a specific time period; there are also studies that propose using stochastic differential equations and generative adversarial networks (GANs) for action prediction. In the prior art, many action prediction methods based on recurrent neural networks (RNNs) have shown excellent performance. Especially when dealing with sequence prediction tasks, the time series processing ability of RNNs has attracted much attention.
[0004] In recent years, the introduction of the Transformer architecture has provided a new solution for human action prediction. By combining spatio-temporal autoregressive or non-autoregressive Transformer models, researchers have achieved short-term accurate prediction and long-term reasonable action sequences. However, the application of the Transformer architecture in the human-robot collaboration drug dispensing scenario still faces multiple technical challenges. First, the Transformer model has high requirements for large-scale data and computing resources. However, it is difficult to obtain labeled data in the medical scenario. In addition, due to its high computational complexity, the Transformer model is difficult to meet the application scenarios with high real-time requirements in the human-robot collaboration drug dispensing process. In contrast, the RNN-based model can gradually update its hidden state, has good real-time processing ability, and can quickly respond to human actions in complex scenarios. In addition, the RNN network structure is simple, easy to implement and deploy, and is particularly suitable for the application environment in medical devices. Compared with the complex Transformer model, the RNN method has less demand for training data and can still provide relatively reliable prediction performance under limited data conditions.
[0005] To solve the above problems, the present invention proposes an adaptive hand trajectory prediction method, system, computer device and storage medium. Summary of the invention
[0006] In view of the above technical problems, the present invention provides an adaptive hand trajectory prediction method, system, computer device and storage medium.
[0007] The technical solution adopted by the present invention to solve the technical problem is: An adaptive hand trajectory prediction method, the method comprising the following steps: S100: Obtain N steps of historical trajectory data of a human hand; S200: Build a hand trajectory prediction model, which includes an exponential gated unit (EGU) and a KAN network. S300: Input the current historical trajectory data into EGU, update the unit state and hidden state through the input gate, forget gate and output gate, and the hidden state of EGU at the last time step is the output of EGU , the output state Pass it to the KAN network; S400: KAN network receives output from EGU , using the activation function Perform multi-layer processing on the input data, gradually extract features from the data, generate high-dimensional feature representations, and obtain preliminary predictions; S500: At each time step k in the overall prediction process, the parameters of the prediction model are adjusted online using the modified Kalman filter MKF to update the predicted model parameters. , according to the feature input of the current time step k , the prediction model parameters adjusted by MKF Generate the prediction result of the current time step, add the prediction result of the current time step to the input, predict the next step, and obtain and output the prediction results of all M steps, where M is the number of steps of the prediction result.
[0008] Preferably, updating the unit state and the hidden state through the input gate, the forget gate and the output gate in S300 includes: By weighting the input Apply activation function tanh to get unit input , through the input gate and forget gate , combined with the previous unit state and current input information Get unit status : ; ; By calculating the sum of the products of the input gate and the forget gate at the current time step and the normalized state at the previous time step, the normalized state at the current time step is obtained: ; The output gate is obtained by applying the sigmoid function to the weighted input . The hidden state is calculated through the output gate and the normalized cell state , where is obtained by dividing the cell state by the normalized state : ; ; ; For the input gate and the forget gate , a stable exponential gating formula is used, and a stable state is introduced through logarithmic and smoothing operations to achieve stable exponential gating. At the same time, in the forward propagation, all the original gating variables and will be replaced by the stabilized versions and : ; Among them, for in the previous formula, , , is expressed as: ; , , , The linear combination results of the gates or states, that is, the original values that have not been processed by the non-linear activation function, , , and respectively represent the input weight vectors connecting the input to the cell input, input gate, forget gate, and output gate. The weight vectors , , and respectively represent the hidden state recursive weight vectors connecting to the cell input, input gate, forget gate, and output gate, and bias terms , , and are the bias terms of the cell input, input gate, forget gate, and output gate.
[0009] Preferably, S400 includes: For each input dimension of the hidden state , it is processed using the activation function to achieve the step-by-step extraction and combination of features: ; wherein, represents a learnable scalar weight coefficient; wherein, the spline function Spline(x) is expressed as a linear combination of B-splines: ; wherein, are trainable parameters, are B-spline basis functions; The overall output of the KAN network is expressed as: ; ; wherein, is the function matrix corresponding to the l-th KAN layer, L is the total number of KAN layers, and the function matrix is composed of the activation functions of all neurons in the l-th layer. Its elements represent the activation functions from the i-th neuron in the l-th layer to the j-th neuron in the l+1-th layer. Through the function matrix , the input activation value X l is mapped to the output activation value X l+1 ; The overall output of the network is the composite result of multiple KAN layers. These KAN layers process and propagate the input data layer by layer to generate a preliminary prediction result , wherein, represents the prediction result of the m-th step at the k-th time step.
[0010] Preferably, S500 includes: Dynamically adjust the parameter step size V k and the covariance matrix Z k through the Kalman gain K k, to optimize the online adaptation ability of the model. Among them, calculating the Kalman gain is specifically as follows: ; is the Kalman gain, which is used to balance the influence of prediction error and input features. is the covariance matrix at time step k - 1. is the noise influence in parameter update. is the input feature. is the identity matrix. The covariance update is specifically as follows: ; ; represents the covariance matrix at the current time step. is the forgetting factor, which is used to control the influence of historical data on the current estimate. is the EMA smoothing factor for covariance update. represents the updated covariance matrix at the current time step. The parameter step size update is specifically as follows: ; ; is the parameter update step size at the current time step. is the current residual. is the true observation value at time step k. is the predicted value at the previous time step. is the smoothing factor of the exponential moving average, which is used to reduce the random fluctuation in the update process. The model parameter update is specifically as follows: ; According to the update step size gradually adjust the model parameters ; represents the model parameters at time step k. represents the model parameters at the previous time step. According to the adjusted model parameters , continuously update to generate the output final prediction result: ; ; Among them, [;] represents vertical concatenation. represents the j - th step prediction result at time k.
[0011] Preferably, before S100, it also includes: Initialize the parameter matrix, covariance matrix of the neural network, and the initial historical trajectory data; initialize the parameters of the exponential gating unit EGU, including the initial weights of the input gate, forget gate, and output gate, and set the initial parameters of the Kalman filter, and set the forgetting factor and the averaging factor of the exponential moving average.
[0012] An adaptive human hand trajectory prediction system, including a historical trajectory data acquisition module, a human hand trajectory prediction model building module, an exponential gating unit, a KAN network, and a human hand trajectory prediction module. The historical trajectory data acquisition module is used to acquire the historical trajectory data of N steps of the human hand. The human hand trajectory prediction model building module is used to build a human hand trajectory prediction model, and the model includes an exponential gating unit EGU and a KAN network. The exponential gating unit receives the current historical trajectory data, updates the unit state and hidden state through the input gate, forget gate, and output gate, and the hidden state of the last time step of the EGU is the output of the EGU. and transmit the output state to the KAN network. The KAN network receives the output from the EGU. and uses the activation function to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction. The human hand trajectory prediction module is used to, at each time step k in the overall prediction process, use the modified Kalman filter MKF to perform online adaptive adjustment on the parameters of the prediction model and update the predicted model parameters. According to the feature input at the current time step k. and the predicted model parameters adjusted by the MKF. generate the prediction result of the current time step, add the prediction result of the current time step to the input, predict the next step until all M-step prediction results are obtained and output, where M is the number of steps of the prediction result.
[0013] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the adaptive human hand trajectory prediction method.
[0014] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the adaptive human hand trajectory prediction method.
[0015] The above-mentioned adaptive human hand trajectory prediction method, system, computer device, and storage medium integrate the KAN network and the exponential gating unit (EGU), and realize the online adaptive optimization of the model through the modified Kalman filter (MKF). It is specially designed for human-machine cooperation in the scenario of human-machine cooperation in dispensing medicine. It is optimized for the situation of insufficient data representation ability or complex hand movement trajectories, and can effectively process complex motion patterns. It is mainly applied to the prediction of wrist movement. By tracking the wrist movement trajectory, it can effectively understand and predict human operation intentions, thereby improving the efficiency of human-machine cooperation in the medical dispensing process. Description of the Drawings
[0016] Figure 1 It is a flowchart of the adaptive human hand trajectory prediction method in an embodiment of the present invention; Figure 2 It is a framework diagram of the prediction model provided by the present invention; Figure 3 It is a comparison diagram of the prediction results of the hand trajectory provided by the present invention; Figure 4 It is a comparison diagram of the prediction results of the hand trajectory of the LSTM method; Figure 5 It is a comparison diagram of the prediction results of the hand trajectory of the WITRAN method; Figure 6 It is a comparison diagram of the prediction results of the hand trajectory of the RNNIK method. Detailed Embodiments
[0017] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0018] In one embodiment, as Figure 1 shown, for the adaptive human hand trajectory prediction method, the method includes the following steps: S100: Obtain the historical trajectory data of N steps of the human hand; S200: Build a human hand trajectory prediction model, the model includes an exponential gating unit EGU and a KAN network; S300: Input the current historical trajectory data into the EGU, update the unit state and the hidden state through the input gate, forget gate, and output gate, and the hidden state of the last time step of the EGU is the output of the EGU , and transfer the output state to the KAN network; S400: The KAN network receives the output from the EGU , and uses the activation function to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction; S500: At each time step k during the overall prediction process, the modified Kalman filter MKF is used to perform online adaptive adjustment on the parameters of the prediction model, and the predicted model parameters are updated , according to the feature input at the current time step k , the prediction model parameters adjusted by MKF Generate the prediction result of the current time step, add the prediction result of the current time step to the input, predict the next step until all M-step prediction results are obtained and output, where M is the number of steps of the prediction result.
[0019] Specifically, human motion is a time series process, and the recurrent neural network (RNN) is very suitable for processing sequence data due to its unique design structure. RNN can capture and utilize time-dependent relationships to predict motion trajectories. However, traditional RNN-based methods have deficiencies in simultaneously capturing and modeling global trends and local details, which leads to a decrease in prediction accuracy. Although RNN-based methods can perform well when the training data is limited, they often perform poorly when dealing with trajectory data with complex patterns and dependencies. This is particularly evident in the scenario of human-robot collaborative dispensing, because hand movements are usually very complex. Traditional RNN-based methods may not be able to accurately predict the complex hand movements of doctors due to limitations in data representation ability or data volume. To solve these problems, this paper proposes a hand trajectory prediction framework based on an improved RNN neural network.
[0020] The present invention discloses a hand trajectory prediction method based on an improved recurrent neural network (RNN), which can effectively improve the human-robot collaboration efficiency of robots in complex scenarios, especially in the medical dispensing scenario. By combining the exponential gating unit (EGU) and the KAN network, the prediction accuracy and the processing ability for complex motion patterns are improved. And through the modified Kalman filter (MKF), online adaptive optimization of the model is realized. The objective of the present invention is to solve the problems existing in the prior art in hand trajectory prediction, such as poor real-time performance, insufficient data representation ability, and poor adaptive ability.
[0021] The present invention adopts an N-to-1 structure. For M-step prediction, the network iteratively adds the new prediction result to the input, and then predicts the next step until all M-step prediction results are obtained. The input data consists of historical trajectory data of N steps, denoted as . First, the current historical trajectory data is input into the EGU module, which can capture short-term and long-term time-dependent information, and the output is the hidden state of the last time step of the EGU i.e., . Then, the KAN module receives , taking advantage of the activation and linear combination of its basis functions to improve prediction accuracy and stability. Finally, adaptive adjustment is performed through the Kalman filter, and the final output is obtained. . Among them, [;] represents vertical concatenation, and the constants M and N represent the number of steps of prediction and historical data respectively. represents the prediction result of the j-th step at the k-th time step.
[0022] Further, in order to enhance the model's ability to capture short-term and long-term motion patterns, the present invention introduces the EGU module. This module is an improvement based on the traditional LSTM, and it introduces a scalar update mechanism. Specifically, through exponential gating and normalization techniques, EGU can maintain high stability and accuracy when processing long sequences and data with fine time variations, and at the same time provide performance comparable to complex models under lower computational complexity. Through these designs, the model can more accurately predict complex human hand trajectories, capture global and local information in trajectory data, especially showing excellent performance in long time series and small motion changes, as Figure 2 shown.
[0023] In one embodiment, updating the cell state and hidden state through the input gate, forget gate, and output gate in S300 includes: By applying the activation function tanh to the weighted input to obtain the cell input , through the input gate and the forget gate , combining the previous cell state and the current input information to obtain the cell state : ; ; By calculating the sum of the products of the input gate and the forget gate at the current time step and the normalized state at the previous time step, the normalized state at the current time step is obtained: ; By applying the sigmoid function to the weighted input to obtain the output gate , through the output gate and the normalized cell state to calculate the hidden state , where is obtained by dividing the cell state by the normalized state obtained as follows: ; ; ; For the input gate and the forget gate , a stable exponential gating formula is used, and a steady state is introduced by performing logarithmic and smoothing operations to achieve stable exponential gating. Meanwhile, during forward propagation, all original gating variables and will be replaced by the stabilized versions and : ; Among them, for in the previous formula, , , is expressed as: ; , , , The linear combination result of the gating or state, that is, the original value that has not been processed by the non-linear activation function, , , and respectively represent the input weight vectors connecting the input to the cell input, input gate, forget gate, and output gate. The weight vectors , , and respectively represent the recurrent weight vectors connecting the hidden state to the cell input, input gate, forget gate, and output gate. The bias terms , , and are the bias terms of the cell input, input gate, forget gate, and output gate.
[0024] Furthermore, the present invention introduces the KAN network. Based on the Kolmogorov-Arnold representation theorem, which states that any continuous multivariate function can be represented as a combination of a finite number of univariate continuous functions and addition operations. In practical applications, the network can be extended to any width and depth. The KAN network enhances the prediction effect under the condition of limited samples through basis function activation and linear combination. The KAN network can decompose complex multi-dimensional data and extract key feature information layer by layer. By applying activation functions and feature extraction layer by layer, it can improve the prediction accuracy of the model, especially maintaining good performance in complex scenarios with scarce data. It enhances the representation and processing ability of motion data.
[0025] Specifically, the KAN network adopts a multi-layer structure and combines the mechanisms of activation functions and linear combinations for feature extraction. In this module, when the output at time t is passed to the KAN layer, the system will apply activation functions to the input data layer by layer for processing. For each input dimension of the hidden state , this process will be executed sequentially and repeated in each layer, thus realizing the gradual extraction and combination of features.
[0026] In one embodiment, S400 includes: For each input dimension of the hidden state , use the activation function for processing, thus realizing the gradual extraction and combination of features: ; where represents a learnable scalar weight coefficient; where the spline function Spline(x) is represented as a linear combination of B-splines: ;
[0027] where are trainable parameters, are B-spline basis functions; this combination method enables the KAN network to flexibly handle different motion patterns and enhances its generalization ability in different application scenarios.
[0028] The GELU activation function performs excellently in time series tasks. Its main advantage is that it can provide smooth gradient flow, reduce the problem of gradient disappearance during training, and thus improve the convergence speed of the model.
[0029] The multi-layer structure of the KAN network generates high-dimensional feature representations through a series of activation and feature extraction operations. These features are gradually transmitted and integrated through the multi-layer processing of the KAN network, thereby generating the final output of the model.
[0030] The overall output of the KAN network is expressed as: ; ; where is the function matrix corresponding to the l-th KAN layer, L is the total number of KAN layers, and the function matrix is composed of the activation functions of all neurons in the l-th layer. Its elements represent the activation functions from the i-th neuron in the l-th layer to the j-th neuron in the l+1-th layer. Through the function matrix , the input activation value X l is mapped to the output activation value X l+1 ; The overall output of the network is the composite result of multiple KAN layers. These KAN layers process and propagate the input data layer by layer to generate preliminary prediction results , where represents the m-th prediction result at the k-th time step.
[0031] This multi-level KAN network not only enhances the model's feature extraction ability in complex scenarios but also improves the model's robustness and prediction accuracy in dealing with data scarcity and diverse motion patterns through effective activation functions and linear combination mechanisms.
[0032] Furthermore, the present invention proposes an online adaptive method based on an improved Kalman filter (MKF), which is specifically used to cope with the time-varying and individual differences of human actions in the human-machine collaborative drug dispensing scenario, and further improve the online adaptive ability of the system. Since human actions are time-varying, for example, a worker may move the entire arm in the initial stage, while only move the wrist in the later stage. In addition, there may be significant differences in the actions of different individuals when performing the same task. Therefore, online adaptability is very important for ensuring the robustness and generality of the system. Most of the existing online adaptive algorithms rely on the stochastic gradient method, and these methods usually cannot ensure obtaining the optimal solution and are less efficient in dealing with complex dynamic scenarios. To solve this problem, the present invention adopts a second-order method to improve the convergence speed and optimization performance. The recursive least squares parameter adaptive algorithm (RLS-PAA) has been proven to be an effective method to achieve optimal adaptation. However, the traditional RLS-PAA fails to consider the noise model, resulting in low efficiency in dealing with noisy data. Therefore, the present invention introduces a smoothing technique in the adaptive process to ensure the stability of prediction and adaptation in a noisy environment.
[0033] To further optimize the adaptive effect and ensure that the latest information has a greater impact on the estimated value, the present invention uses an improved Kalman filter (MKF) for online adaptive adjustment. To prevent the estimated value from saturating, a forgetting factor λ is added to the traditional Kalman filter, and the exponential moving average (EMA) filtering technique is combined to smooth the adaptive process, thereby enhancing the self-adaptability and generalization ability of the prediction system.
[0034] In one embodiment, S500 includes: Dynamically adjust the parameter step size V k and the covariance matrix Z k through the Kalman gain K k and exponential moving average smoothing to optimize the online adaptation ability of the model. Among them, calculating the Kalman gain is specifically: ; is the Kalman gain, which is used to balance the influence of the prediction error and the input feature, is the covariance matrix at time step k-1, is the noise influence in parameter update, is the input feature, is the identity matrix; The covariance update is specifically: ; ; Represents the covariance matrix at the current time step, is the forgetting factor, which is used to control the influence of historical data on the current estimate, is the EMA smoothing factor for covariance update, represents the updated covariance matrix at the current time step; The specific parameter step size update is as follows: ; ; is the parameter update step size at the current time step, is the current residual, is the true observation value at time step k, is the predicted value at the previous time step, is the smoothing factor of the exponential moving average, which is used to reduce the random fluctuations in the update process; The specific model parameter update is as follows: ; According to the update step size gradually adjust the model parameters ; represents the model parameters at time step k, represents the model parameters at the previous time step; According to the adjusted model parameters , continuously update to generate the output final prediction result: ; ; Among them, [;] represents vertical concatenation, represents the j-th step prediction result at time k.
[0035] In one embodiment, before S100, it further includes:
[0036] Initialize the parameter matrix, covariance matrix of the neural network, and the initial historical trajectory data; initialize the parameters of the exponential gating unit EGU, including the initial weights of the input gate, forgetting gate, and output gate, and set the initial parameters of the Kalman filter, set the forgetting factor and the averaging factor of the exponential moving average.
[0037] The overall prediction process is specifically as follows: Initialize model parameters: Initialize the parameter matrices, covariance matrices of the neural network, and the initial historical trajectory data. Initialize the parameters of the Exponential Gated Unit (EGU), including the initial weights of the input gate, forget gate, and output gate, and set the initial parameters of the Kalman filter, set the forgetting factor and the relevant parameters of the Exponential Moving Average (EMA); Historical trajectory input: Obtain the historical trajectory data of the human hand through the realsenseT435 camera ; Input the historical data into the Exponential Gated Unit (EGU) module for processing to capture the short-term and long-term motion features at the current time step t; EGU module processing: In the EGU module, process the input data, update the unit state and hidden state through the input gate, forget gate, and output gate, and pass the processed output state to the KAN network for further processing; KAN network feature extraction: The KAN network receives the output from the EGU , and uses the activation function and the spline function Spline(x) to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction; MKF adaptive adjustment: At each time step k, use the Modified Kalman Filter (MKF) to perform online adaptive adjustment on the model parameters, update the predicted model parameters , according to the feature input at the current time step k , generate the final output prediction result through the model parameters adjusted by the MKF .
[0038] In the scenario of human-robot collaborative dispensing, objects often need to be transferred between the robot and the human. For example, the human may want the robot to pick up a discarded test tube for processing, or the human may reach out to receive a tool provided by the robot. In these cases, the robot must have the ability to predict the trajectory of the human hand to determine the task target point and plan a collision-free path. For this purpose, an experiment was conducted. In the experiment, an experimenter stood in front of the robot arm and used test tubes to transfer objects. Three test tube racks were placed in front of the robot arm, on both sides and in the middle of the experimental table respectively. The experimenter needed to start from the initial position and transfer the test tubes to each test tube rack in turn. This experiment used an Intel RealSense D435 camera to record data and used MediaPipe to extract the wrist trajectory. To quantitatively evaluate the prediction accuracy, the average prediction error of each prediction step was defined as: ; where T represents the total number of time steps, is the actual position at time step i + j, is the predicted position at time step i + j. Experimental results show that the model can accurately predict hand trajectories and, when compared with several popular non-Transformer methods such as LSTM, RNNIK, and WITRAN, the present invention demonstrates superiority in this scenario, as Figures 3 - 6 shown, where x, y, and z are the three dimensions of the Cartesian coordinates of the human hand's spatial position, with the unit being mm. When the step size is 10, the error of the present invention is reduced by 15.15% and 6.67% compared to LSTM and RNNIK respectively; when the step sizes are 20, 30, 40, and 50, the error of the present invention is significantly reduced compared to all other methods, with the maximum reduction reaching 26.67%, as shown in Table 1: Table 1 Comparison of predicted errors of hand trajectories for each method
[0039] An adaptive human hand trajectory prediction system, including a historical trajectory data acquisition module, a human hand trajectory prediction model construction module, an exponential gating unit, a KAN network, and a human hand trajectory prediction module; The historical trajectory data acquisition module is used to acquire the historical trajectory data of N steps of the human hand; The human hand trajectory prediction model construction module is used to construct a human hand trajectory prediction model, and the model includes an exponential gating unit EGU and a KAN network; The exponential gating unit receives the current historical trajectory data, updates the unit state and hidden state through an input gate, a forget gate, and an output gate, and the hidden state of the last time step of the EGU is the output of the EGU , and transmits the output state to the KAN network; The KAN network receives the output from the EGU , and uses an activation function to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction; The human hand trajectory prediction module is used to, at each time step k in the overall prediction process, use a modified Kalman filter MKF to perform online adaptive adjustment on the parameters of the prediction model and update the predicted model parameters , according to the feature input at the current time step k , and generate the prediction result of the current time step through the predicted model parameters adjusted by the MKF , add the prediction result of the current time step to the input, predict the next step until all M-step prediction results are obtained and output, where M is the number of steps of the prediction result.
[0040] For the specific limitations of the adaptive human hand trajectory prediction system, reference may be made to the limitations of the adaptive human hand trajectory prediction method in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned adaptive human hand trajectory prediction system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0041] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the adaptive human hand trajectory prediction method are implemented.
[0042] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the adaptive human hand trajectory prediction method are implemented.
[0043] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0044] The above has introduced in detail the adaptive human hand trajectory prediction method, system, computer device, and storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An adaptive human hand trajectory prediction method, characterized in that, The method includes the following steps: S100: Obtain the historical trajectory data of N steps of the human hand; S200: Build a human hand trajectory prediction model, where the model includes an exponential gating unit EGU and a KAN network; S300: Input the current historical trajectory data into the EGU, and update the cell state and hidden state through the input gate, forget gate, and output gate. The hidden state of the last time step of the EGU is the output of the EGU. Transfer the output state to the KAN network. S400: The KAN network receives the output from the EGU , and uses an activation function to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction; S500: At each time step k during the overall prediction process, the modified Kalman filter MKF is used to perform online adaptive adjustment on the parameters of the prediction model, and the predicted model parameters are updated , according to the feature input at the current time step k , the predicted model parameters adjusted by the MKF generate the prediction result of the current time step, add the prediction result of the current time step to the input, predict the next step until the prediction results of all M steps are obtained and output, where M is the number of steps of the prediction result.
2. The method according to claim 1, characterized in that, Updating the cell state and the hidden state through the input gate, forget gate, and output gate in S300 includes: By applying an activation function tanh to the weighted input to obtain the cell input , through the input gate and the forget gate , combining the previous cell state and the current input information to obtain the cell state : ; ; By calculating the sum of the products of the input gate and the forget gate at the current time step with the normalized state at the previous time step, the normalized state at the current time step is obtained: ; By applying a sigmoid function to the weighted input the output gate is obtained and the hidden state is calculated through the output gate and the normalized cell state where is obtained by dividing the cell state by the normalized state as follows: ; ; ; For the input gate and the forget gate , a stable exponential gating formula is used. A stable state is introduced by performing logarithmic and smoothing operations to achieve stable exponential gating. Meanwhile, during forward propagation, all original gating variables and will be replaced by the stabilized versions and : ; Among them, for in the previous formula, , , is expressed as: ; , , , The linear combination result of the gating or state, i.e., the raw value that has not been processed by the non-linear activation function, , , and respectively represent the input weight vectors that connect the input to the cell input, input gate, forget gate, and output gate. The weight vectors , , and respectively represent the recurrent weight vectors that connect the hidden state to the cell input, input gate, forget gate, and output gate. The bias terms , , and are the bias terms for the cell input, input gate, forget gate, and output gate.
3. The method according to claim 2, characterized in that, S400 includes: For the hidden state of each input dimension , it is processed using the activation function to achieve the gradual extraction and combination of features: ; Among them, represents a learnable scalar weight coefficient; Among them, the spline function Spline(x) is expressed as a linear combination of B-splines: ; Among them, are trainable parameters, are B-spline basis functions; The overall output of the KAN network is expressed as: ; ; Among them, is the function matrix corresponding to the l-th KAN layer, where L is the total number of KAN layers. The function matrix is composed of the activation functions of all neurons in the l-th layer. Its elements represent the activation function from the i-th neuron in the l-th layer to the j-th neuron in the l+1-th layer. Through the function matrix , the input activation value X l is mapped to the output activation value X l+1 ; The overall output of the network is the combined result of multiple KAN layers that process and propagate the input data layer by layer to generate preliminary prediction results. , where represents the prediction result of the m-th step at the k-th time step.
4. The method according to claim 3, characterized in that, S500 includes: Through the Kalman gain K k and exponential moving average smoothing, dynamically adjust the parameter step size V k and the covariance matrix Z k , to optimize the online adaptability of the model. Among them, the calculation of the Kalman gain is specifically as follows: ; is the Kalman gain, which is used to balance the influence of the prediction error and the input features, is the covariance matrix at time step k-1, is the noise influence in parameter update, is the input feature, is the identity matrix; The covariance update is specifically: ; ; represents the covariance matrix at the current time step, is the forgetting factor used to control the influence of historical data on the current estimate, is the EMA smoothing factor for covariance update, represents the covariance matrix at the updated current time step; The parameter step size update is specifically: ; ; is the parameter update step size at the current time step, is the current residual, is the true observation value at time step k, is the predicted value at the previous time step, is the smoothing factor of the exponential moving average, which is used to reduce the random fluctuations in the update process; The model parameter update is specifically: ; According to the update step size Gradually adjust the model parameters ; represents the model parameters at time step k, represents the model parameters of the previous time step; According to the adjusted model parameters , continuously update to generate and output the final prediction result: ; ; where [;] represents vertical concatenation, represents the prediction result of the j-th step at time k.
5. The method according to claim 4, wherein Before S100, it also includes: Initialize the parameter matrix, covariance matrix of the neural network, and the initial historical trajectory data; initialize the parameters of the exponential gating unit EGU, including the initial weights of the input gate, forget gate, and output gate, and set the initial parameters of the Kalman filter, set the forgetting factor and the averaging factor of the exponential moving average.
6. Adaptive human hand trajectory prediction system, characterized in that, It includes a historical trajectory data acquisition module, a human hand trajectory prediction model building module, an exponential gating unit, a KAN network, and a human hand trajectory prediction module; The historical trajectory data acquisition module is used to obtain the historical trajectory data of N steps of the human hand; The human hand trajectory prediction model building module is used to build a human hand trajectory prediction model, where the model includes an exponential gating unit EGU and a KAN network; The exponential gating unit receives the current historical trajectory data, updates the cell state and the hidden state through the input gate, the forget gate, and the output gate, and the hidden state at the last time step of the EGU is the output of the EGU. , and passes the output state to the KAN network; The KAN network receives the output from the EGU , and uses the activation function to perform multi-layer processing on the input data, gradually extract the features in the data, generate a high-dimensional feature representation, and obtain a preliminary prediction; The human hand trajectory prediction module is used to, at each time step k in the overall prediction process, use the modified Kalman filter (MKF) to perform online adaptive adjustment on the parameters of the prediction model and update the predicted model parameters. , according to the feature input at the current time step k , the prediction model parameters adjusted by the MKF Generate the prediction result at the current time step, add the prediction result at the current time step to the input, predict the next step until the prediction results of all M steps are obtained and output, where M is the number of steps of the prediction result.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 5.
Citation Information
Cited By
Dexterous hand multi-mode sensing and control method, system and equipment and medium
CN121552386A