Hand shape recovery method, electronic device, and storage medium
By using neural network models for temporal modeling and cross-finger collaborative modeling, the problem of unstable hand shape recovery in traditional magnetic sensing data acquisition is solved, achieving high-precision and robust hand shape recovery, adapting to different devices and wearing methods, and improving user interaction experience and system reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI TASHI ZHIHANG TECHNOLOGY CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-06-12
AI Technical Summary
Traditional frame-by-frame independent analysis methods are easily affected by environmental disturbances, noise, and minute hand displacements during magnetic sensing data acquisition, resulting in unstable hand shape recovery results, frequent jitter and sudden posture changes, which affect user interaction experience and system reliability.
A neural network model is used for temporal modeling and cross-finger collaborative modeling. By acquiring continuous multi-frame magnetic field observation data, the hand posture is restored by utilizing the temporal context and the correlation between fingers. The model includes a temporal modeling module, a cross-finger collaborative modeling module, and a posture decoding module. The model is trained by combining temporal smoothing loss and hand structure consistency loss, and outputs high-precision hand shape restoration results.
It improves the accuracy and robustness of hand shape recovery, reduces the impact of single-frame random noise, ensures the continuity and reliability of output results, adapts to different devices and wearing methods, and enhances the stability and practicality of the system.
Smart Images

Figure CN122197655A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technology, and in particular to a hand shape recovery method, electronic device, and storage medium. Background Technology
[0002] Hand posture recognition and hand shape reconstruction technologies are widely used in virtual reality, human-computer interaction, remote control, rehabilitation training, and smart wearable devices. Magnetic positioning technology has become an important technical approach for hand motion perception due to its advantages such as not relying on lighting conditions, strong resistance to occlusion, and ability to achieve continuous tracking at close range.
[0003] Currently, a common method involves deploying magnetic sensing units and magnetic sources (such as electromagnetic coils) on the hand or fingers to collect data on magnetic field changes. Then, based on a pre-defined physical model, geometric relationships, or calibration parameters, each frame of data is independently analyzed to reconstruct the current hand joint angles or key point positions. This frame-by-frame analysis method can achieve a certain degree of attitude estimation under conditions of fixed equipment structure and sufficient calibration.
[0004] However, due to the susceptibility of magnetic sensing data to environmental disturbances, magnetic sensing unit noise, hardware jitter, and minute hand movements during acquisition, traditional frame-by-frame independent processing methods lack effective utilization of temporal continuity. When a frame experiences a momentary anomaly or random fluctuation, the output of that frame will show a significant jump, resulting in frequent jitter in the continuously output hand gesture sequence, manifesting as abrupt posture changes, discontinuous movement edges, and noticeable visual flickering. This unstable output not only degrades the user experience but also introduces unnecessary command fluctuations, affecting the system's reliability and control accuracy. Summary of the Invention
[0005] This application provides a hand shape recovery method, electronic device, and storage medium, which can improve the accuracy and robustness of hand shape recovery.
[0006] In a first aspect, this application provides a hand shape recovery method applied to an electronic device. The method includes: acquiring continuous multi-frame magnetic field observation data collected by a magnetic sensing unit disposed on the hand, and constructing a temporal input segment based on the continuous multi-frame magnetic field observation data in chronological order; inputting the temporal input segment into a neural network model, and obtaining the hand shape recovery result output by the neural network model. The neural network model includes a temporal modeling module, a cross-finger collaborative modeling module, and a posture decoding module. The temporal modeling module is used to perform temporal encoding processing based on the temporal input segment and output temporal feature data. The temporal feature data is used to characterize the correlation between consecutive moments in the temporal input segment. The cross-finger collaborative modeling module is used to perform cross-finger attention interaction processing based on the temporal feature data and output collaborative feature data. The collaborative feature data is used to characterize the correlation between each finger. The posture decoding module is used to perform regression mapping processing based on the collaborative feature data and output the hand shape recovery result.
[0007] In the above method, the dynamic change pattern of magnetic field observation data in the time dimension is captured by the temporal modeling module, the spatial dependency relationship between different fingers is learned by the cross-finger collaborative modeling module, and the complete hand shape is obtained by the posture decoding module. It can directly recover the hand posture end-to-end from continuous multi-frame magnetic field data without relying on external optical references or complex hand kinematic prior models, thus improving the accuracy and robustness of hand shape recovery.
[0008] In one possible implementation of the first aspect, the temporal modeling module is one of a long short-term memory network, a gated recurrent unit, a temporal convolutional network, or a Transformer network.
[0009] In one possible implementation of the first aspect, the cross-finite collaborative modeling module is a multi-head self-attention mechanism or a graph neural network.
[0010] In one possible implementation of the first aspect, a time-series input segment is constructed in chronological order based on continuous multi-frame magnetic field observation data, including: using a sliding time window to extract a fixed-length time-series input segment from the continuous time series, wherein the window length of the sliding time window is set according to the sampling frequency of the magnetic sensing unit.
[0011] In one possible implementation of the first aspect, the hand shape restoration result includes the angle values of each joint of the five fingers or the three-dimensional spatial coordinates of key points of the hand.
[0012] In one possible implementation of the first aspect, the method further includes: acquiring a training dataset containing magnetic field observation data of multiple different magnetic sensing units and corresponding hand shape labels; constructing an initial neural network model, the initial neural network model including a temporal modeling module, a cross-finger collaborative modeling module, and a posture decoding module; training the initial neural network model using the training dataset and an objective function to obtain a neural network model; wherein the objective function includes at least a temporal smoothing loss and a hand structure consistency loss, the temporal smoothing loss being used to suppress the differences between outputs at adjacent time points, and the hand structure consistency loss being used to constrain the hand shape topology.
[0013] In one possible implementation of the first aspect, the objective function also includes a cross-device domain alignment loss, which is used to characterize the distributional differences between the feature representations of corresponding actions on different magnetic sensing units.
[0014] In one possible implementation of the first aspect, the temporal smoothing loss is the first-order difference norm between the hand shape recovery results of adjacent frames output by the attitude decoding module; the hand structure consistency loss includes at least one of the following: the deviation between the actual length of each phalanx calculated based on the phalanx endpoint coordinates in the hand shape recovery results output by the attitude decoding module and the preset length, and the deviation between the joint angle in the hand shape recovery results and the preset angle range.
[0015] In a second aspect, this application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of one or more processors of the electronic device, for performing the method described in the first aspect or any possible implementation of the first aspect.
[0016] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect. Attached Figure Description
[0017] Figure 1 An exemplary flow of a hand shape restoration method is provided according to an embodiment of this application;
[0018] Figure 2 According to an embodiment of this application, an exemplary flow of a method for training a neural network model using a hand shape recovery method is provided.
[0019] Figure 3 According to an embodiment of this application, a hardware structure diagram of an electronic device is provided. Detailed Implementation
[0020] The illustrative embodiments of this application include, but are not limited to, a hand shape recovery method, an electronic device, and a storage medium.
[0021] As previously mentioned, with the development of wearable hand interaction devices, data gloves, virtual reality / augmented reality interaction devices, robot teleoperation systems, and human hand motion capture systems, solutions based on electromagnetic field (EMF) for hand posture perception have been widely applied. These solutions typically acquire the position, posture, or related observations of the hand's end-effector or key points in space by deploying electromagnetic sensing units or electromagnetic transmitting / receiving units, and then combine this with a hand kinematic model to reconstruct the hand's movement state.
[0022] In some embodiments, hand shape reconstruction schemes rely on fitting only a single pose or a small number of samples, resulting in insufficient constraints and poor robustness. For example, relying solely on a single static action (such as clenching a fist) to collect positional data for hand shape reconstruction can lead to problems such as unidentifiable parameters, local optima, or non-unique solutions. Furthermore, noise, missing data points, or local drift may occur during the measurement process, resulting in unstable results.
[0023] To address the aforementioned issues, this application provides a method comprising: acquiring continuous multi-frame magnetic field observation data collected by a magnetic sensing unit disposed on the hand, and constructing a temporal input segment based on the continuous multi-frame magnetic field observation data in chronological order; inputting the temporal input segment into a neural network model, and obtaining the hand shape recovery result output by the neural network model. The neural network model includes a temporal modeling module, a cross-finger collaborative modeling module, and a posture decoding module. The temporal modeling module performs temporal encoding processing based on the temporal input segment, outputting temporal feature data, which characterizes the correlation between consecutive moments in the temporal input segment; the cross-finger collaborative modeling module performs cross-finger attention interaction processing based on the temporal feature data, outputting collaborative feature data, which characterizes the correlation between individual fingers; and the posture decoding module performs regression mapping processing based on the collaborative feature data, outputting the hand shape recovery result.
[0024] Thus, by using temporal input segments, the current output can be constrained using the temporal context, reducing the impact of random noise in a single frame. Furthermore, the temporal modeling module can model the dynamic relationships between consecutive moments, avoiding abrupt changes caused by independent frame-by-frame parsing. Therefore, by utilizing the correlation information between frame data, the accuracy and reliability of hand shape reconstruction can be improved.
[0025] The following is combined with Figure 1 This application presents an exemplary process for a time-series learning-based magnetic positioning hand shape recovery method.
[0026] It is understood that the methods provided in this application embodiment can be executed by electronic devices (such as computers, wearable devices, VR controllers, etc.), and this application does not limit the type of electronic device. The main application scenario of this application is embodied intelligent data collection, but it can also be used in fields such as virtual reality, human-computer interaction, and rehabilitation training, and this application does not limit it in this regard.
[0027] like Figure 1 As shown, this exemplary process includes the following steps.
[0028] S101: Acquire continuous multi-frame magnetic field observation data collected by the magnetic sensing unit set on the hand.
[0029] For example, the magnetic sensing unit can be an array of one or more magnetic sensors, or a magnetic detection module integrated into a data glove, wearable device, handheld terminal, or interactive controller.
[0030] For example, the magnetic source can be an electromagnetic coil. The placement of the magnetic sensing unit varies depending on the task: for hand skeleton estimation tasks, the magnetic sensing unit is placed on the fingers; for hand position tracking tasks, the magnetic sensing unit is placed on the wrist. The data acquisition device can be a glove (for hand skeleton estimation) or a wristband (for hand position tracking).
[0031] For example, the data collected could be from users or robots.
[0032] For example, the acquired multi-frame magnetic field observation data may include one or more of the following: multi-axis magnetic field strength values, magnetic field direction components, synchronous magnetic measurement values at different sensing locations, and inertial data, timestamps, or status identifiers acquired synchronously with the magnetic measurements. The sampling frequency can be set according to application requirements, for example, 50Hz, 100Hz, or 200Hz. To enhance the model's generalization ability, data acquisition can cover different devices, different wearing methods, individuals with different hand shapes, different action categories, and different environmental conditions; this application does not impose any limitations in this regard.
[0033] S102: Construct a time-series input segment based on the continuous multi-frame magnetic field observation data in chronological order.
[0034] For example, the collected raw magnetic field observation data is preprocessed, including time synchronization, missing value filling, noise filtering or outlier removal, mean normalization or variance normalization, and multi-channel stitching.
[0035] In addition, preprocessing may include converting the original magnetic signal into a high-dimensional feature vector. For example, when acquiring a 9-axis magnetic signal, it can be converted into a 3-axis included angle and a 9-axis modulus. The converted features are used as part of the network input.
[0036] Then, a fixed-length time sequence input segment is extracted from the continuous time series using a sliding time window. The window length of the sliding time window can be set according to the sampling frequency of the magnetic sensing unit; for example, 10 frames are taken for 50Hz data. Furthermore, the window length can also be dynamically controlled according to the motion speed, so that the displacement within the time window is kept within a certain range, i.e., a dynamic length window. The sliding step size is, for example, 1 frame, to achieve frame-by-frame output. The time sequence input segment can include the current time and several times before it, or it can include several times after it (for example, for offline processing).
[0037] It is understandable that by constructing temporal inputs instead of single-frame inputs, information about the continuity of actions can be provided to subsequent models.
[0038] S103: Input the time sequence input segment into the neural network model and obtain the hand shape recovery result output by the neural network model.
[0039] For example, the neural network model as a whole can be a recurrent neural network (RNN), a long short-term memory network (LSTM), a gated recurrent unit (GRU), a temporal convolutional network (TCN), a Transformer network, a graph neural network (GNN), or a combination of the above networks, and this application does not limit it.
[0040] For example, the neural network model includes a temporal modeling module, a cross-finger collaborative modeling module, and a pose decoding module.
[0041] The temporal modeling module is used to perform temporal encoding processing based on temporal input segments and output temporal feature data. The temporal feature data is used to characterize the correlation between consecutive moments in the temporal input segments.
[0042] For example, the temporal modeling module may employ one of the following: Long Short-Term Memory (LSTM) network, Gated Recurrent Unit (GRU), Temporal Convolutional Network (TCN), or Transformer network.
[0043] The cross-finger collaborative modeling module is used to perform cross-finger attention interaction processing based on temporal feature data and output collaborative feature data, which is used to characterize the relationship between each finger.
[0044] For example, the cross-finite collaborative modeling module can employ a multi-head self-attention mechanism or a graph neural network.
[0045] The pose decoding module is used to perform regression mapping based on collaborative feature data and output hand shape recovery results.
[0046] For example, the pose decoding module may employ a fully connected layer, a hybrid expert network, a hierarchical output head, or a decoder structure incorporating hand kinematic constraints. The hand shape recovery result may include the angle values of each joint of the five fingers (e.g., 21 angles) or the three-dimensional spatial coordinates of key points of the hand (e.g., the three-dimensional coordinates of 30 key points).
[0047] According to some embodiments, the neural network model does not simply infer individual local parameters from single-point magnetic field values, but rather learns the relationship between the overall trends of five-finger movements and the reasonable hand shape. For example, during training, the neural network model learns one or more of the following: the coordinated changes of the thumb, index finger, middle finger, ring finger, and little finger in different movements; the natural coupling relationship between adjacent joints; the temporal continuity during movement transitions; and the inherent physiological reasonable range of the human hand structure. Therefore, even if some channel data is disturbed at a certain moment, the model can still combine the information of the remaining fingers and historical movement trends to recover a hand shape restoration result that is consistent with or approximately consistent with the actual movement, avoiding local anomalies, posture reversals, joint abrupt changes, or overall hand shape distortions commonly found in traditional formula analysis.
[0048] For example, the output results can be directly used to drive downstream applications such as virtual hand display, robot hand control, human-computer interaction command recognition, rehabilitation assessment, or motion scoring.
[0049] This application's embodiments achieve high-precision and high-stability hand pose reconstruction by constructing temporal input segments and combining them with a deep neural network model. Constraining the current output using temporal context information effectively reduces the impact of single-frame random noise; capturing the dynamic correlation between consecutive moments through a temporal modeling module avoids abrupt pose changes caused by independent frame-by-frame parsing, thereby improving the accuracy and reliability of the reconstruction results. Simultaneously, the introduction of a cross-finger collaborative modeling module effectively depicts the spatial collaborative relationships between fingers, making the reconstructed hand shape more consistent with real kinematics, further enhancing the system's robustness and practicality.
[0050] According to some embodiments, to further improve the physiological rationality of the output results during hand shape reconstruction, hand structure rationality constraints are added during both the model training and inference phases. That is, before outputting the hand shape reconstruction result, the rationality of the reconstruction result is determined based on these constraints; if reasonable, the reconstruction result is output. These constraints include joint range of motion limitations, coupling relationships between adjacent joints, consistency of phalangeal lengths, or predefined topological relationships of the hand skeleton. After outputting posture parameters, the model first checks for unreasonable postures through the constraint module, and then performs restrictions, corrections, or remapping. For example, when a joint angle exceeds a reasonable threshold range, or when certain finger joint combinations do not conform to the natural human hand movement patterns, the constraint module corrects the output. Through this implementation, even if the model deviates due to noise at local moments, the overall rationality of the final output hand shape can be guaranteed, improving system usability.
[0051] According to some embodiments, in application scenarios with high real-time requirements, the trained model can be compressed and deployed on an edge processor, wearable device main control unit, or mobile terminal. Through model pruning, quantization, distillation, or lightweight network design, the model can complete real-time inference with less computing resources. After receiving continuously sampled data from the magnetic sensing unit, the edge system (as an example of an electronic device) directly outputs the hand shape reconstruction result without relying on a remote server, meeting the requirements for low-latency interaction while maintaining low jitter and high accuracy.
[0052] The following is combined with Figure 2 This application presents an exemplary process for training a neural network model using a hand shape recovery method. This process can be executed by an electronic device or a server; this application makes no limitation on this.
[0053] like Figure 2 As shown, this exemplary process includes the following steps.
[0054] S201: Obtain the training dataset.
[0055] For example, magnetic field observation data containing multiple different magnetic sensing units and corresponding hand-shaped tags are collected.
[0056] For example, different devices include, but are not limited to, devices with different magnetic sensing unit models, magnetic source layouts, device installation positions, sampling frequencies, or hardware parameters. Data collection covers individuals with different wearing methods, different hand sizes, different types of movements (such as fingers outstretched, fist clenched, thumb against palm, index finger independently bent, etc.) and different environmental conditions.
[0057] For example, the true value of the hand shape label can be obtained through high-precision optical motion capture equipment, mechanical measuring device, manual annotation, simulation mapping or multimodal fusion, including the true value of the angle of each joint of the five fingers, the three-dimensional coordinates of key points of the hand or the posture parameters of the hand skeleton.
[0058] According to some embodiments, training data comes from multiple different magnetic sensing units. These different magnetic sensing units differ in one or more aspects, such as the number of magnetic sensing units, spatial layout, installation tolerances, sampling accuracy, and response curves. For example, data from different devices is first uniformly formatted and then mapped to a standard input structure according to channel definitions, or device difference information is represented through learnable embeddings. During training, the neural network model simultaneously receives data from multiple devices and learns common hand gesture representations through a shared backbone network; cross-device consistency constraints are introduced in the feature layer or loss layer to keep the implicit representations of the same or similar gestures close across different devices. Thus, the neural network model gains cross-device generalization ability. In actual deployment, even with changes in hardware platforms or certain assembly differences, the model can still output stable and reasonable hand shape recovery results.
[0059] S202: Construct the initial neural network model.
[0060] For example, building with Figure 1 The illustrated embodiment has an initial model with the same structure, including a temporal modeling module, a cross-finger collaborative modeling module, and a pose decoding module. For example, in the initial neural network model, the parameters of each module are randomly initialized.
[0061] S203: Using the training dataset and the objective function, train the initial neural network model to obtain the trained neural network model.
[0062] For example, the objective function includes at least the temporal smoothing loss and the hand structure consistency loss.
[0063] For example, the total loss is obtained by weighted summation of pose regression loss, temporal smoothing loss, and hand structure consistency loss. Temporal smoothing loss is used to suppress the differences between outputs at adjacent time points, and can be specifically calculated as the first-order difference norm (or second-order difference norm) between the hand shape recovery results of adjacent frames output by the pose decoding module.
[0064] For example, hand structure consistency loss is used to constrain hand topology, including at least one of the following: the deviation between the actual length of each phalanx calculated from the phalanx endpoint coordinates in the hand shape recovery result output by the pose decoding module and a preset length; and the deviation between the joint angle in the hand shape recovery result and a preset angle range. The preset length is, for example, an average human anatomical value or an individual calibration value, and the preset angle range is, for example, the range of human physiological activity.
[0065] For example, a cross-device domain alignment loss can also be added to the objective function. This loss characterizes the distributional differences between the feature representations of corresponding actions on different magnetic sensing units. For instance, maximum mean difference (MMD) or domain discriminator loss from a domain adversarial neural network can be employed. By minimizing this loss, the features learned by the model are invariant across different devices, thereby improving generalization performance on unseen new devices.
[0066] According to some embodiments, to reduce dependence on device differences, this application further employs one or more of the following strategies: converting the raw magnetic signal into a feature representation weakly correlated with absolute device parameters, such as relative changes, normalized distribution features, temporal difference features, or multi-channel coupling features; introducing multi-device mixed data during the training phase, enabling the model to learn the essential rules of hand movements shared across devices; introducing a device-independent feature extraction layer into the model to reduce device-specific noise and bias; adding a cross-device domain alignment loss to the loss function to reduce the distribution differences between the feature representations of corresponding actions on different devices; and allowing device identity as an optional input, used to assist decoupling during the training phase and to achieve explicit or implicit adaptation during the inference phase. Through these methods, it is unnecessary to re-derive fixed formulas and deep parameter tuning for each device, achieving stronger device compatibility and transferability.
[0067] To fully leverage the characteristic of AI models that continuously improve accuracy with increasing data scale, the following training mechanism can be adopted: constructing a large-scale training dataset covering a wide range of action types and user distributions; performing stratified sampling training according to device, action complexity, and user characteristics; gradually expanding the model capacity, training rounds, or feature dimensions as the number of samples increases; and introducing new device data and new action data through continuous learning or incremental training. Unlike fixed formula analysis methods, the performance improvement of this application does not rely on manually designing more complex solution formulas, but can achieve model capability improvement by increasing data and training scale, possessing good engineering scalability and long-term evolution value. After the system is put into use, data generated by new users, new action types, and new devices can be continuously collected to form an expanded training set. The expanded training set is used periodically to perform incremental training or full retraining on the original model. As the sample scale expands, the action coverage increases, and the device types become more diverse, the model accuracy continuously improves, demonstrating significant benefits from data scale expansion. This implementation method is particularly suitable for magnetic positioning products that require long-term iteration. Compared to the traditional analytical formula method, which requires manual debugging every time a new device or new scenario is added, this application can achieve natural performance growth through model updates, reducing long-term maintenance costs.
[0068] For example, data augmentation can also be performed during training, such as sensor noise perturbation, channel occlusion simulation, time-series pruning or resampling, signal amplitude scaling, and device distribution hybrid training.
[0069] For example, by iteratively optimizing and updating the model parameters, a result usable... Figure 1 The neural network model shown in the embodiment.
[0070] For example, the temporal input segments from the training dataset are fed into the initial neural network model to obtain the predicted hand shape reconstruction results. These reconstruction results are then compared with the actual hand shape labels to determine the current error. Next, the model's internal parameters are adjusted to reduce this error. This process is repeated until the model output stabilizes and the error reaches the expected level, resulting in the trained neural network model.
[0071] In this embodiment, by introducing temporal smoothing loss and hand structure consistency loss into the objective function, the stability and rationality of the model output are improved. Temporal smoothing loss forces the model to learn continuously changing motion patterns during the training phase, effectively suppressing abrupt differences between adjacent frames and ensuring the smoothness of dynamic gestures. Hand structure consistency loss incorporates prior knowledge such as finger bone length and joint range of motion to prevent the generation of deformed gestures that do not conform to physiological structures. The synergistic effect of these two losses enables the model to output smooth, realistic, and structurally accurate hand shape reconstruction results even in complex environments, significantly improving the system's robustness and generalization ability.
[0072] This application is not limited to the specific structure and process described above. The following are alternative implementation methods of this application: In addition to magnetic field data, the input data can also incorporate inertial data, pressure data, bending sensing data, or tactile data; in addition to continuous hand shape parameters, the output results can also output action category, grip type, or control command; the temporal input can use either a fixed-length window or a dynamic-length window; the model can output the hand shape for each frame or output key frame results and then interpolate and smooth; the model can be either a lightweight edge network or a high-precision cloud network; the training method can be either fully supervised training, semi-supervised training, self-supervised pre-training plus supervised fine-tuning, or teacher-student distillation training; the constraint on reasonable hand shape can be achieved through explicit loss function or implicit learning through network structure.
[0073] This application also provides a hand shape restoration system for performing the aforementioned hand shape restoration method. It is understood that this system can be deployed on electronic devices (such as wearable devices, VR controllers, and computers) or cloud servers, and this application does not impose any limitations on this.
[0074] For example, the hand shape recovery system includes the following modules.
[0075] The data acquisition module is used to perform step S101, which involves acquiring continuous multi-frame magnetic field observation data collected by the magnetic sensing unit located on the hand. Specifically, this module is communicatively connected to the magnetic sensing unit on the hand and receives multi-channel magnetic field strength values, direction components, and optional inertial data, timestamps, etc.
[0076] The preprocessing module is used to perform the function of constructing a time-series input segment based on continuous multi-frame magnetic field observation data in chronological order in step S102. Specifically, this module performs time synchronization, noise reduction, and normalization on the raw magnetic data, and uses a sliding time window to extract a fixed-length time-series input segment. The window length can be set according to the sampling frequency of the magnetic sensing unit.
[0077] The model inference module is used to perform the function of inputting the temporal input segment into the neural network model and obtaining the hand shape recovery result in step S103. This module loads a pre-trained neural network model, which includes a temporal modeling module (such as LSTM), a cross-finger collaborative modeling module (such as multi-head self-attention), and a pose decoding module (such as fully connected layers), and outputs the angles of the five finger joints or the coordinates of key points.
[0078] The result constraint module is used to apply hand structure rationality constraints, smoothing constraints, or confidence level filtering to the output hand shape restoration results. For example, it detects whether the joint angles exceed the preset physiological range and whether the finger bone lengths are consistent; if not, it corrects or discards the results. Exemplarily, this step can be used as a processing step in step S103 before obtaining the hand shape restoration results.
[0079] The output interface module is used to perform the function of sending the hand shape recovery result to the host computer, display device, control terminal or external application system in step S103 of the first embodiment, such as driving virtual hand rendering, sending robot control commands or performing rehabilitation assessment.
[0080] Furthermore, the hand restoration system may also include the following optional modules:
[0081] The training management module executes the training steps of the neural network model, such as steps S201 to S204. Specifically, this module manages a training dataset containing magnetic field observation data of multiple different magnetic sensing units and corresponding hand shape labels, constructs an initial neural network model, and trains it using an objective function that includes at least temporal smoothing loss and hand structure consistency loss, and may further introduce cross-device domain alignment loss. After training is complete, the model parameters are deployed to the model inference module.
[0082] The device adaptation module is used to process the differences in data format, channel definition, or sampling parameters of different hardware devices during the preprocessing process in step S102, so that the data acquisition module can be compatible with multiple magnetic sensing units.
[0083] The self-calibration module is used in the preprocessing step S102 to adaptively correct the input distribution based on short-term acquired data, thereby improving inference performance on new devices or new users. For example, it acquires a few seconds of static data for a newly connected device, calculates the magnetic field baseline offset, and performs normalization compensation.
[0084] This application also provides an electronic device, including a memory and a processor. The memory stores instructions executed by one or more processors of the electronic device; the processor is one of one or more processors of the electronic device, used to execute the hand shape recovery method of any of the above embodiments.
[0085] This application also provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the hand shape recovery method of any of the above embodiments.
[0086] Figure 3 According to some embodiments of this application, a hardware structure block diagram of an electronic device 30 for a hand shape recovery method is shown. Figure 3 In the illustrated embodiment, the electronic device 30 may include one or more processors 301, system control logic 302 connected to at least one of the processors 301, system memory 303 connected to the system control logic 302, non-volatile memory (NVM) 304 connected to the system control logic 302, and network interface 306 connected to the system control logic 302.
[0087] In some embodiments, processor 301 may include one or more single-core or multi-core processors. In some embodiments, processor 301 may include any combination of general-purpose processors and special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments where electronic device 30 employs an Evolved Node B (eNB) or Radio Access Network (RAN) controller, processor 301 may be configured to perform various corresponding embodiments. For example, processor 301 may be used to implement a hand-image recovery method.
[0088] In some embodiments, system control logic 302 may include any suitable interface controller to provide any suitable interface to at least one suitable device or component in processor 301 that communicates with system control logic 302.
[0089] In some embodiments, system control logic 302 may include one or more memory controllers to provide an interface to system memory 303. System memory 303 may be used to load and store vector data and / or instructions. For example, system memory 303 may load graph indexes and compressed vectors from embodiments of this application, and may also store received query vectors, etc.
[0090] In some embodiments of the electronic device 30, the system memory 303 may include any suitable volatile memory, such as suitable dynamic random access memory (DRAM).
[0091] NVM memory 304 may include one or more tangible, non-transitory computer-readable media for storing vector data and / or instructions. In some embodiments, NVM memory 304 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.
[0092] NVM memory 304 may include a portion of the storage resources on the device on which electronic device 30 is installed, or it may be accessible by the device, but is not necessarily part of the device. For example, NVM memory 304 may be accessed over a network via network interface 306.
[0093] Specifically, system memory 303 and NVM memory 304 may each include a temporary copy and a permanent copy of instruction 305. Instruction 305 may include, when executed by at least one of processors 301, causing electronic device 30 to implement hand-shaped recovery instructions as described in embodiments of this application. In some embodiments, instruction 305, hardware, firmware, and / or its software components may additionally / alternatively reside in system control logic 302, network interface 306, and / or processor 301.
[0094] Network interface 306 may include a transceiver for providing a radio interface to electronic device 30, thereby enabling communication with any other suitable device (such as a front-end module, antenna, etc.) via one or more networks. In some embodiments, network interface 306 may be integrated into other components of electronic device 30. For example, network interface 306 may be integrated into at least one of processor 301, system memory 303, NVM memory 304, and firmware device (not shown) with instructions, which, when executed by at least one of processor 301, enable electronic device 30 to implement the method as shown in the method embodiments. In embodiments of this application, network interface 306 may be used to receive unstructured vector data, query vectors, query content, etc., sent by client 300.
[0095] Network interface 306 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, network interface 306 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0096] In some embodiments, at least one of the processors 301 may be packaged together with the logic of one or more controllers for system control logic 302 to form a system in a package (SiP). In some embodiments, at least one of the processors 301 may be integrated on the same die with the logic of one or more controllers for system control logic 302 to form a system on a chip (SoC).
[0097] The electronic device 30 may further include an input / output (I / O) device 307. The I / O device 307 may include a user interface enabling a user to interact with the electronic device 30; the peripheral component interface is designed to allow peripheral components to also interact with the electronic device 30. In some embodiments, the electronic device 30 further includes a magnetic sensing unit for determining at least one of environmental conditions and location information related to the electronic device 30.
[0098] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.
[0099] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0100] In some embodiments, the magnetic sensing unit may include, but is not limited to, a gyroscope magnetic sensing unit, an accelerometer, a short-range magnetic sensing unit, an ambient light magnetic sensing unit, and a positioning unit. The positioning unit may also be part of or interact with the network interface 306 to communicate with components of the positioning network (e.g., BeiDou satellites).
[0101] Understandable Figure 3 The illustrated structure does not constitute a specific limitation on the electronic device 30. In other embodiments of this application, the electronic device 30 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented by hardware or software, or a combination of software and hardware.
[0102] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0103] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.
[0104] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0105] One or more aspects of at least one embodiment can be implemented by representational instructions stored on a computer-readable storage medium, the instructions representing various logics in a processor, which, when read by a machine, cause the machine to create logic for performing the techniques described herein. These representations, referred to as “IP cores,” can be stored on a tangible computer-readable storage medium and provided to multiple customers or production facilities for loading into manufacturing machines that actually manufacture the logic or processor.
[0106] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0107] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0108] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0109] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0110] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. A method for restoring hand shape, characterized in that, Applied to electronic devices, the method includes: Acquire continuous multi-frame magnetic field observation data collected by a magnetic sensing unit set on the hand, and construct a time-series input segment based on the continuous multi-frame magnetic field observation data in chronological order; The time-series input segment is fed into a neural network model to obtain the hand shape reconstruction result output by the neural network model. The neural network model includes a temporal modeling module, a cross-finger collaborative modeling module, and a pose decoding module. The timing modeling module is used to perform timing encoding processing based on the timing input segment and output timing feature data. The timing feature data is used to characterize the correlation between consecutive moments in the timing input segment. The cross-finger collaborative modeling module is used to perform cross-finger attention interaction processing based on the temporal feature data and output collaborative feature data, which is used to characterize the relationship between each finger. The posture decoding module is used to perform regression mapping processing based on the collaborative feature data and output the hand shape recovery result.
2. The method according to claim 1, characterized in that, The temporal modeling module is one of the following: Long Short-Term Memory Network, Gated Recurrent Unit, Temporal Convolutional Network, or Transformer Network.
3. The method according to claim 1 or 2, characterized in that, The cross-finger collaborative modeling module is a multi-head self-attention mechanism or a graph neural network.
4. The method according to claim 1, characterized in that, The construction of a time-series input segment based on the continuous multi-frame magnetic field observation data in chronological order includes: A sliding time window is used to extract a fixed-length time sequence input segment from a continuous time series. The window length of the sliding time window is set according to the sampling frequency of the magnetic sensing unit.
5. The method according to claim 1, characterized in that, The hand shape restoration results include the angle values of each joint of the five fingers or the three-dimensional spatial coordinates of key points of the hand.
6. The method according to claim 1, characterized in that, Also includes: Obtain a training dataset containing magnetic field observation data and corresponding hand shape labels from multiple different magnetic sensing units; An initial neural network model is constructed, which includes a temporal modeling module, a cross-finger collaborative modeling module, and a pose decoding module; The initial neural network model is trained using the training dataset and the objective function to obtain the neural network model. The objective function includes at least a temporal smoothing loss and a hand structure consistency loss. The temporal smoothing loss is used to suppress the difference between outputs at adjacent time points, and the hand structure consistency loss is used to constrain the hand topology.
7. The method according to claim 6, characterized in that, The objective function also includes a cross-device domain alignment loss, which is used to characterize the distribution differences between the feature representations of corresponding actions on different magnetic sensing units.
8. The method according to claim 6 or 7, characterized in that, The temporal smoothing loss is the first-order difference norm between the hand shape recovery results of adjacent frames output by the attitude decoding module; The loss of hand structural consistency includes at least one of the following: Based on the finger bone endpoint coordinates in the hand shape recovery result output by the posture decoding module, the actual length of each finger bone calculated, and the deviation between the actual length and the preset length; as well as The deviation between the joint angle in the hand shape restoration result and the preset angle range.
9. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of the electronic device, and a processor, being one of one or more processors of the electronic device, for performing the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1-8.