Method for Handwritten Letter Input on the Back of a Smartphone by Integrating IMU and UWB
Through the smartphone back handwritten letter input method that integrates IMU and UWB data, and using LSTM and TCN models for handwriting recognition, the problem of limited interaction content and hand occlusion on the back of the smartphone is solved, and the effective recognition and extended interaction method of dynamic handwritten input is realized.
Patent Information
- Application Number
- CN202510369996.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the prior art, the human-computer interaction method on the back of the smartphone has limited interaction content, and there are problems such as hand covering the screen, finger covering the space, and watch input inconveniently.
By integrating the inertial measurement unit (IMU) and ultra-wideband ranging (UWB) data of smartphones and smart watches, a communication framework is built, data preprocessing and feature extraction is performed, and handwriting letter recognition is achieved by using long and short-term memory network (LSTM) and time convolution network (TCN) models to realize handwriting input on the back of the smartphone.
It realizes dynamic recognition of handwritten letters on the back of the smartphone, avoids the screen blocking of the hand, expands the interaction method, solves the problem of inconvenient watch input, and provides a better operating experience.
Smart Images

Figure CN119883094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction, and particularly to a method for inputting handwritten letters on the back of a smartphone that integrates IMU and UWB. Background Art
[0002] Human-Computer Interaction (HCI) is an interdisciplinary field that studies the information transfer and collaborative work between humans and computer systems. Its purpose is to improve the operation efficiency, comfort, and experience of users through optimized technology design. It integrates knowledge from fields such as computer science, psychology, design, and cognitive science. The core goal is to make technology more in line with natural human behavior patterns and reduce the usage threshold.
[0003] Common implementation methods of human-computer interaction include: 1. Graphical User Interface (GUI), which is based on visual elements (windows, icons, menus, buttons, etc.). Users directly operate through devices such as mice and touchscreens. For example, the touch interaction of smartphones and the visual operation of desktop systems reduce the difficulty of using technology. 2. Voice interaction, which uses Natural Language Processing (NLP) technology to achieve voice conversations between humans and machines, such as smart speakers and in-vehicle voice assistants. Users control devices through voice commands, which is suitable for multi-task scenarios or barrier-free requirements. 3. Gesture and somatosensory interaction, which captures user actions through cameras and sensors, such as VR controllers, Microsoft Kinect, or gesture recognition screens. Users can interact with the virtual environment by waving or moving their bodies, enhancing the sense of immersion. 4. Virtual and Augmented Reality (VR / AR), where VR constructs a virtual space through a head-mounted display device, and users can explore freely; AR overlays digital information onto the real world and is applied in fields such as education and healthcare to expand the interaction dimension. 5. Wearable devices, such as smartwatches and AR glasses, achieve seamless interaction through biosensors (heart rate, motion monitoring) or environmental perception, emphasizing real-time and convenience. 6. Brain-Computer Interface (BCI), a cutting-edge technology that directly controls external devices by interpreting brain electrical signals, helps disabled people or enables more efficient command transmission. Although it is still in the experimental stage, it has great potential.
[0004] The most common scenario of human-computer interaction is the interaction between humans and smartphones. The implementation methods that can be adopted for the interaction between humans and smartphones are very diverse, and almost all of the above-introduced common human-computer interaction methods can be applied in the interaction between humans and mobile phones. The diverse interaction methods make the interaction between humans and mobile phones very flexible, and the human-computer interaction experience is also very excellent. However, various interaction methods applied above are often face-to-face, that is, frontal interactions.
[0005] If it is possible to achieve back interaction between humans and mobile phones, it will surely greatly enrich the usage scenarios of human-computer interaction and bring more human-computer interaction methods. In addition, there are the following benefits to human-computer interaction on the back of the mobile phone: 1. Enrich operation types: For example, the same upward sliding operation can be designed with different functions according to different application requirements on the front and back of the mobile phone, and it can also be applied to shortcut keys. 2. Protect personal privacy: When operating on the front of the mobile phone in public, others may guess the content you input based on the input gestures, while there will be no such problem with back operation. 3. Increase the visible area: Placing the key input on the back can effectively reduce the reduction of the visible area caused by the occlusion of the keyboard and hands. 4. Improve single-handed operation: For users of large-screen mobile phones, the back touch area is easier to reach during single-handed operation (such as the index finger naturally placed on the back of the mobile phone), reducing the fatigue of excessive thumb extension.
[0006] Regarding the human-computer interaction method on the back of the mobile phone, there are also some current studies. For example, by some specific gestures (such as knocking on the back of the mobile phone with a knuckle), some shortcut functions of the mobile phone can be awakened. However, this interaction method is too simple and the content that can be interacted with is very limited. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems existing in the related art to some extent.
[0008] The purpose of the present invention is to provide a method for inputting handwritten letters on the back of a smart phone that integrates IMU and UWB, realizing handwritten recognition and input on the back of the smart phone, expanding new human-computer interaction methods, and enhancing the interaction experience.
[0009] To achieve the above purpose, on the one hand, the present invention provides a method for inputting handwritten letters on the back of a smart phone that integrates IMU and UWB, including:
[0010] S1. Construct a communication framework between the smart phone and the smart watch, and obtain the IMU sequence data and UWB ranging data in the two devices through the API interfaces of the smart phone and the smart watch;
[0011] S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices;
[0012] S3. Analyze the influence of different features in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization processing;
[0013] S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data;
[0014] S5. Load the trained handwritten letter recognition learning model into the smartphone and / or smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and the smartwatch.
[0015] A further preferred technical solution of the present invention is that in step S2, the IMU sequence data and UWB ranging data in the two devices are preprocessed to align the data in space and time, and the relative motion characteristics of the two devices are obtained. The specific method is as follows:
[0016] S21. Truncate the head and tail of the IMU sequence data of the smartphone and the smartwatch, and interpolate the UWB ranging data to align the time of the two IMU sequence data and the UWB ranging data;
[0017] S22. Set a global coordinate system and map the obtained attitude information of the smartphone and the smartwatch to this coordinate system;
[0018] S23. Construct a set of relative motion characteristics of the two devices, including UWB ranging data, various relative motion characteristics of the smartphone and the smartwatch in their respective local coordinate systems, and various relative motion characteristics of the smartwatch in the coordinate system of the smartphone.
[0019] Preferably, in step S21, the method of truncating the head and tail of the IMU sequence data of the smartphone and the smartwatch is as follows:
[0020] Cut the IMU sequence data of the smartphone and the smartwatch, and only keep the data between the previous timestamp of the first timestamp corresponding to the UWB ranging data and the next timestamp of the last timestamp of the UWB ranging data;
[0021] In step S21, the method of aligning the time of the two IMU sequence data is as follows:
[0022] Let the update frequency of the IMU sequence data of the smartphone and the smartwatch be , when the time interval between the timestamps of the IMU sequence data of the smartphone and the timestamps of the IMU sequence data of the smartwatch is less than , take the middle moment of the two timestamps as the alignment moment, and the corresponding data are the IMU sequence data of the smartphone and the smartwatch adjacent to the alignment moment respectively.
[0023] Preferably, in step S21, the specific method of interpolating the UWB ranging data to align the time of the two IMU sequence data and the UWB ranging data is as follows:
[0024] Let the update frequency of the UWB ranging data be , based on the timestamps of the IMU sequence data of the smartwatch, determine the timestamps of the IMU sequence data of the smartwatch adjacent to a certain timestamp of the UWB ranging data, and combine the two with the IMU sequence data of the smartphone at the corresponding timestamp to form a data group of timestamps;
[0025] At this time, there is no corresponding UWB ranging data for the IMU sequence data of some timestamps. Use the UWB ranging data of two adjacent IMU sequence data with corresponding data for piecewise linear interpolation to fill in the missing UWB ranging data.
[0026] Preferably, in step S22, a global coordinate system is set, and the attitude information of the smartphone and the smartwatch obtained is mapped to this coordinate system. The specific method is as follows:
[0027] First, set a global coordinate system, with its X-axis pointing to the geographic north pole, the Z-axis perpendicular to the ground and upward, and the Y-axis obtained by the right-hand rule;
[0028] Perform the operation of taking the inverse matrix of the rotation matrix corresponding to the smartwatch and then multiplying it by the rotation matrix corresponding to the smartphone to obtain the rotation matrix from the smartwatch coordinate system to the smartphone coordinate system, expressed as:
[0029]
[0030] Among them, is the rotation matrix from the directly obtained global coordinate system to the smartphone coordinate system, is the rotation matrix from the directly obtained global coordinate system to the smartwatch coordinate system, is the rotation matrix from the calculated smartwatch coordinate system to the smartphone coordinate system; T is the transpose operation of the matrix, which is equivalent to taking the inverse matrix.
[0031] Preferably, in the set of relative motion features of the two devices constructed in step S23, some features are directly obtained from the IMU sequence data and the UWB ranging data, and some features are calculated and obtained by combining the data groups at each timestamp with the rotation matrix from the smartwatch coordinate system to the smartphone coordinate system.
[0032] Preferably, in step S3, analyze the influence of different features among the relative motion features of the two devices on the recognition accuracy of handwritten letters; the specific method is as follows:
[0033] S31. Conduct a preliminary experiment to train different models for different features. During the training, use 26 letters as labels, and use one-hot encoding for the 26 labels. Use the categorical cross-entropy as the loss function and use the early stopping method for training;
[0034] S32. Select the required features based on the performance of different models trained with different features in the validation set and the training set before overfitting. During the training of the early stopping method, when a new minimum loss value appears, if no new minimum loss smaller than the previous one appears within 3 rounds of training, it is considered overfitting.
[0035] Preferably, in step S3, the specific method for selecting the required feature combination and performing normalization processing is as follows:
[0036] S33. Select the UWB distance, the acceleration of the smartwatch in the smartphone coordinate system, and the rotation matrix from the smartwatch to the smartphone as the feature combination;
[0037] S34. Fill the sequence length of each sample with 0.0 to 500;
[0038] S35. Use MinMaxScaler in the sklearn.preprocessing module of Python to normalize the data. The formula is as follows:
[0039]
[0040] Where X is the current value of a certain feature, is the minimum value of this feature in the current sample, is the maximum value of this feature in the current sample; is the range to which the data is restricted defined by the user, is its maximum value, is its minimum value, and the default range is (0,1).
[0041] Preferably, the network architecture of the handwritten letter recognition learning model based on LSTM and TCN constructed in step S4 includes:
[0042] Masking layer, used to mask the filled 0.0 in the sample; assuming the sample is , whose shape is (batch_size, timesteps, features), the Masking layer generates a mask matrix M by checking whether the value of each time step is equal to a specified specific value. The shape of this matrix is (batch_size, timesteps), and its elements are:
[0043]
[0044] Where i is the sample index in the batch and j is the time step index;
[0045] A bidirectional TCN network for data convolution; the size of the convolution kernel is set to 6, the number of filters used in the convolutional layer is 128, and skip residual connections are used in each layer; it is set that there is no causality during convolution, that is, the data of the previous timestamp is not used to predict the data of the subsequent timestamp; the dilation list is set to [1, 2, 4, 8]; it is set to return the data of all timestamps; finally, the proportion of the lost input data of the previous layer is set to 0.1;
[0046] A unidirectional LSTM network for transforming the sequence of data into 128 dimensions and outputting the data of the last timestamp;
[0047] Two fully connected layers, where the number of units in the first fully connected layer is 64, and the activation function is ReLU. The calculation formula is:
[0048]
[0049] Among them, is the input value after linear transformation, which is obtained by multiplying the 1×128 data output by the unidirectional LSTM network by a 128×64 weight matrix and then adding a 1×64 bias term;
[0050] The number of units in the second fully connected layer is 26, and the activation function is Softmax. The calculation formula is:
[0051]
[0052] Among them, is the data of the th row of the column vector obtained by multiplying the 1×64 data output by the first fully connected layer by a 64×26 weight matrix and then adding a 1×26 bias term, is the total number of categories, is the index, is the th row of the data corresponding to the index.
[0053] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions enable the computer to execute the above-mentioned method for inputting handwritten letters on the back of a smartphone by fusing IMU and UWB.
[0054] On yet another aspect, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus, and the processor calls the logical instructions in the memory to execute the above-mentioned method for inputting handwritten letters on the back of a smartphone by fusing IMU and UWB.
[0055] In another aspect, the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the above-mentioned method for inputting handwritten letters on the back of a smart phone by integrating IMU and UWB.
[0056] Beneficial effects: The method for inputting handwritten letters on the back of a smart phone by integrating IMU and UWB according to the present invention extracts the IMU and UWB sensor data from commercial mobile phones and wearable watches, fuses the two data, and combines the LSTM and TCN network modules in deep learning to achieve a new interaction method for data input on the back of the mobile phone, thereby realizing an extended way of human-computer interaction. It can effectively avoid the situation where when writing by hand on the front, the hand or finger blocks part of the screen, resulting in discontinuous browsing. When multiple swipes or pinyin input are required, the left and right fingers crowd each other for space. In addition, it solves the problem that the input keyboard of the watch is relatively small compared to the fingers and is not convenient for input. Finally, it can play a role when the front screen of the mobile phone has poor contact or the fingers are stained and it is not convenient to wipe after writing by hand on the front.
[0057] The present invention realizes lightweight dynamic handwritten letter recognition. Compared with traditional handwriting recognition (HWR), which relies on a computer to analyze and process the image of handwritten text, the present invention does not require image recognition, does not use a camera to obtain user privacy, dynamically accepts and processes data, and realizes dynamic handwritten recognition. Brief Description of the Drawings
[0058] Figure 1 It is the overall flowchart of the method for inputting handwritten letters on the back of a smart phone by integrating IMU and UWB according to the present invention;
[0059] Figure 2 It is the confusion matrix of the model on the training set when training stops in Embodiment 1 of the present invention;
[0060] Figure 3 It is the confusion matrix of the model on the test set in Embodiment 1 of the present invention. Detailed Embodiments
[0061] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.
[0062] Before elaborating on the specific embodiments of the present invention in detail, it is necessary to explain some technical terms involved in the present invention.
[0063] IMU is the abbreviation of Inertial Measurement Unit. It is an electronic device used to measure and report the three basic linear motions (accelerations) and three basic angular motions (angular velocities) of an object.
[0064] UWB ranging (Ultra-Wideband ranging) is a high-precision distance measurement method based on ultra-wideband wireless communication technology. Its core principle is to determine the distance by calculating the flight time (TOF) of the signal between two devices multiplied by the speed of light. For example, device A sends a ranging request and records the transmission timestamp (T1), device B records the timestamp (T2) after receiving it and replies with a response, and device A then records the timestamp (T3) of receiving the response. By calculating the time difference (such as T3 - T1 - T2 - T1), the round-trip time of the signal is obtained, and then the one-way flight time is calculated.
[0065] API interface (Application Programming Interface) is a set of predefined rules, protocols, and tools for the interaction between different software systems or components. It defines how to request services, transfer data, and obtain results, while hiding the underlying implementation details, enabling developers to call functions without caring about the internal logic.
[0066] LSTM (Long Short-Term Memory) is a deep learning model commonly used to process sequential data. Compared with the traditional RNN (Recurrent Neural Network), LSTM introduces three gates (input gate, forget gate, output gate) and a cell state, and these mechanisms enable LSTM to better handle long-term dependencies in sequences.
[0067] TCN (Temporal Convolutional Network) is a deep learning model specifically designed for processing sequential data. It combines the parallel processing capabilities of convolutional neural networks (CNNs) and the long-term dependency modeling capabilities of recurrent neural networks (RNNs), making it a powerful tool for sequence modeling tasks.
[0068] Handwriting Recognition (HWR) refers to the technology by which a computer analyzes and processes images of handwritten text and converts it into machine-readable text or data.
[0069] The following combines Figures 1 - 3 to describe the method for inputting handwritten letters on the back of a smartphone that fuses IMU and UWB provided by the present invention.
[0070] Embodiment 1: This embodiment provides a method for inputting handwritten letters on the back of a smartphone that fuses IMU and UWB. The overall process of this method is as Figure 1 shown and includes:
[0071] S1. Construct a communication framework between the smartphone and the smartwatch, and obtain the IMU sequence data and UWB ranging data in the two devices through the API interfaces of the smartphone and the smartwatch;
[0072] S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices;
[0073] S3. Analyze the influence of different features in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization processing;
[0074] S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data;
[0075] S5. Load the trained handwritten letter recognition learning model into the smartphone and / or smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and the smartwatch.
[0076] The following elaborates on each step in detail.
[0077] In step S1, the Nearby Interaction and CoreMotion frameworks are used to obtain the UWB ranging data between the smartphone and the smartwatch, as well as the IMU sequence data between the smartphone and the smartwatch. The specific method is as follows:
[0078] S11. The smartphone and the smartwatch respectively create a WCSession object, and respectively use the activate() method of the object to activate the connection between the smartphone and the smartwatch. The isReachable property can be used to determine whether the connection is successful.
[0079] S12. The smartphone and the smartwatch respectively create a NISession object, respectively obtain their corresponding discoveryTokens, and use the sendMessage method of the WCSession object to send the discoveryTokens of the NISession to each other. After successful reception, configure NIConfiguration and set the value of isCameraAssistanceEnabled to false. Use the run method of the NISession object to start UWB ranging, and obtain the distance through the distance object of NINearbyObject. The principle of obtaining the distance is TOF, and the calculation formula is as follows:
[0080]
[0081] d is the calculated distance, is the timestamp of the transmitted pulse, is the timestamp of the received reflected pulse, and c is the speed of light.
[0082] S13. The smartphone and the smartwatch respectively create a CMDeviceMotion object, and set the deviceMotionUpdateInterval property of the device update frequency to 0.02 seconds. Use the startDeviceMotionUpdates method to obtain the acceleration of the three axes, the angular velocity of the three axes, the Euler angles corresponding to the attitude information, and the rotation matrix (set the global frame to xTrueNorthZVertical).
[0083] S14. Write the data into phone_uwb.txt, phone_imu.txt, watch_uwb.txt, and watch_imu.txt.
[0084] The preprocessing of the IMU sequence data and UWB ranging data in the two devices described in step S2 is to align the data in space and time and obtain the relative motion characteristics of the two devices. The specific method is as follows:
[0085] S21. Since the IMU sequence data starts to be collected before the UWB ranging data begins and ends after the UWB ranging data is obtained, it is necessary to cut the IMU sequence data of the smartphone and the smartwatch, and only retain the data between the previous timestamp of the first timestamp of the corresponding UWB ranging data and the next timestamp of the last timestamp of the UWB ranging data;
[0086] Interpolate the UWB ranging data to align the time of the two IMU sequence data with the UWB ranging data. The method is as follows: Taking the timestamp of the IMU sequence data of the watch as a reference, taking one of the timestamps A as an example, since the update frequencies of the IMUs of the mobile phone and the watch are both 0.02 seconds, if there is a timestamp B of the IMU sequence data of the mobile phone and the time interval between A and B is less than 0.01 seconds, then take the middle moment of A and B, and the corresponding data comes from the data of the mobile phone and the watch corresponding to A and B respectively. For the UWB ranging data, taking the timestamp of the IMU sequence data of the watch as a reference, see which two adjacent watch IMU timestamps the timestamp of the UWB ranging data falls between, and judge which timestamp it is closer to, and combine the two with the IMU sequence data of the corresponding timestamp of the smartphone to form a data group with a timestamp.
[0087] Since the update frequency of the UWB ranging data is about 0.15 seconds and the update frequency of the watch IMU sequence is 0.02 seconds, some IMU sequence data do not have corresponding UWB ranging data to match. Use the evenly divided linear interpolation of two adjacent UWB ranging data with corresponding IMU sequence data to fill in the missing UWB ranging data.
[0088] S22. Set up a global coordinate system and map the obtained attitude information of the smartphone and the smartwatch to this coordinate system. The specific method is as follows:
[0089] First, set up a global coordinate system, with its X-axis pointing to the geographic north pole, the Z-axis perpendicular to the ground upward, and the Y-axis obtained by the right-hand rule;
[0090] Perform the operation of taking the inverse matrix of the rotation matrix corresponding to the smartwatch and then multiplying it by the rotation matrix corresponding to the smartphone to obtain the rotation matrix from the smartwatch coordinate system to the smartphone coordinate system, expressed as:
[0091]
[0092] Among them, is the rotation matrix from the global coordinate system to the smartphone coordinate system directly obtained, is the rotation matrix from the global coordinate system to the smartwatch coordinate system directly obtained, R is the rotation matrix from the calculated smartwatch coordinate system to the smartphone coordinate system; T is the transpose operation of the matrix, which is equivalent to taking the inverse matrix.
[0093] S23. Construct the relative motion feature set of the two devices, including UWB ranging data, various relative motion features of the smartphone and the smartwatch in their respective local coordinate systems, and various relative motion features of the smartwatch in the smartphone coordinate system. In the relative motion feature set of the two devices constructed here, some features are directly obtained from the IMU sequence data and UWB ranging data, and some features are calculated by combining the data of each timestamp with the rotation matrix from the smartwatch coordinate system to the smartphone coordinate system. For example:
[0094] Multiply the acceleration of the smartwatch by the previous rotation matrix to obtain the acceleration of the smartwatch in the smartphone coordinate system.
[0095] Subtract the acceleration of the smartphone in its own coordinate system from the acceleration of the smartwatch in the smartphone coordinate system to obtain the acceleration of the smartwatch relative to the smartphone. Expressed as:
[0096]
[0097]
[0098] In the above formula is the acceleration of the smartphone in its own coordinate system, is the calculated acceleration of the smartwatch in the smartphone coordinate system, is the acceleration of the smartphone in its own coordinate system, a is the calculated acceleration of the smartwatch in the smartphone coordinate system.
[0099] Integrate the acceleration in the watch local coordinate system and the acceleration in the phone coordinate system respectively to obtain the velocity, and the initial velocity is set to 0. Normalize the three-dimensional velocity to obtain the direction of motion. Expressed as:
[0100]
[0101]
[0102] The specific method of step S3 is as follows:
[0103] S31. Conduct a pre-experiment. For different models trained with different features, 26 letters are used as labels during training. The 26 labels are one-hot encoded, and the categorical cross-entropy is used as the loss function. The early stopping method is used for training;
[0104] S32. Select the required features based on the performance of different models trained with different features in the validation set and the training set before overfitting. During the process of training data using early stopping, when a new minimum loss value appears, if no new minimum loss smaller than the previous one appears within 3 rounds of training, it is considered overfitting.
[0105] S33. Select the UWB distance, the acceleration of the smartwatch in the smartphone coordinate system, and the rotation matrix from the smartwatch to the smartphone as the feature combination;
[0106] S34. Fill the sequence length of each sample with 0.0 to 500;
[0107] S35. Normalize the data using MinMaxScaler in the sklearn.preprocessing module of Python. The formula is as follows:
[0108]
[0109] where X is the current value of a certain feature, is the minimum value of this feature in the current sample, is the maximum value of this feature in the current sample; is the range to which the data is restricted defined by the user, is its maximum value, is its minimum value, and the default range is (0, 1).
[0110] The handwritten letter recognition learning model based on LSTM and TCN constructed in step S4 has a network architecture including:
[0111] Masking layer. Use the Masking layer in keras of tensorflow to mask the filled 0.0 in the sample. The Masking layer will detect a specific value (usually 0) in the input tensor and mark it as "invalid" or "masked". These marks will be passed to the subsequent model layers to tell them which time steps should be ignored.
[0112] Assume the sample is , whose shape is (batch_size, timesteps, features). The Masking layer generates a mask matrix M by checking whether the value of each time step is equal to the specified specific value. The shape of this matrix is (batch_size, timesteps), and its elements are:
[0113]
[0114] where i is the sample index in the batch and j is the time step index;
[0115] A bidirectional TCN network is used for data convolution; the size of the convolutional kernel is set to 6, and the number of filters used in the convolutional layer is 128. Here, the number of filters is equivalent to the number of cells in the LSTM; skip residual connections are used in each layer; it is set that there is no causal relationship during convolution, that is, data from previous timestamps are not used to predict data from subsequent timestamps; the dilation list is set to [1, 2, 4, 8], which means whether the timestamp data used during convolution is adjacent, or separated by 1, 3, or 7; it is set to return data for all timestamps; finally, the proportion of lost input data from the previous layer is set to 0.1;
[0116] A unidirectional LSTM network is used to transform the sequence of data into 128 dimensions and output the data of the last timestamp; the principle of the LSTM is expressed as:
[0117] cell state
[0118] hidden state
[0119] cell input
[0120] input gate
[0121] forget gate
[0122] output gate
[0123] Weight vectors , , , are related to the relationships with cell input, input gate, forget gate, and output gate. The weights , , , , are related to the relationships with cell input, input gate, forget gate, and output gate. , , , , are bias terms. and are activation functions (usually tanh) for cell input and hidden state. It is used to normalize or compress the unit state. The activation function of all gates is sigmod, 1 / (1 + exp(-x)). By introducing a gating mechanism, LSTM effectively controls the flow of information. The forget gate controls the discarding of information, the input gate controls the addition of new information, and the output gate controls the output of the hidden state. In this way, LSTM can maintain important information for a long time when processing sequential data, avoiding the vanishing gradient problem of the standard RNN.
[0124] Two fully connected layers, where the number of units in the first fully connected layer is 64, and the activation function is ReLU. The calculation formula is:
[0125]
[0126] where, is the input value after linear transformation, which is obtained by multiplying the 1×128 data output by the unidirectional LSTM network by a 128×64 weight matrix and then adding a 1×64 bias term; the output of ReLU is zero or the input itself, and it is usually used in deep networks, especially in convolutional neural networks (CNNs).
[0127] The number of units in the second fully connected layer is 26, and the activation function is Softmax. The calculation formula is:
[0128]
[0129] where, is the data of the th row of the column vector obtained by multiplying the 1×64 data output by the first fully connected layer by a 64×26 weight matrix and then adding a 1×26 bias term, is the total number of categories, is the index, is the th row of the data corresponding to the index; Softmax maps each output to between [0, 1], and the sum of all outputs is 1, making them interpretable as probabilities.
[0130] To verify the method of this embodiment, the following experiments are carried out:
[0131] Use the smartphone iPhone12 with the system ios18.2; use the smartwatch iWatch seris9 with the system watchOS 11.2.
[0132] Keep the index finger, middle finger, and ring finger of the left hand together and press them against the back of the mobile phone, then hold the phone. Use the thumb of the left hand to click on the screen to start UWB connection or click end after writing the letter. Place the right hand in the area between the index finger of the left hand and the lower frame edge of the mobile phone camera, and write the letter with the middle finger of the right hand.
[0133] Run the experiment using Keras 2.14.0. For the entire model, set the learning rate to 0.0001 and use Adam as the optimizer. Use early stopping to train the data. When a new minimum loss value appears, if no new minimum loss smaller than the previous one appears within 3 rounds of training, it is considered overfitting.
[0134] There are 1040 samples in the training data, with 40 samples for each of the 26 letters. Set 8 orientations when writing by hand, namely east, south, west, north, southeast, southwest, northwest, and northeast. When writing in each orientation, the angle between the mobile phone and the horizontal plane is 0°, 30°, 45°, 75°, 90°. Overfitting occurred during the 93rd round of training of the training data, with 100 samples used in each round of training. The final accuracy on the training set is 98.8%. Specifically, as Figure 2 shown, there is one sample each for A, D, N, T, Y, Z that was not successfully trained, and 3 samples each for G and Q that were not successfully trained.
[0135] There are 208 test samples, with 8 samples for each of the 26 letters. Set 8 angles: 30° east-south, 60° east-south, 30° south-west, 60° south-west, 30° west-north, 60° west-north, 30° north-east, 60° north-east. The corresponding angles between the mobile phone and the horizontal plane are 0°, 90°, 30°, 45°, 45°, 30°, 75°, 0°. The specific prediction situation is as Figure 3 shown, and the accuracy on the test set is 85.57%.
[0136] In addition to this experiment, the present invention conducts input tests on multiple different smartphones and smart watches. The results show that this method can achieve effective recognition of uppercase letter input, and at the same time, the recognition rate is fast enough to meet the real-time interaction requirements.
[0137] Embodiment 2: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored. These computer instructions cause the computer to execute the method for inputting handwritten letters on the back of a smartphone by integrating IMU and UWB. The method includes the following steps:
[0138] S1. Build a communication framework for smartphones and smart watches, and obtain IMU sequence data and UWB ranging data in the two devices through the API interfaces of smartphones and smart watches;
[0139] S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices;
[0140] S3. Analyze the influence of different features in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization processing;
[0141] S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data;
[0142] S5. Load the trained handwritten letter recognition learning model into the smartphone and / or smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and smartwatch.
[0143] Embodiment 3: This embodiment provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The processor can call the logical instructions in the memory to execute the method for inputting handwritten letters on the back of a smartphone by fusing IMU and UWB. The method includes the following steps:
[0144] S1. Construct a communication framework for the smartphone and the smartwatch, and obtain the IMU sequence data and UWB ranging data in the two devices through the API interfaces of the smartphone and the smartwatch;
[0145] S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices;
[0146] S3. Analyze the influence of different features in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization processing;
[0147] S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data;
[0148] S5. Load the trained handwritten letter recognition learning model into the smartphone and / or smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and smartwatch.
[0149] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0150] Embodiment 4: A computer program product provided in this embodiment includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for inputting handwritten letters on the back of a smartphone by integrating an IMU and UWB. The method includes the following steps:
[0151] S1. Construct a communication framework between the smartphone and the smartwatch, and obtain the IMU sequence data and UWB ranging data in the two devices through the API interfaces of the smartphone and the smartwatch;
[0152] S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices;
[0153] S3. Analyze the influence of different features in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization processing;
[0154] S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data;
[0155] S5. Load the trained handwritten letter recognition learning model into the smartphone and / or smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and the smartwatch.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for inputting handwritten letters on the back of a smartphone integrating IMU and UWB, characterized in that, Including: S1. Construct a communication framework between the smartphone and the smartwatch, and obtain the IMU sequence data and UWB ranging data in the two devices through the API interfaces of the smartphone and the smartwatch; S2. Preprocess the IMU sequence data and UWB ranging data in the two devices to align the data in space and time, and obtain the relative motion characteristics of the two devices; Specifically: S21. Truncate the head and tail of the IMU sequence data of the smartphone and the smartwatch, and interpolate the UWB ranging data to align the time of the two IMU sequence data and the UWB ranging data; S22. Set up a global coordinate system and map the obtained attitude information of the smartphone and the smartwatch to this coordinate system; S23. Construct a set of relative motion characteristics of the two devices, including UWB ranging data, various relative motion characteristics of the smartphone and the smartwatch in their respective local coordinate systems, and various relative motion characteristics of the smartwatch in the coordinate system of the smartphone; S3. Analyze the influence of different characteristics in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters, and then select the required feature combination accordingly for normalization; In step S3, analyze the influence of different characteristics in the relative motion characteristics of the two devices on the recognition accuracy of handwritten letters; specifically: S31. Conduct a preliminary experiment. For different models trained with different characteristics, use 26 letters as labels during training. The 26 labels are encoded using one-hot encoding, use the categorical cross-entropy as the loss function, and use the early stopping method for training; S32. Select the required features according to the performance of different models trained with different characteristics in the validation set and the training set before overfitting. During the training of the early stopping method, when a new minimum loss value appears, if no new minimum loss smaller than the former appears within 3 rounds of training, it is considered overfitting; S4. Construct a handwritten letter recognition learning model based on LSTM and TCN, and train it using the feature combination extracted from the IMU sequence data and UWB ranging data; S5. Load the trained handwritten letter recognition learning model into the smartphone and the smartwatch, and recognize the handwritten letters input on the back of the smartphone according to the real-time IMU sequence data and UWB ranging data of the smartphone and the smartwatch.
2. The method for inputting handwritten letters on the back of a smartphone integrating an IMU and UWB according to claim 1, wherein In step S21, the method for truncating the head and tail of the IMU sequence data of the smartphone and the smartwatch is: Cut the IMU sequence data of the smartphone and the smartwatch, and only retain the data between the previous timestamp corresponding to the first timestamp of the UWB ranging data and the next timestamp of the last timestamp of the UWB ranging data; In step S21, the method for aligning the time of the two IMU sequence data is: Set the update frequency of the IMU sequence data of the smart phone and the smart watch to be , when the time interval between the time stamps of the IMU sequence data of the smart phone and the time stamps of the IMU sequence data of the smart watch is less than , take the middle moment of the two time stamps as the alignment moment, and the corresponding data are the IMU sequence data of the smart phone and the smart watch close to the alignment moment respectively.
3. The method for inputting handwritten letters on the back of a smartphone integrating an IMU and UWB according to claim 2, characterized in that, In step S21, the specific method for interpolating the UWB ranging data to align the time of the two IMU sequence data and the UWB ranging data is: Let the update frequency of UWB ranging data be , and based on the timestamp of the IMU sequence data of the smartwatch, determine the timestamp of the IMU sequence data of the smartwatch adjacent to a certain timestamp of the UWB ranging data, and combine the two with the IMU sequence data of the corresponding timestamp of the smartphone to form a data group with a timestamp; At this time, there is no corresponding UWB ranging data for some of the IMU sequence data of the timestamps. The adjacent two UWB ranging data with corresponding IMU sequence data are used for equal linear interpolation to fill in the missing UWB ranging data.
4. The method for inputting handwritten letters on the back of a smartphone integrating an IMU and UWB according to claim 3, wherein In step S22, a global coordinate system is set, and the attitude information of the obtained smart phone and smart watch is mapped to this coordinate system. The specific method is as follows: First, a global coordinate system is set, with its X-axis pointing to the geographical north pole, its Z-axis perpendicular to the ground and upward, and its Y-axis obtained by the right-hand rule; The inverse matrix of the rotation matrix corresponding to the smart watch is taken and then multiplied by the rotation matrix corresponding to the smart phone to obtain the rotation matrix from the smart watch coordinate system to the smart phone coordinate system, expressed as: ; Among them, is the rotation matrix from the directly obtained global coordinate system to the smartphone coordinate system, is the rotation matrix from the directly obtained global coordinate system to the smartwatch coordinate system, is the rotation matrix from the calculated smartwatch coordinate system to the smartphone coordinate system; T is the transpose operation of the matrix, which is equivalent to taking the inverse matrix.
5. The method for inputting handwritten letters on the back of a smartphone integrating IMU and UWB according to claim 4, characterized in that, In the relative motion feature set of the two devices constructed in step S23, some features are directly obtained from the IMU sequence data and UWB ranging data, and some features are calculated by combining the data of each timestamp with the rotation matrix from the smart watch coordinate system to the smart phone coordinate system.
6. The method for inputting handwritten letters on the back of a smartphone integrating IMU and UWB according to claim 1, wherein In step S3, the specific method for selecting the required feature combination and performing normalization processing is as follows: S33. Select the UWB distance, the acceleration of the smart watch in the smart phone coordinate system, and the rotation matrix from the smart watch to the smart phone as the feature combination; S34. Fill the sequence length of each sample with 0.0 to 500; S35. Use MinMaxScaler in the sklearn.preprocessing module of Python to normalize the data. The formula is as follows: ; Wherein, X is the current value of a certain feature, is the minimum value of the feature in the current sample, is the maximum value of the feature in the current sample; is the range to which the data is restricted defined by the user, is its maximum value, is its minimum value.
7. The method for inputting handwritten letters on the back of a smartphone integrating IMU and UWB according to claim 6, wherein The handwritten letter recognition learning model based on LSTM and TCN constructed in step S4 has a network architecture including: Masking layer, used to mask the filled 0.0 in the sample; assuming the sample is , whose shape is (batch_size, timesteps, features). The Masking layer generates a mask matrix M by checking whether the value of each time step is equal to a specified specific value. The shape of this matrix is (batch_size, timesteps), and its elements are: ; Among them, i is the sample index in the batch, and j is the time step index; A bidirectional TCN network for data convolution; the size of the convolution kernel is set to 6, the number of filters used in the convolutional layer is 128, and each layer uses a skip residual connection; it is set that there is no causal relationship during convolution, that is, the data of the previous timestamp is not used to predict the data of the next timestamp; the dilation list is set to [1, 2, 4, 8]; it is set to return the data of all timestamps; finally, the proportion of the lost input data of the previous layer is set to 0.1; A unidirectional LSTM network for converting the sequence of data into 128 dimensions and outputting the data of the last timestamp; Two fully connected layers, where the number of units in the first fully connected layer is 64, and the activation function is ReLU. The calculation formula is: ; Among them, is the input value after linear transformation, which is obtained by multiplying the 1×128 data output by the unidirectional LSTM network by a 128×64 weight matrix and then adding a 1×64 bias term. The number of units in the second fully connected layer is 26, and the activation function is Softmax. The calculation formula is: ; Among them, is the data of 1×64 output by the first fully connected layer, multiplied by a weight matrix of 64×26, and then added with a bias term of 1×26 to obtain the data of the th row of the column vector, is the total number of categories, is the index, is the data of the th row corresponding to the index.
8. A computer program product, characterized in that, The computer program product includes a computer program. The computer program is stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the method for inputting handwritten letters on the back of a smart phone by fusing IMU and UWB according to any one of claims 1-7.
Citation Information
Patent Citations
Robot autonomous localization and navigation based on image-text recognition and semantic meaning
CN107967473A
Detection model training method and device, code detection method and device and related equipment
CN118839721A