Sign language recognition method based on wearable computing

By combining inertial data gloves and neural network models, high-precision recognition of complex and subtle sign language is achieved, solving the problem of poor recognition results in existing technologies, simplifying the processing circuit and improving the recognition rate.

CN115904086BActive Publication Date: 2026-03-27DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing sign language recognition methods are not effective in recognizing complex and delicate dynamic sign language. Systems based on a single sensor are difficult to achieve high-precision recognition, while multi-sensor systems have complex processing circuits and require calibration and correction, making it impossible to effectively recognize complex and delicate hand movements.

Method used

Employing two inertial data gloves containing 16 inertial sensors, combined with a routing device and a human-computer interaction platform, the system achieves precise capture and recognition of hand movements through error correction, posture calculation, and a neural network sign language recognition model.

Benefits of technology

The processing circuitry has been simplified, the recognition rate for complex and nuanced sign language has been improved, the computational load has been reduced, and high-precision recognition performance has been maintained in various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115904086B_ABST
    Figure CN115904086B_ABST
Patent Text Reader

Abstract

The application belongs to the field of sign language recognition, and provides a sign language recognition method based on wearable computing, and a sign language recognition system based on wearable computing, which comprises a pair of inertial data gloves, a routing device and a human-computer interaction platform. Each inertial data glove comprises sixteen inertial sensors and a wireless transmission module, wherein the inertial sensors comprise a three-axis accelerometer, a three-axis gyroscope and a three-axis magnetometer, and the wireless transmission module completes interaction with the human-computer interaction platform through the transmission control protocol with the routing device as the hub. The human-computer interaction module can operate the sign language recognition system, and is integrated with a trained Bi-LSTM model, hand data uploaded by the inertial data gloves is input into the Bi-LSTM model for sign language recognition, and the recognized sign language information is displayed. The system can recognize 20 kinds of Chinese dynamic sign languages, and has the characteristics of high precision, low power consumption, small device size and easy deployment and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sign language recognition, and more specifically to a sign language recognition method based on wearable computing. Background Technology

[0002] Sign language, as a language used by deaf and mute individuals to express information, plays a vital role in their lives. Sign language primarily includes finger language and sign language. Finger language uses the shapes and movements of the fingers to represent letters, while sign language uses hand gestures, facial expressions, and body postures to convey meaning. Aside from those working in related professions, most people do not understand or use sign language, creating a communication barrier between deaf and mute individuals and hearing people. Deaf and mute sign language interpreters bridge this gap, but due to a shortage of qualified interpreters and their high cost, it is difficult to meet current societal needs. Therefore, establishing an accessible communication platform between deaf and mute individuals, and between deaf and mute individuals and hearing people, is crucial. The establishment of a sign language recognition and communication platform can provide deaf and mute individuals with a channel to communicate with hearing people, thereby enjoying more comprehensive public services. This platform not only allows us to enter the lives of deaf and mute individuals but also enables them to integrate into social life more naturally, harmoniously, and comfortably.

[0003] Currently, mainstream sign language recognition methods generally rely on vision or sensors to capture relevant information about the hands, and then combine this with machine learning or deep learning models to determine the meaning of the sign language. Among these, computer vision-based methods primarily use cameras as the main tool. Although using cameras as the core of sign language recognition simplifies the operation, reduces the need for additional sensors, and provides a more comfortable user experience, vision-based sign language recognition systems also have many drawbacks. "P. Kumar, H. Gauba, P. Pratim Roy, and D. Prosad Dogra, 'A multimodal framework for sensor-based sign language recognition,' Neurocomputing, 259, pp. 21-38, (2017)" points out that the main problems include variations in lighting intensity and environmental changes, high computational costs, and poor portability.

[0004] Conversely, sensor-based sign language recognition systems offer advantages such as relatively simple computation, lack of environmental limitations, and low cost, thus compensating for the shortcomings of visual sign language recognition technology. Current wearable sensor devices have also significantly improved user comfort. The research on “Y.Li, N.Yang, L.Li, L.Liu, Y.Yang, Z.Zhang, L.Liu, L.Li, and X.Zhang,”Finger gesture recognition using a smartwatch with integrated motionsensors,”Web Intelligence, 16, pp. 123-129, (2018), and “T.Tai, Y.Jhang, Z.Liao, K.Teng, and W.Hwang,”Sensor-Based Continuous Hand Gesture Recognition by LongShort-Term Memory,”IEEE Sensors Letters, 2, pp. 1-4, (2018), respectively utilizes accelerometers and gyroscopes integrated into smartwatches and smartphones to acquire muscle activity data for gesture recognition. These sensors are not constrained by the environment, making them more suitable for sign language recognition by the deaf and mute. However, for the recognition of complex and subtle dynamic sign language movements, it is difficult to achieve significant recognition results based on a single sensor. Therefore, many studies often combine multiple sensors for sign language recognition, including data gloves. For example, the data glove designed by combining a bending sensor, Hall sensor, and accelerometer in "T. Chouhan, A. Panse, AKVoona, and SMSameer, 'Smartglove with gesture recognition ability for the hearing and speech impaired,' 2014, pp. 105-110" or the data glove integrating a bending sensor and IMU in "Y. Dong, J. Liu and W. Yan, 'Dynamic Hand Gesture Recognition Based on Signals From Specialized Data Glove and Deep Learning Algorithms,' IEEE T. Instrum. Meas., 70, pp. 1-14, (2021)" is an example. However, data gloves based on multiple sensors have problems such as complex processing circuits, the need for calibration and correction for different sensors during the processing stage, and the inability to recognize complex and delicate hand movements. Summary of the Invention

[0005] The purpose of this invention is to provide a sign language recognition method based on wearable computing, which can achieve precise capture of hand movements, thereby recognizing complex and dynamic sign language.

[0006] The technical solution of the present invention is as follows:

[0007] A sign language recognition method based on wearable computing, wherein a sign language recognition system based on wearable computing is used to recognize sign language; the sign language recognition system based on wearable computing includes two inertial data gloves, a routing device, and a human-computer interaction platform;

[0008] Each inertial data glove includes sixteen inertial sensors and a wireless transmission module; the inertial sensors include a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer, used to capture motion and position information of the palm and fingers of a single hand; the wireless transmission module interacts with the human-machine interface platform through a transmission control protocol with a routing device as the central hub.

[0009] The human-computer interaction platform performs error correction, posture calculation, and sign language sample segmentation on the data collected by the inertial data glove, and then inputs it into the neural network sign language recognition model. The recognized gestures are finally displayed on the screen.

[0010] The sign language recognition method based on wearable computing includes the following steps:

[0011] Step 1: Create a sign language dataset;

[0012] Step 1.1: Collect the raw inertial data of the sign language to be recognized.

[0013] Raw inertial data of 20 dynamic sign languages ​​were collected using inertial data gloves;

[0014] Step 1.2: Preprocess the collected raw sign language inertial data.

[0015] Preprocessing includes error correction, attitude calculation, sign language sample segmentation, and data normalization of the collected raw inertial data;

[0016] Step 1.2.1: Use the least squares method to solve the ellipsoid fitting problem to calibrate the triaxial magnetometer;

[0017]

[0018] in, This indicates the measured value of the triaxial magnetometer. This represents the theoretical value indicating that the triaxial magnetometer has no error. k represents the zero bias error of the triaxial magnetometer. x ,k y ,k zk represents the sensitivity error coefficient of each axis of the triaxial magnetometer. xy k yx k zx k zy k xz k yz This represents the coupling coefficient of each axis of the triaxial magnetometer;

[0019] The ellipsoid equation is obtained by fitting using the least squares method, and the error coefficient K and the zero bias error B are solved. 0 ,

[0020]

[0021] in, The matrix represents the shape of the ellipsoid; a, b, c, d, e, f, p, q, r represent the ellipsoidal surface parameters to be solved; H represents the magnetic field strength at a certain location.

[0022] Step 1.2.2: Attitude Calculation. The raw inertial data collected by the inertial data glove is used to extract hand quaternion information using a gradient descent data fusion algorithm, and the following formula is used for iteration:

[0023]

[0024] Where q represents the rotation quaternion, β represents the measurement error of the three-axis gyroscope, which is obtained through a calibration algorithm; t s Indicates the sampling period; F f,m (q,f,m) represents the objective function;

[0025] The error generated by the triaxial accelerometer is expressed by the error function F. f (q,f) is represented as:

[0026]

[0027] in, The measured value is from a triaxial accelerometer; F fx For accelerometer X-axis error; F fy For accelerometer Y-axis error; F fz q0, q1, q2, q3 are the Z-axis error of the accelerometer; q0, q1, q2, q3 are rotational quaternions;

[0028] The error generated by the triaxial magnetometer is expressed by the error function F. m (q,m) is represented as:

[0029]

[0030] in, q0, q1, q2, q3 are the measured values ​​of the triaxial magnetometer; q0, q1, q2, q3 are the rotational quaternions; Fmx For the X-axis error of the triaxial magnetometer; F my For the Y-axis error of the triaxial magnetometer; F mz This refers to the Z-axis error of the triaxial magnetometer.

[0031] Taking the partial derivative of the objective function yields the Jacobian matrix.

[0032]

[0033] The error q is minimized using the Gauss-Newton method. ▽,t This represents the target pose during the calculation process.

[0034]

[0035] Where μ is the step size; ▽ represents the differential operator, represents the rate of change of q, and its relationship with the error vector F f,m The relationship between (q,f,m) is as follows:

[0036]

[0037] Finally, a quaternion representing the attitude is obtained, consisting of two parts: one part is obtained by solving the differential equation using the output values ​​of the three-axis gyroscope; the other part is a corrected quaternion obtained by solving the objective function. Both describe the hand's attitude at the same moment. Let the attitude quaternion obtained by solving the quaternion differential equation be q. ω,t Let α represent the weight of the modified quaternion. Combining these two quaternions yields the final attitude quaternion q. t Represented as;

[0038] q t =αq ▽,t +(1-α)q ω,t

[0039] β represents the measurement error of the three-axis gyroscope, obtained through a calibration algorithm; t s The sampling period is represented by α; the weight α and the step size μ are related as follows:

[0040]

[0041] The final attitude quaternion representation is as follows:

[0042]

[0043] Step 1.2.2.1: Determine the geographic North Pole direction using the calibrated triaxial magnetometer signal from Step 1.2.1, and obtain the initial posture of the hand based on the three-dimensional acceleration signal when the hand is naturally hanging down and the calibrated three-dimensional magnetometer signal from Step 1.2.1, and then perform calibration.

[0044] By solving for the Euler angles relative to the sensor coordinate system and the carrier coordinate system, the pitch angle θ and roll angle γ are obtained using the three-dimensional acceleration signal, as shown in the following formulas:

[0045]

[0046]

[0047] in, These are the outputs of the three axes of the triaxial accelerometer;

[0048] The heading angle is obtained by combining the calibrated three-dimensional magnetometer signal.

[0049]

[0050] in, It is the projection of the X and Y axes of the three-dimensional magnetometer signal corrected to the local plane, with the X axis pointing to the Earth's magnetic north pole;

[0051] Step 1.2.2.2: Use the following transformation formula between the carrier coordinate system and the sensor coordinate system to convert the Euler angles obtained in Step 1.2.2.1 into unit quaternions from the carrier coordinate system to the sensor coordinate system:

[0052]

[0053] Step 1.2.2.3: Based on the motion signal of the hand hanging naturally and the calibrated three-dimensional magnetometer signal, determine the rotation quaternion between the sensor coordinate system and the geographic coordinate system by the direction of the gravity vector and the direction of the Earth's magnetic field north pole, and complete the calibration of the initial hand posture and the initial hand model, including the following process:

[0054] Formula for the coordinate position of joint points in a hand model:

[0055]

[0056] in, Let m be the coordinates of the hand joint at time t in the geographic coordinate system; u is the number of hand limb vectors; O(t) is the coordinates of the starting joint of the hand, referring to the coordinates of the center of the palm.

[0057] The rotation of the hand joint vector in the geographic coordinate system is as follows:

[0058]

[0059] in, The conjugate of the rotation quaternion from the carrier coordinate system to the geographic coordinate system; the hand joint vector is represented in the geographic coordinate system as... The hand joint vector is represented in the carrier coordinate system as: Represents quaternion multiplication;

[0060] Finally, the rotation formula for the hand joint vector in the geographic coordinate system is obtained:

[0061]

[0062] Among them, the initial The gravity vector and the direction of the magnetic field north pole are determined by using three-dimensional acceleration data and calibrated three-dimensional magnetometer data.

[0063] During the initial hand posture correction, the carrier coordinate system is aligned with the geographic coordinate system by aligning the carrier with the magnetic north pole when the hand is naturally hanging down and at rest, thus obtaining the formula at the initial moment:

[0064]

[0065] in, The rotation quaternion represents the rotation from the hand carrier coordinate system to the sensor coordinate system; The rotation quaternion represents the rotation from the hand's geographic coordinate system to the sensor coordinate system;

[0066] Step 1.2.2.4: Finally, the hand motion posture is updated using the gradient descent data fusion algorithm to obtain the real-time updated hand quaternion information;

[0067] Step 1.2.3: Sign language sample segmentation. The raw acceleration, three-axis gyroscope data and the calculated hand posture quaternion information are combined as shown in the following formula; and a high-speed camera is used to perform manual sign language segmentation on the combined sign language time series data and assign specific sign language labels.

[0068] [ax,ay,az,gx,gy,gz,q0,q1,q2,q3]

[0069] Where ax, ay, az are triaxial acceleration signals acquired by the triaxial accelerometer, gx, gy, gz are three-dimensional angular velocity signals acquired by the triaxial gyroscope, and q0, q1, q2, q3 are hand unit quaternions;

[0070] Step 1.2.4: Data normalization will place the segmented data with specific sign language labels within the range of 0 to 1;

[0071] Step 1.3: Divide the preprocessed sign language time series data of each subject into a balanced training set and a test set at a ratio of 9:1;

[0072] Step 2: Establish a neural network sign language recognition model;

[0073] Step 3: Train the neural network sign language recognition model using the sign language dataset;

[0074] Step 4: Perform generalization performance testing on the trained neural network sign language recognition model;

[0075] Step 5: The neural network sign language recognition model that meets the test results is ported to the human-computer interaction platform, and the sign language is recognized in real time through the wearable computing-based sign language recognition system.

[0076] The neural network sign language recognition model is a bidirectional long short-term memory neural network model (Bi-LSTM); step 2 specifically includes the following steps:

[0077] Step 2.1: Determine the type of neural network sign language recognition model

[0078] The Bi-LSTM model was used to detect the processed sign language time series data. The Bi-LSTM model consists of two LSTM layers. The input sequences were fed into the two LSTM layers in both forward and reverse order for feature extraction.

[0079] Step 2.2: Model Structure and Selection of Model Optimization Function

[0080] The Bi-LSTM model consists of a time series input layer, a Bi-LSTM layer, a fully connected layer, a softmax layer, and a classification output layer; the model optimization function is ADMA.

[0081] In step 4, accuracy, precision, recall, and F1 score are used to evaluate the generalization performance of the neural network sign language recognition model.

[0082] The accuracy formula is as follows:

[0083]

[0084] The precision formula is as follows:

[0085]

[0086] The recall rate formula is as follows:

[0087]

[0088] The formula for F1 score is as follows:

[0089]

[0090] TP represents true positive, FP represents false positive, TN represents true negative, FN represents false negative, and F1-score is the harmonic mean of precision and recall.

[0091] Step 5 specifically includes the following steps:

[0092] Step 5.1: Load the trained sign language recognition model into the human-computer interaction platform;

[0093] Step 5.2: Correct the triaxial magnetometer data, same as step 1.2.1; collect data during random figure-eight rotation as samples for triaxial magnetometer correction and perform ellipsoid fitting correction using the least squares method;

[0094] Step 5.2: The subject faces north with their hands hanging naturally for 3-4 seconds to complete the initial posture calibration and begin sign language recognition;

[0095] Step 5.3: Perform gesture calculation, automatic sign language segmentation, and data normalization on the real-time collected sign language data; the automatic sign language segmentation method is as follows:

[0096]

[0097] Where N is the smoothing length, These are the quaternions at the k-th sampling point, obtained by applying Δ qN After setting the threshold, it is used to achieve accurate detection of the start and end points of the gesture;

[0098] Step 5.4: Input the processed data into the loaded Bi-LSTM model for recognition;

[0099] Step 5.5: Output the recognition results.

[0100] The training method in step 3 uses five-fold cross-validation.

[0101] The beneficial effects of this invention are as follows:

[0102] 1. The sign language recognition system of the present invention uses a data glove based on multiple IMUs as the acquisition end, which simplifies the processing circuit, reduces the amount of computation in the data processing stage, and can ensure the acquisition of all posture, movement and position information of the hand, including the fingers, to recognize complex and intricate sign language.

[0103] 2. The sign language recognition system of the present invention uses the least squares method to perform ellipsoidal fitting correction on the magnetometer of the data glove, which solves the problem that the magnetometer data may be interfered with by different magnetic field strengths in the environment, resulting in a decrease in recognition rate.

[0104] 3. The sign language recognition system of the present invention adopts a gradient descent-based posture calculation method, which ensures high accuracy in posture changes when the hand makes large movements and in fine movements, thereby improving the recognition rate of the sign language recognition system for highly similar sign languages. Attached Figure Description

[0105] Figure 1 This is a hardware block diagram of a sign language recognition system based on wearable computing according to the present invention;

[0106] Figure 2 The IMU of the inertial data glove in this invention is located in the hand portion.

[0107] Figure 3 This is the overall algorithm flow of a sign language recognition method based on wearable computing in this invention;

[0108] Figure 4 shows 20 sample sign language actions in the embodiments of the present invention. The samples are selected from 20 dynamic sign languages ​​in Chinese deaf-mute sign language. Figure 4(a) is sorry; Figure 4(b) is angry; Figure 4(c) is sad; Figure 4(d) for you; Figure 4(e) is hello; Figure 4(f) is try; Figure 4(g) is for them; Figure 4(h) is me; Figure 4(i) is thank you; Figure 4(j) is goodbye; Figure 4(k) is protect; Figure 4(l) is dynamic; Figure 4(m) is happy; Figure 4(n) is welcome; Figure 4(o) is getting married; Figure 4(p) is joking; Figure 4(q) is stand up; Figure 4(r) is agree; Figure 4(s) is stop; Figure 4(t) is unity.

[0109] Figure 5(a) is a schematic diagram of the body coordinate system;

[0110] Figure 5(b) is a schematic diagram of the geographic coordinate system;

[0111] Figure 5(c) is a schematic diagram of the sensor coordinate system;

[0112] Figure 6 This is the preprocessed data structure in this invention;

[0113] Figure 7 This is the specific structure of the sign language recognition model in this invention;

[0114] Figure 8 This is the confusion matrix for recognizing various sign language targets in the examples of this invention.

[0115] Figure 9(a) shows the triaxial magnetometer before calibration;

[0116] Figure 9(b) shows the triaxial magnetometer after calibration;

[0117] Figure 10(a) shows the reconstruction effect of the gradient descent data fusion method;

[0118] Figure 10(b) shows the reconstruction effect of the complementary filtering data fusion method; Detailed Implementation

[0119] To make the technical problems solved by the present invention, the technical solutions adopted, and the technical effects achieved clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and not for limiting the present invention. Furthermore, it should be noted that, for ease of description, only parts of the present invention are shown in the accompanying drawings, not all of them.

[0120] according to Figure 1 The present invention provides a sign language recognition system based on wearable computing, including an inertial data glove, a routing device, and a human-computer interaction platform.

[0121] The inertial data glove comprises four parts: a motion acquisition node, a data aggregation node, a data transmission module, and a main control unit for the acquisition system. Each glove's motion acquisition node includes 16 IMUs (Inertial Measurement Units) to capture the movement and position information of the hand and fingers. The data aggregation node combines the acquired hand sensor information and adds frame headers and trailers for subsequent data upload. The data transmission module uses a network communication protocol to achieve data transmission between the inertial data glove and the acquisition software, enabling wireless uploading, storage, and processing of motion data. The main control unit for the acquisition system is responsible for preprocessing the data acquired by the sensors.

[0122] The routing device serves as the system's data communication hub, establishing a local area network to enable data transmission and command control between the human-machine interface platform and the data glove. The routing configuration must ensure that the data glove, routing device, and human-machine interface platform are all within the same IP network segment.

[0123] The human-computer interaction platform is used to operate the sign language recognition system and process data uploaded from the inertial data glove. It mainly consists of a control unit and an analysis unit. The control unit sends acquisition commands to the inertial data glove via the TCP network protocol and monitors the working status of the inertial data glove. The analysis unit processes the raw IMU data stored on the PC, reconstructs hand postures and segments, and recognizes sign language. The above operations include devices such as a keyboard and mouse. Considering the portability of the system, a touchscreen can be used to display results and set operation settings.

[0124] In this embodiment, the inertial data IMU uses the MPU9250. This chip can acquire nine-axis data, including data from a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. It has two optional communication interfaces, IIC and SPI, and its power consumption is in the mA range, which can meet the requirements for real-time acquisition of sign language information. The data transmission module uses an onboard ESP8266 WIFI module, providing a high-speed SPI serial host communication interface with a communication bandwidth of up to megabytes per second, meeting the requirements for uploading sign language data. The main controller of the acquisition system uses the STM32407VGT6 chip with a Cortex-M4 architecture, which integrates a floating-point unit to accelerate the calculation of IMU data. The entire system is centrally powered by a 3.3V lithium battery.

[0125] The inertial data glove distributes 15 IMUs (Inertial Measurement Units) on each finger (thumb, index, middle, ring, and little fingers) to measure finger movement information. The IMU nodes are connected by flexible cables to ensure joint flexibility. One IMU is located on the back of the hand to measure hand movement information.

[0126] Figure 2 The mapping relationship between each sensor and the hand is shown.

[0127] In this embodiment, a 300Mbps router is used as the routing device, whose transmission bandwidth is far higher than that of the inertial data glove, thus meeting the system requirements. This invention does not impose special requirements on the intermediate routing device; in outdoor environments, devices capable of building a local area network, including but not limited to mobile hotspots and wireless modules, can be used.

[0128] In this embodiment, the human-computer interaction platform includes a keyboard, a mouse, and a monitor. The software functions include monitoring the status of the data glove, system control, data uploading and storage, data preprocessing, and sign language recognition. The sign language recognition is accomplished through an embedded sign language recognition neural network model.

[0129] In this embodiment, the construction process of the wearable computing-based sign language recognition model is as follows:

[0130] Step 1: Create a sign language dataset;

[0131] Step 2: Establish a neural network model;

[0132] Step 3: Train the neural network model using the sign language dataset;

[0133] Step 4: Perform generalization performance testing on the trained neural network model;

[0134] Step 5: Transfer the model that meets the test results to the human-computer interaction platform, and use the system to recognize sign language in real time.

[0135] The overall algorithm and usage process are as follows: Figure 3 As shown.

[0136] In this embodiment, the method for creating the sign language dataset is as follows:

[0137] Step 1.1: Collect raw sample data of the sign language to be recognized. The sample data is shown in Figure 4, which contains 20 kinds of dynamic sign language. The specific description of the sign language actions is shown in Table 1. In this embodiment, a total of 140 sets of sample data for each sign language are collected, and each set contains the above 20 sign languages.

[0138] Table 1 Summary of Sign Language Descriptions

[0139]

[0140] Step 1.1.1: Turn on the data glove switch to complete system initialization, connect to the human-computer interaction platform, and prepare to collect sign language movements. The subject should face north and remain still for 3-4 seconds to complete the initial posture calibration.

[0141] The output data of the inertial data glove all come from their respective sensor coordinate systems, while the sign language gestures in this invention are relative to spatial position. This involves multiple commonly used coordinate systems and the transformation relationships between them. Through the transformation relationships between different coordinate systems, the measurement values ​​in the sensor coordinate system can be transformed to the geographic coordinate system. The entire system includes three coordinate systems, as shown in Figure 5, as follows:

[0142] Sensor coordinate system (SCS, O-XsYsZs): The coordinate system of the gyroscope in the MPU9250 datasheet is selected as the sensor coordinate system. When wearing the sensor, aligning the sensor coordinate system with the carrier coordinate system can reduce attitude calculation errors.

[0143] Body coordinate system (BCS, O-XbYbZb): The body coordinate system uses the center of mass of the body as its origin. Commonly used human body coordinate systems include the "right-front-upper" and "front-right-lower" rectangular coordinate systems. In this article, all three coordinate systems point to the front-right-lower.

[0144] Geographic coordinate system (GCS, O-XgYgZg): The geographic coordinate system takes the location of the carrier as its origin. Based on whether the O-Zg axis points upwards or downwards along the local vertical, it can be divided into an East-North-Sky coordinate system and a North-East-Earth coordinate system, both of which satisfy the right-hand rule. In this paper, the three axes of the geographic coordinate system point to North, East, and the Earth's center, respectively. North, as mentioned in this paper, specifically refers to the magnetic North Pole.

[0145] Step 1.1.2: The subject follows the on-screen prompts to complete the sign language actions. After each sign language action is completed, the arm hangs down naturally.

[0146] Step 1.1.3: After completing the collection of a set of sign language data, save the raw sign language inertial data to the designated folder as sample data for subsequent processing.

[0147] Step 1.1.4: Rest for two minutes, then proceed to the next set of data collection. Collect 20 sets of sign language data from each subject.

[0148] Step 1.2: Preprocess the collected sign language sample data, including error correction, posture calculation, sign language sample segmentation, and data normalization;

[0149] Step 1.2.1: Error correction is an indispensable part of inertial sensors. Sensor deviation and error are the main factors affecting the accuracy of attitude fusion. In this embodiment, the IMU used in the inertial data glove is the MPU9250, which is low-cost. However, it exhibits significant gyroscope drift when calculating hand attitude, leading to inaccurate attitude calculations. Typically, triaxial magnetometer data fusion compensation is required. However, triaxial magnetometers are easily affected by various factors in space, resulting in large measurement errors. Therefore, calibration is necessary before use to improve measurement accuracy. In this embodiment, the triaxial magnetometer is calibrated using the figure-eight calibration method. That is, after the glove begins data acquisition, it is rotated in space in a figure-eight pattern for 10 to 15 seconds. This invention uses the least squares method to solve the ellipsoid fitting problem to calibrate the magnetometer, solving the following error model:

[0150]

[0151] in, This indicates the measured value of the triaxial magnetometer. This represents the theoretical value indicating that the triaxial magnetometer has no error. k represents the zero bias error of the triaxial magnetometer. x ,k y ,k z k represents the sensitivity error coefficient of each axis of the triaxial magnetometer. xy k yx k zx k zy k xz k yz This represents the coupling coefficient of each axis of the triaxial magnetometer.

[0152] The ellipsoid equation is obtained by fitting using the least squares method, and the error coefficient K and the zero bias error B are solved. 0 ,

[0153]

[0154] in, denoted by ellipsoid shape matrix; a, b, c, d, e, f, p, q, r represent the ellipsoidal surface parameters to be solved; H represents the magnetic field strength at a certain location.

[0155] The triaxial magnetometer data is corrected according to the magnetic field error model;

[0156] Step 1.2.2: Attitude calculation involves using data fusion algorithms to determine the quaternions or Euler angles of the current carrier from the triaxial accelerometer, triaxial gyroscope, and triaxial magnetometer data collected by the inertial data glove. In this example, a gradient descent-based data fusion method is used to calculate the hand's attitude. Since directly solving the differential equation using gyroscope data introduces significant errors, error correction terms based on accelerometer and magnetometer data are added to eliminate the integral error of the gyroscope data and achieve attitude update. The principle of the gradient descent data fusion method in this embodiment is described below:

[0157] In this algorithm, an error function is defined to reflect the attitude error during the solution process. This error function is the objective function in the gradient descent algorithm. In this embodiment, the error function is a multivariate vector function. The system's attitude error is generated by two parts: one part is generated by the triaxial accelerometer, and the other part is generated by the triaxial magnetometer. Therefore, the error function consists of the triaxial accelerometer calculation error and the triaxial magnetometer calculation error. When the error function is zero, the obtained attitude matrix is ​​considered to have no error.

[0158] The error generated by the accelerometer is expressed by the error function F. f (q,f) can be represented as:

[0159]

[0160] in, These are measurements from a triaxial accelerometer.

[0161] The error generated by the triaxial magnetometer is expressed by the error function F. m (q,m) can be represented as:

[0162]

[0163] in, This is the measurement value from the magnetometer.

[0164] The total error function, i.e., the objective function F f,m (q,f,m) can be represented as

[0165]

[0166] Taking the partial derivative of the objective function, the Jacobian matrix can be expressed as follows:

[0167]

[0168] From an optimization perspective, to minimize the error, we can use the Gauss-Newton method to solve for q. ▽,t The target pose during the calculation process can be represented by the following relationship:

[0169]

[0170] Where μ is the step size; ▽ represents the differential operator, represents the rate of change of q, and represents the sum of the error vector, i.e., F. f,m The relationship between (q,f,m) is as follows:

[0171]

[0172] The algorithm ultimately solves for a quaternion representing the hand's attitude, which consists of two parts. One part is obtained by solving a differential equation using the output values ​​of the three-axis gyroscope; the other part is a corrected quaternion obtained by solving the objective function. Both can describe the hand's attitude at the same moment. Let q be the attitude quaternion obtained by solving the quaternion differential equation. ω,t Let α represent the weight of the modified quaternion. Combining these two quaternions yields the final attitude quaternion q. t It can be represented as

[0173] q t =αq ▽,t +(1-α)q ω,t

[0174] Assuming β represents the measurement error of the gyroscope, the system parameters can be obtained through a calibration algorithm; t s The sampling period is represented by α; the weight α and the step size μ are related as follows:

[0175]

[0176] The final pose quaternion can be represented as:

[0177]

[0178] Step 1.2.2.1: Determine the geographic North Pole direction using the calibrated 3D magnetometer signal from Step 1.2.1, and obtain the initial posture of the hand based on the 3D acceleration signal when the hand is naturally hanging down and the calibrated 3D magnetometer signal from Step 1.2.1, and then perform calibration.

[0179] By solving for the Euler angles relative to the sensor coordinate system and the carrier coordinate system, the pitch angle θ and roll angle γ are obtained using the three-dimensional acceleration signal, as shown in the following formulas:

[0180]

[0181]

[0182] in These are the outputs of the three axes of the accelerometer;

[0183] The heading angle can be obtained by combining the corrected three-dimensional magnetometer signal.

[0184]

[0185] in It is the projection of the X and Y axes of the three-dimensional magnetometer signal corrected to the local plane, with the X axis pointing to the Earth's magnetic north pole;

[0186] Step 1.2.2.2: Use the following transformation formula between the carrier coordinate system and the sensor coordinate system to convert the Euler angles obtained in Step 1.2.2.1 into unit quaternions from the carrier coordinate system to the sensor coordinate system:

[0187]

[0188] Step 1.2.2.3: Based on the motion signal of the hand hanging naturally and the calibrated three-dimensional magnetometer signal, the rotation quaternion between the sensor coordinate system and the geographic coordinate system can be determined by the direction of the gravity vector and the direction of the Earth's magnetic field north pole. This completes the calibration of the initial hand posture and the initial hand model, including the following process:

[0189] Formula for the coordinate position of joint points in a hand model:

[0190]

[0191] in, Let m be the coordinates of the hand joint at time t in the geographic coordinate system; u is the number of hand limb vectors; O(t) is the coordinates of the starting joint of the hand, specifically the coordinates of the center of the palm.

[0192] The rotation of the hand joint vector in the geographic coordinate system can be represented as:

[0193]

[0194] in, The conjugate of the rotation quaternion from the carrier coordinate system to the geographic coordinate system; the hand joint vector is represented in the geographic coordinate system as... The hand joint vector is represented in the carrier coordinate system as: This represents quaternion multiplication.

[0195] Finally, the rotation formula for the hand joint vector in the geographic coordinate system is obtained:

[0196]

[0197] Among them, the initial The gravity vector and the direction of the magnetic field north pole are determined by using three-dimensional acceleration data and three-dimensional magnetometer data.

[0198] During the initial hand posture correction, the carrier coordinate system is aligned with the geographic coordinate system by aligning the carrier with the magnetic north pole when the hand is naturally hanging down and at rest. This allows us to obtain the following at the initial moment:

[0199]

[0200] in, The rotation quaternion represents the rotation from the hand carrier coordinate system to the sensor coordinate system; The rotation quaternion represents the rotation from the hand's geographic coordinate system to the sensor coordinate system;

[0201] Step 1.2.2.4: Finally, the hand motion posture is updated using the gradient descent data fusion algorithm to obtain real-time updated hand quaternion information; the updated quaternion is combined with the original acceleration signal collected by the sensor and the angular velocity signal collected by the gyroscope for further processing. The combined data structure is as follows: Figure 6 As shown.

[0202] Step 1.2.3: Sign language data segmentation involves labeling the collected sign language data with sign language samples for subsequent model training; and segmenting the sign language data by comparing it with high frame rate camera video. Since the data of each node is sampled synchronously, only the hand nodes need to be segmented.

[0203] This embodiment includes the three-dimensional acceleration signal, three-dimensional angular velocity signal, and hand posture unit quaternion for each acquisition node of the hand, as shown in the following formula:

[0204] [ax,ay,az,gx,gy,gz,q0,q1,q2,q3]

[0205] Where ax, ay, az are the three-axis acceleration signals acquired by the accelerometer, gx, gy, gz are the three-dimensional angular velocity signals acquired by the three-axis gyroscope, and q0, q1, q2, q3 are the hand unit quaternions.

[0206] This feature combination method is adopted because the sign language being recognized is dynamic. If Euler angles are used as features, there will be a gimbal lock problem in large-amplitude movements, while rotation quaternions do not have this problem. Acceleration and angular velocity signals are introduced as features because they contain the original dynamic information of hand movements. Compared with using only hand rotation quaternions as features, the accuracy is significantly improved.

[0207] Step 1.2.4: Data normalization involves placing the final segmented data within the range of 0 to 1;

[0208] In this embodiment, the specific steps for constructing the neural network model are as follows:

[0209] Step 2.1: Determine the type of neural network model;

[0210] This embodiment employs a Bi-LSTM model, which combines forward and backward information from the input sequence. In this invention, the vector outputs of two LSTM layers are concatenated and sent to the next layer. The Bi-LSTM model framework is as follows: Figure 7 As shown, the collected sign language data is manually segmented and then directly input into the Bi-LSTM. The model leverages its own characteristics to learn features from the sign language context, finally outputting the recognition result. The main reason for choosing Bi-LSTM is that the inertial sensor signals for sign language are time-dependent, and sign language contains many implicit features in the time domain. Bi-LSTM can better capture the contextual information in the sequence. Secondly, considering the small sample size of self-built datasets, it is difficult to train a good deep learning model, while Bi-LSTM is more suitable for small sample datasets.

[0211] Step 2.2: Model Structure and Selection of Model Optimization Function;

[0212] The Bi-LSTM model consists of a time series input layer, a Bi-LSTM layer, a fully connected layer, a softmax layer, and a classification output layer; the model optimization function is ADMA.

[0213] Training and testing of the Bi-LSTM model can be performed on any platform, such as TensorFlow, PyTouch, or MATLAB. In this embodiment, training, testing, and algorithm comparison are conducted within the MATLAB deep learning framework. The specific training process is as follows:

[0214] Step 3: Train the model using five-fold cross-validation on the sign language dataset divided into training and validation sets in Step 1. Cross-validation is used because severe imbalance in the dataset or overly complex models may lead to overfitting, meaning a well-trained model might perform well on the training set but perform poorly on the validation set. Conversely, a small training set may also lead to underfitting, resulting in poor performance on both the training and validation sets. A good recognition model typically strikes a balance between these two aspects to ensure its usefulness in subsequent engineering. Therefore, this embodiment considers practical considerations and uses five-fold cross-validation, dividing the dataset into five subsets: four for training and one for validation. This process is repeated five times, validating each subset once, and the final prediction result is obtained by averaging the validation results.

[0215] The testing process for the sign language recognition model in this invention is as follows:

[0216] Step 4 uses accuracy, precision, recall, and F1 score to evaluate the model's generalization performance. Accuracy is generally an important indicator for evaluating the performance of multi-class classification models, and the formula is as follows:

[0217]

[0218] Precision is only applied to correctly predicted positive samples; it represents the percentage of truly positive samples out of those predicted as positive. The formula is as follows:

[0219]

[0220] Recall rate represents how many actual positive samples the classifier can predict, and the formula is as follows:

[0221]

[0222] The F1 score is a composite evaluation metric, a harmonic ratio of precision and recall, and is expressed by the following formula:

[0223]

[0224] TP represents a true positive, FP represents a false positive, TN represents a true negative, and FN represents a false negative.

[0225] Currently, sign language recognition models based on wearable technology mainly rely on traditional machine learning models. These models extract features manually from collected sensor data and then pass them to a model recognizer for sign language recognition. This invention selects four classification algorithms—DT, SVM, KNN, and RF—and conducts comparative experiments with the method used in this invention to verify its effectiveness.

[0226] In this embodiment, a five-fold cross-validation method is used to prevent model overfitting, and the average accuracy of five predictions is calculated. We used a total of five algorithms to train on 20 sign languages. The final accuracy of the five trained models for recognizing the 20 sign languages ​​is shown in Table 2.

[0227] Table 2. Model Cross-Validation Accuracy

[0228]

[0229] The model with the highest accuracy after five cross-validations was selected from the five model methods for generalization performance evaluation. A pre-prepared dataset that was not used for training or validation was input into the model, and the final evaluation metrics are shown in Table 3.

[0230] Table 3 Summary of Model Evaluation Indicators

[0231]

[0232]

[0233] Figure 8 The confusion matrix for 20 sign language recognition is given, representing the percentage of predicted sign language labels that are equal to the true sign language labels.

[0234] In this embodiment, the specific steps for sign language recognition using a model are as follows:

[0235] Step 5.1: Load the trained sign language recognition model into the human-computer interaction platform;

[0236] Step 5.2: Correct the magnetometer data, same as step 1.2.1; collect data during random figure-eight rotation as samples for magnetometer correction and perform ellipsoid fitting correction using the least squares method;

[0237] Step 5.2: Same as step 1.1.1;

[0238] Step 5.3: Perform gesture calculation, automatic sign language segmentation, and data normalization on the real-time collected sign language data; the automatic sign language segmentation method is as follows:

[0239]

[0240] Where N is the smoothing length, These are the quaternions at the k-th sampling point.

[0241] By analyzing Δ qNAfter setting an appropriate threshold, the algorithm can accurately detect the start and end points of the gesture. To effectively reduce false detections caused by sudden user jerking, N needs to be set. In this embodiment, with a sampling frequency of 100Hz, the smoothing length is set to 10 and the threshold is set to 3.

[0242] Step 5.4: Input the processed data into the loaded Bi-LSTM model for recognition;

[0243] Step 5.5: Output the recognition results.

[0244] The invention will be further explained and illustrated below with examples:

[0245] Example 1

[0246] The subjects wore data gloves and collected environmental magnetic field data in the current environment by figure-eight motion. Magnetometer calibration was performed, and a single sensor was selected. The experimental results are shown in Figure 9. Before calibration, the center of the sphere fitted by the triaxial magnetometer was not at the origin due to the influence of the magnetic field. After calibration, the center of the sphere of the magnetometer was at the origin and the position was positive.

[0247] Example 2

[0248] Subjects wore data gloves, made a simple upward movement with their four fingers together, and then returned their fingers to be together. The pose calculation effects of the two algorithms were observed. The experimental results are shown in Figure 10. The experimental results show that the gradient descent data fusion method has a better pose calculation effect when the hand makes complex and delicate movements.

Claims

1. A sign language recognition method based on wearable computing, characterized by, The sign language recognition system based on wearable computing implements sign language recognition; the sign language recognition system based on wearable computing includes two inertial data gloves, a routing device and a human-computer interaction platform; Each inertial data glove includes sixteen inertial sensors and a wireless transmission module; the inertial sensors include a three-axis accelerometer, a three-axis gyroscope and a three-axis magnetometer, which are used to capture the motion information and position information of the palm and fingers of a single hand; The wireless transmission module completes interaction with the human-computer interaction platform through the transmission control protocol with the routing device as the hub; The human-computer interaction platform inputs the data collected by the inertial data gloves after error correction, attitude calculation and sign language sample segmentation into a neural network sign language recognition model, and finally displays the recognized gestures on the screen; The sign language recognition method based on wearable computing specifically includes the following steps: Step 1: make a sign language data set; Step 1.1: collect original inertial data of the sign language to be recognized, Use the inertial data gloves to collect original inertial data of 20 dynamic sign languages; Step 1.2: pre-process the collected sign language original inertial data, The pre-processing includes error correction, attitude calculation, sign language sample segmentation and data normalization of the collected original inertial data; Step 1.2.1: use the least squares method to solve the ellipsoid fitting problem to calibrate the three-axis magnetometer; ; wherein, represents a measured value of the tri-axial magnetometer, represents a theoretical value of the tri-axial magnetometer in the absence of errors, represents a zero offset error of the tri-axial magnetometer, represents a sensitivity error coefficient of each axis of the tri-axial magnetometer; represents a coupling coefficient of each axis of the tri-axial magnetometer; The ellipsoid equation is obtained by least square method fitting, and error coefficient is solved and zero bias error , , ; wherein, denotes an ellipsoidal shape matrix; denotes an ellipsoidal surface parameter to be solved; denotes a magnetic field intensity at a certain place; Step 1.2.2: the attitude calculation extracts the hand quaternion information from the original inertial data collected by the inertial data gloves using the gradient descent data fusion algorithm, and uses the following formula for iteration: , wherein, represents a rotation quaternion, represents a measurement error of the three-axis gyroscope, measured by a calibration algorithm; represents a sampling period; represents an objective function; ; Errors generated by tri-axial accelerometers are expressed as an error function is expressed as: , wherein, is a measurement value of a three-axis accelerometer; is an accelerometer X-axis error; is an accelerometer Y-axis error; is an accelerometer Z-axis error; is a rotation quaternion; Errors generated by tri-axial magnetometers using error functions are represented as: ; wherein, is a measurement value of a three-axis magnetometer; is a rotation quaternion; is a three-axis magnetometer X-axis error; is a three-axis magnetometer Y-axis error; is a three-axis magnetometer Z-axis error; Take the partial derivative of the objective function to get the Jacobian matrix ; solving the minimization of errors using the Gauss-Newton method, represents the target pose in the course of the calculation, ; wherein is a step size; denotes a differential operator, denotes a rate of change of the error vector, i.e. the relationship is as follows: , Finally, a quaternion representing the pose is solved, which contains two parts. One part is solved by solving the differential equation using the output value of the three-axis gyroscope. The other part is the modified quaternion solved by the objective function. Both describe the pose of the hand at the same time. Let the pose quaternion solved by solving the differential equation of the quaternion be The weight of the modified quaternion is represented by Combining the two parts of the quaternion, the final pose quaternion is represented as ; denotes the measurement error of the tri-axial gyroscope, measured by a calibration algorithm; denotes the sampling period; the weight has the following relationship with the step size has the following relationship: , The final attitude quaternion is represented as: ; Step 1.2.2.1: determine the geographic north direction using the corrected three-axis magnetometer signal in step 1.2.1, and obtain the initial attitude of the hand according to the three-dimensional acceleration signal when the hand is naturally drooping and the corrected three-dimensional magnetometer signal in step 1.2.1, and correct it; By solving the Euler angles of the sensor coordinate system relative to the carrier coordinate system, the pitch angle is obtained using the three-dimensional acceleration signal and the roll angle , as follows: , , wherein, are the outputs of the three axes of the triaxial accelerometer, respectively; combining the corrected three-dimensional magnetometer signals to obtain a heading angle : , wherein, is the projection of the X and Y axes of the three-dimensional magnetometer signal corrected to the local horizontal plane, the X axis pointing towards the Earth's magnetic North Pole; Step 1.2.2.2: convert the Euler angles obtained in step 1.2.2.1 into the unit quaternion from the carrier coordinate system to the sensor coordinate system using the conversion formula of the two coordinate systems: ; Step 1.2.2.3: determine the rotation quaternion of the sensor coordinate system and the geographic coordinate system according to the motion signal when the hand is naturally drooping and the corrected three-dimensional magnetometer signal, complete the calibration of the initial attitude of the hand and the initial hand model, including the following processes: The hand model joint coordinate position formula is: ; wherein, is the coordinate of the hand joint numbered in the geographical coordinate system at the time instant t; is the number of hand limb vectors; is the coordinate of the hand start joint, which refers to the coordinate of the palm center; The rotation of the hand joint vector in the geographic coordinate system is: , wherein, denotes the conjugate of the rotation quaternion of the body coordinate system to the geographical coordinate system; the hand joint vector is represented in the geographical coordinate system as ; the hand joint vector is represented in the body coordinate system as ; denotes quaternion multiplication; The final rotation formula of the hand joint vector in the geographic coordinate system is obtained: , wherein the initial The gravity vector and the magnetic north direction are determined from the three-dimensional acceleration data and the corrected three-dimensional magnetometer data. When performing initial hand attitude correction, the carrier coordinate system coincides with the geographic coordinate system by standing in the direction of the magnetic north pole when the hand is naturally drooping and stationary, and the formula at the initial moment is obtained: , wherein, represents a rotation quaternion of the hand frame coordinate system to the sensor coordinate system; represents a rotation quaternion of the hand geographic coordinate system to the sensor coordinate system; Step 1.2.2.4: finally update the hand motion attitude using the gradient descent data fusion algorithm to obtain real-time updated hand quaternion information; Step 1.2.3: The original acceleration, three-axis gyroscope data and the calculated hand posture quaternion information are combined as shown in the following formula; and the combined hand sign time series data is manually segmented and given a specific hand sign label by using a high-speed camera; ; wherein, is a three-axis acceleration signal collected by a three-axis accelerometer, is a three-dimensional angular velocity signal collected by a three-axis gyroscope, is a hand unit quaternion; Step 1.2.4: Data normalization is performed on the data given a specific hand sign label after segmentation to be in the 0~1 interval; Step 1.3: The hand sign time series data of each subject after preprocessing is divided into training set and test set in the ratio of 9:1; Step 2: Establish a neural network hand sign recognition model; Step 3: Train the neural network hand sign recognition model using the hand sign data set; Step 4: Test the generalization performance of the trained neural network hand sign recognition model; Step 5: Transplant the neural network hand sign recognition model that meets the test effect standard to the human-computer interaction platform, and recognize the hand sign in real time through the hand sign recognition system based on wearable computing.

2. The sign language recognition method based on wearable computing according to claim 1, characterized in that, The neural network hand sign recognition model is a bidirectional long short-term memory neural network model Bi-LSTM; step 2 specifically includes the following steps: Step 2.1: Determine the type of neural network hand sign recognition model The processed hand sign time series data is detected using the Bi-LSTM model; the Bi-LSTM model is divided into two LSTM layers, and the input sequence is input to the two LSTM layers in forward and reverse order for feature extraction; Step 2.2: Model structure and selection of model optimization function The Bi-LSTM model includes a time series input layer, a Bi-LSTM layer, a fully connected layer, a softmax layer, and a classification output layer; the model optimization function is selected as adma.

3. The sign language recognition method based on wearable computing according to claim 1 or 2, characterized in that, The accuracy, precision, recall and F1 score are used to evaluate the generalization performance of the neural network hand sign recognition model in step 4; The accuracy formula is as follows: ; The precision formula is as follows: ; The recall formula is as follows: ; The F1 score formula is as follows: ; where TP is true positive, FP is false positive, TN is true negative, and FN is false negative, is the harmonic mean of precision and recall.

4. The sign language recognition method based on wearable computing according to claim 3, characterized in that, Step 5 specifically includes the following steps: Step 5.1: Load the trained hand sign recognition model to the human-computer interaction platform; Step 5.2: Correct the three-axis magnetometer data, as in step 1.2.1; collect the data during random rotation in the figure eight shape as the sample for three-axis magnetometer correction, and use the least squares method for ellipsoid fitting correction; Step 5.2: The subject faces north and the hand naturally droops for 3-4 seconds to complete the initial posture calibration, and then the hand sign recognition starts; Step 5.3: The real-time collected hand sign data is subjected to posture calculation, automatic hand sign segmentation and data normalization; the automatic hand sign segmentation method is as follows: ; Wherein, N is the smoothing length, Respectively, the quaternion at the kth sampling point, by After setting the threshold, the gesture starting point and the end point are accurately detected. Step 5.4: The processed data is input into the loaded Bi-LSTM model for recognition; Step 5.5: Output the recognition result.

5. The sign language recognition method based on wearable computing according to claim 1, wherein, The training method in step 3 uses five-fold cross-validation.