Inertial sensor semantic information recognition method and system based on cyclic reconstruction network

By applying a method based on cyclic reconstruction network in inertial sensors, the motion trajectory is reconstructed and restored, and combined with semantic information recognition technology, the problems of low inertial sensor accuracy and large semantic information recognition error are solved, and the accurate recognition of specific trajectory motion is achieved.

CN114611554BActive Publication Date: 2025-05-13HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210223491.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-05-13
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

The accuracy of inertial sensors is poor, especially the inexpensive sensor signals contain a large number of random errors, making it difficult to accurately identify semantic information, especially when there are rich types of semantic information and large differences in motion habits.

Method used

Using a method based on cyclic reconstruction network, a multi-source signal cyclic reconstruction network is constructed, including a convolutional neural network model, a first wavelet encoding network model and a second wavelet encoding network model, the reconstruction of inertial sensor signals and the restoration of motion trajectory are carried out, and semantic information is recognized in combination with trajectory morphology information.

Benefits of technology

Accurate recognition of specific trajectory motions made by users in the air is achieved, and the accuracy and reliability of semantic information recognition of inertial sensors are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611554B_ABST
    Figure CN114611554B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for semantic information recognition of an inertial sensor based on a cyclic reconstruction network. The method comprises: constructing a multi-source signal cyclic reconstruction network; reconstructing the inertial sensor signal through a trained multi-source signal cyclic reconstruction network; restoring the motion trajectory based on the reconstructed inertial sensor signal; merging the restored motion trajectory with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments; merging the position information and time information in the feature vectors at each moment to obtain a coded quaternion; normalizing the coded quaternion to obtain a unit coded quaternion; constructing multiple feature quaternions based on acceleration components, angular velocity components and trajectory components; determining feature coding vectors based on unit coded quaternions and feature quaternions; inputting the feature coding vectors into a Transformer for semantic information recognition. The present invention can accurately identify the movement of a user in a specific trajectory in the air based on an inertial sensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic information recognition of inertial sensors, and in particular to a method and system for semantic information recognition of inertial sensors based on a cyclic reconstruction network. Background Art

[0002] MEMS inertial sensors have the advantages of small size, easy to wear, low power consumption, low cost, and easy mass production. The unit price of common inertial sensors on the market can be as low as 0.2-0.5 yuan. Therefore, human-computer interaction based on inertial sensors is not only more natural (compared with voice, image and other methods, it is not limited by distance and scene), but also has low requirements for hardware configuration, so it has huge market prospects. However, the accuracy of inertial sensors is poor, especially the cheap inertial sensor signals contain a large number of random errors, so it is difficult to accurately recognize the semantic information contained in the signal, especially when the semantic information is rich and the movement habits vary greatly. Using the motion data collected by the inertial sensor to recognize semantic information will produce very large errors. Summary of the invention

[0003] The purpose of the present invention is to provide a method and system for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network, so as to accurately recognize the movement of a user in a specific trajectory made in the air based on the inertial sensor.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network, comprising:

[0006] Constructing a multi-source signal cyclic reconstruction network; the multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model;

[0007] Reconstruct the inertial sensor signal through the trained multi-source signal recurrent reconstruction network;

[0008] Restore the motion trajectory based on the reconstructed inertial sensor signal;

[0009] The restored motion trajectory is combined with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments;

[0010] The position information and time information in the feature vector at each moment are combined to obtain the encoded quaternion;

[0011] Normalizing the coded quaternion to obtain a unit coded quaternion;

[0012] Based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment, multiple feature quaternions are constructed;

[0013] Determining a feature coding vector based on the unit coding quaternion and the feature quaternion;

[0014] The feature encoding vector is input into Transformer for semantic information recognition.

[0015] Optionally, the training process of the multi-source signal cyclic reconstruction network is as follows:

[0016] Determine the optimal wavelet transform mode for inertial sensor multi-axis signal data from multiple information sources;

[0017] Constructing a training set by annotating the corresponding inertial sensor multi-axis signal data through the optimal wavelet transform pattern;

[0018] The multi-source signal cyclic reconstruction network is trained using the training set.

[0019] Optionally, determining the optimal wavelet transform mode of the inertial sensor multi-axis signal data of multiple information sources specifically includes:

[0020] Perform variational mode decomposition on each axis signal of each information source to obtain the IMF component of each axis signal;

[0021] Perform multiple wavelet transforms on each IMF component;

[0022] The original IMF component is replaced by the wavelet transform result, and the original signal is reconstructed based on the replaced IMF component to obtain the reconstruction result of the original signal;

[0023] Performing posture calculation based on the reconstruction result;

[0024] The attitude solution result is compared with the real recorded data to obtain the attitude solution error; and the wavelet transform mode corresponding to the minimum attitude solution error is taken as the optimal wavelet transform mode.

[0025] Optionally, the training the multi-source signal cyclic reconstruction network by using the training set specifically includes:

[0026] The first IMF component decomposed from the first axis signal A1 of the first information source A After connecting it in parallel with the original signal X, connect a zero vector of equal length in parallel An input matrix is ​​formed; the original signal is multi-axis signal data of inertial sensors from multiple information sources;

[0027] The input matrix is ​​input into the convolutional neural network model, and the convolutional neural network model outputs a wavelet selection vector

[0028] The wavelet selection vector After inputting the softmax function, compared with the pre-obtained The cross entropy loss is calculated by the ont-hot encoding of the optimal wavelet transform mode, denoted as

[0029] The wavelet selection vector Input the first wavelet coding network model to get the wavelet coding of the current round Encode the wavelet of the previous round Multiply by weight ω1 and the wavelet code of the current round Add together to get the joint wavelet coding vector

[0030] The joint wavelet coded vector With the next IMF component And the original signal X in parallel to generate a new input matrix, again input the convolutional neural network model, get a new wavelet selection vector

[0031] The wavelet selection vector Input softmax function, and The corresponding optimal wavelet transform mode ont-hot encoding calculates the cross entropy loss, denoted as

[0032] The wavelet selection vector Input the first wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by weight ω2 and add to the wavelet code of the current round Add together to get the joint wavelet coding vector

[0033] Joint Wavelet Coding and the IMF component of the next round and the original signal X in parallel to form the input for the next round;

[0034] Repeat the above process until the last IMF component of information source A;

[0035] Assume that the last IMF component of information source A The corresponding wavelet selection vector is The wavelet selection vector Input softmax function, and The corresponding one-hot encoding of the optimal wavelet transform mode calculates the cross entropy loss, which is recorded as:

[0036] The wavelet selection vector Input the second wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by the weight ω 15 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round

[0037] The joint wavelet coded vector The first IMF component with the information source G And the original signal X in parallel to form an 8×N input matrix, and input into the convolutional neural network model to obtain a new wavelet selection vector

[0038] The wavelet selection vector Input softmax function, and The one-hot encoding of the corresponding optimal wavelet transform model calculates the cross entropy loss, denoted as

[0039] The wavelet selection vector is Input the second wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by the weight ω 16 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round

[0040] The wavelet coded vector The second IMF component of the information source G and the original signal X in parallel to form an 8×N input matrix, which is then input into the convolutional neural network model. The above process is repeated until all IMF components are operated on.

[0041] A corresponding wavelet selection vector and a corresponding loss term are obtained for each IMF component, and all loss terms are added together to obtain the total wavelet selection loss of the multi-source signal cyclic reconstruction network;

[0042] The total loss is back-propagated to implement the training of the multi-source signal cyclic reconstruction network.

[0043] Optionally, after the multi-source signal cyclic reconstruction network is trained using the training set, the method further includes: regularizing the trained first wavelet coding network model and the second wavelet coding network model.

[0044] Optionally, the regularization term R of the first wavelet coding network model is A for:

[0045]

[0046] The auxiliary regularization term of the first wavelet coding network model for:

[0047]

[0048] The regular term R of the second wavelet coding network model is G for:

[0049]

[0050] Auxiliary regularization term of the second wavelet coding network model for:

[0051]

[0052] in, Represents the first wavelet encoding matrix W A The α-order Renyi information entropy of Respectively represent the coding vectors of the i-th and j-th wavelet transform modes of the first wavelet coding matrix, Represents the second wavelet encoding matrix W G The α-order Renyi information entropy of Respectively represent the coding vectors of the i-th and j-th wavelet transform modes of the first wavelet coding matrix

[0053] Optionally, the determining the feature coding vector based on the unit coding quaternion and the feature quaternion specifically includes:

[0054] Multiplying the unit encoding quaternion by a learning matrix to obtain a quaternion;

[0055] Multiplying the quaternion by the feature quaternion to obtain a feature encoding quaternion;

[0056] The feature coding quaternions are concatenated to obtain a feature coding vector.

[0057] Optionally, the determining the feature coding vector based on the unit coding quaternion and the feature quaternion specifically includes:

[0058] The unit coding quaternion is applied to the feature quaternion by quaternion multiplication to obtain a feature coding vector.

[0059] The present invention also provides an inertial sensor semantic information recognition system based on a cyclic reconstruction network, comprising:

[0060] A network construction module is used to construct a multi-source signal cyclic reconstruction network; the multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model;

[0061] A signal reconstruction module is used to reconstruct the inertial sensor signal through a trained multi-source signal recurrent reconstruction network;

[0062] A trajectory restoration module, used to restore the motion trajectory based on the reconstructed inertial sensor signal;

[0063] A first merging module is used to merge the restored motion trajectory with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments;

[0064] The second merging module is used to merge the position information and the time information in the feature vector at each moment to obtain a coded quaternion;

[0065] A normalization processing module, used for performing normalization processing on the coded quaternion to obtain a unit coded quaternion;

[0066] A feature quaternion construction module is used to construct multiple feature quaternions based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment;

[0067] A feature coding vector determination module, used to determine a feature coding vector based on the unit coding quaternion and the feature quaternion;

[0068] The semantic information recognition module is used to input the feature encoding vector into the Transformer for semantic information recognition.

[0069] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0070] The present invention reconstructs signals through a constructed multi-source signal cyclic reconstruction network to restore the motion trajectory, and then integrates the trajectory morphology information into the semantic information recognition task of gesture actions based on inertial sensors, so as to accurately recognize the specific trajectory movements made by the user in the air based on the inertial sensor. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0072] Figure 1 It is a flowchart of a method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to an embodiment of the present invention;

[0073] Figure 2 Reconstruct flow charts for model-driven signals;

[0074] Figure 3 Reconstruct network structure for multi-source signal cycles;

[0075] Figure 4 This is the trajectory restoration result obtained after data enhancement;

[0076] Figure 5 It is a one-dimensional grating position encoding;

[0077] Figure 6 The input data has multimodal characteristics;

[0078] Figure 7 Encode the quaternion position; DETAILED DESCRIPTION

[0079] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0080] The purpose of the present invention is to provide a method and system for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network, so as to accurately recognize the movement of a user in a specific trajectory made in the air based on the inertial sensor.

[0081] Semantic information recognition based on inertial sensors is essentially a time series analysis technology. Most existing time series analysis technologies rely on signal features in the time domain or frequency domain, which can be attributed to "waveform features". In the task of semantic information recognition of actions, the difference in semantic information of different actions is essentially determined by the trajectory morphology corresponding to different actions, so the morphological information of the trajectory is extremely important. High-precision restoration of inertial sensor trajectories and integration of trajectory morphological information into downstream tasks based on inertial sensors, especially semantic information recognition tasks of gesture actions, this strategy will be a very valuable solution in the field of inertial sensors.

[0082] In order to achieve accurate trajectory restoration, the present invention designs a "multi-source signal cyclic reconstruction network", which can significantly improve the accuracy of inertial sensor data in attitude solution and trajectory restoration tasks.

[0083] In order to utilize the trajectory information obtained by solving and the timing information of the inertial sensor, the present invention designs a quaternion position encoding, and thereby improves the Transformer structure to form a new Quaterformer model.

[0084] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0085] like Figure 1 As shown, the present invention provides an inertial sensor semantic information recognition method based on a cyclic reconstruction network, comprising the following steps:

[0086] Step 101: construct a multi-source signal cyclic reconstruction network; the multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model.

[0087] Step 102: Reconstruct the inertial sensor signal through the trained multi-source signal cyclic reconstruction network. The training process of the multi-source signal cyclic reconstruction network is as follows: determine the optimal wavelet transform mode of the inertial sensor multi-axis signal data of multiple information sources; construct a training set by annotating the corresponding inertial sensor multi-axis signal data through the optimal wavelet transform mode; and train the multi-source signal cyclic reconstruction network through the training set.

[0088] Step 103: restore the motion trajectory based on the reconstructed inertial sensor signal.

[0089] Step 104: Combine the restored motion trajectory with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments.

[0090] Step 105: Combine the position information and time information in the feature vector at each moment to obtain a coded quaternion.

[0091] Step 106: normalize the coded quaternion to obtain a unit coded quaternion.

[0092] Step 107: construct multiple feature quaternions based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment.

[0093] Step 108: Determine a feature coding vector based on the unit coding quaternion and the feature quaternion, which specifically includes: multiplying the unit coding quaternion by a learning matrix to obtain a quaternion; multiplying the quaternion by the feature quaternion to obtain a feature coding quaternion; and concatenating the feature coding quaternions to obtain a feature coding vector.

[0094] Step 109: Input the feature encoding vector into Transformer for semantic information recognition.

[0095] Among them, step 102 of using a multi-source signal cyclic reconstruction network to realize inertial sensor signal reconstruction can be divided into two steps: model-driven signal reconstruction (i.e., data annotation to build a training set) and data-driven signal reconstruction (i.e., model training), the former providing the latter with the data set required for training.

[0096] (1) Model-driven signal reconstruction

[0097] Model-driven signal reconstruction schemes such as Figure 2 As shown in the figure, the inertial sensor is placed on a turntable and rotated multiple times. At the end of each movement, the turntable is used to record the changes in the sensor posture during this movement. Assume that during a certain movement, the 6-axis raw signal samples collected by the inertial sensor are in is the acceleration signal, is the angular velocity signal, then the model-driven reconstruction process of signal X is as follows:

[0098] 1) Input multi-axis signal data from multiple information sources, such as Figure 2 As shown, the input 3-axis accelerometer signal And 3-axis angular velocity sensor signal

[0099] 2) For each axis signal of each information source The variational mode decomposition (VMD) is performed separately, and the number of modes is set to J (in this problem, J is set to 5), so the information source The i-th axis signal is decomposed into J intrinsic mode functions (IMF components):

[0100] 3) For each IMF component obtained in turn 15 kinds of wavelet transforms are performed respectively (different wavelet bases are regarded as a wavelet transform mode. The present invention sets 15 different wavelet basis functions, which are: 'bior1.1', 'bior1.3', 'bior1.5', 'bior2.2', 'bior2.4', 'bior2.6', 'bior2.8', 'bior3.1', 'bior3.3', 'bior3.5', 'bior3.7', 'bior3.9', 'bior4.4', 'bior5.5', 'bior6.8'). In addition, not performing wavelet transform on the IMF component is also a result of wavelet transform, such as Figure 2 At this time, the present invention obtains any IMF component 16 wavelet transform results (i=1,2,3; j=1,2,3…J);

[0101] 4) Replace the original IMF components with these 16 wavelet transform results Then the original signals are reconstructed respectively to obtain 16 reconstruction results of the original signals.

[0102] 5) Use the 16 original signal reconstruction results to perform attitude calculations respectively (there are mature algorithms for attitude calculations, see related inertial guidance theories), and obtain 16 attitude calculation results. Compare these 16 attitude calculation results with the actual attitude changes recorded when measuring data, and obtain the attitude calculation errors of the 16 signals;

[0103] 6) The wavelet transform mode β corresponding to the minimum attitude error k (β k ∈{'bior1.1','bior1.3','bior1.5','bior2.2','bior2.4','bior2.6','bior2.8','bior3.1','bior3.3','bior3.5','bior3.7','bior3.9','bior4.4','bior5.5','bior6.8','null'}, k=1,2,3…16) and replace the original IMF components with the processing results of the wavelet on the IMF components;

[0104] 7) Repeat steps 2) to 6) to continue operating on the next IMF component.

[0105] After model-driven signal reconstruction, for a set of dual-source triaxial signals The present invention decomposes a total of 30 IMF components (30 = 2 × 3 × 5, dual-source three-axis signal, each axis decomposes 5 IMF components), each IMF component corresponds to an optimal wavelet transform mode β k (β k ∈{'bior1.1','bior1.3','bior1.5','bior2.2','bior2.4','bior2.6','bior2.8','bior3.1','bior3.3','bior3.5','bior3.7','bior3.9','bior4.4','bior5.5','bior6.8','null'}, k=1,2,3…16). Therefore, these 30 IMF components and the corresponding 30 optimal wavelet transform modes constitute a set of "input-output matching relationships": in Through model-driven signal reconstruction, the present invention obtains a total of 7000 sets of such "input-output matching relationships".

[0106] (2) Data-driven signal reconstruction

[0107] The purpose of model-driven signal reconstruction is to obtain the most ideal wavelet transform type for each IMF component. Data-driven signal reconstruction is to train a deep learning model (here the present invention selects a one-dimensional convolutional neural network 1DCNN) to learn how to select the most appropriate wavelet transform type for different IMF components. The 7000 sets of "input-output matching relationships" obtained previously will be used as training sets to train the neural network model (deep learning model) designed by the present invention.

[0108] In order to achieve accurate prediction of the ideal wavelet patterns of different IMF components and thus better reconstruct the original signal, the present invention designs a "multi-source signal recurrent reconstruction network (RRN)", such as Figure 3 As shown, the dual-source triaxial signal is decomposed into 30 IMF components. The operation for each IMF component (1×N, N is the signal length) is as follows:

[0109] 1) The first IMF component decomposed from the first axis signal A1 of the first information source A After connecting it in parallel with the original signal X(6×N), a zero vector of equal length is connected in parallel. At this point, an 8×N input matrix can be formed. The role of parallel equal-length zero vectors is to ensure that the data structure can be generated in the end as an 8×N matrix, so as to input the subsequent network structure;

[0110] 2) The 8×N matrix is ​​input into a one-dimensional convolutional neural network model (1DCNN), and the model outputs a 1×16 wavelet selection vector This vector represents the importance of the 16 wavelet transform modes determined by the model based on the input IMF components and the original data X;

[0111] 3) Select the wavelet vector After inputting the softmax function, the cross entropy loss is calculated with the ont-hot encoding of the optimal wavelet transform mode of the IMF component obtained in advance by the present invention, which is recorded as

[0112] 4) Select the wavelet vector Input the first wavelet encoding network (a single-layer neural network W A ), that is, wavelet selection vector Multiply by matrix W A (16×N), get the wavelet code of the current round Encode the wavelet of the previous round Multiply it by a learnable weight ω1 and add it to the wavelet code of the current round Add together to get the joint wavelet coding vector before the current round

[0113]

[0114] It is worth noting that: is a sparse vector, so the matrix W A Each row will be trained as the encoding vector corresponding to a different wavelet transform mode. Therefore, the joint wavelet encoding vector The meaning of is: to summarize the wavelet selection results of this round and all previous IMF components in the form of coding;

[0115] 5) The joint wavelet coding vector With the next IMF component And the original signal X(6×N) in parallel to generate a new 8×N input matrix, which is input into the 1DCNN model again to obtain a new wavelet selection vector

[0116] 6) Select the wavelet vector Input the softmax function and calculate the cross entropy loss with the ont-hot encoding of the optimal wavelet transform mode of the IMF component, denoted as

[0117] 7) Select the wavelet vector Input the first wavelet encoding network (a single-layer neural network W A), that is, wavelet selection vector Multiply by matrix W A (16×N), get the wavelet code of the current round The joint wavelet coding vector of the previous round Multiply it by a learnable weight ω2 and add it to the wavelet code of the current round Add together to get the joint wavelet coding vector before the current round The joint wavelet code of the current round and the IMF component of the next round and the original signal X in parallel to form the input for the next round;

[0118] 8) Repeat the above process until the last IMF component of information source A is

[0119] 9) Let the last IMF component of information source A be The corresponding wavelet selection vector is The wavelet selection vector Input the softmax function and calculate the cross entropy loss with the one-hot encoding of the optimal wavelet transform mode of the IMF component, which is recorded as:

[0120] 10) Select the wavelet vector Input the second wavelet encoding network W G (A single-layer neural network W G ), that is, wavelet selection vector Multiply by matrix W G (16×N), get the wavelet code of the current round The joint wavelet coding vector of the previous round Multiply by a learnable weight ω 15 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round

[0121] 11) The joint wavelet coding vector The first IMF component with the information source G And the original signal X in parallel to form an 8×N input matrix, which is input into the 1DCNN model to obtain a new wavelet selection vector

[0122] 12) Select the wavelet vector Input the softmax function and calculate the cross entropy loss with the one-hot encoding of the optimal wavelet transform model of the IMF component, denoted as

[0123] 13) The wavelet selection vector is Input wavelet coding network W G , that is, wavelet selection vector Multiply by matrix W G (16×N), get the wavelet code of the current round The joint wavelet coding vector of the previous round Multiply by a learnable weight ω 16 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round

[0124] 14) Continue to transform the wavelet encoding vector The second IMF component of the information source G And the original signal X are connected in parallel to form an 8×N input matrix, which is input into the 1DCNN model. The above process is repeated until all IMF components are operated on;

[0125] 15) Finally, for each IMF component A corresponding wavelet selection vector can be obtained and a corresponding loss term Adding up all 30 loss items, we can get the total wavelet selection loss Loss of the cyclic reconstruction network. classification :

[0126]

[0127] 16) The loss is back-propagated (existing technology), which can realize the three sub-networks in the cyclic reconstruction network (one-dimensional convolutional neural network 1DCNN, wavelet coding network W A , wavelet coding network W G ) training.

[0128] In order to make the two wavelet coding networks W A With W G The 16 wavelet transform modes can be better encoded in training, and the present invention needs to regularize them. "Better encoding" should mainly meet two requirements:

[0129] (1) The encoding of different wavelet transform modes should reflect the differences between different modes as much as possible, so different encoding vectors should be as orthogonal as possible;

[0130] (2) Each encoding vector should store as much information as possible;

[0131] Below is W ATake the example to illustrate the design of the regularization term. In order to make the trained wavelet coding network (wavelet coding matrix) W A In order to meet the above requirements as much as possible, the present invention uses the α-order Renyi information entropy of the matrix To measure W A Since each row of the wavelet coding matrix represents a coding vector of wavelet transform, and They represent the kth row of the two matrices, i.e., the encoding vector of the kth wavelet transform mode, k = 1, 2, 3, ..., 16. Thus, the matrix W is obtained. A The α-order Renyi information entropy expression:

[0132]

[0133]

[0134] in, Representation Matrix The i-th eigenvalue of Representation Matrix In similar problems, α is usually taken as 2.

[0135] It should be noted that The larger the matrix W is, the A The more information it contains, the more and The smaller the inner product of the two, the changes of which are exactly the same as the wavelet coding network (matrix) W of the present invention. A Therefore, the present invention changes the regularization term R A Set to:

[0136]

[0137] When R A Reduced, the wavelet coding matrix W A The amount of information increases, that is, each row of wavelet coding vector The amount of information contained increases. At the same time, different rows of the matrix (i.e. different wavelet coding vectors) and The inner product between them just decreases, and they become more orthogonal to each other.

[0138] However, in the actual training process, the regular term R A It is difficult to reduce to the ideal level according to the concept of the present invention, so the present invention continues to add auxiliary regularization terms to the matrix

[0139]

[0140] It can be found that the auxiliary regularization term Essentially a matrix L1 norm of non-diagonal elements. This operation can reduce the inner product between different rows of the wavelet coding matrix at a faster speed and can reduce it to an ideal degree.

[0141] Similarly, the matrix W G Information entropy S α (W G ) is expressed as follows:

[0142]

[0143]

[0144] The regularization term corresponding to the information source G is:

[0145]

[0146]

[0147] In summary, the final loss function of the cyclic reconstruction network is Loss Final The expression is as follows:

[0148]

[0149] where γ h (h=1, 2, 3, 4) are artificially set weights used to control the proportion of the four regularization terms in the total loss function.

[0150] After training, the recurrent encoding network can output the optimal wavelet change pattern corresponding to each IMF component after variational mode decomposition of different axon signals of different source original signals. After performing the optimal wavelet transform on each IMF component, the original signal can be reconstructed. The reconstructed inertial sensor signal It can achieve more accurate motion trajectory restoration, such as Figure 4 shown.

[0151] The ultimate goal of the multi-source signal recurrent reconstruction network in the previous section is to obtain accurate motion trajectories through data enhancement. The reason for obtaining accurate motion trajectories as much as possible is that in the task of semantic information recognition of gesture actions, the core and intuitive information is the trajectory corresponding to different gesture actions. Different actions are defined by the geometric shapes of different trajectories, rather than by the original acceleration and angular velocity waveforms measured by inertial sensors. Therefore, in order to achieve high-precision motion semantic information recognition, it is far from enough to simply input acceleration and angular velocity data into the classifier. Combining the motion trajectory information with the acceleration and angular velocity data measured by the sensor is the standard paradigm for future inertial sensor motion semantic recognition and human-computer interaction control.

[0152] Therefore, the present invention realizes data enhancement by designing a multi-source signal cyclic reconstruction network, and restores the three-dimensional motion trajectory based on the enhanced data with the help of inertial guidance theory. The present invention connects the three-dimensional motion trajectory S in parallel with the six-axis inertial sensor signal X to form a 9-axis time series signal That is, at each time t (t = 1, 2, 3, ..., T), there is a 1×9 eigenvector f t , which includes 3D acceleration, 3D angular velocity and 3D trajectory coordinates.

[0153] To identify the signal The semantic information contained in the signal is used in this invention. The Transformer structure is used as a classifier. However, when the signal When the Transformer structure is input, the Transformer structure converts the feature vector f at each moment in the signal t The quaternion position encoding is designed and integrated with the Transformer structure to form a Quaterformer model, which significantly improves the Transformer structure's ability to recognize semantic information of inertial sensor data.

[0154] Among them, step 105 to step 108 specifically include:

[0155] Traditional position coding can only provide one-dimensional position coding for features. However, this position coding has obvious defects. First, for two-dimensional or even higher-dimensional features, one-dimensional position coding obviously cannot provide enough position information for them. In fact, using one-dimensional position coding to encode high-dimensional data features will introduce obvious errors, such as Figure 5 As shown. For the input data, A 11 With A 12 , A 11 With A 21The spatial distances of A and B are all 1, but the position codes they obtain are "1", "2" and "5" respectively. Under this code, A 11 With A 12 , A 11 With A 21 The spatial distances are 1 and 4 respectively, which is obviously inconsistent with the actual situation.

[0156] The second major problem with traditional positional encoding is that it can only focus on a specific type of relationship between features. Figure 5 The one-dimensional grating position encoding shown is used as an example to illustrate that this encoding method can only focus on the neighbor relationship of the input data in terms of position (there are a lot of wrong encodings only in this neighbor relationship, as mentioned above). If there are other types of relationships between the input features in addition to the position relationship, such as time relationship, density relationship, energy relationship, and sparsity relationship, then this encoding method can only encode one of them, such as Figure 6 This is the second major flaw in this encoding.

[0157] In order to solve this problem, the present invention proposes quaternion position encoding. A quaternion can represent a posture in three-dimensional space or a posture rotation mode in three-dimensional space. Using quaternions for feature encoding is essentially to give the feature a spatial posture, and features with close positions have similar postures. The posture information in three-dimensional space can accommodate multiple categories of feature relationships, and the relative relationship between each feature can be maintained during the quaternion encoding process.

[0158] The coded information known in the present invention has two categories: one-dimensional time information t (t = 1, 2, 3, ..., N), three-dimensional space information [S x , S y , S z ]. They can just form a set of quaternions. After normalizing them, the present invention obtains the "unit coded quaternion" q = [q0, q1, q2, q3] generated by the known coding information, which will act on the feature vector f to be encoded in the form of quaternion multiplication. t .

[0159] The feature vector f to be encoded t It is mainly composed of three signals: acceleration signal f t [1:3]=[A x , A y , A z ], angular velocity signal f t [4:6]=[G x , G y , G z]; Fill "1" in front of each signal vector to form three feature quaternions. Apply the "unit coded quaternion" to the three feature quaternions in the form of "quaternion multiplication" to obtain three "feature coded quaternions", which can be concatenated to obtain the final 1×12 feature coded vector.

[0160] However, the multiplication rules of quaternions are relatively complicated. Suppose two quaternions are q = [q0, q1, q2, q3] and f = [f0, f1, f2, f3]. Apply q to f using quaternion multiplication ☆, and the result can be expressed as:

[0161]

[0162] It can be found that the calculation process is relatively complicated, and when there is a deviation in the encoding information, the encoding method is difficult to make adaptive adjustments. Therefore, the present invention sets a learnable weight matrix W between the two quaternions, and replaces the quaternion multiplication with matrix multiplication, that is, q☆f=q×W×f, where "☆" represents quaternion multiplication and "×" represents matrix multiplication. Figure 7 shown.

[0163] In summary, for a set of raw 6-axis data from an inertial sensor containing specific semantic information The specific identification process can be summarized as follows:

[0164] (1) The original 6-axis data X of the inertial sensor is enhanced through the first part (multi-source signal cyclic reconstruction network), and based on the enhanced inertial sensor data Perform motion trajectory restoration (existing technology), and the obtained 3D trajectory information is recorded as S;

[0165] (2) The calculated 3D motion trajectory S is combined with the enhanced inertial sensor data Merge and transpose the 9×T signal to obtain a T×9 signal.

[0166] (3) The signal Input the Quaterformer proposed in the present invention, and the specific calculation process is as follows:

[0167] 1) Construct the encoding information: For each feature vector f t , the location information contained in [S x , S y , S z ] Take it out and combine it with the time information t to form the encoded quaternion

[0168] 2) Encoded quaternion Normalize to obtain the unit coded quaternion q = [q0, q1, q2, q3];

[0169] 3) The feature vector f t The acceleration component f t [1:3]=[A x , A y , A z ], angular velocity component f t [4:6]=[G x , G y , G z ], and the trajectory component f t [7:9]=[S x , S y , S z ] are taken out separately, and "1" is filled in front of each vector to form three characteristic quaternions: q A =[1,A x , A y , A z ], q G =[1, G x , G y , G z ], q S =[1, S x , S y , S z ];

[0170] 4) To encode the quaternion q, multiply it by three learnable matrices W A , W G , W S , multiply the three quaternions obtained by the corresponding characteristic quaternion q A ,q G ,q S , get 3 feature encoding quaternions

[0171] 5) Encode the three features into quaternions Splice to get the final feature encoding vector at time t

[0172] 6) Calculate the feature vectors at T moments according to the operations described in 1) to 5) respectively, and input the obtained results into the traditional Transformer to complete the final semantic information recognition.

[0173] The main advantages of the present invention are as follows:

[0174] 1. A general research paradigm in the field of inertial sensor research is proposed, namely:

[0175] Data enhancement → trajectory restoration → signal waveform and trajectory morphological feature fusion → downstream tasks

[0176] The traditional research approach is:

[0177] Waveform signal features → downstream tasks

[0178] This scheme introduces the focus on trajectory morphology information, which is often the core of downstream tasks. For example, in the downstream task of action semantic information recognition, the difference in semantic information mainly comes from the difference in the trajectory morphology of different actions. Therefore, this research paradigm has certain promotion value.

[0179] 2. A multi-source signal recurrent reconstruction network was designed. This model can significantly enhance the performance of multi-source time series signals in specific tasks and is also highly valuable in other tasks besides inertial sensor semantic information recognition;

[0180] 3. The quaternion position encoding was designed and integrated with the Transformer structure to form the Quaterformer model, which significantly improved the Transformer model's ability to extract features from inertial sensor data. The quaternion position encoding breaks the disadvantage that the existing position encoding method in the field of deep learning research is only effective for a single type of low-dimensional feature relationship, and has wide application value in many tasks.

[0181] The present invention also provides an inertial sensor semantic information recognition system based on a cyclic reconstruction network, comprising:

[0182] A network construction module is used to construct a multi-source signal cyclic reconstruction network; the multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model;

[0183] A signal reconstruction module is used to reconstruct the inertial sensor signal through a trained multi-source signal recurrent reconstruction network;

[0184] A trajectory restoration module, used to restore the motion trajectory based on the reconstructed inertial sensor signal;

[0185] A first merging module is used to merge the restored motion trajectory with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments;

[0186] The second merging module is used to merge the position information and the time information in the feature vector at each moment to obtain a coded quaternion;

[0187] A normalization processing module, used for performing normalization processing on the coded quaternion to obtain a unit coded quaternion;

[0188] A feature quaternion construction module is used to construct multiple feature quaternions based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment;

[0189] A feature coding vector determination module, used to determine a feature coding vector based on the unit coding quaternion and the feature quaternion;

[0190] The semantic information recognition module is used to input the feature encoding vector into the Transformer for semantic information recognition.

[0191] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0192] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for semantic information recognition of inertial sensors based on a cyclic reconstruction network, characterized in that: include: Construct a multi-source signal cyclic reconstruction network; The multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model; Reconstruct the inertial sensor signal through the trained multi-source signal recurrent reconstruction network; Restore the motion trajectory based on the reconstructed inertial sensor signal; The restored motion trajectory is combined with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments; The position information and time information in the feature vector at each moment are combined to obtain the encoded quaternion; Normalizing the coded quaternion to obtain a unit coded quaternion; Based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment, multiple feature quaternions are constructed; Determining a feature coding vector based on the unit coding quaternion and the feature quaternion; The feature encoding vector is input into Transformer for semantic information recognition.

2. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 1, characterized in that: The training process of the multi-source signal cyclic reconstruction network is as follows: Determine the optimal wavelet transform mode for inertial sensor multi-axis signal data from multiple information sources; Constructing a training set by annotating the corresponding inertial sensor multi-axis signal data through the optimal wavelet transform pattern; The multi-source signal cyclic reconstruction network is trained using the training set.

3. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 2, characterized in that: The determining of the optimal wavelet transform mode of the inertial sensor multi-axis signal data of the multiple information sources specifically includes: Perform variational mode decomposition on each axis signal of each information source to obtain the IMF component of each axis signal; Perform multiple wavelet transforms on each IMF component; The original IMF component is replaced by the wavelet transform result, and the original signal is reconstructed based on the replaced IMF component to obtain the reconstruction result of the original signal; Performing posture calculation based on the reconstruction result; The attitude solution result is compared with the real recorded data to obtain the attitude solution error; and the wavelet transform mode corresponding to the minimum attitude solution error is taken as the optimal wavelet transform mode.

4. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 2, characterized in that: The training of the multi-source signal cyclic reconstruction network by using the training set specifically includes: The first IMF component decomposed from the first axis signal A1 of the first information source A After connecting it in parallel with the original signal X, connect a zero vector of equal length in parallel An input matrix is ​​formed; the original signal is multi-axis signal data of inertial sensors from multiple information sources; The input matrix is ​​input into the convolutional neural network model, and the convolutional neural network model outputs a wavelet selection vector The wavelet selection vector After inputting the softmax function, compared with the pre-obtained The cross entropy loss is calculated by the ont-hot encoding of the optimal wavelet transform mode, denoted as The wavelet selection vector Input the first wavelet coding network model to get the wavelet coding of the current round Encode the wavelet of the previous round Multiply by weight ω1 and the wavelet code of the current round Add together to get the joint wavelet coding vector The joint wavelet coded vector With the next IMF component And the original signal X in parallel to generate a new input matrix, again input the convolutional neural network model, get a new wavelet selection vector The wavelet selection vector Input softmax function, and The corresponding optimal wavelet transform mode ont-hot encoding calculates the cross entropy loss, denoted as The wavelet selection vector Input the first wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by weight ω2 and add to the wavelet code of the current round Add together to get the joint wavelet coding vector Joint Wavelet Coding and the IMF component of the next round and the original signal X in parallel to form the input for the next round; Repeat the above process until the last IMF component of information source A; Assume that the last IMF component of information source A The corresponding wavelet selection vector is The wavelet selection vector Input softmax function, and The corresponding one-hot encoding of the optimal wavelet transform mode calculates the cross entropy loss, which is recorded as: The wavelet selection vector Input the second wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by the weight ω 15 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round The joint wavelet coded vector The first IMF component with the information source G And the original signal X in parallel to form an 8×N input matrix, and input into the convolutional neural network model to obtain a new wavelet selection vector The wavelet selection vector Input softmax function, and The one-hot encoding of the corresponding optimal wavelet transform model calculates the cross entropy loss, denoted as The wavelet selection vector is Input the second wavelet coding network model to get the wavelet coding of the current round The joint wavelet coding vector of the previous round Multiply by the weight ω 16 Wavelet coding of the next and current rounds Add together to get the joint wavelet coding vector before the current round The wavelet coded vector The second IMF component of the information source G and the original signal X in parallel to form an 8×N input matrix, which is then input into the convolutional neural network model. The above process is repeated until all IMF components are operated on. A corresponding wavelet selection vector and a corresponding loss term are obtained for each IMF component, and all loss terms are added together to obtain the total wavelet selection loss of the multi-source signal cyclic reconstruction network; The total loss is back-propagated to implement the training of the multi-source signal cyclic reconstruction network.

5. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 2, characterized in that: After the multi-source signal cyclic reconstruction network is trained by the training set, the method further includes: regularizing the trained first wavelet coding network model and the second wavelet coding network model.

6. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 5, characterized in that: The regular term R of the first wavelet coding network model is A for: The auxiliary regularization term of the first wavelet coding network model for: The regular term R of the second wavelet coding network model is G for: Auxiliary regularization term of the second wavelet coding network model for: in, Represents the first wavelet encoding matrix W A The α-order Renyi information entropy of Respectively represent the coding vectors of the i-th and j-th wavelet transform modes of the first wavelet coding matrix, Represents the second wavelet encoding matrix W G The α-order Renyi information entropy of Respectively represent the coding vectors of the i-th and j-th wavelet transform modes of the second wavelet coding matrix.

7. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 1, characterized in that: The determining of the feature coding vector based on the unit coding quaternion and the feature quaternion specifically includes: Multiplying the unit encoding quaternion by a learning matrix to obtain a quaternion; Multiplying the quaternion by the feature quaternion to obtain a feature encoding quaternion; The feature coding quaternions are concatenated to obtain a feature coding vector.

8. The method for recognizing semantic information of an inertial sensor based on a cyclic reconstruction network according to claim 1, characterized in that: The determining of the feature coding vector based on the unit coding quaternion and the feature quaternion specifically includes: The unit coding quaternion is applied to the feature quaternion by quaternion multiplication to obtain a feature coding vector.

9. An inertial sensor semantic information recognition system based on a cyclic reconstruction network, characterized in that: include: A network construction module, used to construct a multi-source signal cyclic reconstruction network; The multi-source signal cyclic reconstruction network includes: a convolutional neural network model, a first wavelet coding network model and a second wavelet coding network model; A signal reconstruction module is used to reconstruct the inertial sensor signal through a trained multi-source signal recurrent reconstruction network; A trajectory restoration module, used to restore the motion trajectory based on the reconstructed inertial sensor signal; A first merging module is used to merge the restored motion trajectory with the reconstructed inertial sensor signal to obtain feature vectors at multiple moments; The second merging module is used to merge the position information and the time information in the feature vector at each moment to obtain a coded quaternion; A normalization processing module, used for performing normalization processing on the coded quaternion to obtain a unit coded quaternion; A feature quaternion construction module is used to construct multiple feature quaternions based on the acceleration component, angular velocity component and trajectory component in the feature vector at each moment; A feature coding vector determination module, used to determine a feature coding vector based on the unit coding quaternion and the feature quaternion; The semantic information recognition module is used to input the feature encoding vector into the Transformer for semantic information recognition.

Citation Information

Patent Citations

  • Football action recognition and evaluation system and method based on wearable equipment and machine learning

    CN111744156A

  • Accelerometer fault diagnosis method based on convolutional neural network

    CN112329650A