A motion analysis system based on inertial-visual signal enhancement and fusion
By introducing the Kolmogorov-Arnold network and a deep learning model for generating illusion entropy, combined with the feature alignment method of Sharma-MIttal entropy and optimal transport theory, the problems of inertial sensor signal saturation and multimodal data fusion were solved, achieving high-precision motion analysis and capture.
Patent Information
- Application Number
- CN202411322461.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-09-23
AI Technical Summary
In highly dynamic environments, signal saturation of inertial sensors leads to reduced signal reconstruction accuracy and reliability, and the challenges of multimodal sensor data fusion affect motion capture accuracy.
A deep learning model based on Kolmogorov-Arnold network for signal reconstruction and generation of illusion entropy is adopted. It combines the rare memory enhancement module of Sharma-MIttal entropy and the multimodal feature alignment method of optimal transport theory. Signal processing and feature alignment are performed through an inertial-visual information fusion module, and motion trajectory reconstruction is constrained by KANIadakIs entropy regularization.
It improves the accuracy and reliability of inertial sensor signal reconstruction, optimizes multimodal sensor data fusion, and achieves high-precision motion capture and analysis.
Smart Images

Figure CN119245634B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent sensing and fusion, in particular to a motion analysis system based on inertial-visual signal enhancement and fusion. BACKGROUND
[0002] Inertial sensors play a crucial role in modern technology applications, such as in automotive, aerospace, robotics, and golfing. Inertial sensors are devices that can measure an object's acceleration, angular velocity, and magnetic field strength, typically including accelerometers, gyroscopes, and magnetometers. They capture information about an object's motion in space and are widely used in motion capture, navigation systems, and attitude control. In motion capture, inertial sensors are particularly important as they can track and record an object's motion trajectory in real-time, providing key data for analyzing motion behavior and constructing motion models. Through inertial sensors, researchers and engineers can obtain precise motion data, enabling highly realistic motion capture in virtual reality, animation production, and biomechanics research.
[0003] In high dynamic environments, inertial sensors often face signal saturation problems, which limit the application range of low-cost inertial sensors. Traditional signal processing methods usually treat signals as discrete point sequences, resulting in significant information loss during signal reconstruction, especially in the aspect of super-range reconstruction. In addition, deep learning models may produce "hallucinations" that do not conform to actual situations during signal reconstruction, affecting the reliability of the model. In applications such as human key point detection, the detection and memory of sparse features are also a challenge. At the same time, multi-modal sensor data fusion, such as the alignment of inertial sensor and visual sensor data, is crucial to improving motion capture accuracy, but there are technical difficulties. SUMMARY
[0004] The purpose of the present application is to provide a motion analysis system based on inertial-visual signal enhancement and fusion, which can effectively solve the problem of inertial sensor signal saturation in high dynamic environments, improve the accuracy and reliability of signal reconstruction, and optimize multi-modal sensor data fusion to improve the accuracy of motion capture.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions:
[0006] In a first aspect, the present application provides a motion analysis system based on inertial-visual signal enhancement and fusion, comprising: an inertial sensor module, a visual sensor module, a signal processing module, an inertial-visual information fusion module, and a motion analysis module.
[0007] The inertial sensor module is configured to collect inertial signals of the moving object, and the inertial signals include acceleration, angular velocity and magnetic field intensity of a designated part of the moving object.
[0008] The visual sensor module is a rare memory enhancement module based on Sharma-Mittal entropy, and is configured to collect human key point information of the moving object.
[0009] The signal processing module is configured to perform enhancement processing on the inertial signals by using an improved deep learning model, and reconstruct the super-range signals in the inertial signals based on a Kolmogorov-Arnold network (KAN).
[0010] The inertial-visual information fusion module is configured to perform fusion processing on the human key point information and the processed inertial signals based on an optimal transport theory-based multi-modal feature alignment method and a SInkhorn algorithm, to obtain the motion trajectory of the moving object.
[0011] The motion analysis module is configured to analyze the motion mode and posture change of the moving object based on the motion trajectory of the moving object.
[0012] According to the embodiments provided in the present application, the following technical effects are disclosed.
[0013] The present application provides a motion analysis system based on inertial-visual signal enhancement and fusion. The system is based on an inertial sensor module and a visual sensor module, and can more accurately collect inertial signals and human key point information of a moving object, thereby improving the accuracy of motion capture. The signal processing module is configured to perform enhancement processing on the inertial signals by using an improved deep learning model, and reconstruct the super-range signals in the inertial signals based on a KAN. The inertial-visual information fusion module is configured to fuse the collected human key point information and the processed inertial signals, so that the system can more accurately restore the motion trajectory of the moving object. The motion trajectory of the moving object can be accurately described based on the system. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0015] Figure 1This is a schematic diagram of the structure of a motion analysis system based on inertial-visual signal enhancement and fusion in one embodiment of this application.
[0016] Figure 2 This is a schematic diagram of the overall framework of a motion capture and intelligent analysis system based on inertial-visual signal enhancement and fusion, provided as an embodiment of this application.
[0017] Figure 3 This is a schematic diagram of overrange signal reconstruction based on KAN, provided as an embodiment of this application.
[0018] Figure 4 This is a schematic diagram of the calculation process for Generated Illusion Entropy (GHE) provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Example 1
[0022] like Figure 1 As shown, this embodiment provides a motion analysis system based on inertial-visual signal enhancement and fusion, including: an inertial sensor module, a visual sensor module, a signal processing module, an inertial-visual information fusion module, and a motion analysis module.
[0023] The inertial sensor module is used to collect inertial signals of a moving object; the inertial signals include the acceleration, angular velocity, and magnetic field strength of a designated part of the moving object.
[0024] The visual sensor module is a rare memory enhancement module based on Sharma-MIttal entropy, used to collect key point information of moving objects.
[0025] The signal processing module is used to: enhance the inertial signal using an improved deep learning model, and reconstruct the overrange signal in the inertial signal based on the Kolmogorov-Arnold network; the improved deep learning model is a deep learning model that introduces the generation of illusion entropy into the loss function of the deep learning model.
[0026] The inertial-visual information fusion module is configured to fuse human key point information and signal-processed inertial signals based on an optimal transport theory-based multi-modal feature alignment method and a SInkhorn algorithm to obtain a motion trajectory of the moving object.
[0027] The motion analysis module is configured to analyze a motion mode and a posture change of the moving object based on the motion trajectory of the moving object.
[0028] In some embodiments, the inertial sensor module is configured to acquire inertial signals of the moving object. The inertial signals include various physical quantities of a designated part of the moving object, specifically, acceleration, angular velocity and magnetic field intensity. Through acquisition of the data, dynamic characteristics of the moving object during motion can be comprehensively understood and analyzed.
[0029] In some embodiments, the visual sensor module is a Sharma-Mittal entropy-based rare memory enhancement module configured to acquire human key point information of the moving object.
[0030] The visual sensor module includes a feature extraction submodule, an attention enhancement submodule, a Sharma-Mittal entropy-based rare memory enhancement submodule, a two-dimensional key point detection submodule and a three-dimensional visual reconstruction submodule.
[0031] Specifically, a normal mobile phone shot picture has about several hundred thousand to several million pixels, however, in such a picture, human key points are several sparse and rare pixels. In order to better detect human key points in the picture, the embodiment proposes a Sharma-Mittal entropy-based rare memory enhancement (RME) module, and constructs a deep learning-based visual sensor module (human key point detector), SM-Detector. The SM-Detector mainly includes the following four modules.
[0032] The feature extraction submodule (Feature EXtraction) is configured to extract multi-channel features in the input image.
[0033] The attention enhancement submodule (AttentIon AugmentatIon) is configured to enhance attention to sparse features through an attention mechanism.
[0034] The Sharma-Mittal entropy-based rare memory enhancement submodule (RME) is configured to regularize the feature layer to enhance the memory of the model to the sparse features.
[0035] The two-dimensional key point detection submodule (KeypoIntDetectIon) outputs the positions of the two-dimensional key points of the human body.
[0036] The three-dimensional vision reconstruction submodule (three-dimensional VIsIonReconstructIon) outputs the positions of the three-dimensional key points of the human body.
[0037] The specific calculation process of the SM-Detector is as follows:
[0038] 1) Feature extraction submodule: First, the embodiment uses a ResNet50 network to extract multi-channel features from the input image: F = ResNet50(I).
[0039] In the formula, I is the input image, and the dimension is H0 and W0 are the height and width of the input image, respectively, and F is the extracted feature, and the dimension is H and W are the height and width of the feature map, respectively, and c is the number of feature channels.
[0040] 2) Attention enhancement submodule:
[0041] In order to enhance the attention to sparse key points, the embodiment adopts an attention mechanism. This module calculates the weight of each position and uses these weights to weight the features. The attention weight matrix A is calculated:
[0042]
[0043] In the formula, Q = W q F, Q is the query matrix, W q is the first learnable linear transformation matrix, and the dimension is K = W k F, K is the key-value matrix, W k is the second learnable linear transformation matrix, and the dimension is d k is the dimension of the key-value matrix, A is the attention weight matrix, which is used to represent the importance of different positions on the feature map, and the dimension is V is the value matrix, W v is the third learnable linear transformation matrix, and the dimension is F' is the weighted feature, and the dimension is
[0044] After obtaining the attention weight, the embodiment uses the attention weight to weight the feature to highlight the features of important positions:
[0045] F' = A·V, V = W v F.
[0046] where V is the value matrix, W v is the learnable linear transformation matrix with dimension F' is the weighted feature with dimension
[0047] 2) Rare memory enhancement module based on Sharma-Mittal entropy:
[0048] To enhance the model's memory of sparse features, this embodiment introduces Sharma-Mittal entropy. This module acts on the weighted feature layer F' and realizes regularization by calculating the probability distribution of each position. For each channel c of the feature map F' c , this embodiment directly calculates the probability distribution p i,j,c of each position in two-dimensional space:
[0049]
[0050] where F' i,j,c is the feature value of the feature map at position (i, j). p i,j,c is the probability of the feature value at position (i, j) of the cth channel. is the sum of all feature values on the feature map, used to normalize the feature value and ensure it is a valid probability distribution.
[0051] For each channel c, based on the calculated probability distribution p i,j,c , this embodiment calculates the Sharma-Mittal entropy S r,q (F' c ):
[0052]
[0053] where r and q are two parameters of Sharma-Mittal entropy. By selecting r > 1 and q < 1, the model's memory of sparse features can be enhanced. S r,q (F' c ) is the Sharma-Mittal entropy on the cth channel, which acts on the two-dimensional feature map to enhance the model's attention to certain sparse features in the feature map.
[0054] Finally, the Sharma-Mittal entropy regularization term is the sum of the entropy values of all channels:
[0055]
[0056] where λ is the regularization coefficient, controlling the weight of Sharma-Mittal entropy in the total loss function, S r,q (F' c) is the Sharma-Mittal entropy, c denotes the c-th channel, and c denotes the number of channels. This regularization term helps to enhance the model's response to sparse features, especially in focusing on low-probability events. A smaller Sharma-Mittal entropy indicates a more concentrated feature distribution, allowing the model to focus on certain key points and thus enhancing its ability to remember sparse features.
[0057] 4) Two-dimensional keypoint detection sub-module:
[0058] The features enhanced by attention and rare memory are input into the keypoint detection module to predict the positions of human key points:
[0059]
[0060] where W y and b y are learnable parameters of linear transformation, used to map features to the coordinate space of key points. is the predicted position of the two-dimensional human key point, the final result output by the two-dimensional keypoint detection module, and F vision are the features enhanced by attention and rare memory.
[0061] 5) Three-dimensional visual reconstruction sub-module:
[0062] In applications such as motion analysis and virtual reality, the three-dimensional coordinates of human key points are crucial. Therefore, in the overall network architecture, this embodiment designs a three-dimensional human key point reconstruction algorithm, which aims to reconstruct three-dimensional human key points using the aforementioned features enhanced by "attention enhancement" and "rare memory enhancement" and the detected two-dimensional key points.
[0063] Considering that the features are for two-dimensional keypoint detection, this embodiment needs to be properly processed to serve three-dimensional keypoint reconstruction. To this end, this embodiment uses a deformable attention mechanism to act on multi-channel features. For each joint j and each channel feature F c , a deformable attention operation is performed on the feature map to obtain the feature
[0064]
[0065] where, is the two-dimensional coordinate of joint j. The specific deformable attention DeformAttn formula is as follows:
[0066]
[0067] where M is the number of attention heads, K is the number of sampling points for each attention head, and Δp c,m,kis the offset of the k-th sampling point of the m-th attention head relative to the origin, A c,m,k is the corresponding attention weight.
[0068] Subsequently, the embodiment fuses the feature processed by the deformable convolution with the two-dimensional coordinate embedding of the joint to form a feature vector with rich spatial context information. Specifically, the embodiment stacks the feature of each joint with the joint coordinate embedding:
[0069]
[0070] wherein, is the fused feature vector of joint j.
[0071] Then, the stacked fused vector is input into the self-attention layer for feature fusion:
[0072]
[0073] Finally, the fused feature vector is obtained:
[0074]
[0075] After obtaining the feature representation of each joint, the embodiment also needs to model the spatial dependency between different joints, so that the model can better understand the overall structure of the human skeleton. Therefore, the embodiment globally models the fused joint features through a TransfoRMEr encoder. Specifically, the embodiment inputs the feature of each joint into the TransfoRMEr encoder for spatial modeling:
[0076]
[0077] The output obtained encodes the spatial dependency information between all joints. Therefore, the embodiment maps the spatially modeled features to a three-dimensional coordinate space to generate the final three-dimensional joint coordinates:
[0078]
[0079] wherein, is the predicted three-dimensional key point coordinate.
[0080] In some embodiments, the signal processing module specifically includes the following content:
[0081] The signal processing module is configured to perform enhancement processing on the inertial signal by using an improved deep learning model. The improved deep learning model is a deep learning model in which a generated hallucination entropy is introduced into a loss function of the deep learning model, and is configured to reconstruct an over-range signal in the inertial signal based on a Kolmogorov-Arnold network.
[0082] wherein the input inertial signal is extracted by using a 1D-ResNet to obtain a feature vector; each feature in the feature vector is converted by using a learnable first spline parameterization function to generate a first layer intermediate node; the first layer intermediate node is processed by using a second spline parameterization function to generate a second layer intermediate node; the process of generating the intermediate node is repeatedly executed, and the reconstructed signal is generated through multi-layer spline function transformation.
[0083] Specifically, since the traditional signal processing method often regards the signal as a series of discrete point sequence, this discretization processing method has obvious limitations in capturing the continuity and dynamic change characteristics of the signal, especially when dealing with high dynamic nonlinear signals, the effect is usually not ideal.
[0084] Therefore, the embodiment proposes to introduce the Kolmogorov-Arnold network (KAN) into the inertial signal processing. The KAN has a built-in learnable spline function, which is suitable for curve fitting. Its core idea is to use a multi-layer network structure embedded with spline functions to express complex multivariate continuous functions. Compared with the traditional multilayer perceptron (MLP), KAN can more effectively capture the nonlinear characteristics and subtle fluctuations of the signal. This network structure is suitable for processing inertial signals that are prone to saturation in high dynamic environments. The schematic diagram of the KAN-based over-range signal reconstruction is shown in Figure 3 .
[0085] Specifically, the reconstruction process is as follows:
[0086] First, the input inertial signal is extracted by using a 1D-ResNet to map the signal to a feature vector F = [F1, F2, …, F m ], which provides high-quality input for the subsequent KAN network. For each feature F j , a learnable spline parameterization function is used to convert it to generate a first layer intermediate node The mathematical expression is:
[0087]
[0088] wherein g ij is a learnable first spline parameterization function, which has the ability to capture nonlinear relationships in the signal, is the first layer intermediate node; m is the number of feature vectors, F j is the jth feature vector extracted by using a 1D-ResNet from the input inertial signal.
[0089] Through learning, these functions can dynamically adjust to more accurately fit the complex curves in the signal. Next, these intermediate nodes are further processed by another set of spline parameterization functions, generating the next layer of intermediate nodes The mathematical expression is:
[0090]
[0091] where f ki is a learnable second spline parameterization function, is the second layer of intermediate nodes; I is the Ith feature vector of the second layer, and p is the number of second layer feature vectors.
[0092] In the multi-layer propagation process, the output of each layer will be used as the input of the next layer, and the final output layer generates the reconstructed signal. The motivation behind this layer-by-layer progressive structure design is to gradually approximate the complex nonlinear characteristics of the original signal through layer-by-layer spline function transformation, thereby achieving high-precision reconstruction of the signal.
[0093] Therefore, the KAN network overcomes the limitations of traditional discrete point sequence models by embedding continuous spline functions, and can better capture and reconstruct nonlinear dynamic changes in the signal. Especially in the processing of inertial signals in high dynamic environments, KAN network can effectively reduce the information loss caused by discretization, and achieve more accurate signal reconstruction. In addition, by introducing learnable spline parameterization functions into the network, KAN can adaptively adjust the parameters of each layer to adapt to signal features of different complexities, which makes it have strong generalization ability in diversified application scenarios.
[0094] Specifically, in the application of deep learning models, especially in the field of signal processing, the so-called "hallucination" problem is often encountered, that is, the model generates chaotic or false output that does not conform to the actual situation when reconstructing the signal. This phenomenon is particularly unacceptable in high-risk fields such as aerospace, medical devices, and autonomous driving, because even a small error can have serious consequences.
[0095] To this end, as Figure 4 shown, the present embodiment proposes a new metric way - generated hallucination entropy (GHE) to quantitatively evaluate and reduce the hallucination phenomenon of the model, ensuring the stability and reliability of signal reconstruction. The calculation process of GHE can be divided into the following steps:
[0096] (1) Multi-resolution down-sampling:
[0097] N times down-sampling is performed on the input saturated signal to obtain a series of down-sampled signals of different lengths This step simulates the performance of signals at different scales by generating multiple resolution signal versions, enabling the GHE to detect the consistency of model reconstructions at different resolutions.
[0098] (2) Reconstruction of the KAN model:
[0099] These multi-resolution signals are input into the KAN model for reconstruction, resulting in reconstructed signals
[0100] (3) Pearson correlation coefficient calculation:
[0101] All reconstructed signals are uniformly downsampled to simplify computational complexity, and then the Pearson correlation coefficients P between these signals are calculated. ij The Pearson correlation coefficient is used to quantify the similarity between signals at different resolutions, thereby evaluating the consistency of model output.
[0102] (4) Constructing GHE:
[0103] The final GHE is calculated by the following formula:
[0104]
[0105] where σ is the SIgmoId function, used to normalize the Pearson correlation coefficient, and the weight ensures that signals at similar scales are more similar. GHE By weighting the consistency of output signals for similar inputs, the hallucination phenomenon generated by the model is reduced. ij P is the Pearson correlation coefficient, len (j) is the data of the jth feature vector, the data of the initial feature vector, len (i) is the data of the Ith feature vector.
[0106] Then, the GHE is introduced into the loss function, and by minimizing the GHE, the hallucination phenomenon of the deep learning model in signal reconstruction is reduced, ensuring that the signals generated by the model are more stable and reliable. Compared to traditional loss functions, GHE not only considers the difference between reconstructed signals and original signals, but also pays special attention to the consistent performance of the model at different resolutions, which makes it more applicable and effective in signal processing tasks. The advantage of GHE is that it can effectively punish model outputs that show inconsistency on multi-resolution signals, thereby forcing the model to continuously optimize its reconstruction ability during training, ultimately achieving high-precision reconstruction of overloaded signals.
[0107] The I-KAN model proposed in this embodiment successfully solves the signal overload problem of inertial sensors in high dynamic environments by introducing the Kolmogorov-Arnold network and generating hallucination entropy. The KAN network performs well in fitting complex curves, and the introduction of GHE greatly improves the output stability and reliability of the model. Experimental verification shows that the I-KAN model not only can effectively reconstruct the overload signal, but also can significantly expand the working range of low-cost inertial sensors.
[0108] In some embodiments, the inertial-visual information fusion module is configured to fuse the human key point information and the processed inertial signals to obtain the motion trajectory of the moving object, and specifically includes:
[0109] 1) Inertial sensor feature extraction:
[0110] The inertial sensor captures acceleration and angular velocity data, which are then repaired and enhanced by the I-KAN architecture designed in this embodiment, thereby avoiding signal overload caused by range limitation. The final 3D acceleration data obtained is A t = [a x (t), a y (t), a z (t)], and the angular velocity data is Ω t = [ω x (t), ω y (t), ω z (t)], where t represents the time step. This embodiment uses a 1D convolutional neural network (conv1D) to extract inertial signal features:
[0111] F inertial = Conv1D(A t , Ω t ).
[0112] where F inertial is the extracted inertial sensor feature, T is the number of time steps, and d is the feature dimension.
[0113] 2) Visual feature extraction:
[0114] In the "2 AI rare memory enhancement for computer vision" section, this embodiment has obtained the features F' c enhanced by attention and rare memory enhancement. Here, this embodiment aligns and fuses it with the inertial features.
[0115] 3) Multi-modal feature alignment based on optimal transport theory (OT-FA):
[0116] The embodiment first maps the features of the inertial sensor and the optical sensor to a unified feature space through self-attention mechanism:
[0117]
[0118] wherein, and have the same feature dimension. Considering that they record the same set of motions, the two modal features should have potential consistency. However, due to their modal differences, this consistency is difficult to be mined and constrained. Therefore, the embodiment proposes OptImal Transport based Feature AlIgn (OT-FA) based on optimal transport theory. The optimal transport theory finds the optimal mapping between two different distributions so as to minimize the transmission cost, which is very suitable for mining the potential consistency of multi-modal features.
[0119] Suppose the distribution of the inertial sensor features is μ inertial and the distribution of the optical sensor features is μ vision . The goal of the optimal transport problem is to find the optimal mapping γ ∈ Γ(μ inertial , μ vision ) that can minimize the transmission cost between the two distributions. The optimal transport problem can be formalized as the following optimization problem:
[0120]
[0121] where Γ(μ inertial , μ vision ) represents the set of all coupled distributions, f i and f i represent the features of the inertial and optical sensors respectively, and c(f j , f i ) is the cost function. The embodiment uses the Euclidean distance between the features:
[0122] c(f j , f i ) = ||f j -f 2 ||.
[0123] To solve this optimization problem, the embodiment uses the SInkhorn algorithm for approximate calculation. This algorithm can efficiently calculate the approximate optimal transport mapping P * . Further, the embodiment can map the inertial signal features and the visual signal features to an aligned feature space:
[0124]
[0125] Subsequently, the embodiment can constrain the consistency of the multi-modal signals in the aligned feature space:
[0126]
[0127] In summary, the OT-FA method can effectively align the features of different modalities by utilizing the optimal transport theory, and maximize the supervised information obtained from unpaired data, thereby ensuring more stable training of the model and more accurate feature alignment.
[0128] 4) Time correlation constraint based on KANIadakIs entropy regularization:
[0129] The multi-modal feature interaction based on the optimal transport theory can utilize the potential correlation between modalities, thereby providing more available information for each modality of motion capture. Based on these high-quality features, the embodiment can achieve 3D motion trajectory reconstruction with the help of VQ-VAE (Vector Quantized Variational Autoencoder) and TransfoRMEr modules. Assuming that the motion tracking and trajectory reconstruction results based on inertial sensors are The visual human key point trajectory reconstruction result is Considering that the utilization of time correlation information in the reconstruction process is not sufficient, the embodiment designs a time correlation constraint based on KANIadakIs entropy regularization, which utilizes the continuity and local smoothness of motion to constrain the trajectory generation result, thereby forcing the model to optimize the feature extraction and reasoning prediction process.
[0130] KANIadakIs entropy is a form of asymmetric entropy that can control the nonlinear effects and heterogeneity of a system, and is particularly suitable for handling non-equilibrium state problems in complex systems. The calculation formula of KANIadakIs entropy is as follows:
[0131]
[0132] where κ is the deformation parameter of KANIadakIs entropy. When κ = 0, KANIadakIs entropy degenerates into the classical Shannon entropy. i is the probability distribution of the trajectory point. The embodiment uses kernel density estimation (KDE, Kernel Density Estimation) to estimate the probability density function p or of the trajectory point i :
[0133]
[0134] where `n` is a kernel function (e.g., a Gaussian kernel) used to measure the similarity between trajectory points. `n` represents the total number of trajectory points.
[0135] By Adding it as a regularization term to the total loss function can penalize overly complex or non-smooth trajectory outputs, thereby forcing the model to relearn the training parameters.
[0136] In this embodiment, the motion analysis module is used to analyze the motion pattern and posture changes of the moving object based on its motion trajectory, as follows:
[0137] like Figure 2 As shown, the motion analysis module can be applied to golf motion analysis.
[0138] In this application scenario, this embodiment applies the aforementioned multi-sensor fusion and trajectory reconstruction method to golf motion analysis, specifically calculating the X-factor in the golf swing. The X-factor is an important parameter in golf kinematics, reflecting the difference in rotational angle between the shoulder and pelvis during the swing, and is typically used to assess an athlete's trunk rotation ability and efficiency at impact. The specific process is as follows:
[0139] First, data collection is performed:
[0140] (1) Objective: To accurately capture the body motion data of golfers during the swing, including the three-dimensional position information of key joints and limbs.
[0141] (2) Sensor settings:
[0142] Inertial sensors (IMUs):
[0143] Location: Installed on key parts of the golfer and club, including: Torso: sternum, lumbar spine; Limbs: upper arm, forearm, thigh, lower leg; Head: back of the head; Club: clubhead.
[0144] Sampling frequency: Set above 100Hz to capture details of fast motion.
[0145] Measurement data: acceleration, angular velocity, magnetic field strength.
[0146] Optical sensor (camera system): either monocular or multi-view camera is acceptable.
[0147] (3) Data collection process:
[0148] Preparation phase: Install and calibrate all sensors. Have the golfer warm up to get used to the presence of the sensors.
[0149] Data recording: Let the golfer perform multiple standardized swings, covering different force and speed. During each swing, all sensors record data simultaneously.
[0150] Data storage: Safely store the collected raw data, preparing for subsequent processing and analysis.
[0151] Then, inertial-visual signal enhancement, fusion and motion capture:
[0152] Using the previously designed multi-modal data fusion model, the data of inertial sensors and optical sensors are fused to obtain high-precision three-dimensional trajectories.
[0153] Then, the calculation of X-factor is performed, as follows:
[0154] X-factor refers to the difference in rotation angle between the shoulder and the hip during the golf swing. A larger X-factor is usually associated with greater swing power and longer ball distance. The calculation steps are as follows:
[0155] (1) Key point extraction:
[0156] Shoulder key points: Left shoulder point: S L = [x SL , y SL , z SL ]. Right shoulder point: S L = [x SR , y SR , z SR ]. Hip key points: Left hip point: H L = [x HL , y HL , z HL ]. Right hip point: H R = [x HR , y HR , z HR ].
[0157] (2) Define shoulder and hip vectors:
[0158] Shoulder vector: V shoulder = S R -S L ; Hip vector: V hip = H R -H L .
[0159] (3) Project to the horizontal plane:
[0160] In order to calculate the horizontal rotation angle, this embodiment projects the above vectors to the horizontal plane (usually the XY plane): shoulder vector projection: where vsx = x SR - x SL , v sy = y SR - y SL .
[0161] Hip vector projection: where v hx = x HR - x HL , v hy = y HR - y HL .
[0162] (4) Calculate the angle of the vector:
[0163] Shoulder angle: θ shoulder = arctan2(v sy , v sx ); Hip angle: θ hip = arctan2(v hy , v hx ).
[0164] (5) Calculate the X-factor:
[0165] X-factor angle: X factor = θ shoulder - θ hip ; Positive values indicate the shoulder rotates backward relative to the hip, and negative values indicate forward rotation.
[0166] (6) Time series analysis:
[0167] For each time frame t, calculate the corresponding X factor , and obtain the X-factor change curve throughout the swing. At the same time, calculate the X-factor at the key moment: Top of Backswing: usually the moment when the X-factor reaches the maximum value. Impact: observe the regression degree of X-factor, and evaluate the energy release efficiency.
[0168] Finally, conduct motion analysis, as follows:
[0169] Through the foregoing, the embodiment can obtain accurate motion capture results, and calculate indicators such as X-factor based on the motion capture results. Based on these indicators, the embodiment can carry out more detailed analysis on the motion process. For example, draw the curve of X-factor changing with time, intuitively display the rotation difference in the swing process; compare the X-factor curves of different swing attempts or different players; train the model to predict the optimal X-factor range, improve the hitting effect; identify abnormal patterns, provide real-time feedback and correction suggestions. Actual applications include but are not limited to:
[0170] (1) Technical evaluation and improvement:
[0171] Power output: By analyzing the maximum value and rate of change of X-factor, the power output ability of the golfer is evaluated. Motion coordination: Observe the synchronization of shoulder and hip rotation to find potential motion coordination problems. Swing efficiency: By the change of X-factor at different stages, the efficiency of energy transfer and release is evaluated.
[0172] (2) Injury prevention:
[0173] Over-rotation risk: Excessive X-factor may increase the risk of lower back injury, and timely adjustment can prevent injury. Asymmetry detection: Find the difference between left and right rotation, guide the golfer to conduct targeted training and balance.
[0174] (3) Personalized training program development:
[0175] Targeted training: According to the analysis results of X-factor, develop training plans to improve core strength and flexibility. Progress tracking: By continuously monitoring the changes of X-factor, evaluate the training effect and progress level.
[0176] By applying multi-sensor fusion and trajectory reconstruction technology to golf motion analysis, this embodiment can capture and analyze the details of the golfer's motion with high precision, especially the key X-factor indicators. This provides scientific basis and effective means for golf technical training, performance evaluation and injury prevention, and promotes the intelligentization and data-driven of golf motion analysis.
[0177] In summary, the present application has the following technical effects:
[0178] 1) I-KAN-based over-range inertial signal reconstruction: This application introduces Kolmogorov-Arnold network (KAN) for inertial signal processing, especially for the signal saturation problem of inertial sensors in high dynamic environment. KAN network has built-in learnable spline function, which can approximate complex nonlinear functions layer by layer through multi-layer structure, so as to realize high-precision reconstruction of signals. Compared with traditional methods, this network can regard signals as continuous curves rather than discrete point sequences, so it can effectively capture the continuous characteristics of signals and reduce the information loss caused by discretization processing, which is suitable for over-range reconstruction of sensor signals and significantly expands the working range of low-cost inertial sensors. In fact, this application first applies KAN architecture to sensor signal processing.
[0179] 2) Generate hallucination entropy (GHE) for reducing the phenomenon of deep learning model generating hallucinations: In view of the "hallucination" phenomenon that may occur in the signal reconstruction process of the deep learning model (i.e. generating chaotic output that does not conform to the actual situation), the present application proposes to generate hallucination entropy (GHE) as a quantitative evaluation index. GHE quantifies the reconstruction stability of the model at different resolutions through multi-resolution downsampling, reconstruction signal consistency evaluation and Pearson correlation coefficient calculation. By introducing GHE into the loss function, the hallucination phenomenon of the model can be effectively reduced by minimizing GHE, ensuring the reliability of signal reconstruction, which is of great significance especially in high-risk fields such as aerospace, medical equipment and autonomous driving.
[0180] 3) Sharma-Mittal entropy for rare memory enhancement of deep learning model: In human key point detection, to deal with the problem of sparse feature detection in images, the present application proposes a rare memory enhancement (RME) module based on Sharma-Mittal entropy. By introducing the regularization processing of Sharma-Mittal entropy on the feature layer, this module can significantly enhance the memory ability of the model for low probability and sparse features. Combined with the multi-channel feature extraction of ResNet50 network and the feature weighting of attention mechanism, the SM-Detector of the present application has shown high accuracy in 2D key point detection and 3D reconstruction, especially in the extraction and memory of sparse human features.
[0181] 4) Multi-modal feature alignment based on optimal transport theory: To solve the problem of multi-modal feature alignment between inertial sensors and visual sensors, the present application adopts the optimal transport theory and performs efficient calculation through SInkhorn algorithm to realize accurate alignment of feature space. The OT-FA method can maximize the supervised information of unpaired data by minimizing the transmission cost between modalities while preserving the potential consistency of inertial and visual signals. This alignment method improves the stability and consistency of inertial and visual data fusion, ensuring more accurate and reliable signal fusion in motion capture process.
[0182] 5) Time correlation constraint based on KANIadakIs entropy regularization: Considering the problem of insufficient utilization of time correlation information in the process of motion trajectory reconstruction, the present application designs a time correlation constraint based on KANIadakIs entropy regularization. KANIadakIs entropy is particularly suitable for handling non-equilibrium problems by controlling the nonlinear effects and heterogeneity of the system. Combined with the kernel density estimation method to estimate the probability density function of the trajectory points, it is included as a regularization term in the loss function, forcing the model to optimize the feature extraction and reasoning process, thereby generating smooth and time-correlated 3D motion trajectories.
[0183] 6) The present application proposes an innovative multi-modal signal processing architecture that deeply integrates inertial signal processing with visual signal processing, forming a holistic motion capture and intelligent analysis system. By integrating the Kolmogorov-Arnold Network (KAN) for signal reconstruction, generating the Ghost Entropy (GHE) for quality control, and modules such as Rare Memory Enhancement (RME) and Multi-modal Feature Alignment (OT-FA) in the same architecture, the present application achieves comprehensive and multi-level processing of motion data. This modular architecture not only improves the robustness of the system, but also significantly enhances the collaborative working ability between different modal data, enabling the entire system to achieve accurate and stable motion capture and analysis in complex environments.
[0184] 7) The system architecture of the present application adopts a modular design, and each functional module can be independently run or combined with other modules. This design greatly improves the scalability of the system, allowing it to be flexibly adjusted according to specific application requirements. For example, different sensor data types can be flexibly introduced or replaced with corresponding feature extraction and alignment modules. This scalability not only extends the life cycle of the system, but also increases its applicability in different application scenarios, ensuring the system's sustained competitiveness in future technological development.
[0185] 8) From the application perspective, the present application applies the integration of inertial signal processing and visual signal processing to the precise and intelligent analysis of golf movements. The system architecture is particularly optimized for the swing motion in golf, combining multi-modal data processing and motion trajectory reconstruction technology to achieve comprehensive analysis of the golfer's swing trajectory and posture. This system not only captures the three-dimensional motion trajectory of the club through inertial sensors, but also analyzes the golfer's posture changes in combination with visual signals to generate key biomechanical parameters such as X-Factor and Tightening Factor. Through the fusion and analysis of multi-modal data, the system can provide precise technical guidance for golf beginners and help coaches develop personalized training programs, greatly improving the training efficiency and scientific nature of golf.
[0186] The database involved in each embodiment provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on blockchain, etc., without being limited thereto. The processor involved in each embodiment provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0187] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, it should be understood that the application encompasses all possible combinations of the technical features described above.
[0188] The principles and implementation manners of the present application are described herein by using specific examples, and the above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, according to the idea of the present application, the specific implementation manners and application scopes will be changed by those skilled in the art. In conclusion, the content of the present specification should not be understood as a limitation of the present application.
Claims
1. A motion analysis system based on inertial-visual signal augmentation and fusion, characterized by, The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method.
2. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 1, characterized in that, The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method.
3. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 2, characterized in that, The application relates to a motion analysis system and a motion analysis method. wherein g ij is a first spline parameterization function that is learnable, is a first layer intermediate node; m is the number of first layer feature vectors, F j is the jth feature vector after feature extraction of the input inertial signal by the 1D-ResNet.
4. The motion analysis system based on fusion and enhancement of inertial-visual signals according to claim 3, characterized in that, The application relates to a motion analysis system and a motion analysis method. wherein f ki is a second spline parameterization function that is learnable, is a second layer intermediate node; I is the Ith feature vector of the second layer, and p is the number of feature vectors of the second layer.
5. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 1, characterized in that, The application relates to a motion analysis system and a motion analysis method. where σ is a SIgmoId function to normalize the Pearson correlation coefficient, H GHE By weighting the penalty on the consistency of the output signals under similar inputs, P ij is the Pearson correlation coefficient, len (j) is the data of the jth feature vector, len (0) is the data of the initial feature vector, len (i) is the data of the Ith feature vector.
6. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 1, wherein, The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to a motion analysis system and a motion analysis method. The application relates to The three-dimensional visual reconstruction submodule is configured to output the position of the three-dimensional key point of the human body in the input image.
7. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 6, characterized in that, The feature extraction submodule is configured to extract multi-channel features from the input image using a ResNet50 network, and the extraction formula is F = ResNet50 (I). In the formula, I is an input image, and the dimension is H0 and W0 are the height and width of the input image respectively, F is the extracted feature, and the dimension is H and W are the height and width of the feature map, and c is the number of feature channels.
8. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 6, characterized in that, The attention enhancement submodule is configured to enhance the attention to the sparse features in the multi-channel features through an attention mechanism. In terms of attention, the attention enhancement submodule is configured to: According to the formula The importance of each position on the feature map is calculated to determine the position of the sparse feature; according to the formula F'=A·V, V=W v F, the sparse feature is weighted; where Q = W q F, W q is a learnable first linear transformation matrix with dimension K = W k F, W k is a learnable second linear transformation matrix with dimension d k is the dimension of the key-value matrix, A is an attention weight matrix used to represent the importance of different positions on the feature map, with dimension V is a value matrix, W v is a learnable third linear transformation matrix with dimension F' is the weighted feature with dimension 9. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 6, wherein, In terms of enhancing the memory of the visual detection model to the sparse features by regularizing the feature layer, the formula expression of the Sharma-Mittal entropy regularization term of the Sharma-Mittal entropy-based rare memory enhancement submodule is where λ is a regularization coefficient, S r,q (F c ) is the Sharma-Mittal entropy, c denotes the c-th channel, and c denotes the number of channels.
10. The motion analysis system based on fusion of inertial-visual signals enhancement according to claim 6, characterized in that, In the aspect of outputting the positions of the human body two-dimensional key points in the input image, the two-dimensional key point detection submodule is configured to determine the positions of the human body two-dimensional key points in the input image according to the formula predict the positions of the human body two-dimensional key points; where W y and b y are learnable parameters of a linear transformation that maps the features to the coordinate space of the keypoints, is the predicted position of the two-dimensional keypoints of the human body, F vision is the feature after attention enhancement and rare memory enhancement.
Citation Information
Patent Citations
Detection device, detection system, motion analysis system, recording medium, and analysis method
CN107106900A
Pedestrian inertial navigation system and method assisted by machine learning algorithms and models
CN108168548A