A method, system, device and storage medium for adaptive evaluation of posture motion

By combining the information of movement and breathing heartbeat, using neural network models to extract features and perform interactive fusion, the problem of insufficient guidance in the existing technology is solved, and the efficiency and quality of exercise teaching is improved.

CN114973411BActive Publication Date: 2025-05-16HUAZHONG NORMAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210604517.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-05-16
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve targeted guidance in physical education, fitness and dance learning and training guidance, teachers have a large workload and students have low enthusiasm for learning.

Method used

By combining the movement and respiratory heartbeat of the object to be detected, the video sequence and respiratory heartbeat echo signals are collected using visible light cameras and millimeter wave radars, and the respiratory heartbeat and 3D human posture characteristics are extracted using neural network models, and interactively fusion is performed to output action scores and respiratory state prediction results.

Benefits of technology

It improves the efficiency and quality of exercise teaching and training, can more comprehensively judge the degree of movement standards, and provides training guidance through breathing status prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973411B_ABST
    Figure CN114973411B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system, device and storage medium for adaptive evaluation of posture-type motion. The method comprises: collecting a video sequence and a respiratory heartbeat echo signal of an object to be detected, and preprocessing the respiratory heartbeat echo signal; inputting the preprocessed respiratory heartbeat echo signal into a trained first network model to obtain respiratory heartbeat features; inputting the video sequence into a trained second network model to obtain 3D human posture features; interactively fusing the respiratory heartbeat features with the 3D human posture features to obtain fused interactive features, and outputting action scores and respiratory state prediction results according to the fused interactive features; predicting 3D human posture according to the 3D human posture features, and calculating the similarity between the predicted 3D human posture and the standard action. The present invention is helpful to improve the efficiency and quality of sports teaching and training by combining the action and respiratory heartbeat of the object to be detected for evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of posture-based motion adaptive evaluation, and more specifically, to a posture-based motion adaptive evaluation method, system, device and storage medium. Background Art

[0002] At present, artificial intelligence technology has been widely used in many fields. In sports, fitness and dance learning and training guidance, artificial intelligence plays a very important role due to its convenient technology. On campus or other fitness venues, teachers or coaches are usually unable to provide one-on-one targeted guidance to each student, and for teachers, demonstrating and correcting students' mistakes over and over again will not only greatly increase the workload of teachers, but also reduce students' enthusiasm for learning due to this boring method. Summary of the invention

[0003] In response to at least one defect or improvement need in the prior art, the present invention provides a method, system, device and storage medium for adaptive evaluation of posture-based motion, which helps to improve the efficiency and quality of sports teaching and training by combining the movements, breathing and heartbeat of the object to be detected for evaluation.

[0004] To achieve the above object, according to a first aspect of the present invention, a method for adaptively evaluating posture-based motion is provided, comprising:

[0005] Using a visible light camera and a millimeter wave radar to collect a video sequence and a respiratory and heartbeat echo signal of the object to be detected, and preprocessing the respiratory and heartbeat echo signal;

[0006] Inputting the preprocessed respiratory and heartbeat echo signals into the trained first network model to obtain respiratory and heartbeat features;

[0007] Inputting the video sequence into the trained second network model to obtain 3D human body posture features;

[0008] Interactively fusing the respiratory and heartbeat features with the 3D human body posture features to obtain fused interactive features, and outputting action scores and respiratory state prediction results according to the fused interactive features;

[0009] The 3D human body posture is predicted according to the 3D human body posture feature, and the similarity between the predicted 3D human body posture and the standard action is calculated.

[0010] Furthermore, the preprocessing includes:

[0011] The respiratory and heartbeat echo signals are mixed with the transmission signal of the millimeter-wave radar and then low-pass filtered to obtain an intermediate frequency signal;

[0012] Performing a fast Fourier transform on the intermediate frequency signal to obtain a signal frequency domain energy spectrum, and obtaining target distance information from the signal frequency domain energy spectrum;

[0013] Solve the phase information based on the distance information to obtain a respiratory and heartbeat waveform diagram;

[0014] The respiratory and heartbeat waveform diagram is expanded into a time-frequency spectrum diagram through Fourier transformation.

[0015] Furthermore, the first network model includes:

[0016] A ResNet50 backbone convolutional neural network is used to extract the image features of each frame from the preprocessed respiratory and heartbeat echo signals;

[0017] Long short-term memory network, used to establish feature fusion of context information in the time domain and obtain enhanced features;

[0018] Normalization layer, used to normalize the enhanced features;

[0019] The multi-layer perceptron is used to perform feature conversion on the normalized features to obtain the breathing and heartbeat features.

[0020] Furthermore, the second network model includes:

[0021] A multi-hypothesis posture generation module, used for generating a plurality of initialized posture hypotheses according to the video sequence;

[0022] A temporal information embedding module, which is used to embed the temporal position encoding into the feature representation of the hypothesized posture;

[0023] Single hypothesis feature enhancement module, used to enhance the features within a single pose hypothesis;

[0024] Multi-hypothesis feature fusion module, used to achieve feature information fusion between multiple enhanced posture hypotheses;

[0025] The 3D posture regression module is used to apply linear transformation operation regression to obtain the 3D human body posture feature.

[0026] Furthermore, the generating of the multiple initialization posture hypotheses includes:

[0027] Extract the 2D posture sequence X∈R of the human body in each frame from the video sequence N×J×2 , where R N×J×2 represents an N×J×2 vector, where N represents the total number of input frames, J represents the total number of human joints, and (x, y) represents the coordinates of the joints. The coordinates (x, y) of the 2D pose sequence are concatenated into

[0028] Using learnable position embeddings The position information of each joint point is retained, and the embedding result is used as the input of the Transformer encoder to extract its features, and then residual connections are performed to obtain multiple initialized posture hypotheses.

[0029] Furthermore, the features of enhancing the internal structure of a single posture hypothesis include:

[0030] First, each pose hypothesis is layer-normalized and then self-attention is calculated; a new feature block is obtained after residual connection; then a multi-layer perceptron is used to mix the different channel information of a single pose hypothesis to further enhance the features.

[0031] Furthermore, the interactive fusion includes:

[0032] The 3D human body posture feature and the breathing and heartbeat feature are first converted into a d-dimensional vector through a fully connected layer. The vector after the 3D human body posture feature conversion is represented as P = (p 1 ,…,p d ), the vector after the 3D human posture feature conversion is expressed as Q = (q 1 ,…,q d );

[0033] Construct circulant matrices A and B using projection vectors:

[0034]

[0035] Multiply the circulant matrix and vectors P and Q to obtain F and G, F = PA, G = QB;

[0036] Through a d×k-dimensional projection matrix W, F and G are converted into k-dimensional fused interaction features M.

[0037] According to a second aspect of the present invention, there is also provided a posture-based motion adaptive evaluation system, comprising:

[0038] A signal acquisition and preprocessing module, used to acquire the video sequence and respiratory and heartbeat echo signals of the object to be detected, and preprocess the respiratory and heartbeat echo signals;

[0039] A respiratory and heartbeat feature extraction module, inputting the pre-processed respiratory and heartbeat echo signals into a first network model to obtain respiratory and heartbeat features;

[0040] A 3D human posture feature extraction module, used for inputting the video sequence into a trained second network model to obtain 3D human posture features;

[0041] A binary feature loop interaction module is used to interactively fuse the respiratory heartbeat feature with the 3D human posture feature to obtain a fused interactive feature, and output an action score and a respiratory state prediction result according to the fused interactive feature;

[0042] The evaluation module is used to predict the 3D human posture according to the 3D human posture features, and calculate the similarity between the predicted 3D human posture and the standard action.

[0043] According to the third aspect of the present invention, there is also provided an electronic device, comprising at least one processor and at least one storage module, wherein the storage module stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of the above methods.

[0044] According to a fourth aspect of the present invention, there is also provided a storage medium storing a computer program executable by a processor, wherein when the computer program runs on the processor, the processor executes the steps of any one of the above methods.

[0045] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0046] (1) In the present invention, 3D human posture information and breathing and heartbeat information are integrated to predict the results, and the standard degree of the action is judged more comprehensively. Compared with the breathing and heartbeat data under standard conditions, when the breathing and heartbeat are too slow, it may indicate that the learner is not proficient in the movement; when the breathing and heartbeat are too intense, it may be necessary to give an abnormal reminder to slow down the learner's movement training. Therefore, the judgment method of binary feature cyclic interaction adopted by the present invention will be more helpful in guiding learners to conduct reasonable training.

[0047] (2) The present invention uses a neural network model to extract the internal features of respiratory and heartbeat information, and uses the time-frequency spectrum as the input of the model, which further simplifies the fixed and cumbersome steps of manual feature extraction, and realizes the automatic extraction of its multiple essential features end-to-end, thereby effectively improving the accuracy and efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1A schematic diagram of a flow chart of a method for adaptively evaluating posture-based motion provided by an embodiment of the present invention;

[0050] Figure 2 A schematic diagram of the principle of the posture-based motion adaptive evaluation method provided by an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of the first network model, the second network model and the dual-attention loop interaction network structure provided in an embodiment of the present invention;

[0052] Figure 4 A schematic diagram of the principle of 3D human posture estimation provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0054] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally includes steps or modules that are not listed, or optionally includes other steps or modules inherent to these processes, methods, products or devices.

[0055] like Figure 1 and Figure 2 As shown, a posture motion adaptive evaluation method according to an embodiment of the present invention comprises the following steps:

[0056] S101, using a visible light camera and a millimeter wave radar to collect a video sequence and a respiratory and heartbeat echo signal of the object to be detected in real time, and preprocessing the respiratory and heartbeat echo signal.

[0057] The specific sub-steps include:

[0058] (1) The human target is located within the detection area of ​​the visible light and millimeter wave radar sensors; in this embodiment, the human target is a learner who is undergoing sports training;

[0059] (2) Use visible light cameras to obtain RGB video sequences and millimeter wave radars to receive echo signals;

[0060] (3) Preprocess the echo signal to enhance relevant information and eliminate useless information.

[0061] Furthermore, the echo preprocessing includes the following sub-steps:

[0062] (1) Obtain the intermediate frequency signal from the echo signal. The specific processing process is as follows:

[0063] The millimeter wave radar transmits a linear frequency modulated continuous wave to the target user's chest area. The delay of the received echo will change with the movement of the chest cavity. The radar received signal is mixed and filtered to obtain an intermediate frequency signal, which contains the movement information of the target chest breathing and heartbeat. The signal transmitted in one frequency modulation cycle of the frequency modulated continuous wave used by the millimeter wave radar is:

[0064]

[0065] Among them, A T is the amplitude of the transmitted signal, f c is the center frequency, W is the bandwidth, T m is the signal frequency modulation period, and τ represents the extension time from the transmitted signal to the received signal. After being reflected by the target and the environment, the echo signal is obtained.

[0066]

[0067] Among them, A R is the amplitude of the echo signal, Δt is the time delay, Δf d Represents Doppler frequency shift. The transmitted signal and the echo signal are mixed and processed and low-pass filtered to obtain the intermediate frequency signal:

[0068] S IF (t) = S T (t)S R (t)≈A T A R exp{j2π[f c Δt+(f I -Δf d )t]},(7)

[0069] in, Represents the frequency of the intermediate frequency signal at time t.

[0070] (2) The target distance is obtained through FFT processing. The specific processing process is as follows:

[0071] The method of collecting vital signs signals through radar is mainly to measure the phase at the corresponding distance of the target. Since the chest movement of the linear frequency modulation continuous wave radar is coupled with the distance of the target, the signal needs to be preprocessed before the phase information is obtained to obtain the distance of the target. In order to obtain the target distance range, it is necessary to use fast Fourier transform (FFT) to obtain the frequency domain energy spectrum of the signal. The distance range of the target corresponds to the distance bin with the largest energy. However, due to the limitation of frequency resolution, this is only a distance range and cannot be accurately located, so the two adjacent distance bins with the largest energy are selected for principal component analysis, and then the first principal component is extracted as the actual distance of the target.

[0072] (3) Solve the target phase information and obtain the expression of respiratory and heartbeat signs. The specific processing process is as follows:

[0073] Millimeter wave radar can obtain I / Q dual-path signals. Since there is a mismatch between the two branches, the present invention uses the Differential and Cross-Multiply (DACM) method to restore the phase information of the signal. The process of the DACM algorithm is as follows:

[0074]

[0075] in, represents the recovered signal phase information, t represents time, Q(t) represents the Q channel signal at time t, I(t) represents the I channel signal at time t, Q ′ (t) represents the derivative of Q(t) with respect to t, and I′(t) represents the derivative of I(t) with respect to t. According to the definition of derivative, the above formula can be further transformed into:

[0076]

[0077] Where Δt represents a time interval that tends to 0.

[0078] Since differentiation will increase the interference of high-frequency noise, the obtained More sensitive to noise interference. In order to suppress noise, let Δt = 1 in the above formula, and integrate the signal to obtain the final phase information

[0079]

[0080] In this way, the actual phase information of the object is obtained. The relationship between the phase of the intermediate frequency signal and the distance to the object is as follows:

[0081]

[0082] in, represents the phase change of the intermediate frequency signal, and Δd represents the displacement change caused by the heart or chest cavity.

[0083] Therefore, it can be found that the human body's respiratory and heartbeat waveform information is related to the phase of the distance unit where the human target is located. By extracting the phase information of the distance unit where the target is located within a period of time, the respiratory and heartbeat waveform information of the human target can be obtained.

[0084] (4) The waveform diagram is converted into a time-frequency spectrum diagram through Fourier expansion. The specific processing process is as follows:

[0085] Feature extraction is an essential process for inputting signals into the neural network. In order to embody the end-to-end concept as much as possible, we consider converting the respiratory and heartbeat waveforms into time-frequency spectrograms as the input data of the neural network, and use convolutional neural networks and long short-term memory networks to predict whether the physiological information of the target user is abnormal, so as to guide the user to make adjustments.

[0086] From the above analysis, we can see that the phase reflects the movement changes of the chest cavity, and the movement changes of the chest cavity are caused by breathing and heartbeat. Therefore, the breathing and heartbeat waveform with time as the horizontal axis can be converted into a time-frequency spectrum with time as the horizontal axis, frequency as the vertical axis, and color representing amplitude by wavelet transform. This is more conducive to preserving the required features, so that it can be used as the input of the neural network. After a lot of training, the model will be able to identify the standardization or rationality of the target user's physiological data.

[0087] S102, inputting the preprocessed respiratory and heartbeat echo signals into the trained first network model to obtain respiratory and heartbeat features.

[0088] like Figure 3 As shown in FIG. 1 , the first network model includes a ResNet50 backbone convolutional neural network, a long short-term memory network, a normalization layer, and a multi-layer perceptron (MLP). The preprocessed respiratory and heartbeat echo signals, that is, the spectrogram sequence, are used as the input of this network model. The image features of each frame are first extracted through the ResNet50 backbone convolutional neural network, and then the long short-term memory network is used to characterize the time correlation information of the sequence data; then, the enhanced features are layer-normalized and input into the MLP for feature conversion again; finally, the required respiratory and heartbeat features B∈R are obtained. o .

[0089] The ResNet50 network is mainly composed of convolutional layers, pooling layers, fully connected layers, and residual connections. Convolution is a mathematical calculation, and in neural networks it can be understood as an operation used to extract image features. The convolution operation gradually increases the receptive field to extract high-level features of the image; pooling is a spatial operation that reduces the height and length directions. Its function is to reduce the dimension of the feature map while maintaining the most important information, thereby compressing the number of data and parameters, reducing overfitting, and improving the fault tolerance of the model; the fully connected layer enables the output of each result to be determined by all inputs; the residual connection is used to solve the problem of feature information loss as the network deepens. The feature extraction can be obtained after the frequency spectrum of each frame passes through the network.

[0090] The long short-term memory network is a special recursive neural network that can effectively use the time association between the data at time t, time t-1, and time t+1 to obtain rich features of the data at time t. Compared with general recursive neural networks, the long short-term memory network can solve the problem of useful information from a long time ago being ignored and the gradient disappearing. Through the long short-term memory network, the feature fusion of the time domain context information can be established between the time-frequency spectrum graphs to obtain a richer feature representation Y i (i∈[1,…,S]).

[0091] It is further preferred that cross-attention calculation is performed on the multiple features obtained above to obtain different degrees of attention to each feature, so as to achieve the goal of paying more attention to the main information and less attention to the secondary information. Specifically, the feature matrix Y i (i∈[1,…,S]) first obtains Q, K, V through linear mapping, and then performs the following calculations:

[0092] Sim(Q,K)=Q·K T ,(12)

[0093] A=Softmax*Sim(Q,K)),(13)

[0094] Attention(Q)=A·V.(14)

[0095] Formula (12) calculates the similarity Sim(Q, K) between features by dot product; Formula (13) performs Softmax normalization on the similarity score; Formula (14) uses A as the weight coefficient for weighted summation to obtain the final attention score result matrix.

[0096] In this way, the degree of attention to each feature is obtained, and then weighted fusion is performed, layer normalization is performed, and then MLP processing is performed to obtain the final respiratory heartbeat features. MLP is used to enhance the features. It contains two linear layers and one activation layer:

[0097] MLP(x)=σ(xw 1 +b 1 )W 2 +b 2 ,(15)

[0098] Among them, σ represents the GELU activation function, and b 2 ∈R d Represent the weights and biases of the two linear layers respectively.

[0099] S103, input the video sequence into the trained second network model to obtain 3D human body posture features.

[0100] like Figure 3 and Figure 4 As shown in the figure, the second network model uses a multi-posture hypothesis interaction network model (MPHInteraction), including: a multi-hypothesis posture generation module, which is used to generate multiple initialized posture hypotheses according to the video sequence; a time information embedding module, which is used to embed the time position code into the feature representation of the hypothesized posture; a single hypothesis feature enhancement module, which is used to enhance the features within a single posture hypothesis; a multi-hypothesis feature fusion module, which is used to realize the feature information fusion between multiple enhanced posture hypotheses; a 3D posture regression module, which applies linear transformation operations to regress and obtain 3D human posture features.

[0101] The details are as follows:

[0102] (1) Multi-hypothesis posture generation, the specific processing process is as follows:

[0103] Let (x, y) represent the coordinates of a joint point on a certain frame, N represents the total number of input frames, and J represents the total number of joint points of the person. First, use OpenPose to extract a string of 2D poses X∈R N×J×2 , which is spliced ​​into Then use a learnable position embedding The position information of each joint point is retained; then it is input into the Transformer encoder for processing. The whole process can be expressed as:

[0104]

[0105] Among them, L 1 (l∈[1,...,L 1 ]) represents the number of Transformer encoders used in the multi-hypothesis pose generation module. In addition, X m Represents the generated mth hypothetical posture, these M hypothetical postures are generated by L 1The final representation after the encoder of the layer is Indicates X m The result after position encoding embedding, represents the feature representation of the m-th posture at the (l-1) layer, and Indicates the Transformer encoder generated during calculation LN represents the layer normalization operation. MSA represents the multi-head self-attention operation, which converts the input x∈R n×d The linear mapping is QueryQ∈R n×d ,KeyK∈R n×d And ValueV∈R n×d , where n represents the sequence length and d represents the dimension, the attention is calculated as follows:

[0106]

[0107] MSA will perform the above operations in parallel with h heads, and finally concatenate the output results of these h attention heads.

[0108] (2) Time information embedding: The specific processing process is as follows:

[0109] The previous position encoding of joint points belongs to the spatial domain, so the features obtained are not rich enough. In order to utilize the temporal information, we consider converting the spatial domain features to the temporal domain. The features extracted by the posture hypothesis obtained for each frame Use one conversion operation and one linear embedding to obtain high-dimensional features Where C represents the embedding dimension. Then, a learnable temporal position encoding To preserve the time information between frames. This process can be expressed as:

[0110]

[0111] In the formula, for Embedded temporal position encoded representation.

[0112] (3) Single hypothesis feature enhancement: the specific processing process is as follows:

[0113] First, self-attention is calculated for each pose hypothesis, and then the different channel information of a single pose hypothesis is mixed through a multi-layer perceptron to further enhance the features. Specifically, the embedded features of different pose hypotheses are As input to multiple MSA blocks in parallel:

[0114]

[0115] Among them, l∈[1,…,L 2 ] represents the index of the SHR layer, It represents the feature representation of the m-th hypothesis posture at the (l-1) layer. In this way, the features of the posture hypothesis are enhanced, which is helpful for judging the final result.

[0116] (4) Multi-hypothesis feature fusion. The specific processing process is as follows:

[0117] Furthermore, the features of multiple hypotheses are concatenated as the input of the MLP:

[0118]

[0119] Among them, Concat(·) represents the concatenation operation. Represents the result of concatenating multiple hypothetical features.

[0120] These aggregated features are evenly divided into non-overlapping blocks along the channel dimension, which results in a mixture of the relationships between channels of different hypotheses.

[0121] In order to make the hypothesized postures interact, cross-attention calculation is required. Specifically, let the m-th posture hypothesis be the feature of the l-th layer Perform the following calculations:

[0122]

[0123] Among them, l∈[1,…,L 3 ] represents the index of the CHI layer, m 1 ,m 2 are the other two posture assumptions, represents the feature representation of the m-th pose hypothesis at the (l-1) layer. MCA(Q,K,V) represents multi-head cross-attention, which is the same as the Attention calculation process of the multi-layer perceptron in the first network model.

[0124] After the attention calculation, the information between different channels needs to be mixed through a multi-layer perceptron:

[0125]

[0126] Finally, the segmentation operation is no longer performed, so that the aggregated features can finally be synthesized into a single pose hypothesis representation: Z M ∈RN×(C·M) .

[0127] (5) 3D posture regression, the specific processing process is as follows:

[0128] Z M Apply a linear transformation layer to regress the 3D pose sequence Finally, from Select the 3D pose of the center frame as the final prediction result.

[0129] The network is trained in an end-to-end manner, and the loss function used is the mean (per) joint position error (MPJPE). The goal of model training is to minimize the error between the predicted value and the true value:

[0130]

[0131] in, and They represent the predicted 3D coordinates and true values ​​of the i-th joint point in the n-th frame respectively.

[0132] S104, interactively fuse the respiratory and heartbeat features with the 3D human posture features to obtain the fused interactive features, predict the 3D human posture according to the 3D human posture features, calculate the similarity between the predicted 3D human posture and the standard action, and output the action score and respiratory state prediction results according to the fused interactive features.

[0133] The binary feature loop interaction model is used to perform interactive fusion of features. The final model regression obtains the action standard degree score and breathing state prediction, which are described in detail as follows:

[0134] Given a 3D human body posture feature Z M ∈R N×(C·M) With the respiratory and heartbeat features B∈R o , first through the fully connected layer to convert it into vector data of the same dimension, represented as P(p 1 ,…,p d )∈R d ,Q(q 1 ,…,q d )∈R d ; Then use the projection vector P∈R d ,Q∈R d Construct a circulant matrix A∈R d×d ,B∈R d×d :

[0135]

[0136]

[0137] In order to make full use of the elements in the projection vector and the circulant matrix, multiply the circulant matrix and the projection vector:

[0138] F=PA,G=QB. (31)

[0139] Finally, through a projection matrix W∈R d×k Let F∈R d ,G∈R d Transformed into target vector M∈R k After the model is trained with a large amount of data, it can predict the score scoreX of the action standard degree based on the target vector. The MSE loss function is used during training:

[0140]

[0141] Among them, scoreY i It represents the score of manual labeling of the ith sample, and m is the total number of samples. is a regularization term, and λ is called a regularization parameter, which controls the trade-off between two different objectives to avoid overfitting. The training goal is to make the loss function as small as possible.

[0142] In a specific example, the classification network can use a softmax classifier to predict the respiratory and heartbeat states, and classify them into three classification results: rapid, stable, and bradycardic. The cross entropy is used as the loss function for training to obtain the best prediction result.

[0143] S105, predicting a 3D human body posture according to the 3D human body posture features, calculating the similarity between the predicted 3D human body posture and the standard action, and outputting an action score according to the fused interaction features.

[0144] According to the regression result value of the learner's 3D human posture estimation, the similarity score between each part and the standard action is calculated, and then the average similarity score is obtained.

[0145] In a specific example, the dynamic time warping algorithm (3DPose-DTW) based on 3D posture coordinate feature differences can be used to match action postures. This algorithm can overcome the matching problem of different sequence lengths and solve the problem of advanced or delayed actions. In general, training is performed based on background music or background guidance prompts, so first of all, the dynamic programming idea is used to nonlinearly warp the time series information to be identified and the time series information of the standard template, and then the best corresponding point between the two sequences is found based on the sound, and then the Euclidean distance between the learner's posture feature vector and the standard action feature vector in the video frame of the corresponding time is calculated. Regression obtains the learner's 3D posture feature vector Assume that the 3D posture feature vector obtained by regressing the video frame of the standard action is After a fully connected layer, it is expanded into a one-dimensional vector Then the Euclidean distance is calculated as follows:

[0146]

[0147] The Euclidean distance is used to represent the similarity between two posture features. The smaller the distance, the more similar they are, and the larger the distance, the smaller the similarity. In an ideal situation, Indicates that the two posture features are exactly the same, thus judging that the two actions are exactly the same. In order to use the percentage system as the scoring standard, the following conversion is performed to obtain the final score

[0148]

[0149] The learners are evaluated and guided based on the average similarity scores of the movement parts, the movement standard degree scores, and the breathing state prediction results.

[0150] In a specific example, as a preferred solution, the weighted sum of the average similarity score of the action parts and the action standard degree score is used to obtain the final score, which is used as a quantitative evaluation of the learner. In order to qualitatively evaluate the practice of each part of the learner, according to the action similarity score of the video frame obtained, when the score is less than 60 points, the video frame image is given, and the part with larger deviation action is displayed. At the same time, according to the feature analysis of the respiratory and heartbeat data and the prediction of the respiratory state, guidance on the learner's breathing adjustment is given.

[0151] An adaptive evaluation system for posture-based motion according to an embodiment of the present invention includes:

[0152] The signal acquisition and preprocessing module is used to acquire the video sequence and respiratory and heartbeat echo signals of the object to be detected, and preprocess the respiratory and heartbeat echo signals;

[0153] A respiratory and heartbeat feature extraction module inputs the preprocessed respiratory and heartbeat echo signals into the first network model to obtain respiratory and heartbeat features;

[0154] A 3D human posture feature extraction module is used to input the video sequence into the trained second network model to obtain 3D human posture features;

[0155] The dual-element feature loop interaction module is used to interactively fuse the respiratory and heartbeat features with the 3D human posture features to obtain fused interactive features, and output the action score and respiratory state prediction results based on the fused interactive features;

[0156] The evaluation module is used to predict the 3D human posture according to the 3D human posture features and calculate the similarity between the predicted 3D human posture and the standard action.

[0157] The implementation principle of the system is the same as the above method and will not be repeated here.

[0158] This embodiment also provides an electronic device, which includes at least one processor and at least one memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor executes any step of the above-mentioned posture motion adaptive evaluation method. The specific steps are shown in the method embodiment and will not be repeated here. In this embodiment, the types of processor and memory are not specifically limited. For example, the processor can be a microprocessor, a digital information processor, an on-chip programmable logic system, etc. The memory can be a volatile memory, a non-volatile memory, or a combination thereof.

[0159] The present application also provides a storage medium, which stores a computer program executable by a processor, and when the computer program is run on the processor, the processor executes any one of the steps of the above-mentioned posture-based motion adaptive evaluation method. Wherein, the computer-readable storage medium may include but is not limited to any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a micro drive and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nano system (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0160] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0161] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0162] In the several embodiments provided in the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are only schematic, such as the division of the modules, which is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, indirect coupling or communication connection of systems or modules, which can be electrical or other forms.

[0163] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or software functional modules.

[0165] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program code.

[0166] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by entering a program to instruct the relevant hardware. The program may be stored in a computer-readable memory, and the memory may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0167] The above is only an exemplary embodiment of the present disclosure, and the scope of the present disclosure cannot be limited thereto. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure here, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field not recorded in the present disclosure. The description and examples are regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

[0168] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0169] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for adaptive evaluation of posture-based motion, characterized in that: include: Using a visible light camera and a millimeter wave radar to collect a video sequence and a respiratory and heartbeat echo signal of the object to be detected, and preprocessing the respiratory and heartbeat echo signal; Inputting the preprocessed respiratory and heartbeat echo signals into the trained first network model to obtain respiratory and heartbeat features; Inputting the video sequence into the trained second network model to obtain 3D human body posture features; Interactively fusing the respiratory and heartbeat features with the 3D human body posture features to obtain fused interactive features, and outputting action scores and respiratory state prediction results according to the fused interactive features; Predicting a 3D human body posture according to the 3D human body posture feature, and calculating the similarity between the predicted 3D human body posture and a standard action; Wherein, the second network model includes: The multi-hypothesis posture generation module is used to generate multiple initial posture hypotheses based on the video sequence; specifically, let (x, y) represent the coordinates of the joint points on a certain frame, N represents the total number of input frames, and J represents the total number of joint points of the person; extract the 2D posture sequence X∈R of the human body in each frame image from the video sequence N×J×2 , which is spliced ​​into Using a learnable position embedding The position information of each joint point is retained and input into the Transformer encoder for processing; The temporal information embedding module is used to embed the temporal position code into the feature representation of the assumed posture; specifically, a conversion operation and a linear embedding are performed on the features extracted from the posture hypothesis obtained in each frame to obtain high-dimensional features. C represents the embedding dimension; using a learnable temporal position encoding To preserve the time information between frames; The single hypothesis feature enhancement module is used to enhance the features within a single posture hypothesis; specifically, the embedded features of different posture hypotheses are As input to multiple MSAs in parallel; The multi-hypothesis feature fusion module is used to realize the feature information fusion between multiple enhanced posture hypotheses. Specifically, the embedded features of multiple posture hypotheses are concatenated as the input of MLP, cross-attention calculation is performed, and the information between different channels is mixed through the multi-layer perceptron to synthesize a single posture hypothesis Z. M ∈R N×(C·M) ; The 3D posture regression module is used to apply linear transformation operation regression to obtain the 3D human posture feature; specifically, for a single posture hypothesis Z M ∈R N×(C·M) Apply a linear transformation layer to regress the 3D pose sequence 2. The method for adaptive evaluation of posture-based motion according to claim 1, characterized in that: The pre-processing comprises: The respiratory and heartbeat echo signals are mixed with the transmission signal of the millimeter-wave radar and then low-pass filtered to obtain an intermediate frequency signal; Performing a fast Fourier transform on the intermediate frequency signal to obtain a signal frequency domain energy spectrum, and obtaining target distance information from the signal frequency domain energy spectrum; Solve the phase information based on the distance information to obtain a respiratory and heartbeat waveform diagram; The respiratory and heartbeat waveform diagram is expanded into a time-frequency spectrum diagram through Fourier transformation.

3. The method for adaptive evaluation of posture-based motion according to claim 1, characterized in that: The first network model includes: A ResNet50 backbone convolutional neural network is used to extract the image features of each frame from the preprocessed respiratory and heartbeat echo signals; Long short-term memory network, used to establish feature fusion of context information in the time domain and obtain enhanced features; Normalization layer, used to normalize the enhanced features; The multi-layer perceptron is used to perform feature conversion on the normalized features to obtain the breathing and heartbeat features.

4. The method for adaptive evaluation of posture-based motion according to claim 1, characterized in that: The multiple posture hypotheses for generating initialization include: Extract the 2D posture sequence X∈R of the human body in each frame from the video sequence N×J×2 , where R N×J×2 represents an N×J×2 vector, where N represents the total number of input frames, J represents the total number of human joints, and (x, y) represents the coordinates of the joints. The coordinates (x, y) of the 2D pose sequence are concatenated into Using learnable position embeddings The position information of each joint point is retained, and the embedding result is used as the input of the Transformer encoder to extract its features, and then residual connections are performed to obtain multiple initialized posture hypotheses.

5. The method for adaptive evaluation of posture-based motion according to claim 1, characterized in that: The features of enhancing the internal structure of a single posture hypothesis include: First, each pose hypothesis is layer-normalized and then self-attention is calculated; a new feature block is obtained after residual connection; then a multi-layer perceptron is used to mix the different channel information of a single pose hypothesis to further enhance the features.

6. The method for adaptive evaluation of posture-based motion according to claim 1, characterized in that: The interactive fusion includes: The 3D human body posture feature and the breathing and heartbeat feature are first converted into a d-dimensional vector through a fully connected layer. The vector after the 3D human body posture feature conversion is represented as P = (p1, ..., p d ), the vector after the 3D human posture feature conversion is expressed as Q = (q1, ..., q d ); Construct circulant matrices A and B using projection vectors: Multiply the circulant matrix and vectors P and Q to obtain F and G, F = PA, G = QB; Through a d×k-dimensional projection matrix W, F and G are converted into k-dimensional fused interaction features M.

7. A posture-based motion adaptive evaluation system, the system being used to execute the steps of the method according to any one of claims 1 to 6, characterized in that: include: A signal acquisition and preprocessing module, used to acquire the video sequence and respiratory and heartbeat echo signals of the object to be detected, and preprocess the respiratory and heartbeat echo signals; A respiratory and heartbeat feature extraction module, inputting the pre-processed respiratory and heartbeat echo signals into a first network model to obtain respiratory and heartbeat features; A 3D human posture feature extraction module, used for inputting the video sequence into a trained second network model to obtain 3D human posture features; A binary feature loop interaction module is used to interactively fuse the respiratory heartbeat feature with the 3D human posture feature to obtain a fused interactive feature, and output an action score and a respiratory state prediction result according to the fused interactive feature; The evaluation module is used to predict the 3D human posture according to the 3D human posture feature, and calculate the similarity between the predicted 3D human posture and the standard action.

8. An electronic device, characterized in that: The method comprises at least one processor and at least one storage module, wherein the storage module stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The device stores a computer program, and when the computer program is run on a processor, the processor is enabled to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Human face pose estimation method fusing manual design descriptor and depth features

    CN109858342A

  • Method and device of detecting human vital signs based on ultra-wideband radar

    CN109965858A