Martial arts action recognition and action quality evaluation method based on artificial intelligence
By fusing visible light and infrared images to recognize martial arts movements and extracting and smoothing the human skeleton from point cloud data, the problem of low accuracy of existing martial arts movement recognition and evaluation methods is solved, and intelligent and quantitative evaluation of martial arts movements is achieved.
Patent Information
- Application Number
- CN202510514903.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The accuracy of existing martial arts movement recognition and evaluation methods is low, and it is greatly affected by subjective factors of the evaluator, and there is a lack of accurate quantitative evaluation methods.
Combining visible light images and infrared images is combined for fusion processing, martial arts actions are identified through the action recognition network, and the human skeleton is extracted from the point cloud data for smoothing processing, and the action quality evaluation is performed using the skeleton characteristics.
The accuracy of martial arts movement recognition and the objectivity of the evaluation of action quality are improved, the dependence on the subjective factors of the evaluator is reduced, and an intelligent and quantitative action evaluation is realized.
Smart Images

Figure CN120375474A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of martial arts action recognition and action quality evaluation, and particularly relates to a method for martial arts action recognition and action quality evaluation based on artificial intelligence. Background Art
[0002] Exercise is the main way for residents to exercise in their daily lives. At the same time, sports technology teaching is also an important part of college physical education. Through teaching, students can effectively master the techniques, skills and knowledge of different projects, comprehensively improve their physical fitness, develop the habit of exercising, and thus form a healthy lifestyle. With the strong promotion of martial arts in recent years, the number of martial arts enthusiasts has been increasing. At the same time, with the strong promotion of using school martial arts education to assist in the inheritance of the national pulse in China, martial arts has gradually entered the college physical education classroom and become a very common college physical education course. Then, how to comprehensively, scientifically and reasonably evaluate the quality of martial arts actions has become the focus.
[0003] However, the traditional methods for martial arts action recognition and evaluation are too affected by the subjective factors of the evaluators and lack an accurate measurement scale. Therefore, there is an urgent need for a more precise quantification and effective objectivity method like a computer to assist in the evaluation of martial arts actions. To achieve intelligent, quantitative and accurate evaluation of the action quality of martial arts routine projects, intelligent recognition and action evaluation of martial arts actions are required.
[0004] In terms of action recognition, with the development of deep learning technology, methods for martial arts action recognition and evaluation based on deep learning have gradually emerged. Among them:
[0005] The convolutional neural network obtains an image containing the human body from the video, and then sends the obtained image into the convolutional neural network. After feature extraction, the action recognition result is obtained. For example, in the patent application with the publication number CN114419505A, a method for martial arts action recognition based on human pose estimation is proposed.
[0006] The recurrent neural network models the obtained human key point data into a sequence, takes the time dimension as the sequence dimension and the key point dimension as the feature dimension, uses the recurrent neural network to extract features from it, and finally performs classification to obtain the recognition result.
[0007] In terms of action evaluation, existing methods can extract the human skeleton based on deep learning technology, and then perform action evaluation according to the human skeleton features. However, generally speaking, the accuracy of existing martial arts action recognition and evaluation methods is still low and needs to be further improved. Summary of the Invention
[0008] The object of the present invention is to solve the problem of low accuracy of existing martial arts movement recognition and evaluation methods, and a martial arts movement recognition and movement quality evaluation method based on artificial intelligence is proposed.
[0009] The technical solution adopted by the present invention to solve the above technical problems is: a martial arts movement recognition and movement quality evaluation method based on artificial intelligence, and the method specifically includes the following steps:
[0010] Step 1, at each acquisition time point, synchronously acquire visible light images, infrared images and point cloud data of the martial arts movement;
[0011] Step 2, for the visible light image and the infrared image synchronously acquired at any time point t, fuse the visible light image and the infrared image to obtain a fused image corresponding to the time point t;
[0012] Similarly, obtain the fused image corresponding to each acquisition time point;
[0013] Step 3, take the fused images from the first acquisition time point to the T-th acquisition time point as the input of the action recognition network, and output the martial arts movement recognition result through the action recognition network;
[0014] Step 4, extract the human skeleton from the point cloud data obtained at each acquisition time point respectively, and smooth the extracted skeleton to obtain the human skeleton corresponding to each acquisition time point;
[0015] Step 5, evaluate the quality of the martial arts movement according to the finally obtained human skeleton in Step 4 and the martial arts movement recognition result in Step 3.
[0016] Further, the process of fusing the visible light image and the infrared image to obtain a fused image corresponding to the time point t is as follows:
[0017] Take the visible light image and the infrared image synchronously acquired at the time point t as the input of the image fusion network, and the working process in the image fusion network is as follows:
[0018] Take the infrared image as the input of the first convolutional layer with a convolutional kernel size of 3×3, and take the output of the first convolutional layer as the input of the second convolutional layer with a convolutional kernel size of 1×1;
[0019] Take the output of the second convolutional layer as the input of the first max-pooling layer, and take the output of the first max-pooling layer as the input of the third convolutional layer with a convolutional kernel size of 3×3;
[0020] Take the output of the third convolutional layer as the input of the fourth convolutional layer with a convolutional kernel size of 1×1, add the output of the first max-pooling layer and the output of the fourth convolutional layer, and take the added result as the input of the second max-pooling layer;
[0021] Use the output of the second max pooling layer as the input to the fifth convolutional layer with a convolutional kernel size of 3×3, and use the output of the fifth convolutional layer as the input to the sixth convolutional layer with a convolutional kernel size of 1×1;
[0022] Add the output of the second max pooling layer to the output of the sixth convolutional layer, and use the added result as the input to the third max pooling layer;
[0023] Use the output of the third max pooling layer as the input to the seventh convolutional layer with a convolutional kernel size of 3×3, and use the output of the seventh convolutional layer as the input to the eighth convolutional layer with a convolutional kernel size of 1×1;
[0024] Use the visible light image as the input to the ninth convolutional layer with a convolutional kernel size of 3×3, and use the output of the ninth convolutional layer as the input to the tenth convolutional layer with a convolutional kernel size of 1×1;
[0025] Use the output of the tenth convolutional layer as the input to the fourth max pooling layer, and use the output of the fourth max pooling layer as the input to the eleventh convolutional layer with a convolutional kernel size of 3×3;
[0026] Use the output of the eleventh convolutional layer as the input to the twelfth convolutional layer with a convolutional kernel size of 1×1, add the output of the fourth max pooling layer to the output of the twelfth convolutional layer, and use the added result as the input to the fifth max pooling layer;
[0027] Use the output of the fifth max pooling layer as the input to the thirteenth convolutional layer with a convolutional kernel size of 3×3, and use the output of the thirteenth convolutional layer as the input to the fourteenth convolutional layer with a convolutional kernel size of 1×1;
[0028] Add the output of the fifth max pooling layer to the output of the fourteenth convolutional layer, and use the added result as the input to the sixth max pooling layer;
[0029] Use the output of the sixth max pooling layer as the input to the fifteenth convolutional layer with a convolutional kernel size of 3×3, and use the output of the fifteenth convolutional layer as the input to the sixteenth convolutional layer with a convolutional kernel size of 1×1;
[0030] Use the output of the eighth convolutional layer and the output of the sixteenth convolutional layer as the input to the attention mechanism fusion module, use the output of the attention mechanism fusion module as the input to the seventeenth convolutional layer with a convolutional kernel size of 1×1, and then use the output of the seventeenth convolutional layer as the input to the first feature extractor;
[0031] Use the infrared image as the input of the second feature extractor, use the visible light image as the input of the third feature extractor, use the outputs of the first feature extractor, the second feature extractor, and the third feature extractor as the input of the CNN network, and use the binary image output by the CNN network as the fused image of the infrared image and the visible light image.
[0032] Furthermore, the working process of the attention mechanism fusion module is as follows:
[0033] Perform weighted summation on the output of the eighth convolutional layer and the output of the sixteenth convolutional layer to obtain a weighted summation result a;
[0034] Connect the output of the eighth convolutional layer with the weighted summation result a to obtain a connection result a1;
[0035] Use the connection result a1 as the input of the eighteenth convolutional layer with a kernel size of 1×1, and use the output of the eighteenth convolutional layer as the input of the first sigmoid activation function layer;
[0036] Use the connection result a1 as the input of the first spatial attention unit, add the output of the first sigmoid activation function layer and the output of the first spatial attention unit to obtain an addition result b, and then multiply the addition result b by the output of the eighth convolutional layer to obtain a multiplication result b1;
[0037] Connect the output of the sixteenth convolutional layer with the weighted summation result a to obtain a connection result a2;
[0038] Use the connection result a2 as the input of the nineteenth convolutional layer with a kernel size of 1×1, and use the output of the nineteenth convolutional layer as the input of the second sigmoid activation function layer;
[0039] Use the connection result a2 as the input of the second spatial attention unit, add the output of the second sigmoid activation function layer and the output of the second spatial attention unit to obtain an addition result c, and then multiply the addition result c by the output of the sixteenth convolutional layer to obtain a multiplication result c1;
[0040] Then add the multiplication result b1 and the multiplication result c1, and use the addition result as the output of the attention mechanism fusion module.
[0041] Furthermore, the working process of the first feature extractor is as follows:
[0042] Inside the first feature extractor, use the input of the first feature extractor as the input of the first branch and the second branch respectively, where:
[0043] The first branch sequentially includes the 20th convolutional layer with a 3×3 convolutional kernel, the first ReLU activation function layer, the 21st convolutional layer with a 3×3 convolutional kernel, the second ReLU activation function layer, the 22nd convolutional layer with a 3×3 convolutional kernel, and the third ReLU activation function layer, and the output of the third ReLU activation function layer is taken as the output of the first branch.
[0044] The second branch sequentially includes the 23rd convolutional layer with a 3×3 convolutional kernel and the 24th convolutional layer with a 3×3 convolutional kernel, and the output of the 24th convolutional layer is taken as the output of the second branch.
[0045] The output of the first branch and the output of the second branch are added, and the added result is taken as the input of the third branch and the input of the fourth branch respectively, where:
[0046] The third branch sequentially includes the 25th convolutional layer with a 3×3 convolutional kernel, the fourth ReLU activation function layer, the 26th convolutional layer with a 3×3 convolutional kernel, the fifth ReLU activation function layer, the 27th convolutional layer with a 3×3 convolutional kernel, and the sixth ReLU activation function layer, and the output of the sixth ReLU activation function layer is taken as the output of the third branch.
[0047] The fourth branch sequentially includes the 28th convolutional layer with a 3×3 convolutional kernel and the 29th convolutional layer with a 3×3 convolutional kernel, and the output of the 29th convolutional layer is taken as the output of the fourth branch.
[0048] The output of the third branch and the output of the fourth branch are added, and the added result is taken as the output of the first feature extractor.
[0049] Further, the action recognition network includes T spatio-temporal feature extraction modules, and the working process of the action recognition network is as follows:
[0050] The fused image at the t-th acquisition time point is taken as the input of the t-th spatio-temporal feature extraction module, where t = 1, 2, …, T; and the output of the first spatio-temporal feature extraction module is taken as the input of the first GRU module.
[0051] The output of the t-th spatio-temporal feature extraction module and the output of the (t - 1)-th GRU module are taken as the input of the t-th GRU module.
[0052] The outputs of the first GRU module to the T-th GRU module are averaged, and then the averaged result passes through a fully connected layer. Finally, the output of the fully connected layer is taken as the input of the Softmax layer, and the action recognition result is output through the Softmax layer.
[0053] Further, the working process of the first spatio-temporal feature extraction module is as follows:
[0054] Within the first spatio-temporal feature extraction module, the input fused image is first used as the input of the thirtieth convolutional layer, and then the output of the thirtieth convolutional layer is used as the input of the first spatio-temporal feature extraction unit. Next, the output of the first spatio-temporal feature extraction unit is used as the input of the second spatio-temporal feature extraction unit, and the output of the second spatio-temporal feature extraction unit is used as the input of the third spatio-temporal feature extraction unit.
[0055] The output of the thirtieth convolutional layer is added to the output of the third spatio-temporal feature extraction unit, and the added result is used as the output of the first spatio-temporal feature extraction module.
[0056] Further, the working process of the first spatio-temporal feature extraction unit is as follows:
[0057] Within the first spatio-temporal feature extraction unit, the input of the first spatio-temporal feature extraction unit is used as the input of the thirty-first convolutional layer, and then the output of the thirty-first convolutional layer is used as the input of the first BN layer. The output of the first BN layer is used as the input of the seventh ReLU activation function layer.
[0058] The output of the seventh ReLU activation function layer is used as the input of the thirty-second convolutional layer, and then the output of the thirty-second convolutional layer is used as the input of the second BN layer. The output of the second BN layer is used as the input of the eighth ReLU activation function layer, and the output of the eighth ReLU activation function layer is used as the input of the first channel attention unit.
[0059] The input of the first spatio-temporal feature extraction unit is used as the input of the thirty-third convolutional layer, and then the output of the thirty-third convolutional layer is used as the input of the third BN layer. The output of the third BN layer is used as the input of the ninth ReLU activation function layer.
[0060] The output of the ninth ReLU activation function layer is used as the input of the thirty-fourth convolutional layer, and then the output of the thirty-fourth convolutional layer is used as the input of the fourth BN layer. The output of the fourth BN layer is used as the input of the tenth ReLU activation function layer.
[0061] The input of the first spatio-temporal feature extraction unit is used as the input of the thirty-fifth convolutional layer, and then the output of the thirty-fifth convolutional layer is used as the input of the fifth BN layer. The output of the fifth BN layer is used as the input of the eleventh ReLU activation function layer.
[0062] The output of the eleventh ReLU activation function layer is used as the input of the thirty-sixth convolutional layer, and then the output of the thirty-sixth convolutional layer is used as the input of the sixth BN layer. The output of the sixth BN layer is used as the input of the twelfth ReLU activation function layer, and the output of the twelfth ReLU activation function layer is used as the input of the third spatial attention unit.
[0063] Multiply the output of the tenth ReLU activation function layer by the output of the first channel attention unit to obtain a multiplication result e;
[0064] Multiply the output of the tenth ReLU activation function layer by the output of the third spatial attention unit to obtain a multiplication result f;
[0065] Add the multiplication result e and the multiplication result f to obtain an addition result g;
[0066] Multiply the output of the third spatial attention unit by the output of the first channel attention unit to obtain a multiplication result h;
[0067] Add g and h to obtain an addition result k, and use the addition result k as the output of the first spatio-temporal feature extraction unit.
[0068] Further, extracting the human skeleton from the point cloud data obtained at each acquisition time point specifically includes:
[0069] Taking the point cloud data obtained at any one acquisition time point as an example
[0070] Step 1: Initialize the iteration number n = 1;
[0071] Step 2: Initialize l = 1;
[0072] Step 3: For the l-th point in the point cloud data, calculate the search radius r centered on the l-th point:
[0073]
[0074] where l k represents the k-th neighborhood point of point l, and dist(l, l k ) represents the distance between point l and point l k , and k = 1, 2,..., K;
[0075] Search for the point set within the region centered on point l with radius r, calculate the center of all points in the point set, and then move point l to the position of the calculated center;
[0076] Step 4: Determine whether l has traversed all points in the point cloud data;
[0077] If not all points in the point cloud data have been traversed, let l = l + 1, and return to execute Step 3;
[0078] If all points in the point cloud data have been traversed, continue to execute Step 5;
[0079] Step 5: Determine whether n = N is satisfied, where N represents the set maximum number of iterations:
[0080] If n < N, then let n = n + 1, and return to execute step two;
[0081] If n = N, then continue to return to execute step six;
[0082] Step six: Perform human skeleton extraction on the point cloud data obtained after the Nth iteration.
[0083] Furthermore, smooth the extracted skeleton to obtain the human skeleton corresponding to each acquisition time point; the specific process is as follows:
[0084]
[0085] Among them, x t′,j represents the coordinate of the jth joint point in the human skeleton corresponding to the t'-th acquisition time point, and x t′,j represents the coordinate of the jth joint point in the human skeleton corresponding to the t'-th acquisition time point after smoothing processing.
[0086] Even further, the specific process of step 5 is as follows:
[0087] Step 51: Calculate the direction vectors of each joint according to the coordinates of each joint point of the human skeleton, and obtain the standard actions from the database according to the martial arts action recognition result in step 3;
[0088] Step 52: Calculate the cosine similarity between the joint direction vectors calculated in step 51 and the direction vectors of the standard actions;
[0089] Step 53: Obtain the final evaluation result of the martial arts action according to the calculation result of step 52.
[0090] The beneficial effects of the present invention are:
[0091] The present invention combines visible light images and infrared images to recognize martial arts actions, which can take more comprehensive information into account during the recognition process. Through visible light images and infrared images at multiple consecutive acquisition time points, the continuous state characteristics of martial arts actions can be perfectly matched. The present invention can ensure the accuracy of martial arts action recognition through the fusion processing of various image information and the joint application of image data at multiple acquisition time points. According to the martial arts action recognition result, the standard actions to be compared with the point cloud data can be determined. Then, after processing the point cloud data using the skeleton extraction method designed by the present invention, the human skeleton can be accurately extracted, and the human skeleton extraction result can be compared with the standard actions to improve the accuracy of evaluating the quality of martial arts actions. Brief Description of the Drawings
[0092] Figure 1 is a flowchart of a method for recognizing martial arts actions and evaluating action quality based on artificial intelligence according to the present invention;
[0093] Figure 2(a) is the first visible light image;
[0094] Figure 2(b) is the infrared image corresponding to Figure 2(a);
[0095] Figure 2(c) is the point cloud data corresponding to Figure 2(a);
[0096] Figure 2(d) is the human skeleton diagram corresponding to Figure 2(c);
[0097] Figure 3(a) is the second visible light image;
[0098] Figure 3(b) is the infrared image corresponding to Figure 3(a);
[0099] Figure 3(c) is the point cloud data corresponding to Figure 3(a);
[0100] Figure 3(d) is the human skeleton diagram corresponding to Figure 3(c);
[0101] Figure 4(a) is the third visible light image;
[0102] Figure 4(b) is the infrared image corresponding to Figure 4(a);
[0103] Figure 4(c) is the point cloud data corresponding to Figure 4(a);
[0104] Figure 4(d) is the human skeleton diagram corresponding to Figure 4(c). Detailed implementation manners
[0105] Detailed implementation manner one: Combine Figure 1 to illustrate this implementation manner. A method for martial arts action recognition and action quality evaluation based on artificial intelligence described in this implementation manner specifically includes the following steps:
[0106] Step 1, at each acquisition time point, synchronously acquire the visible light image, infrared image, and point cloud data of the martial arts action;
[0107] Since the martial arts action is not an instantaneous action but a continuous action in time, therefore, in the present invention, a plurality of continuous acquisition time points are set within the duration of a martial arts action, and at each acquisition time point, a visible light image, an infrared image, and point cloud data are acquired;
[0108] Step 2, for the visible light image and infrared image synchronously acquired at any time point t, use the visible light image and infrared image synchronously acquired at time point t as the input of the image fusion network. The working process within the image fusion network is:
[0109] Take the infrared image as the input of the first convolutional layer with a convolutional kernel size of 3×3, and take the output of the first convolutional layer as the input of the second convolutional layer with a convolutional kernel size of 1×1;
[0110] Take the output of the second convolutional layer as the input of the first max pooling layer, and take the output of the first max pooling layer as the input of the third convolutional layer with a convolutional kernel size of 3×3;
[0111] Take the output of the third convolutional layer as the input of the fourth convolutional layer with a convolutional kernel size of 1×1, add the output of the first max pooling layer and the output of the fourth convolutional layer, and take the added result as the input of the second max pooling layer;
[0112] Take the output of the second max pooling layer as the input of the fifth convolutional layer with a convolutional kernel size of 3×3, and take the output of the fifth convolutional layer as the input of the sixth convolutional layer with a convolutional kernel size of 1×1;
[0113] Add the output of the second max pooling layer and the output of the sixth convolutional layer, and take the added result as the input of the third max pooling layer;
[0114] Take the output of the third max pooling layer as the input of the seventh convolutional layer with a convolutional kernel size of 3×3, and take the output of the seventh convolutional layer as the input of the eighth convolutional layer with a convolutional kernel size of 1×1;
[0115] Take the visible light image as the input of the ninth convolutional layer with a convolutional kernel size of 3×3, and take the output of the ninth convolutional layer as the input of the tenth convolutional layer with a convolutional kernel size of 1×1;
[0116] Take the output of the tenth convolutional layer as the input of the fourth max pooling layer, and take the output of the fourth max pooling layer as the input of the eleventh convolutional layer with a convolutional kernel size of 3×3;
[0117] Take the output of the eleventh convolutional layer as the input of the twelfth convolutional layer with a convolutional kernel size of 1×1, add the output of the fourth max pooling layer and the output of the twelfth convolutional layer, and take the added result as the input of the fifth max pooling layer;
[0118] Take the output of the fifth max pooling layer as the input of the thirteenth convolutional layer with a convolutional kernel size of 3×3, and take the output of the thirteenth convolutional layer as the input of the fourteenth convolutional layer with a convolutional kernel size of 1×1;
[0119] Add the output of the fifth max pooling layer and the output of the fourteenth convolutional layer, and take the added result as the input of the sixth max pooling layer;
[0120] Take the output of the sixth max pooling layer as the input of the fifteenth convolutional layer with a convolutional kernel size of 3×3, and take the output of the fifteenth convolutional layer as the input of the sixteenth convolutional layer with a convolutional kernel size of 1×1;
[0121] Take the output of the eighth convolutional layer and the output of the sixteenth convolutional layer as the inputs of the attention mechanism fusion module. The working process of the attention mechanism fusion module is as follows:
[0122] Perform weighted summation on the output of the eighth convolutional layer and the output of the sixteenth convolutional layer to obtain the weighted summation result a;
[0123] Concatenate the output of the eighth convolutional layer and the weighted summation result a to obtain the concatenation result a1;
[0124] Take the concatenation result a1 as the input of the eighteenth convolutional layer with a kernel size of 1×1, and take the output of the eighteenth convolutional layer as the input of the first sigmoid activation function layer;
[0125] Take the concatenation result a1 as the input of the first spatial attention unit. Add the output of the first sigmoid activation function layer and the output of the first spatial attention unit to obtain the addition result b, and then multiply the addition result b by the output of the eighth convolutional layer to obtain the multiplication result b1;
[0126] Concatenate the output of the sixteenth convolutional layer and the weighted summation result a to obtain the concatenation result a2;
[0127] Take the concatenation result a2 as the input of the nineteenth convolutional layer with a kernel size of 1×1, and take the output of the nineteenth convolutional layer as the input of the second sigmoid activation function layer;
[0128] Take the concatenation result a2 as the input of the second spatial attention unit. Add the output of the second sigmoid activation function layer and the output of the second spatial attention unit to obtain the addition result c, and then multiply the addition result c by the output of the sixteenth convolutional layer to obtain the multiplication result c1;
[0129] Then add the multiplication result b1 and the multiplication result c1, and take the addition result as the output of the attention mechanism fusion module;
[0130] Take the output of the attention mechanism fusion module as the input of the seventeenth convolutional layer with a kernel size of 1×1, and then take the output of the seventeenth convolutional layer as the input of the first feature extractor (the first feature extractor can extract the common features of infrared images and visible light images);
[0131] Take the infrared image as the input of the second feature extractor (the second feature extractor can extract the unique features of the infrared image), take the visible light image as the input of the third feature extractor (the third feature extractor can extract the unique features of the visible light image), take the output of the first feature extractor, the output of the second feature extractor, and the output of the third feature extractor as the input of the CNN network, and take the binary image output by the CNN network as the fused image of the infrared image and the visible light image (in the binary image, the pixel value corresponding to the human body area is 1, and the pixel value corresponding to the background area is 0). When performing subsequent classification tasks based on the fused image, since the accurate position and posture of the human body are already reflected in the fused image, therefore, based on the fused images at consecutive multiple moments, the accuracy of martial arts action recognition can be significantly improved.
[0132] Among them: The working process of the first feature extractor is as follows:
[0133] Inside the first feature extractor, take the input of the first feature extractor as the input of the first branch and the input of the second branch respectively, where:
[0134] Inside the first branch, there are successively the 20th convolutional layer with a convolution kernel size of 3×3, the first Relu activation function layer, the 21st convolutional layer with a convolution kernel size of 3×3, the second Relu activation function layer, the 22nd convolutional layer with a convolution kernel size of 3×3, and the third Relu activation function layer, that is, take the output of the third Relu activation function layer as the output of the first branch;
[0135] Inside the second branch, there are successively the 23rd convolutional layer with a convolution kernel size of 3×3 and the 24th convolutional layer with a convolution kernel size of 3×3, that is, take the output of the 24th convolutional layer as the output of the second branch;
[0136] Add the output of the first branch and the output of the second branch, and take the added result as the input of the third branch and the input of the fourth branch respectively, where:
[0137] Inside the third branch, there are successively the 25th convolutional layer with a convolution kernel size of 3×3, the fourth Relu activation function layer, the 26th convolutional layer with a convolution kernel size of 3×3, the fifth Relu activation function layer, the 27th convolutional layer with a convolution kernel size of 3×3, and the sixth Relu activation function layer, that is, take the output of the sixth Relu activation function layer as the output of the third branch;
[0138] Inside the fourth branch, there are successively the 28th convolutional layer with a convolution kernel size of 3×3 and the 29th convolutional layer with a convolution kernel size of 3×3, that is, take the output of the 29th convolutional layer as the output of the fourth branch;
[0139] Add the output of the third branch and the output of the fourth branch, and use the added result as the output of the first feature extractor respectively.
[0140] In the present invention, the structures of the second feature extractor and the third feature extractor are the same as that of the first feature extractor.
[0141] Step 3: Use the fused images at the 1st acquisition time point to the fused images at the Tth acquisition time point as the input of the action recognition network, and output the martial arts action recognition result through the action recognition network;
[0142] The action recognition network includes T spatio-temporal feature extraction modules, and the working process of the action recognition network is as follows:
[0143] Use the fused image at the tth acquisition time point as the input of the tth spatio-temporal feature extraction module, where t = 1, 2, …, T; and use the output of the 1st spatio-temporal feature extraction module as the input of the 1st GRU module;
[0144] The working process of the 1st spatio-temporal feature extraction module is as follows:
[0145] In the 1st spatio-temporal feature extraction module, the input fused image is first used as the input of the 30th convolutional layer, and then the output of the 30th convolutional layer is used as the input of the first spatio-temporal feature extraction unit. In the first spatio-temporal feature extraction unit, the input of the first spatio-temporal feature extraction unit is used as the input of the 31st convolutional layer, and then the output of the 31st convolutional layer is used as the input of the first BN layer, and the output of the first BN layer is used as the input of the seventh ReLU activation function layer;
[0146] Use the output of the seventh ReLU activation function layer as the input of the 32nd convolutional layer, and then the output of the 32nd convolutional layer is used as the input of the second BN layer, and the output of the second BN layer is used as the input of the eighth ReLU activation function layer, and the output of the eighth ReLU activation function layer is used as the input of the first channel attention unit;
[0147] Use the input of the first spatio-temporal feature extraction unit as the input of the 33rd convolutional layer, and then the output of the 33rd convolutional layer is used as the input of the third BN layer, and the output of the third BN layer is used as the input of the ninth ReLU activation function layer;
[0148] Use the output of the ninth ReLU activation function layer as the input of the 34th convolutional layer, and then the output of the 34th convolutional layer is used as the input of the fourth BN layer, and the output of the fourth BN layer is used as the input of the tenth ReLU activation function layer;
[0149] Take the input of the first spatio-temporal feature extraction unit as the input of the 35th convolutional layer, then take the output of the 35th convolutional layer as the input of the 5th BN layer, and take the output of the 5th BN layer as the input of the 11th ReLU activation function layer;
[0150] Take the output of the 11th ReLU activation function layer as the input of the 36th convolutional layer, then take the output of the 36th convolutional layer as the input of the 6th BN layer, take the output of the 6th BN layer as the input of the 12th ReLU activation function layer, and take the output of the 12th ReLU activation function layer as the input of the third spatial attention unit;
[0151] Multiply the output of the 10th ReLU activation function layer by the output of the first channel attention unit to obtain the multiplication result e;
[0152] Multiply the output of the 10th ReLU activation function layer by the output of the third spatial attention unit to obtain the multiplication result f;
[0153] Add the multiplication result e and the multiplication result f to obtain the addition result g;
[0154] Multiply the output of the third spatial attention unit by the output of the first channel attention unit to obtain the multiplication result h;
[0155] Add g and h to obtain the addition result k, and take the addition result k as the output of the first spatio-temporal feature extraction unit.
[0156] Then take the output of the first spatio-temporal feature extraction unit as the input of the second spatio-temporal feature extraction unit, and take the output of the second spatio-temporal feature extraction unit as the input of the third spatio-temporal feature extraction unit;
[0157] Add the output of the 30th convolutional layer to the output of the third spatio-temporal feature extraction unit, and take the addition result as the output of the 1st spatio-temporal feature extraction module. Through the continuous processing of three spatio-temporal feature extraction units, the present invention can extract spatio-temporal features of different scales and ensure the accuracy of martial arts action recognition;
[0158] Take the output of the t-th spatio-temporal feature extraction module and the output of the (t - 1)-th GRU module as the input of the t-th GRU module;
[0159] Take the average of the outputs of the 1st GRU module to the T-th GRU module, then pass the average result through the fully connected layer, and finally take the output of the fully connected layer as the input of the Softmax layer, and output the action recognition result through the Softmax layer. The present invention can further ensure the accuracy of martial arts action recognition by recognizing martial arts actions based on multiple images in a continuous action process.
[0160] Step 4: Extract the human skeleton from the point cloud data obtained at each acquisition time point, and smooth the extracted skeleton to obtain the human skeleton corresponding to each acquisition time point. Specifically:
[0161] Take the point cloud data obtained at any one acquisition time point as an example
[0162] Step 1: Initialize the iteration number n = 1;
[0163] Step 2: Initialize l = 1;
[0164] Step 3: For the l-th point in the point cloud data, calculate the search radius r centered on the l-th point:
[0165]
[0166] where l k represents the k-th neighborhood point of point l, and dist(l, l k ) represents the distance between point l and point l k , and k = 1, 2,..., K;
[0167] In the present invention, the distances between each of the remaining points and the l-th point are calculated respectively, and then the K points with the smallest distances are selected, and the selected points are then involved in the calculation of the search radius;
[0168] Search the point set within the region centered on point l with a radius of r, calculate the center of all the points in the point set, and then move point l to the position of the calculated center;
[0169] Step 4: Determine whether l has traversed all the points in the point cloud data;
[0170] If not all the points in the point cloud data have been traversed, let l = l + 1, and return to execute Step 3;
[0171] If all the points in the point cloud data have been traversed, continue to execute Step 5;
[0172] Step 5: Determine whether n = N is satisfied, where N represents the set maximum number of iterations:
[0173] If n < N, let n = n + 1, and return to execute Step 2;
[0174] If n = N, continue to return to execute Step 6;
[0175] Step 6: By moving the point cloud data as close as possible to the final skeleton, the accuracy of human skeleton extraction can be improved. The specific process of smoothing the extracted skeleton is as follows:
[0176]
[0177] Among them, \(x\) t′,j represents the coordinate of the \(j\)-th joint point in the human body skeleton corresponding to the \(t'\)-th acquisition time point, and \(x\) t′,j represents the smoothed coordinate of the \(j\)-th joint point in the human body skeleton corresponding to the \(t'\)-th acquisition time point.
[0178] By performing smoothing processing on the joint points of the human body skeleton, the coordinate error of the joint points caused by single-frame point cloud data can be minimized, and the accuracy of martial arts action evaluation can be improved.
[0179] Step 5: Evaluate the quality of the martial arts action according to the finally obtained human body skeleton in Step 4 and the martial arts action recognition result in Step 3:
[0180] Step 51: Calculate the direction vectors of each joint according to the coordinate of each joint point of the human body skeleton, and obtain the standard action from the database according to the martial arts action recognition result in Step 3;
[0181] Step 52: Calculate the cosine similarity between the joint direction vector calculated in Step 51 and the direction vector of the standard action;
[0182] Step 53: Obtain the final martial arts action evaluation result according to the calculation result in Step 52.
[0183] In martial arts actions, in order to ensure the standardization of actions, a fixed angle needs to be formed between joint and joint. Therefore, the present invention takes direction as the main evaluation index. The action at one time point may correspond to multiple joint direction vectors. At this time, the average value of the cosine similarities corresponding to the multiple joint direction vectors can be used as the evaluation index. The whole set of recognition and evaluation methods of the present invention does not require any manual participation and can automatically output the recognition and evaluation results of martial arts actions.
[0184] Experimental part
[0185] As shown in FIGS. 2(a), 2(b) and 2(c), the first group of visible light images, infrared images and point cloud data collected are shown, and the corresponding extracted human body skeleton is shown in FIG. 2(d). As shown in FIGS. 3(a), 3(b) and 3(c), the second group of visible light images, infrared images and point cloud data collected are shown, and the corresponding extracted human body skeleton is shown in FIG. 3(d). As shown in FIGS. 4(a), 4(b) and 4(c), the third group of visible light images, infrared images and point cloud data collected are shown, and the corresponding extracted human body skeleton is shown in FIG. 4(d).
[0186] The above calculation examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or modifications derived from the technical solution of the present invention still fall within the protection scope of the present invention.
Claims
1. A method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence, characterized in that, The method specifically includes the following steps: Step 1: At each acquisition time point, synchronously acquire visible light images, infrared images, and point cloud data of the martial arts movements; Step 2: For the visible light image and the infrared image synchronously acquired at any time point t, fuse the visible light image and the infrared image to obtain a fused image corresponding to time point t; Similarly, obtain the fused image corresponding to each acquisition time point; Step 3: Use the fused images from the first acquisition time point to the T-th acquisition time point as the input of the action recognition network, and output the martial arts action recognition result through the action recognition network; Step 5: Extract the human skeleton from the point cloud data obtained at each acquisition time point, and perform smoothing processing on the extracted skeleton to obtain the human skeleton corresponding to each acquisition time point; Step 6: Evaluate the quality of the martial arts action according to the finally obtained human skeleton in Step 5 and the martial arts action recognition result in Step 3.
2. The method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 1, wherein The process of fusing the visible light image and the infrared image to obtain a fused image corresponding to time point t is as follows: Use the visible light image and the infrared image synchronously acquired at time point t as the input of the image fusion network. The working process within the image fusion network is as follows: Use the infrared image as the input of the first convolutional layer with a convolutional kernel size of 3×3, and use the output of the first convolutional layer as the input of the second convolutional layer with a convolutional kernel size of 1×1; Use the output of the second convolutional layer as the input of the first max-pooling layer, and use the output of the first max-pooling layer as the input of the third convolutional layer with a convolutional kernel size of 3×3; Use the output of the third convolutional layer as the input of the fourth convolutional layer with a convolutional kernel size of 1×1, add the output of the first max-pooling layer and the output of the fourth convolutional layer, and use the added result as the input of the second max-pooling layer; Use the output of the second max-pooling layer as the input of the fifth convolutional layer with a convolutional kernel size of 3×3, and use the output of the fifth convolutional layer as the input of the sixth convolutional layer with a convolutional kernel size of 1×1; Add the output of the second max-pooling layer and the output of the sixth convolutional layer, and use the added result as the input of the third max-pooling layer; Use the output of the third max-pooling layer as the input of the seventh convolutional layer with a convolutional kernel size of 3×3, and use the output of the seventh convolutional layer as the input of the eighth convolutional layer with a convolutional kernel size of 1×1; Use the visible light image as the input of the ninth convolutional layer with a convolutional kernel size of 3×3, and use the output of the ninth convolutional layer as the input of the tenth convolutional layer with a convolutional kernel size of 1×1; Use the output of the tenth convolutional layer as the input of the fourth max-pooling layer, and use the output of the fourth max-pooling layer as the input of the eleventh convolutional layer with a convolutional kernel size of 3×3; Use the output of the eleventh convolutional layer as the input of the twelfth convolutional layer with a convolutional kernel size of 1×1, add the output of the fourth max-pooling layer and the output of the twelfth convolutional layer, and use the added result as the input of the fifth max-pooling layer; Use the output of the fifth max-pooling layer as the input of the thirteenth convolutional layer with a convolutional kernel size of 3×3, and use the output of the thirteenth convolutional layer as the input of the fourteenth convolutional layer with a convolutional kernel size of 1×1; Add the output of the fifth max - pooling layer to the output of the fourteenth convolutional layer, and use the added result as the input to the sixth max - pooling layer; Use the output of the sixth max - pooling layer as the input to the fifteenth convolutional layer with a kernel size of 3×3, and use the output of the fifteenth convolutional layer as the input to the sixteenth convolutional layer with a kernel size of 1×1; Use the output of the eighth convolutional layer and the output of the sixteenth convolutional layer as the input to the attention - mechanism - based fusion module, use the output of the attention - mechanism - based fusion module as the input to the seventeenth convolutional layer with a kernel size of 1×1, and then use the output of the seventeenth convolutional layer as the input to the first feature extractor; Use the infrared image as the input to the second feature extractor, use the visible - light image as the input to the third feature extractor, use the output of the first feature extractor, the output of the second feature extractor, and the output of the third feature extractor as the input to the CNN network, and use the binary image output by the CNN network as the fused image of the infrared image and the visible - light image.
3. The method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 2, wherein The working process of the attention - mechanism - based fusion module is as follows: Perform weighted summation on the output of the eighth convolutional layer and the output of the sixteenth convolutional layer to obtain a weighted summation result a; Connect the output of the eighth convolutional layer and the weighted summation result a to obtain a connection result a1; Use the connection result a1 as the input to the eighteenth convolutional layer with a kernel size of 1×1, and use the output of the eighteenth convolutional layer as the input to the first sigmoid activation function layer; Use the connection result a1 as the input to the first spatial attention unit, add the output of the first sigmoid activation function layer and the output of the first spatial attention unit to obtain an added result b, and then multiply the added result b by the output of the eighth convolutional layer to obtain a multiplication result b1; Connect the output of the sixteenth convolutional layer and the weighted summation result a to obtain a connection result a2; Use the connection result a2 as the input to the nineteenth convolutional layer with a kernel size of 1×1, and use the output of the nineteenth convolutional layer as the input to the second sigmoid activation function layer; Use the connection result a2 as the input to the second spatial attention unit, add the output of the second sigmoid activation function layer and the output of the second spatial attention unit to obtain an added result c, and then multiply the added result c by the output of the sixteenth convolutional layer to obtain a multiplication result c1; Then add the multiplication result b1 and the multiplication result c1, and use the added result as the output of the attention - mechanism - based fusion module.
4. The method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 3, characterized in that, The working process of the first feature extractor is as follows: Inside the first feature extractor, use the input of the first feature extractor as the input to the first branch and the second branch respectively, where: The first branch sequentially includes the twentieth convolutional layer with a kernel size of 3×3, the first Relu activation function layer, the twenty - first convolutional layer with a kernel size of 3×3, the second Relu activation function layer, the twenty - second convolutional layer with a kernel size of 3×3, and the third Relu activation function layer, that is, use the output of the third Relu activation function layer as the output of the first branch; The second branch sequentially includes a twenty-third convolutional layer with a kernel size of 3×3 and a twenty-fourth convolutional layer with a kernel size of 3×3, that is, the output of the twenty-fourth convolutional layer is used as the output of the second branch; The output of the first branch and the output of the second branch are added, and the added result is used as the input of the third branch and the input of the fourth branch respectively, where: The third branch sequentially includes a twenty-fifth convolutional layer with a kernel size of 3×3, a fourth Relu activation function layer, a twenty-sixth convolutional layer with a kernel size of 3×3, a fifth Relu activation function layer, a twenty-seventh convolutional layer with a kernel size of 3×3, and a sixth Relu activation function layer, that is, the output of the sixth Relu activation function layer is used as the output of the third branch; The fourth branch sequentially includes a twenty-eighth convolutional layer with a kernel size of 3×3 and a twenty-ninth convolutional layer with a kernel size of 3×3, that is, the output of the twenty-ninth convolutional layer is used as the output of the fourth branch; The output of the third branch and the output of the fourth branch are added, and the added result is used as the output of the first feature extractor.
5. The method for identifying martial arts movements and evaluating movement quality based on artificial intelligence according to claim 4, wherein, The action recognition network includes T spatio-temporal feature extraction modules. The working process of the action recognition network is as follows: The fused image at the t-th acquisition time point is used as the input of the t-th spatio-temporal feature extraction module, where t = 1, 2, …, T; and the output of the first spatio-temporal feature extraction module is used as the input of the first GRU module; The output of the t-th spatio-temporal feature extraction module and the output of the (t - 1)-th GRU module are used as the input of the t-th GRU module; The outputs of the first GRU module to the T-th GRU module are averaged, and then the averaged result passes through a fully connected layer. Finally, the output of the fully connected layer is used as the input of the Softmax layer, and the action recognition result is output through the Softmax layer.
6. The method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 5, wherein The working process of the first spatio-temporal feature extraction module is as follows: In the first spatio-temporal feature extraction module, the input fused image is first used as the input of the thirtieth convolutional layer, and then the output of the thirtieth convolutional layer is used as the input of the first spatio-temporal feature extraction unit. Then, the output of the first spatio-temporal feature extraction unit is used as the input of the second spatio-temporal feature extraction unit, and the output of the second spatio-temporal feature extraction unit is used as the input of the third spatio-temporal feature extraction unit; The output of the thirtieth convolutional layer and the output of the third spatio-temporal feature extraction unit are added, and the added result is used as the output of the first spatio-temporal feature extraction module.
7. A method for identifying martial arts movements and evaluating movement quality based on artificial intelligence according to claim 6, wherein The working process of the first spatio-temporal feature extraction unit is as follows: In the first spatio-temporal feature extraction unit, the input of the first spatio-temporal feature extraction unit is used as the input of the thirty-first convolutional layer, and then the output of the thirty-first convolutional layer is used as the input of the first BN layer. The output of the first BN layer is used as the input of the seventh ReLU activation function layer; The output of the seventh ReLU activation function layer is used as the input of the thirty-second convolutional layer, and then the output of the thirty-second convolutional layer is used as the input of the second BN layer. The output of the second BN layer is used as the input of the eighth ReLU activation function layer, and the output of the eighth ReLU activation function layer is used as the input of the first channel attention unit; Take the input of the first spatio-temporal feature extraction unit as the input of the thirty-third convolutional layer, then take the output of the thirty-third convolutional layer as the input of the third BN layer, and take the output of the third BN layer as the input of the ninth ReLU activation function layer; Take the output of the ninth ReLU activation function layer as the input of the thirty-fourth convolutional layer, then take the output of the thirty-fourth convolutional layer as the input of the fourth BN layer, and take the output of the fourth BN layer as the input of the tenth ReLU activation function layer; Take the input of the first spatio-temporal feature extraction unit as the input of the thirty-fifth convolutional layer, then take the output of the thirty-fifth convolutional layer as the input of the fifth BN layer, and take the output of the fifth BN layer as the input of the eleventh ReLU activation function layer; Take the output of the eleventh ReLU activation function layer as the input of the thirty-sixth convolutional layer, then take the output of the thirty-sixth convolutional layer as the input of the sixth BN layer, take the output of the sixth BN layer as the input of the twelfth ReLU activation function layer, and take the output of the twelfth ReLU activation function layer as the input of the third spatial attention unit; Multiply the output of the tenth ReLU activation function layer by the output of the first channel attention unit to obtain the multiplication result e; Multiply the output of the tenth ReLU activation function layer by the output of the third spatial attention unit to obtain the multiplication result f; Add the multiplication result e and the multiplication result f to obtain the addition result g; Multiply the output of the third spatial attention unit by the output of the first channel attention unit to obtain the multiplication result h; Add g and h to obtain the addition result k, and take the addition result k as the output of the first spatio-temporal feature extraction unit.
8. A method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 7, characterized in that The specific method of extracting the human skeleton from the point cloud data obtained at each acquisition time point is as follows: Take the point cloud data obtained at any one acquisition time point as an example Step 1: Initialize the iteration number n = 1; Step 2: Initialize l = 1; Step 3: For the l-th point in the point cloud data, calculate the search radius r centered on the l-th point: where, l k represents the k-th neighboring point of point l, and dist(l, l k ) represents the distance between point l and point l k , where k = 1, 2, …, K; Search for the point set within the region centered on point l with radius r, calculate the center of all points in the point set, and then move point l to the position of the calculated center; Step 4: Determine whether l has traversed all points in the point cloud data; If not all points in the point cloud data have been traversed, let l = l + 1, and return to execute Step 3; If all points in the point cloud data have been traversed, continue to execute Step 5; Step 5: Determine whether n = N is satisfied, where N represents the set maximum number of iterations: If n < N, let n = n + 1, and return to execute Step 2; If n = N, continue to return to execute Step 6; Step 6: Extract the human skeleton from the point cloud data obtained after the N-th iteration.
9. A method for martial arts action recognition and action quality evaluation based on artificial intelligence according to claim 8, characterized in that, The specific process of smoothing the extracted skeleton to obtain the human skeleton corresponding to each acquisition time point is as follows: where, x t′,j represents the coordinate of the j-th joint point in the human skeleton corresponding to the t'-th acquisition time point, and x t′,j represents the smoothed coordinate of the j-th joint point in the human skeleton corresponding to the t'-th acquisition time point.
10. The method for recognizing martial arts movements and evaluating movement quality based on artificial intelligence according to claim 9, characterized in that, The specific process of Step 5 is as follows: Step 51: Calculate the direction vectors of each joint according to the coordinates of each joint point of the human skeleton, and obtain the standard actions from the database according to the martial arts action recognition result in Step 3; Step 52: Calculate the cosine similarity between the joint direction vectors calculated in Step 51 and the direction vectors of the standard actions; Step 53: Obtain the final evaluation result of the martial arts movement according to the calculation result of Step 52.
Citation Information
Patent Citations
Wushu action recognition method based on human body posture estimation
CN114419505A
SAR vessel detection method based on multi-scale feature extraction and FRFT convolution
CN118884439A
Rock mass trace automatic extraction method and system based on point cloud
CN119169304A
Point cloud skeleton extraction method and apparatus
WO2014187046A1