A Video Action Recognition Method and System for Adaptive Movement Rhythm Learning
By splitting and identifying teaching videos, combining the data and information of athletes, adjusting the playback speed and rest time, the problem of athletes being unable to keep up with the rhythm is solved and the fitness effect is improved.
Patent Information
- Application Number
- CN202410904043.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-08
AI Technical Summary
In the prior art, athletes are prone to fail to keep up with the rhythm when following fitness teaching videos, resulting in inaccurate movements, affecting the exercise effect, and lack an effective feedback mechanism to adjust the video playback speed and rest time.
By splitting the teaching video, identifying the action marking data set, and decomposing the exercise videos of athletes frame by frame, calculating the action completion, standards and fluency coefficients, adjusting the playback speed and rest time, and weighting calculations based on personal information and motion data.
Accurately adjusting the video playback speed and rest time improves the learning and exercise effects of athletes, ensuring the accuracy of movement and exercise safety.
Smart Images

Figure CN118823876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video action recognition, and specifically to a video action recognition method and system for adaptive exercise rhythm learning. Background Art
[0002] Learning fitness actions through fitness teaching videos for exercise is a common current mode, which has low requirements for the time and space of exercisers, good operability and fitness effects. The content of fitness teaching videos includes dumbbell exercises, yoga, aerobics, etc., effectively expanding the content and methods of fitness.
[0003] In the process of exercisers just starting to contact fitness teaching videos, due to unfamiliarity with the actions and processes, problems such as not being able to keep up with the video rhythm and inaccurate actions are likely to occur; this not only dampens the enthusiasm of the exercisers, but also fails to achieve the due exercise effect; therefore, it is necessary to recognize the actions of exercisers and adjust the fitness videos according to the results to improve the exercise effect.
[0004] In the prior art, the patent with the application number CN202410146629.0 proposed a video action recognition method and system based on adaptive exercise rhythm learning; after fusing the compressed video features and the forward input features to obtain the first output feature, the motion-related tensor and the first output feature are fused in the form of element addition to obtain the second output feature; based on the second output feature, the action category to which the video belongs is recognized; improving the performance of the video-based action recognition method. The patent with the application number CN202210517629.8 proposed an action recognition method based on a spatio-temporal smooth feature network; the network model reads video data through a server and performs equally spaced frame division; uses an action detector to extract features from the video information, uses a spatio-temporal smooth feature fusion method to smooth the features in the time domain and the space domain to complete feature extraction, and uses a deep learning method to comprehensively analyze the features to judge the target action and accurately detect the target action. The patent with the application number CN202311338628.8 proposed a human action recognition method, system and storage medium based on video sequences; uses a motion branch and a spatial branch and the fusion of the two branches to realize the feature fusion of motion information, appearance information, and multi-frequency domain information, and adds an adaptive multi-frequency domain self-attention cross-fusion module during the fusion process to improve the frequency adaptability in a more flexible way and improve the recognition effect.
[0005] The prior art mainly recognizes the actions of exercisers, without establishing a linkage and feedback mechanism between exercisers and fitness videos, and cannot well solve the situation of poor exercise effects caused by exercisers not being able to keep up with the video rhythm.
[0006] To this end, a video action recognition method and system for adaptive exercise rhythm learning are proposed. Summary of the Invention
[0007] The purpose of the present invention is to provide a video action recognition method and system for adaptive exercise rhythm learning, which splits a teaching video into teaching sub-videos; recognizes the teaching sub-videos to obtain an action marking data set; decomposes the exercise sub-videos of the exerciser frame by frame to obtain an exercise image set; recognizes the exercise image set through the action marking data set to obtain an action completion coefficient, an action standard coefficient, and an action fluency coefficient, and weights the playback speed to obtain a corrected playback speed; obtains exercise data through the exercise sub-videos, determines a threshold according to the personal information of the exerciser, and obtains a fatigue coefficient according to the exercise data and the threshold; weights and calculates the rest time according to the fatigue coefficient to obtain a corrected rest time; the present invention effectively improves the learning and exercise effects of the exerciser by adjusting the playback speed and rest time of the teaching video.
[0008] To achieve the above object, the present invention provides a video action recognition method for adaptive exercise rhythm learning, including:
[0009] Decompose the actions of the teaching video to obtain a teaching video data set; the teaching video data set includes teaching sub-videos and rest times; perform action recognition and marking on the teaching sub-videos to obtain an action marking data set;
[0010] Collect the personal information of the exerciser and store it; when the teaching sub-video starts to play, obtain a search image set through a shooting device; identify and determine the position of the exerciser according to the search image set and the personal information;
[0011] Track and shoot the actions of the exerciser, and obtain the exercise video of the exerciser during the period from the start to the end of the teaching sub-video to obtain an exercise sub-video;
[0012] Recognize the exercise sub-video through the action marking data set of the teaching sub-video to obtain the action completion coefficient, the action standard coefficient, and the action fluency coefficient of the exerciser; calculate an action recognition coefficient according to the action completion coefficient, the action standard coefficient, and the action fluency coefficient; adjust the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain a corrected playback speed;
[0013] Collect data on the exerciser in the exercise sub-video to obtain an exercise data set; calculate a fatigue coefficient through the exercise data set and the personal information; adjust the rest time according to the fatigue coefficient of the exerciser to obtain a corrected rest time.
[0014] The personal information includes the gender, age, height, weight, past medical history, and facial image of the athlete; the sports data set includes the heart rate, respiratory rate data, and video screenshots of the athlete in the sports video.
[0015] The steps for obtaining the action completion coefficient, action standard coefficient, and action fluency coefficient are as follows:
[0016] Obtain a teaching sub-video, identify the landmark actions in the teaching sub-video; obtain an action marking data set based on the landmark actions, and the action marking data set includes teaching marking images and teaching marking time points;
[0017] Obtain a practice sub-video corresponding to the teaching sub-video; perform frame-by-frame segmentation on the practice sub-video to obtain a set of practice images; traverse the set of practice images in chronological order to obtain the action similarity between the practice images and the first teaching marking image; when the action similarity is greater than the threshold, determine that the practice image is the first practice marking image, and record the time point as the first practice marking time point;
[0018] Update the set of practice images, delete the motion images at and before the first practice marking time point to obtain a second set of practice images; traverse the second set of practice images according to the second teaching marking image, identify the second practice marking images, and update to obtain a third set of practice images;
[0019] Perform traversal recognition and update of the set of practice images in the above method. When the action similarity of the practice images is less than the threshold or all the teaching marking images have been recognized, terminate the recognition, and obtain the finally recognized practice marking images and practice marking time points;
[0020] Obtain the action completion coefficient according to the finally recognized practice marking time point and the total time of the teaching sub-video; the calculation formula for the action completion degree coefficient is:
[0021]
[0022] where act i represents the action completion coefficient of the i-th teaching sub-video; t j is the teaching marking time point corresponding to the finally recognized practice marking time point; T i is the total time of the i-th teaching sub-video;
[0023] Identify the action standard coefficient according to the action similarity between the teaching marking image and the practice marking image;
[0024]
[0025] Among them, sta i represents the action standard coefficient of the i-th teaching sub-video; sim() represents the action similarity recognition model; Mv c represents the c-th teaching marker image; Fv c represents the c-th practice marker image; j represents the total number of practice marker images;
[0026] Through the teaching marker time point and the practice marker time point, the action fluency coefficient is recognized;
[0027]
[0028] Among them, flu i represents the action fluency coefficient of the i-th teaching sub-video; t d represents the teaching marker time point where the d-th teaching marker image is located; t o d represents the practice marker time point where the d-th practice marker image is located.
[0029] The calculation formula for correcting the playback speed is:
[0030] A i =α1*exp(act i -1)+α2*sta i +α3*flu i ;
[0031]
[0032] Among them, A i represents the action recognition coefficient of the i-th practice sub-video; α1 represents the first weight; act i represents the action completion coefficient of the i-th practice sub-video; α2 represents the second weight; sta i represents the action standard coefficient of the i-th practice sub-video; α3 represents the third weight; flu i represents the action fluency coefficient of the i-th practice sub-video; Play i+1 represents the corrected playback speed of the (i + 1)-th teaching sub-video; P max represents the upper limit of the playback speed of the teaching sub-video; P min represents the lower limit of the playback speed of the teaching sub-video; PH represents the action recognition coefficient threshold; exp represents the exponential function with the natural constant as the base.
[0033] The calculation process of the fatigue coefficient is as follows: Obtain the personal information of the athlete; Set the exercise threshold according to the personal information of the athlete; Calculate the fatigue coefficient of the athlete based on the exercise data and the exercise threshold;
[0034]
[0035] Among them, Ftg i is the fatigue coefficient of the i-th exercise sub-video; breathe i represents the breathing abnormality coefficient of the i-th exercise sub-video; sweat i represents the dehydration prediction; ST represents the dehydration threshold; heart i represents the maximum heart rate; HT represents the heart rate threshold.
[0036] Adjust the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time;
[0037] Rest i =Re i *ln(e + Ftg i );
[0038] Among them, Rest i represents the i-th corrected rest time; Re i the i-th rest time; Ftg i represents the fatigue coefficient of the i-th exercise sub-video; e represents the natural constant; ln represents the logarithmic function with the natural constant as the base.
[0039] A video action recognition system for adaptive exercise rhythm learning, comprising:
[0040] A teaching video decomposition module that decomposes the actions of the teaching video to obtain a teaching video dataset; the teaching video dataset includes teaching sub-videos and rest times; performs action recognition and marking on the teaching sub-videos to obtain an action marking dataset;
[0041] An athlete positioning module that collects and stores the personal information of the athlete; when the teaching sub-video starts to play, obtains a search image set through a shooting device; determines the position of the athlete according to the search image set and the personal information;
[0042] A motion video acquisition module that tracks and shoots the actions of the athlete to obtain a motion video of the athlete during the period from the start to the end of the teaching sub-video, and obtains an exercise sub-video;
[0043] A playback speed adjustment module that identifies the exercise sub-video through the action marking dataset of the teaching sub-video to obtain the action completion coefficient, action standard coefficient, and action fluency coefficient of the athlete; calculates an action recognition coefficient according to the action completion coefficient, the action standard coefficient, and the action fluency coefficient; adjusts the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain a corrected playback speed;
[0044] The rest time adjustment module collects data on the exerciser in the exercise sub-video to obtain an exercise data set; calculates a fatigue coefficient based on the exercise data set and the personal information; and adjusts the rest time according to the fatigue coefficient of the exerciser to obtain a corrected rest time.
[0045] The formula for calculating the corrected playback speed is:
[0046] A i =α1*exp(act i -1)+α2*sta i +α3*flu i ;
[0047]
[0048] Wherein, A i represents the action recognition coefficient of the i-th exercise sub-video; α1 represents the first weight; act i represents the action completion coefficient of the i-th exercise sub-video; α2 represents the second weight; sta i represents the action standard coefficient of the i-th exercise sub-video; α3 represents the third weight; flu i represents the action fluency coefficient of the i-th exercise sub-video; Play i+1 represents the corrected playback speed of the (i + 1)-th teaching sub-video; P max represents the upper limit of the playback speed of the teaching sub-video; P min represents the lower limit of the playback speed of the teaching sub-video; PH represents the action recognition coefficient threshold; exp represents the exponential function with the natural constant as the base.
[0049] The calculation process of the fatigue coefficient is as follows: obtain the personal information of the exerciser; set an exercise threshold according to the personal information of the exerciser; calculate the fatigue coefficient of the exerciser based on the exercise data and the exercise threshold;
[0050]
[0051] Wherein, Ftg i is the fatigue coefficient of the i-th exercise sub-video; breathe i represents the breathing abnormality coefficient of the i-th exercise sub-video; sweat i represents the predicted dehydration amount; ST represents the dehydration amount threshold; heart i represents the maximum heart rate; HT represents the heart rate threshold.
[0052] Adjust the rest time according to the fatigue coefficient of the exerciser to obtain a corrected rest time;
[0053] Rest i =Re i *ln(e + Ftg i );
[0054] Among them, Rest i represents the i-th corrected rest time; Re i the i-th rest time; Ftg i represents the fatigue coefficient of the i-th exercise sub-video; e represents the natural constant; ln represents the logarithmic function with the natural constant as the base.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] 1. The present invention splits the teaching video to obtain teaching sub-videos; identifies and splits the teaching sub-videos to obtain an action marker dataset; decomposes the obtained exercise sub-videos frame by frame to obtain an exercise image set; identifies the exercise image set through the action marker dataset to obtain exercise marker images and exercise marker time points; and then calculates to obtain an action completion coefficient, an action standard coefficient, and an action fluency coefficient; accurately measuring the actions of the exercise sub-videos.
[0057] 2. The present invention obtains the upper and lower limits of the playback speed of the fitness teaching video, and weights the playback speed through the obtained action completion coefficient, action standard coefficient, and action fluency coefficient to obtain a corrected playback speed; accurately adjusting the video speed according to the action conditions of the current exerciser.
[0058] 3. Obtain the breathing abnormality coefficient of the exerciser through the breathing frequency data; obtain the dehydration prediction through the video screenshot, and obtain the heart rate of the exerciser through the detection device; determine the threshold of the exercise according to the personal information and data of the exerciser; accurately calculate the fatigue coefficient of the exerciser according to the obtained exercise data and exercise threshold; weight and calculate the rest time according to the fatigue coefficient to obtain a corrected rest time, and accurately adjust the rest time according to the exercise conditions of the exerciser. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a schematic flow chart of a video action recognition method for adaptive exercise rhythm learning of the present invention;
[0060] Figure 2 is a schematic flow chart of the recognition of the exercise sub-video of the present invention;
[0061] Figure 3 is a schematic diagram of the teaching marker image of the body rotation movement of the broadcast gymnastics of the present invention;
[0062] Figure 4Schematic diagram of the structure of a video action recognition system for adaptive exercise rhythm learning according to the present invention. Detailed implementation manners
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] In the process of a sports person just starting to watch a fitness teaching video, due to unfamiliarity with the actions and processes, problems such as not being able to keep up with the video rhythm and inaccurate actions are likely to occur; this not only dampens the enthusiasm of the sports person, but also fails to achieve the desired exercise effect; therefore, it is necessary to identify the actions of the sports person and adjust the fitness video according to the results to improve the exercise effect.
[0065] For this reason, a video action recognition method and system for adaptive exercise rhythm learning are proposed.
[0066] Embodiment 1
[0067] Taking the teaching video of broadcast gymnastics as an example, the teaching video of broadcast gymnastics contains a total of nine sections, namely marching in place, stretching exercise, chest expanding exercise, kicking exercise, side body exercise, body turning exercise, whole body exercise, jumping exercise and finishing exercise; a 10-second rest time is set between each section.
[0068] The process of the video action recognition method for adaptive exercise rhythm learning is as Figure 1 shown, and includes:
[0069] Decompose the actions of the broadcast gymnastics teaching video to obtain a teaching video data set; the teaching video data set includes teaching sub-videos and rest times, denoted as {Dz1, Re1, Dz2, Re2,..., Dz i , Re i ,..., Dz n-1 , Re n-1 , Dz n}; where Dz i represents the i-th teaching sub-video in the teaching video; Re i represents the i-th rest time in the teaching video; n takes the value of 9.
[0070] Collect the personal information of the sports person and store it; where the personal information is uploaded and updated by the sports person, and includes the gender, age, height, weight, past medical history and facial image of the sports person.
[0071] The present invention obtains and stores personal information of a sports person, preliminarily understands and evaluates the physical condition of the sports person through the personal information, and sets corresponding exercise thresholds; provides data support for the subsequent calculation of the fatigue coefficient according to the exercise thresholds and exercise data sets.
[0072] When starting to learn the teaching video, the teaching sub-videos are played in sequence; when the teaching video starts to be played, a search image set is obtained by shooting with a shooting device; the position of the sports person is identified according to the search image set and the personal information; wherein, the search image set is obtained by shooting at a specified sports area, and the position of the target person is determined by comparing the search image set with the facial image of the sports person.
[0073] The actions of the sports person are tracked and photographed, and the exercise video of the sports person during the period from the start to the end of the teaching sub-video is obtained to get the practice sub-video.
[0074] The practice sub-video is identified through the action marking data set of the teaching sub-video to obtain the action completion coefficient, action standard coefficient and action fluency coefficient of the sports person; the action recognition coefficient is calculated according to the action completion coefficient, the action standard coefficient and the action fluency coefficient; the video playback speed of the next teaching sub-video is adjusted according to the action recognition coefficient to obtain the corrected playback speed.
[0075] The obtaining steps of the action completion coefficient, action standard coefficient and action fluency coefficient are as Figure 2 shown;
[0076] Obtain the teaching sub-video, identify the marked actions in the teaching sub-video to determine the marked actions; obtain the teaching marked images and teaching marked time points according to the marked actions to get the action marking data set;
[0077] The action marking data set is denoted as {(Dz i , T i )|(Mv1, t1),..., (Mv j , t j ),..., (Mv m , t m )}; where, Dz i represents the i-th teaching sub-video; T i represents the duration of the i-th teaching sub-video; Mv j represents the j-th teaching marked image of the i-th teaching sub-video; t j is the teaching marked time point where the j-th teaching marked image is located;
[0078] Obtain the i-th practice sub-video and perform frame-by-frame segmentation on it to obtain a set of practice images;
[0079] Traverse the set of practice images in chronological order to obtain the action similarity between the practice images and the first teaching marker images; when the action similarity is greater than the threshold, determine that the practice image is the first practice marker image, and record the time point where the first practice marker image is located as the first practice marker time point; at this time, update the set of practice images, delete the practice images before and including the first practice marker time point, and obtain a second set of practice images;
[0080] Traverse the second set of practice images in chronological order to obtain the action similarity between the practice images in the second set of practice images and the second teaching marker images. When the action similarity is greater than the threshold, determine that the motion image is the second practice marker image, and record the time point where the second practice marker image is located as the second practice marker time point; update the second set of practice images, delete the practice images before and including the second practice marker time point, and obtain a third set of practice images;
[0081] Perform traversal and recognition by the above method. When the action similarity of the practice images is less than the threshold or all the teaching marker images have been recognized, terminate the recognition, and obtain the finally recognized practice marker images and practice marker time points;
[0082] Obtain the corresponding teaching marker time points according to the finally recognized practice marker time points, and obtain the action completion coefficient according to the progress ratio of the teaching marker time points in the total time of the i-th teaching sub-video;
[0083] The calculation formula of the action completion coefficient is:
[0084]
[0085] where, act i represents the action completion coefficient of the i-th practice sub-video; t j is the corresponding teaching marker time point obtained from the finally recognized practice marker time point; T i is the total time of the i-th teaching sub-video.
[0086] Identify the action standard coefficient according to the action similarity between the teaching marker images and the practice marker images;
[0087]
[0088] where, sta i represents the action standard coefficient of the i-th practice sub-video; sim() represents the action similarity recognition model; Mvc represents the c-th teaching marker image; Fv c represents the c-th practice marker image; j represents the total number of practice marker images.
[0089] Among them, the recognition of action similarity is as follows: The action is described as the movement trajectory or pose image of skeletal joints, and the skeletal joints of the moving person are described by setting joint nodes; then a convolutional neural network is used for feature extraction and similarity measurement to obtain the action similarity; the action similarity recognition model is trained and optimized based on the convolutional neural network.
[0090] The action smoothness coefficient is recognized through the time points corresponding to the action marker image and the motion marker image;
[0091]
[0092] where flu i represents the action smoothness coefficient of the i-th practice sub-video; t d represents the teaching marker time point at which the d-th teaching marker image is located; t o d represents the practice marker time point at which the d-th practice marker image is located.
[0093] Combined with Figure 3 , taking the body rotation movement as an example, the teaching marker image is explained; the body rotation movement includes a total of four identical eight-beat rhythms, and a complete eight-beat movement is as follows:
[0094] The initial posture is the attention posture, as shown in a of Figure 3 ;
[0095] In the first beat, the left leg steps to the left, slightly wider than the shoulders, and at the same time, both arms are raised horizontally to the sides, with the palms facing down, as shown in b of Figure 3 ;
[0096] In the second beat, the lower part of the waist maintains the posture of the first beat, and the upper part of the waist turns 90 degrees to the left. At the same time, the hands clap twice in front of the chest, as shown in c of Figure 3 ;
[0097] In the third beat, the upper part of the waist turns 180 degrees to the right, and at the same time, the arms are straightened and raised to the upper sides, with the palms facing inwards, as shown in d of Figure 3 ;
[0098] In the fourth beat, the left foot returns to the attention posture, and at the same time, the body turns straight, and the two arms return to the sides of the body through the sides, as shown in e of Figure 3 ;
[0099] In the fifth beat, the right leg steps to the right, slightly wider than the shoulders, and at the same time, both arms are raised horizontally to the sides, with the palms facing down;
[0100] For the sixth beat, keep the position from the first beat below the waist, turn the upper body to the right by 90 degrees, and clap the hands twice in front of the chest at the same time;
[0101] For the seventh beat, turn the upper body to the left by 180 degrees, and at the same time extend the arms straight up to the sides with the palms facing inwards;
[0102] For the eighth beat, restore the right foot to the standing-at-attention position, and at the same time, turn the body straight, and the two arms return to the sides of the body through the sides;
[0103] Among them, the movements of the fifth, sixth, seventh, and eighth beats are the same as those of the first, second, third, and fourth beats respectively, but in the opposite direction; the above nine movements are set as teaching marker images.
[0104] Set the time of one complete eight-beat cycle to 8 seconds, and the time of one beat to 1 second; the action marker data of the first eight beats of the body rotation exercise sub-video are shown in the following table; among them, the threshold of action similarity is set to 0.70.
[0105] Table 1 Partial recognition data table of the body rotation exercise video in the broadcast gymnastics
[0106]
[0107] The present invention splits the teaching video to obtain teaching sub-videos; identifies and splits the teaching sub-videos to obtain an action marker data set; decomposes the obtained practice sub-videos frame by frame to obtain a practice image set; identifies the practice image set through the action marker data set to obtain practice marker images and practice marker time points; and then calculates the action completion coefficient, action standard coefficient, and action fluency coefficient; effectively and accurately measures the practice sub-videos.
[0108] Calculate the action recognition coefficient according to the action completion coefficient, action standard coefficient, and action fluency coefficient; adjust the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain the corrected playback speed;
[0109] The calculation formula for the corrected playback speed is:
[0110] A i =α1*exp(act i -1)+α2*sta i +α3*flu i ;
[0111]
[0112] Among them, A i represents the action recognition coefficient of the i-th practice sub-video; α1 represents the first weight; act irepresents the action completion coefficient of the i-th practice sub-video; α2 represents the second weight; sta i represents the action standard coefficient of the i-th practice sub-video; α3 represents the third weight; flu i represents the action fluency coefficient of the i-th practice sub-video; Play i+1 represents the corrected playback speed of the (i + 1)-th teaching sub-video; P max represents the upper limit of the playback speed of the teaching sub-video; P min represents the lower limit of the playback speed of the teaching sub-video; PH represents the threshold of the action recognition coefficient; exp represents the exponential function with the natural constant as the base.
[0113] Among them, the first weight, the second weight, and the third weight are obtained through verification of historical data. The following table shows the recognition data of some teaching sub-videos of radio calisthenics. The upper and lower limits of the playback speed are obtained by comprehensively considering the safety and exercise effect of fitness actions.
[0114] Table 2 Recognition data table of some videos in radio calisthenics
[0115]
[0116] The present invention obtains the upper and lower limits of the playback speed of the fitness teaching video, and weights the playback speed through the obtained action completion coefficient, action standard coefficient, and action fluency coefficient to obtain the corrected playback speed; accurately adjusts the video speed according to the action conditions of the current exerciser.
[0117] Collect data on the exerciser in the practice sub-video to obtain a motion data set; calculate the fatigue coefficient through the motion data set and the personal information; adjust the rest time according to the fatigue coefficient of the exerciser to obtain the corrected rest time.
[0118] The motion data set includes the heart rate, respiratory rate data, and video screenshots of the exerciser in the motion video.
[0119] The calculation process of the fatigue coefficient is as follows:
[0120] Obtain the personal information of the exerciser; set a motion threshold according to the personal information of the exerciser; calculate the fatigue coefficient of the exerciser according to the motion data and the motion threshold;
[0121] The calculation formula of the fatigue coefficient is:
[0122]
[0123] Among them, Ftg i represents the fatigue coefficient of the i-th practice sub-video; breathe iDenote the respiratory abnormality coefficient of the i-th exercise sub-video; sweat i Denote the dehydration prediction; ST denotes the dehydration threshold; heart i Denote the maximum heart rate; HT denotes the heart rate threshold. Wherein, the dehydration threshold and the heart rate threshold are confirmed according to the personal information of the athlete.
[0124] Wherein, the respiratory abnormality coefficient is calculated through the breathing amplitude and breathing frequency of the athlete.
[0125] The present invention determines the exercise threshold according to the personal information and data of the athlete; meanwhile, obtains the exercise sub-video of the athlete; obtains the respiratory abnormality coefficient of the athlete through the breathing frequency data; obtains the dehydration prediction through the video screenshot, and obtains the heart rate of the athlete through the detection device; accurately calculates the fatigue coefficient of the athlete according to the obtained exercise data and exercise threshold.
[0126] Adjust the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time;
[0127] The calculation formula of the corrected rest time is:
[0128] Rest i =Re i *ln(e + Ftg i );
[0129] Wherein, Rest i Denotes the i-th corrected rest time; Re i The i-th rest time; Ftg i Denotes the fatigue coefficient of the i-th exercise sub-video; e denotes the natural constant; ln denotes the logarithmic function with the natural constant as the base.
[0130] The present invention performs weighted calculation on the rest time according to the calculated fatigue coefficient of the athlete to obtain the corrected rest time, and accurately adjusts the rest time according to the exercise condition of the athlete.
[0131] According to a video action recognition method for adaptive exercise rhythm learning described in the present invention, segment the teaching video of the broadcast gymnastics to obtain multiple teaching sub-videos and the intermediate rest time; then identify and mark the actions in the teaching sub-videos, identify the obtained exercise sub-videos, so as to adjust the playback speed; meanwhile, adjust the intermediate rest time according to the fatigue coefficient of the athlete, effectively improving the learning efficiency and exercise effect of the broadcast gymnastics.
[0132] The present invention splits a teaching video to obtain teaching sub - videos; performs recognition and splitting on the teaching sub - videos to obtain an action marker dataset; decomposes the obtained practice sub - videos frame by frame to obtain a practice image set; uses the action marker dataset to recognize the practice image set to obtain practice marker images and practice marker time points; then calculates an action completion coefficient, an action standard coefficient, and an action smoothness coefficient; weights the playback speed using the obtained action completion coefficient, action standard coefficient, and action smoothness coefficient to obtain a corrected playback speed, accurately adjusts the video speed according to the action situation of the current exerciser; determines a motion threshold based on the personal information and data of the exerciser; accurately calculates the fatigue coefficient of the exerciser based on the obtained motion data and motion threshold; weights the rest time according to the fatigue coefficient to obtain a corrected rest time, and accurately adjusts the rest time according to the exercise situation of the exerciser.
[0133] Embodiment 2
[0134] Taking the teaching video of dumbbell exercise as an example, the video action recognition system for adaptive exercise rhythm learning according to the present invention is used to recognize and adjust the learning of the teaching video of dumbbell exercise.
[0135] The structure of the video action recognition system for adaptive exercise rhythm learning is as Figure 4 shown, and includes a teaching video decomposition module, an exerciser positioning module, an exercise video acquisition module, a playback speed adjustment module, and a rest time adjustment module.
[0136] The teaching video decomposition module decomposes the actions of the teaching video to obtain a teaching video dataset; the teaching video dataset includes teaching sub - videos and rest times; performs action recognition and marking on the teaching sub - videos to obtain an action marker dataset;
[0137] The exerciser positioning module collects the personal information of the exerciser and stores it; when the teaching sub - video starts to play, obtains a search image set through a shooting device; determines the position of the exerciser according to the search image set and the personal information;
[0138] The exercise video acquisition module tracks and shoots the actions of the exerciser, and obtains the exercise video of the exerciser during the period from the start to the end of the teaching sub - video to obtain a practice sub - video;
[0139] The playback speed adjustment module identifies the practice sub-video through the action marker data set of the teaching sub-video to obtain the action completion coefficient, action standard coefficient, and action fluency coefficient of the athlete; calculates the action recognition coefficient based on the action completion coefficient, action standard coefficient, and action fluency coefficient; adjusts the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain the corrected playback speed.
[0140] The rest time adjustment module collects data on the athlete in the practice sub-video to obtain a motion data set; calculates the fatigue coefficient based on the motion data set and the personal information; adjusts the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time.
[0141] Further, the steps for obtaining the action completion coefficient, action standard coefficient, and action fluency coefficient are as follows:
[0142] Obtain the teaching sub-video, identify the landmark actions within the teaching sub-video; obtain the teaching marker images and teaching marker time points based on the landmark actions to obtain the action marker data set.
[0143] Obtain the practice sub-video corresponding to the teaching sub-video; perform frame-by-frame segmentation on the practice sub-video to obtain a set of practice images; traverse the set of practice images in chronological order to obtain the action similarity between the practice images and the first teaching marker image; when the action similarity is greater than the threshold, determine that the practice image is the first practice marker image, and the time point is recorded as the first practice marker time point.
[0144] Update the set of practice images, delete the motion images before and including the first practice marker time point to obtain a second set of practice images; traverse the second set of practice images according to the second teaching marker image to identify the second practice marker image.
[0145] Perform traversal recognition and update of the set of practice images in the above method, and terminate the recognition when the action similarity of the practice images is less than the threshold or all the teaching marker images have been recognized, and obtain the finally recognized practice marker images and practice marker time points.
[0146] Obtain the action completion coefficient based on the finally recognized practice marker time point and the total time of the teaching sub-video; the calculation formula for the action completion degree coefficient is:
[0147]
[0148] where act i represents the action completion coefficient of the i-th teaching sub-video; tj Obtain the corresponding teaching mark time point for the finally recognized practice mark time point; T i Total time of the i-th teaching sub-video.
[0149] Identify the action standard coefficient according to the image similarity between the teaching mark image and the practice mark image;
[0150] The calculation formula for the action standard coefficient is:
[0151]
[0152] where, sta i represents the action standard coefficient of the i-th teaching sub-video; sim() represents the action similarity recognition model; Mv c represents the c-th teaching mark image; Fv c represents the c-th practice mark image; j represents the total number of practice mark images.
[0153] Identify the action fluency coefficient through the teaching mark time point and the practice mark time point;
[0154] The calculation formula for the action fluency coefficient is:
[0155]
[0156] where, flu i represents the action fluency coefficient of the i-th teaching sub-video; t d represents the teaching mark time point at which the d-th teaching mark image is located; t o d represents the practice mark time point at which the d-th practice mark image is located.
[0157] Furthermore, calculate the action recognition coefficient according to the action completion coefficient, action standard coefficient and action fluency coefficient; adjust the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain the corrected playback speed;
[0158] The calculation formula for the corrected playback speed is:
[0159] A i = α1 * exp(act i - 1)+ α2 * sta i + α3 * flu i ;
[0160]
[0161] where, A iDenote the action recognition coefficient of the $i$-th practice sub-video; $\alpha_1$ represents the first weight; act i Denote the action completion coefficient of the $i$-th practice sub-video; $\alpha_2$ represents the second weight; sta i Denote the action standard coefficient of the $i$-th practice sub-video; $\alpha_3$ represents the third weight; flu i Denote the action fluency coefficient of the $i$-th practice sub-video; Play i+1 Denote the corrected playback speed of the $(i + 1)$-th teaching sub-video; P max Denote the upper limit of the playback speed of the teaching sub-video; P min Denote the lower limit of the playback speed of the teaching sub-video; PH represents the threshold of the action recognition coefficient; exp represents the exponential function with the natural constant as the base.
[0162] The following table shows the recognition data of some teaching sub-videos in the dumbbell exercise teaching video.
[0163] Table 3 Recognition data table of some videos in the dumbbell exercise teaching video
[0164]
[0165] Furthermore, the calculation process of the fatigue coefficient is as follows:
[0166] Obtain the personal information of the athlete; set the exercise threshold according to the personal information of the athlete; calculate the fatigue coefficient of the athlete based on the exercise data and the exercise threshold;
[0167]
[0168] Among them, Ftg i Denote the fatigue coefficient of the $i$-th practice sub-video; breathe i Denote the abnormal breathing coefficient of the $i$-th practice sub-video; sweat i Denote the predicted dehydration amount; ST represents the dehydration amount threshold; heart i Denote the maximum heart rate; HT represents the heart rate threshold.
[0169] Adjust the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time;
[0170] Rest i =Re i *ln(e + Ftg i );
[0171] Among them, Rest i Denote the $i$-th corrected rest time; Re i The $i$-th rest time; Ftg iIt represents the fatigue coefficient of the $i$-th practice sub-video; $e$ represents the natural constant; $\ln$ represents the logarithmic function with the natural constant as the base.
[0172] The present invention splits the teaching video to obtain teaching sub-videos, identifies the teaching sub-videos to obtain an action marking data set, decomposes the practice sub-videos of the athlete frame by frame to obtain a set of practice images, identifies the set of practice images through the action marking data set to obtain an action completion coefficient, an action standard coefficient, and an action fluency coefficient, and weights the playback speed to obtain a corrected playback speed. The fatigue coefficient is obtained based on the motion data obtained from the practice sub-videos and the threshold determined according to the personal information of the athlete. The rest time is weighted and calculated based on the fatigue coefficient to obtain a corrected rest time, effectively improving the learning and exercise effects of the athlete.
[0173] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A video action recognition method for adaptive exercise rhythm learning, characterized in that, Including: Decompose the actions in the teaching video to obtain a teaching video dataset; the teaching video dataset includes teaching sub-videos and rest times; Perform action recognition and marking on the teaching sub-videos to obtain an action marking dataset; Collect the personal information of the exerciser and store it; When the teaching sub-video starts playing, obtain a search image set through a shooting device; determine the position of the exerciser based on the search image set and the personal information; Track and shoot the actions of the exerciser to obtain a motion video of the exerciser during the period from the start to the end of the teaching sub-video, and obtain a practice sub-video; Identify the practice sub-video through the action marking dataset of the teaching sub-video to obtain the action completion coefficient, action standard coefficient, and action fluency coefficient of the exerciser; calculate an action recognition coefficient based on the action completion coefficient, the action standard coefficient, and the action fluency coefficient; adjust the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain a corrected playback speed; The calculation formula for the corrected playback speed is: A i = α1 * exp(act i - 1) + α2 * sta i + α3 * flu i ; Among them, A i represents the action recognition coefficient of the i-th practice sub-video; α1 represents the first weight; act i represents the action completion coefficient of the i-th practice sub-video; α2 represents the second weight; sta i represents the action standard coefficient of the i-th practice sub-video; α3 represents the third weight; flu i represents the action fluency coefficient of the i-th practice sub-video; Play i+1 represents the corrected playback speed of the (i + 1)-th teaching sub-video; P max represents the upper limit of the playback speed of the teaching sub-video; P min represents the lower limit of the playback speed of the teaching sub-video; PH represents the threshold of the action recognition coefficient; exp represents the exponential function with the natural constant as the base; Collect data on the exerciser in the practice sub-video to obtain a motion dataset; calculate a fatigue coefficient based on the motion dataset and the personal information; adjust the rest time according to the fatigue coefficient of the exerciser to obtain a corrected rest time.
2. The video action recognition method for adaptive exercise rhythm learning according to claim 1, wherein: The personal information includes the gender, age, height, weight, past medical history, and facial image of the exerciser; the motion dataset includes the heart rate, respiratory rate data, and video screenshots of the exerciser in the motion video.
3. The video action recognition method for adaptive exercise rhythm learning according to claim 1, wherein: The steps for obtaining the action completion coefficient, action standard coefficient, and action fluency coefficient are: Obtain a teaching sub-video, identify the landmark actions in the teaching sub-video; obtain an action marking dataset based on the landmark actions, and the action marking dataset includes teaching marking images and teaching marking time points; Obtain the practice sub-video corresponding to the teaching sub-video; perform frame-by-frame segmentation on the practice sub-video to obtain a set of practice images; traverse the set of practice images in chronological order to obtain the action similarity between the practice image and the first teaching marking image; when the action similarity is greater than the threshold, determine that the practice image is the first practice marking image, and the time point where it is located is recorded as the first practice marking time point; Update the set of practice images, delete the motion images before and including the first practice marking time point to obtain a second set of practice images; Traverse the second set of practice images according to the second teaching marking image, identify the second practice marking image, and update to obtain a third set of practice images; Traverse and identify the update of the image set through the above method, and terminate the identification when the action similarity of the practice images is less than the threshold or all the teaching marked images are recognized, and obtain the finally recognized practice marked image and the practice marking time point; Obtain the action completion coefficient according to the finally recognized practice marking time point and the total time of the teaching sub-video; the calculation formula of the action completion degree coefficient is: Among them, act i represents the action completion coefficient of the i-th teaching sub-video; t j is the corresponding teaching marker time point obtained from the finally recognized exercise marker time point; T i The total time of the i-th teaching sub-video; Identify the action standard coefficient according to the action similarity between the teaching marked image and the practice marked image; Among them, sta i represents the action standard coefficient of the i-th teaching sub-video; sim() represents the action similarity recognition model; Mv c represents the c-th teaching marker image; Fv c represents the c-th practice marker image; j represents the total number of practice marker images; Identify the action fluency coefficient through the teaching marking time point and the practice marking time point; Among them, flu i represents the action fluency coefficient of the i-th teaching sub-video; t d represents the teaching mark time point where the d-th teaching mark image is located; t o d represents the practice mark time point where the d-th practice mark image is located.
4. A video action recognition method for adaptive exercise rhythm learning according to claim 1, characterized in that: The calculation process of the fatigue coefficient is as follows: Obtain the personal information of the athlete; set the exercise threshold according to the personal information of the athlete; calculate the fatigue coefficient of the athlete according to the exercise data and the exercise threshold; Among them, Ftg i represents the fatigue coefficient of the i-th exercise sub-video; breathe i represents the breathing abnormality coefficient of the i-th exercise sub-video; sweat i represents the dehydration prediction; ST represents the dehydration threshold; heart i represents the maximum heart rate; HT represents the heart rate threshold.
5. A video action recognition method for adaptive exercise rhythm learning according to claim 1, characterized in that: Adjust the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time; Rest i =Rei*ln(e+Ftg i ); Among them, Rest i represents the i-th corrected rest time; Re i the i-th rest time; Ftg i represents the fatigue coefficient of the i-th exercise sub-video; e represents the natural constant; ln represents the logarithmic function with the natural constant as the base.
6. A video action recognition system for adaptive learning of motion rhythm, characterized in that, Including: A teaching video decomposition module that decomposes the actions of the teaching video to obtain a teaching video data set; the teaching video data set includes teaching sub-videos and rest times; Perform action recognition and marking on the teaching sub-video to obtain an action marking data set; An athlete positioning module that collects and stores the personal information of the athlete; when the teaching sub-video starts to play, obtain a search image set through a shooting device; identify and determine the position of the athlete according to the search image set and the personal information; A motion video acquisition module that tracks and shoots the actions of the athlete to obtain a motion video of the athlete during the period from the start to the end of the teaching sub-video, and obtain a practice sub-video; A playback speed adjustment module that identifies the practice sub-video through the action marking data set of the teaching sub-video to obtain the action completion coefficient, action standard coefficient, and action fluency coefficient of the athlete; calculate the action recognition coefficient according to the action completion coefficient, the action standard coefficient, and the action fluency coefficient; adjust the video playback speed of the next teaching sub-video according to the action recognition coefficient to obtain the corrected playback speed; The calculation formula of the corrected playback speed is: A i = α1 * exp(act i - 1)+ α2 * sta i + α3 * flu i ; Among them, A i represents the action recognition coefficient of the i-th practice sub-video; α1 represents the first weight; act i represents the action completion coefficient of the i-th practice sub-video; α2 represents the second weight; sta i represents the action standard coefficient of the i-th practice sub-video; α3 represents the third weight; flu i represents the action fluency coefficient of the i-th practice sub-video; Play i+1 represents the corrected playback speed of the (i + 1)-th teaching sub-video; P max represents the upper limit of the playback speed of the teaching sub-video; P min represents the lower limit of the playback speed of the teaching sub-video; PH represents the action recognition coefficient threshold; exp represents the exponential function with the natural constant as the base; A rest time adjustment module that collects data on the athlete in the practice sub-video to obtain a motion data set; calculates the fatigue coefficient through the motion data set and the personal information; adjusts the rest time according to the fatigue coefficient of the athlete to obtain the corrected rest time.
7. A video action recognition system for adaptive exercise rhythm learning according to claim 6, characterized in that: The calculation process of the fatigue coefficient is as follows: Obtain the personal information of the athlete; Set a motion threshold according to the personal information of the athlete; Calculate the fatigue coefficient of the athlete based on the motion data and the motion threshold; Among them, Ftg i represents the fatigue coefficient of the i-th exercise sub-video; breathe i represents the abnormal breathing coefficient of the i-th exercise sub-video; sweat i represents the predicted dehydration amount; ST represents the dehydration amount threshold; heart i represents the maximum heart rate; HT represents the heart rate threshold.
8. An adaptive motion rhythm learning video action recognition system according to claim 6, characterized in that: Adjust the rest time according to the fatigue coefficient of the athlete to obtain a corrected rest time; Rest i = Re i *ln(e + Ftg i ) Among them, Rest i represents the i-th corrected rest time; Re i the i-th rest time; Ftg i represents the fatigue coefficient of the i-th practice sub-video; e represents the natural constant; ln represents the logarithmic function with the natural constant as the base.
Citation Information
Patent Citations
Action recognition method based on space-time smooth feature network
CN114926761A
Human body action recognition method and system based on video sequence, and storage medium
CN117079352A
Video motion recognition method and system based on adaptive motion rhythm learning
CN117994845A
Motion teaching system based on AI visual perception technology
CN110751050A
Human motion posture recognition and evaluation method and system
CN111401270A