Video content analysis method based on human bone point recognition action

By fragmenting and segmenting exercise videos and using skeletal algorithms to identify them, exercise evaluation reports are generated, solving the problem of difficulty in evaluating fitness effects without AI devices and improving the user's exercise experience.

CN117238028BActive Publication Date: 2026-05-12SHENZHEN KONKA ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN KONKA ELECTRONIC TECH CO LTD
Filing Date
2023-08-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In the absence of AI-powered fitness equipment, it is difficult to effectively evaluate the effects of exercise and fitness.

Method used

By acquiring video files, the video source is fragmented and segmented. A skeletal algorithm is used to perform action recognition and evaluation on the segmented video segments, generating action evaluation data and outputting a video analysis and evaluation report.

Benefits of technology

It enables the evaluation of exercise effects even without AI fitness equipment, thus improving the user's exercise and fitness experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238028B_ABST
    Figure CN117238028B_ABST
Patent Text Reader

Abstract

The application discloses a video content analysis method based on human skeleton point identification action, and the method comprises the following steps: acquiring a video file, and performing fragmentation cutting on the video source of the video file to obtain a cutting video segment; performing action identification and evaluation on the cutting video segment based on a skeleton algorithm to obtain action evaluation data; and outputting a video analysis evaluation report based on the action evaluation data. The application can realize the motion effect evaluation without the intelligent fitness equipment by first performing fragmentation cutting on the video file, then performing motion effect analysis and evaluation on the cutting video segment based on the skeleton algorithm, and obtaining the motion evaluation result, thereby improving the user motion fitness experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sports and health technology, and in particular to a video content analysis method based on human skeletal point recognition of motion. Background Technology

[0002] With the advancement of technology, using AI algorithms to improve the standardization of exercise movements has become a trend. Current solutions for standardizing exercise movements using AI algorithms involve embedding motion recognition points on the human skeleton within tutorial videos. These points determine the position, number, and action name of the human skeleton points in a specific playback area of ​​the video, thus judging whether the actions represented by the human skeleton points captured by the user's camera are standard. However, sometimes users exercise not in front of a device with AI-powered exercise tutorials, but through other means of video recording. How to evaluate exercise results without such a device remains a problem to be solved.

[0003] Therefore, existing technologies still need improvement. Summary of the Invention

[0004] The problem to be solved by this invention is to provide a video content analysis method based on human skeletal point recognition of movements, which addresses the above-mentioned deficiencies of the prior art and aims to solve the problem that it is inconvenient to evaluate the exercise effect in the absence of fitness equipment.

[0005] The technical solution adopted by this invention to solve the problem is as follows:

[0006] In a first aspect, embodiments of the present invention provide a video content analysis method based on human skeleton point recognition of actions, wherein the method includes:

[0007] Obtain a video file and perform fragmented segmentation of the video file to obtain segmented video segments;

[0008] Based on the skeletal algorithm, the cut video segment is subjected to motion recognition and evaluation of skeletal points to obtain motion evaluation data;

[0009] Based on the motion evaluation data, a video analysis and evaluation report is output.

[0010] In one implementation, the step of acquiring a video file and fragmenting the video file to obtain segmented video segments includes:

[0011] Obtain the video file and read its duration;

[0012] If the duration of the video file exceeds the preset duration, the video codec is invoked to decode the video file;

[0013] Obtain preset cutting rules, and perform fragmented cutting of the decoded video file based on the preset cutting rules to obtain the cut video segments. The preset cutting rules include the number of actions input by the user or the preset cutting duration.

[0014] In one implementation, the action evaluation data includes: action recognition results, action matching degree, and action evaluation results. The step of performing action recognition and evaluation on the segmented video based on a skeletal algorithm to obtain the action evaluation data includes:

[0015] The action recognition result is obtained by performing motion recognition on the skeletal points of the cut video segment based on the skeletal algorithm.

[0016] The action recognition result is matched with a preset standard action to obtain the action matching degree;

[0017] Motion evaluation is performed on the action recognition result based on the action matching degree to obtain the action evaluation result.

[0018] In one implementation, the method further includes:

[0019] Action recognition is performed on the segmented video based on a skeletal algorithm to obtain action recognition results;

[0020] The number of actions in the segmented video is obtained based on the action recognition results;

[0021] If the number of actions is greater than one, the segmented video is repeatedly fragmented based on the action recognition result.

[0022] In one implementation, after repeatedly fragmenting the video segment based on the action recognition result if the number of actions is greater than one, the method further includes:

[0023] The video segments are labeled with actions based on the action recognition results.

[0024] The video segments are time-marked based on the action recognition results;

[0025] When the video file is played, the action marker and the duration marker are displayed in a window.

[0026] In one implementation, the method further includes:

[0027] Based on the action recognition results, determine whether the cut video segments contain the same action;

[0028] If the cut video segments contain the same action, then the cut video segments are marked with associated actions;

[0029] Based on the action association tags, perform a year-on-year action comparison analysis on the segmented video.

[0030] In one implementation, the step of evaluating the motion recognition result based on the motion matching degree to obtain the motion evaluation result includes:

[0031] If the action evaluation result is unsatisfactory, the cut video segment is marked.

[0032] In one implementation, the method further includes:

[0033] Facial data in the video file is obtained based on a facial recognition algorithm;

[0034] Based on the facial data, determine whether the video file contains multiple facial data;

[0035] If the video file contains multiple facial data, the action recognition result is used to determine whether the facial data contain the same action;

[0036] If the facial data contains the same action, the video segments containing the same action will be played synchronously, and the action evaluation results corresponding to the video segments will be displayed.

[0037] Secondly, the present invention also provides a terminal device, including a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs including a video content analysis method based on human skeleton point recognition as described in any of the above.

[0038] Thirdly, embodiments of the present invention also provide a non-transitory computer-readable storage medium, wherein when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device is able to perform the video content analysis method based on human skeleton point recognition motion as described in any of the above.

[0039] The beneficial effects of this invention are as follows: Compared with the prior art, this invention provides a video content analysis method based on human skeletal point recognition of motion. This invention first acquires a video file and then fragments the video file into video segments. Next, based on a skeletal algorithm, it performs motion recognition and evaluation of the skeletal points in the fragmented video segments to obtain motion evaluation data. Finally, based on the motion evaluation data, it outputs a video analysis and evaluation report. This invention, by fragmenting the video file and performing motion analysis and evaluation on the fragmented video segments using a skeletal algorithm, obtains motion evaluation results, enabling motion effect evaluation even without AI fitness equipment, thus improving the user's exercise and fitness experience. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a flowchart illustrating the video content analysis method based on human skeleton point recognition motion provided in an embodiment of the present invention.

[0042] Figure 2 This is an operation flowchart of the video content analysis method based on human skeleton point recognition motion provided in the embodiments of the present invention.

[0043] Figure 3 This is a flowchart of human skeleton point recognition action provided in an embodiment of the present invention.

[0044] Figure 4 This is a schematic diagram of human-computer interaction provided in an embodiment of the present invention.

[0045] Figure 5 This is a block diagram illustrating the internal structure of the terminal device provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0047] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0048] Currently, the evaluation of exercise and fitness effects requires the use of AI-powered exercise and fitness devices. Without these devices, the evaluation is not convenient, causing inconvenience for users.

[0049] To address the problems in existing technologies, this embodiment provides a video content analysis method based on human skeletal point recognition of actions. This method enables the evaluation of fitness effects from exercise videos even without AI-powered fitness equipment. In practice, the video file is first acquired and then fragmented into video segments. Next, a skeletal algorithm is used to identify and evaluate the actions of skeletal points in the segmented video footage, yielding action evaluation data. Finally, a video analysis and evaluation report is output based on the action evaluation data. Therefore, this invention, by fragmenting video files and then using a skeletal algorithm to identify and analyze actions within the segmented video footage to obtain an exercise evaluation report, enables the evaluation of fitness effects without AI-powered fitness equipment, thereby improving the user's fitness experience.

[0050] For example, such as Figure 2 As shown, the system first acquires user-recorded workout videos, then fragments and segments them into video segments. Next, a fitness application equipped with a skeletal algorithm performs motion recognition and evaluation on these video segments, obtaining an evaluation result for each exercise within each segment. For example, if a user's exercise in a segment is identified as a Russian twist, the application matches this exercise with standard exercise data from a standard exercise library to obtain the user's motion evaluation data, such as "Russian twist - 90 points - Standard." Finally, the motion evaluation data from each video segment is combined to generate a video analysis and evaluation report, which is then displayed in a window. Through this report, users can clearly see the evaluation results for all exercises in their workout videos. Therefore, even without AI-powered fitness equipment, users can evaluate their fitness results, greatly simplifying the process and improving their overall fitness experience.

[0051] Exemplary methods

[0052] This embodiment provides a video content analysis method based on human skeleton point recognition of motion, which can be applied to terminal devices. Specifically, as follows... Figure 1 As shown, the method includes:

[0053] Step S100: Obtain the video file and perform fragmentation cutting on the video file to obtain the cut video segments.

[0054] In this embodiment, the video file is a user-recorded exercise video, i.e., a video file that requires exercise evaluation and analysis. The segmented video file is obtained by fragmenting the video source, and motion recognition and analysis are performed on these segmented video files. Fragmenting the video file to obtain segmented video files effectively reduces the difficulty of video analysis, thereby increasing the analysis speed and accuracy of the analysis results.

[0055] In practical implementation, firstly, the video file is acquired, and its duration is read. Then, the duration of the video file is determined; if the duration exceeds a preset limit, a video codec is invoked to decode the video file. Finally, a preset segmentation rule is obtained, and the decoded video file is fragmented based on this rule to obtain the segmented video segments. The preset segmentation rule includes the number of user-input actions or a preset segmentation duration. By determining the duration of the acquired video file, some short videos that do not require fragmentation can be effectively filtered out, saving time and improving the efficiency of video analysis. If the video file is shorter than the preset duration, it indicates that the video file is a short video and does not require fragmentation; action recognition and analysis can be performed immediately. For example, if the preset duration is set to 5 minutes, video files shorter than 5 minutes are considered short videos and do not require fragmentation, while video files longer than 5 minutes are considered long videos and require fragmentation to obtain segmented video segments. User-recorded fitness videos are typically quite long, such as 15 minutes, half an hour, or an hour. These videos are uploaded to a cloud server or terminal device for fragmentation and then segmented for motion recognition and analysis. The preset duration can be set by the user. For video files requiring fragmentation, a video codec is first used to decode the video file. Then, the video file is fragmented based on the preset segmentation rules. These rules can be the number of actions in the recorded video file, input by the user (e.g., 5 or 10 actions), resulting in the number of video segments equal to the total video length divided by the number of actions input. Alternatively, the preset segmentation rules can be user-preset segmentation durations, where the video file is segmented at fixed time intervals. For example, a preset segmentation duration of 3 minutes results in each segment being 3 minutes long, or 5 minutes results in each segment being 5 minutes long. By fragmenting the video file to obtain segmented video segments, and then performing action recognition and analysis on the segmented video segments, the difficulty of action recognition and evaluation can be greatly reduced, while also improving processing efficiency and helping to obtain more accurate analysis results.

[0056] In another implementation, motion recognition is performed on the video file using a skeletal algorithm to obtain motion recognition results. Based on these results, the video file is then automatically segmented. For example, if the first motion in the video file is identified as 120 seconds and the second as 100 seconds, the video file is automatically segmented to obtain video segments of 120 seconds and 100 seconds in length, respectively. Based on this, the video file is segmented into multiple video segments containing complete motion sequences.

[0057] Step S200: Based on the skeletal algorithm, perform motion recognition and evaluation of the skeletal points of the cut video segment to obtain motion evaluation data.

[0058] In one implementation, the process of action recognition based on the skeletal algorithm is as follows: Figure 3 As shown, the pre-stored standard movement set includes various training movement sets such as high-intensity muscle-building training. Each training movement set has corresponding movements, including movement 1 / 2 / 3, and a skeletal point database of the standard movements. After obtaining the standard movement set, the movements in the video file can be matched and compared for analysis. First, the training movements and skeletal point data in the video file are identified and extracted using a skeletal algorithm. Then, the extracted training movements, dynamic movements, and the skeletal point database of the standard movements are compared and analyzed to output the movement recognition result based on the skeletal algorithm. The movement evaluation data includes: movement recognition result, movement matching degree, and movement evaluation result. The movement recognition result is the name of the identified movement, such as squat, Russian twist, kneeling push-up, etc. The movement matching degree is the score obtained by matching the movement in the segmented video with the preset standard movement, and the movement matching degree reflects the standardization of the user's movement. The movement evaluation result is the user's training result for the movement.

[0059] In specific implementation, after obtaining the segmented video, firstly, motion recognition of skeletal points is performed on the segmented video based on a skeletal algorithm to obtain the motion recognition result; then, the motion recognition result is matched with a preset standard motion to obtain the motion matching degree; finally, the motion recognition result is evaluated based on the motion matching degree to obtain the motion evaluation result. For example, if the motion recognition result obtained based on the skeletal algorithm is a kneeling push-up, and this motion is matched with a preset standard kneeling push-up in the standard motion library, the motion matching degree is 90%, then the motion evaluation result is standard. The motion matching degree is a score, divided into standard, average, and substandard. For example, a motion matching degree of 80% or higher is considered standard, 60%-80% is average, and below 60% is substandard. Alternatively, the calorie consumption value corresponding to the motion can be obtained, and the motion matching degree can be calculated according to the percentage of conformity. By matching with preset standard motions to obtain a score-based motion matching degree, the exercise and fitness effect can be effectively quantified, helping users clearly understand the training effect of the motion.

[0060] In one implementation, if the action evaluation result is substandard, the segmented video is marked. In practice, if the action evaluation result is substandard, a pop-up window prompts the user and marks the segment. The user can also manually favorite or follow the segmented video with substandard action evaluation results; for example, segment 1 - action 1 - 30 seconds - 50 seconds, standard / substandard, favorite. Users can mark the segmented video in a timely manner for focused training on that action.

[0061] In another implementation, if the action matching degree corresponding to the action recognition result in the segmented video is lower than a preset value, a prompt is given to the segmented video, and a corresponding exercise and fitness tutorial is recommended simultaneously. The exercise and fitness tutorial includes explanations of training essentials and standard action videos. Users can click on the standard action videos to watch and learn, and then perform intelligent follow-up exercises.

[0062] In one implementation, users can customize training improvement plans to focus on training movements in video segments where the evaluation result is substandard. Once the movements in these substandard video segments are trained to meet the standard, they are marked as having met the standard. Timely marking of video segments as met allows users to know immediately that they have mastered the training techniques for that movement and can move on to the next movement. In practice, by clicking on a substandard video segment, a link to its associated standard movement video will automatically pop up. Users can choose to play the standard movement video to learn and practice the substandard movement. When the movement is trained to meet the standard, a training achievement report is generated and simultaneously updated in the user's training improvement plan, which is a submodule of the motion analysis function module. Customized training improvement plans effectively help users update their training results in a timely manner, allowing them to gain a sense of accomplishment in their exercise.

[0063] Step S300: Output a video analysis and evaluation report based on the motion evaluation data.

[0064] In this embodiment, after performing segmented action recognition and evaluation on the video file, the action evaluation data is obtained, and a video analysis and evaluation report of the video file is output, such as... Figure 4 As shown. The video analysis and evaluation report includes: the total number of movements, the overall movement pass rate, the movement duration in each video segment, the movement evaluation results in each video segment, and a list of video segments where movements did not meet the standards. The list of video segments where movements did not meet the standards includes the substandard video segments and their corresponding standard training videos. Users can click the corresponding video link in the list to play and watch the videos. Clicking "More" on the interface will pause the video and bring up the "Motion Analysis" interface, which leads to the video analysis and evaluation report interface for that video file.

[0065] In one implementation, the displayed results may include the number and location markers of skeletal points, the display of skeletal point number markers in the standard training video, audio explanations of substandard movements, and point location markers and reminders corresponding to the skeletal points. For example, if the user's elbow joint skeletal point is found to be substandard, a red hollow circle will be displayed at the elbow joint position in the video and flash. The key points of the substandard movement will also be explained. This function can be selected to be turned on or off according to the user's needs, and is on by default. When the key point explanation function is on, the video file will be muted during playback, while the audio of the standard movement video will be on. The key point explanation data comes from the key point data in the movement training database of the exercise fitness application, and the key point data is stored in a cloud database.

[0066] In one implementation, motion recognition is performed on the segmented video based on a skeletal algorithm to obtain motion recognition results; the number of motions in the segmented video is then determined based on the motion recognition results; if the number of motions is greater than one, the segmented video is repeatedly fragmented based on the motion recognition results. This repeated fragmentation based on the motion recognition results ensures that each segmented video for motion evaluation analysis contains only one motion, facilitating matching and evaluation with standard motions. After repeated fragmentation, the segmented video is marked with motion and duration based on the motion recognition results; during video file playback, the motion and duration markers are displayed in a window.

[0067] In practice, after decoding the video file, the video codec segments the video file into fragments and immediately mark each segment with an action name and duration. That is, after obtaining the segment, it is marked according to its time within the complete video file, i.e., a segment is marked along its timeline within the complete video file. For example, segment 1, with the action name "burpee," has a duration of 000 seconds to 180 seconds in the video file; segment 2, with the action name "kneeling push-up," has a duration of 181 seconds to 360 seconds, and so on, marking all segments with action and duration. After marking, the action and duration markers are displayed on the progress bar during video playback. By marking each segment with action and duration and displaying them on the progress bar, users can clearly understand the specific details of the action.

[0068] In one implementation, the action recognition result is used to determine whether the segmented video contains the same action; if the segmented video contains the same action, the segmented video is associated with an action tag; and the segmented video is subjected to a year-on-year action comparison analysis based on the action association tag.

[0069] In practical implementation, action recognition results are obtained based on skeletal algorithms. If multiple video segments containing the same action are found in the video file based on the action recognition results, these video segments containing the same action are compared and presented. For example, if multiple video segments containing Russian twists are found in a video file, these multiple video segments are associated and marked. If the action evaluation results of these multiple associated video segments are all substandard, the associated video segments and standard action videos in the standard video library are presented simultaneously through multiple small windows. If among all associated video segments, some video segments have standard actions, some have average actions, and some have substandard actions, these standard video segments, average video segments, and substandard video segments are distinguished and marked, and presented through multiple small windows. By analyzing the association of video segments containing the same action, it is possible to effectively remind users to distinguish whether the substandard action is caused by user error or by a lack of mastery of the action technique, helping users to improve their training in a timely manner.

[0070] In one implementation, facial data in the video file is obtained based on a facial recognition algorithm; it is determined whether the video file contains multiple facial data based on the facial data; if the video file contains multiple facial data, the action recognition result is called to determine whether the facial data contains the same action; if the facial data contains the same action, the video segments containing the same action are played synchronously, and the action evaluation result corresponding to the video segments is displayed.

[0071] In practical implementation, the facial recognition algorithm is used to identify facial data in the video file, so that the facial data and action evaluation data can be correlated and analyzed. First, the facial recognition algorithm determines whether there is facial data in the video file. If facial data is identified, it is then determined whether there are multiple facial data. If multiple facial data are identified, it is determined whether the video segments corresponding to different facial data contain the same action. That is, whether there are multiple facial data corresponding to the same action in a single video file. If multiple facial data containing the same action are identified, the video segments containing the same action are output synchronously, presenting the analysis results of different people performing the same action in a single window. For example, video segment 1 has a duration of 000 seconds to 180 seconds in the video file, and the identified face ID is ID_RL_202309053030000180A, which is the face ID of user A. Video segment 2 has a duration of 181 seconds to 360 seconds in the video file, and the identified face ID is ID_RL_202309053030181360B, which is the face ID of user B. If the video segments corresponding to users A and B contain the same action, then these two video segments are associated and displayed synchronously in the window.

[0072] In another implementation, users can select video segments of facial data from multiple video files and display them in a window for comparing the training effects of multiple training videos. For example, users can view the video analysis and evaluation reports of yesterday's training videos and today's training videos to obtain a difference report, which can help users know in a timely manner whether they have made progress in their daily training.

[0073] Based on the above embodiments, the present invention also provides a terminal device, the schematic diagram of which can be as follows: Figure 5 As shown. The terminal device may include one or more processors 100 ( Figure 5 (Only one is shown in the diagram), memory 101, and computer program 102 stored in memory 101 and executable on one or more processors 100, such as a program for a video content analysis method based on human skeleton point recognition of motion. When one or more processors 100 execute computer program 102, they can implement various steps in the embodiments of the video content analysis method based on human skeleton point recognition of motion, which are not limited herein.

[0074] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0075] In one embodiment, memory 101 may be an internal storage unit of an electronic device, such as a hard drive or RAM. Memory 101 may also be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, memory 101 may include both internal and external storage units. Memory 101 is used to store computer programs and other programs and data required by the terminal device. Memory 101 can also be used to temporarily store data that has been output or will be output.

[0076] Those skilled in the art will understand that Figure 5 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, operational databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual operating data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0078] In summary, this invention discloses a video content analysis method based on human skeletal point recognition of motion. The method includes: acquiring a video file and fragmenting the video file into video segments; performing motion recognition and evaluation of skeletal points on the fragmented video segments based on a skeletal algorithm to obtain motion evaluation data; and outputting a video analysis and evaluation report based on the motion evaluation data. This invention, by fragmenting video files and performing motion analysis and evaluation on the fragmented video segments based on a skeletal algorithm, obtains motion evaluation results, enabling motion effect evaluation even without AI fitness equipment, thus improving the user's exercise and fitness experience.

[0079] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A video content analysis method based on human skeleton point recognition of actions, characterized in that, The method includes: Obtain a video file and perform fragmented segmentation of the video file to obtain segmented video segments; Based on the skeletal algorithm, the cut video segment is subjected to motion recognition and evaluation of skeletal points to obtain motion evaluation data; Output a video analysis and evaluation report based on the motion evaluation data; The method further includes: Action recognition is performed on the segmented video based on a skeletal algorithm to obtain action recognition results; The number of actions in the segmented video is obtained based on the action recognition results; If the number of actions is greater than one, the video segment is repeatedly fragmented based on the action recognition result. The method further includes: Facial data in the video file is obtained based on a facial recognition algorithm; Based on the facial data, determine whether the video file contains multiple facial data; If the video file contains multiple facial data, the action recognition result is used to determine whether the facial data contain the same action; If the facial data contains the same action, the video segments containing the same action will be played synchronously, and the action evaluation results corresponding to the video segments will be displayed. Video segments of facial data from multiple video files will be selected and displayed in a window to compare the training effects of multiple training recording videos.

2. The video content analysis method based on human skeleton point recognition of action according to claim 1, characterized in that, The process of acquiring a video file and fragmenting the video file to obtain segmented video segments includes: Obtain the video file and read its duration; If the duration of the video file exceeds the preset duration, the video codec is invoked to decode the video file; Obtain preset cutting rules, and perform fragmented cutting of the decoded video file based on the preset cutting rules to obtain the cut video segments. The preset cutting rules include the number of actions input by the user or the preset cutting duration.

3. The video content analysis method based on human skeleton point recognition of action according to claim 1, characterized in that, The motion evaluation data includes: motion recognition results, motion matching degree, and motion evaluation results. The motion recognition and evaluation of the segmented video segments based on a skeletal algorithm to obtain motion evaluation data includes: The action recognition result is obtained by performing motion recognition on the skeletal points of the cut video segment based on the skeletal algorithm. The action recognition result is matched with a preset standard action to obtain the action matching degree; Motion evaluation is performed on the action recognition result based on the action matching degree to obtain the action evaluation result.

4. The video content analysis method based on human skeleton point recognition of action according to claim 1, characterized in that, After repeatedly fragmenting the video segment based on the action recognition result if the number of actions is greater than one, the method further includes: The video segments are labeled with actions based on the action recognition results. The video segments are time-marked based on the action recognition results; When the video file is played, the action marker and the duration marker are displayed in a window.

5. The video content analysis method based on human skeleton point recognition of action according to claim 1, characterized in that, The method further includes: Based on the action recognition results, determine whether the cut video segments contain the same action; If the cut video segments contain the same action, then the cut video segments are marked with associated actions; Based on the action association tags, perform a year-on-year action comparison analysis on the segmented video.

6. The video content analysis method based on human skeleton point recognition of action according to claim 3, characterized in that, The step of evaluating the motion recognition result based on the motion matching degree to obtain the motion evaluation result includes: If the action evaluation result is unsatisfactory, the cut video segment is marked.

7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a program for video content analysis based on human skeleton point recognition motion stored in the memory and executable on the processor. When the processor executes the program for video content analysis based on human skeleton point recognition motion, it implements the steps of the video content analysis method based on human skeleton point recognition motion as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for a video content analysis method based on human skeleton point recognition of motion. When the program for the video content analysis method based on human skeleton point recognition of motion is executed by a processor, it implements the steps of the video content analysis method based on human skeleton point recognition of motion as described in any one of claims 1-6.